CDC / RDC · All levels

Reset Debug Playbook: Expanded Case Study

Expanded Case Study for Reset Debug Playbook.

Extended case study

Milestone review flags reset-related failure reproduction time, first-pass root-cause hit rate around Reset Debug Playbook.

Background

Team believed crossings were stable until stress mode exposed intermittent anomalies.

Symptoms observed

  • reset-related failure reproduction time, first-pass root-cause hit rate regression

  • Mismatch between structural report and dynamic behavior

  • Waiver debate under schedule pressure

Investigation timeline

  1. Freeze design, reset, and tool configuration tags.

  2. Reproduce failing scenario with minimized stimulus.

  3. Map path/protocol/reset dependencies.

  4. Classify root cause and containment options.

  5. Execute minimal safe fix.

  6. Run stress regression and review board.

Root cause

Root cause links to Reset Debug Playbook: Reset bugs are often timing-sequence interactions across power, clocks, and reset domains; debug must isolate order, polarity, and release windows.

Fix and validation

  • Targeted RTL/protocol/reset correction

  • Evidence refresh

  • Signoff board decision

Lessons learned

  • Treat waivers as temporary risk contracts

  • Track owner and expiry

  • Regression before closure

diagram
CASE STUDY — Reset Debug Playbook
critical issues before / after / signoff

Crossing sequence under stress

diagram
CROSSING FLOW — Reset Debug Playbook

source clock domain -> launch signal -> crossing structure -> destination sample
      |                    |                 |                    |
   source FF           protocol           sync / fifo         destination FF

Key metric: reset-related failure reproduction time, first-pass root-cause hit rate

CDC/RDC deep dive

Reset release ordering is a first-order reliability contract.

Concept diagram

diagram
RESET RELEASE FLOW

assert global -> clocks stable -> sync release per domain -> first transaction

Metric graph

diagram
BOOT STABILITY

passes per 1k boots: 920 -> 980 -> 999

Reports and artifacts

  • reset dependency matrix

  • RDC warning classes

  • boot stress logs

  • waiver aging

Mini case study

Domain B released before producer A was valid, causing rare startup deadlock.

Debug branches

  • Correlate reset and clock timelines

  • verify async assert/sync release

  • exercise skewed release tests

Senior review question

Ask: what evidence proves this risk is closed for silicon, not just tool-clean?

Key takeaways

  • State crossing class, assumptions, and owner with every issue.

  • Run structural and dynamic regressions after each fix.

Common pitfalls

  • Treating all warnings as equivalent risk.

  • Waiving issues without containment evidence.

  • Skipping reset and reconvergence stress after CDC fixes.

Principal CDC/RDC review addendum

Reset bugs are often timing-sequence interactions across power, clocks, and reset domains; debug must isolate order, polarity, and release windows.

Metric: reset-related failure reproduction time, first-pass root-cause hit rate