CDC / RDC · All levels

Reset Debug Playbook

Reset Domain Crossing: Reset bugs are often timing-sequence interactions across power, clocks, and reset domains; debug must isolate order, polarity, and release windows.

What this topic teaches

Reset Debug Playbook focuses on closing CDC/RDC risk with mechanism-level reasoning. Reset bugs are often timing-sequence interactions across power, clocks, and reset domains; debug must isolate order, polarity, and release windows. Senior signoff depends on proving behavior with targeted evidence, not just clearing tool warnings.

The senior-engineer question

When reset-related failure reproduction time, first-pass root-cause hit rate regresses, can you classify the hazard, identify accountable owners, and choose the smallest fix or waiver backed by evidence?

diagram
CDC/RDC SIGNOFF FLOW — Reset Debug Playbook

crossing inventory + reset map
          |
          v
crossing classification (level/pulse/bus/reset)
          |
          v
structure + protocol + reset checks
          |
          v
critical issues + waiver review
          |
          v
fix / validate / regress / signoff

Picture the crossing behavior

Draw the behavior before touching tools. These visuals are the expected whiteboard baseline for reviews and interviews.

Reset failure timeline

diagram
t0 reset assert
t1 clock ungated
t2 reset release domain A
t3 reset release domain B
t4 first traffic

Correlate failure to exact release ordering.

Crossing sequence

diagram
CROSSING FLOW — Reset Debug Playbook

source clock domain -> launch signal -> crossing structure -> destination sample
      |                    |                 |                    |
   source FF           protocol           sync / fifo         destination FF

Key metric: reset-related failure reproduction time, first-pass root-cause hit rate

Ownership layers

diagram
CDC/RDC OWNERSHIP LAYERS — Reset Debug Playbook

layer                 owns                            common failure
------------------    -----------------------------   -----------------------------
design intent         crossing architecture           wrong topology selected
protocol semantics    req/ack, fifo, ordering        liveness/deadlock bugs
reset behavior        assert/deassert sequencing      boot instability
analysis setup        tool rules + waivers            false confidence
signoff governance    risk acceptance + dashboard     stale critical waivers

Evidence to collect

  • Primary metric: reset-related failure reproduction time, first-pass root-cause hit rate.

  • Primary artifact: boot waveform bundle, reset event timeline, root-cause memo.

  • Owners to involve: integration lead, verification lead, RDC owner.

  • At least one reproducer tied to mode/reset/traffic context.

  • Decision record: fix, waive, or escalate with rationale.

Ownership map

diagram
OWNERSHIP MAP — Reset Debug Playbook

artifact                  owner
----------------------    -------------------------
design intent           integration lead
verification evidence   verification lead
signoff decision        RDC owner

Every open CDC/RDC issue needs one accountable owner before waiver or fix.

Subpages in this topic

Each topic includes mechanism, I/O contract, metrics, debug, worked example, pitfalls, interview drills, checklist, theory, design tradeoffs, expanded case study, walkthrough, comparison matrix, software view, and silicon impact.

Key takeaways

  • Classify crossing/reset hazards before proposing fixes.

  • Pair structural results with protocol/reset behavioral proof.

  • Treat waivers as bounded risk contracts, not cleanup shortcuts.

Common pitfalls

  • Mass-waiving warnings near tapeout.

  • Assuming local IP cleanliness guarantees SoC behavior.

  • Skipping reconvergence and reset stress after CDC fixes.

CDC/RDC deep dive

Reset release ordering is a first-order reliability contract.

Concept diagram

diagram
RESET RELEASE FLOW

assert global -> clocks stable -> sync release per domain -> first transaction

Metric graph

diagram
BOOT STABILITY

passes per 1k boots: 920 -> 980 -> 999

Reports and artifacts

  • reset dependency matrix

  • RDC warning classes

  • boot stress logs

  • waiver aging

Mini case study

Domain B released before producer A was valid, causing rare startup deadlock.

Debug branches

  • Correlate reset and clock timelines

  • verify async assert/sync release

  • exercise skewed release tests

Senior review question

Ask: what evidence proves this risk is closed for silicon, not just tool-clean?

Key takeaways

  • State crossing class, assumptions, and owner with every issue.

  • Run structural and dynamic regressions after each fix.

Common pitfalls

  • Treating all warnings as equivalent risk.

  • Waiving issues without containment evidence.

  • Skipping reset and reconvergence stress after CDC fixes.