CDC / RDC · All levels

Reset Debug Playbook: Worked Example

Worked Example for Reset Debug Playbook.

Worked example

Worked Example for Reset Debug Playbook focuses on reset-related failure reproduction time, first-pass root-cause hit rate. The goal is to convert issue observations into mechanism-backed closure decisions.

A milestone review shows reset-related failure reproduction time, first-pass root-cause hit rate. Teams disagree on severity. The right move is to isolate one representative issue, prove mechanism class, and decide fix or waiver with explicit residual risk.

Crossing under inspection

diagram
CROSSING FLOW — Reset Debug Playbook

source clock domain -> launch signal -> crossing structure -> destination sample
      |                    |                 |                    |
   source FF           protocol           sync / fifo         destination FF

Key metric: reset-related failure reproduction time, first-pass root-cause hit rate

Reset failure timeline

diagram
t0 reset assert
t1 clock ungated
t2 reset release domain A
t3 reset release domain B
t4 first traffic

Correlate failure to exact release ordering.
  1. Capture warning, waveform, and owning module context.

  2. Tag mode/reset/traffic state for the failure.

  3. Validate assumptions against spec and assertions.

  4. Compare outcome with boot waveform bundle, reset event timeline, root-cause memo.

  5. Choose one reversible action and define regression upfront.

Did the action work?

diagram
BEFORE / AFTER — Reset Debug Playbook

open critical issues
  ^
  |  o baseline
  |     o after fix batch
  |         o after protocol proof
  |             o signoff-ready
  +---------------------------------> closure iteration

Track issue burn-down with evidence quality, not only count.

CDC/RDC deep dive

Reset release ordering is a first-order reliability contract.

Concept diagram

diagram
RESET RELEASE FLOW

assert global -> clocks stable -> sync release per domain -> first transaction

Metric graph

diagram
BOOT STABILITY

passes per 1k boots: 920 -> 980 -> 999

Reports and artifacts

  • reset dependency matrix

  • RDC warning classes

  • boot stress logs

  • waiver aging

Mini case study

Domain B released before producer A was valid, causing rare startup deadlock.

Debug branches

  • Correlate reset and clock timelines

  • verify async assert/sync release

  • exercise skewed release tests

Senior review question

Ask: what evidence proves this risk is closed for silicon, not just tool-clean?

Key takeaways

  • State crossing class, assumptions, and owner with every issue.

  • Run structural and dynamic regressions after each fix.

Common pitfalls

  • Treating all warnings as equivalent risk.

  • Waiving issues without containment evidence.

  • Skipping reset and reconvergence stress after CDC fixes.

Principal CDC/RDC review addendum

Reset bugs are often timing-sequence interactions across power, clocks, and reset domains; debug must isolate order, polarity, and release windows.

Metric: reset-related failure reproduction time, first-pass root-cause hit rate