CDC / RDC · All levels

Handshake Protocols Across Domains: Debug Playbook

Debug Playbook for Handshake Protocols Across Domains.

Debug playbook

Debug Playbook for Handshake Protocols Across Domains focuses on req/ack completion reliability, deadlock risk, backpressure behavior. The goal is to convert issue observations into mechanism-backed closure decisions.

CDC/RDC debug is about finding the earliest violated assumption. Start with intent and context before touching low-level signal traces.

Root-cause tree

diagram
ROOT-CAUSE TREE — Handshake Protocols Across Domains

crossing failure observed
        |
   reproducible?
     /        \
   no          yes
   |            |
stress mode   classify issue
expansion       /      |      \
            synchronizer protocol reset/reconvergence
                 |         |           |
            MTBF fit    liveness    release ordering
  1. Freeze RTL/config/tool tags for reproducibility.

  2. Reproduce in smallest mode/reset/traffic scenario.

  3. Classify mechanism: synchronizer, protocol, reset, reconvergence, or governance.

  4. Collect one decisive artifact that proves the class.

  5. Pick minimal fix or bounded waiver.

  6. Run targeted and full-regression matrices before closure.

Review memo template

diagram
STAFF CDC/RDC REVIEW MEMO — CDC Protocols & Handshakes / Handshake Protocols Across Domains

1. Symptom
   - Failing metric: req/ack completion reliability, deadlock risk, backpressure behavior
   - Context: <mode, traffic, reset state, corner>
   - Risk class: <critical/high/medium/low>
   - Database tags: <rtl, config, assertions, tool setup>

2. Mechanism hypothesis
   - Primary mechanism: Two-way handshakes guarantee delivery across asynchronous clocks when both request and acknowledge paths are synchronized and protocol invariants are enforced.
   - Competing hypothesis: <false warning / protocol bug / reset order / reconvergence>
   - Missing evidence: <assertion, waveform, formal proof, stress replay>

3. Proposed action
   - Minimal reversible change: <sync/protocol/reset/waiver decision>
   - Expected metric movement: <critical count delta>
   - Regression risk: throughput, boot, latency, mode interaction

4. Signoff
   - Re-run artifact: handshake timing spec, assertion suite, latency histogram
   - Required owners: RTL owner, verification owner, CDC owner
   - Final decision: fix, bounded waiver, or escalate

CDC/RDC deep dive

Protocol correctness is the bridge between structural clean and functional safe.

Concept diagram

diagram
PROTOCOL FLOW

intent -> transport protocol -> synchronization -> destination acceptance

Metric graph

diagram
PROTOCOL ISSUE BURNDOWN

open issues: 20 -> 11 -> 5 -> 0

Reports and artifacts

  • FIFO pointer proofs

  • req/ack liveness

  • pulse miss checks

  • protocol assertions

Mini case study

Async FIFO empty/full logic looked correct until gray decode mismatch appeared during reset overlap.

Debug branches

  • Pointer sync audit

  • formal liveness checks

  • reset interaction review

Senior review question

Ask: what evidence proves this risk is closed for silicon, not just tool-clean?

Key takeaways

  • State crossing class, assumptions, and owner with every issue.

  • Run structural and dynamic regressions after each fix.

Common pitfalls

  • Treating all warnings as equivalent risk.

  • Waiving issues without containment evidence.

  • Skipping reset and reconvergence stress after CDC fixes.

Principal CDC/RDC review addendum

Two-way handshakes guarantee delivery across asynchronous clocks when both request and acknowledge paths are synchronized and protocol invariants are enforced.

Metric: req/ack completion reliability, deadlock risk, backpressure behavior