Interface Protocols · All levels
Training & Timing Modes: Debug Playbook
Debug Playbook for Training & Timing Modes.
Debug playbook
Debug Playbook for Training & Timing Modes focuses on training margin, eye width, boot failure rate. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.
Protocol debug is a search for the FIRST deviation, not the loudest symptom. Timeouts and error flags are usually many cycles downstream of the real cause.
Root-cause tree
ROOT-CAUSE TREE — Training & Timing Modes
training margin, eye width, boot failure rate looks wrong
|
reproducible?
/ \
no yes
| |
flaky env same first transaction every time?
/ seed / \
yes no
| |
protocol rule timing/reset/PVT
or config bug or load-dependentFreeze the failing seed, firmware tag, and spec revision.
Find the first bad transaction, not the loudest downstream timeout.
Map the transaction to signals and VIP monitor events.
Classify the failure: protocol rule, integration config, timing/reset, or performance pressure.
Prove the mechanism with one reduced sequence.
Patch the smallest owner-controlled boundary and rerun compliance plus workload traffic.
Review memo template
STAFF PROTOCOL REVIEW MEMO — Memory Interfaces (DDR / LPDDR / HBM) / Training & Timing Modes
1. Symptom
- Watched metric: training margin, eye width, boot failure rate
- Failing interface: <master/slave/endpoint/controller/PHY>
- Transaction identity: <ID/tag/address/endpoint/lane>
- Repro setup: <sim/emulation/FPGA/silicon + firmware tag>
2. Mechanism hypothesis
- Primary mechanism: training aligns DQS/DQ timing and voltage margins so digital transfers survive PVT and board/package variation.
- Competing hypothesis: <timing, reset, bridge, ordering, firmware, or VIP issue>
- Missing evidence: <waveform, analyzer trace, counter, spec clause, or log>
3. Proposed action
- Minimal reversible change: <RTL, register setting, bridge config, scheduler, VIP check>
- Expected metric movement: <delta and workload>
- Regression risk: ordering, compatibility, performance, power, area, or timing
4. Signoff
- Re-run artifact: training transcript, margin report, mode register dump
- Required owners: bring-up owner, PHY owner, board owner
- Final decision: fix, waive, document limitation, or escalate to architectureProtocol deep dive
DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.
Concept diagram
MEMORY PATH
masters -> controller scheduler -> PHY -> DRAM banks
| |
refresh/QoS training/margin
Scheduler sees transactions; PHY sees picoseconds.Metric graph
BANDWIDTH LOSS WATERFALL
peak ████████████████████████
refresh █████████████████████
turnaround ██████████████████
row miss ██████████████
effective ██████████████
Quote the bottom bar in reviews.Metrics and artifacts to collect
effective BW
row hit rate
refresh stall %
training margin
ECC error log
Mini case study
Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.
Debug branches
If ECC errors, check training margin and address interleave first.
If BW low with high row hit, suspect port arbitration not DRAM.
If boot fail, stop at training step in transcript.
Senior review question
Ask: what is the first transaction that deviates, and which spec rule does it test?
Key takeaways
Connect every protocol claim to a transaction identity and measurable metric.
Store the artifact (waveform, log, counter) next to every signoff decision.
Common pitfalls
Debugging timeouts without finding the first bad transaction.
Quoting peak bus width without payload efficiency and retry overhead.
Treating VIP compliance as a substitute for system integration replay.
Principal review addendum
Re-read Training & Timing Modes against one concrete product workload, not a synthetic directed test.
training aligns DQS/DQ timing and voltage margins so digital transfers survive PVT and board/package variation.