Interface Protocols · All levels
Training & Timing Modes: Worked Example
Worked Example for Training & Timing Modes.
Worked example
Worked Example for Training & Timing Modes focuses on training margin, eye width, boot failure rate. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.
A product workload shows training margin, eye width, boot failure rate. The first review mistake is to blame the whole interface. A better review starts by pinning one transaction, proving where protocol progress stopped, and checking whether the observed behavior is legal for Training & Timing Modes.
Sequence under inspection
SEQUENCE — Training & Timing Modes
initiator interconnect/PHY target
| request (id) -------> | |
| | forward ----------> |
| | | work
| | <---- response ---- |
| <----- complete ------ | |
|
metric captured here: training margin, eye width, boot failure rateRead eye diagram
READ DATA EYE (sample in the center of the opening)
voltage
^ ____________
| / \ <- wider eye = more margin
| / sample \
| | . |
| \ /
| \____________/
+-------------------------> time (DQS phase)
^ ^
left edge right edge
center = (left+right)/2 -> training picks this pointCapture the failing waveform and transaction log.
Tag the request ID, address, endpoint, or lane.
Find the first response, retry, stall, or missing completion.
Compare against training transcript, margin report, mode register dump.
Choose one reversible fix and write the regression list before editing RTL or firmware.
Did the fix work?
BEFORE / AFTER — Training & Timing Modes
failing target
metric | ● ┄┄┄┄┄┄┄
| \
| \___ ● bounded fix
| \
| ● validated
+-------------------------------> change set
Prove the mechanism moved the metric; one good dot is not proof.Protocol deep dive
DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.
Concept diagram
MEMORY PATH
masters -> controller scheduler -> PHY -> DRAM banks
| |
refresh/QoS training/margin
Scheduler sees transactions; PHY sees picoseconds.Metric graph
BANDWIDTH LOSS WATERFALL
peak ████████████████████████
refresh █████████████████████
turnaround ██████████████████
row miss ██████████████
effective ██████████████
Quote the bottom bar in reviews.Metrics and artifacts to collect
effective BW
row hit rate
refresh stall %
training margin
ECC error log
Mini case study
Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.
Debug branches
If ECC errors, check training margin and address interleave first.
If BW low with high row hit, suspect port arbitration not DRAM.
If boot fail, stop at training step in transcript.
Senior review question
Ask: what is the first transaction that deviates, and which spec rule does it test?
Key takeaways
Connect every protocol claim to a transaction identity and measurable metric.
Store the artifact (waveform, log, counter) next to every signoff decision.
Common pitfalls
Debugging timeouts without finding the first bad transaction.
Quoting peak bus width without payload efficiency and retry overhead.
Treating VIP compliance as a substitute for system integration replay.
Narrative walkthrough
A team sees training margin, eye width, boot failure rate drop 40% after a seemingly small change near Training & Timing Modes.
They almost widen the interface. Instead they capture id=7 read burst and find W beats never matched AW len.