Interface Protocols · All levels

Training & Timing Modes: Interview Drills

Interview Drills for Training & Timing Modes.

Interview drills

Interview Drills for Training & Timing Modes focuses on training margin, eye width, boot failure rate. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

diagram
PROMPT
You see training margin, eye width, boot failure rate on Training & Timing Modes. Walk through root cause and fix.

STRONG ANSWER
1. Names the layer and transaction identity.
2. Explains training aligns DQS/DQ timing and voltage margins so digital transfers survive PVT and board/package variation.
3. Requests training transcript, margin report, mode register dump.
4. Proposes one reduced sequence and one system regression.

WEAK ANSWER
Jumps to widening the interface, increasing FIFO depth, or blaming firmware without evidence.

Diagram to draw on the whiteboard

Read eye diagram

diagram
READ DATA EYE (sample in the center of the opening)

voltage
  ^      ____________
  |     /            \        <- wider eye = more margin
  |    /   sample     \
  |   |      .         |
  |    \              /
  |     \____________/
  +-------------------------> time (DQS phase)
        ^           ^
     left edge   right edge
   center = (left+right)/2  -> training picks this point

Root-cause tree to narrate

diagram
ROOT-CAUSE TREE — Training & Timing Modes

training margin, eye width, boot failure rate looks wrong
        |
   reproducible?
     /        \
   no          yes
   |            |
 flaky env   same first transaction every time?
 / seed         /            \
              yes             no
               |               |
        protocol rule     timing/reset/PVT
        or config bug     or load-dependent

Protocol deep dive

DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.

Concept diagram

diagram
MEMORY PATH

masters -> controller scheduler -> PHY -> DRAM banks
              |                      |
         refresh/QoS            training/margin

Scheduler sees transactions; PHY sees picoseconds.

Metric graph

diagram
BANDWIDTH LOSS WATERFALL

peak              ████████████████████████
refresh           █████████████████████
turnaround        ██████████████████
row miss          ██████████████
effective         ██████████████

Quote the bottom bar in reviews.

Metrics and artifacts to collect

  • effective BW

  • row hit rate

  • refresh stall %

  • training margin

  • ECC error log

Mini case study

Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.

Debug branches

  • If ECC errors, check training margin and address interleave first.

  • If BW low with high row hit, suspect port arbitration not DRAM.

  • If boot fail, stop at training step in transcript.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

Interview whiteboard

Draw layers first, then place the failing transaction on the diagram.