Interface Protocols · All levels

Refresh & Bandwidth Efficiency: Interview Drills

Interview Drills for Refresh & Bandwidth Efficiency.

Interview drills

Interview Drills for Refresh & Bandwidth Efficiency focuses on effective bandwidth, row-hit rate, refresh stall percentage. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

diagram
PROMPT
You see effective bandwidth, row-hit rate, refresh stall percentage on Refresh & Bandwidth Efficiency. Walk through root cause and fix.

STRONG ANSWER
1. Names the layer and transaction identity.
2. Explains refresh, bank conflicts, turnaround, and command scheduling reduce useful bandwidth below headline bus width.
3. Requests bandwidth efficiency stack, bank conflict histogram, traffic class report.
4. Proposes one reduced sequence and one system regression.

WEAK ANSWER
Jumps to widening the interface, increasing FIFO depth, or blaming firmware without evidence.

Diagram to draw on the whiteboard

Where DDR bandwidth is lost

diagram
EFFECTIVE BANDWIDTH BREAKDOWN

peak bus            ████████████████████████  100%
- refresh stalls    ██████████████████████     ~92%
- read/write turn   ███████████████████        ~78%
- row miss penalty  ██████████████             ~58%
= effective         ██████████████             ~58%

Fix targets: better interleave, batch same-direction traffic, page policy.

Root-cause tree to narrate

diagram
ROOT-CAUSE TREE — Refresh & Bandwidth Efficiency

effective bandwidth, row-hit rate, refresh stall percentage looks wrong
        |
   reproducible?
     /        \
   no          yes
   |            |
 flaky env   same first transaction every time?
 / seed         /            \
              yes             no
               |               |
        protocol rule     timing/reset/PVT
        or config bug     or load-dependent

Protocol deep dive

DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.

Concept diagram

diagram
MEMORY PATH

masters -> controller scheduler -> PHY -> DRAM banks
              |                      |
         refresh/QoS            training/margin

Scheduler sees transactions; PHY sees picoseconds.

Metric graph

diagram
BANDWIDTH LOSS WATERFALL

peak              ████████████████████████
refresh           █████████████████████
turnaround        ██████████████████
row miss          ██████████████
effective         ██████████████

Quote the bottom bar in reviews.

Metrics and artifacts to collect

  • effective BW

  • row hit rate

  • refresh stall %

  • training margin

  • ECC error log

Mini case study

Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.

Debug branches

  • If ECC errors, check training margin and address interleave first.

  • If BW low with high row hit, suspect port arbitration not DRAM.

  • If boot fail, stop at training step in transcript.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

Interview whiteboard

Draw layers first, then place the failing transaction on the diagram.