Interface Protocols · All levels

Memory Interface Debug: Reports & Metrics

Reports & Metrics for Memory Interface Debug.

Reports and metrics

Reports & Metrics for Memory Interface Debug focuses on ECC error rate, read timeout count, bandwidth regression. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

The job of a report is to turn ECC error rate, read timeout count, bandwidth regression into a decision. A single average number is almost never enough; you need the distribution, the traffic class breakdown, and a clear gap between legal maximum and product target.

Metric movement

diagram
METRIC GRAPH — ECC error rate, read timeout count, bandwidth regression

throughput / success
  ^
  |                         target
  |                       - - - - - - -
  |                  o after bounded fix
  |              o
  |         o baseline
  |    o failing run
  +--------------------------------------> experiment
    config A     isolated root cause     accepted change

Readout:
  - compare identical payload, clock, reset, traffic seed, and firmware setup
  - separate headline bandwidth from useful payload bandwidth
  - explain why the protocol mechanism moved the metric

Latency distribution

diagram
LATENCY HISTOGRAM — Memory Interface Debug

count
  |               ███
  |             ███████
  |          █████████████
  |        █████████████████        <- long tail = the real complaint
  |      ████████████████████████████
  +------------------------------------> latency
   p50      p90    p95       p99  (watch p99, not the average)

Average hides the tail; product pain lives at p95/p99.
  • Track ECC error rate, read timeout count, bandwidth regression by traffic class, payload size, and clock/reset mode.

  • Report p50/p95/p99 latency when user-visible stalls matter.

  • Include legal maximums and product targets; they are not the same thing.

  • Always store the metric next to the artifact that produced it.

Protocol deep dive

DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.

Concept diagram

diagram
MEMORY PATH

masters -> controller scheduler -> PHY -> DRAM banks
              |                      |
         refresh/QoS            training/margin

Scheduler sees transactions; PHY sees picoseconds.

Metric graph

diagram
BANDWIDTH LOSS WATERFALL

peak              ████████████████████████
refresh           █████████████████████
turnaround        ██████████████████
row miss          ██████████████
effective         ██████████████

Quote the bottom bar in reviews.

Metrics and artifacts to collect

  • effective BW

  • row hit rate

  • refresh stall %

  • training margin

  • ECC error log

Mini case study

Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.

Debug branches

  • If ECC errors, check training margin and address interleave first.

  • If BW low with high row hit, suspect port arbitration not DRAM.

  • If boot fail, stop at training step in transcript.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

How to read the numbers

ECC error rate, read timeout count, bandwidth regression must be split by traffic class, payload size, and reset mode.