Interface Protocols · All levels

Bandwidth & Latency Budgeting: Debug Playbook

Debug Playbook for Bandwidth & Latency Budgeting.

Debug playbook

Debug Playbook for Bandwidth & Latency Budgeting focuses on sustained bandwidth, p99 latency, utilization, head-of-line blocking. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

Protocol debug is a search for the FIRST deviation, not the loudest symptom. Timeouts and error flags are usually many cycles downstream of the real cause.

Root-cause tree

diagram
ROOT-CAUSE TREE — Bandwidth & Latency Budgeting

sustained bandwidth, p99 latency, utilization, head-of-line blocking looks wrong
        |
   reproducible?
     /        \
   no          yes
   |            |
 flaky env   same first transaction every time?
 / seed         /            \
              yes             no
               |               |
        protocol rule     timing/reset/PVT
        or config bug     or load-dependent
  1. Freeze the failing seed, firmware tag, and spec revision.

  2. Find the first bad transaction, not the loudest downstream timeout.

  3. Map the transaction to signals and VIP monitor events.

  4. Classify the failure: protocol rule, integration config, timing/reset, or performance pressure.

  5. Prove the mechanism with one reduced sequence.

  6. Patch the smallest owner-controlled boundary and rerun compliance plus workload traffic.

Review memo template

diagram
STAFF PROTOCOL REVIEW MEMO — Protocol Fundamentals / Bandwidth & Latency Budgeting

1. Symptom
   - Watched metric: sustained bandwidth, p99 latency, utilization, head-of-line blocking
   - Failing interface: <master/slave/endpoint/controller/PHY>
   - Transaction identity: <ID/tag/address/endpoint/lane>
   - Repro setup: <sim/emulation/FPGA/silicon + firmware tag>

2. Mechanism hypothesis
   - Primary mechanism: burst length, outstanding depth, arbitration, and packet overhead convert interface width into real workload throughput.
   - Competing hypothesis: <timing, reset, bridge, ordering, firmware, or VIP issue>
   - Missing evidence: <waveform, analyzer trace, counter, spec clause, or log>

3. Proposed action
   - Minimal reversible change: <RTL, register setting, bridge config, scheduler, VIP check>
   - Expected metric movement: <delta and workload>
   - Regression risk: ordering, compatibility, performance, power, area, or timing

4. Signoff
   - Re-run artifact: bandwidth budget sheet, latency histogram, traffic replay summary
   - Required owners: SoC architect, performance owner, integration owner
   - Final decision: fix, waive, document limitation, or escalate to architecture

Protocol deep dive

Before naming AXI or PCIe, engineers must master layering, handshakes, ordering, and bandwidth math. These four ideas explain 80% of integration bugs.

Concept diagram

diagram
FUNDAMENTALS STACK

software intent
     |
transaction (ID, addr, len, attr, order)
     |
link/channel (handshake, credit, retry)
     |
physical (clock, reset, lanes, PHY)

Debug golden rule: never change layers without carrying transaction identity.

Metric graph

diagram
STALL BREAKDOWN EXAMPLE

ready stalls      ████████████████  42%
credit wait       ██████████        26%
ordering block    ██████            16%
reset/config      ████              10%
other             ██                6%

If ready stalls dominate, widening the bus will not help.

Metrics and artifacts to collect

  • transaction latency by class

  • ready stall cycles

  • outstanding depth utilization

  • payload efficiency vs headline width

  • retry and error rate

Mini case study

A team widened a 64-bit interface to 128-bit but throughput rose only 8% because ready stalls from a slow slave dominated. Fixing slave acceptance and FIFO depth moved the metric; width did not.

Debug branches

  • If latency spikes but bandwidth flat, check outstanding limits and ordering.

  • If throughput collapses at high load, draw the knee curve — you are past queue stability.

  • If intermittent, compare reset release order and clock domain boundaries.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

Principal review addendum

Re-read Bandwidth & Latency Budgeting against one concrete product workload, not a synthetic directed test.

burst length, outstanding depth, arbitration, and packet overhead convert interface width into real workload throughput.