Interface Protocols · All levels

Bandwidth & Latency Budgeting

Protocol Fundamentals: burst length, outstanding depth, arbitration, and packet overhead convert interface width into real workload throughput.

What this topic teaches

Bandwidth & Latency Budgeting is about converting a protocol rule into a measurable silicon contract. burst length, outstanding depth, arbitration, and packet overhead convert interface width into real workload throughput. The hard part is never the happy-path diagram; it is proving, under real traffic, which layer and which transaction broke the contract.

The senior-engineer question

When sustained bandwidth, p99 latency, utilization, head-of-line blocking moves, can you identify the transaction, the protocol layer, the responsible owner, and the smallest experiment that proves the root cause?

diagram
PROTOCOL STACK VIEW — Bandwidth & Latency Budgeting

software / firmware intent
        |
        v
transaction semantics: address, ID, length, attributes, ordering
        |
        v
link / channel behavior: handshake, credits, backpressure, retries
        |
        v
physical or timing layer: clocking, reset, pins, lanes, PHY
        |
        v
observability: waveform, VIP transaction, counter, analyzer trace

Debug rule: never jump layers without carrying the transaction identity with you.

Picture the protocol

Start every study session by drawing the behavior before reading signals. The diagrams below are the mental models to reproduce on a whiteboard.

Bandwidth vs offered load (knee curve)

diagram
LATENCY vs OFFERED LOAD

latency
  ^                                   *
  |                                 *
  |                              *
  |                          *  <- knee: queues build fast
  |                  *  *
  |   *  *  *  *
  +--------------------------------------> offered load (% of peak)
   0%        50%        80%   90%  100%

Lesson: usable bandwidth ends at the knee, not at 100% peak.

Payload efficiency stack

diagram
WHERE HEADLINE BANDWIDTH GOES

raw link            ████████████████████████  100%
- protocol overhead ██████████████████████    ~92%
- turnaround/idle   ██████████████████        ~75%
- retries/refresh   ████████████████          ~66%
= useful payload    ████████████████          ~66%

Always quote the bottom bar, not the top bar.

Transaction sequence

diagram
SEQUENCE — Bandwidth & Latency Budgeting

  initiator            interconnect/PHY            target
      |  request (id) ------->  |                     |
      |                         |  forward ----------> |
      |                         |                     | work
      |                         |  <---- response ---- |
      |  <----- complete ------ |                     |
      |
   metric captured here: sustained bandwidth, p99 latency, utilization, head-of-line blocking

Who owns which layer

diagram
LAYER RESPONSIBILITY — Bandwidth & Latency Budgeting

layer          owns                         common failure
-----------    --------------------------   -----------------------
software       intent, ordering needs       wrong assumption
transaction    id/addr/len/attributes       ordering / outstanding
link/channel   handshake, credits, retry    backpressure / deadlock
physical       clock/reset/lanes/PHY        timing / training / SI
observability  waveform/log/counter         missing evidence

Evidence to collect

  • Primary metric: sustained bandwidth, p99 latency, utilization, head-of-line blocking.

  • Primary artifact: bandwidth budget sheet, latency histogram, traffic replay summary.

  • Owners to bring into review: SoC architect, performance owner, integration owner.

  • Spec clause or requirement ID for every claim.

  • One traffic replay that fails and one reduced sequence that isolates the rule.

Ownership map

diagram
OWNERSHIP MAP — Bandwidth & Latency Budgeting

evidence type        owner who reads it
-----------------    ---------------------------
waveform/RTL        SoC architect
spec/VIP            performance owner
firmware/system     integration owner

Rule: every metric must have a named owner before a review starts.

Subpages in this topic

Each topic is taught across mechanism, inputs/outputs, reports, debug, worked example, pitfalls, interview, checklist, theory, design space, expanded case study, walkthrough, comparison matrix, software view, and silicon PPA impact.

Key takeaways

  • Carry transaction identity across waveform, log, counter, and spec view.

  • Separate protocol violation, integration configuration, and performance bottleneck before proposing a fix.

  • Draw the diagram first; the waveform should confirm the picture, not replace it.

Common pitfalls

  • Debugging only one channel or layer.

  • Treating a VIP error message as root cause instead of evidence.

  • Quoting peak interface bandwidth without payload efficiency.

Protocol deep dive

Before naming AXI or PCIe, engineers must master layering, handshakes, ordering, and bandwidth math. These four ideas explain 80% of integration bugs.

Concept diagram

diagram
FUNDAMENTALS STACK

software intent
     |
transaction (ID, addr, len, attr, order)
     |
link/channel (handshake, credit, retry)
     |
physical (clock, reset, lanes, PHY)

Debug golden rule: never change layers without carrying transaction identity.

Metric graph

diagram
STALL BREAKDOWN EXAMPLE

ready stalls      ████████████████  42%
credit wait       ██████████        26%
ordering block    ██████            16%
reset/config      ████              10%
other             ██                6%

If ready stalls dominate, widening the bus will not help.

Metrics and artifacts to collect

  • transaction latency by class

  • ready stall cycles

  • outstanding depth utilization

  • payload efficiency vs headline width

  • retry and error rate

Mini case study

A team widened a 64-bit interface to 128-bit but throughput rose only 8% because ready stalls from a slow slave dominated. Fixing slave acceptance and FIFO depth moved the metric; width did not.

Debug branches

  • If latency spikes but bandwidth flat, check outstanding limits and ordering.

  • If throughput collapses at high load, draw the knee curve — you are past queue stability.

  • If intermittent, compare reset release order and clock domain boundaries.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.