Interface Protocols · All levels

High-Speed I/O Debug: Worked Example

Worked Example for High-Speed I/O Debug.

Worked example

Worked Example for High-Speed I/O Debug focuses on BER, link retrain count, throughput under real traffic. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

A product workload shows BER, link retrain count, throughput under real traffic. The first review mistake is to blame the whole interface. A better review starts by pinning one transaction, proving where protocol progress stopped, and checking whether the observed behavior is legal for High-Speed I/O Debug.

Sequence under inspection

diagram
SEQUENCE — High-Speed I/O Debug

  initiator            interconnect/PHY            target
      |  request (id) ------->  |                     |
      |                         |  forward ----------> |
      |                         |                     | work
      |                         |  <---- response ---- |
      |  <----- complete ------ |                     |
      |
   metric captured here: BER, link retrain count, throughput under real traffic

BER vs equalization

diagram
BIT ERROR RATE vs EQ SETTING

BER (log)
 1e-3 |*                         *
 1e-6 |  *                    *
 1e-9 |     *             *
1e-12 |        *  *  *  *        <- usable window
      +-------------------------> EQ / tap setting
Pick the center of the low-BER window, not the edge.
  1. Capture the failing waveform and transaction log.

  2. Tag the request ID, address, endpoint, or lane.

  3. Find the first response, retry, stall, or missing completion.

  4. Compare against link monitor log, eye/margin report, packet analyzer capture.

  5. Choose one reversible fix and write the regression list before editing RTL or firmware.

Did the fix work?

diagram
BEFORE / AFTER — High-Speed I/O Debug

           failing        target
metric  |    ●              ┄┄┄┄┄┄┄
        |     \
        |      \___ ● bounded fix
        |           \
        |            ● validated
        +-------------------------------> change set
Prove the mechanism moved the metric; one good dot is not proof.

Protocol deep dive

USB/Ethernet/MIPI failures cross MAC counters, PCS framing, PHY adaptation, and channel SI.

Concept diagram

diagram
HIGH-SPEED STACK

app -> MAC/framing -> PCS/encoding -> SerDes/PHY -> channel

CRC errors often mean PCS/PHY/channel, not TCP.

Metric graph

diagram
BER vs EQ SETTING

BER
1e-3 |*
1e-6 |  *
1e-9 |     **** usable window
1e-12|          *
     +-----------------> EQ tap

Metrics and artifacts to collect

  • CRC error rate

  • retrain count

  • frame drop

  • lane error

  • BER

Mini case study

Ethernet link up at 100G but lossy: equalization margin on one lane narrow after package change. Digital counters were clean; PHY margin was not.

Debug branches

  • If link up but lossy, PHY margin and retrain.

  • If enumeration OK but throughput low, check packet size and DMA batching.

  • If MIPI frame drops, blanking budget and lane polarity.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

Narrative walkthrough

A team sees BER, link retrain count, throughput under real traffic drop 40% after a seemingly small change near High-Speed I/O Debug.

They almost widen the interface. Instead they capture id=7 read burst and find W beats never matched AW len.