Interface Protocols · All levels

PCIe / CXL Debug: Debug Playbook

Debug Playbook for PCIe / CXL Debug.

Debug playbook

Debug Playbook for PCIe / CXL Debug focuses on link degrade event count, completion timeout rate, poison/error log. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

Protocol debug is a search for the FIRST deviation, not the loudest symptom. Timeouts and error flags are usually many cycles downstream of the real cause.

Root-cause tree

diagram
ROOT-CAUSE TREE — PCIe / CXL Debug

link degrade event count, completion timeout rate, poison/error log looks wrong
        |
   reproducible?
     /        \
   no          yes
   |            |
 flaky env   same first transaction every time?
 / seed         /            \
              yes             no
               |               |
        protocol rule     timing/reset/PVT
        or config bug     or load-dependent
  1. Freeze the failing seed, firmware tag, and spec revision.

  2. Find the first bad transaction, not the loudest downstream timeout.

  3. Map the transaction to signals and VIP monitor events.

  4. Classify the failure: protocol rule, integration config, timing/reset, or performance pressure.

  5. Prove the mechanism with one reduced sequence.

  6. Patch the smallest owner-controlled boundary and rerun compliance plus workload traffic.

Review memo template

diagram
STAFF PROTOCOL REVIEW MEMO — PCIe & CXL / PCIe / CXL Debug

1. Symptom
   - Watched metric: link degrade event count, completion timeout rate, poison/error log
   - Failing interface: <master/slave/endpoint/controller/PHY>
   - Transaction identity: <ID/tag/address/endpoint/lane>
   - Repro setup: <sim/emulation/FPGA/silicon + firmware tag>

2. Mechanism hypothesis
   - Primary mechanism: debug starts at link health, then checks packet credit, ordering, firmware resource allocation, and endpoint behavior.
   - Competing hypothesis: <timing, reset, bridge, ordering, firmware, or VIP issue>
   - Missing evidence: <waveform, analyzer trace, counter, spec clause, or log>

3. Proposed action
   - Minimal reversible change: <RTL, register setting, bridge config, scheduler, VIP check>
   - Expected metric movement: <delta and workload>
   - Regression risk: ordering, compatibility, performance, power, area, or timing

4. Signoff
   - Re-run artifact: protocol analyzer capture, AER log, LTSSM history, credit graph
   - Required owners: debug lead, firmware owner, controller owner
   - Final decision: fix, waive, document limitation, or escalate to architecture

Protocol deep dive

PCIe is reliable packet delivery over unreliable links; debug flows PHY -> DLL -> TLP -> firmware.

Concept diagram

diagram
PCIe DEBUG TOP-DOWN

L0 link healthy?  -> credits OK?  -> TLP completes?  -> driver happy?

Skip a layer and you will mis-own the bug.

Metric graph

diagram
LINK DEGRADE EXAMPLE

target x4 Gen4  ---- ---- ---- ----
actual   x4 Gen4  ---- ---- ---- ----   (eval board)
actual   x1 Gen3  -                   (product board)

Package/SI often shows up as width downgrade, not hard fail.

Metrics and artifacts to collect

  • link width/speed

  • replay count

  • completion timeout

  • AER error log

  • LTSSM history

Mini case study

Endpoint enumerated but DMA timed out: completion credits exhausted because a switch port was misconfigured in firmware, not because the endpoint was broken.

Debug branches

  • If degrade at width/speed, PHY/SI before driver.

  • If replay storm, link layer before transaction layer.

  • If CXL coherency bug, separate .io vs .cache vs .mem traffic.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

Principal review addendum

Re-read PCIe / CXL Debug against one concrete product workload, not a synthetic directed test.

debug starts at link health, then checks packet credit, ordering, firmware resource allocation, and endpoint behavior.