AMS Interface · All levels

High-Speed SI Awareness: Debug Playbook

Debug Playbook for High-Speed SI Awareness.

Debug playbook

Debug Playbook for High-Speed SI Awareness focuses on SI-induced retries, lane margin, intermittent link drop rate. The goal is to connect observed symptom to boundary mechanism, ownership, and signoff risk.

AMS debug is a hunt for first divergence, not downstream symptom management. Most costly delays come from wrong-owner first actions.

Root-cause tree

diagram
ROOT-CAUSE TREE — High-Speed SI Awareness

SI-induced retries, lane margin, intermittent link drop rate regressed
        |
same silicon / run tags?
   /            \
 no              yes
 |                |
env mismatch    boundary contract or
tag mismatch    true physical issue
 /   \             |
clk   reset      isolate first failing
map   sequence   boundary transition
  1. Freeze reproducer: mode, firmware/config, and evidence tags.

  2. Find the first boundary signal that diverges.

  3. Map divergence to contract clause and owner.

  4. Classify failure: contract, sequencing, coupling, package, or tool-view mismatch.

  5. Prove mechanism with one reduced reproducer.

  6. Apply smallest reversible fix and rerun cross-domain regressions.

Review memo template

diagram
STAFF AMS REVIEW MEMO — SerDes & High-Speed I/O / High-Speed SI Awareness

1. Symptom
   - Watched metric: SI-induced retries, lane margin, intermittent link drop rate
   - Failing mode/condition: <power/clock/temp/workload>
   - Boundary under suspicion: <macro/wrapper/interface/lane/island>
   - Repro setup: <sim/emulation/lab + firmware/config tags>

2. Mechanism hypothesis
   - Primary mechanism: At multi-Gbps speeds, package/channel discontinuities and return-path quality convert into jitter and eye closure that appear as digital protocol instability.
   - Competing hypothesis: <contract gap, sequencing, physical coupling, package, tooling>
   - Missing evidence: <waveform, report, scope/analyzer capture, dashboard snapshot>

3. Proposed action
   - Minimal reversible change: <RTL/config/layout/policy>
   - Expected metric movement: <delta and conditions>
   - Regression risk: timing, noise, power, performance, compatibility

4. Signoff
   - Re-run artifact: channel compliance report, lane margin scan, SI waiver sheet
   - Required owners: package owner, SI/PI owner, SerDes integration owner
   - Final decision: fix, waive with controls, or escalate

AMS deep dive

SerDes closure requires protocol, training, and SI evidence together.

Concept diagram

diagram
SERDES FLOW

training -> equalization -> lane margin -> protocol stability

Metric graph

diagram
BER VS EQ

BER
 ^
 |  high  low  high
 +-----------------> EQ setting

Reports and artifacts

  • lane BER

  • training state transitions

  • EQ sweep report

  • link retry counters

Mini case study

Link looked protocol-clean but lane margin collapsed under thermal sweep.

Debug branches

  • Correlate LTSSM and lane metrics

  • Check SI margins

  • Review firmware timeout assumptions

Senior review question

Ask: what boundary condition proves this topic is actually closed?

Key takeaways

  • State boundary, mode, and evidence tag with every claim.

  • Always align analog, digital, and physical owners before signoff decisions.

Common pitfalls

  • Fixing averages while tails still fail.

  • Skipping package/supply evidence in jitter or SerDes issues.

  • Shipping with waivers that lack owner and expiration criteria.

Principal AMS review addendum

At multi-Gbps speeds, package/channel discontinuities and return-path quality convert into jitter and eye closure that appear as digital protocol instability.

Metric: SI-induced retries, lane margin, intermittent link drop rate