AMS Interface · All levels
Link Training: Theory Deep Dive
Theory Deep Dive for Link Training.
Foundational theory
Link Training is central to SerDes & High-Speed I/O. Training protocols establish lane polarity, timing, equalization, and protocol states; digital state machines must synchronize with PHY indications and timeout policies. Senior AMS owners always tie observed failure to boundary assumptions, ownership, and measurable evidence before changing RTL or layout.
Core concepts explained
Training protocols establish lane polarity, timing, equalization, and protocol states; digital state machines must synchronize with PHY indications and timeout policies.
Primary metric: training pass rate, time-to-L0, fallback mode frequency
Primary artifact: training state trace, firmware log, protocol analyzer capture
Owners: firmware owner, PHY owner, system validation lead
Boundary and mode context are mandatory for any claim.
Treat lock/ready/valid bits as evidence, not proof of health.
Why this matters at signoff
At tapeout and bring-up, Link Training escapes are expensive to fix. SerDes reliability emerges from protocol + PHY + SI alignment. Wrong diagnosis burns schedule across analog, digital, and package teams.
Mental model
detect -> poll -> config EQ -> align -> active/L0
^ |
+------------- retry/fallback -----+Worked intuition
Name boundary and product mode where failure appears.
Open training pass rate, time-to-L0, fallback mode frequency and identify worst scenario.
Trace clocks/resets/config from analog macro to digital consumer.
Verify wrapper and handoff assumptions on the failing path.
Collect training state trace, firmware log, protocol analyzer capture and freeze evidence tags.
Classify root cause: contract gap, physical coupling, sequencing bug, or tool-view mismatch.
Propose minimal bounded change plus cross-domain regression.
Common misconceptions
Lock high means clock quality is automatically good.
Boundary cells are one-time checklist items, not runtime risks.
SerDes training failure is always firmware.
If average metric is healthy, there is no silicon risk.
Visual reinforcement
Training state path
detect -> poll -> config EQ -> align -> active/L0
^ |
+------------- retry/fallback -----+Layer responsibilities
AMS OWNERSHIP LAYERS — Link Training
layer owns failure mode
------------------ ---------------------------------- --------------------------
spec contract clocks/resets/interfaces hidden assumption drift
wrapper logic synchronizers/framing/flags silent data corruption
physical integration floorplan/isolation/power coupled noise and droop
signoff governance waivers/checklists/dashboard release with blind spots
closure debug order + regression fix regresses another modeAMS deep dive
SerDes closure requires protocol, training, and SI evidence together.
Concept diagram
SERDES FLOW
training -> equalization -> lane margin -> protocol stabilityMetric graph
BER VS EQ
BER
^
| high low high
+-----------------> EQ settingReports and artifacts
lane BER
training state transitions
EQ sweep report
link retry counters
Mini case study
Link looked protocol-clean but lane margin collapsed under thermal sweep.
Debug branches
Correlate LTSSM and lane metrics
Check SI margins
Review firmware timeout assumptions
Senior review question
Ask: what boundary condition proves this topic is actually closed?
Key takeaways
State boundary, mode, and evidence tag with every claim.
Always align analog, digital, and physical owners before signoff decisions.
Common pitfalls
Fixing averages while tails still fail.
Skipping package/supply evidence in jitter or SerDes issues.
Shipping with waivers that lack owner and expiration criteria.
Theory reinforcement
SerDes reliability emerges from protocol + PHY + SI alignment.