Interface Protocols · All levels
DDR Controller / PHY Split
Memory Interfaces (DDR / LPDDR / HBM): the controller schedules memory commands while the PHY handles electrical timing, calibration, and lane alignment.
What this topic teaches
DDR Controller / PHY Split is about converting a protocol rule into a measurable silicon contract. the controller schedules memory commands while the PHY handles electrical timing, calibration, and lane alignment. The hard part is never the happy-path diagram; it is proving, under real traffic, which layer and which transaction broke the contract.
The senior-engineer question
When command efficiency, PHY training pass rate, read/write turnaround loss moves, can you identify the transaction, the protocol layer, the responsible owner, and the smallest experiment that proves the root cause?
PROTOCOL STACK VIEW — DDR Controller / PHY Split
software / firmware intent
|
v
transaction semantics: address, ID, length, attributes, ordering
|
v
link / channel behavior: handshake, credits, backpressure, retries
|
v
physical or timing layer: clocking, reset, pins, lanes, PHY
|
v
observability: waveform, VIP transaction, counter, analyzer trace
Debug rule: never jump layers without carrying the transaction identity with you.Picture the protocol
Start every study session by drawing the behavior before reading signals. The diagrams below are the mental models to reproduce on a whiteboard.
Controller / PHY responsibility split
MEMORY STACK
[ requestors ] --AXI/CHI--> [ MEMORY CONTROLLER ]
| schedule, reorder, refresh
v
[ PHY ]
| DQS/DQ timing, training, calibration
v
[ DRAM ] banks / rows / columns
Controller thinks in transactions; PHY thinks in picoseconds.Bank/row/column access
DRAM ACCESS = ACTIVATE -> READ/WRITE -> PRECHARGE
ACT row ──> row open in sense amps
| row hit -> fast column access (good)
| row miss -> precharge + activate again (slow)
RD/WR col
PRE ──> close row
ROW-HIT RATE GRAPH
hit% 90|████████ random-friendly layout
60|█████
30|██ pointer-chasing / bad interleave
+----------------------------------> workloadTransaction sequence
SEQUENCE — DDR Controller / PHY Split
initiator interconnect/PHY target
| request (id) -------> | |
| | forward ----------> |
| | | work
| | <---- response ---- |
| <----- complete ------ | |
|
metric captured here: command efficiency, PHY training pass rate, read/write turnaround lossWho owns which layer
LAYER RESPONSIBILITY — DDR Controller / PHY Split
layer owns common failure
----------- -------------------------- -----------------------
software intent, ordering needs wrong assumption
transaction id/addr/len/attributes ordering / outstanding
link/channel handshake, credits, retry backpressure / deadlock
physical clock/reset/lanes/PHY timing / training / SI
observability waveform/log/counter missing evidenceEvidence to collect
Primary metric: command efficiency, PHY training pass rate, read/write turnaround loss.
Primary artifact: controller command trace, PHY training log, timing mode table.
Owners to bring into review: memory controller owner, PHY owner, SoC integration owner.
Spec clause or requirement ID for every claim.
One traffic replay that fails and one reduced sequence that isolates the rule.
Ownership map
OWNERSHIP MAP — DDR Controller / PHY Split
evidence type owner who reads it
----------------- ---------------------------
waveform/RTL memory controller owner
spec/VIP PHY owner
firmware/system SoC integration owner
Rule: every metric must have a named owner before a review starts.Subpages in this topic
Each topic is taught across mechanism, inputs/outputs, reports, debug, worked example, pitfalls, interview, checklist, theory, design space, expanded case study, walkthrough, comparison matrix, software view, and silicon PPA impact.
Key takeaways
Carry transaction identity across waveform, log, counter, and spec view.
Separate protocol violation, integration configuration, and performance bottleneck before proposing a fix.
Draw the diagram first; the waveform should confirm the picture, not replace it.
Common pitfalls
Debugging only one channel or layer.
Treating a VIP error message as root cause instead of evidence.
Quoting peak interface bandwidth without payload efficiency.
Protocol deep dive
DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.
Concept diagram
MEMORY PATH
masters -> controller scheduler -> PHY -> DRAM banks
| |
refresh/QoS training/margin
Scheduler sees transactions; PHY sees picoseconds.Metric graph
BANDWIDTH LOSS WATERFALL
peak ████████████████████████
refresh █████████████████████
turnaround ██████████████████
row miss ██████████████
effective ██████████████
Quote the bottom bar in reviews.Metrics and artifacts to collect
effective BW
row hit rate
refresh stall %
training margin
ECC error log
Mini case study
Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.
Debug branches
If ECC errors, check training margin and address interleave first.
If BW low with high row hit, suspect port arbitration not DRAM.
If boot fail, stop at training step in transcript.
Senior review question
Ask: what is the first transaction that deviates, and which spec rule does it test?
Key takeaways
Connect every protocol claim to a transaction identity and measurable metric.
Store the artifact (waveform, log, counter) next to every signoff decision.
Common pitfalls
Debugging timeouts without finding the first bad transaction.
Quoting peak bus width without payload efficiency and retry overhead.
Treating VIP compliance as a substitute for system integration replay.