Interface Protocols · All levels

DDR Controller / PHY Split: Mechanism

Mechanism for DDR Controller / PHY Split.

Mechanism to understand

Mechanism for DDR Controller / PHY Split focuses on command efficiency, PHY training pass rate, read/write turnaround loss. The goal is to connect the observable symptom to protocol mechanism, ownership, and regression risk.

the controller schedules memory commands while the PHY handles electrical timing, calibration, and lane alignment. Think of it as a contract enforced at boundaries: the sender promises stability and legality, the receiver promises forward progress, and the fabric in between promises not to silently change identity or ordering.

  • Identify the transaction boundary: request, data, response, completion, or retry.

  • Identify the flow-control boundary: valid/ready, grant, credit, FIFO depth, or lane state.

  • Identify what the receiver is allowed to assume and what the sender must hold stable.

Layered view

diagram
PROTOCOL STACK VIEW — DDR Controller / PHY Split

software / firmware intent
        |
        v
transaction semantics: address, ID, length, attributes, ordering
        |
        v
link / channel behavior: handshake, credits, backpressure, retries
        |
        v
physical or timing layer: clocking, reset, pins, lanes, PHY
        |
        v
observability: waveform, VIP transaction, counter, analyzer trace

Debug rule: never jump layers without carrying the transaction identity with you.

Controller / PHY responsibility split

diagram
MEMORY STACK

  [ requestors ] --AXI/CHI--> [ MEMORY CONTROLLER ]
                                |  schedule, reorder, refresh
                                v
                              [ PHY ]
                                |  DQS/DQ timing, training, calibration
                                v
                              [ DRAM ]  banks / rows / columns

Controller thinks in transactions; PHY thinks in picoseconds.

Bank/row/column access

diagram
DRAM ACCESS = ACTIVATE -> READ/WRITE -> PRECHARGE

ACT row ──> row open in sense amps
   |          row hit  -> fast column access (good)
   |          row miss -> precharge + activate again (slow)
RD/WR col
PRE     ──> close row

ROW-HIT RATE GRAPH
hit% 90|████████  random-friendly layout
     60|█████
     30|██     pointer-chasing / bad interleave
       +----------------------------------> workload

Layer responsibilities

diagram
LAYER RESPONSIBILITY — DDR Controller / PHY Split

layer          owns                         common failure
-----------    --------------------------   -----------------------
software       intent, ordering needs       wrong assumption
transaction    id/addr/len/attributes       ordering / outstanding
link/channel   handshake, credits, retry    backpressure / deadlock
physical       clock/reset/lanes/PHY        timing / training / SI
observability  waveform/log/counter         missing evidence

Protocol deep dive

DDR bandwidth is scheduler + PHY: rows, banks, refresh, and turnarounds eat headline data rate.

Concept diagram

diagram
MEMORY PATH

masters -> controller scheduler -> PHY -> DRAM banks
              |                      |
         refresh/QoS            training/margin

Scheduler sees transactions; PHY sees picoseconds.

Metric graph

diagram
BANDWIDTH LOSS WATERFALL

peak              ████████████████████████
refresh           █████████████████████
turnaround        ██████████████████
row miss          ██████████████
effective         ██████████████

Quote the bottom bar in reviews.

Metrics and artifacts to collect

  • effective BW

  • row hit rate

  • refresh stall %

  • training margin

  • ECC error log

Mini case study

Video workload lost half effective bandwidth after firmware enabled aggressive low-power refresh. Scheduler and firmware QoS had to be co-designed.

Debug branches

  • If ECC errors, check training margin and address interleave first.

  • If BW low with high row hit, suspect port arbitration not DRAM.

  • If boot fail, stop at training step in transcript.

Senior review question

Ask: what is the first transaction that deviates, and which spec rule does it test?

Key takeaways

  • Connect every protocol claim to a transaction identity and measurable metric.

  • Store the artifact (waveform, log, counter) next to every signoff decision.

Common pitfalls

  • Debugging timeouts without finding the first bad transaction.

  • Quoting peak bus width without payload efficiency and retry overhead.

  • Treating VIP compliance as a substitute for system integration replay.

Mechanism deep dive

the controller schedules memory commands while the PHY handles electrical timing, calibration, and lane alignment.

Walk the transaction forward: request accepted → data moves → response completes → software visible effect.