DRAM & Memory Design · All levels
FR-FCFS, Row-Buffer Locality, and Page Policy Control: Reports and Metrics
Reports and Metrics for FR-FCFS, Row-Buffer Locality, and Page Policy Control.
Reports and metrics
Reports and Metrics for FR-FCFS, Row-Buffer Locality, and Page Policy Control focuses on Row-hit rate, effective command efficiency, and average activate/precharge overhead per request.. The purpose is to turn memory observations into mechanism-backed actions with explicit owners and release-safe validation.
Reports should explain why Row-hit rate, effective command efficiency, and average activate/precharge overhead per request. moved, not simply that it moved. Require evidence that links the movement to command behavior, queue policy, PHY margin, or reliability controls.
Before/after trend
BEFORE / AFTER GRAPH - FR-FCFS, Row-Buffer Locality, and Page Policy Control
metric quality
^
| o target band
| o post-fix sweep
| o
| o baseline (failing)
+----------------------------------------------> iteration
evidence capture fix applied closure run
Use this view to prove improvement is causal, not accidental.Evidence matrix
DRAM EVIDENCE MATRIX - FR-FCFS, Row-Buffer Locality, and Page Policy Control
+-------------------------------+--------------------------------+--------------------------------+---------------------------+
| Evidence | Tells you | Does not prove | Next action |
+-------------------------------+--------------------------------+--------------------------------+---------------------------+
| row-hit/miss + ACT/PRE mix | locality and row-state cost | lane-level capture integrity | inspect training margins |
| queue age + class breakdown | fairness and starvation risk | command legality details | parse command timeline |
| JEDEC legality + bus timeline | timing-window pressure | root cause by itself | correlate with traffic map|
| eye / Vref / skew snapshots | PHY margin and drift behavior | controller policy quality | pair with schedule logs |
| CE/UE + scrub telemetry | reliability trajectory | immediate perf bottleneck only | map to hotspot addresses |
+-------------------------------+--------------------------------+--------------------------------+---------------------------+Track p50/p95/p99 latency and effective bandwidth together.
Include command and queue context alongside high-level counters.
Tag reports with firmware, timing profile, and thermal state.
Call out contradictory evidence instead of hiding it.
DRAM deep dive
Controller policy decides whether DRAM serves locality, fairness, and QoS targets simultaneously.
Concept diagram
CONTROLLER SCHEDULING LOOP
request queues -> row-policy + priority -> command issue -> bank state updateMetric graph
QUEUE PRESSURE MIX
row-hit preference bias ██████
aging/fairness pressure █████
QoS override cost ███Reports and artifacts
scheduler policy comparison
queue age distribution
starvation/fairness incident report
QoS latency percentile dashboard
Mini case study
FR-FCFS tuning improved bulk throughput but starved latency-critical traffic until age caps and class quotas were added.
Debug branches
Measure queue age tails by traffic class
Separate row-hit gains from fairness regressions
Stress policy under mixed burst and random streams
Senior review question
Ask: which latency, bandwidth, and reliability evidence proves this DRAM topic is closed under real traffic?
Key takeaways
Always tie controller and PHY counter shifts to application latency and throughput outcomes.
Lock firmware timing profile, thermal condition, and DIMM state before comparing DRAM captures.
Common pitfalls
Chasing peak bandwidth while ignoring p99 latency and fairness tails.
Changing timing guardbands without separating SI noise from scheduling issues.
Declaring closure without reliability gates, fault injection, and regression replay.
Report interpretation
FR-FCFS (First-Ready, First-Come-First-Serve) prioritizes commands that are timing-ready now, and among those typically prefers older arrivals; in practice this strongly favors row hits because an open-row access can issue quickly while a row miss requires PRECHARGE plus ACTIVATE latency. The policy boosts throughput by harvesting row-buffer locality, but can also bias service toward hot rows and penalize streams that repeatedly miss. Page policy selection (open-page, close-page, or adaptive hybrids) determines whether the controller keeps a row open after service or proactively closes it to reduce future conflict cost. Open-page favors bursty locality workloads, while close-page limits row-conflict penalties and can stabilize latency under random access. Adaptive implementations monitor hit/miss patterns, bank-level contention, and command bus pressure, then adjust close timing or row-retention heuristics per bank. The controller must reconcile this with timing constraints such as tRAS minimum, tFAW power windows, and bank-group turnaround rules, because aggressive row management can improve one metric while degrading global fairness or power integrity. DRAM inefficiency is multiplicative: one extra ACTIVATE, one unnecessary turnaround, one weak lane margin, or one refresh collision repeated across billions of accesses can dominate product tail latency and power.
Use Row-hit rate, effective command efficiency, and average activate/precharge overhead per request. as the opening signal, not the conclusion. A metric move only becomes actionable when paired with workload context, command traces, training telemetry, and evidence artifacts such as Row-buffer analytics report: FR-FCFS issue decisions, row-hit/miss timeline, and adaptive page-policy state transitions..
Memory-controller quality is measured by throughput and tail predictability under mixed traffic, not average bandwidth alone. Senior review quality comes from proving a complete chain: request pattern -> memory-state transition -> bottleneck mechanism -> smallest owner fix -> regression-safe validation.
For FR-FCFS, Row-Buffer Locality, and Page Policy Control, reports should explain why Row-hit rate, effective command efficiency, and average activate/precharge overhead per request. moved: fewer row misses, lower turnaround waste, better refresh placement, or stronger lane margin stability.
Strong reports include consistency checks: scheduler narrative matches command logs; PHY narrative matches margin sweeps; reliability narrative matches CE/UE trajectories.