DRAM & Memory Design · All levels
Rows, Columns, and Subarray Granularity: Reports and Metrics
Reports and Metrics for Rows, Columns, and Subarray Granularity.
Reports and metrics
Reports and Metrics for Rows, Columns, and Subarray Granularity focuses on Effective tRCD/tRAS/tRP versus bitline length, wordline length, and local row size per subarray.. The purpose is to turn memory observations into mechanism-backed actions with explicit owners and release-safe validation.
Reports should explain why Effective tRCD/tRAS/tRP versus bitline length, wordline length, and local row size per subarray. moved, not simply that it moved. Require evidence that links the movement to command behavior, queue policy, PHY margin, or reliability controls.
Before/after trend
BEFORE / AFTER GRAPH - Rows, Columns, and Subarray Granularity
metric quality
^
| o target band
| o post-fix sweep
| o
| o baseline (failing)
+----------------------------------------------> iteration
evidence capture fix applied closure run
Use this view to prove improvement is causal, not accidental.Evidence matrix
DRAM EVIDENCE MATRIX - Rows, Columns, and Subarray Granularity
+-------------------------------+--------------------------------+--------------------------------+---------------------------+
| Evidence | Tells you | Does not prove | Next action |
+-------------------------------+--------------------------------+--------------------------------+---------------------------+
| row-hit/miss + ACT/PRE mix | locality and row-state cost | lane-level capture integrity | inspect training margins |
| queue age + class breakdown | fairness and starvation risk | command legality details | parse command timeline |
| JEDEC legality + bus timeline | timing-window pressure | root cause by itself | correlate with traffic map|
| eye / Vref / skew snapshots | PHY margin and drift behavior | controller policy quality | pair with schedule logs |
| CE/UE + scrub telemetry | reliability trajectory | immediate perf bottleneck only | map to hotspot addresses |
+-------------------------------+--------------------------------+--------------------------------+---------------------------+Track p50/p95/p99 latency and effective bandwidth together.
Include command and queue context alongside high-level counters.
Tag reports with firmware, timing profile, and thermal state.
Call out contradictory evidence instead of hiding it.
DRAM deep dive
Cell-array and subarray organization determines bitline delay, sensing margin, and locality-sensitive energy cost.
Concept diagram
ARRAY ORGANIZATION VIEW
rows x columns -> mats/subarrays -> local sense amps -> global I/O
physical distance shapes timing and energyMetric graph
ARRAY ACCESS COST SHARE
bitline settle delay ██████
sense/restore time █████
global routing overhead ███Reports and artifacts
subarray toggle heatmap
sense-amplifier utilization report
bitline RC delay audit
wordline coupling checklist
Mini case study
A dense address remap increased long-bitline activations, creating extra tRCD guardband and persistent tail-latency drift.
Debug branches
Map hot addresses to mats and subarray boundaries
Inspect sense-margin behavior under temperature corners
Evaluate row-mapping changes before voltage retuning
Senior review question
Ask: which latency, bandwidth, and reliability evidence proves this DRAM topic is closed under real traffic?
Key takeaways
Always tie controller and PHY counter shifts to application latency and throughput outcomes.
Lock firmware timing profile, thermal condition, and DIMM state before comparing DRAM captures.
Common pitfalls
Chasing peak bandwidth while ignoring p99 latency and fairness tails.
Changing timing guardbands without separating SI noise from scheduling issues.
Declaring closure without reliability gates, fault injection, and regression replay.
Report interpretation
A DRAM array is physically tiled into subarrays so each local wordline and bitline segment stays within a manageable RC envelope. Longer rows increase row-buffer capacity but lengthen wordline propagation and bitline loading, which increases ACTIVATE latency and sensing energy. Narrower subarrays improve local timing and noise margin but add peripheral overhead (local decoders, isolation devices, sense resources), reducing area efficiency. Column muxing then trades pin bandwidth and internal burst granularity against peripheral complexity. The final row/column partition is therefore not an abstract addressing choice; it is a first-order physical design knob that sets access latency, activation current profile, and manufacturability. DRAM inefficiency is multiplicative: one extra ACTIVATE, one unnecessary turnaround, one weak lane margin, or one refresh collision repeated across billions of accesses can dominate product tail latency and power.
Use Effective tRCD/tRAS/tRP versus bitline length, wordline length, and local row size per subarray. as the opening signal, not the conclusion. A metric move only becomes actionable when paired with workload context, command traces, training telemetry, and evidence artifacts such as Subarray sizing tradeoff sheet: row length, bitline RC, timing deltas, and die-area overhead..
Array organization sets the geometry of latency, bandwidth, and power before scheduler policy is even considered. Senior review quality comes from proving a complete chain: request pattern -> memory-state transition -> bottleneck mechanism -> smallest owner fix -> regression-safe validation.
For Rows, Columns, and Subarray Granularity, reports should explain why Effective tRCD/tRAS/tRP versus bitline length, wordline length, and local row size per subarray. moved: fewer row misses, lower turnaround waste, better refresh placement, or stronger lane margin stability.
Strong reports include consistency checks: scheduler narrative matches command logs; PHY narrative matches margin sweeps; reliability narrative matches CE/UE trajectories.