Computer Architecture · All levels
Memory Locality and Hierarchy Co-Design — Software / Programmer View
Software / Programmer View for Memory Locality and Hierarchy Co-Design (Accelerator Architectures).
Software and programmer view
Compiler, runtime, and memory layout determine whether peak OPS becomes throughput.
What programmers feel
Tail latency under contention
Layout-sensitive cliffs
Rare ordering bugs
API / ABI / runtime implications
Alignment/allocation
Pinning/domain awareness
Fence semantics
Compiler and runtime interaction
Prefetch sensitivity
Padding/layout
Allocator behavior
Software-side mitigations
Improve locality
Reduce false sharing
Expose counters
SOFTWARE — Memory Locality and Hierarchy Co-Design
per-core counters beat shared hot counters for coherence trafficArchitecture deep dive
Accelerators win on locality and bandwidth contracts, not peak OPS alone.
Concept diagram
ACCELERATOR DATAFLOW
Host CPU ── commands ──► Queue / scheduler
▲ │
│ completion ▼
Coherent memory ◄── DMA ── Local SRAM ──► Compute array
▲ │
└ tiles ┘
Peak TOPS matters only when data reaches the array at the needed rate.Metric graph
UTILIZATION BREAKDOWN
compute active ██████████████████ 58%
DMA wait ██████████ 31%
host sync █████ 15%
cache/coherency ████ 12%
idle bubbles ███████ 22%
Low utilization is usually a system integration problem.Metrics and artifacts
accelerator utilization
DMA bandwidth
kernel launch overhead
coherency invalidation rate
Mini case study
NPU met TOPs target but end-to-end inference slow — DMA and weight fetch dominated. Architecture added on-chip SRAM tile and double-buffering.
Debug branches
If util low, check launch overhead and host sync first.
If BW high, examine weight layout and sparsity support.
Senior review question
Ask: what single metric would prove this concept is working or failing on your workload?
Key takeaways
Connect every architecture claim to a workload and measurable metric.
State verification and PPA impact before proposing design changes.
Common pitfalls
Feature-driven design without MPKI/IPC/bandwidth evidence.
Ignoring coherency and NoC traffic in cache and accelerator sizing.
Study notes
Re-read this topic with one concrete workload.