Computer Architecture · All levels

Memory Locality and Hierarchy Co-Design — Software / Programmer View

Software / Programmer View for Memory Locality and Hierarchy Co-Design (Accelerator Architectures).

Software and programmer view

Compiler, runtime, and memory layout determine whether peak OPS becomes throughput.

What programmers feel

  • Tail latency under contention

  • Layout-sensitive cliffs

  • Rare ordering bugs

API / ABI / runtime implications

  • Alignment/allocation

  • Pinning/domain awareness

  • Fence semantics

Compiler and runtime interaction

  • Prefetch sensitivity

  • Padding/layout

  • Allocator behavior

Software-side mitigations

  • Improve locality

  • Reduce false sharing

  • Expose counters

diagram
SOFTWARE — Memory Locality and Hierarchy Co-Design
per-core counters beat shared hot counters for coherence traffic

Architecture deep dive

Accelerators win on locality and bandwidth contracts, not peak OPS alone.

Concept diagram

diagram
ACCELERATOR DATAFLOW

Host CPU ── commands ──► Queue / scheduler
   ▲                         │
   │ completion              ▼
Coherent memory ◄── DMA ── Local SRAM ──► Compute array
                         ▲       │
                         └ tiles ┘

Peak TOPS matters only when data reaches the array at the needed rate.

Metric graph

diagram
UTILIZATION BREAKDOWN

compute active   ██████████████████  58%
DMA wait         ██████████          31%
host sync        █████               15%
cache/coherency  ████                12%
idle bubbles     ███████             22%

Low utilization is usually a system integration problem.

Metrics and artifacts

  • accelerator utilization

  • DMA bandwidth

  • kernel launch overhead

  • coherency invalidation rate

Mini case study

NPU met TOPs target but end-to-end inference slow — DMA and weight fetch dominated. Architecture added on-chip SRAM tile and double-buffering.

Debug branches

  • If util low, check launch overhead and host sync first.

  • If BW high, examine weight layout and sparsity support.

Senior review question

Ask: what single metric would prove this concept is working or failing on your workload?

Key takeaways

  • Connect every architecture claim to a workload and measurable metric.

  • State verification and PPA impact before proposing design changes.

Common pitfalls

  • Feature-driven design without MPKI/IPC/bandwidth evidence.

  • Ignoring coherency and NoC traffic in cache and accelerator sizing.

Study notes

Re-read this topic with one concrete workload.