AI Accelerator Design · All levels
Scratchpad vs Cache: Managed Locality Tradeoffs: Reports and Metrics
Reports and Metrics for Scratchpad vs Cache: Managed Locality Tradeoffs.
Reports and metrics
Reports and Metrics for Scratchpad vs Cache: Managed Locality Tradeoffs is anchored on Delivered throughput change and miss or spill overhead when moving a kernel between cache-managed and scratchpad-managed execution.. Convert measurements into mechanism-backed decisions with clear owner accountability.
A useful report explains why movement happened, not only that movement happened.
Evidence matrix
EVIDENCE MATRIX - Scratchpad vs Cache: Managed Locality Tradeoffs
+-----------------------------+--------------------------------+--------------------------------+---------------------------+
| Evidence | Tells you | Does not prove | Next action |
+-----------------------------+--------------------------------+--------------------------------+---------------------------+
| occupancy + timeline traces | where utilization is lost | precise root cause | map to memory and schedule|
| cache/SRAM/bandwidth stats | data movement pressure | model-level quality impact | correlate with quality run|
| counter + profile alignment | bottleneck class confidence | rollout safety | run full regression matrix|
| thermal/power telemetry | sustained operating envelope | correctness closure | pair with verification |
| before/after scenario pack | mitigation movement | long-tail stability | execute guardrail replay |
+-----------------------------+--------------------------------+--------------------------------+---------------------------+Track Delivered throughput change and miss or spill overhead when moving a kernel between cache-managed and scratchpad-managed execution. on representative production workloads.
Include build/runtime metadata in every report header.
Correlate throughput, latency, and quality before rollout decisions.
Call out contradictory evidence explicitly.
AI accelerator deep dive
Memory hierarchy discipline sets the practical compute ceiling for AI accelerators.
Concept diagram
MEMORY HIERARCHY VIEW
register/SRAM -> shared buffers -> NoC -> HBM
locality quality decides how long compute stays fedMetric graph
MEMORY WALL SIGNALS
HBM near-saturation ███████████
NoC backpressure ███████
compute idle fraction █████Metrics and artifacts to collect
SRAM hit ratio
HBM utilization timeline
bank-conflict hotspots
NoC queue pressure
Mini case study
HBM channels saturated under burst traffic while compute occupancy dropped, proving a memory-bound regime.
Debug branches
Separate locality vs bandwidth limits
Quantify bank conflicts
Tune tiling before resizing compute arrays
Senior review question
Ask: which first-principles bottleneck class explains the symptom, and what artifact proves it reproducibly?
Key takeaways
Tie every accelerator claim to a reproducible workload slice and one primary metric trend.
Prefer bounded fixes with clear owner and rollback boundary over broad tuning bundles.
Common pitfalls
Optimizing synthetic kernels without production-shape validation.
Reading average latency while ignoring p95 and p99 behavior.
Declaring sparse or precision wins without fallback and quality evidence.
Report interpretation
Scratchpads expose explicit software control over placement, prefetch, and eviction, enabling predictable latency when access patterns are regular and compiler scheduling is mature. Hardware caches reduce software complexity and handle irregular reuse patterns automatically, but they can introduce nondeterministic misses and contention under multi-kernel interference. Most production systems blend both: critical tiles are pinned or staged through scratchpads while less predictable data uses cache paths. The choice should be based on measured reuse distance, synchronization pattern, and development cost, not ideology. A useful explanation always ties observed symptom to a repeatable path where useful work was blocked, delayed, or diluted by overhead.
Use Delivered throughput change and miss or spill overhead when moving a kernel between cache-managed and scratchpad-managed execution. as an alarm, then anchor action using hard evidence such as Kernel locality decision matrix comparing cache and scratchpad policy by operator class..
Memory hierarchy quality determines whether compute remains fed or sits idle behind bandwidth walls. Senior reviews expect a chain of proof: workload intent -> mapping -> hardware behavior -> product impact.
For Scratchpad vs Cache: Managed Locality Tradeoffs, reports should explain why Delivered throughput change and miss or spill overhead when moving a kernel between cache-managed and scratchpad-managed execution. moved and which path consumed budget first.