AI Accelerator Design · All levels

Scratchpad vs Cache: Managed Locality Tradeoffs: Inputs and Outputs

Inputs and Outputs for Scratchpad vs Cache: Managed Locality Tradeoffs.

Inputs and outputs contract

Inputs and Outputs for Scratchpad vs Cache: Managed Locality Tradeoffs is anchored on Delivered throughput change and miss or spill overhead when moving a kernel between cache-managed and scratchpad-managed execution.. Convert measurements into mechanism-backed decisions with clear owner accountability.

diagram
INPUTS
  - workload profile and SLA target
  - model precision and quality thresholds
  - compiler/runtime/firmware metadata
  - hardware operating envelope assumptions

OUTPUTS
  - evidence-backed bottleneck classification
  - owner-signed mitigation proposal
  - validation matrix with rollback triggers
  - release recommendation

Ownership split

diagram
OWNERSHIP LAYERS - Scratchpad vs Cache: Managed Locality Tradeoffs

+----------------------+--------------------------------+--------------------------------+
| Team                 | Primary responsibility         | Closure artifact               |
+----------------------+--------------------------------+--------------------------------+
| compiler/runtime architect | mechanism and architecture intent| design rationale + tradeoffs   |
| kernel performance engineer | mapping, runtime, and execution   | profile traces + bottleneck map|
| hardware cache designer | correctness, risk, and signoff    | test report + closure memo     |
+----------------------+--------------------------------+--------------------------------+

AI accelerator deep dive

Memory hierarchy discipline sets the practical compute ceiling for AI accelerators.

Concept diagram

diagram
MEMORY HIERARCHY VIEW

register/SRAM -> shared buffers -> NoC -> HBM
  locality quality decides how long compute stays fed

Metric graph

diagram
MEMORY WALL SIGNALS

HBM near-saturation   ███████████
NoC backpressure      ███████
compute idle fraction █████

Metrics and artifacts to collect

  • SRAM hit ratio

  • HBM utilization timeline

  • bank-conflict hotspots

  • NoC queue pressure

Mini case study

HBM channels saturated under burst traffic while compute occupancy dropped, proving a memory-bound regime.

Debug branches

  • Separate locality vs bandwidth limits

  • Quantify bank conflicts

  • Tune tiling before resizing compute arrays

Senior review question

Ask: which first-principles bottleneck class explains the symptom, and what artifact proves it reproducibly?

Key takeaways

  • Tie every accelerator claim to a reproducible workload slice and one primary metric trend.

  • Prefer bounded fixes with clear owner and rollback boundary over broad tuning bundles.

Common pitfalls

  • Optimizing synthetic kernels without production-shape validation.

  • Reading average latency while ignoring p95 and p99 behavior.

  • Declaring sparse or precision wins without fallback and quality evidence.

Handoff explanation

Inputs should include workload profile, model revision, compiler/runtime versions, and platform power mode.

Outputs must include actionable interpretation of Delivered throughput change and miss or spill overhead when moving a kernel between cache-managed and scratchpad-managed execution., required artifacts (Kernel locality decision matrix comparing cache and scratchpad policy by operator class.), owner, and validation scope.

The ideal handoff packet is reproducible: fixed seeds, explicit baseline, and rejected alternatives.