Computer Architecture · All levels

Accelerator Design Patterns — Design Space Exploration

Design Space Exploration for Accelerator Design Patterns (Accelerator Architectures).

Design space exploration

For Accelerator Design Patterns, senior architects do not pick one answer — they map the design space, estimate metric movement, and choose based on product constraints.

Option A — conservative

  • Conservative: helps lower risk

  • Risk: less upside

  • Validate with: baseline suite

Option B — balanced

  • Balanced: helps good perf/watt

  • Risk: may miss peak

  • Validate with: multi-workload sweep

Option C — aggressive

  • Aggressive: helps peak wins

  • Risk: PPA/DV risk

  • Validate with: stress suite

Option D — software-first

  • Software-first: helps low silicon

  • Risk: fragile

  • Validate with: controlled apps

diagram
DESIGN SPACE — Accelerator Design Patterns
low risk -> balanced -> aggressive
with software-first as alternate axis

Common pitfalls

  • Aggressive hardware before workload proof

  • Balanced by habit without numbers

Architecture deep dive

Accelerators win on locality and bandwidth contracts, not peak OPS alone.

Concept diagram

diagram
ACCELERATOR DATAFLOW

Host CPU ── commands ──► Queue / scheduler
   ▲                         │
   │ completion              ▼
Coherent memory ◄── DMA ── Local SRAM ──► Compute array
                         ▲       │
                         └ tiles ┘

Peak TOPS matters only when data reaches the array at the needed rate.

Metric graph

diagram
UTILIZATION BREAKDOWN

compute active   ██████████████████  58%
DMA wait         ██████████          31%
host sync        █████               15%
cache/coherency  ████                12%
idle bubbles     ███████             22%

Low utilization is usually a system integration problem.

Metrics and artifacts

  • accelerator utilization

  • DMA bandwidth

  • kernel launch overhead

  • coherency invalidation rate

Mini case study

NPU met TOPs target but end-to-end inference slow — DMA and weight fetch dominated. Architecture added on-chip SRAM tile and double-buffering.

Debug branches

  • If util low, check launch overhead and host sync first.

  • If BW high, examine weight layout and sparsity support.

Senior review question

Ask: what single metric would prove this concept is working or failing on your workload?

Key takeaways

  • Connect every architecture claim to a workload and measurable metric.

  • State verification and PPA impact before proposing design changes.

Common pitfalls

  • Feature-driven design without MPKI/IPC/bandwidth evidence.

  • Ignoring coherency and NoC traffic in cache and accelerator sizing.

Study notes

Re-read this topic with one concrete workload.