Computer Architecture · All levels

Accelerator Design Patterns — Silicon & PPA Impact

Silicon & PPA Impact for Accelerator Design Patterns (Accelerator Architectures).

Silicon, power, area, and timing impact

SRAM tiles, DMA engines, and NoC pressure dominate accelerator silicon economics.

Area drivers

  • Buffers/tables/SRAM

  • Bypass and issue width wiring

  • Coherency metadata

Power drivers

  • Activity factor

  • SRAM energy

  • Wake-up bursts

Timing and frequency impact

  • Critical path movement

  • Macro distance

  • Frequency pressure

PD and floorplan consequences

  • Place hot structures near consumers

  • Macro placement constraints

  • NoC congestion

Verification burden

  • More states/policies

  • Ordering regressions

  • Traceable workload proof

diagram
PPA — Accelerator Design Patterns
area/power/timing/verif all workload-dependent

Key takeaways

  • No architecture signoff without PPA statement

  • PD latency budget can force architecture change

Architecture deep dive

Accelerators win on locality and bandwidth contracts, not peak OPS alone.

Concept diagram

diagram
ACCELERATOR DATAFLOW

Host CPU ── commands ──► Queue / scheduler
   ▲                         │
   │ completion              ▼
Coherent memory ◄── DMA ── Local SRAM ──► Compute array
                         ▲       │
                         └ tiles ┘

Peak TOPS matters only when data reaches the array at the needed rate.

Metric graph

diagram
UTILIZATION BREAKDOWN

compute active   ██████████████████  58%
DMA wait         ██████████          31%
host sync        █████               15%
cache/coherency  ████                12%
idle bubbles     ███████             22%

Low utilization is usually a system integration problem.

Metrics and artifacts

  • accelerator utilization

  • DMA bandwidth

  • kernel launch overhead

  • coherency invalidation rate

Mini case study

NPU met TOPs target but end-to-end inference slow — DMA and weight fetch dominated. Architecture added on-chip SRAM tile and double-buffering.

Debug branches

  • If util low, check launch overhead and host sync first.

  • If BW high, examine weight layout and sparsity support.

Senior review question

Ask: what single metric would prove this concept is working or failing on your workload?

Key takeaways

  • Connect every architecture claim to a workload and measurable metric.

  • State verification and PPA impact before proposing design changes.

Common pitfalls

  • Feature-driven design without MPKI/IPC/bandwidth evidence.

  • Ignoring coherency and NoC traffic in cache and accelerator sizing.

Study notes

Re-read this topic with one concrete workload.