AI for VLSI · All levels

On-Chip Memory Hierarchy

Accelerator Architecture: Register files, local SRAM, and shared buffers determine data reuse efficiency and dominate accelerator energy profile.

What this topic teaches

On-Chip Memory Hierarchy turns AI concepts into VLSI-ready engineering decisions. Register files, local SRAM, and shared buffers determine data reuse efficiency and dominate accelerator energy profile. The practical challenge is proving value with reproducible evidence, bounded risk, and explicit ownership.

The senior-engineer question

When SRAM hit ratio, buffer occupancy, and energy per data movement moves, can you identify the failing layer, the mechanism, the artifact, and the owner who can close risk with a measurable fix?

diagram
AI-VLSI FLOW — On-Chip Memory Hierarchy

problem framing
      |
      v
data + model definition
      |
      v
training / optimization
      |
      v
compute-hardware mapping
      |
      v
deployment + validation

Primary metric: SRAM hit ratio, buffer occupancy, and energy per data movement

Picture the system

Start each review with an architecture sketch before opening dashboards. These diagrams are designed for design reviews and interview whiteboards.

On-chip memory tiers

diagram
MEMORY HIERARCHY

register file -> local SRAM -> shared SRAM -> DRAM

Closer memory: lower latency/energy
Farther memory: larger capacity, higher cost.

Tensor and data path

diagram
TENSOR / PIPELINE MAP — On-Chip Memory Hierarchy

feature source -> preprocessing -> tensorized input
      |                             |
      +---- shape + scale checks ---+
                    |
                    v
            model execution / inference

Shape and scaling discipline decides correctness and portability.

Training and update loop

diagram
TRAINING TIMELINE — On-Chip Memory Hierarchy

time --->
data batch      __/--/--/--/--/--/--/--
forward pass    ____/--/--/--/--/--/---
backward pass   ________/--/--/--/-----
optimizer step  ____________/--/--/----
eval checkpoint _____________/--/-------

Convergence depends on stable loop timing and signal quality.

Compute limit lens

diagram
ROOFLINE LENS — On-Chip Memory Hierarchy

performance
   ^
   |                compute bound region
   |               /
   |              /
   |-------------/---------------- memory bound region
   +----------------------------------------------> operational intensity

Use this to decide compute optimization vs memory optimization.

Ownership layers

diagram
AI-VLSI OWNERSHIP LAYERS — On-Chip Memory Hierarchy

layer                    owns                               typical failure
---------------------    --------------------------------   ----------------------------
problem framing          metric + acceptance criteria       wrong objective target
model + training         representation + optimization      unstable or biased model
hardware mapping         dataflow + memory + precision      bandwidth stalls / mismatch
deployment stack         runtime + firmware + drivers       latency jitter / incompatibility
governance               monitoring + rollback + signoff    silent drift in production

Evidence to collect

  • Primary metric: SRAM hit ratio, buffer occupancy, and energy per data movement.

  • Primary artifact: memory hierarchy diagram, access heatmap, and buffer sizing report.

  • Owners to bring into review: memory architect, physical design owner, performance team.

  • One workload slice where behavior regressed and one where it held.

  • One profile view that separates model issue from runtime/hardware issue.

Ownership map

diagram
OWNERSHIP MAP — On-Chip Memory Hierarchy

artifact focus         owner
------------------     ----------------------------
modeling          memory architect
architecture      physical design owner
integration       performance team

Production issues happen when ownership is assumed, not declared.

Subpages in this topic

Each topic is taught across mechanism, inputs/outputs, reports, debug, worked example, pitfalls, interview, checklist, theory, design space, expanded case study, walkthrough, comparison matrix, software view, and silicon impact.

Key takeaways

  • Always map ML metrics to engineering decisions and release risk.

  • Separate data/model issues from hardware/runtime bottlenecks before fixing.

  • Use reproducible artifacts and owner signoff for every rollout decision.

Common pitfalls

  • Benchmark wins with no signoff correlation.

  • Ignoring calibration and drift when deploying quantized models.

  • Shipping without a rollback and ownership matrix.

AI-VLSI deep dive

Dataflow and memory hierarchy choices dominate practical accelerator performance.

Concept diagram

diagram
ACCELERATOR STACK

MAC array -> dataflow map -> SRAM hierarchy -> NoC/DRAM

Metric graph

diagram
BOTTLENECK MIX

array underutilization ████
SRAM pressure          ███████
NoC/DRAM congestion    ██████

Reports and artifacts

  • array utilization

  • reuse factor report

  • buffer hit analysis

  • NoC/DRAM bottleneck trace

Mini case study

A dataflow change improved reuse but exposed NoC congestion, requiring route-policy updates.

Debug branches

  • Map reuse factors

  • Inspect memory stall sources

  • Profile NoC arbitration

Senior review question

Ask: what evidence connects this ML claim to a concrete VLSI workflow decision and owner signoff?

Key takeaways

  • Every AI claim should map to a measurable engineering outcome.

  • Validate both model quality and hardware/runtime feasibility before adoption.

Common pitfalls

  • Optimizing benchmark metrics that do not correlate with signoff goals.

  • Ignoring data drift and calibration after deployment.

  • Shipping ML workflows without clear rollback ownership.

Execution drill pack 1

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 1

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>

Execution drill pack 2

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 2

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>

Execution drill pack 3

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 3

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>

Execution drill pack 4

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 4

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>

Execution drill pack 5

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 5

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>

Execution drill pack 6

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 6

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>

Execution drill pack 7

Use this pack to rehearse AI-for-VLSI decision making on ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy: metric framing, mechanism proof, hardware implications, and release safety.

Evidence checklist

  • Metric context includes workload, dataset slice, and revision tags.

  • Mechanism explanation links model behavior to observed outcome.

  • Hardware/runtime feasibility is profiled, not assumed.

  • Owner and rollback path are documented before rollout.

Review prompts

  1. Which decision will this model output influence?

  2. What is the first failing layer when metric regresses?

  3. Which owner applies the smallest reversible fix?

  4. What validation matrix is required before deployment?

Evidence capsule

diagram
AI-VLSI EVIDENCE CAPSULE 7

PATH: ai-vlsi/accelerator-architecture/on-chip-memory-hierarchy
WORKLOAD SLICE: <name>
PRIMARY METRIC: <value/trend>
FIRST FAILING LAYER: <data/model/runtime/hardware>
OWNER: <name>
PRIMARY ARTIFACT: <report/profile/dashboard>
DECISION: <ship / rollback / escalate>