AI Accelerator Design · All levels
Input-Stationary Dataflow: Expanded Case Study
Expanded Case Study for Input-Stationary Dataflow.
Expanded case study
Expanded Case Study for Input-Stationary Dataflow is anchored on Activation reuse factor and input-buffer read traffic per tera-operations for representative convolution and GEMM layers.. Convert measurements into mechanism-backed decisions with clear owner accountability.
Use this page to rehearse incident closure: symptom intake, mechanism split, evidence request, owner assignment, bounded fix, and release decision.
Incident memo
ACCELERATOR REVIEW MEMO - Dataflow Architectures / Input-Stationary Dataflow
1. Symptom
- Failing metric: Activation reuse factor and input-buffer read traffic per tera-operations for representative convolution and GEMM layers.
- Workload or traffic slice: <name>
- First failing layer or stage: <operator, schedule, memory, runtime>
- Build and runtime tags: <compiler/firmware/runtime/hardware>
2. Mechanism hypothesis
- Primary mechanism: Input-stationary execution keeps activation tiles resident in local buffers or PE-adjacent storage while weights and partial sums stream through the compute fabric. This reduces repeated activation fetches from higher memory levels, which is valuable when feature-map bandwidth dominates energy. The design challenge is balancing local activation capacity with timely weight delivery and reduction routing so compute units remain occupied. Practical schedulers tune tile shape, preload depth, and multicast strategy to prevent contention that can erase expected reuse gains.
- Competing hypotheses: <dataflow mismatch, memory stalls, precision drift, thermal limits>
- Missing evidence: <counter packet, trace, replay, signoff data>
3. Proposed action
- Smallest reversible change: <mapping/runtime/policy/config>
- Expected movement: <throughput, p99 latency, perf-per-watt>
- Regression risk: correctness, quality, thermal, software compatibility
4. Signoff
- Required artifact: Traffic decomposition sheet showing activation residency, refill cadence, and bottleneck attribution by layer.
- Required owners: accelerator architect, memory hierarchy owner, compiler mapping owner, runtime scheduling owner
- Final decision: ship, bounded rollout, rollback, or escalateAI accelerator deep dive
Dataflow choices are durable architecture decisions that shape memory and scheduling cost.
Concept diagram
DATAFLOW DECISION
input/output/weight stationary
-> locality pattern
-> movement cost
-> throughput and powerMetric graph
DATAFLOW COST MIX
activation traffic ███████
weight traffic █████
partial-sum traffic ██████Metrics and artifacts to collect
reuse factor map
buffer pressure profile
NoC traffic mix
shape sensitivity analysis
Mini case study
A dataflow that won for convolution lost on attention-heavy batches due to activation movement pressure.
Debug branches
Segment by model family
Compare reuse vs movement
Re-check mapping assumptions under batch variance
Senior review question
Ask: which first-principles bottleneck class explains the symptom, and what artifact proves it reproducibly?
Key takeaways
Tie every accelerator claim to a reproducible workload slice and one primary metric trend.
Prefer bounded fixes with clear owner and rollback boundary over broad tuning bundles.
Common pitfalls
Optimizing synthetic kernels without production-shape validation.
Reading average latency while ignoring p95 and p99 behavior.
Declaring sparse or precision wins without fallback and quality evidence.
Principal accelerator review addendum
Input-Stationary Dataflow should be framed as a full-system behavior, not an isolated kernel trick. Production outcomes are set by model shape mix, compiler choices, runtime queueing policy, memory hierarchy limits, and silicon delivery margins.
Input-stationary execution keeps activation tiles resident in local buffers or PE-adjacent storage while weights and partial sums stream through the compute fabric. This reduces repeated activation fetches from higher memory levels, which is valuable when feature-map bandwidth dominates energy. The design challenge is balancing local activation capacity with timely weight delivery and reduction routing so compute units remain occupied. Practical schedulers tune tile shape, preload depth, and multicast strategy to prevent contention that can erase expected reuse gains. A useful explanation always ties observed symptom to a repeatable path where useful work was blocked, delayed, or diluted by overhead.
Use Activation reuse factor and input-buffer read traffic per tera-operations for representative convolution and GEMM layers. as an alarm, then anchor action using hard evidence such as Traffic decomposition sheet showing activation residency, refill cadence, and bottleneck attribution by layer..
Dataflow selection governs reuse, movement cost, and predictability across model shapes. Senior reviews expect a chain of proof: workload intent -> mapping -> hardware behavior -> product impact.
Use this addendum to force explicit owner assignment, bounded fixes, and reproducible evidence before declaring closure.