AI Accelerator Design · All levels

Hybrid Dataflow Selection

Dataflow Architectures: Hybrid policies choose different dataflows by operator class or shape regime instead of enforcing one stationary style globally. For example, output-stationary may win on deep reductions while weight-stationary performs better on high-filter-reuse layers, and input-stationary can help bandwidth-bound feature maps. The key is selecting switch points that account for retile overhead, control complexity, and compiler/runtime transition cost. Teams usually rely on profiling-guided heuristics or cost models that include both kernel efficiency and orchestration penalty.

What this topic teaches

Hybrid Dataflow Selection converts accelerator architecture concepts into release-ready engineering decisions. Hybrid policies choose different dataflows by operator class or shape regime instead of enforcing one stationary style globally. For example, output-stationary may win on deep reductions while weight-stationary performs better on high-filter-reuse layers, and input-stationary can help bandwidth-bound feature maps. The key is selecting switch points that account for retile overhead, control complexity, and compiler/runtime transition cost. Teams usually rely on profiling-guided heuristics or cost models that include both kernel efficiency and orchestration penalty.

Senior-engineer framing question

When End-to-end joules per inference and latency improvement from per-layer dataflow switching versus single-style baseline. regresses, can you isolate the first failing execution mechanism, collect decisive evidence, assign owners, and close with rollback-safe validation?

diagram
ACCELERATOR EXECUTION FLOW - Hybrid Dataflow Selection

request ingress and model metadata
      |
      v
graph lowering and kernel selection
      |
      v
tile/dataflow scheduling and memory placement
      |
      v
tensor execution + synchronization barriers
      |
      v
result assembly + quality/SLA validation
      |
      v
release decision and rollback guardrails

Evidence to collect

  • Primary metric: End-to-end joules per inference and latency improvement from per-layer dataflow switching versus single-style baseline..

  • Primary artifact: Dataflow policy playbook with layer-wise recommendations, transition rules, and expected gains..

  • Owners to include: system architect, compiler/runtime owner, performance modeling owner, production inference lead.

  • One reproducible failing workload and one stable comparator run.

  • One fixed-metadata run with compiler/runtime/hardware tags locked.

Bandwidth lens

diagram
BANDWIDTH LENS - Hybrid Dataflow Selection

working-set pressure
  ^
  |                saturation zone
  |          ----------------------------
  |      o   unstable tail latency
  |   o      tuning candidate
  | o        baseline behavior
  +-------------------------------------> optimization iteration

Primary metric tracked:
End-to-end joules per inference and latency improvement from per-layer dataflow switching versus single-style baseline.

Ownership layers

diagram
OWNERSHIP LAYERS - Hybrid Dataflow Selection

+----------------------+--------------------------------+--------------------------------+
| Team                 | Primary responsibility         | Closure artifact               |
+----------------------+--------------------------------+--------------------------------+
| system architect | mechanism and architecture intent| design rationale + tradeoffs   |
| compiler/runtime owner | mapping, runtime, and execution   | profile traces + bottleneck map|
| performance modeling owner | correctness, risk, and signoff    | test report + closure memo     |
+----------------------+--------------------------------+--------------------------------+

Key takeaways

  • Start with mechanism classification before changing tuning knobs.

  • Use one proving artifact for each major claim in review discussions.

  • Close with explicit owners, validation matrix, and rollback criteria.

Common pitfalls

  • Optimizing only peak throughput while p99 latency or quality regresses.

  • Mixing evidence captured from mismatched runtime or thermal conditions.

  • Declaring closure without production-like replay and guardrail checks.

AI accelerator deep dive

Dataflow choices are durable architecture decisions that shape memory and scheduling cost.

Concept diagram

diagram
DATAFLOW DECISION

input/output/weight stationary
  -> locality pattern
  -> movement cost
  -> throughput and power

Metric graph

diagram
DATAFLOW COST MIX

activation traffic   ███████
weight traffic       █████
partial-sum traffic  ██████

Metrics and artifacts to collect

  • reuse factor map

  • buffer pressure profile

  • NoC traffic mix

  • shape sensitivity analysis

Mini case study

A dataflow that won for convolution lost on attention-heavy batches due to activation movement pressure.

Debug branches

  • Segment by model family

  • Compare reuse vs movement

  • Re-check mapping assumptions under batch variance

Senior review question

Ask: which first-principles bottleneck class explains the symptom, and what artifact proves it reproducibly?

Key takeaways

  • Tie every accelerator claim to a reproducible workload slice and one primary metric trend.

  • Prefer bounded fixes with clear owner and rollback boundary over broad tuning bundles.

Common pitfalls

  • Optimizing synthetic kernels without production-shape validation.

  • Reading average latency while ignoring p95 and p99 behavior.

  • Declaring sparse or precision wins without fallback and quality evidence.