Computer Architecture · All levels

Bottleneck Analysis Framework — Pitfalls & Red Flags

Pitfalls & Red Flags for Bottleneck Analysis Framework (Performance Analysis).

Common mistakes

  • Optimizing without naming workload, metric, and model setup

  • Local fix that regresses neighboring metrics

  • Skipping documented checklist before architecture review

Red flags in reviews

  • Cannot explain worst report line

  • No regression list after proposed fix

  • Waiver requested without cluster analysis

Failure modes seen in real product programs

  • A performance win is accepted on one benchmark while product workloads regress.

  • A simulation result is trusted without matching PMU counter definitions.

  • A microarchitecture knob hides a workload-specific issue but creates verification and PPA debt.

  • A local improvement in Bottleneck Analysis Framework regresses Architecture spec updates and performance closure schedule..

How a senior engineer recovers

  1. Freeze the evidence: workload, model/RTL tag, counter setup, trace, and simulator switches.

  2. Name the real owner and approval path.

  3. Convert the lesson into a checklist item, regression, or methodology guardrail.

Pitfall map

diagram
TRADEOFF MATRIX — Bottleneck Analysis Framework

+----------------------+----------------------+----------------------+----------------------+
| Option               | Helps                | Can hurt             | Validation needed    |
+----------------------+----------------------+----------------------+----------------------+
| Larger / wider block | peak perf, miss rate | area, power, timing  | workload sweep       |
| Smarter policy       | hit rate, QoS, IPC   | verification risk    | corner cases + PMU   |
| More buffering       | latency tails, stalls| deadlock, leakage    | stress traffic tests |
| Software contract    | locality, ordering   | portability, APIs    | production workload  |
+----------------------+----------------------+----------------------+----------------------+

Senior rule: pick the smallest change that proves or disproves the mechanism.

Architecture deep dive

PMU evidence beats intuition for architecture decisions.

Concept diagram

diagram
TOP-DOWN PERFORMANCE METHOD

Total cycles
 ├─ Retiring useful work
 ├─ Frontend bound
 ├─ Bad speculation
 ├─ Backend core bound
 └─ Backend memory bound

Only after classification should you propose cache, branch, pipeline, or NoC changes.

Metric graph

diagram
ROOFLINE SKETCH

Performance
  ^
  |                     compute roof
  |-------------------------------
  |                   /
  |                 /
  |               /   ● workload A (compute-bound)
  |             /
  |   ● workload B (memory-bound)
  +---------------------------------> arithmetic intensity
        memory bandwidth slope

Metrics and artifacts

  • PMU event sets

  • roofline chart

  • top-down stall breakdown

  • workload sensitivity matrix

Mini case study

Team proposed wider SIMD but roofline showed memory-bound kernel — bandwidth upgrade and locality fix delivered 2× speedup at lower area cost.

Debug branches

  • If counters disagree with sim, align workload and warmup.

  • If bottleneck unclear, use top-down method before microarch tweaks.

Senior review question

Ask: what single metric would prove this concept is working or failing on your workload?

Key takeaways

  • Connect every architecture claim to a workload and measurable metric.

  • State verification and PPA impact before proposing design changes.

Common pitfalls

  • Feature-driven design without MPKI/IPC/bandwidth evidence.

  • Ignoring coherency and NoC traffic in cache and accelerator sizing.

Study notes

Re-read this topic with one concrete workload.