Physical Design · All levels

Power Planning Ring and Stripes — Debug Playbook

Debug Playbook for Power Planning Ring and Stripes (Floorplanning Expanded).

On-call / interview prompt

Power Planning Ring and Stripes looks wrong — walk your first five debug steps.

diagram
CLOSURE CHAIN

1. METRIC     — name the failing report line (WNS, DRV, DRC count, IR %)
2. HYPOTHESIS — 2–3 likely causes ordered by probability
3. EXPERIMENT — one cheap check (corner, clock, report, map)
4. FIX        — minimal physical or constraint change
5. REGRESSION — what you re-run and what must not regress

Reference workflow

diagram
1. Locate droop clusters and classify edge versus interior behavior.
2. Sweep ring width and strap pitch in limited what-if runs.
3. Verify power-domain tie continuity and cut-layer via health.
4. Check routing blockage impact of tighter strap plan.
5. Select minimum change that clears IR threshold with acceptable route cost.

Mechanism to narrate

  • Separate symptom from root cause

  • Fix systematic clusters before one-offs

Common pitfalls

  • Random optimization without metric

  • Skipping regression after local fix

Staff-level debug discipline

For Power Planning Ring and Stripes, senior debug is branch-and-bound: reduce the search space quickly, keep experiments reversible, and avoid hiding a systematic issue behind one local fix.

Debug decision tree

  1. Reproduce the failure on the tagged database and exact analysis setup.

  2. Classify the failure as data issue, constraint issue, physical implementation issue, tool/methodology issue, or true design limitation.

  3. Run one cheap experiment that can falsify the leading hypothesis.

  4. Prefer a fix that improves a cluster over one that only hides the worst line.

  5. After the fix, re-check Early IR drop map with ring/stripe sensitivity and the likely regression surface: Placement feasibility, CTS insertion freedom, and route closure quality depend on this balance..

Escalation triggers

  • The failure crosses team ownership boundaries: RTL, synthesis, CAD, IP, package, or foundry.

  • The local fix consumes margin that another signoff domain needs.

  • The issue repeats across blocks, suggesting methodology or library root cause.

  • The remaining risk is silicon-facing: Undersized grids cause late IR/EM ECO churn; oversized grids create route/timing congestion penalties..

Deep dive: how this shows up in real closure

Floorplan quality is the earliest predictor of place-and-route pain.

Reports and artifacts to inspect

  • floorplan summary: core area, macro area, std-cell utilization

  • macro/channel review: pin-facing sides, halos, routing channels

  • power plan preview: ring width, strap pitch, follow-pin connectivity

  • trial route congestion: overflow around macro corners and pin fields

Mini case study

A 2 MB SRAM cluster is placed with pins facing the die edge. Trial route shows red overflow along the north edge. A senior PD answer is to rotate or mirror the SRAM, open the channel, and re-run trial route before attempting timing optimization.

Debug branches

  • If congestion is local to macro corners, inspect pin sides and halo width before reducing global utilization.

  • If IR is weak at the core edge, widen the ring or add edge straps before adding random decaps.

  • If timing paths cross the whole block, review pin assignment and macro orientation before post-route ECO.

Senior review question

Ask yourself: what single report line would prove this page's concept is either passing or failing?

What changes at 10+ years

  • You are expected to predict what your fix can break before running it.

  • You should recognize when the issue is methodology, not one block's implementation.

  • You should communicate risk in tapeout language: owner, evidence, impact, mitigation, and decision date.

Principal-level review bar

Deep subpage pages in this course should be read like real closure review material. For a 10+ year PD engineer, the bar is not remembering terminology; it is making a release-quality decision under ambiguity.

What excellent looks like

  • Names the failing metric, corner/mode, database tag, and analysis switches before proposing a fix.

  • Separates data, constraint, physical, tool, and methodology root causes instead of treating all failures as optimization problems.

  • Chooses experiments by information gain and reversibility, not by habit.

  • States regression blast radius across timing, route, power, PV, DFT, package, and tapeout manifest.

  • Turns recurring failures into methodology guardrails, dashboards, or checklist items.

Closure note template

diagram
STAFF / PRINCIPAL CLOSURE NOTE

Context:
  stage: <pre-CTS | post-CTS | post-route | post-fill | signoff>
  tag: <database / netlist / SDC / library stack>
  failing metric: <exact report line>
  affected scope: <block / hierarchy / path group / power domain / region>

Hypotheses:
  H1: <most likely physical or constraint mechanism>
  H2: <competing explanation>
  H3: <methodology or input-data issue>

Decision:
  next experiment: <cheap check that can falsify H1>
  fix candidate: <minimal reversible change>
  rollback trigger: <metric that says the fix is wrong>
  regression set: <timing / route / power / PV / DFT / package>
  escalation owner: <team or reviewer>

Tradeoffs a senior engineer must discuss

Technical tradeoff

Floorplan quality is the earliest predictor of place-and-route pain. Explain not only the preferred fix, but what margin or schedule you are spending to get it.

Cross-team tradeoff

  • What must RTL, synthesis, CAD, STA, DFT, package, IP, or foundry agree to before this decision is final?

  • Which artifact becomes the source of truth after the decision: report, waiver, manifest, ECO script, or methodology deck?

  • What is the cost of being wrong: one rerun, ECO churn, mask risk, performance loss, or silicon escape?

Leadership communication

diagram
"The current blocker is <metric> in <corner/mode/stage>. The leading cause is <mechanism>. I recommend <fix> because it is bounded and reversible. The regression surface is <domains>. If it fails, we escalate to <owner> with <evidence>."

Key takeaways

  • Always connect the concept back to a measurable signoff artifact.

  • A fix is not complete until you can name the regression checks.

Common pitfalls

  • Optimizing by habit instead of reading the current report.

  • Forgetting that a local fix can regress timing, routing, power, or PV elsewhere.