Physical Design · All levels

Via Planning and Via Ladder Reliability

Via Planning and Via Ladder Reliability — physical design implementation and signoff.

On-call / interview prompt

EM analysis flags high-current stripes with sparse vias. How do you decide between local via ladders, strap redesign, or current redistribution?

diagram
CLOSURE CHAIN

1. METRIC     — name the failing report line (WNS, DRV, DRC count, IR %)
2. HYPOTHESIS — 2–3 likely causes ordered by probability
3. EXPERIMENT — one cheap check (corner, clock, report, map)
4. FIX        — minimal physical or constraint change
5. REGRESSION — what you re-run and what must not regress

Topic overview

Engineer via strategy for resistance, EM robustness, and manufacturable vertical connectivity.

Mechanism to narrate

  • Section: Routing

  • Primary artifact: see Reports subpage

  • Downstream dependency: Power integrity and lifetime reliability signoff depend on via robustness.

Staff/principal ownership model

Own Via Planning and Via Ladder Reliability as a release decision, not a page of notes. A senior PD engineer names the metric, the physical mechanism, the cross-team dependency, and the smallest evidence-producing experiment.

diagram
STAFF REVIEW MEMO — Routing / Via Planning and Via Ladder Reliability

1. Current state
   - Failing / watched metric: Routing closure dashboard
   - Database tag, corner/mode, tool version: <fill before review>
   - Physical scope: block, hierarchy, macro region, clock domain, or net class

2. Root-cause hypothesis
   - Most likely mechanism: <name physical or constraint mechanism>
   - Competing hypothesis: <name the second plausible cause>
   - Evidence still missing: <report/map/schematic/check>

3. Proposed action
   - Minimal reversible fix: <physical, constraint, ECO, or methodology change>
   - Expected improvement: <metric delta>
   - Regression risk: Under-designed via structures can cause EM failures and silicon reliability issues.

4. Regression and signoff
   - Re-run: Routing closure dashboard
   - Must not regress: Power integrity and lifetime reliability signoff depend on via robustness.
   - Decision owner: PD owner

Sub-lessons in this topic

  1. mechanism — Mechanism

  2. inputs-outputs — Inputs & Outputs

  3. reports — Reports & Metrics

  4. debug-playbook — Debug Playbook

  5. worked-example — Worked Example

  6. pitfalls — Pitfalls & Red Flags

  7. interview — Interview Drills

  8. checklist — Review Checklist

Related topics

Key takeaways

  • Master Via Planning and Via Ladder Reliability through reports, not GUI habit.

Deep dive: how this shows up in real closure

Routing decides whether the placed design is manufacturable and extractable.

Reports and artifacts to inspect

  • global route overflow by layer and region

  • detailed route DRC/DRV summary grouped by rule

  • antenna report: gate area, metal area, ratio, suggested diode

  • post-route extraction delta on critical nets

Mini case study

Detailed route completes with 4,000 spacing violations around a macro channel. The design is not 'routed'. Cluster by rule and region, fix the channel or layer plan, then re-route locally.

Debug branches

  • If antenna violations cluster on long lower-metal routes, jump to upper metal earlier or insert diode cells.

  • If via violations dominate, inspect via arrays and preferred via stacks, especially on power transitions.

  • If crosstalk worsens timing, space, shield, or layer-promote the victim nets.

Senior review question

Ask yourself: what single report line would prove this page's concept is either passing or failing?

What changes at 10+ years

  • You are expected to predict what your fix can break before running it.

  • You should recognize when the issue is methodology, not one block's implementation.

  • You should communicate risk in tapeout language: owner, evidence, impact, mitigation, and decision date.

Principal-level review bar

Lesson pages in this course should be read like real closure review material. For a 10+ year PD engineer, the bar is not remembering terminology; it is making a release-quality decision under ambiguity.

What excellent looks like

  • Names the failing metric, corner/mode, database tag, and analysis switches before proposing a fix.

  • Separates data, constraint, physical, tool, and methodology root causes instead of treating all failures as optimization problems.

  • Chooses experiments by information gain and reversibility, not by habit.

  • States regression blast radius across timing, route, power, PV, DFT, package, and tapeout manifest.

  • Turns recurring failures into methodology guardrails, dashboards, or checklist items.

Closure note template

diagram
STAFF / PRINCIPAL CLOSURE NOTE

Context:
  stage: <pre-CTS | post-CTS | post-route | post-fill | signoff>
  tag: <database / netlist / SDC / library stack>
  failing metric: <exact report line>
  affected scope: <block / hierarchy / path group / power domain / region>

Hypotheses:
  H1: <most likely physical or constraint mechanism>
  H2: <competing explanation>
  H3: <methodology or input-data issue>

Decision:
  next experiment: <cheap check that can falsify H1>
  fix candidate: <minimal reversible change>
  rollback trigger: <metric that says the fix is wrong>
  regression set: <timing / route / power / PV / DFT / package>
  escalation owner: <team or reviewer>

Tradeoffs a senior engineer must discuss

Technical tradeoff

Routing decides whether the placed design is manufacturable and extractable. Explain not only the preferred fix, but what margin or schedule you are spending to get it.

Cross-team tradeoff

  • What must RTL, synthesis, CAD, STA, DFT, package, IP, or foundry agree to before this decision is final?

  • Which artifact becomes the source of truth after the decision: report, waiver, manifest, ECO script, or methodology deck?

  • What is the cost of being wrong: one rerun, ECO churn, mask risk, performance loss, or silicon escape?

Leadership communication

diagram
"The current blocker is <metric> in <corner/mode/stage>. The leading cause is <mechanism>. I recommend <fix> because it is bounded and reversible. The regression surface is <domains>. If it fails, we escalate to <owner> with <evidence>."

Key takeaways

  • Always connect the concept back to a measurable signoff artifact.

  • A fix is not complete until you can name the regression checks.

Common pitfalls

  • Optimizing by habit instead of reading the current report.

  • Forgetting that a local fix can regress timing, routing, power, or PV elsewhere.