Physical Design · All levels
Electromigration Analysis — Review Checklist
Review Checklist for Electromigration Analysis (Power Signoff).
Review gate
Worst EM paths have root-cause category
Wire/via ECOs validated by re-analysis
Clock and power trunks checked separately
Temperature assumptions match reliability policy
Residual EM waivers are bounded and approved
Smoke check (5 minutes)
Every checklist item has an owner
Failed items have mitigation or waiver ID
Definition of done for a senior owner
The exact database tag and analysis setup are recorded.
The primary metric is clean or accepted by waiver: EM ratio histogram with worst instance list.
The fix is explained by mechanism, not by tool folklore.
Regression coverage includes the obvious downstream domains: timing, routing, power, PV, DFT, package, or tapeout signoff.
Residual risk has an owner, approval path, and expiration date.
The lesson is captured as a methodology guardrail if it can recur.
Smoke check (5 minutes)
Could another engineer reproduce the conclusion from the notes alone?
Would you sign this off if the design came from another team?
Deep dive: how this shows up in real closure
Power signoff validates whether the physical grid supports real activity.
Reports and artifacts to inspect
static IR and dynamic IR maps with max drop percentage
EM ratio report by wire/via segment
activity source: vectorless, SAIF/VCD, or scenario-specific waveform
low-power cell placement: isolation, level shifter, retention, power switch
Mini case study
Dynamic IR fails near a clock-gated compute island during wake-up. Adding a far-away ring is not enough. Strengthen local straps/vias, add decap, and check wake sequencing.
Debug branches
If dynamic IR fails but static passes, correlate with switching vectors and clock domains.
If EM fails on vias, add via ladders or parallel straps rather than only widening wire.
If leakage fails, use HVT swaps on non-critical paths before reducing performance mode.
Senior review question
Ask yourself: what single report line would prove this page's concept is either passing or failing?
What changes at 10+ years
You are expected to predict what your fix can break before running it.
You should recognize when the issue is methodology, not one block's implementation.
You should communicate risk in tapeout language: owner, evidence, impact, mitigation, and decision date.
Principal-level review bar
Deep subpage pages in this course should be read like real closure review material. For a 10+ year PD engineer, the bar is not remembering terminology; it is making a release-quality decision under ambiguity.
What excellent looks like
Names the failing metric, corner/mode, database tag, and analysis switches before proposing a fix.
Separates data, constraint, physical, tool, and methodology root causes instead of treating all failures as optimization problems.
Chooses experiments by information gain and reversibility, not by habit.
States regression blast radius across timing, route, power, PV, DFT, package, and tapeout manifest.
Turns recurring failures into methodology guardrails, dashboards, or checklist items.
Closure note template
STAFF / PRINCIPAL CLOSURE NOTE
Context:
stage: <pre-CTS | post-CTS | post-route | post-fill | signoff>
tag: <database / netlist / SDC / library stack>
failing metric: <exact report line>
affected scope: <block / hierarchy / path group / power domain / region>
Hypotheses:
H1: <most likely physical or constraint mechanism>
H2: <competing explanation>
H3: <methodology or input-data issue>
Decision:
next experiment: <cheap check that can falsify H1>
fix candidate: <minimal reversible change>
rollback trigger: <metric that says the fix is wrong>
regression set: <timing / route / power / PV / DFT / package>
escalation owner: <team or reviewer>Tradeoffs a senior engineer must discuss
Technical tradeoff
Power signoff validates whether the physical grid supports real activity. Explain not only the preferred fix, but what margin or schedule you are spending to get it.
Cross-team tradeoff
What must RTL, synthesis, CAD, STA, DFT, package, IP, or foundry agree to before this decision is final?
Which artifact becomes the source of truth after the decision: report, waiver, manifest, ECO script, or methodology deck?
What is the cost of being wrong: one rerun, ECO churn, mask risk, performance loss, or silicon escape?
Leadership communication
"The current blocker is <metric> in <corner/mode/stage>. The leading cause is <mechanism>. I recommend <fix> because it is bounded and reversible. The regression surface is <domains>. If it fails, we escalate to <owner> with <evidence>."Key takeaways
Always connect the concept back to a measurable signoff artifact.
A fix is not complete until you can name the regression checks.
Common pitfalls
Optimizing by habit instead of reading the current report.
Forgetting that a local fix can regress timing, routing, power, or PV elsewhere.