Physical Design · All levels

Die Size and Utilization Targets — Interview Drills

Interview Drills for Die Size and Utilization Targets (Floorplanning Expanded).

Interview drills

Practice aloud for Floorplanning Expanded → Die Size and Utilization Targets. Use METRIC → HYPOTHESIS → FIX → REGRESSION.

Why can 68% utilization still fail routing?

diagram
[INT][PD][TOPIC]

Q: Why can 68% utilization still fail routing?

A:
Because utilization is spatially uneven. Macro channels and pin hotspots can exceed capacity while average utilization looks healthy.

FOLLOW-UP TRAP: Treating global utilization as a sufficient routing predictor.

How do you justify die size in a review meeting?

diagram
[INT][PD][TOPIC]

Q: How do you justify die size in a review meeting?

A:
Show area composition, guard-band assumptions, hotspot evidence, and the impact on timing/IR if reduced.

FOLLOW-UP TRAP: Saying 'industry norm is 70%' without project evidence.

When is a die shrink acceptable late in schedule?

diagram
[INT][PD][TOPIC]

Q: When is a die shrink acceptable late in schedule?

A:
Only if overflow, IR headroom, and timing sensitivity remain within signed thresholds after comparative trial floorplans.

FOLLOW-UP TRAP: Approving shrink based on area pressure only.

10+ year interview answer bar

At senior/principal level, the interviewer is testing ownership judgment more than vocabulary. Answer Die Size and Utilization Targets through failure mode, evidence, tradeoff, and release decision.

You inherit a late-stage Die Size and Utilization Targets failure one week before release. What do you do in the first hour?

diagram
[INT][PD][STAFF]

Q: You inherit a late-stage Die Size and Utilization Targets failure one week before release. What do you do in the first hour?

A:
Freeze the database tag, name the failing metric (Utilization by region and congestion predictor report), confirm corner/mode/stage, cluster the issue, assign the first experiment, and publish a regression/owner plan before making broad tool changes.

FOLLOW-UP TRAP: Jumping directly to optimization knobs without preserving evidence.

When would you stop trying to improve Die Size and Utilization Targets and escalate?

diagram
[INT][PD][STAFF]

Q: When would you stop trying to improve Die Size and Utilization Targets and escalate?

A:
Escalate when the remaining risk crosses ownership boundaries, consumes shared margin, changes signed-off assumptions, or threatens Placement quality, buffer insertion freedom, and route convergence directly depend on this decision.. Bring exact report lines and options, not vague concern.

FOLLOW-UP TRAP: Escalating without data or continuing alone after a cross-team decision is needed.

Deep dive: how this shows up in real closure

Floorplan quality is the earliest predictor of place-and-route pain.

Reports and artifacts to inspect

  • floorplan summary: core area, macro area, std-cell utilization

  • macro/channel review: pin-facing sides, halos, routing channels

  • power plan preview: ring width, strap pitch, follow-pin connectivity

  • trial route congestion: overflow around macro corners and pin fields

Mini case study

A 2 MB SRAM cluster is placed with pins facing the die edge. Trial route shows red overflow along the north edge. A senior PD answer is to rotate or mirror the SRAM, open the channel, and re-run trial route before attempting timing optimization.

Debug branches

  • If congestion is local to macro corners, inspect pin sides and halo width before reducing global utilization.

  • If IR is weak at the core edge, widen the ring or add edge straps before adding random decaps.

  • If timing paths cross the whole block, review pin assignment and macro orientation before post-route ECO.

Senior review question

Ask yourself: what single report line would prove this page's concept is either passing or failing?

What changes at 10+ years

  • You are expected to predict what your fix can break before running it.

  • You should recognize when the issue is methodology, not one block's implementation.

  • You should communicate risk in tapeout language: owner, evidence, impact, mitigation, and decision date.

Principal-level review bar

Deep subpage pages in this course should be read like real closure review material. For a 10+ year PD engineer, the bar is not remembering terminology; it is making a release-quality decision under ambiguity.

What excellent looks like

  • Names the failing metric, corner/mode, database tag, and analysis switches before proposing a fix.

  • Separates data, constraint, physical, tool, and methodology root causes instead of treating all failures as optimization problems.

  • Chooses experiments by information gain and reversibility, not by habit.

  • States regression blast radius across timing, route, power, PV, DFT, package, and tapeout manifest.

  • Turns recurring failures into methodology guardrails, dashboards, or checklist items.

Closure note template

diagram
STAFF / PRINCIPAL CLOSURE NOTE

Context:
  stage: <pre-CTS | post-CTS | post-route | post-fill | signoff>
  tag: <database / netlist / SDC / library stack>
  failing metric: <exact report line>
  affected scope: <block / hierarchy / path group / power domain / region>

Hypotheses:
  H1: <most likely physical or constraint mechanism>
  H2: <competing explanation>
  H3: <methodology or input-data issue>

Decision:
  next experiment: <cheap check that can falsify H1>
  fix candidate: <minimal reversible change>
  rollback trigger: <metric that says the fix is wrong>
  regression set: <timing / route / power / PV / DFT / package>
  escalation owner: <team or reviewer>

Tradeoffs a senior engineer must discuss

Technical tradeoff

Floorplan quality is the earliest predictor of place-and-route pain. Explain not only the preferred fix, but what margin or schedule you are spending to get it.

Cross-team tradeoff

  • What must RTL, synthesis, CAD, STA, DFT, package, IP, or foundry agree to before this decision is final?

  • Which artifact becomes the source of truth after the decision: report, waiver, manifest, ECO script, or methodology deck?

  • What is the cost of being wrong: one rerun, ECO churn, mask risk, performance loss, or silicon escape?

Leadership communication

diagram
"The current blocker is <metric> in <corner/mode/stage>. The leading cause is <mechanism>. I recommend <fix> because it is bounded and reversible. The regression surface is <domains>. If it fails, we escalate to <owner> with <evidence>."

Key takeaways

  • Always connect the concept back to a measurable signoff artifact.

  • A fix is not complete until you can name the regression checks.

Common pitfalls

  • Optimizing by habit instead of reading the current report.

  • Forgetting that a local fix can regress timing, routing, power, or PV elsewhere.