Physical Design · All levels
CTS Goals: Skew, Latency, and Transition — Debug Playbook
Debug Playbook for CTS Goals: Skew, Latency, and Transition (Clock Tree Synthesis).
On-call / interview prompt
CTS Goals: Skew, Latency, and Transition looks wrong — walk your first five debug steps.
CLOSURE CHAIN
1. METRIC — name the failing report line (WNS, DRV, DRC count, IR %)
2. HYPOTHESIS — 2–3 likely causes ordered by probability
3. EXPERIMENT — one cheap check (corner, clock, report, map)
4. FIX — minimal physical or constraint change
5. REGRESSION — what you re-run and what must not regressReference workflow
1. Confirm active mode/corner and propagated clock assumptions.
2. Compare skew, latency, and transition against domain-specific targets.
3. Locate violating sink clusters and inspect trunk/leaf depth imbalance.
4. Adjust CTS constraints or buffering strategy with minimal topology churn.
5. Re-run targeted analysis and verify no collateral hold/SI regression.Mechanism to narrate
Separate symptom from root cause
Fix systematic clusters before one-offs
Common pitfalls
Random optimization without metric
Skipping regression after local fix
Staff-level debug discipline
For CTS Goals: Skew, Latency, and Transition, senior debug is branch-and-bound: reduce the search space quickly, keep experiments reversible, and avoid hiding a systematic issue behind one local fix.
Debug decision tree
Reproduce the failure on the tagged database and exact analysis setup.
Classify the failure as data issue, constraint issue, physical implementation issue, tool/methodology issue, or true design limitation.
Run one cheap experiment that can falsify the leading hypothesis.
Prefer a fix that improves a cluster over one that only hides the worst line.
After the fix, re-check Clock Tree Synthesis closure dashboard and the likely regression surface: Routing reserve strategy and hold-fixing burden depend on this objective setup..
Escalation triggers
The failure crosses team ownership boundaries: RTL, synthesis, CAD, IP, package, or foundry.
The local fix consumes margin that another signoff domain needs.
The issue repeats across blocks, suggesting methodology or library root cause.
The remaining risk is silicon-facing: Poor target balancing leads to expensive power growth and unstable post-route timing..
Deep dive: how this shows up in real closure
CTS changes the timing problem by making clocks real.
Reports and artifacts to inspect
clock tree summary: skew, latency, buffer count, sink count
clock transition and capacitance violations
post-CTS setup/hold WNS by clock domain
ICG enable timing and test-mode clock exceptions
Mini case study
Pre-CTS setup is green, but post-CTS hold WNS is -60 ps on short register-to-register paths. This is expected: clock arrival spread is now real. Add data delay or adjust skew carefully, then check setup regression.
Debug branches
If skew is high, inspect clock root, macro stops, NDR, and buffer placement channels.
If hold explodes, classify short paths by domain and add targeted delay, not global padding.
If ICG enable fails, debug enable path timing separately from clock skew.
Senior review question
Ask yourself: what single report line would prove this page's concept is either passing or failing?
What changes at 10+ years
You are expected to predict what your fix can break before running it.
You should recognize when the issue is methodology, not one block's implementation.
You should communicate risk in tapeout language: owner, evidence, impact, mitigation, and decision date.
Principal-level review bar
Deep subpage pages in this course should be read like real closure review material. For a 10+ year PD engineer, the bar is not remembering terminology; it is making a release-quality decision under ambiguity.
What excellent looks like
Names the failing metric, corner/mode, database tag, and analysis switches before proposing a fix.
Separates data, constraint, physical, tool, and methodology root causes instead of treating all failures as optimization problems.
Chooses experiments by information gain and reversibility, not by habit.
States regression blast radius across timing, route, power, PV, DFT, package, and tapeout manifest.
Turns recurring failures into methodology guardrails, dashboards, or checklist items.
Closure note template
STAFF / PRINCIPAL CLOSURE NOTE
Context:
stage: <pre-CTS | post-CTS | post-route | post-fill | signoff>
tag: <database / netlist / SDC / library stack>
failing metric: <exact report line>
affected scope: <block / hierarchy / path group / power domain / region>
Hypotheses:
H1: <most likely physical or constraint mechanism>
H2: <competing explanation>
H3: <methodology or input-data issue>
Decision:
next experiment: <cheap check that can falsify H1>
fix candidate: <minimal reversible change>
rollback trigger: <metric that says the fix is wrong>
regression set: <timing / route / power / PV / DFT / package>
escalation owner: <team or reviewer>Tradeoffs a senior engineer must discuss
Technical tradeoff
CTS changes the timing problem by making clocks real. Explain not only the preferred fix, but what margin or schedule you are spending to get it.
Cross-team tradeoff
What must RTL, synthesis, CAD, STA, DFT, package, IP, or foundry agree to before this decision is final?
Which artifact becomes the source of truth after the decision: report, waiver, manifest, ECO script, or methodology deck?
What is the cost of being wrong: one rerun, ECO churn, mask risk, performance loss, or silicon escape?
Leadership communication
"The current blocker is <metric> in <corner/mode/stage>. The leading cause is <mechanism>. I recommend <fix> because it is bounded and reversible. The regression surface is <domains>. If it fails, we escalate to <owner> with <evidence>."Key takeaways
Always connect the concept back to a measurable signoff artifact.
A fix is not complete until you can name the regression checks.
Common pitfalls
Optimizing by habit instead of reading the current report.
Forgetting that a local fix can regress timing, routing, power, or PV elsewhere.