Silicon Bring-up · All levels
Test Program Bring-up: From Characterization Script to Screening Flow: Debug Playbook
Debug Playbook for Test Program Bring-up: From Characterization Script to Screening Flow.
Debug playbook
Debug Playbook for Test Program Bring-up: From Characterization Script to Screening Flow is anchored on First-pass test-program pass rate on known-good silicon, escaped-defect proxy rate, and debug turnaround time per failing test block.. Convert observed behavior into mechanism-backed and owner-bound actions.
Freeze setup metadata and preserve first-failure state.
Locate first persistent boundary where behavior diverges.
Classify mechanism: dependency, margin, protocol, software, or silicon.
Apply one focused reproducer and one bounded fix.
Re-run replay, corner, and soak confidence matrix.
Review memo template
BRING-UP REVIEW MEMO - ATE Correlation & Test / Test Program Bring-up: From Characterization Script to Screening Flow
1. Symptom
- Failing metric: First-pass test-program pass rate on known-good silicon, escaped-defect proxy rate, and debug turnaround time per failing test block.
- Trigger context: <board/firmware/corner/test window>
- First failing boundary: <power/reset/clock/interface/firmware>
2. Mechanism hypothesis
- Candidate mechanism: Early test programs are usually stitched from characterization snippets, but production-worthy bring-up requires conversion into deterministic, restart-safe, and diagnosable test methods. Engineers sequence tests to control thermal history and avoid pattern interactions, define guardbands from measured process spread rather than single-die behavior, and instrument datalogs so each fail can be traced to setup, pattern, timing edge, or limit decision. Known-good and known-bad vehicles are both required: known-good validates overkill risk, while seeded-failure or marginal parts validate detection sensitivity and diagnostic specificity. Program maturity also depends on robust site-to-site behavior in multisite execution, where shared resources, tester timing skew, and handler effects can create false yield loss. A disciplined bring-up phase therefore treats reproducibility and diagnosability as equal to pass/fail correctness.
- Competing hypotheses: setup, dependency, margin, software path, silicon defect
- Missing evidence: <trace/scope/register/report>
3. Proposed action
- Smallest reversible change: <setup/script/config/firmware>
- Expected movement: <repro rate/latency/pass trend>
- Regression risk: stability, safety, release timeline, ownership handoff
4. Signoff
- Required artifact: Bring-up checklist with test-order rationale, guardband derivation notes, reproducibility report, and fail-log decode map.
- Required owners: product test engineer, test program developer, yield engineering owner, DFT representative, manufacturing test operations owner
- Final decision: ship, bounded rollout, rollback, respin escalationSilicon bring-up deep dive
Correlation succeeds when tester and bench experiments share identical conditions and evidence expectations.
Concept diagram
CORRELATION LADDER
ATE fail bin -> extract pattern -> reproduce on bench -> reconcile deltasMetric graph
CORRELATION CONFIDENCE
unmatched signatures █████
partial matches ████
full context matches ███████Metrics and artifacts to collect
ATE-to-bench signature match ratio
pattern replay fidelity score
environment mismatch incident rate
yield-impact closure tracker
Mini case study
Correlation speed improved dramatically after enforcing shared metadata headers and one replay protocol across tester and lab.
Debug branches
Normalize V/F/T and pattern-window metadata first.
Audit fixture and probing assumptions before silicon blame.
Require repeatable signature in both environments before closure.
Senior review question
Ask: what is the first failing boundary, which artifact proves it, and who owns bounded closure?
Key takeaways
Tie every bring-up claim to one reproducible setup state and one proving artifact.
Prefer bounded fixes with clear owner and rollback trigger over broad multi-variable edits.
Common pitfalls
Running parallel uncontrolled experiments and losing causality.
Declaring closure without replaying across representative corners.
Escalating severity before bench/setup hypotheses are disproven.
Debug ladder
Sequence: reproduce -> classify -> isolate -> instrument -> bounded fix -> replay.
Avoid parallel broad edits before first root-cause class is proven.