Formal Verification · All levels
Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search: Expanded Case Study
Expanded Case Study for Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search.
Extended case study
A formal regression involving Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search reopens late in the release cycle after RTL and constraint updates.
Background
Earlier runs were stable, but model assumptions drifted and property intent was not re-audited after implementation changes.
Symptoms observed
non-vacuous closure rate, counterexample turnaround time, and requirement-level residual risk trend trends worsen while status dashboards look superficially stable.
counterexample patterns recur across related properties.
reviewers disagree on whether failures are real bugs or modeling artifacts.
Investigation timeline
Hour 0: freeze RTL, assumptions, and tool settings for reproducibility.
Hour 1: classify failures into bug, model mismatch, or weak-property buckets.
Hour 2: isolate first divergence and map to requirement intent.
Hour 3: apply one constrained change and rerun focused property set.
Hour 4: confirm reachability and vacuity quality did not regress.
Hour 5: replay representative traces in simulation or equivalent flow.
Hour 6: publish closure memo with residual risk classification.
Root cause
Root cause traced to Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search: Simulation validates behavior for sampled traces created by directed and constrained-random stimulus, so its confidence depends on test quality, scenario coverage, and seed diversity.
Fix and validation
Correct assumption/property scope to preserve legal behavior.
Add targeted helper checks that expose key intermediate invariants.
Update runbook and requirement traceability for future regression stability.
Lessons learned
Status color is not proof quality; audit supporting evidence.
First-divergence classification outperforms broad trace inspection.
Constraint and abstraction governance must be versioned and reviewed.
CASE STUDY - Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search
closure slope / vacuity trend / inconclusive aging / replay confidenceFormal deep dive
FPV foundations are reliable only when assumptions, reset semantics, and requirement intent are explicitly modeled and audited.
Concept diagram
FPV FOUNDATION LOOP
requirements -> property set -> assumptions and reset model -> prove/fail traces -> closure auditMetric graph
FOUNDATION HEALTH
vacuous passes ████
reachable proofs ███████
inconclusive backlog █████
reopened properties ███Metrics and artifacts to collect
assumption traceability matrix
vacuity and reachability status
proof core relevance summary
counterexample classification trend
Mini case study
A green-looking run was invalidated after legal-mode covers failed, exposing assumptions that removed realistic traffic.
Debug branches
Validate requirement-to-property mapping before tuning runtime.
Check legal scenario reachability after every assumption change.
Classify first divergence as model issue or RTL bug.
Senior review question
Ask: which requirement intent is proven, under which assumptions, and what residual risk remains?
Key takeaways
Tie each proof claim to assumption boundaries and reachability evidence.
Prefer minimal reversible fixes and preserve legal behavior visibility.
Common pitfalls
Treating runtime reduction as proof-quality improvement without audits.
Declaring closure while critical covers remain unreachable.
Using broad waivers instead of first-divergence root-cause ownership.
Principal formal review addendum
Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search should be reviewed as a requirement-evidence workflow, not a single status report.
Use non-vacuous closure rate, counterexample turnaround time, and requirement-level residual risk trend as the monitoring lens and formal closure packet: assumptions audit, proof status matrix, counterexample classification, and requirement traceability as closure proof.
Formal foundations are strongest when assumptions, reset semantics, and requirement intent are all explicit and reviewable. Strong teams preserve legal reachability while improving convergence.