Formal Verification · All levels

Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search: Expanded Case Study

Expanded Case Study for Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search.

Extended case study

A formal regression involving Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search reopens late in the release cycle after RTL and constraint updates.

Background

Earlier runs were stable, but model assumptions drifted and property intent was not re-audited after implementation changes.

Symptoms observed

  • non-vacuous closure rate, counterexample turnaround time, and requirement-level residual risk trend trends worsen while status dashboards look superficially stable.

  • counterexample patterns recur across related properties.

  • reviewers disagree on whether failures are real bugs or modeling artifacts.

Investigation timeline

  1. Hour 0: freeze RTL, assumptions, and tool settings for reproducibility.

  2. Hour 1: classify failures into bug, model mismatch, or weak-property buckets.

  3. Hour 2: isolate first divergence and map to requirement intent.

  4. Hour 3: apply one constrained change and rerun focused property set.

  5. Hour 4: confirm reachability and vacuity quality did not regress.

  6. Hour 5: replay representative traces in simulation or equivalent flow.

  7. Hour 6: publish closure memo with residual risk classification.

Root cause

Root cause traced to Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search: Simulation validates behavior for sampled traces created by directed and constrained-random stimulus, so its confidence depends on test quality, scenario coverage, and seed diversity.

Fix and validation

  • Correct assumption/property scope to preserve legal behavior.

  • Add targeted helper checks that expose key intermediate invariants.

  • Update runbook and requirement traceability for future regression stability.

Lessons learned

  • Status color is not proof quality; audit supporting evidence.

  • First-divergence classification outperforms broad trace inspection.

  • Constraint and abstraction governance must be versioned and reviewed.

diagram
CASE STUDY - Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search
closure slope / vacuity trend / inconclusive aging / replay confidence

Formal deep dive

FPV foundations are reliable only when assumptions, reset semantics, and requirement intent are explicitly modeled and audited.

Concept diagram

diagram
FPV FOUNDATION LOOP

requirements -> property set -> assumptions and reset model -> prove/fail traces -> closure audit

Metric graph

diagram
FOUNDATION HEALTH

vacuous passes         ████
reachable proofs       ███████
inconclusive backlog   █████
reopened properties    ███

Metrics and artifacts to collect

  • assumption traceability matrix

  • vacuity and reachability status

  • proof core relevance summary

  • counterexample classification trend

Mini case study

A green-looking run was invalidated after legal-mode covers failed, exposing assumptions that removed realistic traffic.

Debug branches

  • Validate requirement-to-property mapping before tuning runtime.

  • Check legal scenario reachability after every assumption change.

  • Classify first divergence as model issue or RTL bug.

Senior review question

Ask: which requirement intent is proven, under which assumptions, and what residual risk remains?

Key takeaways

  • Tie each proof claim to assumption boundaries and reachability evidence.

  • Prefer minimal reversible fixes and preserve legal behavior visibility.

Common pitfalls

  • Treating runtime reduction as proof-quality improvement without audits.

  • Declaring closure while critical covers remain unreachable.

  • Using broad waivers instead of first-divergence root-cause ownership.

Principal formal review addendum

Formal vs Simulation: Exhaustive Proof and Stimulus-Based Search should be reviewed as a requirement-evidence workflow, not a single status report.

Use non-vacuous closure rate, counterexample turnaround time, and requirement-level residual risk trend as the monitoring lens and formal closure packet: assumptions audit, proof status matrix, counterexample classification, and requirement traceability as closure proof.

Formal foundations are strongest when assumptions, reset semantics, and requirement intent are all explicit and reviewable. Strong teams preserve legal reachability while improving convergence.