Formal Verification · All levels

Safety vs Liveness, Strong vs Weak: Expanded Case Study

Expanded Case Study for Safety vs Liveness, Strong vs Weak.

Extended case study

A formal regression involving Safety vs Liveness, Strong vs Weak reopens late in the release cycle after RTL and constraint updates.

Background

Earlier runs were stable, but model assumptions drifted and property intent was not re-audited after implementation changes.

Symptoms observed

  • non-vacuous closure rate, counterexample turnaround time, and requirement-level residual risk trend trends worsen while status dashboards look superficially stable.

  • counterexample patterns recur across related properties.

  • reviewers disagree on whether failures are real bugs or modeling artifacts.

Investigation timeline

  1. Hour 0: freeze RTL, assumptions, and tool settings for reproducibility.

  2. Hour 1: classify failures into bug, model mismatch, or weak-property buckets.

  3. Hour 2: isolate first divergence and map to requirement intent.

  4. Hour 3: apply one constrained change and rerun focused property set.

  5. Hour 4: confirm reachability and vacuity quality did not regress.

  6. Hour 5: replay representative traces in simulation or equivalent flow.

  7. Hour 6: publish closure memo with residual risk classification.

Root cause

Root cause traced to Safety vs Liveness, Strong vs Weak: Safety properties state that something bad never happens (for example mutual exclusion: `assert property (@(posedge clk) !(wr && rd));`), while liveness properties state that something good eventually happens (for example progress: `req |-> s_eventually gnt`).

Fix and validation

  • Correct assumption/property scope to preserve legal behavior.

  • Add targeted helper checks that expose key intermediate invariants.

  • Update runbook and requirement traceability for future regression stability.

Lessons learned

  • Status color is not proof quality; audit supporting evidence.

  • First-divergence classification outperforms broad trace inspection.

  • Constraint and abstraction governance must be versioned and reviewed.

diagram
CASE STUDY - Safety vs Liveness, Strong vs Weak
closure slope / vacuity trend / inconclusive aging / replay confidence

Formal deep dive

SVA scales when temporal intent, clock sampling, and reset gating are precise enough to be replayed and reviewed.

Concept diagram

diagram
SVA INTENT CHAIN

timing contract -> sequence composition -> property implication -> sampled failure trace

Metric graph

diagram
ASSERTION QUALITY SIGNALS

non-vacuous hit rate    ████████
clock/reset mismatches  ████
false-positive churn    ███

Metrics and artifacts to collect

  • assertion trigger hit-rate

  • implication timing mismatch bucket

  • reset-window noise ratio

  • assertion decomposition quality score

Mini case study

A protocol failure vanished after correcting `|->` vs `|=>` semantics and reset masking boundaries.

Debug branches

  • Confirm antecedent trigger at sampled clock edges.

  • Verify implication operator matches protocol timing contract.

  • Split monolithic properties into stage-local checks.

Senior review question

Ask: which requirement intent is proven, under which assumptions, and what residual risk remains?

Key takeaways

  • Tie each proof claim to assumption boundaries and reachability evidence.

  • Prefer minimal reversible fixes and preserve legal behavior visibility.

Common pitfalls

  • Treating runtime reduction as proof-quality improvement without audits.

  • Declaring closure while critical covers remain unreachable.

  • Using broad waivers instead of first-divergence root-cause ownership.

Principal formal review addendum

Safety vs Liveness, Strong vs Weak should be reviewed as a requirement-evidence workflow, not a single status report.

Use non-vacuous closure rate, counterexample turnaround time, and requirement-level residual risk trend as the monitoring lens and formal closure packet: assumptions audit, proof status matrix, counterexample classification, and requirement traceability as closure proof.

SVA quality comes from temporal precision, correct sampling semantics, and disciplined vacuity control. Strong teams preserve legal reachability while improving convergence.