Low Power Verification · All levels

Detecting Unintended State Loss Scenarios: Debug Playbook

Debug Playbook for Detecting Unintended State Loss Scenarios.

Debug playbook

Debug Playbook for Detecting Unintended State Loss Scenarios is anchored on Escaped state-loss incident rate per power mode and observability coverage of non-retained critical state.. Convert observations into mechanism-backed and owner-bound actions.

  1. Freeze seed, metadata, and boundary under investigation.

  2. Locate first persistent low-power phase divergence.

  3. Classify mechanism: setup, transition, boundary, retention, or X-prop class.

  4. Apply one focused reproducer and one bounded fix.

  5. Re-run determinism and broader regression matrix.

Review memo template

diagram
LPV REVIEW MEMO - Retention & Restore / Detecting Unintended State Loss Scenarios

1. Symptom
   - Failing metric: Escaped state-loss incident rate per power mode and observability coverage of non-retained critical state.
   - Trigger context: <seed/mode/sequence>
   - First failing phase: <entry/off/exit/boundary>

2. Mechanism hypothesis
   - Candidate mechanism: Unintended state loss often hides in corner transitions where retention assumptions are invalidated by reset, clock, or software sequencing. Verification should classify state into retained, recomputed, checkpointed, and software-reinitialized categories, then prove each category behaves correctly across all supported power modes. Scenario design must include rapid power cycling, nested domain dependencies, retention bypass modes, and warm-reset during restore to expose cases where logic appears alive but architectural state is silently stale or zeroed. Checkers should detect both direct value loss and derived symptoms such as illegal FSM state, inconsistent cache tags, or protocol context mismatches after wake. End-to-end scoreboards that compare pre-sleep intent with post-wake architectural invariants are essential to catch losses that do not manifest as immediate signal mismatches.
   - Competing hypotheses: setup, transition race, boundary bug, retention drift, X-prop noise
   - Missing evidence: <trace/assertion/report>

3. Proposed action
   - Smallest reversible change: <intent/RTL/checker/flow>
   - Expected movement: <failure trend/replay stability>
   - Regression risk: compatibility, coverage, signoff delay

4. Signoff
   - Required artifact: State survivability campaign report covering mode matrix, invariant checks, and residual risk signoff decisions.
   - Required owners: system validation owner, low-power verification owner, software bring-up owner, quality/signoff owner
   - Final decision: ship, bounded rollout, rollback, or escalate

Low-power verification deep dive

Retention closure requires proving end-to-end state lifecycle through save, off, and restore windows.

Concept diagram

diagram
RETENTION LIFECYCLE

save request -> state capture -> power off -> power on -> restore -> traffic resume

Metric graph

diagram
RETENTION STABILITY

restore mismatch       █████
save timing defects    ████
stable wake cycles     ███████

Metrics and artifacts to collect

  • retention save/restore timing report

  • pre/post state diff matrix

  • multi-cycle retention stress summary

  • state-loss bug trend by mode

Mini case study

A corruption issue persisted until retention checks compared multi-cycle state snapshots rather than single wake events.

Debug branches

  • Track save acknowledgement against actual state capture.

  • Validate restore completion before functional traffic resumes.

  • Run repeated sleep/wake cycles to expose drift.

Senior review question

Ask: what exact low-power transition boundary failed first, and which artifact proves the closure claim reproducibly?

Key takeaways

  • Tie each LPV claim to a concrete transition boundary and one proving artifact.

  • Prefer minimal reversible fixes with explicit owner and rollback criteria.

Common pitfalls

  • Treating power-aware failures as random before boundary classification.

  • Waiving X-prop failures before proving impact and root cause.

  • Declaring closure without deterministic replay across key modes.

Debug ladder

Sequence: reproduce -> classify -> isolate boundary -> prove mechanism -> bounded fix.

Avoid mixed fixes before first-principles classification.