Low Power Verification · All levels

Detecting Unintended State Loss Scenarios: Theory Deep Dive

Theory Deep Dive for Detecting Unintended State Loss Scenarios.

Foundational theory

Detecting Unintended State Loss Scenarios is core to Retention & Restore. Treat each power behavior change as a correctness and signoff risk decision.

Core concepts explained

  • Unintended state loss often hides in corner transitions where retention assumptions are invalidated by reset, clock, or software sequencing. Verification should classify state into retained, recomputed, checkpointed, and software-reinitialized categories, then prove each category behaves correctly across all supported power modes. Scenario design must include rapid power cycling, nested domain dependencies, retention bypass modes, and warm-reset during restore to expose cases where logic appears alive but architectural state is silently stale or zeroed. Checkers should detect both direct value loss and derived symptoms such as illegal FSM state, inconsistent cache tags, or protocol context mismatches after wake. End-to-end scoreboards that compare pre-sleep intent with post-wake architectural invariants are essential to catch losses that do not manifest as immediate signal mismatches.

  • Primary metric: Escaped state-loss incident rate per power mode and observability coverage of non-retained critical state.

  • Primary artifact: State survivability campaign report covering mode matrix, invariant checks, and residual risk signoff decisions.

  • Owners: system validation owner, low-power verification owner, software bring-up owner, quality/signoff owner

  • Power intent and RTL behavior must stay aligned through transitions

  • Proof quality beats broad waive strategies in low-power closure

Why this matters in low-power signoff

Retention closure requires proving save, off, and restore phases as one lifecycle with explicit handshake timing. Teams that enforce this reduce false alarms and real escapes.

Mental model

diagram
RETENTION SAVE / RESTORE FLOW

PMU            RET CTRL             RET FLOPS              DOMAIN
 | save_req ---> |                     |                     |
 |               |--- capture ----->   | latch state         |
 | <--- save_ack |                     |                     |
 | power_off --->|---------------------X--------------------> OFF
 | power_on  --->|------------------------------------------> RAMP
 |               |--- restore ----->   | load state          |
 | <--- rst_done |                     |                     |

Verification focus:
- no data loss across save/restore window
- restore completes before functional traffic resumes

Worked intuition

  1. Classify symptom first: illegal transition, corruption, isolation break, retention drift, or X-prop ambiguity.

  2. Pinpoint first phase boundary where expected low-power behavior diverges.

  3. Quantify movement in Escaped state-loss incident rate per power mode and observability coverage of non-retained critical state. before broad refactors.

  4. Collect State survivability campaign report covering mode matrix, invariant checks, and residual risk signoff decisions. with fixed run metadata and mode sequencing.

  5. Apply one bounded fix and replay both targeted and broader scenarios.

  6. Publish owner-signed closure note with rollback trigger.

Common misconceptions

  • Passing nominal ON/OFF smoke proves transition correctness.

  • UPF compile clean means all intent semantics are correct.

  • All X-prop failures indicate real product escapes.

  • Retention behavior can be trusted without multi-cycle restore stress.

Low-power verification deep dive

Retention closure requires proving end-to-end state lifecycle through save, off, and restore windows.

Concept diagram

diagram
RETENTION LIFECYCLE

save request -> state capture -> power off -> power on -> restore -> traffic resume

Metric graph

diagram
RETENTION STABILITY

restore mismatch       █████
save timing defects    ████
stable wake cycles     ███████

Metrics and artifacts to collect

  • retention save/restore timing report

  • pre/post state diff matrix

  • multi-cycle retention stress summary

  • state-loss bug trend by mode

Mini case study

A corruption issue persisted until retention checks compared multi-cycle state snapshots rather than single wake events.

Debug branches

  • Track save acknowledgement against actual state capture.

  • Validate restore completion before functional traffic resumes.

  • Run repeated sleep/wake cycles to expose drift.

Senior review question

Ask: what exact low-power transition boundary failed first, and which artifact proves the closure claim reproducibly?

Key takeaways

  • Tie each LPV claim to a concrete transition boundary and one proving artifact.

  • Prefer minimal reversible fixes with explicit owner and rollback criteria.

Common pitfalls

  • Treating power-aware failures as random before boundary classification.

  • Waiving X-prop failures before proving impact and root cause.

  • Declaring closure without deterministic replay across key modes.

Theory reinforcement

Theory matters when it predicts concrete failure signatures and closure boundaries.

Translate LPV semantics into reproducible verification outcomes.