Silicon Bring-up · All levels

On-Chip Trace and Embedded Logic Analyzer: Debug Playbook

Debug Playbook for On-Chip Trace and Embedded Logic Analyzer.

Debug playbook

Debug Playbook for On-Chip Trace and Embedded Logic Analyzer is anchored on Trigger hit fidelity, useful trace-window depth, and root-cause localization latency for intermittent boot and timing failures.. Convert observed behavior into mechanism-backed and owner-bound actions.

  1. Freeze setup metadata and preserve first-failure state.

  2. Locate first persistent boundary where behavior diverges.

  3. Classify mechanism: dependency, margin, protocol, software, or silicon.

  4. Apply one focused reproducer and one bounded fix.

  5. Re-run replay, corner, and soak confidence matrix.

Review memo template

diagram
BRING-UP REVIEW MEMO - Debug Interfaces & Observability / On-Chip Trace and Embedded Logic Analyzer

1. Symptom
   - Failing metric: Trigger hit fidelity, useful trace-window depth, and root-cause localization latency for intermittent boot and timing failures.
   - Trigger context: <board/firmware/corner/test window>
   - First failing boundary: <power/reset/clock/interface/firmware>

2. Mechanism hypothesis
   - Candidate mechanism: On-chip trace infrastructure and embedded logic analyzers (ELA) provide time-correlated visibility into internal protocol signals, state transitions, and event timelines that cannot be reconstructed from software logs alone. Effective bring-up configures trigger conditions around critical boundaries such as reset deassertion, clock-domain handshakes, boot-ROM branching, and fabric timeout events, then captures pre-trigger and post-trigger context to expose the first divergence point. Because trace bandwidth and SRAM depth are constrained, teams must prioritize semantic signals, use compression/selective funneling, and align trace clocks/timestamps across blocks to avoid false causality. The strongest debug flows tie ELA captures to known boot phases and expected invariants, enabling fast distinction between control-flow bugs, CDC effects, and analog-timing sensitivity.
   - Competing hypotheses: setup, dependency, margin, software path, silicon defect
   - Missing evidence: <trace/scope/register/report>

3. Proposed action
   - Smallest reversible change: <setup/script/config/firmware>
   - Expected movement: <repro rate/latency/pass trend>
   - Regression risk: stability, safety, release timeline, ownership handoff

4. Signoff
   - Required artifact: Trace observability plan with trigger catalog, signal-priority list, timestamp alignment rules, and standard decode templates for bring-up incidents.
   - Required owners: silicon validation owner, SoC integration owner, clock/reset owner, debug instrumentation owner
   - Final decision: ship, bounded rollout, rollback, respin escalation

Silicon bring-up deep dive

Debug interfaces are useful only when access paths are trusted, minimally intrusive, and synchronized to failure context.

Concept diagram

diagram
DEBUG ACCESS STACK

physical probes -> debug transport -> trace/scan capture -> correlated analysis

Metric graph

diagram
OBSERVABILITY MATURITY

access failures          ████
partial captures         █████
actionable captures      ███████

Metrics and artifacts to collect

  • JTAG/SWD access success rate

  • trace trigger hit coverage

  • scan dump decode turnaround time

  • observability gap backlog

Mini case study

A misdiagnosed silicon issue was cleared after TAP chain validation revealed a board-level debug domain assumption error.

Debug branches

  • Validate access-layer prerequisites before deep protocol decode.

  • Correlate trace timestamps with software checkpoints.

  • Treat missing evidence as an observability gap, not closure.

Senior review question

Ask: what is the first failing boundary, which artifact proves it, and who owns bounded closure?

Key takeaways

  • Tie every bring-up claim to one reproducible setup state and one proving artifact.

  • Prefer bounded fixes with clear owner and rollback trigger over broad multi-variable edits.

Common pitfalls

  • Running parallel uncontrolled experiments and losing causality.

  • Declaring closure without replaying across representative corners.

  • Escalating severity before bench/setup hypotheses are disproven.

Debug ladder

Sequence: reproduce -> classify -> isolate -> instrument -> bounded fix -> replay.

Avoid parallel broad edits before first root-cause class is proven.