Cache Coherency · All levels

Resilience, ECC Poison, and Recovery: Expanded Case Study

Expanded Case Study for Resilience, ECC Poison, and Recovery.

Expanded case study

Expanded Case Study for Resilience, ECC Poison, and Recovery explains how to reason from coherency invariant to measurable engineering decision.

Review a realistic incident end-to-end: symptom capture, mechanism isolation, corrective action, and long-tail prevention.

Evidence pack

diagram
STAFF REVIEW MEMO — Power, Physical, and Reliability / Resilience, ECC Poison, and Recovery

1) Symptom
   - Tracked metric: poison propagation depth and recovery completion time
   - Workload and mode: <explicitly named>
   - First failing evidence: <artifact ID and timestamp>

2) Mechanism hypothesis
   - Candidate mechanism: Error containment must prevent poisoned lines from violating ownership and forward-progress expectations.
   - Alternative explanations: ordering, backpressure, metadata staleness, or software misuse
   - Missing evidence required for decision: <list>

3) Action plan
   - Smallest reversible fix: <RTL, firmware, policy, or tooling>
   - Expected movement: <numeric trend expectation>
   - Risk of regression: performance, power, compatibility, or timing

4) Signoff gates
   - Primary artifact: error-injection campaign report
   - Owners: reliability, verification, firmware
   - Decision: fix now, bounded waiver, or escalate

Cache coherency deep dive

Cache coherence is a correctness contract across caches, interconnect, and software ordering.

Concept diagram

diagram
requester -> coherence fabric -> owner or memory -> state update

Metric graph

diagram
traffic mix across request, snoop, response, data

Metrics and artifacts to collect

  • coherence latency

  • invalidation rate

  • retry rate

  • stale-read incidents

Mini case study

Anchor debug to first stale read and the exact line state transition.

Debug branches

  • Track ownership

  • Track ordering

  • Track evidence

Senior review question

Ask: what is the first line state transition that deviates, and which ordering rule does it break?

Key takeaways

  • Tie every coherency claim to one cache line, one transaction identity, and one measurable counter.

  • Keep proof artifacts from simulation and silicon replay aligned by address, state, and ordering event.

Common pitfalls

  • Chasing bandwidth regressions without checking false sharing and line bouncing first.

  • Assuming coherence correctness implies memory consistency correctness.

  • Declaring closure without litmus, stress, and post-silicon replay evidence.