PCIe/CXL Deep Dive · All levels

Poisoned TLPs and ECRC Protection: Inputs and Outputs

Inputs and Outputs for Poisoned TLPs and ECRC Protection.

Inputs and outputs contract

Inputs and Outputs for Poisoned TLPs and ECRC Protection focuses on Poisoned TLP count, ECRC mismatch rate, and containment success rate. The purpose is to turn memory observations into mechanism-backed actions with explicit owners and release-safe validation.

Use this contract for architecture, controller firmware, PHY, and validation handoffs. Missing inputs create expensive late-stage rework and inconclusive debug loops.

diagram
INPUTS
  - workload distribution and QoS target
  - firmware revision, controller policy profile, timing registers
  - data-rate / voltage / temperature operating state
  - training snapshot and reliability policy status

OUTPUTS
  - bottleneck classification with command-level evidence
  - owner-signed mitigation proposal
  - before/after trend for latency, bandwidth, and reliability
  - regression matrix with rollback triggers

Ownership split

diagram
OWNERSHIP LAYERS - Poisoned TLPs and ECRC Protection

layer              owner
-----------------  ----------------
protocol/RTL       reliability owner
PHY/SI             PHY + SI/PI owner
firmware/OS        FW + driver owner
validation         compliance + post-silicon

PCIe/CXL deep dive

RAS closure maps AER, poison, and surprise-down events to bounded containment and recovery actions.

Concept diagram

diagram
RAS ESCALATION

detect -> classify -> contain -> recover -> validate

Metric graph

diagram
RAS EVENT MIX

correctable trend   ███████
uncorrectable       ██
surprise-down       █

Reports and artifacts

  • AER register dump

  • poison injection log

  • surprise-down timeline

  • containment action record

Mini case study

Masked correctable errors accumulated until a surprise-down during peak traffic forced unplanned failover.

Debug branches

  • Separate CE trend from UE containment paths

  • Validate poison handling end-to-end

  • Test surprise-down drain and driver recovery

Senior review question

Ask: which latency, bandwidth, and reliability evidence proves this PCIe/CXL topic is closed under real traffic?

Key takeaways

  • Always tie controller and PHY counter shifts to application latency and throughput outcomes.

  • Lock firmware timing profile, thermal condition, and DIMM state before comparing PCIe/CXL captures.

Common pitfalls

  • Chasing peak bandwidth while ignoring p99 latency and fairness tails.

  • Changing timing guardbands without separating SI noise from scheduling issues.

  • Declaring closure without reliability gates, fault injection, and regression replay.

Handoff explanation

Inputs extend beyond timing registers. PCIe/CXL analysis inputs include traffic distribution, address map, queue policy, training state, SI/PI condition, thermal state, and firmware version.

Outputs must be action-ready: Poisoned TLP count, ECRC mismatch rate, and containment success rate, artifact packet (Poison injection log, ECRC error trace, and containment action record), bottleneck class, owner, expected gain, and rollback scope. "Bandwidth improved" without this packet is not signoff-ready.

The safest handoff is a before/after evidence set: environment tags, traces, hypothesis, chosen fix, rejected alternatives, and regression criteria.