PCIe/CXL Deep Dive · All levels

Completions, Tags, and Split-Transaction Tracking

PCIe Transactions and DMA: Non-posted requests allocate tags tracked until completions return. Split completions for large reads must assemble in order; tag leaks stall future requests and mimic performance cliffs.

What this topic teaches

Completions, Tags, and Split-Transaction Tracking turns PCIe/CXL theory into production-grade review decisions. Non-posted requests allocate tags tracked until completions return. Split completions for large reads must assemble in order; tag leaks stall future requests and mimic performance cliffs.

The main objective is to identify where the first loss starts in the memory service path, prove it with reproducible traces, and close with the smallest owner-controlled fix.

Senior PCIe/CXL work is less about isolated register tuning and more about cross-layer causality: traffic shape, TLP legality, credit accounting, LTSSM stability, PHY margin, and field reliability must agree before signoff.

Senior-engineer framing question

When Outstanding tag utilization, completion timeout rate, and split completion assembly errors regresses, can you prove whether the first failure is locality collapse, timing-window pressure, scheduler fairness loss, lane-margin drift, or reliability policy overhead?

diagram
PCIe/CXL PROTOCOL STACK - Completions, Tags, and Split-Transaction Tracking

[Application / Driver]
        |
        v
[Transaction Layer]  TLP headers, routing, ordering, completions
        |
        v
[Data Link Layer]    seq/ack, LCRC, replay buffer
        |
        v
[Physical Layer]     encoding, scrambling, LTSSM, lanes
        |
        v
[Link Partner]

Focus: link physical state changes to service-level latency and bandwidth outcomes
Metric tracked: Outstanding tag utilization, completion timeout rate, and split completion assembly errors

Architecture and timing visuals

Draw the mechanism before tuning knobs. These visuals are optimized for design reviews, bring-up triage, and interview whiteboards.

TLP type mix (Completions And Tags)

diagram
TLP MIX (typical DMA workload)

MemRd  ████████
MemWr  ██████
CplD   ████████
CfgRd  █
Atomic ██

Malformed types and alignment errors surface as UR/CA completions.

Tag tracking (Completions And Tags)

diagram
OUTSTANDING TAG POOL

tags free: [3,7,9,...]
in-flight: tag5 MemRd -> waiting CplD
           tag8 MemRd -> split completion 1/2

Tag leaks exhaust pool and stall new non-posted requests.

P2P vs host staging (Completions And Tags)

diagram
DMA PATH OPTIONS

GPU A ---> switch P2P ---> GPU B   (preferred)
GPU A ---> host memory ---> GPU B  (bounce fallback)

ACS + IOMMU policy can force fallback silently.

Array hierarchy context

diagram
PCIe TOPOLOGY MAP - Completions, Tags, and Split-Transaction Tracking

[Root Complex]
    |
    +-- Root Port 0 ---- [Switch] ---- [Endpoint A]
    |                      |
    |                      +---- [Endpoint B]
    +-- Root Port 1 ---- [CXL Type 3 Expander]

BDF routing + bridge windows + HDM decode define reachability.

Command timing context

diagram
LTSSM TIMELINE - Completions, Tags, and Split-Transaction Tracking

time --->  t0      t1       t2        t3       t4
state      Detect  Polling  Config    L0       Recovery
ordered    -       TS1      TS2       TLP/DLLP TS1/TS2
service    down    train    align     active   retrain

Key checks:
- Detect -> Polling timeout
- Config completion before L0
- Recovery trigger correlation with errors

Controller queue context

diagram
CREDIT FLOW VIEW - Completions, Tags, and Split-Transaction Tracking

VC0 posted credits:     [####------] 4/10 available
VC0 non-posted credits: [######----] 6/10 available
VC0 completion credits: [###-------] 3/10 available

Stall signature:
- posted credit exhaustion -> write TLP backpressure
- completion credit exhaustion -> read latency cliff

Ownership layers

diagram
OWNERSHIP LAYERS - Completions, Tags, and Split-Transaction Tracking

layer              owner
-----------------  ----------------
protocol/RTL       PCIe architect
PHY/SI             PHY + SI/PI owner
firmware/OS        FW + driver owner
validation         compliance + post-silicon

Evidence to collect before changing knobs

Fast closure comes from complete evidence packets, not from isolated counter wins. Every recommendation should carry a metric, artifact, owner, and rollback-safe validation plan.

  • Primary metric: Outstanding tag utilization, completion timeout rate, and split completion assembly errors.

  • Primary artifact: Tag pool timeline, completion latency histogram, and timeout register dump.

  • Owners to include: PCIe architect, RTL owner, firmware owner, validation owner.

  • One reproducible failing traffic slice plus one stable comparator capture.

  • One command legality timeline that isolates first failing transition.

  • One margin or reliability packet when PHY or RAS behavior is implicated.

Bandwidth-latency operating lens

diagram
BANDWIDTH/LATENCY CURVE - Completions, Tags, and Split-Transaction Tracking

throughput
    ^
    |     ****  (peak Gen5 x16)
    |   **    **
    |  *        *  <- tail latency inflation
    +----------------> offered load

Metric: Outstanding tag utilization, completion timeout rate, and split completion assembly errors

Root-cause decision tree

diagram
ROOT CAUSE TREE - Completions, Tags, and Split-Transaction Tracking

symptom: Outstanding tag utilization, completion timeout rate, and split completion assembly errors
  |-- LTSSM / PHY margin
  |-- credit / ordering stall
  |-- coherency / HDM config
  |-- RAS / poison handling
  |-- enumeration / resource conflict

Key takeaways

  • Prove first failing transition before touching broad tuning policies.

  • Tie command-level behavior to application-visible QoS outcomes.

  • Close with accountable owner, rollback criteria, and corner validation.

Common pitfalls

  • Optimizing average GB/s while p99 latency and fairness degrade.

  • Comparing traces without fixed firmware, timing profile, and thermal tags.

  • Declaring closure without reliability and retrain robustness checks.

PCIe/CXL deep dive

Transaction patterns (tags, atomics, DMA, P2P) dominate performance and correctness beyond raw link speed.

Concept diagram

diagram
TRANSACTION LIFECYCLE

MemRd -> tag alloc -> completion(s) -> tag free

Metric graph

diagram
TRANSACTION LOSS MIX

tag exhaustion      █████
P2P fallback        ████
atomic retry        ███

Reports and artifacts

  • TLP type histogram

  • tag pool timeline

  • atomic trace

  • P2P path verification matrix

Mini case study

Tag leaks after split-completion stress stalled non-posted traffic while the link remained in L0.

Debug branches

  • Track outstanding tags and completion latency

  • Verify P2P with ACS/IOMMU policy matrix

  • Run coherency litmus for atomics and ordering attrs

Senior review question

Ask: which latency, bandwidth, and reliability evidence proves this PCIe/CXL topic is closed under real traffic?

Key takeaways

  • Always tie controller and PHY counter shifts to application latency and throughput outcomes.

  • Lock firmware timing profile, thermal condition, and DIMM state before comparing PCIe/CXL captures.

Common pitfalls

  • Chasing peak bandwidth while ignoring p99 latency and fairness tails.

  • Changing timing guardbands without separating SI noise from scheduling issues.

  • Declaring closure without reliability gates, fault injection, and regression replay.