GPU Design · All levels

L1 Cache & Texture Path: Inputs & Outputs

Inputs & Outputs for L1 Cache & Texture Path.

Inputs and outputs contract

Inputs & Outputs for L1 Cache & Texture Path centers on L1 hit rate, texture cache efficiency, and cache-thrashing incidents. The objective is to connect profiler evidence to root-cause mechanism and release-safe action.

Use this as the handoff contract across architecture, kernel, compiler, and silicon teams. Ambiguity here creates expensive late-stage rework.

diagram
INPUTS
  - workload definition and expected KPI target
  - kernel launch geometry, compiler flags, and toolchain versions
  - hardware assumptions: SM count, memory type, clocks, thermal envelope
  - acceptance criteria for throughput, latency tail, and stability

OUTPUTS
  - counter and timeline report with reproducible tags
  - bottleneck classification (compute, memory, fabric, thermal)
  - owner-signed mitigation proposal
  - benchmark rerun summary and release recommendation

Hierarchy map

diagram
GPU MEMORY HIERARCHY — L1 Cache & Texture Path

                [ Registers ]
              latency:   1-2 cycles
                     |
                [ Shared/L1 ]
              latency:  20-40 cycles
                     |
                    [ L2 ]
              latency: 150-250 cycles
                     |
             [ HBM/GDDR VRAM ]
              latency: 300ns+ effective

Optimization lens: capacity vs latency

Scheduler map

diagram
WARP SCHEDULER VIEW — L1 Cache & Texture Path

cycle ->      0    1    2    3    4
eligible   [W1,W2,W5] [W2] [W2,W7] [W7] [W3,W7]
issued         W1      W2    W7      W7    W3
stall reason    -    dep wait  -   mem wait  -

Scheduler objective: keep issue slots non-empty.
Focus: eligible warp quality

GPU deep dive

Bandwidth wins come from coalescing and locality discipline, not peak-memory specs alone.

Concept diagram

diagram
MEMORY HIERARCHY

register -> shared/L1 -> L2/LLC -> HBM/GDDR
access pattern quality decides latency

Metric graph

diagram
BANDWIDTH UTILIZATION

requested BW  ███████████
effective BW  ████████
wasted BW     ███

Reports and artifacts

  • L1/L2 hit-rate report

  • HBM efficiency counters

  • coalescing transaction log

  • shared-memory bank audit

Mini case study

Stencil kernel sat at 43% of peak HBM due to uncoalesced loads; layout rewrite recovered 1.6x effective bandwidth.

Debug branches

  • Check transactions per request at warp granularity

  • Classify cache-thrash versus true DRAM saturation

  • Audit shared-memory bank conflicts before algorithm rewrites

Senior review question

Ask: which metric and benchmark pairing proves this topic is truly closed in production context?

Key takeaways

  • Always pair micro-kernel metrics with end-to-end workload impact.

  • Lock toolchain, driver, and launch metadata before comparing performance results.

Common pitfalls

  • Optimizing occupancy without checking memory-system saturation.

  • Comparing profiler captures from different driver or compiler builds.

  • Declaring wins without reproducible accuracy and performance gates.

Handoff explanation

Inputs are not only API parameters or RTL configuration bits. For GPU design, inputs include workload distribution, launch geometry, shader/compiler form, memory layout, clocks, thermal state, SKU target, and runtime policy. Missing any of these makes the same counter mean different things.

Outputs must be decision-ready: L1 hit rate, texture cache efficiency, and cache-thrashing incidents, the artifact set (cache hit/miss profile, access stride study, and texture-path latency report), a bottleneck class, owner, expected effect, and regression scope. A handoff that says only "performance improved" is not enough for architecture or silicon signoff.

The safest handoff format is a before/after packet: workload, revisions, counters, traces, root-cause hypothesis, chosen change, rejected alternatives, and rollback criteria.