GPU Design · All levels
L2 & Last-Level Cache: Interview Drills
Interview Drills for L2 & Last-Level Cache.
Interview drills
Interview Drills for L2 & Last-Level Cache centers on L2 hit ratio, eviction pressure, and inter-SM coherence traffic. The objective is to connect profiler evidence to root-cause mechanism and release-safe action.
PROMPT
You see L2 hit ratio, eviction pressure, and inter-SM coherence traffic on L2 & Last-Level Cache. Walk through root cause and release decision.
STRONG ANSWER
1. Names failing workload/scene and first broken metric.
2. Explains A shared L2/LLC smooths off-chip traffic and inter-client sharing, but contention and policy choices can shift bottlenecks across engines.
3. Requests L2 traffic breakdown, eviction reason chart, and bandwidth pressure summary.
4. Proposes bounded fix + owner + validation matrix.
WEAK ANSWER
Suggests generic tuning without SIMT, warp, cache, or interconnect evidence.Whiteboard diagram
Shared cache choke points
GPU MEMORY HIERARCHY — L2 & Last-Level Cache
[ Registers ]
latency: 1-2 cycles
|
[ Shared/L1 ]
latency: 20-40 cycles
|
[ L2 ]
latency: 150-250 cycles
|
[ HBM/GDDR VRAM ]
latency: 300ns+ effective
Optimization lens: explain cross-SM traffic concentration at L2/LLCDebug tree to narrate
ROOT-CAUSE TREE — L2 & Last-Level Cache
L2 hit ratio, eviction pressure, and inter-SM coherence traffic regressed
|
reproducible on replay?
/ \
no yes
| |
env/test noise counter triage
|
compute-bound or memory-bound?
/ \
compute memory/interconnect
issue stalls cache/NoC/DRAM stalls
Stop at first failing mechanism, then patch.GPU deep dive
Bandwidth wins come from coalescing and locality discipline, not peak-memory specs alone.
Concept diagram
MEMORY HIERARCHY
register -> shared/L1 -> L2/LLC -> HBM/GDDR
access pattern quality decides latencyMetric graph
BANDWIDTH UTILIZATION
requested BW ███████████
effective BW ████████
wasted BW ███Reports and artifacts
L1/L2 hit-rate report
HBM efficiency counters
coalescing transaction log
shared-memory bank audit
Mini case study
Stencil kernel sat at 43% of peak HBM due to uncoalesced loads; layout rewrite recovered 1.6x effective bandwidth.
Debug branches
Check transactions per request at warp granularity
Classify cache-thrash versus true DRAM saturation
Audit shared-memory bank conflicts before algorithm rewrites
Senior review question
Ask: which metric and benchmark pairing proves this topic is truly closed in production context?
Key takeaways
Always pair micro-kernel metrics with end-to-end workload impact.
Lock toolchain, driver, and launch metadata before comparing performance results.
Common pitfalls
Optimizing occupancy without checking memory-system saturation.
Comparing profiler captures from different driver or compiler builds.
Declaring wins without reproducible accuracy and performance gates.
Interview answer expansion
A strong interview answer for L2 & Last-Level Cache starts with the workload and metric, then states the mechanism in plain language: A shared L2/LLC smooths off-chip traffic and inter-client sharing, but contention and policy choices can shift bottlenecks across engines.
Then it gives a measurement plan. Good answers name lane masks, issue slots, cache/transaction counters, memory-controller state, NoC congestion, thermal/DVFS telemetry, or stage queues depending on the topic.
Finally, it proposes one bounded fix and explains regression risk. GPU interviews reward tradeoff ownership: what improves, what may regress, and how you would know before tapeout or release.