GPU Design · All levels

L2 & Last-Level Cache: Interview Drills

Interview Drills for L2 & Last-Level Cache.

Interview drills

Interview Drills for L2 & Last-Level Cache centers on L2 hit ratio, eviction pressure, and inter-SM coherence traffic. The objective is to connect profiler evidence to root-cause mechanism and release-safe action.

diagram
PROMPT
You see L2 hit ratio, eviction pressure, and inter-SM coherence traffic on L2 & Last-Level Cache. Walk through root cause and release decision.

STRONG ANSWER
1. Names failing workload/scene and first broken metric.
2. Explains A shared L2/LLC smooths off-chip traffic and inter-client sharing, but contention and policy choices can shift bottlenecks across engines.
3. Requests L2 traffic breakdown, eviction reason chart, and bandwidth pressure summary.
4. Proposes bounded fix + owner + validation matrix.

WEAK ANSWER
Suggests generic tuning without SIMT, warp, cache, or interconnect evidence.

Whiteboard diagram

Shared cache choke points

diagram
GPU MEMORY HIERARCHY — L2 & Last-Level Cache

                [ Registers ]
              latency:   1-2 cycles
                     |
                [ Shared/L1 ]
              latency:  20-40 cycles
                     |
                    [ L2 ]
              latency: 150-250 cycles
                     |
             [ HBM/GDDR VRAM ]
              latency: 300ns+ effective

Optimization lens: explain cross-SM traffic concentration at L2/LLC

Debug tree to narrate

diagram
ROOT-CAUSE TREE — L2 & Last-Level Cache

L2 hit ratio, eviction pressure, and inter-SM coherence traffic regressed
        |
  reproducible on replay?
      /              \
    no                yes
    |                  |
env/test noise    counter triage
                   |
             compute-bound or memory-bound?
                /                  \
             compute            memory/interconnect
             issue stalls       cache/NoC/DRAM stalls

Stop at first failing mechanism, then patch.

GPU deep dive

Bandwidth wins come from coalescing and locality discipline, not peak-memory specs alone.

Concept diagram

diagram
MEMORY HIERARCHY

register -> shared/L1 -> L2/LLC -> HBM/GDDR
access pattern quality decides latency

Metric graph

diagram
BANDWIDTH UTILIZATION

requested BW  ███████████
effective BW  ████████
wasted BW     ███

Reports and artifacts

  • L1/L2 hit-rate report

  • HBM efficiency counters

  • coalescing transaction log

  • shared-memory bank audit

Mini case study

Stencil kernel sat at 43% of peak HBM due to uncoalesced loads; layout rewrite recovered 1.6x effective bandwidth.

Debug branches

  • Check transactions per request at warp granularity

  • Classify cache-thrash versus true DRAM saturation

  • Audit shared-memory bank conflicts before algorithm rewrites

Senior review question

Ask: which metric and benchmark pairing proves this topic is truly closed in production context?

Key takeaways

  • Always pair micro-kernel metrics with end-to-end workload impact.

  • Lock toolchain, driver, and launch metadata before comparing performance results.

Common pitfalls

  • Optimizing occupancy without checking memory-system saturation.

  • Comparing profiler captures from different driver or compiler builds.

  • Declaring wins without reproducible accuracy and performance gates.

Interview answer expansion

A strong interview answer for L2 & Last-Level Cache starts with the workload and metric, then states the mechanism in plain language: A shared L2/LLC smooths off-chip traffic and inter-client sharing, but contention and policy choices can shift bottlenecks across engines.

Then it gives a measurement plan. Good answers name lane masks, issue slots, cache/transaction counters, memory-controller state, NoC congestion, thermal/DVFS telemetry, or stage queues depending on the topic.

Finally, it proposes one bounded fix and explains regression risk. GPU interviews reward tradeoff ownership: what improves, what may regress, and how you would know before tapeout or release.