GPU Design · All levels
Instruction Issue & Scoreboard: Interview Drills
Interview Drills for Instruction Issue & Scoreboard.
Interview drills
Interview Drills for Instruction Issue & Scoreboard centers on issue slot utilization, dependency stall ratio, and scoreboard wait depth. The objective is to connect profiler evidence to root-cause mechanism and release-safe action.
PROMPT
You see issue slot utilization, dependency stall ratio, and scoreboard wait depth on Instruction Issue & Scoreboard. Walk through root cause and release decision.
STRONG ANSWER
1. Names failing workload/scene and first broken metric.
2. Explains Scoreboards track data hazards and memory readiness so schedulers issue only safe instructions; scoreboard pressure directly limits ILP extraction.
3. Requests scoreboard state timeline, issue reason breakdown, and stall attribution report.
4. Proposes bounded fix + owner + validation matrix.
WEAK ANSWER
Suggests generic tuning without SIMT, warp, cache, or interconnect evidence.Whiteboard diagram
Issue slots vs dependency stalls
WARP SCHEDULER VIEW — Instruction Issue & Scoreboard
cycle -> 0 1 2 3 4
eligible [W1,W2,W5] [W2] [W2,W7] [W7] [W3,W7]
issued W1 W2 W7 W7 W3
stall reason - dep wait - mem wait -
Scheduler objective: keep issue slots non-empty.
Focus: map scoreboard wait reasons to lost issue opportunitiesDebug tree to narrate
ROOT-CAUSE TREE — Instruction Issue & Scoreboard
issue slot utilization, dependency stall ratio, and scoreboard wait depth regressed
|
reproducible on replay?
/ \
no yes
| |
env/test noise counter triage
|
compute-bound or memory-bound?
/ \
compute memory/interconnect
issue stalls cache/NoC/DRAM stalls
Stop at first failing mechanism, then patch.GPU deep dive
Shader-core throughput is gated by issue policy, register-bank access, and pipeline hazard behavior.
Concept diagram
SM CORE LOOP
warp schedulers -> issue ports -> ALU/FPU/Tensor pipelines
scoreboard + register file gate progressMetric graph
SM BOTTLENECK MIX
dependency stalls ███████
bank conflicts ████
pipeline bubbles ███Reports and artifacts
SM IPC dashboard
issue stall taxonomy
register-bank conflict log
shader unit utilization
Mini case study
Compiler register allocation shifted operand banking, doubling RF conflicts and causing a 14% shader regression.
Debug branches
Inspect scoreboard wait-depth trends
Track RF conflicts by instruction class
Separate front-end issue loss from backend saturation
Senior review question
Ask: which metric and benchmark pairing proves this topic is truly closed in production context?
Key takeaways
Always pair micro-kernel metrics with end-to-end workload impact.
Lock toolchain, driver, and launch metadata before comparing performance results.
Common pitfalls
Optimizing occupancy without checking memory-system saturation.
Comparing profiler captures from different driver or compiler builds.
Declaring wins without reproducible accuracy and performance gates.
Interview answer expansion
A strong interview answer for Instruction Issue & Scoreboard starts with the workload and metric, then states the mechanism in plain language: Scoreboards track data hazards and memory readiness so schedulers issue only safe instructions; scoreboard pressure directly limits ILP extraction.
Then it gives a measurement plan. Good answers name lane masks, issue slots, cache/transaction counters, memory-controller state, NoC congestion, thermal/DVFS telemetry, or stage queues depending on the topic.
Finally, it proposes one bounded fix and explains regression risk. GPU interviews reward tradeoff ownership: what improves, what may regress, and how you would know before tapeout or release.