GPU Design · All levels

Branch Divergence & Predication: Interview Drills

Interview Drills for Branch Divergence & Predication.

Interview drills

Interview Drills for Branch Divergence & Predication centers on divergence rate, reconvergence delay, and branch efficiency. The objective is to connect profiler evidence to root-cause mechanism and release-safe action.

diagram
PROMPT
You see divergence rate, reconvergence delay, and branch efficiency on Branch Divergence & Predication. Walk through root cause and release decision.

STRONG ANSWER
1. Names failing workload/scene and first broken metric.
2. Explains Divergent control flow serializes paths under lane masks; predication can reduce branch overhead but may execute extra instructions.
3. Requests branch mask timeline, reconvergence stack trace, and predication tradeoff report.
4. Proposes bounded fix + owner + validation matrix.

WEAK ANSWER
Suggests generic tuning without SIMT, warp, cache, or interconnect evidence.

Whiteboard diagram

Divergence and reconvergence masks

diagram
SIMT EXECUTION — Branch Divergence & Predication

warp 0 lanes:  0 1 2 3 4 5 6 7 ... 31
active mask :  1 1 1 1 0 0 1 1 ...  1
instruction :  IF branch taken on active lanes

cycle 10: issue warp 0
cycle 11: issue warp 3
cycle 12: warp 0 reconverges

Focus: trace active-lane loss across branch paths and reconvergence point
Metric tracked: divergence rate, reconvergence delay, and branch efficiency

Debug tree to narrate

diagram
ROOT-CAUSE TREE — Branch Divergence & Predication

divergence rate, reconvergence delay, and branch efficiency regressed
        |
  reproducible on replay?
      /              \
    no                yes
    |                  |
env/test noise    counter triage
                   |
             compute-bound or memory-bound?
                /                  \
             compute            memory/interconnect
             issue stalls       cache/NoC/DRAM stalls

Stop at first failing mechanism, then patch.

GPU deep dive

Warp scheduling quality determines whether latency hiding survives real control-flow and memory variance.

Concept diagram

diagram
WARP SCHEDULING LOOP

ready warp? -> issue -> dependency wait -> reconverge -> issue

Metric graph

diagram
STALL REASON SHARE

long scoreboard    ███████
divergence replay  █████
barrier wait       ███

Reports and artifacts

  • eligible warp ratio

  • stall reason histogram

  • barrier wait cycles

  • scheduler fairness report

Mini case study

A barrier-heavy kernel looked occupancy-safe, but warp arrival imbalance turned sync points into dominant stalls.

Debug branches

  • Compare scheduler policy traces under bursty workloads

  • Measure reconvergence delay and predication side effects

  • Quantify barrier idle time before tuning launch size

Senior review question

Ask: which metric and benchmark pairing proves this topic is truly closed in production context?

Key takeaways

  • Always pair micro-kernel metrics with end-to-end workload impact.

  • Lock toolchain, driver, and launch metadata before comparing performance results.

Common pitfalls

  • Optimizing occupancy without checking memory-system saturation.

  • Comparing profiler captures from different driver or compiler builds.

  • Declaring wins without reproducible accuracy and performance gates.

Interview answer expansion

A strong interview answer for Branch Divergence & Predication starts with the workload and metric, then states the mechanism in plain language: Divergent control flow serializes paths under lane masks; predication can reduce branch overhead but may execute extra instructions.

Then it gives a measurement plan. Good answers name lane masks, issue slots, cache/transaction counters, memory-controller state, NoC congestion, thermal/DVFS telemetry, or stage queues depending on the topic.

Finally, it proposes one bounded fix and explains regression risk. GPU interviews reward tradeoff ownership: what improves, what may regress, and how you would know before tapeout or release.