DevOpsInterviewPrep logo
AI & GPU Infrastructure / 14
expertNewNVIDIADatabricksMeta

You have eight 80GB GPUs and a 70B model. Tensor parallel across all eight, or two replicas of four? Defend the choice.

Both configurations fit the weights. Compare cache capacity, collective overhead and replica placement under the actual traffic shape; neither layout wins every workload.

Updated Sep 2026 · Grounded in researched DevOps, SRE and platform engineering interview loops, written to a senior-engineer editorial bar, and never padded to hit a word count.

Both configurations fit the weights. Compare cache capacity, collective overhead and replica placement under the actual traffic shape; neither layout wins every workload.

an account raises the per-topic limit · no card
UP NEXT ON YOUR JOURNEY
DISCUSSION · 0

Nothing here yet. Say how you would answer it.