cluster/drill

AI & GPU cluster interview prep

Practise on a cluster you're allowed to break.

150 troubleshooting questions, 10 incident labs, and a simulated 4-node HGX cluster that answers real commands — InfiniBand, RoCE, NCCL, GPU, Lustre and Linux. It runs in your browser. There is nothing to install and no hardware to borrow.

Free tier: an email address, no card 1 month of access, one payment Nothing auto-renews

gpu01.cluster.local — ssh
root@gpu01:~#

Live, not a recording. Try or type your own.


Who this is for

Specialist content, not fundamentals. If you already know what a Xid code is, you're in the right place.

HPC engineers going to industry

You've run Slurm and Lustre at a university or national centre. The AI labs want the same instincts, expressed in NCCL and NVLink vocabulary.

SREs moving into GPU infrastructure

Strong Linux and networking, thin on RDMA. The gap in the interview is almost always the fabric: GIDs, PFC, lossless Ethernet, subnet managers.

Anyone with a loop booked

An on-site at an AI infrastructure team is mostly diagnostic scenarios. This is a rehearsal room for exactly that format.

What's inside

Built around the failures that actually page you, not trivia about protocol history.

150 troubleshooting questions

Every one is a symptom you have to read: a command's real output, and four plausible readings of it. The wrong answers are wrong for reasons worth knowing.

10 incident labs

A fault is injected and not named. You diagnose it in the shell, in sequence, the way the interviewer will ask you to think out loud.

The sandbox shell

The cluster above, yours to explore. Inject a random fault, hunt it, reset. ibstat, dcgmi, lctl, perfquery and friends all answer.

Explanations that teach

Each answer explains why the distractors fail too — the port state machine, the counter that was misread, the layer the symptom actually came from.

Weak-area drill

The app tracks what you get wrong by domain and can hand you only those. Your gaps, not a fixed sequence.

Timed interview sim

30 questions against the clock, mixed across every domain, for the week before the loop.

Two questions. One of them is free.

Both are live — answer them here, no account. The gap between them is the product. All 10 free questions are readable in full, no sign-up.

Free tier · fundamentals One of the 10 you get without paying. Single layer: read the output, know the state machine.

You are on call. A compute node was recabled during maintenance and MPI jobs refuse to start on it. ibstat on the node shows the output below. What is the most likely cause?

CA 'mlx5_0'
	CA type: MT4129
	Number of ports: 1
	Firmware version: 28.39.2048
	Port 1:
		State: Initializing
		Physical state: LinkUp
		Rate: 400
		Base lid: 65535
		LMC: 0
		SM lid: 0
		Port GUID: 0x88e9a40300c85e10
		Link layer: InfiniBand

This is what the 50 hardest look like. Every fabric check is green and the cluster is still broken — the layer that is failing is not the layer that is complaining.

A 4-node, 32-GPU all-reduce that sustained 185 GB/s bus bandwidth last week now runs at 22 GB/s. Nothing on the fabric was changed. ibstat reports every port Active at 400 Gb/s, perfquery shows no counter climbing on any rail, and nvidia-smi topo -m still shows the expected NV18 between GPUs inside each node. The job's NCCL_DEBUG=INFO output contains the lines below. What happened?

gpu01:31427:31509 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB
                                        [2]mlx5_2:1/IB [3]mlx5_3:1/IB
gpu01:31427:31509 [0] NCCL INFO Channel 00/0 : 0[07000] -> 1[0b000] via P2P/IPC
gpu01:31427:31509 [0] NCCL INFO Channel 00/0 : 7[c7000] -> 8[07000] [send] via NET/IB/3
gpu01:31427:31509 [0] NCCL INFO GPU Direct RDMA Disabled for HCA 0 'mlx5_0'
gpu01:31427:31509 [0] NCCL INFO GPU Direct RDMA Disabled for HCA 3 'mlx5_3'

How hard does it get?

The whole bank, by domain and by depth. These numbers are counted out of the question files at build time, not written by hand.

DomainFundamentalsDo you know the vocabulary and the tools?Working knowledgeRead real output, decide what it accuses.Production-downMulti-layer: the symptom is not where the cause is.All
Linux71 free121 free625
InfiniBand81 free141 free830
RoCE51 free10520
GPU71 free121 free625
NCCL61 free131 free625
Storage7121 free625
All domains407337150

Every one of the 10 free questions sits in the two left-hand columns. The right-hand column is not sampled — and it is the column interviews are actually decided in.

Six domains

Weighted the way the failures are — the fabric gets the most, because the fabric breaks the most.

InfiniBand

Port states, subnet managers, counters, cables, routing.

30 questions

Linux

OOM and the killer, NUMA, PCIe, systemd, cgroups, limits.

25 questions

NCCL

Transports, env vars, bus bandwidth, init hangs, ring topology.

25 questions

Storage

Lustre, GPFS, NFS, NVMe, quotas, fio and stripe layout.

25 questions

GPU

Xid codes, ECC and row remapping, throttling, topology, DCGM.

25 questions

RoCE

GIDs, PFC and ECN, DSCP, lossless Ethernet, cross-subnet RDMA.

20 questions

Pricing

One product, one payment, 1 month of access. No tiers to choose between, no auto-renewal, no card on file, nothing to cancel.

What you are actually buying is the right-hand column of the grid. The 10 free questions are fundamentals and working knowledge — enough to judge the format and the writing. The 150 you pay for include the 50 hardest: the 3 a.m., production-down, multi-layer faults where the symptom is in one layer and the cause is in another. Those are the ones interviews are decided on, and none of them are in the free tier.

ClusterDrill

Everything. There is no cheaper version that leaves out the hard parts.

€17.90 €7.99 · 1 month

Promotional price until 15 September 2026. The regular price is €17.90 and it goes back to that afterwards.

  • All 150 questions across all six domains
  • All 10 incident labs
  • The 50 hardest questions, including the multi-layer faults
  • The full sandbox shell with fault injection
  • Weak-area drill and the timed interview sim
  • Every question added during your window
Get ClusterDrill · €7.99

Free, forever

Free account, no card, no expiry.

€0
  • 10 questions — the same format and the same writing, pitched at fundamentals
  • One full incident lab
  • The entire simulated cluster, unrestricted
  • None of the 50 hardest questions — those are the paid tier
  • Stays available after a paid window ends
Sign up free — no card

Prices include VAT. Payments are handled by Stripe, who act as merchant of record and issue your invoice — your statement will show LINK.COM*. Access ends on its own — we never charge you a second time.

Who made this

c/d

Impossible Labs — Independent software studio · Netherlands

ClusterDrill is built and maintained by Impossible Labs. The question bank and the simulator are written against the tools themselves — the output you read in the labs is modelled on what the real commands print.

Questions

Do I need access to real hardware?

No. That's the reason this exists. The cluster is simulated in JavaScript and runs entirely in your browser — four HGX nodes, eight H100s each, four NDR InfiniBand rails, two 200GbE RoCE ports and a Lustre filesystem. The commands return output modelled on the real tools.

Is the simulator actually interactive, or a video?

Interactive. The terminal at the top of this page is the same engine the product runs on, executing whatever you type. If it were a recording, the demo could promise more than the product delivers — this way it can't.

Is this affiliated with NVIDIA, or with any employer?

No. It is an independent study aid. Product names are used descriptively, and nothing here is endorsed by or sourced from any vendor or company.

Will it get me the job?

It will not, on its own, and anyone promising that is selling you something. What it does is let you rehearse diagnostic reasoning on hardware you probably can't get access to, so the first time you read a confusing ibstat under pressure isn't in the interview.

Is there a subscription?

No. One product, one payment, 1 month of access. Nothing renews, no card is kept on file, and there is no cancellation to remember — access simply ends. We chose a fixed window rather than a subscription because the question bank changes over time, and because you almost certainly want this for one interview, not forever.

What happens when my access ends?

The paid questions and labs close, and the 10 free questions and the sandbox shell keep working exactly as before. If you need it again — a second loop, a year later — you buy another window at the same price.

What if it isn't for me?

Email info@impossible-labs.io and we refund you, no argument — see the refund policy for the window and the detail. Try the free tier first — 10 questions and a lab, free account, no card — so you know what you're buying before you pay. Or answer the two questions on this page without registering at all.

Contact us

A question before you buy, something broken, an invoice you need re-issued, or a question you think belongs in the bank — this reaches us directly. We answer within two working days.

We use your address to reply and nothing else — no list, no newsletter. See the privacy policy. You can also just email info@impossible-labs.io.

Start with the free tier

10 questions, one incident lab, and the whole sandbox. Free account, no card.