Two questions. One of them is free.
Both are live — answer them here, no account. The gap between them is the product. All 10 free questions are readable in full, no sign-up.
Free tier · fundamentals One of the 10 you get without paying. Single layer: read the output, know the state machine.
You are on call. A compute node was recabled during maintenance and MPI jobs refuse to start on it. ibstat on the node shows the output below. What is the most likely cause?
CA 'mlx5_0'
CA type: MT4129
Number of ports: 1
Firmware version: 28.39.2048
Port 1:
State: Initializing
Physical state: LinkUp
Rate: 400
Base lid: 65535
LMC: 0
SM lid: 0
Port GUID: 0x88e9a40300c85e10
Link layer: InfiniBand
Why B
State: Initializing combined with Physical state: LinkUp is the classic no-SM signature: the PHY trained fine (LinkUp, Rate 400 negotiated), but a port only progresses Down → Initializing → Armed → Active when a Subnet Manager assigns it a LID and activates it. Base lid 65535 (unassigned) and SM lid 0 confirm no SM has touched the port. A bad cable would leave Physical state at Polling, not LinkUp; and a port in Ethernet mode would show Link layer: Ethernet, not an InfiniBand state machine at all. Firmware is irrelevant when the link already trained at full rate.
Paid tier · production-down This is what the 50 hardest look like. Every fabric check is green and the cluster is still broken — the layer that is failing is not the layer that is complaining.
A 4-node, 32-GPU all-reduce that sustained 185 GB/s bus bandwidth last week now runs at 22 GB/s. Nothing on the fabric was changed. ibstat reports every port Active at 400 Gb/s, perfquery shows no counter climbing on any rail, and nvidia-smi topo -m still shows the expected NV18 between GPUs inside each node. The job's NCCL_DEBUG=INFO output contains the lines below. What happened?
gpu01:31427:31509 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB
[2]mlx5_2:1/IB [3]mlx5_3:1/IB
gpu01:31427:31509 [0] NCCL INFO Channel 00/0 : 0[07000] -> 1[0b000] via P2P/IPC
gpu01:31427:31509 [0] NCCL INFO Channel 00/0 : 7[c7000] -> 8[07000] [send] via NET/IB/3
gpu01:31427:31509 [0] NCCL INFO GPU Direct RDMA Disabled for HCA 0 'mlx5_0'
gpu01:31427:31509 [0] NCCL INFO GPU Direct RDMA Disabled for HCA 3 'mlx5_3'
Why C
GPU Direct RDMA Disabled is the whole answer, and it is the only line in the output that is not healthy. When ACS (Access Control Services) is enabled on the PCIe switch ports above a GPU and a NIC, every peer-to-peer transaction between them is redirected up to the root complex for translation — which defeats GPUDirect RDMA, so NCCL falls back to bouncing each transfer through a host bounce buffer. Roughly an order of magnitude of inter-node bandwidth disappears and every fabric diagnostic stays clean, because nothing on the fabric is wrong: the loss is on the PCIe path between the GPU and its HCA. Confirm with lspci -vvv | grep -i acsctl and look for SrcValid+ on the upstream ports, then disable ACS in the BIOS or with setpci. The distractors all describe fabric-layer faults, and the fabric layer is the one part of this cluster that is provably fine — a degraded rail would show a lower Rate in ibstat, a routing collision would show PortXmitWait climbing in perfquery, and an algorithm choice cannot switch GPUDirect off.
Pricing
One product, one payment, 1 month of access. No tiers to choose between, no auto-renewal, no card on file, nothing to cancel.
What you are actually buying is the right-hand column of the grid. The 10 free questions are fundamentals and working knowledge — enough to judge the format and the writing. The 150 you pay for include the 50 hardest: the 3 a.m., production-down, multi-layer faults where the symptom is in one layer and the cause is in another. Those are the ones interviews are decided on, and none of them are in the free tier.
ClusterDrill
Everything. There is no cheaper version that leaves out the hard parts.
€17.90 €7.99 · 1 month
Promotional price until 15 September 2026. The regular price is €17.90 and it goes back to that afterwards.
- All 150 questions across all six domains
- All 10 incident labs
- The 50 hardest questions, including the multi-layer faults
- The full sandbox shell with fault injection
- Weak-area drill and the timed interview sim
- Every question added during your window
Get ClusterDrill · €7.99
Free, forever
Free account, no card, no expiry.
€0
- 10 questions — the same format and the same writing, pitched at fundamentals
- One full incident lab
- The entire simulated cluster, unrestricted
- None of the 50 hardest questions — those are the paid tier
- Stays available after a paid window ends
Sign up free — no card
Prices include VAT. Payments are handled by Stripe, who act as merchant of record and issue your invoice — your statement will show LINK.COM*. Access ends on its own — we never charge you a second time.
Questions
Do I need access to real hardware?
No. That's the reason this exists. The cluster is simulated in JavaScript and runs entirely in your browser — four HGX nodes, eight H100s each, four NDR InfiniBand rails, two 200GbE RoCE ports and a Lustre filesystem. The commands return output modelled on the real tools.
Is the simulator actually interactive, or a video?
Interactive. The terminal at the top of this page is the same engine the product runs on, executing whatever you type. If it were a recording, the demo could promise more than the product delivers — this way it can't.
Is this affiliated with NVIDIA, or with any employer?
No. It is an independent study aid. Product names are used descriptively, and nothing here is endorsed by or sourced from any vendor or company.
Will it get me the job?
It will not, on its own, and anyone promising that is selling you something. What it does is let you rehearse diagnostic reasoning on hardware you probably can't get access to, so the first time you read a confusing ibstat under pressure isn't in the interview.
Is there a subscription?
No. One product, one payment, 1 month of access. Nothing renews, no card is kept on file, and there is no cancellation to remember — access simply ends. We chose a fixed window rather than a subscription because the question bank changes over time, and because you almost certainly want this for one interview, not forever.
What happens when my access ends?
The paid questions and labs close, and the 10 free questions and the sandbox shell keep working exactly as before. If you need it again — a second loop, a year later — you buy another window at the same price.
What if it isn't for me?
Email info@impossible-labs.io and we refund you, no argument — see the refund policy for the window and the detail. Try the free tier first — 10 questions and a lab, free account, no card — so you know what you're buying before you pay. Or answer the two questions on this page without registering at all.