Table of Contents
A friend of mine, a solid data center network engineer with about twelve years behind him, walked out of an AI infrastructure engineer interview last month looking a bit shell-shocked. He told me he’d prepared for BGP, EVPN and spine-leaf, which he knows cold. What he got was forty minutes on why one slow link could drag down a training job running on two thousand GPUs.
He didn’t fail. But he said it felt like sitting an exam for a subject he’d only half enrolled in.
That story comes up a lot right now. The title “AI infrastructure engineer” is everywhere: banks, telcos, colo providers, SaaS companies that have decided renting GPUs by the hour is getting too expensive. And nobody has quite agreed on what an AI infrastructure engineer interview should look like. So let me walk you through what people are actually running into.

First, Figure out Which Job it really is
This sounds obvious, but it’s where most people go wrong. The same title is being used for three fairly different jobs.
Some teams want a fabric person. They own the back-end GPU network, the front-end network and the storage network, and they’ll grill you on RoCEv2, InfiniBand, congestion control and topology.
Some want a cluster or platform person. That’s nodes, drivers, Kubernetes or Slurm, GPU scheduling and observability. Networking still comes up, but Linux and automation carry more weight.
And some want an architect. That’s someone who can take “we need 2,000 GPUs by next year” and turn it into a design that fits the building, the power budget and the finance team’s patience.
| If the job description talks about… | Expect the interview to lean toward… |
| NCCL, RDMA, collective communication, lossless Ethernet | Deep networking: PFC, ECN, topology, packet-level troubleshooting |
| Researcher productivity, utilisation, scheduling, tooling | Linux, Kubernetes/Slurm, GPU Operator, automation |
| Capacity planning, build vs rent, vendor selection, power | System design, trade-offs, cost and facilities |
Read the first five bullets of the JD carefully. They usually tell you more than the title does.
The AI Infrastructure Engineer Interview loop is longer than you’d think
Most processes I’ve seen run somewhere between five and seven conversations. There’s a recruiter or hiring manager chat, a technical screen, a deep-dive on networking or platforms, a design round, a troubleshooting round, some kind of coding exercise and a behavioural round. Some companies merge a couple of these. Very few skip the design round.
It’s a lot. Pace your preparation accordingly.
The Early Screens: Have you Actually touched this?
The first technical conversation is short, and it’s really just testing one thing: is your GPU experience real, or have you read a few vendor blog posts?
You’ll get questions like these. Why is AI training traffic so different from normal enterprise traffic? What’s RDMA and why should anyone care? What’s the difference between the front-end and back-end network in a GPU cluster? Where does NVLink fit in?
Take that first question. The lazy answer is “there’s a lot more bandwidth.” That’s true, but it misses the point. What the interviewer wants to hear is that training traffic is synchronised. Thousands of GPUs finish a step and then all try to swap results at exactly the same moment. The whole job moves at the pace of the slowest exchange. So one congested link or one flaky optic doesn’t just slow down one flow. It slows down everyone.
If you can say that in your own words in under a minute, you’re usually through.
The Networking Deep-dive: Where it gets Real
This is the round my friend got caught out in, and it’s where network engineers either really shine or quietly unravel.
The questions go a level deeper than most people prepare for:
- How does RoCEv2 stay lossless? What are PFC and ECN each doing?
- What’s DCQCN, and what goes wrong when ECN thresholds are set badly?
- What’s a PFC storm or deadlock, and how do you avoid one?
- InfiniBand or RoCEv2: which would you pick, and why?
- What’s a rail-optimised topology?
- Why does ECMP struggle when you’ve got a handful of enormous flows?
- What is Ultra Ethernet trying to fix?
That last one has become fairly common since the Ultra Ethernet Consortium published its 1.0 spec in June 2025. Nobody expects you to have read it cover to cover. But you should be able to explain the idea: take the things that used to need InfiniBand or a proprietary fabric, like better transport and smarter congestion handling, and make them work on open, multi-vendor Ethernet.
On InfiniBand vs RoCEv2, please don’t pick a side like it’s a football match. Interviewers can’t stand that. InfiniBand gives you mature congestion control and adaptive routing out of the box, but you’re in a smaller vendor world and you need people who know how to run it. RoCEv2 lets you use the Ethernet skills, tools and optics you already have, and some very large clusters have proven it scales. The catch is that you have to get PFC, ECN and buffers tuned properly, and when it goes wrong, it goes wrong in confusing ways. Lay out both sides, then say what you’d choose for their situation. That last part is what they’re really listening for.
On rail-optimised designs, grab the whiteboard if there is one. The short version: GPU 0 in every server goes to one leaf switch, GPU 1 in every server goes to another, and so on. Most of the heavy traffic then stays within one “rail” and doesn’t have to hop through the spine. Bonus points if you mention the downside too. The cabling gets complicated, and losing one rail switch hurts every job using that rail.
The Design Round: Slow Down Before you Draw
This is where seniority really shows, and honestly, it’s my favourite round to watch.
The prompts sound simple. “Design the network for a 1,024-GPU training cluster.” “We’ve got 256 GPUs today and want 4,096 in two years.” “Design storage for a cluster that checkpoints every few minutes.”
Weaker candidates start drawing switches straight away. Stronger ones ask questions first. Is this training, inference or both? Those behave very differently. How are the models split across GPUs? That decides how much traffic leaves the server and how much stays on NVLink. Is it one team or many? Multi-tenancy brings a whole set of isolation problems with it. And what can the data centre actually give us in power and cooling per rack? Dense GPU racks often run out of power long before they run out of network.
Once you’ve got answers, walk through the layers calmly: the server, the scale-up domain, the scale-out fabric, front-end, storage, management. Talk about oversubscription, failure domains and how you’d grow the thing without re-cabling half the hall.
And don’t forget storage. I’ve seen otherwise good candidates design a beautiful fabric and then say nothing about checkpoints. A cluster that trains brilliantly but freezes every time it saves its progress isn’t a good design. Interviewers know this, and they’ll poke at it.
Troubleshooting: They’re Grading your Method, not your Answer
This round sorts people who’ve run clusters from people who’ve only drawn them.
You’ll hear things like: “Training throughput dropped 30% overnight and nobody changed anything.” Or: “This job is fine on 64 GPUs but crawls at 512.” Or: “PFC pause frames keep climbing on a few leaf ports.”
There’s rarely a single right answer, and that’s the point. They want to hear how you think. I’d go about it something like this. Start by narrowing it down: is it one job or all of them, one rack or everywhere, and since when? Then check the boring stuff, because the boring stuff is usually the culprit. Look for link flaps, CRC errors, a NIC that came up at the wrong speed, a GPU that’s overheating and throttling.
After that, isolate. Run NCCL tests across smaller groups of nodes until the slow one gives itself away. Only then go hunting in the fabric for ECN marks, PFC counters and uneven ECMP. And finish by saying how you’d confirm the fix and what alert would catch it next time.
If you can connect it back to that “everyone waits for the slowest GPU” idea from earlier, even better. It shows you understand why the symptom looks the way it does.
Yes, There’s Usually a Coding Round
Even networking-heavy roles now include one. It’s rarely LeetCode-style puzzles. It’s more practical than that. Parse some switch telemetry and flag ports with rising errors. Check a list of nodes for driver versions and link speeds and report anything that’s drifted. Validate a cabling plan against the intended topology.
Python is the safe choice. Keep it readable, handle errors and think out loud about scale. “This works for 20 nodes, but at 2,000 I’d do it differently” is exactly the kind of thing they want to hear.
The Behavioural Round has an AI flavour
It looks like any other behavioural interview, but the questions tend to circle around one reality: GPUs are expensive and everybody wants more of them.
So expect things like “Tell me about a time you had to push back on a research team” or “How did you handle a capacity request you couldn’t meet?” They want to see that you can say no without making enemies, explain trade-offs to people who don’t speak network, and keep cost in mind. Real stories with real numbers beat polished generalities every time.
What They’re really Listening for
If I had to boil an AI infrastructure engineer interview down to its core, interviewers keep coming back to four things.
Do you understand the workload, or just the hardware? Can you talk in trade-offs instead of absolutes? Have you actually operated something, even a small lab with eight GPUs? And can you connect the dots from the fabric to the scheduler to storage to the training job itself?
If Your Interview is a Month Away
Here’s roughly how I’d spend the time.
Use the first week to get comfortable with RDMA, RoCEv2, PFC, ECN and DCQCN, well enough to explain them to a non-network friend. In week two, study GPU cluster topologies and practise one full design out loud, ideally to someone who’ll interrupt you. Spend week three on troubleshooting scenarios and learning what NCCL tests actually tell you. Save the final week for one coding exercise, four or five behavioural stories and a proper mock interview.
And if you can get your hands on real GPUs at all, whether that’s a cloud instance for a weekend or a test rack at work, do it. It changes how you answer almost everything above.
One Last Thing
If you’re a network engineer with an AI infrastructure engineer interview coming up and you’re feeling a bit nervous, don’t be. Most of what these interviews test is stuff you already know: routing, switching, QoS, data centre design, and the patience to troubleshoot without panicking. What’s new is the workload and a handful of protocols. Learn why AI traffic behaves the way it does, and the rest starts falling into place surprisingly fast.



