An unstructured technical interview process explains about 4% of later job performance. A structured one explains about 26%.
Same hours, same people, same candidates.
Most loops leak that signal in the same 4 places.
The question is designed last. Nobody writes anything down.
The loop runs too long. The bar lives in people’s heads.
Your loop rejects engineers at close to random. From the inside it looks fine.
I have run more than 100 technical interviews and hired over 10 engineers.
I defined the hiring process used across 27 markets at Zalando. Almost every broken loop I have seen fails in the same 4 places.
The number that should end the argument
Unstructured interviews explain about 4% of later job performance. Structured ones explain about 26%.
6 times more of the outcome is explained when the interview has a structure. Same hours, same people, same candidates.
The 4% means this. Interview 2 candidates, pick one, and you mostly measured 3 things.
The interviewer’s mood, the candidate’s nerves, and whether the 2 clicked.
Most teams read that number and agree with it. Then they go back to a loop where every interviewer asks whatever they feel like.
What does a normal loop actually measure?
Four things, none of them the job. Comfort under observation, and shared vocabulary.
Whoever interviewed most recently, and a view formed in the first 4 minutes.
None of it compares between candidates, and none of it predicts the actual work.
Write down the last 5 interviews your team ran. What was each interviewer trying to find out?
If the answer is “whether they are any good”, the hour measured nothing comparable.
Interview fluency. Comfort being watched while thinking. A real skill, and not the one you are hiring for.
Shared vocabulary. “Eventual consistency” outscores the same idea in plain words. You are testing whether they read your panel’s blog posts.
Recency. Last week’s candidate is compared to this morning’s. The older memory has already flattened.
The first 4 minutes. Interviewers form a view early, then spend the hour looking for support. That is what people do without a rubric.
Where does the signal leak out?
In 4 places, and almost every broken loop I have seen leaks in all 4.
The question gets designed after the job ad instead of before it. Nobody writes anything down until the debrief.
The loop runs 6 rounds when 4 would do. And the bar lives in people’s heads rather than on a page.
Leak 1: the question was designed after the job ad
Most teams write the job ad, post it, then think about questions.
The week the first candidate books.
By then the question is chosen for being good. Not for producing evidence about this job.
You ask a graph problem of someone whose year is a data pipeline.
The order that works runs the other way. Write down 3 to 5 things this person must do in their first year.
For each, write what would convince you they can do it. Only then design the question.
There is a whole post on this: write the scorecard before the job ad.
It takes an afternoon and it changes every interview afterwards.
Leak 2: nobody writes anything down until the debrief
Interviewers arrive at the debrief with no written score?
Then the debrief is decided by whoever speaks first, and loudest.
I have watched a strong hire become a no-hire in 11 minutes.
The most senior person said “I just did not feel it”. Nobody wanted to argue.
Scorecards submitted before the debrief fix this at zero cost. Everyone commits in writing, alone, then you open the room.
Disagreement becomes visible and useful.
Leak 3: the loop is too long
Every extra round loses candidates and adds almost no information.
Rounds 4, 5 and 6 mostly re-measure what rounds one to 3 already told you.
Meanwhile the candidate is 3 weeks into your process, with 2 other offers.
4 rounds is enough for almost every engineering role.
A screen, a coding or practical round, a design or depth round.
And a hiring manager conversation.
If you cannot decide after those 4, another round will not help. Your criteria are unclear.
Leak 4: the bar lives in people’s heads
Ask 3 interviewers on your panel what “strong hire” means for this role. Write down the 3 answers.
They will differ. Then your loop has 3 bars, depending on who the candidate drew.
That is the randomness people mistake for rigour.
Calibration is the fix, and it is unglamorous. Everyone scores the same recorded candidate, alone.
Then you compare and argue about why the scores differ. The disagreement is the whole point.
Why do you never find out about a false negative?
A rejected candidate does not send a report in 18 months explaining what you missed.
The cost of rejecting good people stays invisible. A bad hire is loud and remembered.
So teams add rounds, raise the bar vaguely, and reject more good people.
Nothing in the feedback loop tells them to stop.
You see false negatives only when the criteria are written down.
Then a rejection traces to a line. You can ask whether that line was fair.
What changed when I did this at Zalando
We cut the loop and wrote scorecards before the questions.
Then required them submitted before every debrief.
3 things moved. Time to decision dropped, because the debrief stopped being an argument about impressions.
Panel disagreement went up at first. That felt like a problem. It was the process working.
The hires from year 2 needed less rescuing than the ones from year 1.
I cannot give you a clean percentage on hiring quality. Anyone who does is selling something.
The decisions became explainable. Explainable decisions are the ones you can improve.
What does this look like on Monday?
One afternoon and one uncomfortable meeting. Write the 3 to 5 things your next hire does in year one.
Write what evidence would convince you of each. Delete any round that produces no evidence against that list.
Require written scores before every debrief. Then run one calibration session on a real candidate.
- Take the role you are hiring for right now.
- That is your scorecard.
- Delete any interview round that does not produce evidence against a line on it.
- Require written scores before the debrief. No exceptions for senior people.
- Run 1 calibration session where the whole panel scores the same candidate.
Step 5 is the one teams skip. It is also the one that shows you how far apart your interviewers actually are.
Where do teams give up?
On 2 objections, and both come up every time.
“Structure will make us miss exceptional people.” That gets it backwards.
Unstructured loops favour people who present well.
“Our engineers will not fill in forms.” True of a 14-field rubric.
False of a 4-line one that replaces a meeting they hate.
Present-well is a narrower group than good-at-the-job, and an unstructured loop cannot tell them apart.
Common questions
How many interview rounds should an engineering loop have?
Four for almost every role. A screen. A coding or practical round. A design or depth round. A hiring manager conversation. Rounds 5 and 6 re-measure what the first 3 showed. Meanwhile the candidate collects offers elsewhere. If 4 rounds cannot decide it, the criteria are unclear.
Does structure work for senior and staff hires too?
Yes, and the scorecard changes rather than disappearing. At staff the evidence is influence without authority. Choosing what not to build. Handling ambiguity. Those are harder to write down than coding ability. That is exactly why unstructured senior loops go so wrong.
What is the smallest change worth making first?
Written scores submitted before the debrief. It costs nothing and needs no new tooling. It stops the loudest person in the room setting the outcome. Most teams see the effect in the first debrief. The disagreement becomes visible instead of being quietly resolved.
How do you run a calibration session?
Everyone scores the same recording alone, then submits. Only then do you open the room and compare. Spend the hour on why the scores differ, not on the candidate. 2 sessions a year keeps a panel aligned. The first one is always the uncomfortable one.
The short version
Unstructured interviews explain 4% of performance. Structure takes it to 26%.
Write the scorecard before the question. Cut the loop to 4 rounds. Collect written scores before the debrief.
Calibrate the panel on a real candidate.
None of this needs budget. It needs someone to say out loud that the loop is not working.
Want it done with you rather than by you?
I run hiring audits and interviewer calibration through SmartHiring.
The first call is 20 minutes and free. I will say if your loop is already fine.
