Skip to content
Anil Thapa
Leading data teams

An interview that asks for a decision

Most data interviews test whether a candidate can answer the question as asked. The job is to change a decision. Hand the candidate a bounded decision with incomplete evidence, ask for a recommendation, an alternative and a condition, and listen to the reasoning.

7 min read

The interview loop most data teams run tests whether a candidate can answer a question as asked. A SQL exercise with a correct result. A case with a tidy dataset and a known conclusion. A take-home that rewards polish. Each measures something real, and together they select for the skill the job needs least urgently: producing a correct answer to a well-posed question.

The job, past the first year, is to change a decision. That means naming the decision when nobody has, asking what the number is for, saying what the evidence cannot support, and bringing a recommendation someone can act on. None of the usual loop asks for that, so the people who are good at it look the same in the loop as the people who are not. Building a data function from zero calls this hiring for the shape of the problem. This is the interview I use to do it.

One decision, forty-five minutes

The format is a work sample, and it is deliberately small. No take-home. No preparation. One decision, a page of context, three numbers, a stakeholder with a stated preference, and forty-five minutes.

The page describes a situation the candidate can understand without knowing the business: a team of a given size, a request from a named role, a constraint, and a deadline. The three numbers are enough to reason with and not enough to settle it. The stakeholder’s preference is stated because real decisions arrive with one, and how a candidate handles a preference that the evidence does not support is most of what the interview is for.

An illustrative version, as I would hand it over:

The operations lead wants the dispatch report refreshed every five minutes instead of hourly, before next week’s peak. The report feeds a planning meeting and the dispatch queue. The platform team has two days of capacity this week. Three numbers: the current refresh cost, the number of missed dispatch cutoffs last month, and the share of report opens that happen during the planning meeting. The operations lead has already told her director it is happening.

It is the same decision as in What a faster refresh actually buys, because it is small enough to carry into a room. The interviewer answers questions from a sheet. The sheet covers about half of what a candidate might ask; for the rest it says that nobody has measured that, because handling an unknown is part of the test. By the end the candidate has to produce three things: a recommendation, an alternative they would also accept, and the condition that would change their view. That is the shape in Recommendations that go nowhere, and a candidate who has never seen it can still produce it when asked.

What to listen for

The answer matters less than the route to it. Seven things, roughly in the order a strong candidate reaches them; the last can arrive anywhere.

Do they name the decision and its owner? “Whether to change the refresh schedule, and the operations lead decides” is a different start from “let me look at the numbers.” The first candidate knows who they are writing for.

Do they ask what the number is for? The one question that separates answers from charts. Who reads the report, when, and what do they do differently with a fresher one? A candidate who asks it is applying, unprompted, the sentence Built, launched, ignored asks for before anything is drawn: who decides, when, and what changes with the answer.

Do they split the request? The planning meeting and the dispatch queue need different things, and the request bundles them. Noticing that is the analytical move of the exercise, and it is the one the strongest candidates make without prompting.

Do they offer an alternative they would actually accept? A straw option is easy to spot. A real one, “five minutes for the dispatch view only, hourly for the rest,” shows they have weighed the stakeholder’s constraint and not only the evidence.

Do they state a condition? “If a cutoff is missed in the first month, restore it” turns a recommendation into something the owner can say yes to safely. Candidates who hedge instead, “it depends, we’d have to monitor it,” are describing the condition without committing to it.

How do they handle the preference? The operations lead has committed in public. A candidate who ignores that is not ready for the room; one who folds to it is not useful in the room. The answer I want is the one that gives the lead a way to keep her commitment and change the detail, which is what the alternative is for.

And the seventh, which can appear anywhere: do they say “I don’t know” with a plan attached? “I’d want to know how long a dispatch change takes to carry out, and I’d ask the dispatch lead” is the sentence a senior person says constantly, and a junior one is afraid of.

Score it the same way every time

The format only works if every interviewer scores it against the same rubric, written before the first interview. Seven rows, one for each thing above, with the “I don’t know” row scored from wherever in the conversation it appears; three levels each, with an example of what each level sounds like. Interviewers score independently before they talk. The conversation afterward is about disagreements between scores, which is where the interviewers learn what the team actually values.

Frank Schmidt and John Hunter’s meta-analysis of 85 years of selection research, in Psychological Bulletin in 1998, reported validity coefficients for predicting job performance of about .54 for work sample tests and .51 for structured interviews, against about .38 for unstructured interviews. Google’s re:Work guide to structured interviewing describes its own finding that the same questions and rubric for every candidate predicted performance better than unstructured interviews across functions and levels. Those are averages across many jobs, not a measurement of this exercise, so they support the structure rather than the specific questions.

The rubric does a second job. If three interviewers cannot agree on what a good answer to the exercise sounds like, the team does not yet know what it wants from the role, and that is worth finding out before the first offer.

The same exercise, inside the team

The decision interview is a promotion conversation with the names changed. Someone who wants the next role can be handed the same page and asked for the same three things, and the reasoning says more about readiness than a year of completed tickets. It is also the bounded decision I would use to develop judgment in When every decision still needs you: one decision, a recommendation, a review of the reasoning before supplying the answer.

Run inside the team, it has a cost the external version does not: the person knows the business, so the exercise has to be about a decision they have not already seen made.

What it costs

Building and calibrating the case. A good page with three numbers that do not settle the question takes an afternoon to write and two or three practice runs with current team members to calibrate. The numbers have to be invented so that no one has an advantage from knowing the business, and checked so that they do not accidentally make one answer obvious.

Candidates who came to show depth. Someone who has prepared for a SQL round and gets a decision instead can feel short-changed. Keep one technical screen, say in advance that the loop has a decision exercise, and tell them what the three outputs are. The exercise should never be a surprise; the decision is the test, not the format.

It screens out good executors. A strong engineer who is not interested in the decision will score poorly, and that is correct only if the role needs judgment. For a role that needs execution, say so and run a different loop. The failure is using this interview for every role because it is the one the hiring manager likes.

Interviewer time. Independent scoring and a calibration conversation take longer than a debrief where the loudest interviewer decides. That time is the price of the format being fair, and of the team learning what it values.