Skip to main content
The journal
Field notesSeptember 202610 min

The take-home test has become a prompt. We rebuilt the assessment and the trade-offs are now visible.

The take-home test has not disappeared, but the signal it once produced has degraded. Karat's 2026 Engineering Interview Trends survey of four hundred engineering leaders found that sixty-two per cent of candidates use AI during assessments even in formats that prohibit it. We replaced the take-home with a structured live pairing session. The new format costs more in panel time and narrows the candidate pool. It predicts delivery better.

By
Graham Head
Chief Executive
The take-home test has become a prompt. We rebuilt the assessment and the trade-offs are now visible.

Six months into this year, our senior engineering hire was reviewing a full batch of take-home submissions for a mid-level role on the payments team. She came to me with a problem: a cluster of the submissions looked, in her words, as though they had been written by the same person. They had not been. The candidates were in four different countries. When we compared submissions against output from a standard AI coding tool, a substantial portion showed structural similarity we could not dismiss as coincidence. The remaining submissions were harder to classify. The take-home had not failed in the sense of producing nothing useful. It had failed in the sense that we could not, in good conscience, trust what it appeared to show.

I want to be precise about what that failure means, because the imprecise version leads to the wrong conclusion. This is not a question about candidates cheating. We stopped treating Stack Overflow searches as cheating around 2012 and have not revisited that position since. The question is not whether AI tools are permitted. The question is whether the assessment still measures what you designed it to measure. By the middle of last year, ours was not.

#02What the test was measuring

The take-home test worked because it separated candidates along a dimension that correlates reasonably well with engineering competence: the ability to produce working software, alone, in a defined time window, starting from nothing, without access to a colleague who could fill in the gaps. The signal was always imperfect. Candidates who worked slowly but systematically looked like candidates who guessed and got lucky. Candidates who polished presentation over completeness looked better than counterparts who delivered working software with rough edges. These distortions were known and managed by reviewers who had seen enough submissions to weight them appropriately.

What has changed is the cost of traversing the dimension the test was designed to measure, with rather less underlying competence than the test was intended to detect. The person who uses an AI coding tool to produce a working solution they do not fully understand is now considerably harder to distinguish from the person who uses the same tool to accelerate a solution they understand completely. The submission looks similar. The review time is similar. The reviewers, in our experience, cannot reliably tell the difference without a follow-up conversation that the take-home format does not provide.

Karat's 2026 Engineering Interview Trends report, drawn from a survey of four hundred engineering leaders in the US, India, and China, found that seventy-one per cent consider AI to have made technical skills harder to assess. Sixty-two per cent of candidates use AI tools during assessments even in formats that explicitly prohibit them. I am citing these figures not because survey data of this kind is precise, but because they are the closest published approximation to what we were observing internally. The direction is not in doubt. The exact numbers are.

#03What we tried first

We ran live coding interviews for six months before arriving at the format we now use. The result was instructive in ways we had not anticipated. Completion rates fell. Not dramatically, but enough to make the selection effect visible in our offer-stage data: the candidates willing to complete a timed, observed coding session on a video call were not drawn at random from the pool of candidates who would complete an asynchronous take-home. We did not have a clean read on which direction that bias cut. Senior engineers on the panel privately held different theories. The theories were not resolved.

The panel time per hire increased substantially. At our current hiring volumes, that was manageable. It would not remain so at larger scale. We noted this as a constraint, not a reason to stop, because the signal improvement was real enough to justify continuing the experiment.

We also ran a narration variant: candidates were asked to verbalise their reasoning as they worked through the problem. The signal improvement over silent live coding was meaningful. The format created a systematic disadvantage for non-native English speakers and for candidates who do their best technical thinking quietly, without external commentary. Both groups are not randomly distributed across the engineering candidate population. We recorded this and have not solved it.

#04The format we run now

The session we now run is ninety minutes. AI tools are permitted and the candidate is told to use them. The problem is drawn from a real product context: not a standard algorithm exercise, not a variant of a platform coding test, but a simplified version of the kind of ambiguous, constraint-bearing problem the candidate would encounter in the role. Payments infrastructure. Rate-limiting design. Schema migration under a dual-write constraint. Problems with more than one correct approach and real trade-offs between them.

The session divides into three phases. The first twenty minutes are problem understanding and scoping: we are watching how the candidate asks questions, what assumptions they name, and where they push back on the specification. An engineer who treats a spec as a requirements document to be executed and an engineer who treats it as a starting point for negotiation are not the same engineer. The middle fifty minutes are implementation. The final twenty minutes are retrospective.

The retrospective is where most of the assessment happens. A candidate who used AI to produce something they do not understand cannot answer questions about what they would change, what the edge conditions are, or why they made the choices they did. A candidate who used AI to accelerate a solution they understand completely can answer all of these directly, and can usually anticipate the questions before they are asked. You are not evaluating what they built. You are evaluating whether they know what they built.

This is also where seniority signal lives most clearly. The question we use is simple: what would you change about this before putting it in production? The answers range from a blank look to a precise account of what the current implementation does not handle, how you would instrument it, what the performance characteristics look like under load, and how the migration story changes if the requirements shift. That range existed in take-home responses. Observed live, with follow-up, it is considerably more reliable.

You are not evaluating what they built. You are evaluating whether they know what they built.

#05The counter-argument we have not solved

The case against this format is real and we carry it. A ninety-minute live session, with two engineers on the panel plus scheduling coordination, costs substantially more in senior engineering time per candidate than reviewing a take-home that one person can assess in under an hour. The format requires candidates to be synchronously available, which is a filter that is not neutral. Time zones, caregiving responsibilities, and the working preferences of candidates who deliver their best output asynchronously all affect who can participate. We have extended our scheduling window and introduced some choice of problem type. We have not eliminated the narrowing effect.

The format also creates a natural self-selection effect at the scheduling stage, which compresses the funnel. We run fewer assessments per hire than we did when the take-home was the primary screen. Some of the reduction represents lower-quality applications that low-barrier formats attract, and which are largely absent from the live session pool. Some of the reduction may represent candidates we would have wanted, who declined when they saw the time commitment. We do not know the ratio.

The commercial argument for running this format is not that it is cheaper than what it replaced. It is not. The argument is that running a cheaper version of an assessment that no longer measures what you need it to measure is a poor use of the money spent on hiring. Engineering hires at mid-level and above take, in our experience, between six and twelve months to reach the contribution level that justifies their cost, accounting for onboarding, the reduced capacity of colleagues supporting them, and the time to become fluent in the specific system. A poor hire discovered at month four is not made inexpensive by the fact that the assessment producing it was efficient to run. We have counted the cost of both outcomes. We are running the more expensive format.

Two changes would cause us to revisit. The first is a reliable asynchronous evaluation method that preserves signal in the presence of AI tools. No format we have reviewed does this at the fidelity we need. The second is a development in AI capability that makes the engineering judgement we are now trying to assess also largely automatable, at which point the hiring decision changes shape entirely, and the format of the technical screen becomes one of the smaller items to reconfigure.

About the author
Graham Head
Chief Executive

Every piece in the Journal is written personally by a senior practitioner, drawing on the engagement that motivated it. No ghostwriters, no content team, no models. If a paragraph here resonates with a problem you are looking at, the author is the person to reply to — direct lines beat anonymous inboxes.

Get in touch with the practice