The Take-Home Test Was a Small Team’s Cheapest Filter. AI Just Made It Impossible to Trust.
For years the work sample was the honest test: hand people a slice of the real job and see what comes back. Then Karat surveyed 400 engineering leaders across the U.S., India, and China, and 71% said AI has made it harder to assess a candidate’s real skill. The reflex is to lock the test down with proctoring and AI bans. For a small team that’s the losing move. The take-home was quietly doing two jobs, AI only broke one of them, and knowing which one changes everything.
The take-home assignment was the fairest thing a small team had. No whiteboard theater, no trick questions, just a slice of the real work: build this small feature, write this landing page, clean up this messy dataset, and let’s see what you send back. It rewarded people who could actually do the job over people who interviewed well. For a founder or a solo recruiter with no assessment budget, it was the cheapest high-signal step in the process.
That step is now quietly broken, and most teams haven’t updated the ritual to match. They still send the assignment, still read the polished result, still feel reassured by it. The reassurance is the problem. The output looks better than ever and tells you less than ever, because you no longer know whose work you’re looking at.
The Take-Home Was Doing Two Jobs
Here is the thing almost nobody names: a take-home was never one test. It was two, bolted together and pretending to be one. The first job was volume triage. A real assignment takes effort, so it thinned a long shortlist down to the people motivated enough to finish. The second job was skill assessment: of the people who finished, who is actually good?
Those two jobs pulled in opposite directions the whole time, and we ignored it because the test was cheap. As a volume filter, the take-home was always quietly terrible. It filters on free time, not talent. The strongest candidate, the one with a demanding current job and two kids and three other offers in hand, is exactly the person who declines an unpaid four-hour exercise and takes the company that respects their time. The recent graduate with an empty calendar polishes for a weekend. You told yourself you were measuring ability. You were often measuring availability.
AI Broke the One That Mattered
The skill-assessment job was the one worth keeping, and it’s the one AI dismantled. When a model can produce a clean feature, a competent case study, or a tidy analysis in minutes, the finished artifact stops telling you who can do the work. It tells you who has access to the same tools everyone else has. This is not a fringe worry. In Karat’s 2026 engineering interview report, built on a survey of 400 engineering leaders across the U.S., India, and China, 71% said AI has made it harder to assess a candidate’s technical skill. A number that sat comfortably below a third a couple of years ago now describes seven leaders in ten.
Notice what actually broke. AI didn’t make candidates worse. It collapsed the distance between the median applicant and the strong one, at least on the surface of a deliverable produced alone, off-camera, with unlimited time and every tool open. The take-home measures the finished thing. The finished thing is now the least informative part of the whole exercise.
Why Policing It Is the Losing Move
The industry’s first instinct is to defend the old test: lockdown browsers, proctoring software, AI-detection scanners, a stern line in the brief that says no AI allowed. A big company can throw a compliance budget at that arms race and still lose it. A small team cannot, and shouldn’t try. Detection tools flag confident writers and non-native speakers as false positives, surveillance makes your best candidates feel accused before they’ve started, and the ban is unenforceable anyway. You’re asking people to prove they can work without the tools they will use every single day on the job. It’s a test of a world that no longer exists.
The tell is in who abandoned the format first. Karat’s data shows companies moving away from asynchronous take-homes and automated code tests toward live sessions where they watch how someone thinks and works, with firms in China moving fastest of all. They didn’t double down on catching cheaters. They changed what they were looking at, from the artifact to the process that produced it.
What a Small Team Should Test Instead
Stop testing whether someone can produce clean output alone. Everyone can now. Start testing the thing AI hasn’t touched: judgment. Assume the tools are on the table and watch how the person uses them. Give a short, live, paid working session, thirty to forty-five minutes, with a deliberately ambiguous problem and AI explicitly allowed. The signal is no longer in the answer. It’s in the questions they ask before they start, what they choose not to automate, whether they catch the confident-but-wrong suggestion the model hands them, and how they explain a tradeoff out loud. A weak candidate pastes the prompt and ships whatever comes out. A strong one uses the model as a fast intern and stays the editor. You can see that difference in real time. You could never see it in a finished file.
That only works if you protect it, because a live session is expensive and doesn’t scale. You cannot run one for two hundred applicants, which is exactly why the take-home was carrying that volume-filter job in the first place. So split the two jobs back apart. Handle volume upstream, where reading fit against the role, not counting who had a free weekend, gets you to a genuine shortlist. Then spend a real human evaluation on the handful who make it. That upstream sort is the part we built Kynto to carry: score the flood against what the role actually needs, so the short live session you keep is reserved for people worth the hour, and stays human where it counts.
Key Takeaways
- The take-home was two tests in one: a volume filter and a skill test. AI broke the skill test. In Karat’s 2026 survey of 400 engineering leaders across the U.S., India, and China, 71% say AI has made assessing real skill harder.
- Policing the old test is a losing arms race for a small team. Proctoring and AI bans flag the wrong people, push your best candidates away, and test for a work environment that no longer exists.
- Test judgment, not output. Run a short, live, paid session with AI allowed and watch how someone works. Split the volume job back out and handle it upstream, so the human evaluation is spent only on the few who earn it.
The take-home didn’t die because candidates got dishonest. It died because it was always measuring the wrong thing, and AI just made that impossible to ignore. The teams that win the next few years won’t be the ones who policed the old test hardest. They’ll be the ones who stopped grading the artifact and started reading the person.
Table of Contents
A real, human evaluation only works when you can afford to give it to the right few. Kynto scores the application flood against what the role actually needs, so the short live session you keep is reserved for candidates worth the hour.
See how Kynto scores the application flood