For us, the most unexpected thing was that at the end, the three top candidates the AI interview had chosen were exactly the three the CEO had chosen.
This does not mean they are the best candidates for the role. Our system does not work that way. What it shows is that these are the three candidates who fit best with this CEO and with the environment he runs. That is what our system focuses on — an aspect no other automated hiring tool considers, and one only experienced professional recruiters take into account.
It was a deep satisfaction for us to see that, after interviewing everyone tirelessly, the CEO picked exactly the three we had picked.
The CV said one thing. The conservative CEO and our AI, working independently, said another. They agreed.
Finding 3: 100% of the candidates who refused the AI interview were unsuitable
Three of the 18 candidates we contacted refused to do the AI interview, we started with 15 candidates but when someone refused we took the next in line. I insisted personally on meeting them in person, because we wanted to understand the reason.
Here is what those four meetings produced:
– Two of the Three told us, when we asked to meet, that they were no longer interested in the position. Their stated reason for refusing the AI was a justification, not a real preference, they had already decided not to pursue the job.
– One of the four had described his previous role as sales on the CV. In conversation, it became clear it was not sales at all. He had no sales process knowledge. He likely refused the AI because he sensed it would surface the gap.
That is a sample of four, not four hundred. We are not making a universal claim. But the pattern is clean: in this test, every candidate who refused the AI interview turned out to be either uncommitted or misrepresented. The refusal acted as a self-filter for unseriousness. None of them was a viable hire being lost to the technology.
This is the finding we did not expect, and it is the one that has shifted how we think about deployment.
The simplest explanation for the alignment between our AI and the CEO is that we did not use a traditional job description.
A traditional JD is a list of tasks and required skills. It is written for keyword matching. It tells a candidate what they will do, not what the company actually needs.
For this role, we built what we now call the
Hiring Standards. It includes four sections that a traditional JD almost never contains:
1. Winning standards. What does the team leader define as success in this role? Not titles or KPIs in the abstract — the specific behaviors and outcomes that, if observed, would mean the hire was correct.
2. Failing standards. What disqualifies a candidate? What would the team leader walk away from in an interview?
3. Day-one context. What is the candidate stepping into on their first day? Who do they report to, what is the state of the team, what are the unwritten rules?
4. 30- and 90-day outcomes. What does the team leader expect this person to have produced in their first 30 and first 90 days?
This brief is what our AI optimizes against. It is also, when you read it back, what the CEO was implicitly using to evaluate candidates in his own head. The reason the AI and the CEO converged on the same top three is that they were evaluating against the same definition of fit, and that definition lived in the brief, not in the JD.
Most recruiters resist building this kind of brief. It takes longer than rewriting last year’s JD, and it requires the hiring manager to articulate things they normally hold in their head. The recruiters who do build it find that everything downstream sourcing, screening, interviewing gets easier.
What changed afterward
We expected the CEO to evaluate the results and tell us whether the system was good enough.
What he actually said was: “Can we run the next position now?”
He was not ready to hire for the next role yet. But the cost of interviewing had dropped to a point where he was willing to start the process anyway, to widen the funnel, to see more people, to take a chance on candidates he would normally have filtered out at the CV stage.
This is the behavior change we did not design for. The AI interview did not just produce a better ranking. It made the CEO willing to consider more people, because the cost of being wrong had become trivial. Ten hours of his time, worth about 400 USD, replaced for the price of a structured conversation he did not have to attend. The unlock was not the accuracy. The unlock was the courage to broaden the funnel.
What we are tracking next
We are committed to the principle that the only thing that measures our success is the result we produce.
The candidates the CEO hired from this test are now in their first weeks on the job. We will publish a follow-up in 90 days with their performance against the standards we set at the start. If our ranking was correct, the top hires should be producing against the 30- and 90-day outcomes in the brief. If we were wrong, we will say so.
We do not yet have proof that the system works at scale. We have one parallel test, in one industry, with one role. We are going to run more.
If you run a small team in hospitality or retail: we will run the test for free
We are looking for the next two parallel tests, and we are willing to fund them.
If you are a hospitality or retail company hiring for a sales, operations, or front-line role in the next 60 days, we will run the parallel test against your own judgment at no cost. We will give you 2,000 USD in credits on the platform. In return, we ask for two things: permission to publish the results in an anonymized case study, and access to your hires’ performance data at 90 days.
If that interests you,
contact us. We will tell you in one call whether your role is a fit for the test.
Resume Ranking vs. AI Candidate Screening: FAQ
Q1. Does resume ranking predict the best hire?
Not reliably. In this test, a CEO let his shortlist for a Junior Sales Developer role be ranked two ways by resume and by a 15-minute AI voice conversation. When we compared the resume-only ranking against his final picks after he personally interviewed every candidate, the resume ranking missed the best-fit hire 14 out of 15 times. A resume shows where someone has been; it does not show how they think or whether they fit the role.
Q2. Is AI candidate screening more accurate than resume screening?
In this test, yes. Teracrowd’s AI voice screening ranked the same candidates by how they actually answered for the role, and its top three matched the CEO’s own top three in the same order after he interviewed all of them for 30 minutes each. The resume-based ranking, by contrast, was wrong 14 out of 15 times. AI candidate screening worked because it evaluated each person against the role’s real success criteria, not the formatting of their resume.
Q3. What happens to candidates who refuse an AI interview?
In this test, every candidate who refused the AI voice interview turned out to be unsuitable for the role — 100% of them. The people who opted out were not the strong candidates; they were the wrong-fit ones. Refusing the conversation was itself a signal.
*The CEO in this story trusted his own judgment. He still does. What changed is that he now has a second pair of ears in the room, one that does not get tired, does not have a bad day, and does not skim CVs. He hired the same people he would have hired without us. He just hired them faster, with less effort, and with more candidates seen.*
That is the result we will be measured on.