AI screening tools for recruitment are supposed to filter for quality. Software engineer Dan Kinsky recently put that claim to the test. In a detailed analysis published on Substack, he ran the same resume through HackerRank’s open-source ATS 100 times. The scores ranged from 66 to 99 out of 100. If the company’s cutoff sat at 85, he failed 65% of the time. Same resume. Same tool. Different luck.

That is not a filtering system. That is a coin flip with extra steps.

The findings went viral on Hacker News, LinkedIn, and Reddit within days of publication. They confirmed something recruitment professionals have quietly suspected for a while. AI screening tools produce wildly inconsistent results when asked to make subjective judgments about candidates. Recruitment leader Greg Savage has flagged the same pattern from the other side. As reported in Recruiterflow’s 2026 trends analysis, the same AI screening tool run twice on identical candidate data produced only 14% shortlist overlap.

These are not edge cases. They are the default behaviour of systems that thousands of companies rely on to decide which candidates make it past the first round.

AI screening tool score distribution across 100 evaluations of the same resume showing scores ranging from 66 to 99 out of 100
Score distribution across 100 evaluations of the same resume using HackerRank’s open-source ATS. Source: Dan Kinsky, danunparsed.com

Where AI Screening Tools for Recruitment Break Down

Kinsky’s analysis revealed something important about why AI screening tools for recruitment produce inconsistent results. It comes down to what the AI is being asked to do.

Technical skills scored 8 out of 10 in 98 out of 100 runs. Nearly perfect consistency. The reason is simple. Technical skills are a checklist. Someone either knows Python or they do not. There is nothing subjective for the model to interpret.

Project scores told a completely different story. Massive variation across every run. Sometimes the same projects were described as lacking architectural complexity. Other times they were praised for demonstrating real-world deployment. The model was rolling dice on every evaluation.

Experience scoring was the most concerning finding. A junior engineer with one internship received 25 out of 25. A principal engineer with a decade of distributed systems experience also received 25 out of 25. The scoring rubric was two lines long with no examples, no anchors for what separates a 15 from a 25, and no way for the model to meaningfully differentiate between candidates.

The pattern is clear. AI is excellent at structured extraction and checklist matching. It is unreliable at subjective judgment. Switching to a more powerful model tightened the score range slightly but did not solve the fundamental problem. Project evaluation remained inconsistent regardless of which model was used.

The Legal Landscape Is Catching Up

The inconsistency problem is not just a quality issue. It is becoming a legal one.

De jury van de Workday discrimination lawsuit has reached a critical stage in 2026. A federal court allowed the case to proceed and ruled that software vendors can be held liable as agents of the employer. The tool rejected applications numbering in the billions. The lead plaintiff was rejected from over 100 positions despite being qualified. The court’s ruling means recruitment agencies cannot simply point to the technology vendor if the screening tool produces discriminatory outcomes.

A separate class action against Eightfold AI, analysed by Ogletree Deakins, alleges that the company scraped personal data on over one billion workers, scored applicants on a zero-to-five scale, and discarded low-ranked candidates before any human ever reviewed their application. The case was brought by a former EEOC chair and does not claim the algorithm was biased. It claims the algorithm existed in secret, which is a different and potentially larger legal exposure.

The regulatory environment is tightening in parallel. AI bias in recruitment is now a focus area for legislators on both sides of the Atlantic. New York City’s Local Law 144 already requires annual bias audits for automated employment decision tools. The EU AI Act classifies AI hiring tools as high-risk systems. Transparency rules take effect in August 2026, with full employment obligations set for December 2027.

For recruitment agencies evaluating AI tools, the question is no longer just “does this tool save time?” It is “can I demonstrate that this tool does not make automated decisions that expose my agency to liability?”

AI Is Great at Capture. It Is Unreliable at Judgment.

The data from these studies and lawsuits points to a clear dividing line in how AI should be used in recruitment.

What AI does well in recruitment

  • Parsing a resume or transcript into structured fields
  • Extracting specific data points like salary expectations, notice periods, and availability
  • Matching keywords and skills against a checklist of requirements
  • Transcribing conversations accurately across languages
  • Generating structured summaries from unstructured conversation data
  • Automating data entry into an ATS so recruiters do not have to do it manually

What AI does poorly in recruitment

  • Judging whether a candidate’s experience is “worth” a certain score
  • Evaluating subjective qualities like cultural fit, communication style, or motivation
  • Making pass/fail decisions on candidate quality
  • Ranking candidates in a way that produces consistent results across multiple runs
  • Replacing the judgment that comes from an experienced recruiter actually talking to a person

As Kinsky put it in his analysis, use an LLM to parse a resume into structured data and that works perfectly. Use one to check whether someone knows a specific programming language and that works too. Use one to judge whether a candidate’s experience is worth 18 points or 24 points and you get a vibe check. Something HR teams and bar raisers have spent decades trying to avoid.

This distinction matters because the recruitment industry is moving in two very different directions at once. One direction is fully automated screening where AI makes the decision about which candidates advance. The other direction is interview intelligence where AI captures, structures, and organises the data from real conversations so the recruiter can make better decisions with full context.

The first approach produces inconsistent results and growing legal exposure. The second approach uses AI for what it is genuinely good at and leaves the judgment where it belongs.

What Recruitment Agencies Should Actually Look For

According to data compiled by SSR from Greenhouse’s 2026 AI Hiring Report, 91% of recruiters and hiring managers have spotted or suspected candidate deception. 63% have seen AI-generated resume exaggeration. 40% of tech candidates are believed to have meaningfully inflated their resumes, as reported by Recruiterflow. Deepfake interviews are now flagged as an emerging threat in Deloitte’s 2026 talent acquisition report.

When the CV cannot be fully trusted, the interview becomes the most reliable signal in the hiring process. But only if you capture it properly.

The agencies getting this right in 2026 are not using AI to replace the interview. They are using it to make the interview more valuable. The recruiter talks to the candidate, forms their own judgment based on a real conversation, and the AI handles everything that follows. The transcript is generated automatically. Key data points like salary, availability, and motivations are extracted and structured. The summary is formatted and pushed into the ATS without manual data entry.

The result is that recruiters spend their time on conversations instead of admin, and hiring managers receive structured interview reports within minutes rather than days.

According to Bullhorn’s GRID 2026 Industry Trends Report, top-performing staffing firms are four times more likely to be using AI. But the important detail is what they are using it for. The revenue correlation comes from reducing admin overhead and increasing the number of conversations recruiters can have per day, not from automating candidate decisions.

The practical question for any recruitment agency evaluating AI tools is straightforward. Does this tool make decisions about candidates? Or does it capture data so recruiters can make better decisions themselves? The tools in the first category are the ones producing inconsistent scores, facing lawsuits, and creating compliance risk. The tools in the second category are the ones that firms in the GRID report’s top-performing bracket are actually using.

The gap between a 66 and a 99 on the same resume is not a minor calibration issue. It is a fundamental design flaw in how AI screening tools for recruitment approach candidate evaluation. A tool that cannot differentiate is not filtering for quality. It is just filtering. Recruitment agencies that understand this distinction now will avoid the compliance risk and the quality problems that come with letting AI make judgment calls it was never designed to make.