95% Require AI Skills. 59% Still Make the Wrong AI Hire. Here's Why.

Share
95% Require AI Skills. 59% Still Make the Wrong AI Hire. Here's Why.

More Than Half of Hiring Leaders Now Prioritize AI Skills Over Domain Expertise; Yet 59% of Organizations Have Already Made the Wrong AI Hire

The hiring landscape has reached a turning point: 53% of hiring leaders now prioritize AI capability over deep domain expertise. Yet despite this shift, 59% of organizations report having hired candidates who appeared AI-proficient during the hiring process but failed to apply AI effectively on the job.

These findings come from TestGorilla's 2026 report, The State of Hiring for AI Fluency 2026. Codepresso, an AI capability assessment and training company, analyzed the report's key findings and explores what they mean for organizations preparing for AI-driven work.

Key Findings

  • Survey respondents: 1,928 hiring leaders across the United States and the United Kingdom, representing 15 industries including technology, healthcare, finance, consulting, and education.
  • 95% of organizations now list AI capability as an official hiring requirement, yet only 26% actually require candidates to demonstrate AI use and validate AI-generated outputs during the hiring process.
  • 37% of organizations set their minimum AI requirement at simply knowing AI tools exist and what they do. Knowing about AI is not the same as being able to use it effectively.
  • Codepresso's perspective: Interviews measure how well candidates can explain what they know. Organizations, however, need people who can execute. Bridging that gap requires practical, job-based AI capability assessments.

What Is AI Fluency?

AI Fluency is the practical ability to integrate AI into real work, produce meaningful outcomes, verify AI-generated results, and adapt as AI technologies evolve.

It is not simply knowing the names of tools such as ChatGPT or Claude. Instead, it reflects a person's working behavior, how they collaborate with AI to solve real business problems. As TestGorilla describes it, AI Fluency is "a pattern of behavior, not a state of knowledge."

The rapid pace of AI adoption makes this distinction increasingly important.

According to Microsoft's and LinkedIn's 2025 Work Trend Index, 75% of knowledge workers already use AI at work, and AI adoption nearly doubled within just six months.

TestGorilla's research reinforces the same trend. Fifty-three percent of hiring leaders said they would rather hire a candidate with strong AI fluency than one with deeper domain expertise. AI capability has quickly become one of the most important hiring criteria.

The challenge is that there is still no universal definition of AI Fluency.

The report notes that if you ask ten hiring managers what AI Fluency means, you are likely to receive ten different answers. When organizations use the same term but expect different capabilities, inconsistent hiring standards become inevitable.

The same issue is emerging across Korea, where terms such as Generative AI Assessment, AI Literacy Assessment, and AI Skills Testing are often used interchangeably. Although the names differ, they all attempt to measure the same underlying capability: whether someone can effectively use AI to perform real work.

Why Do Organizations Fail to Hire AI Talent?

The problem is not a lack of infrastructure, it is that organizations are measuring the wrong signals.

According to TestGorilla:

  • 71% of organizations have formally defined AI Fluency.
  • 50% have established internal evaluation criteria.
  • 95% include AI capability as a hiring requirement.

Yet 59% still report hiring candidates who successfully passed interviews but were unable to use AI effectively once they joined the organization.

The report calls this contradiction the Infrastructure Paradox.

The root cause lies in the hiring proxies organizations have relied on for decades.

Signals such as academic degrees, years of experience, and confident interview responses are easy to evaluate, but they are poor predictors of actual job performance. TestGorilla cites a meta-analysis by Van Iddekinge and colleagues showing that years of work experience correlate with job performance at only about 0.06—essentially no meaningful relationship.

As AI capability becomes a new hiring requirement, many organizations continue filtering candidates through these outdated indicators, widening the gap between hiring decisions and workplace performance.

The report identifies three primary causes of hiring failure.

1. The Awareness Trap

Thirty-seven percent of organizations set their minimum standard at simply knowing which AI tools exist.

This measures exposure, not AI fluency.

2. The Subjectivity Trap

Nineteen percent of organizations leave AI capability assessment entirely to individual recruiters.

Without standardized evaluation criteria, assessments become subjective impressions, rewarding candidates who communicate confidently rather than those who can actually perform.

3. Confusing Explanation with Execution

Thirty-one percent of hiring professionals say they struggle to distinguish candidates who genuinely understand AI from those who merely use AI terminology convincingly.

This is not surprising.

Interviews are designed to evaluate communication—not execution.

As TestGorilla Co-founder and CEO Wouter Durville observed during the company's February 2026 event:

"Ninety-five percent of companies have made AI capability a hiring requirement, but they still haven't figured out how to evaluate it."

Two Candidates Looked Identical in the Interview; Until the Practical Assessment

One example illustrates what TestGorilla calls the gap between explaining AI and using AI effectively.

During the interview, both candidates appeared equally qualified.

One described using ChatGPT and Claude every day and confidently discussed their AI experience. The other also explained how they regularly incorporated generative AI into their work.

On the surface, both candidates seemed equally capable.

The difference only became clear during a practical job simulation.

The first candidate submitted AI-generated work with minimal modification. There was little evidence that they had verified the output, refined their prompts, or improved the quality of the final result.

The second candidate approached the task very differently.

They broke the problem into smaller components, worked iteratively with AI, verified each response, asked follow-up questions to improve quality, and clearly explained why some AI-generated suggestions were accepted while others were revised manually.

Both candidates knew how to use AI.

Only one demonstrated the ability to produce reliable business outcomes with it.

That difference could not be identified through traditional interview questions or discussions about previous experience. It only became visible when candidates were asked to complete work that closely resembled the job itself.

The real cost appears after hiring.

Imagine an organization that hires solely based on candidates' claims of AI experience.

New employees may quickly generate reports, presentations, and meeting materials, but they repeatedly fail to verify data sources, overlook AI-generated inaccuracies, or unknowingly introduce hallucinated information into business documents.

The result is predictable.

Managers spend far more time reviewing, correcting, and retraining employees than expected.

The cost of failing to assess AI judgment during hiring is ultimately paid across the entire organization after the employee joins.

If you'd like, I can continue with the next section in the same publication-ready U.S. style so the entire article reads consistently from start to finish.


Where Are Organizations Setting the Bar for AI Capability?

The research shows that organizations are far from aligned on what constitutes the minimum level of AI capability. Instead of converging around a single standard, companies are spread across four distinct levels of AI Fluency—each representing a very different expectation of what employees should be able to do.

AI Capability Level Definition Organizations
Awareness Understands what AI tools exist and where they can be applied. 37%
Exploration Has experimented with AI for simple tasks such as drafting emails or summarizing notes. 28%
Functional Fluency Can independently complete core job tasks with AI while validating the accuracy of AI-generated outputs. 26%
Strategic Fluency Can redesign workflows with AI and identify opportunities for automation and business transformation. 8%

Overall, 65% of organizations define AI capability at the exposure level (Awareness or Exploration), while only 34% expect employees to demonstrate execution and validation capabilities (Functional or Strategic Fluency).

This gap matters because the outcomes organizations ultimately care about, higher productivity, improved efficiency, and business impact, depend on judgment, verification, and execution, not simply familiarity with AI tools.

At a minimum, Functional Fluency should represent the baseline expectation for AI-ready employees. Yet only 26% of organizations currently set that standard, making it the exception rather than the norm.

Why Standards Matter: Comparing the U.S. and the U.K.

The report also compares hiring practices in the United States and the United Kingdom, demonstrating how organizations' AI standards directly influence business outcomes.

Metric United States United Kingdom
Employees experiencing AI-related work errors due to overreliance on AI (past six months) 33% 13%
Organizations setting Awareness as the minimum AI standard 45% 29%
Hiring leaders struggling to define AI requirements by role 62% 46%
Organizations lacking internal alignment on AI capability expectations 57% 41%

Organizations in the United States reported 2.5 times more AI-related workplace errors than those in the United Kingdom.

The report suggests that U.S. organizations are generally more likely to adopt lower minimum standards for AI capability while assuming the definition of AI Fluency has already been established. The result is significantly higher rates of AI-related mistakes in day-to-day work.

The lesson is straightforward.

The more loosely an organization defines the AI capabilities required before hiring, the more likely it is to pay the price after hiring through lower-quality work, increased oversight, and costly errors.

In other words, hiring quality is determined long before the interview begins—it starts with defining the right hiring standard.


What AI Capability Assessment Frameworks Exist Today?

Over the past two years, IBM, Zapier, and TestGorilla have each introduced frameworks for evaluating AI capability.

Although their approaches differ, they share three core principles:

  • AI capability should not be measured by familiarity with specific tools.
  • Assessment should focus on observable behaviors rather than theoretical knowledge.
  • Evaluation criteria should remain relevant even as AI models and tools continue to evolve.
Framework Released Structure Key Characteristics
IBM AI Skills Framework 2025 AI Awareness → AI Practitioner → AI Expert Built on reskilling more than 250,000 employees. Emphasizes that AI capability should be evaluated relative to each job role.
Zapier AI Fluency Framework Q1 2026 Three capability dimensions across four maturity levels A practical framework designed for organizations beginning AI adoption and AI-enabled hiring.
TestGorilla Five-Pillar Framework 2026 Applied AI, Learning Agility, Systems Thinking, Responsible AI, Human-AI Collaboration Built on industrial-organizational psychology and validated against job performance data. Measures not only individual capability but also organizational impact.

The TestGorilla framework shifts hiring conversations away from simply asking "Can this person use AI?" and toward three more meaningful questions:

  • Can they produce business results with AI? (Applied AI & Learning Agility)
  • Can they be trusted to use AI responsibly? (Systems Thinking & Responsible AI)
  • Can they help others succeed with AI? (Human-AI Collaboration)

One particularly interesting finding emerged from the research.

When approximately 750 hiring leaders were asked to rank the five pillars by importance, responses were distributed fairly evenly. However, Responsible and Ethical AI Use received the highest number of first-place rankings at 20%.

The message is clear: organizations are looking beyond technical knowledge. They increasingly value sound judgment, responsible decision-making, and the ability to work effectively with AI.


How Should Organizations Measure Employees' AI Capability?

Effective AI assessment should collect verifiable evidence, not simply rely on self-reported claims or interview responses.

Today, organizations generally choose from four primary assessment approaches, each with its own strengths and limitations.

Assessment Method What It Measures Best Use Cases Limitations
Behavioral Interviews & Experience-Based Questions AI experience, communication skills, and work history Evaluating cultural fit, motivation, and small-scale hiring Cannot reliably distinguish between explaining AI and applying AI effectively. (31% of hiring leaders report this challenge.)
Certification Exams & CBT Assessments (e.g., AICE) Standardized AI knowledge and foundational skills Company-wide benchmarking and minimum hiring requirements Limited ability to reflect real job contexts and role-specific performance.
Simulations & Capability Surveys Decision-making in hypothetical situations and self-perceived capability Quickly assessing large workforces Subject to self-report bias and does not produce evidence of actual work performance.
Job-Based Practical Assessments Real-world execution, problem-solving, and AI validation habits Hiring decisions, pre- and post-training evaluation More resource-intensive to design and score.

These approaches are not mutually exclusive. In practice, organizations often combine multiple methods depending on their hiring or development objectives.

However, if the goal is to eliminate the hiring mistakes highlighted in the TestGorilla report—specifically, confusing explanation with execution—practical, job-based assessments must be part of the evaluation process.

As demonstrated in the earlier example, it was the practical assignment—not the interview—that ultimately distinguished the stronger candidate.

Rather than recommending a complete overhaul of hiring processes, TestGorilla proposes three practical improvements.

  1. Ask better questions. Replace "Which AI tools do you use?" with questions such as:"Describe a workflow you recently redesigned using AI. What changed? What didn't work? How did you validate the results?"Only candidates with genuine hands-on experience can answer these questions convincingly.
  2. Automate resume screening so recruiters can spend more time evaluating real capability. As AI-generated resumes become increasingly common, recruiters should invest more time in stages where human judgment matters most.
  3. Start with a pilot role. Introduce structured, job-based AI assessments for a single position, compare hiring outcomes, and expand the approach once its effectiveness has been validated.

Codepresso believes these recommendations closely align with the principles we have consistently advocated.

Being taught is not the same as mastering a skill. Exposure is not the same as capability.

The report's "Awareness Trap" extends beyond hiring. Organizations that measure AI training success solely by course completion rates risk making the same mistake during workforce development.


How Codepresso Assesses AI Capability

AI Fluent is Codepresso's AI capability assessment platform, designed to objectively measure employees' ability to apply AI in real work and provide organizations with data-driven insights into their AI Transformation (AX) readiness.

The platform directly reflects TestGorilla's central recommendation: evaluate evidence of execution—not just explanations.

Our assessment methodology aligns with this principle in three key ways.

1. Real Business Scenarios Using Real Business Data

Candidates complete tasks that closely mirror their own work environments across functions such as marketing, HR, management, customer service, and R&D.

Rather than answering theoretical questions about AI tools, they must define problems, collaborate with AI, and produce practical business deliverables.

2. Multi-Dimensional Assessment That Evaluates the Entire Process

We evaluate more than the final output.

Assessment includes prompt histories, planning documents (such as PRDs), source code, intermediate deliverables, and other evidence generated throughout the problem-solving process.

This reflects TestGorilla's emphasis on metacognition—evaluating not only what candidates produce, but how they arrive at the result.

3. A Unified Enterprise-Wide Assessment Standard

Although candidates complete role-specific practical tasks, every assessment is evaluated using a common competency framework.

This allows organizations to compare AI capability consistently across departments and job functions while generating reliable data to support enterprise-wide AI transformation strategies.

AI Fluent offers structured assessment tracks across areas such as AI Vibe Coding, AI Agent utilization, and Prompt Engineering, with multiple proficiency levels tailored to different job functions.

After completing the assessment, each participant receives a personalized AI Capability Report that helps individuals identify development opportunities while enabling organizations to measure training effectiveness and ROI.

For Codepresso, assessment is only the beginning.

Our approach creates a continuous improvement loop:

  • Assess current capability through AI Fluent and SkillCertify.
  • Develop capability through SkillCamp, SkillPath, and SkillFit.
  • Reassess to quantify measurable improvement.

Training completion rates alone do not demonstrate business value.

What matters is measurable capability growth before and after learning.

That is the foundation of Codepresso's approach to fair AI capability assessment—and the starting point for establishing standardized AI literacy across organizations.

Book a Demo
Book a Demo →

Frequently Asked Questions (FAQ)

What is AI Fluency?

AI Fluency is the ability to integrate AI into real work, produce meaningful outcomes, validate AI-generated results, and adapt as AI technologies evolve. It is not simply knowing the names of AI tools—it is a behavioral capability demonstrated through everyday work. Organizations such as TestGorilla, IBM, and Zapier have all introduced frameworks for measuring AI Fluency.

What is an AI capability assessment?

An AI capability assessment measures how effectively employees or job candidates can apply AI in real work using observable, verifiable evidence. Rather than relying on interviews or self-reported surveys, effective assessments evaluate performance through realistic job-based tasks and the process used to produce results.

Are AI Literacy Assessments, Generative AI Assessments, and AI Skills Tests different?

Although these terms are often used interchangeably, they all aim to measure an individual's ability to apply generative AI in the workplace. The more important distinction lies in how capability is measured. Some assessments focus primarily on knowledge through quizzes or certification exams, while others evaluate real workplace performance through practical tasks.

What should organizations look for when selecting an AI assessment platform?

Organizations should evaluate five key criteria:

  • Does it measure behavioral capability rather than knowledge of specific AI tools?
  • Does it include practical, job-based tasks?
  • Is it tailored to different job functions?
  • Does it assess both the final output and the complete work process, including prompts and intermediate deliverables?
  • Is the scoring methodology validated, transparent, and explainable?

Assessments centered on the usage of a specific AI tool quickly become outdated as AI technology evolves.

If we're already providing AI training, do we still need AI assessments?

Yes.

Training alone cannot demonstrate effectiveness.

According to TestGorilla, 73% of organizations with clearly defined AI capability standards believe upskilling existing employees is more effective than hiring externally. However, meaningful upskilling depends on measurement.

Organizations should establish a baseline through pre-training assessments and then measure capability improvements through post-training reassessments. Only then can they demonstrate the true ROI of AI learning and development.


References

  • TestGorilla. The State of Hiring for AI Fluency 2026 (Survey of 1,928 hiring leaders across the United States and the United Kingdom).
  • Microsoft & LinkedIn. 2025 Work Trend Index.