Hiring insights
Why ATS Keyword Matching Is Not Enough
Why ATS keyword matching is useful for retrieval but unreliable as a true evaluation method for specialist hiring decisions.

Keyword matching is the default logic behind most resume screening. It is fast, scalable, and feels objective. The problem is that it measures something different from what hiring teams actually need to know.
Matching a resume against a list of words from a job description tells you whether the candidate used the right terminology. It does not tell you whether they can do the job. For routine, high-volume roles that distinction may not matter much. For specialist roles with layered requirements, it is the difference between a shortlist built on evidence and one built on vocabulary.
Key Takeaways
- Keyword matching is a retrieval tool, not an evaluation method. It surfaces candidates who used the right words, not necessarily candidates who have the right capability.
- A joint study by Harvard Business School and Accenture found that 88% of employers acknowledged their systems screen out qualified, high-skilled candidates who do not exactly match the language in job descriptions.
- Specialist candidates are disproportionately affected: deep practitioners often describe their experience in domain-specific language that does not mirror generic job ad phrasing.
- The alternative is requirement-level evaluation: assessing what candidates have actually done against each specific criterion, rather than counting term overlap.
- Talentranx is built on this model, scoring candidates against explicit requirements rather than keyword frequency.
Why keyword matching became the default
The appeal of keyword matching is genuine. When a role attracts 200 applications, someone has to narrow that field before human reviewers can engage meaningfully. Keyword filters do that quickly. They are consistent in the sense that the same rules apply to every resume, and they are auditable in that the criteria can be inspected.
For high-volume, entry-level roles where the job requirements map cleanly to a short list of skills and the talent pool is large, keyword matching does useful work. It is a reasonable first triage when the cost of missing individual candidates is relatively low and speed matters most.
ATS platforms were built to solve that problem: managing pipeline volume, tracking stages, and keeping applications organised. Keyword matching sits naturally inside that workflow as a way to move large numbers of applications through a funnel. The problem is not that ATS platforms use keyword logic. It is what happens when that logic gets applied to roles it was not designed for.
What keyword matching actually measures
When an ATS or a manual reviewer scans for keyword overlap, the score reflects how closely a candidate’s resume mirrors the language of the job description. That is a measure of linguistic alignment, not capability.
These are not the same thing, and the gap between them widens considerably for specialist roles.
A candidate with ten years of deep risk governance experience may describe that work as “controls assurance,” “regulatory oversight,” or “second-line risk management,” depending on the industry, the firm, and the era in which they built that experience. A job description asking for “risk governance experience” may not match any of those phrases. The capability is there. The keyword overlap is not.
The reverse problem is equally common and arguably more damaging. A candidate who has learned to optimise their resume for ATS systems will mirror the exact language of the job ad, whether or not their actual experience runs deep. Keyword density rewards vocabulary familiarity. It does not test the substance behind the words.
Harvard Business School and Accenture documented this dynamic at scale in their 2021 “Hidden Workers” study, which surveyed more than 2,250 employers across the US, UK, and Germany. They found that 88% of employers acknowledged their automated systems screen out qualified, high-skilled candidates because those candidates do not exactly match the language in job descriptions. The study’s lead researcher, Professor Joseph Fuller, described the result directly: the effort to make hiring efficient was creating a significant portion of the talent shortage that employers were simultaneously complaining about.
The filtering is not random. It disproportionately affects candidates with non-linear career paths, industry changers, and deep specialists whose domain-specific language does not translate cleanly to generic job ad vocabulary. These are often the candidates specialist roles need most.
Why specialist roles make this worse
Keyword matching has a structural weakness: it treats all words as equally weighted signals of fit. In a specialist role with eight or ten distinct requirements, that assumption breaks down fast.
Consider what keyword overlap actually captures in that context. A candidate might match seven of your ten required terms while being genuinely weak on the three that matter most to role performance. Another candidate might match only five terms while having deep, evidenced experience across your four most critical requirements. Keyword scoring will rank the first candidate higher. Evidence-based evaluation would reverse that ranking.
The synonym problem compounds this further. Specialist fields develop their own vocabulary over time, and that vocabulary varies across industries, firms, and career generations. “Stakeholder engagement,” “relationship management,” and “client advisory” may describe the same underlying capability depending on whether the candidate came from consulting, financial services, or the public sector. A keyword filter built from one job ad will only match one of those framings.
Rigid filters create a related problem at the other end. When knockout criteria are applied as boolean conditions, asking whether the resume contains this exact phrase, candidates who meet the requirement but express it differently are eliminated before any human reviewer sees them. The Harvard-Accenture study found that 49% of employers were filtering out candidates for lacking bachelor’s degrees in roles that did not functionally require them. The filter enforced a stated criterion regardless of whether that criterion was genuinely necessary for the role.
For a role where the talent pool is already narrow and every missed candidate represents real cost, these failure modes are not acceptable margins of error. They are structural problems that compound with each requirement added to the role.
What the keyword problem looks like from the hiring team’s side
The damage from over-reliance on keyword matching is not always visible at shortlisting. It shows up later.
A shortlist arrives that looks strong on paper. All the candidates used the right terminology, hit the required experience years, and came from recognisable employers. Interviews proceed. And gradually it becomes clear that the polished linguistic match does not correspond to the depth of capability the role requires.
Meanwhile, two candidates who could have done the job well were screened out at the keyword stage because one described their project management experience as “programme delivery” and the other worked in an adjacent industry with a slightly different vocabulary.
The problem is hard to diagnose because the candidates who were wrongly screened out are invisible. You do not know they existed. You run the interview process with the shortlist you have, make the best decision available, and attribute a later performance problem to candidate quality rather than to a screening process that measured the wrong thing.
What better evaluation looks like
Keyword matching asks the wrong question. “Does this resume contain the terms we specified?” is a retrieval question. What hiring teams actually need to answer is: “What has this candidate done, and how well does that evidence map to what this role specifically needs?”
The practical difference is that evaluation starts with the role broken into discrete, assessable requirements rather than a list of terms to match. Each candidate is then assessed against every requirement individually, with the score anchored to what the resume actually shows: specific outcomes, scope of responsibility, decisions made, environments navigated. A candidate who describes risk governance as “controls assurance” gets credit for that experience if the evidence is there, regardless of whether the exact phrase appears in the job description.
This also handles adjacent experience more honestly. Where keyword matching either counts a term or does not, requirement-level evaluation can give partial credit for related experience while making that distinction explicit. The reviewer can see that a candidate partially meets a requirement and decide whether that gap matters, rather than having the system decide silently through a boolean filter.
Where Talentranx fits
Most ATS platforms are well suited to what they were designed for: organising pipeline volume, tracking application stages, and managing workflow. Keyword matching serves that infrastructure function adequately. As a note on the wider picture, the existing article Why Your ATS Isn’t Enough covers the distinction between ATS as workflow infrastructure and the separate problem of final-stage decision quality.
Talentranx operates at that final-stage layer, after applications have been collected and before interview time is committed. When a job description is uploaded, the system extracts each requirement individually rather than generating a keyword list. Hiring managers review and adjust the requirement framework before any scoring begins, which means the criteria reflect what the role actually demands rather than which terms appeared most frequently in the job ad.
Each candidate is then scored against every requirement separately, with the evidence drawn from what the resume contains. The output shows not just a rank but the reasoning behind each score, so a reviewer can see where a candidate is strong, where they partially meet a criterion, and where the evidence is missing. For specialist roles, that transparency is what produces a shortlist worth defending.
Measuring the wrong thing has a cost
Keyword matching is a retrieval tool that got promoted into an evaluation role it was not built for. It works well enough when requirements are simple, the talent pool is large, and vocabulary is consistent. In specialist hiring, none of those conditions hold.
The consequence is shortlists that reward candidates who learned to game the vocabulary rather than those with the deepest relevant experience. Getting that wrong is expensive. The cost rarely shows up at the screening stage where the decision was made — it shows up six months later when the hire is not working out and the search has to restart.
Sources
- Fuller, J. B. & Raman, M. (2021). Hidden Workers: Untapped Talent. Harvard Business School and Accenture. https://www.hbs.edu/managing-the-future-of-work/research/Pages/hidden-workers.aspx
- Fuller, J. B. (2021). How to tap the talent automated HR platforms miss. Harvard Business School Working Knowledge. https://www.library.hbs.edu/working-knowledge/how-to-tap-the-talent-automated-hr-platforms-miss
- HR Dive (2021). Report: Applicant tracking systems may exclude whole segments of candidates. https://www.hrdive.com/news/report-applicant-tracking-systems-may-exclude-whole-segments-of-candidates/606614/