Resumap Research · 2026-07-20 · study #2
All 36 of our templates went through an AI recruiter's screening. With the right content, they scored 73–100%.
Round two of our ATS series: the 36-template corpus from the Zoho Recruit parsing study went into Manatal, an ATS with AI candidate scoring, against a job opening that mirrors the resume. When the resume states its evidence explicitly, every one of the 36 templates parsed cleanly and scored 73–100% (average 87.6%), including a perfect 100%. The same templates carrying a vaguer wording of the same true facts scored 34–78%. The template carries the parsing; the words carry the score — and this page publishes both runs, the verbatim AI verdicts and both corpora so you can re-run everything.
Methodology
- Same corpus, second ATS. The 36 post-fix template PDFs from our Zoho study (one fictional senior-engineer CV, identical content, unique phone per file) were imported into a Manatal trial via its bulk resume upload.
- A mirrored job opening. We posted a "Senior Software Developer" job whose description mirrors the CV; Manatal's AI extracted 8 required and 2 preferred criteria from it — the exact stack the resume carries.
- Base run. All 36 candidates were attached to the job and the AI match score shown on each candidate card was recorded, together with the Required/Preferred/Missing counters and the per-requirement AI verdicts where the trial allowed generating them.
- Evidence rewrite.We rewrote the resume so every requirement the AI had rated "Potential Match" carries explicit evidence in the experience bullets — "in production on Kubernetes (EKS)", "10 years (2016–2026)", "deep PostgreSQL (row-level security, partitioning, query-plan tuning)", "3 years of hands-on React frontend development". Same three employers, same dates, same projects — zero invented facts.
- Tuned run + repeatability. The tuned content was rendered in all 36 templates and imported; 5 of those templates had already been imported with identical tuned content in a pilot batch, giving five same-content pairs to compare across runs.
We publish what the product UI displayed — scores, counters and verdict texts — and make no claims about how the vendor's scoring is implemented internally.
Results, part 1 — parsing: what the ATS extracted
Finding: in the evidence-tuned run, 36 of 36 Resumap templates delivered every job title, employer and date range to Manatal's parser, and 70–100% of a 20-item skill list.
- The compact comma-separated skill list parsed far better than the longer grouped list of the base run, which lost up to a third of its entries in the profiles we inspected.
- One base-run PDF parsed to an empty profile (no name, no jobs, no skills) — and the same template parsed perfectly in the tuned run and scored 22/22 in our Zoho study. Parse failures can be transient, which is an argument for testing a resume in more than one system.
- The education line "M.Eng., Computer Science, Distributed Systems" was split differently across templates (four distinct variants) and the degree itself never reached the degree field — worth knowing if a job filters on degree.
Results, part 2 — match scores on baseline wording
Finding: with baseline wording — every required skill present, but evidence left implicit — the same resume scored 34% to 78% (average 48.1%) across templates. Template was the only variable.
The AI extracted 8 required and 2 preferred criteria from the job description and judged each against the parsed resume. On baseline wording it credited requirements only where the bullets spelled the evidence out; skills that appeared in the skills list or in passing were rated "Potential Match" — which counts as missing. The full per-template baseline scores are in the paired table below and the CSV download.
Results, part 3 — match scores on evidence wording
Finding: rewriting the same facts with explicit evidence lifted every template by +11 to +55 points — to 73–100% (average 87.6%), topping out at a perfect 100%. Nothing was invented: same employers, same dates, same projects.
Full paired results, sorted by tuned score. "Skills" is how many of the tuned resume's 20-item list the parser captured for that template.
| # | Template | Baseline score | Evidence score | Gain | Skills (of 20) |
|---|---|---|---|---|---|
| 1 | Vellum | 50% | 100% 🏆 | +50 | 20/20 |
| 2 | Boardroom | 68% | 95% | +27 | 20/20 |
| 3 | Aurora | 61% | 89% | +28 | 14/20 |
| 4 | Axis | 44% | 89% | +45 | 20/20 |
| 5 | Beacon | 39% | 89% | +50 | 20/20 |
| 6 | Bloom | 50% | 89% | +39 | 16/20 |
| 7 | Carbon | 45% | 89% | +44 | 19/20 |
| 8 | Chronicle | 50% | 89% | +39 | 19/20 |
| 9 | Classic Professional | 61% | 89% | +28 | 17/20 |
| 10 | Cobalt | 50% | 89% | +39 | 20/20 |
| 11 | Codex | 45% | 89% | +44 | 19/20 |
| 12 | Compact Engineer | 44% | 89% | +45 | 20/20 |
| 13 | Herald | 39% | 89% | +50 | 20/20 |
| 14 | Modern Professional | 39% | 89% | +50 | 19/20 |
| 15 | Modern Minimalist | 44% | 89% | +45 | 19/20 |
| 16 | Offset | 44% | 89% | +45 | 19/20 |
| 17 | Portrait Premium | 50% | 89% | +39 | 19/20 |
| 18 | Prism | 44% | 89% | +45 | 19/20 |
| 19 | Quarto | 78% | 89% | +11 | 20/20 |
| 20 | Sidebar Mono | 34% | 89% | +55 | 19/20 |
| 21 | Slate | 50% | 89% | +39 | 16/20 |
| 22 | Sovereign | 66% | 89% | +23 | 20/20 |
| 23 | Spine | 39% | 89% | +50 | 19/20 |
| 24 | Syntax | 39% | 89% | +50 | 19/20 |
| 25 | Terminal | 45% | 89% | +44 | 19/20 |
| 26 | Vertex | 44% | 89% | +45 | 19/20 |
| 27 | Atelier | 55% | 84% | +29 | 20/20 |
| 28 | Atrium | 34% | 84% | +50 | 19/20 |
| 29 | Bulletin | 61% | 84% | +23 | 19/20 |
| 30 | Folio | 39% | 84% | +45 | 19/20 |
| 31 | Gazette | 50% | 84% | +34 | 20/20 |
| 32 | Lattice | 50% | 84% | +34 | 20/20 |
| 33 | Linen | 45% | 84% | +39 | 19/20 |
| 34 | Studio | not scored* | 84% | — | 19/20 |
| 35 | Lumen | 39% | 78% | +39 | 19/20 |
| 36 | Premium Editorial | 50% | 73% | +23 | 20/20 |
* The base-run parse failure described in part 1 — nothing for the AI to score.


Results, part 4 — what the AI's verdicts actually said
Four base-run candidates where the trial let us generate the full AI overview. Every quote below judges the same resume content, parsed from different templates.
The evidence bar: skills lists don't count
Atelier · 55%- Matched Experience: "9 years of professional experience, exceeding the requirement"
- Potential Go: "has some experience but lacks specific production experience with Go"
- Potential PostgreSQL: "has some knowledge but lacks deep understanding of PostgreSQL"
- Potential Kubernetes/Docker: "some experience with Kubernetes and Docker but not in production"
- Potential React: "some experience with React but lacks extensive frontend work"
"Potential Match" scores zero. The resume's bullets do mention Go services, PostgreSQL row-level security and a React rewrite — but without the explicit "in production / deep / extensive" framing, the AI discounted them.
The self-contradiction
Atrium · 34%- Not a Match Experience: "Candidate has 9 years of experience but scoring indicates no evidence for this specific requirement"
- Matched Microservices: "proven experience with microservices and Kafka as indicated by their profile"
- Potential GraphQL: "potential ... but lacks concrete examples or evidence"
The same sentence acknowledges 9 years of experience and rejects the experience requirement. On the atelier parse above, the identical text passed.
The flip: same bullets, opposite verdicts
Aurora · 61% vs Atelier · 55%- Matched Kubernetes/Docker (aurora): "experience with Kubernetes and Docker in production environments"
- Potential Kubernetes/Docker (atelier): "some experience with Kubernetes and Docker but not in production"
Identical resume bullets. One parse earned 'in production environments', the other earned 'not in production'.
After the rewrite
tuned run · 73–100%With the evidence written out, the objections above stopped costing points: 34 of 36templates scored 84% or higher and the top uploads reached 100%. (The trial's insight limit was exhausted before the tuned run, so for these candidates we recorded the scores the cards displayed, not fresh per-requirement verdicts.)
Results, part 5 — repeatability
Finding: 2 of 5 byte-identical resumes received different scores on a second upload — one dropped from a perfect 100% to 89%. A single AI match score is a sample, not a grade.
Five templates were imported twice with byte-identical tuned content (only the contact phone digits differ). Same file, same job, same day:
| Template | Upload 1 | Upload 2 | Verdict |
|---|---|---|---|
| Atrium | 84% | 84% | same |
| Atelier | 84% | 84% | same |
| Quarto | 89% | 89% | same |
| Axis | 95% | 89% | differs (−6) |
| Aurora | 100% | 89% | differs (−11) |
A resume that scored a perfect 100% scored 89% on its second, identical upload. If you re-upload the same resume and see a different number, it is not necessarily anything you changed.
What we learned
1. The templates carry the parsing. In the evidence-tuned run, all 36 Resumap templates delivered every job, every date range and the full evidence text to the AI — the text-layer work from our parsing study holds up on a second, AI-first ATS. A template can't win the interview, but it decides whether the AI even sees your experience.
2. The words carry the score: for AI screening, a skill without evidence does not exist. Every requirement the AI rejected was present in the resume — as a skills-list entry or an implicit bullet. Rewriting the same facts with the evidence explicit (what ran in production, for how long, at what depth) lifted every template by +11…+55 points and the average from 48.1% to 87.6%. Nothing was invented.
3. Compact, comma-separated skill lists parse best. The tuned resume's 20-item list parsed at 70–100% across all templates; the base run's longer grouped list lost up to a third of its entries in the profiles we inspected.
4. AI scores are noisy — don't chase a single number. The same bullets earned "in production environments" and "not in production" from the same tool; 2 of 5 identical uploads scored differently (one went 100% then 89%), and the one base-run parse failure did not reproduce. Identical base content spread 44 points across templates; evidence-rich content narrowed that to 27. Treat any single AI match score as a sample — strong evidence raises the whole distribution, which is what actually matters.
Data & reproduction
Everything needed to verify or re-run this study. CC BY 4.0 — cite this page.
- Paired scores (CSV) — base and tuned score per template.
- Evidence-tuned corpus (ZIP, 36 PDFs) — the rewritten resume in all 36 templates.
- Base corpus (ZIP, 36 PDFs) — the same files used in the Zoho study's run 2.
Limitations — read before citing
- One product (Manatal, trial account), one job description, one fictional CV, tested 2026-07-20. Scores were read from the product UI; we make no claims about the underlying implementation, and vendor models change over time.
- Mostly single runs per condition — only 5 templates have same-content repeats. The non-determinism finding rests on those 5 pairs plus the cross-run parser flip; treat effect sizes as indicative, not precise.
- The per-requirement AI verdicts were available for 4 base-run candidates before the trial's generation limit; the Required/Preferred/Missing counters were visible for all scored candidates.
- The tuned rewrite was optimized against feedback from this specific tool — gains may differ on other AI screeners (that cross-tool test is a planned next round).
- Resumap sells resume tooling, including AI tailoring — an obvious interest. Both corpora and every number are downloadable: don't trust us, re-run it.
Questions this data answers
What exactly is an AI match score in an ATS?
Modern ATS products run an AI model over the parsed resume and the job's requirements and produce a percentage plus per-requirement verdicts (matched / potential / not a match). Recruiters use it to rank inboxes. In this test the tool extracted 8 required and 2 preferred criteria from our job description and scored every candidate against them — we recorded exactly what its UI displayed.
Why did identical resumes get different scores?
Two reasons we could observe. First, parsing differs by template — in the base-run profiles we inspected, the same ~50-item skills list parsed to as few as 32 entries, and one template parsed to an empty profile. Second, the AI judge itself is not deterministic: when we uploaded the SAME tuned content twice, 2 of 5 templates scored differently across runs (e.g. Axis 95% then 89%, Aurora 100% then 89%).
What actually raised the score from ~48% to ~88% average?
Evidence, not keywords. The base resume already contained every required skill — but many appeared only in the skills list or without explicit context. The AI marked those 'Potential Match' with comments like 'lacks specific production experience'. The tuned version states the same facts with the evidence spelled out: 'operate 40+ Go and TypeScript microservices running in production on Kubernetes (EKS)', '10 years (2016–2026)', 'deep PostgreSQL expertise (row-level security, partitioning, query-plan tuning)'. Nothing was invented — same employers, same dates, same projects.
Is it fair to tune a resume for an AI reviewer?
Rephrasing your own true experience so its evidence is explicit is exactly what good resume writing has always been; the AI just punishes vagueness harder than humans do. What is NOT fair (and does not survive an interview) is inventing skills or experience — our tuned version added zero new facts, and that constraint is what makes the +40-point average lift meaningful.
Does a 100% match mean the candidate gets the job?
No. The score only ranks how well the parsed resume text evidences the extracted requirements. It decides which resumes a recruiter looks at first — a human still reads, interviews and decides. But at 34% you may simply never be seen, which is why the parsing and evidence layers matter.
Can I reproduce this test?
Yes. Both corpora are downloadable below (36 base PDFs and 36 evidence-tuned PDFs, identical fictional persona). Create a trial on an AI-scoring ATS, import them against a matching job description, and compare the scores the UI shows you. A full run takes about an hour.
The rewrite that scored 100% is what CV tailoring is for
Resumap's tailoring does exactly what this experiment did by hand: it rewrites your real experience so the evidence for each job requirement is explicit — and it is built to invent nothing. The same true story went from 48.1% to 87.6% average here.