Automated vs Manual Accessibility Testing
No automated accessibility scanner — including AllyProof's — can evaluate 100% of WCAG on its own, and any tool claiming otherwise should be treated skeptically.
Before the breakdown, one distinction that causes most of the confusion in this space. There are two different things people count:
- Issue volume — what share of the accessibility problems on a real site a scanner finds. Deque's Automated Accessibility Coverage Report measured this across ~2,000 audits and found automated testing identified 57.38% of recorded issues. This is where the widely quoted "57%" comes from.
- Criteria — what share of WCAG success criteria a scanner can say anything about. The same report found automated issues for only 16 of the 50 WCAG 2.1 A/AA success criteria in its dataset.
Automated rules score much better on volume than on criteria because the failures they check for are also the ones that occur most often. Neither number means a clean scan establishes conformance: automated rules detect failure conditions, they never confirm a criterion is satisfied.
The criteria-level picture is a three-way split. The counts below are AllyProof's own classification of the 55 active WCAG 2.2 Level A and AA success criteria.
Automated failure detection available (12 of 55, ~22%)
These have observable rule conditions, such as a missing attribute or a text-contrast calculation. Inspect the recorded evidence and context when reviewing a finding. A rule that did not report a failure does not establish that the entire criterion is satisfied. See axe-core Explained.
Examples: Insufficient Color Contrast (1.4.3), Missing Page Language (3.1.1), Missing Page Title (2.4.2), Invalid ARIA Attributes (4.1.2).
Partially testable with automation (13 of 55, ~24%)
A scanner can detect a pattern that's often (but not always) a real problem, needing human confirmation. Whether an alt attribute is present is mechanically checkable; whether the alt text is a good description requires a person to read it and judge — which is why Non-text Content (1.1.1) sits here rather than in the category above. Heading structure existing is checkable; whether headings are descriptive (per Headings and Labels) needs human judgment.
Manual evaluation required (30 of 55, ~55%)
More than half of the active A/AA criteria genuinely require a human — no current automated approach can assess these at all:
- Whether audio description accurately conveys a video's visual content
- Whether a page's reading order is genuinely logical (versus just structurally consistent)
- Whether keyboard focus order matches a page's intended logical flow
- Whether error messages are genuinely helpful, not just technically present
Why disclosure matters more than the automation percentage itself
The exact percentage matters far less than being transparent about which criteria fall into which bucket for a given scan, and about what a clean result in each bucket actually establishes. A report that clearly separates "confirmed violations," "needs manual review," and "no automated rule — requires manual audit" is more trustworthy and more useful than one that either overclaims coverage or bundles everything into a single "score" with no methodology disclosed. It is also the difference the FTC's April 2025 order against accessiBe turned on: the problem was claims the evidence did not substantiate, not the use of automation.
What this means practically
Treat an automated scan as a strong first pass that reliably catches the most commonly cited, highest-litigation-risk issues (see Web Accessibility Lawsuit Trends — the top violations are almost all ones automation detects well) — but not as a substitute for a full manual audit if WCAG conformance is the goal, particularly for anything customer-facing at real legal or reputational stakes.
Common questions
- What percentage of WCAG can automated tools test?
- Two different numbers get quoted and they measure different things. By criteria: automated rules detect common failure conditions for about 22% of the 55 active WCAG 2.2 A/AA success criteria, partially cover another 24%, and do not touch the remaining 55%. By issue volume: Deque measured that automated testing identified 57.38% of the accessibility issues recorded across its audit dataset. Neither number means a clean scan establishes conformance.
- Can an automated scanner make a site conform to WCAG?
- No. No automated tool can evaluate all of WCAG, and a tool claiming otherwise should be treated skeptically. More than half of the success criteria have no automated rule at all, and even where rules exist they detect failures rather than confirming success.
- What is automated accessibility testing good for?
- As a reliable first pass that catches the most common, highest-litigation-risk issues — missing alt text, low contrast, missing labels — before a manual audit covers the rest.
Related articles
Want to see how your own site scores?
Run a free accessibility scan