Do AI website audits hallucinate findings?
An audit holds two kinds of finding: things a browser measured and things a model judged. Only one of them can be invented, and the other is less repeatable than you would hope.
"It reads my URL and invents a list." Reasonable suspicion, and about some tools it is the correct one. But an audit holds two kinds of finding: things a browser measured, and things a model judged. Only the second kind can be made up. The first kind has a different problem.
Are AI website audits accurate?
They are accurate in two different ways, which is why one answer never fits. Measured findings, form field counts, contrast ratios, load timings, are read off your rendered page and checkable by hand. Judged findings, such as whether your headline is clear, are a model's opinion. Those can be invented.
So "does it hallucinate" is the wrong question. The narrower one is better: which findings carry a measurement, which carry an opinion, and does the report tell you which is which. A tool that presents both in the same confident voice has hidden the only thing you needed to know.
Which parts of the report could a model invent?
Only the parts that required a judgement. Revslip checks roughly 200 conversion signals, grouped into five areas, and they do not all work the same way. Two of the groups are counting and timing. Two are opinion. One is both, which makes it the group worth reading slowest.
- Technical and tracking. Measured. Load timings, redirects, whether your GA4 events fire.
- UX. Mostly measured. Field counts, tap target sizes, contrast ratios, heading order.
- Trust. Both. Whether a badge is present is counted. Whether it reassures anyone is judged.
- Copy, offer and pricing. Judged. Every finding here is a model's reading of your page.
The two minute version on your own report: take any three findings and ask whether a stopwatch, a ruler or a count would settle it. If yes, verify it. If no, treat it as a hypothesis with your name on the decision.
Worth checking before you trust any of it: Revslip labels each finding with what it measured.
Why does the same audit change between runs?
Because a model is stable where one answer dominates and unstable where two answers are close. That is arithmetic, not a bug. When the model is nearly certain, small numerical wobbles cannot move the outcome. When it is torn between two readings of your headline, the same wobble flips the answer.
Researchers at Politecnico di Milano, Universidad Politecnica de Madrid and UESTC ran four open models 50 times each on four different GPUs and watched the token probabilities rather than the text. The pattern held everywhere: variation was negligible when a probability sat near 0 or 1, and significant between 0.2 and 0.8. Bigger batches made it worse. Model size made no difference.
Read that onto your report. The confident finding repeats. The marginal one is the one that moves.
How strong is the evidence on repeat runs?
Strong, and pointing two directions at once. One study says the instability is tiny. Another says it moves task accuracy by up to 15%. Both are right, and why they disagree is the most useful thing on this page, because it tells you which findings to re-run.
- 90% of token probabilities moved under 0.01% across 50 runs, and about 1% moved by 5%
- 15% accuracy swing across 10 runs of five models on eight tasks, best to worst gap up to 70%
- Certain is how Google rates browser nondeterminism in every Lighthouse environment
The first figure is the token-probability study above. The second is Atil and colleagues at Penn State, who configured five models for deterministic output, ran eight tasks ten times, and found none returned repeatable accuracy, let alone identical text.
Why the gap? The first study counts every token, and most tokens are near-certain, so the average looks calm. The second counts whether the final answer was right, and the final answer is decided by the handful of borderline tokens. Average stability and outcome stability are not the same measurement.
Now the part that surprises people. The measured half is not perfectly repeatable either. Google's own Lighthouse variability documentation rates browser nondeterminism as certain in every environment, says no throttling strategy mitigates page nondeterminism, and advises that the median of five runs is twice as stable as one. Your load timing wobbles for reasons that have nothing to do with AI.
What founders say when two runs disagree
The reaction is rarely "this is inaccurate." It is closer to "so which one do I believe," which is a better instinct than it sounds. People notice the copy notes reworded while the field count sat still. Then they ask whether a finding that moved is worth acting on. Colour rather than evidence, and it is what the next section answers.
How to separate measurement from judgement
Sort your report into two piles before you fix anything. Verify the measured pile yourself, then decide the judged pile with what you know about your customers. Five steps, about twenty minutes, and it works on any tool's output including ours.
- Run it three times. Same URL, same goal, spaced an hour apart.
- Mark what repeated. Findings identical all three runs go in the measured pile. Anything reworded or missing goes in the judged pile.
- Spot-check the measured pile. Count the form fields. Time the load in your own browser. Two minutes settles it.
- Interrogate the judged pile. For each one, ask what a customer would have to believe for this to be true. If you cannot answer, park it.
- Fix the overlap first. A finding that repeated every run and survived your spot-check is the safest euro you will spend.
If you want the piles pre-sorted, run a free audit and re-run it twice.
The claim we have not published Revslip has not released a repeat-run consistency score for its own audits. Until we do, the three-run test above is the honest way to check us, and we would rather hand you the method than a number you cannot audit.
When an AI audit is the wrong purchase
When the thing you need checked is not on a page a crawler can reach. Revslip reads rendered public pages. It does not test your product, your prices against your market, your logged-in flows, your emails, or your ad targeting. Buy it for the page, something else for the rest.
Two flat noes. No automated scan produces a legal accessibility conformance statement, and treating one as though it does costs more than the audit saved. And if your traffic is too low to tell a real change from noise, read A/B testing on low traffic first.
Our own evidence has a shape worth naming too. Revslip has audited 134 businesses, and every one arrived because somebody already suspected a problem. That sample cannot tell you how common a finding is on the web, only what turns up on sites whose owners were already worried.
Questions people ask about AI audit accuracy
Three come up every time, and all three ask the same thing: how much of this can I check myself?
Can an AI audit measure my page speed reliably?
Once, no. Google's own guidance is to take the median of five runs, because browser and network variability move the number without your page changing. Treat a single speed score as a direction, not a fact.
Does asking ChatGPT for a second opinion help?
Not for the measured half. A chat model pasted a URL cannot run your JavaScript or time anything, which is the argument in why pasting your URL into ChatGPT is not an audit. It helps for arguing with a judged finding.
Should I trust the euro figure attached to each finding?
Only as far as you can rebuild it. Revslip publishes the formula, traffic multiplied by conversion gap, order value and mobile weight, so you can rerun it with your own target.