
Empirank spent August checking not whether a search-enabled OpenAI configuration mentioned businesses, but whether what it said about them held up. The answer engine optimisation research firm tested 1,257 factual claims the model made about 182 businesses it had recommended, and found 24 of them contradicted by the available evidence, according to a study published on 18 August.
Twenty of those 182 businesses, or 11.0%, carried at least one contradicted claim. Five carried at least one error the study graded high severity. The bigger number was 479 claims, 38.1% of the total, that the evidence could neither confirm nor refute.
The audit lands in a measurement market built almost entirely around presence. An IAB framework published on 3 August found that only 16% of brands systematically track AI visibility, while more than 20 vendors sell tools that disagree with one another about the same brand. Cloudflare released a dashboard three days later separating how often an assistant names a brand from how often it cites one. Empirank’s numbers add a third axis to that pair: a brand can be named, cited, and still described wrongly.
Empirank started with a pool of 3,000 sampled businesses and narrowed it to those the model recommended in response to local queries, dropping nine recommendations where the business identity could not be resolved safely. The surviving 182 businesses produced 1,515 statements, of which 1,257 were factual claims capable of being tested against evidence. The rest was descriptive or evaluative language that no evidence set can confirm or deny.
Each testable claim received one of four labels. The evidence directly supported 749 claims, or 59.6%, and contradicted 24, or 1.9%. On another 479, or 38.1%, it neither confirmed nor contradicted, and five more, 0.4%, were worded or evidenced too ambiguously to decide either way. No claim was classified as outdated.
The study is filed under the internal identifier AIV-009 and written by Empirank founder James Tandy. It was sent to the trade title PPC Land, which reported it on 24 August. No first-party Empirank report was publicly available at the time of writing, so the figures rest on that account. Empirank discloses two limits of its own in the study: quality control ran through independent AI agents rather than human reviewers, and the response set was captured from a single OpenAI configuration in August 2026 and then frozen.
Empirank reports a 1.9% error rate and an 11.0% error rate from the same audit, and the gap between them is the argument the study is built on. The first counts wrong claims against all tested claims. The second counts affected businesses against all tested businesses. Both are accurate, and they answer different questions.
Restricting the comparison to the 773 claims where the evidence pointed decisively one way or the other lifts the accuracy figure to 96.9%. The study declines to let that stand for the whole, because it excludes by construction the 479 claims the evidence could not settle. It refuses the opposite move as well. Counting unsupported claims as incorrect would produce an inaccuracy rate above 40%, which Empirank argues would overstate model error and blur the difference between a brand that needs a correction and a brand that needs better public evidence about itself.
“Unsupported does not mean false,” Tandy wrote in an email to PPC Land on 24 August, describing the 479 unresolved claims as a limit of the research rather than a charge against the model.
The denominator matters commercially because agencies and in-house teams do not manage claim totals. They manage named accounts, and one wrong sentence attached to one client is a problem whatever it looks like as a percentage of a pool.
Empirank tested local businesses answering local queries, not B2B software vendors, so the percentages should not be lifted across. The mechanism behind them travels better than the numbers do.
Accreditation is the clearest case. Claims about certifications, memberships and awards were confirmed only 19.2% of the time across the 26 tested, and the study states plainly that the low rate came mostly from missing evidence rather than from accreditations being disproved. A certification that no accessible page documents is not one an audit can call false. It is one the audit cannot call anything at all. B2B software carries the same category in different dress: SOC 2 and ISO 27001 attestations, customer counts, integration lists, uptime figures and named partnerships, all of which do commercial work in a shortlist and many of which live only on a slide.
G2’s Answer Economy report, based on a March 2026 survey of 1,076 B2B software buyers, found 51% starting research in an AI chatbot more often than in Google, and 69% choosing a different vendor than they had planned on the strength of chatbot guidance. Demandbase data released on 12 August recorded ChatGPT referrals to B2B sites rising 303% year on year. G2’s later Buyer Behaviour report, published in July, then found review sites edging back ahead of chatbots as the top influence on shortlists, 38% to 37%, so the weight of the channel is contested inside one vendor’s own data.
Empirank has a commercial interest in the answer being more auditing, since it sells research in the category its study describes. The study does hold a line against its own convenience on one point. Claims supported only by first-party material were confirmed 75.2% of the time, against 57.5% for claims supported only by third-party material, and Empirank stops short of treating that as proof that owned sources produce more accurate answers, citing the difference in size and content between the two groups.
Empirank’s proposed remedy is a sequence rather than a product: check the facts attached to every monitored mention, start with identity and location details, align credentials and commercial claims across the company website and authoritative third-party profiles, escalate high-severity errors, then repeat the prompts to see whether the answer changes. The study does not report the outcome of that last step. It measured one frozen set of responses, once.