Growth & Strategy

ChatGPT’s crawlers read the review pages humans skip, log study finds

Written by
Full Name
July 28, 2026
LANY’s analysis of 9.4 million access records from Japanese comparison site IT Trend, published on 13 July, found GPTBot sent 50.5% of its requests to user-review pages against 2.9% of human pageviews, and pricing pages were 18.6 times over-represented for ChatGPT’s answer-time crawler.

A year of server logs from a Japanese software comparison site has given B2B marketers their first page-level view of what OpenAI’s crawlers actually fetch, and it separates sharply from what people read. On IT Trend, a business software comparison site running since 2007, GPTBot sent 50.5% of its requests to user-review pages. Human visitors spent 2.9% of their pageviews there.

The finding arrives while marketing teams are still deciding where to move budget that used to buy search rankings. Most advice on optimising for AI answers has been inferred from citation scraping, which samples AI responses and counts which domains appear in them. Access logs are a different class of evidence: they record what a crawler requested rather than what a model chose to cite, and they come from the site owner rather than being reconstructed from outside.

What did the IT Trend logs show?

LANY, a Tokyo digital marketing agency, published the analysis on 13 July through its LANY LLMO LAB research unit, using logs supplied by Innovation & Co., which operates IT Trend and lists more than 8,400 products. The dataset covers December 2024 to November 2025 and combines 9.4 million records: human behaviour from Google Analytics 4 and requests from OpenAI’s three named crawlers.

GPTBot, the agent OpenAI uses to gather training data, concentrated on review pages. Its 50.5% share includes requests to deleted URLs; strip those out and review pages account for 72.3% of its valid requests. Set against the 2.9% of human pageviews that landed on reviews, that is a ratio of roughly 17 to one. The figure compares shares of attention rather than raw traffic: reviews took 17 times more of GPTBot’s request mix than of the human mix.

ChatGPT-User, which fetches pages in response to a live prompt, showed the sharpest divergence of all on pricing. Pricing pages took 10.7% of its requests against 0.6% of human pageviews, a share 18.6 times higher. Product detail pages took 23.1% against 5.8%. The inversion runs the other way too: category landing pages took 1.7% of the crawler’s requests against 18.9% of human pageviews, and award and feature content 0.7% against 19.5%. The pages a comparison site builds to win human browsing were close to invisible to the crawler; the dry ones marketers rarely fuss over were not.

Why do OpenAI’s three crawlers behave differently?

OpenAI runs three separate agents and, by its own documentation, gives each a different job. GPTBot collects content that may be used to train its models. OAI-SearchBot indexes pages so they can surface in ChatGPT’s search results. ChatGPT-User visits a page when a user’s question sends it there, is not used for automatic crawling, and — because the action is user-initiated — may not follow robots.txt rules. The controls are independent, so a site can allow OAI-SearchBot to stay visible in ChatGPT search while disallowing GPTBot to keep its content out of training.

LANY’s correlations map onto that division of labour. GPTBot tracked internal linking, at a Spearman coefficient of 0.428: pages with no internal links averaged 1.28 requests, pages with more than 100 averaged 12.32. It was indifferent to search demand, averaging 10.3 requests on pages with monthly search volume between one and 99 and 12.7 on pages above 1,000. OAI-SearchBot indexed flat, with pages ranking outside the top 100 averaging 1.78 requests and pages ranking in the top three averaging 4.31, both hitting a ceiling around six. ChatGPT-User followed pageviews most closely of the three, at 0.591.

For a marketing team, that reads as three different tasks rather than one. Internal linking appears to govern what the training crawler revisits. Page coverage governs what the search index picks up. Conventional search performance still governs what gets pulled in at the moment a buyer asks a question.

How far does one Japanese comparison site’s data travel?

The study covers one site, in one market, in one category, and LANY publishes the limits plainly. Its correlation figures come from a single month, October 2025, when GPTBot and OAI-SearchBot activity ran at roughly 289% and 356% of their monthly averages. Outliers were stripped using the interquartile range method, removing 16.6% of GPTBot records. The GPTBot correlations were calculated on the 599 records out of 744,980 that carried a search volume above zero; across the full dataset, the internal-link correlation falls to 0.183. LANY states that correlation does not establish causation and that crawler behaviour may change whenever OpenAI changes its specifications.

Both parties also have something to sell. LANY sells optimisation consultancy for large language models, and Innovation & Co. sells listings on IT Trend, with the report closing on a link to its enquiry form. Western evidence points the same way and carries the same caveat: G2 has published an analysis of roughly 35,000 ChatGPT citation URLs captured in the tracking tool Profound, finding that review platforms take a larger share of citations as buyer intent strengthens. G2 agreed in January 2026 to buy Capterra, Software Advice and GetApp from Gartner, and says the four properties together account for 84% of citations in the review-platform category.

What the logs do not support is a rush to be listed everywhere. LANY’s own conclusion is that a listing on its own earns nothing, because what the crawlers fetched was specific material: the text of individual reviews, product pricing, feature detail. That is the same material most B2B vendors treat as an administrative chore on directories they have already paid for.

The report does not say whether the analysis will be repeated on other sites, or against the crawlers run by Google, Anthropic and Perplexity, which is what would show whether the pattern belongs to comparison sites in general or to IT Trend in particular.

Subscribe to our newsletter

By subscribing you agree to with our Privacy Policy
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Share article

Recommended Reading