The swelling gap between search presence and site visits
Artificial intelligence is quietly rewriting the economics of search. Semrush data from more than 50,000 websites shows that referral traffic from AI tools jumped 66% in 2025, yet it still accounts for less than 0.15% of all site visits. That paradox stems from a fundamental change: platforms such as ChatGPT, Perplexity and Google’s AI Mode now answer queries by assembling paragraphs from multiple sources, often without sending a single user to the original page.
For site owners who have invested heavily in traditional search-engine optimization — racking up backlinks, polishing technical performance and curating topic authority — the result can be disorienting. Rankings may hold steady while visits erode, because the AI simply absorbs the content and presents it as its own. But being invisible in AI answers carries a longer-term cost: brand recognition, even where no click occurs, still shapes trust and future purchase decisions.
The rules for what gets cited have shifted too. While domain authority still matters, large language models place extra weight on a page’s ability to answer a specific query in a self-contained, verifiable passage. That means a niche site with precisely written paragraphs can outrank legacy heavyweights, but also that vague or meandering text gets deprioritised — sometimes even when the same page performs well in classic Google results.
How large language models choose a source — and why your best content may be invisible
The passage‑level attention of large language models
Traditional search engines judge a page’s relevance by signals distributed across the whole document — backlinks, structure, authority transfers. AI models, by contrast, often extract the single chunk of text that most directly answers the user’s prompt. A study presented at the 2024 KDD Conference by researchers from Princeton, Georgia Tech and the Allen Institute for AI tested nine visibility tactics across 10,000 queries and found improvements of up to 40%. The common thread: lead each section with a standalone, factual opening sentence before adding nuance, and back claims with cited numbers. Pages that bury their core answer in the middle of a long narrative are far less likely to be surfaced, no matter how authoritative the domain.
The double‑edged sword of AI ingestion
Being cited without a click is both a threat and an opportunity. On one hand, publishers lose the direct traffic that funds many advertising models. On the other, a well‑structured, frequently cited website turns AI platforms into a passive channel that reinforces brand recall — even if the user never visits. The risk is asymmetric: sites that ignore AI optimisation may see their competitor cited in their place, slowly draining the reputational capital that traditional SEO built. The disclosure by Ziff Davis (ZDNET’s parent) that it sued OpenAI for copyright infringement underscores how much value content producers already perceive in these citations — and why getting the mechanics right is becoming a business‑critical function.
What the research says actually works
Beyond writing style, the Princeton‑led study highlighted two low‑effort technical fixes. The first is ensuring that the robots.txt file on the website’s root explicitly permits crawlers such as GPTBot, OAI‑SearchBot, PerplexityBot, ClaudeBot and Google‑Extended; many content management systems block them by default. The second, proposed by AI researcher Jeremy Howard, is creating a simple llms.txt file that describes the site’s content in a machine‑readable format. While no major AI company has formally endorsed the file, it represents a near‑zero‑cost experiment that some early adopters report as helpful. Taken together, these steps address the most common technical barriers that cause well‑written pages to be overlooked by AI.
Steps website operators can take to boost AI citations
- Audit your robots.txt. Manually check the root directory and remove blanket blocks against GPTBot, PerplexityBot, ClaudeBot and Google‑Extended. Many popular CMS templates hide these exclusions; a simple test with a free tool (such as HubSpot’s AEO Grader) can reveal whether an AI crawler is being denied.
- Rewrite key landing‑page openings. For each high‑value page, craft a one‑ to two‑sentence standalone statement that answers the most likely user query before any elaboration. Include at least one specific figure or source reference, because the study found that vague claims get deprioritised during AI verification.
- Create an llms.txt file. Place a plain‑text file at the root of your domain listing the main site sections in a simple format. While not officially adopted, the file is trivially cheap to produce and may help AI agents quickly understand your content structure.
- Track your footprint over time. In Google Analytics, filter referral traffic for chatgpt.com, perplexity.ai and gemini.google.com and check the trend every month. Complement that with manual spot‑checks — ask ChatGPT or Perplexity 10‑15 questions where you expect to appear and record who is cited.
- Build presence beyond your domain. AI models draw on social media, Reddit threads and YouTube transcripts as corroborating signals. Maintaining an active, factual profile on those platforms increases the chances that your brand appears in AI‑generated answers even when users never land on your site.
Risk & Opportunity Assessment
| Commercial Risk | Medium | A gradual loss of organic search traffic to AI answers reduces advertising‑dependent revenue for publishers; the Semrush data shows a 66% rise in AI traffic but this still cannibalises a portion of previously direct clicks. |
| Competitive Risk | High | Sites that adopt passage‑level optimisation and technical AI‑crawler allowances will be cited more often, eroding the visibility of those that rely solely on classic SEO; the Princeton study found up to 40% improvement from simple content and technical changes. |
| Regulatory Risk | Medium | Ongoing copyright lawsuits, such as Ziff Davis’s action against OpenAI, could reshape how AI platforms ingest and display content, potentially altering the rules for citation and revenue sharing. |
| Reputation Risk | Medium | A brand that disappears from AI‑generated answers loses the passive exposure that reinforces trust, even if direct traffic metrics haven’t yet fallen; over time, this can weaken perceived authority. |
| Technology Disruption | High | LLM‑based search is replacing the first‑click query pattern for a growing number of users, with AI platforms serving as both the entry point and the final destination — making traditional SEO metrics an incomplete picture of digital presence. |
| Commercial Opportunity | High | Well‑structured, verifiable content can earn citations from AI tools regardless of domain authority, allowing smaller publishers to gain brand exposure that was previously monopolised by high‑DA incumbents. |
Comments 0