In July 2026 the practice acquired something it had been sold without: a systematic review of its own evidence base. The review covered 45 studies and concluded that no technique it examined shows a stable, longitudinal, cross-platform causal effect. That finding does not make the category worthless. It makes most of what is currently sold under its name unsupported, and it points clearly at the few things that hold.
What generative engine optimisation claims to do
Generative engine optimisation is a set of content and technical practices intended to raise the probability that an AI answer engine cites a given page. In its strong form, vendors claim it works like search engine optimisation did in 2010: apply the recipe, gain the visibility. That framing is the source of most of the confusion, because the two systems fail in different places.
An answer engine does two separable things. It retrieves a candidate set of documents, then it composes an answer from that set and decides which documents to credit. Almost all published GEO research operates on the second step, with the candidate set already fixed. The techniques are then reported as though they raised visibility, when what they raised was the odds of being quoted from a shortlist the technique had no part in reaching.
For the question of whether a brand should fund GEO or traditional search optimisation, and how the two disciplines divide, see our analysis of GEO versus SEO and what the data says a brand should fund now. This article asks a narrower and more awkward question: of the techniques sold under the GEO banner, which ones survive testing?
The 40% figure everyone quotes comes from one experiment
Almost every GEO pitch deck in circulation carries a version of the same number: optimisation lifts visibility by around 40%. The July 2026 survey traced it. The figure comes from a single 2024 benchmark configuration, in which one visibility metric rose from 19.3 to 27.2 when quotations were added to a document. That is a relative gain of roughly 41% on one metric, in one experimental setup, for one technique.
The survey's evidence hierarchy lists the claim that GEO increases visibility by 40% as rejected as a general claim, with the rationale that the figure is a relative maximum on one metric under a specific configuration. It is not a finding about brands, categories or platforms. It is a finding about what happens to a scoring function when a quotation is pasted into a document that a system has already decided to look at.
This is the pattern that repeats throughout the literature. The mechanism is real. The generalisation is not. Clinical research has a word for this gap: a compound that works in vitro has not been shown to work in patients, and the history of medicine is largely the history of that distinction being learned expensively. GEO's foundational results are in vitro. The trials came later, and they came back weaker.
At this stage, claims about GEO return on investment clearly outstrip the academic evidence.
That sentence is the survey's own, and it is the most honest line published about the category this year.
Forty-five studies found no technique with a stable effect
The survey's central conclusion is worth quoting in full, because it is routinely softened in summary: already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behaviour.
Read that carefully. Already-retrieved content. The evidence supports influencing what happens to a document after a system has found it. It does not support the claim that these techniques get a document found.
The survey grades individual techniques by the strength of the evidence behind them. Exhibit 1 reproduces those gradings and the conditions the paper attaches to each. The conditions are not decoration. In several rows they reverse the practical advice.
| Technique | Evidence grade | Condition or caution |
|---|---|---|
| Query–document relevance | Strong in controlled settings | Primary determinant of the first citation; addresses genuine information needs |
| Position in the context window | Strong | Conditional on the document already being retrieved |
| Extractable evidence (quotations, statistics) | Moderate to strong | Verifiable figures, definitions and comparisons; numbers are the unit of citation |
| Recency, prices and dates | Moderate | Useful for time-sensitive or commercial queries, not universal |
| Document structure | Moderate and heterogeneous | Test headings, tables and fields without assuming the direction of effect |
| Fluency and simplification | Weak to moderate | Domain and engine specific; optimise for the user first |
| Authoritative tone | Weak and unstable | May conflict with credibility; do not conflate confidence with evidence |
| Keyword stuffing | Null or negative | Measurably worse than baseline across multiple benchmarks |
| Formatting alone, or fixed recipes | Poor generalisation | Occasional local gains; requires matched, multi-engine testing |
The row that matters most is the second one. Position in the context window is one of the strongest effects in the literature, and it is entirely conditional on the document already having been retrieved. Every hour spent on it is an hour spent on a step that only exists if a prior step has already succeeded.
The last two rows are the ones the industry does not quote. Fixed recipes generalise poorly, which is the paper's assessment of the GEO checklist as a product. And authoritative tone, the register that half the category's content advice pushes towards, grades weak and unstable, with an explicit caution not to confuse sounding credible with being credible.
Optimising the page can make it harder to find
The most uncomfortable result in the survey comes from a 2026 study that did what almost no earlier work had done: it modelled the full pipeline, including retrieval and reranking, rather than starting from a fixed candidate set.
As reported in the survey, that study covered 171,003 documents and 2,700 queries. Optimising only the body copy of a document reduced its average presence in the top 20 by about 9%, reduced its top 10 presence after reranking by 16%, and reduced final citation by 6%.
The techniques did not fail to help. They actively hurt. Writing that reads as optimised to a retrieval system looks less like the plain, relevant prose that retrieval favours, and the document loses more at the retrieval stage than it gains at the citation stage.
A second study reported in the survey points the same way. Across roughly 1,900 queries and 16,360 documents, only three of 54 method and domain combinations were significantly positive, and none of them were in question answering. Question answering is, of course, what an answer engine does.
Both of these figures come from the survey's summary of the underlying papers rather than from reading those papers directly, and should be weighted accordingly. But the direction is consistent, and it is the opposite of the direction the category sells.
Almost nothing reads your llms.txt
The llms.txt file, a plain-text summary of a site placed at its root for AI systems to read, has become a standard recommendation. Ahrefs tested whether anything reads it, using server log data from 137,210 domains in May 2026.
Of those domains, 28% published an llms.txt file. Of the files published, 97% received zero requests during the month. Among the small share that received any traffic at all, only 19.5% of requests came from named AI tools, and bots specifically identified as AI retrieval agents accounted for around 1% of requests.
The most telling line in the study is the quietest. No requests arrived for llms.txt files that did not exist. Nothing goes looking. The file is read only when something has already been directed to the site by other means, which makes it a convenience for systems that have already found you rather than a mechanism for being found.
Ahrefs note that their sample skews technical and SEO-aware, so 28% adoption should be treated as an upper bound on the wider web. The zero-request finding is unaffected by that skew.
Ranking in the top ten no longer buys the citation
The assumption underneath a great deal of GEO advice is that AI answers draw from the top of the conventional search results, so conventional ranking remains the lever. That relationship has weakened sharply.
Ahrefs analysed the result pages for 863,000 keywords and 4 million URLs cited in Google AI Overviews, in research published in March 2026. Only 37.9% of cited URLs also appeared in the conventional top ten. In July 2025 that overlap had been around 76%.
The driver is query fan-out. Rather than answering the question as asked, the engine decomposes it into sub-queries, runs them separately, and preferentially cites pages that recur across several of them. A page that ranks eleventh for the headline query but appears across four sub-queries beats a page that ranks first for one.
This changes what a page has to be. Being the best answer to a single question is worth less than being a reliable answer to a cluster of adjacent ones. For how to track this in practice, see how to measure brand visibility in AI search.
In professional services, the citations are not on your website
Everything above concerns what happens to a page. The most consequential finding concerns whether the page matters at all.
DeltaV Digital collected 21,075 AI engine responses daily between 14 April and 13 July 2026 across ChatGPT, Perplexity, Gemini, AI Overviews and AI Mode, analysing 25,337 citations across eight brands. For the business technology services brand in that set, third-party listicles earned 61% of citations. The brand's own domain earned 0.0%. The single most-cited domain for that brand was LinkedIn.
| Source of citation | Share |
|---|---|
| Third-party listicles and rankings | 61% |
| The brand's own domain | 0.0% |
| Most-cited single domain | LinkedIn (736 citations) |
The same dataset carries the resolution, and it is not to abandon your website. Own domains that do get retrieved convert better than any other domain type, at 1.51 citations per retrieval once a higher-education outlier is excluded, against 1.43 for editorial domains. The honest reading of that margin is that it is thin. The constraint is not that owned content lacks credibility once an engine reaches it. The constraint is reaching it.
Which returns to the survey's finding by a different route. Retrieval is the binding step. Rewriting a page you already own does not solve retrieval. Being described accurately on the surfaces that do get retrieved might.
This is the strategic case beneath the tactical work of making sure AI describes your brand correctly, and it qualifies rather than contradicts the practical routes set out in how to get your brand mentioned by ChatGPT.
What the evidence does support
Four things, and they are unglamorous.
Be relevant, in the ordinary sense. Query-document relevance grades strong and is the primary determinant of the first citation. It is also the least fashionable recommendation in the category, because it cannot be productised.
Publish extractable specifics. Quotations, statistics, prices, dates and named figures grade moderate to strong. Numbers are the unit of citation. A page that contains a checkable figure gives an engine something to credit; a page of well-composed generality gives it nothing to hold.
Get described accurately by other people. In professional services the citation surface is third-party. Directories, listicles, professional profiles and editorial coverage carry the description of a brand into the answer, and correcting a wrong description at source is worth more than another page on the owned site.
Answer clusters, not single questions. Fan-out rewards pages that recur across related sub-queries. Depth on a genuine topic beats precision on a single keyword. On the mechanism by which engines select brands in the first place, see how AI engines decide which brands to recommend.
What the evidence does not support is the recipe. Formatting alone generalises poorly. Authoritative tone is weak and unstable. Keyword stuffing is measurably worse than doing nothing. And llms.txt, on current log evidence, is read by almost nothing.
One further piece of proportion. Similarweb data reported by Search Engine Journal in July 2026 found that only 6.8% of US ChatGPT answers included a link to an external source as of May 2026, up from roughly 1% a year earlier. Search still draws around 3.3 billion average monthly unique visitors worldwide against 655 million for AI chatbots, and 461 million of ChatGPT's 494 million users also used Google in the same window. AI answers are a real and rapidly growing surface. They are not yet a replacement for the one underneath them, and a brand being told to reallocate wholesale is being oversold.
What this means for a brand in Southeast Asia
Two regional points follow, and neither is currently being made in this market.
The first is about who is doing the advising. A Vantage review of generative engine optimisation queries in August 2026, using a general web and answer index as a directional proxy rather than querying ChatGPT, Perplexity or Gemini directly, returned only agencies selling the service, all of them operating in other markets, and no sceptical source of any kind. A Southeast Asian brand evaluating GEO is therefore being briefed largely by vendors, with little regional evidence available to it. That is not a conspiracy. It is what a thin corpus looks like, and it is a reason to discount the confidence of what you are being told rather than the substance.
The second is sharper. If third-party surfaces carry the citation, then the density of those surfaces determines how hard a wrong description is to shift. In Southeast Asia the corpus is thinner than in the United States or Europe. Fewer directories, fewer regional listicles, less editorial coverage of the professional services categories. A thin corpus propagates fast and resists correction, because there are not enough competing descriptions to dilute a wrong one. A single inaccurate profile can set how an entire market's answer engines describe a firm, and it can hold that position for a long time.
The practical consequence is an inversion of the usual regional advice. In a thick corpus, publishing more is a reasonable strategy because volume eventually shifts the average. In a thin one, corroboration beats publication. Auditing and correcting how a brand is described on the handful of surfaces that actually get retrieved is higher-value work than adding pages to a site that may not be retrieved at all.
There is an obvious tension in an article making that argument on its own domain. It is worth naming rather than hiding. Owned pages that do get retrieved convert at the highest rate of any domain type, so the case is not that publishing is pointless. It is that publishing without corroboration is a bet on the step the evidence says is the constraint. Both pieces of work run together, and a consultancy that told you otherwise while publishing constantly would be worth less of your attention, not more.
Vantage is a Singapore brand consultancy specialising in brand research, strategy, and identity design for ambitious organisations across Southeast Asia, with particular depth in healthcare, finance, government, and cultural-institution branding.