Where the AI Discovery Stack Breaks Down

The AI Discovery Stack is the model I use to explain how AI systems decide which businesses to retrieve, cite and recommend. I still think it is the most useful map of that process I have. But a map that only flatters its author is not worth much, so this piece does the opposite. It documents three places where the Stack, taken too literally, will mislead you, using behaviour observed live on real client pages in June 2026. None of this retires the model. It marks its edges, which is where the honest work is. And underneath all three edges is a single shift: the Stack was built as if AI visibility were deterministic, and the evidence says it is probabilistic.

The Stack, briefly, is a sequence of layers. A page has to be discoverable as an entity before it can be understood, understood before it can be selected, selected before it can be recommended or acted on. The companion Four-Floor walkthrough compresses that into Entity Foundation, Content Extractability, Trust and Selection, and Agentic Execution. The model is clean, sequential and dependency-ordered. Each of those three properties is also where it bends.

Break one: retrieval is not a position, it is a probability

The Stack invites you to think of the bottom floor as a gate you pass or fail. Build the entity foundation correctly, the thinking goes, and you are retrievable. Retrievable is then treated as a settled fact about the page, the way a ranking is a position you hold.

It is not a position. Retrieval for a given query is non-deterministic, and the variance is large enough to break the mental model of a gate.

On 24 June 2026 I ran one informational query against Google AI Mode six times inside forty minutes, pointed at a client page that holds the featured snippet for the same query on standard Google search. The page is built to the CITATE standard, on-topic, and the direct answer in the organic snippet box. By any reading of the Stack it has passed the bottom floor decisively. Yet across those six AI Mode runs it was cited in four and absent in two. Same page, same query, no change to anything on the page, forty minutes.

This is not an anomaly I went looking for. It is the documented behaviour of the platform. SISTRIX, tracking 82,619 prompts over seventeen weeks across six countries, found that Google AI Mode replaces roughly 56% of its cited sources every week, with a small stable core of sources that persist and a much larger carousel that rotates. Profound, analysing 240 million ChatGPT citations, found that 40 to 60% of cited domains change month to month for identical queries. A single citation is a sample of a moving distribution, not a position you hold.

So the honest correction to the Stack is this. The bottom floor does not output a boolean. It outputs odds. A strong page is not retrieved or not retrieved; it is retrieved with some probability that varies run to run. The right question stops being “am I in the answer” and becomes “am I in the stable core or the rotating carousel, and how often do I appear across repeated checks”. Rankings are positions. AI retrieval is a distribution. The model implies a light switch; the data shows a dimmer that flickers.

Break two: AI visibility is not page-centric, it is query-entity-context centric

The second assumption the Stack smuggles in is that the floors are climbed in order, and that the order is about the page. Get the foundation right, then the extractability, then trust. The page works its way up. This is the assumption with the deepest roots, because it is inherited straight from classic SEO, where the page is the unit you optimise and a page has a value you can move.

The evidence says the page is not the unit. A page does not have a single retrievability that you raise or lower. It has many, one per query intent, and they can point in opposite directions at the same moment.

A B2B software client of mine is the clear example. On queries that match the profile the AI systems hold for it, automated PGP encryption and affordable HIPAA-compliant file transfer, the product is returned unprompted, near the top, with its differentiators preserved almost word for word. On generic enterprise queries, top managed-file-transfer platform, top compliance platform, the same product is absent and the enterprise incumbents own the answer. Nothing about the page changed between those queries. It did not climb higher on one and fall lower on the other. The same page was, in the same session, both highly retrievable and barely retrievable, decided entirely by how the query framed the entity.

That is a deeper correction than volatility, and it is worth stating plainly. The Stack models the page. It does not model the query, and it does not model the entity the system holds for you. In the field, those two decide first. Retrieval is conditional on the match between the query intent, the entity the AI systems have built for you from everything they have seen, and the context of the search, and the page is only assessed once that match clears. A page can be flawless on every floor and never enter the building, because the search implies a kind of business the models do not believe you are. You can be a Citation Magnet and Citation Invisible for the same page on the same day, depending only on what was asked. The unit of AI visibility is not the page. It is the query, the entity and the context together. That is a harder idea than “optimise the page”, and it is the one classic SEO thinking is least equipped for.

Break three: citation and recommendation are different gates

The third break is the one most likely to cost a client real money, because it hides inside the word “selected”. The Stack puts selection and recommendation near the top, and the vertical metaphor encourages you to read them as the reward for climbing. Get cited, the implication runs, and the recommendation follows.

It does not. Being the cited source and being the recommended choice are separate gates, and they can diverge inside a single answer.

A local client of mine was the lead citation card in a Google AI Mode answer, named as the source and used for the opening definition and the structured detail the answer was built on. In the same answer, the “top rated providers” block that actually recommended businesses to the user named two competitors instead. My client was the authority the answer leaned on and not the business the answer suggested. The page supplied the trust signal for the topic and lost the selection to firms with more reviews.

That is the gap the Stack’s tidy vertical hides. Citation asks whether your content is good enough to ground the answer. Recommendation asks whether you are the safest entity to put in front of a user, and that turns on signals like review volume and ratings that have little to do with how well your page is built. You can win the floor below and lose the floor above to someone who built a worse page and collected more reviews. Measure the two as different outcomes, because the lever for the second one often sits entirely off the page.

What the Stack still gets right

Having spent three sections on its limits, here is why I still open with it. The Stack is not a predictor of any single AI answer, and it was a mistake to ever let anyone read it as one. It is a diagnostic vocabulary, and as that it holds up.

When a client is invisible for a query they should own, the Stack tells me where to look. Is this a foundation problem, the entity is not corroborated enough to be retrieved at all, or a selection problem, the page is cited but a competitor is recommended? Those two failures look identical from the outside and have opposite fixes, and without the layered model you cannot tell them apart. The three breaks above do not refute that. They refine what the model is for. It is a map of where a problem lives, not a promise about any given search.

The Stack survives all three of these. It just stops pretending to be a crystal ball and goes back to being a map.

What to do differently

Three changes follow directly from the three breaks.

First, measure persistence, not snapshots. One check tells you almost nothing, because the platform is too volatile for a single sample to mean anything. Run the same query repeatedly over days, and record how many consecutive checks you appear in. Persistence is the signal. A single citation is noise dressed as a result.

Second, stop optimising pages and start mapping query-entity matches. Find the queries where your entity is retrieved and the ones where it is not, and look for the profile boundary between them. Do not try to win queries whose intent implies an entity you are not; you will spend effort fighting the model’s entire picture of you. Win the queries that match the entity you actually are, and build the entity outward deliberately if you want the boundary to move. The page is downstream of that match, not upstream of it.

Third, separate citation from recommendation in how you measure and where you act. If you are cited but not recommended, the page is doing its job and the gap is a trust-and-selection problem, often review volume, that no amount of further page work will close. Diagnose which gate you are failing before you decide what to fix.

The throughline is one idea. Version one of AI visibility thinking, the version most of the industry still runs on, treats it as deterministic: optimise the page, earn the position, hold it. The evidence in the observed-outcomes register points the other way. AI visibility is probabilistic, conditional and multi-gated. The Stack is still the best map I have of the terrain. It just describes a landscape that moves, and saying so out loud is the difference between a framework that matures and one that quietly stops being true.

Related topics:

ai-discovery-stack ai-mode ai-visibility citation-drift geo
Sean Mullins

Founder of SEO Strategy Ltd with 20+ years in SEO, web development and digital marketing. Specialising in healthcare IT, legal services and SaaS — from technical audits to AI-assisted development.