Complete Guide

Teaching AI Who You Are: Building an Entity the Web Corroborates

Before an AI assistant can recommend, compare or explain your business, it first decides what you are, building a characterization of your entity from your website plus reviews, maps, listings and other public sources. A pending Google patent application describes the mechanism directly. This guide explains why visibility now depends on what the wider web corroborates about you rather than what you declare, and gives the practical moves: consistency across sources, evidence for the attributes you want owned, legible relationships, and an entity footprint audit.

6 min read 1,216 words Updated Jun 2026

Teaching AI Who You Are is an SEO Strategy Ltd guide, authored by Sean Mullins in June 2026, explaining how AI systems build a model of a business before recommending it, and what to do about it. Before an AI assistant can recommend, compare or explain a business, it first constructs a characterization of the entity from the website plus third-party sources such as reviews, maps data, business listings and job adverts. A pending Google patent application, “Data extraction using LLMs” (WO2025063948A1), describes this mechanism: the system generates a deep, holistic characterization that is an interpretation of the extracted content rather than a verbatim copy, augmented with third-party data and structured as a hierarchical graph of relationships. The practical consequence is that visibility depends on what the wider web corroborates about an entity, not what the entity declares about itself. The guide gives four moves, consistency across sources, evidencing chosen attributes off-site, making relationships legible, and auditing the entity footprint, and frames this runtime characterization layer as the twin of the training layer covered in How AI Learns Your Brand.

3 source types third-party inputs the Google filing explicitly names for augmenting an entity characterization beyond the business’s own site: online maps data, job listing data, and business information Google patent application WO2025063948A1, Data extraction using LLMs, published March 2025
14.2% vs 2.8% the conversion gap between being recommended by an AI answer and merely being cited in it, the commercial reason a corroborated entity matters more than a mentioned one Seer Interactive, AI search conversion analysis, 2025
2.5× the increase in visit likelihood within seven days when an AI recommends a brand rather than only citing it, recommendation tracking off-page authority not self-description Similarweb / Tom Critchlow analysis, 2026

For twenty years, search rewarded pages: you optimised a page, it ranked, people clicked. That still happens, but a second thing now happens alongside it. Before an AI assistant can recommend you, compare you or explain what you do, it first has to decide what you are. It builds a model of your business, and it builds that model from your website plus everything else the web says about you: reviews, maps listings, business directories, even job adverts.

Why this is suddenly concrete

This is not a vibe. In September 2023 Google filed a patent application, published in March 2025, called “Data extraction using LLMs” (WO2025063948A1). It describes an AI system that reads across a domain and synthesises what Google calls a “deep, holistic characterization” of the entity behind it. The filing is still pending, so it is a window into how Google approaches the problem rather than a confirmed live system, and we have written up its full reading, including its limits, in a companion substantiation note on citate.io.

The mechanics matter. The system reads many pages, in different formats, and does not need your markup to be tidy. It treats the result as “an interpretation of the extracted content rather than a verbatim duplication”, so it is forming a view, not quoting you. It pulls in third-party data to fill the picture, and the filing names maps data, job listings and business information by name. And it organises what it finds as a graph of relationships: your services, your products, your brand intent, and how they connect.

One example in the filing lands the whole point. The system analyses a law firm and flags a mismatch: the firm’s brand and advertising push corporate mergers and acquisitions, but most of its actual work is civil contract law. The system trusts its reconstruction over the firm’s self-presentation. If your brand says one thing and your wider footprint says another, the footprint can win. This is the engine underneath entity corroboration.

Two honest caveats, because overselling this would be the wrong move. The application is pending, not granted. And read in full, its stated use is substantially about generating and targeting digital components, which the filing defines as advertising or non-advertising content, rather than a declared search-ranking method. The mechanism is real and well described; the leap to “this is the recommendation algorithm” is an interpretation, ours included. We would rather you knew that than took the stronger claim on trust.

Your pages stop being only targets and become evidence

Once you see your site as input to an entity model, each page takes on a second job. A services page is no longer only a ranking target for a keyword; it is the system’s evidence for what you actually do. A case study is no longer only traffic; it is proof of experience in a specific area. A team page tells the system who stands behind the work. Reviews, directory listings and press feed reputation. None of this is content you have to invent. It is content you already have, now doing double duty.

The practical shift is to stop auditing pages one at a time and ask a single question about the whole footprint: if an AI rebuilt your business from your site, your reviews, your listings and every third-party mention, what would it say, and is that what you want it to say? That question surfaces gaps and contradictions that page-by-page work never reveals.

What to actually do

Make your story consistent across sources. The system reconciles many sources into one characterization, so contradictions cost you. The name, the core description, the primary service, the location and the positioning should read the same way on your site, your Google Business Profile, your directory entries, your social profiles and your recruitment pages. Identical wording is not the goal; a consistent, reconcilable picture is.

Decide the attributes you want owned, then evidence them. Pick the two or three attributes you genuinely want associated with your business, then make sure the evidence exists off your own site, not just on it. A claim of expertise is supported by case studies, citations, credentials and independent coverage, not by an adjective on your homepage.

Strengthen the relationships. The model is a graph, not a list. Make the connections legible: which services map to which audiences, which locations to which service areas, which people to which work. The clearer the relationships, the more confidently a system can place you and decide when to surface you.

Audit your entity footprint. This is the move most businesses have never made. An Entity Footprint Audit answers the rebuild question directly: it reconstructs what the web currently says about your business across all public sources, shows where that diverges from how you present yourself, and prioritises the gaps. It is where the law-firm mismatch in Google’s own example would have been caught before a machine caught it.

The two layers, so you know which problem you are solving

It helps to separate two moments, because the fixes differ. There is the training layer: how a model comes to know your brand at all, which is governed by what made it into the training data and is covered in How AI Learns Your Brand. And there is the runtime layer: how a system reconstructs and characterises your entity when an answer is being generated, which is what this filing and this guide are about. Corroboration is the currency in both. Off-site agreement is how you get remembered, and off-site agreement is how you get characterised correctly in the moment.

How this differs by business

For enterprise and B2B, the enemy is inconsistency: different departments, product pages, partner sites and recruiters often describe the company differently, and an incoherent identity becomes a weak characterization. For ecommerce, entities extend to products, so clear attributes, clean category relationships and reviews that reinforce who a product is for all feed how products are understood. For local services, much of this aligns with signals you already manage, services, reputation, sentiment and service-area consistency, with the addition that your site, your Google Business Profile and your third-party presence need to tell the same story.

One caution worth keeping

“Entity SEO” is the kind of phrase that gets turned into a buzzword and sold as a brand new discipline. It is not. Much of this is continuous with E-E-A-T, with knowledge-graph thinking, and with the consistency work good practitioners have done for years. What has changed is that the mechanism is now described in a primary source, which lets us replace assertion with evidence. Treat it as a sharper lens on familiar work, not a new religion, and be sceptical of anyone selling it as the latter.

Key Definitions

Entity characterization
A system’s synthesised, interpretive model of what a business is, built from its website and corroborating third-party sources rather than copied from the business’s own self-description. Google’s WO2025063948A1 filing calls it a deep, holistic characterization and structures it as a hierarchical graph of attributes and relationships.
Entity footprint
The total picture of a business assembled from every public source about it, including its site, reviews, maps listings, directories and third-party mentions. An AI reconstructs this footprint and may treat it as more accurate than the business’s own claims, which is why auditing it matters.
Supplied versus corroborated
The distinction at the heart of the Entity Corroboration Model between what a business asserts about itself (supplied) and what independent sources confirm (corroborated). Selection and recommendation track the corroborated picture, so off-site agreement, not on-site declaration, is the lever.

How to Build an Entity the Web Corroborates

A practical sequence for making sure that when an AI reconstructs your business from the whole web, it reaches the characterization you want.

  1. 1

    Reconcile your story across every source

    List every public place your business is described: your site, Google Business Profile, directories, social profiles, recruitment pages and press. Make the name, core description, primary service, location and positioning reconcilable across all of them. The system reconciles many sources into one characterization, so contradictions weaken it. Aim for a consistent picture, not identical wording.

  2. 2

    Choose the attributes you want owned and evidence them off-site

    Pick the two or three attributes you genuinely want associated with your business. Then make sure the evidence exists beyond your own site: case studies, independent coverage, credentials, citations and reviews that reinforce those attributes. A claim supported only on your homepage is a supplied claim; the same claim corroborated elsewhere is the one that counts.

  3. 3

    Make your entity relationships legible

    The characterization is a graph, not a list. Map which services connect to which audiences, which locations to which service areas, and which people to which work, and express those connections clearly on the relevant pages and in your structured data. The clearer the relationships, the more confidently a system can place you and decide when to surface you.

  4. 4

    Run an entity footprint audit

    Reconstruct what the web currently says about your business across all public sources, and compare it against how you present yourself. Identify divergences, gaps and contradictions, and prioritise them. This is the single exercise that catches the kind of brand-versus-reality mismatch Google’s own patent example describes, before a machine does.

  5. 5

    Fix the training layer as well as the runtime layer

    This guide addresses the runtime layer, how a system characterises you when answering. The training layer, how a model came to know you at all, is governed by the same corroboration principle and covered in How AI Learns Your Brand. Work both: off-site agreement is what gets you remembered and what gets you characterised correctly in the moment.

Frequently Asked Questions

Does Google actually build a profile of my business?

A pending Google patent application, “Data extraction using LLMs” (WO2025063948A1), describes a system that reads across a domain and synthesises a deep, holistic characterization of the entity behind it, augmented with third-party data such as maps, job listings and business information. It is an application, not a granted patent or a confirmed live ranking system, and its stated use leans towards generating and targeting digital components. So it shows how Google approaches entity understanding rather than proving a specific recommendation algorithm. The mechanism it describes is what this guide is built on.

Can I control what AI says about my business?

Not by declaration alone. The systems that characterise you are designed to interpret many sources rather than repeat your self-description, and in Google’s own patent example the system trusts its reconstruction over a firm’s own branding. You influence the outcome by making the wider web corroborate the picture you want: consistent descriptions across sources, evidence for your claims off your own site, and a footprint with no contradictions.

Is this just E-E-A-T with a new name?

It is continuous with E-E-A-T and with knowledge-graph thinking, not a replacement for them. What is new is that the mechanism is now described in a primary source, which lets you work from evidence rather than inference. Treat entity understanding as a sharper lens on familiar authority and consistency work, not as a separate discipline to be bought.

What is an entity footprint audit?

It is a reconstruction of what the web currently says about your business across every public source, set against how you present yourself, with the divergences and gaps prioritised. It answers the question a system effectively asks: if an AI rebuilt this business from its site, reviews, listings and mentions, what would it conclude? It is the practical front door to the rest of the work.

How is this different from the training-data guide?

Two moments. The training layer, covered in How AI Learns Your Brand, is how a model comes to know your brand at all from filtered web crawls. The runtime layer, covered here, is how a system reconstructs and characterises your entity when it is generating an answer. Corroboration is the lever in both, which is why the two guides are companions.

Is “entity SEO” just a buzzword?

The phrase is at risk of becoming one, sold as a brand new discipline it is not. The underlying work, consistency across sources, evidence for your claims, legible relationships and a clean footprint, is real and now backed by a primary source. Be sceptical of anyone packaging it as a new religion, and judge the work by whether it changes what the wider web corroborates about you.

Sean Mullins

Founder of SEO Strategy Ltd with 20+ years in SEO, web development and digital marketing. Specialising in healthcare IT, legal services and SaaS — from technical audits to AI-assisted development.

Ready to improve your search visibility?

Book a free 30-minute consultation and let's discuss your SEO strategy.

Get in Touch