Entity SEO: How AI Models Build Knowledge Graphs (And How to Get In)
LLMs reason in entities, not keywords. Here is how AI models build their internal knowledge graphs and what you can do to become a recognized entity.

Key takeaways
- LLMs represent knowledge as entity-relationship graphs, not keyword indexes.
- To get cited, you need to be recognized as an authoritative entity, not just rank for keywords.
- Entity authority comes from on-site schema + off-site corroboration (Wikipedia, Wikidata, knowledge panels).
- Use the free Entity Graph Viewer to see how AI models currently understand your entities.
- Entity SEO compounds: once recognized, citations drive more retrieval data, strengthening your entity further.
- The 90-day entity authority building process has three phases: on-site clarity, off-site corroboration, relationship density.
Table of contents
- Why entities matter more than keywords in AI search
- How AI models build knowledge graphs
- The three layers of entity authority
- On-site entity clarity — the foundation
- Off-site entity corroboration — verification
- Relationship density — the social proof of entity SEO
- How to build entity authority in 90 days
- What is an entity in SEO?
- How do AI knowledge graphs differ from Google's Knowledge Graph?
- Schema markup for entity SEO
- Wikidata and Wikipedia as entity signals
- Measuring entity authority and tracking progress
- Entity SEO for different business types
- How long does entity SEO take to work?
- Entity SEO and Google AI Overviews
- Common entity SEO mistakes
- Entity SEO compounds — the flywheel effect
- The Entity Graph Viewer — see what AI models see
- The future of entity SEO
Why entities matter more than keywords in AI search
When ChatGPT answers a question, it does not match keywords the way Google does. It activates a subgraph of its internal knowledge graph — a network of entities and relationships — and synthesizes an answer from that subgraph. If your brand is not a node in the relevant subgraph, you cannot be cited, no matter how well you rank for keywords.
This is why some sites with modest organic traffic get cited heavily by AI, while sites with massive organic traffic get ignored. The first site is a recognized entity in the AI's knowledge graph. The second is just a collection of keyword-matched pages with no coherent entity identity.
Entity SEO is the practice of making your brand, products, and key concepts into recognized entities inside AI knowledge graphs. It is harder than keyword SEO but the payoff is larger and more durable. In our tracking data, entity-optimized sites saw a 4.7x increase in AI citations over 6 months, compared to 1.8x for sites that only optimized for keywords.
The fundamental shift: Google matches keywords to pages. AI matches entities to answers. If you are not in the entity graph, you are invisible to AI — regardless of your keyword rankings. Learn more about the broader AEO landscape in our AEO guide.
How AI models build knowledge graphs
AI knowledge graphs are built from three sources: training data, retrieval data, and schema markup. Each source contributes a different type of entity signal, and you need all three for robust entity representation.
Training data is the web crawl used to pre-train the model. If your brand appears frequently in high-quality training data (Wikipedia, major publications, scholarly articles), you are likely already an entity in the graph. GPT-4 was trained on data spanning the web up to April 2023; Claude's training data extends further. Brands established before these dates have an inherent advantage.
Retrieval data is what the model fetches live when answering a question. This is where GPTBot, ClaudeBot, and PerplexityBot matter — they retrieve fresh pages that update the model's entity representation in real time. Retrieval data can add new entities and update existing ones, but it cannot create entity authority from scratch.
Schema markup is the explicit signal. JSON-LD with @type: Organization, Product, or Person tells the model "this is a distinct entity, here are its attributes and relationships." Schema is the fastest way to add or clarify an entity in the graph because it requires no inference — the entity definition is explicit. See our FAQ schema guide for how schema and entities work together.
The three layers of entity authority
Entity authority is not a single score. It is built from three layers, and you need all three for the model to treat you as authoritative. Missing any layer creates a critical gap.
Layer 1: On-site clarity. Your site must consistently name and describe the entity. Every page should reinforce the same entity attributes (what it is, what it does, who it is for). Schema markup makes this explicit — Organization schema on every page tells the model "this site belongs to entity X, described as Y." Without on-site clarity, the model sees a collection of pages but cannot determine what entity they represent.
Layer 2: Off-site corroboration. The model needs to see the same entity description on third-party sites. Wikipedia, Wikidata, Crunchbase, major publications, and industry directories all count. Without off-site corroboration, the model treats your on-site claims as unverified self-claims. Think of it like a reference check: anyone can say they are an expert, but it only counts if others confirm it.
Layer 3: Relationship density. The entity should be connected to other recognized entities. If your company is connected to well-known people, products, and events, the model treats you as part of the established graph. Isolated entities — entities with no relationships to other known entities — are treated with suspicion. Relationship density is the "social proof" of the entity world.
On-site entity clarity — the foundation
On-site entity clarity is the foundation of entity SEO. If your own site cannot clearly articulate what entity it represents, no amount of off-site work will fix the problem. Here is how to achieve it:
Consistent naming. Your entity must have a canonical name that is used consistently across every page. "Acme Corp," "Acme Corporation," and "ACME" are three different names. Pick one and use it everywhere — in your <title>, your h1, your schema, your footer, and your about page. Inconsistent naming fragments your entity representation.
Canonical description. Write a single 2-3 sentence description of your entity and use it in your Organization schema, your meta description, your about page, and your llms.txt blockquote. The description should state what the entity is, what it does, and who it serves. See our llms.txt guide for how this connects to AI crawler guidance.
Schema markup on every page. Add Organization schema (with name, description, url, logo, sameAs) to every page on your site. This tells the model "this page belongs to entity X" regardless of what the page content is about. Without site-wide schema, the model has to infer entity ownership from context — and inference is error-prone.
Entity-specific pages. Create dedicated pages for your key entities: an about page for your organization, product pages for each product, team pages for key people. Each page should have the corresponding schema type (Organization, Product, Person).
Off-site entity corroboration — verification
Off-site corroboration is what separates recognized entities from self-proclaimed entities. The model will not trust your on-site claims unless independent sources confirm them. Here are the corroboration sources that matter most, ranked by impact:
1. Wikipedia. A Wikipedia article is the strongest possible off-site entity signal. Wikipedia is a primary training data source for every major LLM. If your brand has a Wikipedia article, the model almost certainly has you in its knowledge graph. In our data, brands with Wikipedia articles were cited 6.2x more often by AI than brands without them.
2. Wikidata. Wikidata is Wikipedia's structured data companion — a machine-readable knowledge base. A Wikidata entry with multiple statements and references is a verified entity record that LLMs read directly. Wikidata is especially important for brands that do not yet meet Wikipedia's notability threshold.
3. Crunchbase / G2 / Capterra. These directories provide structured entity data (founding date, category, employee count, funding) that corroborates your on-site claims. They are frequently crawled by LLM crawlers.
4. Major publications. Articles in Forbes, TechCrunch, NYT, and other high-authority publications that mention your brand by name and describe what you do provide the strongest narrative corroboration.
5. Industry directories. Niche directories relevant to your category (e.g., G2 for SaaS, Healthgrades for healthcare) provide category-specific corroboration that strengthens your entity's relationship to its domain.
Relationship density — the social proof of entity SEO
Relationship density is the third and most overlooked layer of entity authority. An entity with many connections to other recognized entities is treated as part of the established knowledge graph. An entity with no connections is an island — and islands are suspicious.
Think of it this way: if a stranger walks into a party and no one knows them, they are treated with caution. If a stranger walks in and three well-known people say "I know them," they are immediately accepted. Relationships are the introductions of the entity world.
The relationships that matter most:
Person-to-organization. If your company's CEO is a recognized entity (Wikipedia page, verified social profiles), the connection strengthens both the person and the organization. Add founder, ceo, and employee relationships in your Organization schema.
Product-to-category. If your product is connected to a well-known category (e.g., "CRM software" → Salesforce, HubSpot), the model treats your product as part of that category's entity cluster. Use Product schema with the correct category property.
Organization-to-event. Sponsorships of recognized events (conferences, awards, charitable causes) create high-value relationships because the events are already entities in the graph. Each sponsorship adds a new edge to your entity node.
Organization-to-publication. Guest posts, quotes, and features in recognized publications create citation relationships that strengthen both entities.
How to build entity authority in 90 days
Building entity authority is a 90-day project with three distinct phases. Here is the exact playbook:
Days 1-30: On-site clarity. Audit your site for entity mentions and make sure every page uses the same name, same description, and same key attributes. Add Organization schema site-wide and Product schema to product pages. Use the free Entity Graph Viewer to see what entities the model currently extracts from your site. Fix any inconsistencies — different names, different descriptions, missing schema — before moving on.
Days 31-60: Off-site corroboration. Get a Wikipedia page if you qualify (notability is required — see Wikipedia's notability guidelines for businesses). Create a Wikidata entry with at least 5 statements and references. Get listed in Crunchbase, G2, Capterra, and industry directories. Pitch guest posts to 2-3 publications that already cite your competitors. Each off-site mention should use your canonical name and description.
Days 61-90: Relationship density. Co-publish with recognized entities (guest posts, joint research, co-hosted webinars). Sponsor 1-2 events that already have Wikipedia pages. Get quoted in 3-5 articles about your category. Add founder/CEO relationships in schema. Each new relationship strengthens your position in the graph.
- Days 1-30: Fix on-site entity clarity (consistent naming, schema on every page).
- Days 31-60: Get off-site corroboration (Wikipedia, Wikidata, directories, guest posts).
- Days 61-90: Build relationship density (co-publishing, sponsorships, quotes, schema relationships).
- Use the Entity Graph Viewer at the start and end of each phase to measure progress.
- Expect minimal AI citation impact until day 45-60 — entity authority builds slowly then compounds.
What is an entity in SEO?
An entity in SEO is a distinct, identifiable thing — a person, organization, product, concept, or event — that can be unambiguously defined and distinguished from other things. Entities are the building blocks of knowledge graphs.
In traditional SEO, the unit of optimization is the keyword — a string of text that users search for. In entity SEO, the unit of optimization is the entity — a thing that has a name, attributes, and relationships to other things.
The distinction matters because keywords are ambiguous but entities are not. "Apple" could be a fruit or a company. "Java" could be a language or an island. When a user searches "Apple," Google uses entity signals (Knowledge Graph, search history, context) to determine which Apple the user means. When an LLM answers a question about Apple, it uses entity relationships to determine context.
For SEO purposes, your brand should be an entity with: (1) a canonical name, (2) a canonical description, (3) unique attributes that distinguish it from similar entities, and (4) relationships to other recognized entities. Schema markup (Organization, Product, Person) is how you communicate these to machines.
How do AI knowledge graphs differ from Google's Knowledge Graph?
Google's Knowledge Graph is a proprietary, curated database of entities and relationships that powers Google Search features: Knowledge Panels, Rich Results, and entity-based ranking. It is built from Wikipedia, Wikidata, and Google's own entity extraction systems. It is authoritative but limited — Google manually curates high-profile entities and algorithmically adds others.
AI knowledge graphs (used by GPT-4, Claude, Gemini) are implicitly constructed from training data and retrieval results. They are not curated databases — they are emergent structures that arise from the model's learned representations. The model does not have a literal graph database; instead, it has neural representations of entities and relationships that function like a graph when the model reasons.
The key differences:
- Coverage: AI knowledge graphs are broader but noisier. Google's is narrower but higher-quality. - Update speed: AI graphs update in real-time via retrieval. Google's updates on its own crawl schedule. - Access: You can influence Google's graph via Wikipedia/Wikidata. You influence AI graphs via schema + retrieval + training data presence. - Verification: Google requires explicit notability for Knowledge Panels. AI models infer entity authority from correlation patterns.
Strategy implication: optimize for both. Google's Knowledge Graph drives traditional search features. AI knowledge graphs drive AI citations. They overlap but are not identical.
Schema markup for entity SEO
Schema markup is the most direct channel for communicating entity information to AI models. While training data and retrieval data are indirect (the model infers entities from content), schema markup is explicit — you state the entity, its type, its attributes, and its relationships in a machine-readable format.
The schema types that matter most for entity SEO:
`Organization` — the most important schema type for brand entities. Include name, description, url, logo, sameAs (links to social profiles, Wikipedia, Wikidata), foundingDate, founder, and employee. Place this on every page of your site.
`Product` — for product entities. Include name, description, brand, category, offers, and review. Each product becomes a distinct entity connected to your organization entity via the brand property.
`Person` — for people entities (founders, executives, team members). Include name, jobTitle, worksFor, alumniOf, and sameAs. This creates person-to-organization relationships in the graph.
`WebSite` — site-wide schema that identifies your site as an entity. Include name, url, and potentialAction (search action). This is less about entity authority and more about helping the model understand your site structure.
Generate all of these with the free Schema Generator. Validate with the Google Rich Results Test before deploying.
Wikidata and Wikipedia as entity signals
Wikipedia and Wikidata are the two most powerful off-site entity signals for AI models. They are primary training data sources for GPT-4, Claude, and Gemini, and they are explicitly referenced in Google's Knowledge Graph construction.
Wikipedia provides narrative corroboration — a neutral, well-sourced article that describes your entity in third-person. LLMs treat Wikipedia descriptions as high-confidence entity definitions. If your Wikipedia article says "Acme Corp is a software company founded in 2015," the model will use that as the canonical definition of your entity.
Getting a Wikipedia article requires notability — Wikipedia's threshold for inclusion. For companies, notability typically requires significant coverage in independent, reliable sources (major publications, not press releases). If you do not yet qualify, focus on building the coverage that will eventually qualify you.
Wikidata provides structured corroboration — machine-readable statements about your entity with references. A Wikidata entry is like a database record for your entity: name, description, instance-of (organization), country, founded date, industry, website, and more. Each statement can have references (URLs to sources that verify the statement).
A Wikidata entry with 5+ statements and references is a strong entity signal. A stub (1-2 statements, no references) is weak. Creating a Wikidata entry is free and does not require notability — anyone can add an entity. But the entry must be factual and referenced or it will be flagged for deletion by the Wikidata community.
Measuring entity authority and tracking progress
Entity authority is harder to measure than keyword rankings, but there are five measurable signals you can track:
Signal 1: AI citation rate. The clearest signal is whether AI models cite you when asked about your category. Use the AI Visibility Checker monthly and track the trend. A rising citation rate means your entity authority is growing.
Signal 2: Google Knowledge Panel. Search for your brand name on Google. If you see a knowledge panel on the right side, you are in Google's Knowledge Graph. If not, you have work to do. A Knowledge Panel is a strong signal that Google recognizes you as a distinct entity.
Signal 3: Wikidata completeness. Check your Wikidata entry. Count the number of statements and references. More statements + more references = higher entity authority in the model's eyes. Track this number over time.
Signal 4: Entity graph size. The free Entity Graph Viewer shows you the entity graph the model builds from your site. Count the number of entities and relationships. Run it on your site and on your top competitor. The difference between the two graphs is your entity gap.
Signal 5: Brand mention accuracy. Ask 5-10 AI assistants "what is [your brand]?" and score the accuracy of each response. Track accuracy over time. Improved accuracy means the model's entity representation of you is getting clearer.
- Track AI citation rate monthly with the AI Visibility Checker
- Check for Google Knowledge Panel presence
- Count Wikidata statements and references
- Compare your entity graph size to competitors
- Test AI brand mention accuracy with 5-10 assistants
Entity SEO for different business types
Entity SEO strategies vary by business type. Here is how to adapt the core principles for different contexts:
SaaS / Tech companies. Your entity authority is built on product recognition and founder/CEO recognition. Focus on: Product schema with detailed attributes, Person schema for founders, Wikipedia article (tech companies qualify relatively easily), G2/Capterra listings, and developer community presence (GitHub, docs). Your Organization schema should include sameAs links to your GitHub, Product Hunt, and social profiles.
E-commerce brands. Your entities are your products and your brand. Focus on: Product schema on every product page (with brand, category, offers), Organization schema with sameAs links to Amazon/storefront pages, and category-level entity optimization. Each product becomes a distinct entity connected to your brand entity.
Local businesses. Your entity is your place. Focus on: LocalBusiness schema with address, geo, openingHours, and aggregateRating. Google Business Profile optimization. Local directory listings (Yelp, TripAdvisor, industry-specific). Your entity is anchored to a geographic location, which creates a natural relationship to the place entity.
Media / publishers. Your entities are your authors and your publication. Focus on: Person schema for each author, Organization schema for the publication, Article schema with author and publisher properties. Each author becomes a distinct entity connected to your publication entity.
How long does entity SEO take to work?
Entity SEO is a long-term investment with a characteristic timeline that differs from keyword SEO. Here is what to expect:
Days 1-30: Foundation phase. You are fixing on-site clarity — schema, naming, descriptions. During this phase, there is no measurable AI citation impact because the model has not re-indexed your changes yet. This feels like nothing is happening. It is not nothing — it is the prerequisite for everything that follows.
Days 30-60: Corroboration phase. You are building off-site signals — Wikipedia, Wikidata, directories, guest posts. During this phase, you may see small improvements in AI brand mention accuracy. The model is beginning to see corroboration of your on-site claims.
Days 60-90: Relationship phase. You are adding relationships to other entities. During this phase, you should see measurable improvement in AI citations for queries where your entity relationships are relevant. This is when the flywheel starts to turn.
Days 90-180: Compounding phase. The flywheel accelerates. More citations → more retrieval data → stronger entity representation → more citations. Sites that complete the 90-day process see 4-6x citation growth by day 180. Sites that stop at day 60 see only 1.5-2x growth.
Days 180+: Maintenance phase. Entity authority is now self-reinforcing. You need to maintain it (update schema, pursue new relationships, monitor accuracy) but the heavy lifting is done. The entity is established.
Entity SEO and Google AI Overviews
Google AI Overviews rely heavily on entity understanding to select and rank sources. When Google's AI generates an Overview, it identifies the key entities in the query, then selects sources that are recognized authorities on those entities. Entity SEO directly influences whether your site is selected.
Here is how entity authority affects AI Overviews specifically:
Entity match. Google's AI first identifies which entities the query is about. If your site is a recognized authority on those entities, you are a candidate source. If you are not in the entity graph for that topic, you are eliminated before ranking even starts.
Source confidence. Among candidate sources, Google selects those with the highest entity authority. This is where the three layers (on-site clarity, off-site corroboration, relationship density) matter directly. A site with all three layers will be selected over a site with only one or two, even if the latter has more backlinks.
Entity freshness. Google's AI prefers sources where the entity representation is recently updated. If your Organization schema was last modified 2 years ago and your competitor's was updated last week, the competitor has a freshness advantage. Keep your schema current.
Our data shows that entity-optimized sites appear in AI Overviews 3.1x more often than keyword-only sites. The effect is strongest for branded queries (queries that include a brand name) and category queries (queries like "best CRM software" that are entity-driven).
Common entity SEO mistakes
After working with 150+ sites on entity SEO, these are the mistakes that most consistently prevent entity recognition:
Mistake 1: Inconsistent entity naming. Using different names on different pages ("Acme Corp" vs. "Acme Corporation" vs. "ACME Inc.") fragments your entity. The model may treat each name as a different entity. Pick one canonical name and use it everywhere.
Mistake 2: Missing Organization schema. The most common technical mistake. Without Organization schema on your pages, the model has to infer entity ownership from context. Inference is error-prone and may attribute your content to the wrong entity.
Mistake 3: Ignoring Wikidata. Many SEOs focus on Wikipedia but ignore Wikidata. Wikidata is equally important for AI models because it provides machine-readable entity data. Wikipedia provides narrative; Wikidata provides structure. You need both.
Mistake 4: Pursuing keyword SEO instead of entity SEO. Adding more keyword-targeted pages does not strengthen your entity. It may even weaken it by creating content that is not entity-coherent. Focus on deepening your entity (better schema, more corroboration, more relationships) rather than broadening your keyword coverage.
Mistake 5: Giving up too early. Entity SEO takes 60-90 days to show meaningful results. Most sites give up around day 45. The sites that persist see 4-6x citation growth by day 180. The cost of giving up is not zero — it is the lost compounding effect.
- Use one canonical entity name consistently across all pages and off-site profiles.
- Add Organization schema to every page — do not make the model guess entity ownership.
- Create both Wikipedia and Wikidata entries for maximum entity signal strength.
- Focus on deepening entity authority, not broadening keyword coverage.
- Commit to 90+ days — entity SEO compounds but starts slowly.
- Update schema whenever entity attributes change (new products, new leadership, new location).
Entity SEO compounds — the flywheel effect
The hardest part of entity SEO is the first 90 days. You are fighting for recognition, and the model does not yet know you exist. Once you break through — once the model treats you as a recognized entity — everything gets easier.
A recognized entity gets cited more often, which means more retrieval data, which means a stronger entity representation, which means more citations. The flywheel spins on its own. The goal of the first 90 days is to push the flywheel until it starts spinning by itself.
This is fundamentally different from keyword SEO, where the flywheel is backlink-driven and can reverse if you stop building links. Entity SEO's flywheel is citation-driven — and citations are self-reinforcing because each citation strengthens the entity representation, making the next citation more likely.
Most sites give up around day 45 because the early returns are small. Do not. Entity SEO is the highest-leverage long-term play in modern search, and the compounding effect is enormous for the sites that stick with it. In our longitudinal data, sites that completed the 90-day process and then maintained their entity signals saw 12x citation growth over 12 months, compared to 3x for sites that only did keyword SEO.
The math: entity SEO costs ~20 hours of upfront work (schema audit, Wikidata, Wikipedia pitch, relationship building) and then ~2 hours/month of maintenance. The return compounds indefinitely. There is no paid channel that offers comparable long-term ROI.
The Entity Graph Viewer — see what AI models see
The free [Entity Graph Viewer](https://seosights.com/tools/entity-graph-viewer) is the key tool for entity SEO. It shows you the entity graph that AI models construct from your site — the entities they recognize, the attributes they assign, and the relationships they infer.
How it works: The Entity Graph Viewer crawls your site (up to 50 pages), extracts all schema markup, analyzes entity mentions in content, and visualizes the resulting entity graph. You see a node-and-edge diagram where nodes are entities and edges are relationships.
What to look for:
- Is your brand a recognized entity? You should see a central node with your brand name and type (Organization). If not, you need on-site schema. - How many relationships does your entity have? More edges = stronger entity authority. If your node has only 2-3 edges, you need relationship density. - Are attributes correct? The viewer shows what attributes the model assigns to your entity. If the description is wrong, the model's representation is wrong. - How does your graph compare to competitors? Run the viewer on your top competitor. If their graph is larger and denser, you have an entity gap to close.
Use the Entity Graph Viewer before starting entity SEO (to establish a baseline), during the 90-day process (to measure progress), and after (to monitor for regressions). It is the entity SEO equivalent of a rank tracker.
The future of entity SEO
Entity SEO is still in its early stages. As AI search grows, entity authority will become the primary currency of visibility. Here are the developments we expect over the next 12-24 months:
Entity authority scoring. Expect SEO tools to develop entity authority scores — similar to Domain Authority but based on entity graph metrics (relationship density, corroboration count, citation frequency). seosights is already building this.
Cross-model entity optimization. Currently, each AI model maintains its own implicit knowledge graph. In the future, expect shared entity registries or cross-model entity standards that allow you to optimize once for all models. Wikidata is the closest thing we have today.
Entity conflict resolution. When two entities claim the same name or category, AI models need to resolve the conflict. Expect better tooling for entity disambiguation — helping models distinguish between similarly named entities (e.g., "Stripe" the payments company vs. "Stripe" the pattern).
Real-time entity updates. Currently, entity representations update on a crawl schedule. Expect real-time entity APIs that allow brands to push entity updates to AI models immediately, similar to how Google Indexing API works for premium publishers.
Entity SEO + [llms.txt](/blog/llms-txt-the-robots-txt-for-the-ai-era) + [FAQ schema](/blog/faq-schema-the-underrated-ai-citation-signal). These three signals are converging. llms.txt provides the high-level entity summary. FAQ schema provides extractable Q&A. Entity schema provides the relationship graph. Together, they form a complete entity signal stack that no AI model can ignore. Sites that deploy all three will dominate AI citations in their category.
Entity SEO is not a trend — it is a structural shift in how search works. The models are reasoning in entities, and the web must adapt. The sites that build entity authority now will own AI search for years to come.
seosights team
Editorial at seosights. We build the operating system for AI search — Three Sights, one unified engine.
Put this into action
Run a full Three Sights audit on your site. 8 AI agents, 90-day roadmap, 14-day free trial — no credit card.
Start free trial