AI Visibility Score Methodology
How VerisAI measures whether a company website is open to AI crawlers, readable without JavaScript, and structured enough to be used as a source.
Version 2.7.1 (2026-08-31) – Technical framework for AI access, website facts, SEO foundation, citation readiness, Knowledge Diff, and priority risk scoring.
How to read this methodology
This page explains the scoring model behind the AI Readiness baseline. It is written for technical reviewers, web teams, SEO specialists, and governance owners.
What the score measures
Whether the site declares access for AI crawlers, serves the right company facts in its HTML, and carries the signals that support accurate AI answers.
What the score does not guarantee
The score measures readiness signals. It does not guarantee that any AI platform will cite the website in every answer.
What teams should do with it
Use the score to prioritize website fixes: access, indexability, structured data, content clarity, and citation-readiness gaps.
Overall Score Calculation
The overall score combines three areas: AI access and understanding, SEO foundation, and citation-readiness signals.
VerisAI L1 checks what a website declares about AI crawlers in its robots.txt, and separately whether the page could be read at all. L2-L4 check whether they can understand it. L7 estimates whether the page is structured enough to be citation-ready; L9 applies a capped adjustment to that final citation-readiness number.
Formula: Overall Score = (AI Readiness + SEO + Citation Readiness) / 3
AI Readiness (33.3%)
Layer 1-4: Is the site open to AI crawlers, and are the core company facts present and extractable?
Components: Gateway, SSR, Indexability, Content Quality
SEO Foundation (33.3%)
Layer 5-6: Does the website have the technical and on-page foundation needed for discovery?
Components: Technical SEO, On-Page SEO
Citation Readiness (33.3%)
Layer 7 plus L9 adjustment: Does the website expose signals that support citation readiness, without negative quality signals?
Components: GPT, Gemini, Claude, Perplexity, Grok, and Meta readiness scoring; capped L9 adjustment
Thresholds:
- ≥70 = PASS (strong readiness signals)
- 40-69 = WARN (partial readiness, visible gaps remain)
- <40 = FAIL (critical blocking issues or weak signals)
Component 1: AI Readiness Score (Layer 1-4)
Formula: AI Readiness = (Content Score × 0.6) + (Indexability Score × 0.4)
Content Score = Layer 4 Quality × SSR Factor
Indexability Score = 100 + Layer 3 Penalties
Layer 1: Gateway (Binary Gate)
Observatory aggregation rule. In EU AI Readiness Observatory Report 01 v1.2, every company website remains in the full market sample. A latest retained scan with l1_score = 0 is counted as not crawler-accessible in that measured visit, regardless of the diagnostic cause. Those websites are not removed from the report. Only L2–L8 averages and cross-layer relationships use l1_score > 0, because no downstream page content is available to analyse at L1=0. This is the report's aggregation rule; it does not change how a new incomplete scanner observation is persisted.
Status: PASS, BLOCKED, or UNMEASURED — the third when robots.txt itself could not be read (no response, 429, 5xx, unreadable body; a 404 or 410 is a readable answer meaning “no restrictions”). UNMEASURED is a statement about the scan, not a claim about the site, and is never presented as a block.
Layer 1 separates two things that are often confused: what a site declares about named AI crawlers, and whether a crawler can read the page at all. They have different owners inside a company and different fixes, so VerisAI reports them separately and never merges them into one verdict.
How the assessment is performed. VerisAI requests the page once, under its own published identity (VerisAI-IndexScanner, declared at verisai.eu/crawler). It does not send other companies' crawler User-Agent strings and does not claim to be GPTBot, ClaudeBot, Google-Extended, PerplexityBot, OAI-SearchBot or GrokBot. What each of those crawlers would be permitted to do is then determined internally, from the site's own robots.txt, by applying that bot's robots.txt rules to the requested URL under an RFC 9309 matcher. That establishes what the site declares for the bot, not the provider's complete crawling behaviour.
Why the method changed (28 August 2026). Until that date part of the assessment sent each provider's crawler User-Agent and scored the response as that provider's access. A User-Agent string does not authenticate a crawler: providers publish IP ranges, and often reverse DNS, precisely so their crawlers can be identified by something a third party cannot copy. A request carrying a provider's name from VerisAI's own network is therefore VerisAI's request, and what it measures is VerisAI's access — not that provider's. Bot-protection products can also treat an unverifiable provider name differently from an ordinary client, in either direction, so the result was not even a stable measurement of our own access. VerisAI now reports two things and keeps them apart: what the site’s robots.txt declares for each named crawler, and what happened to the one VerisAI request. Neither is a claim about how a provider's verified crawler is treated, because that cannot be established from outside.
Scoring (94 bot-weight pts + 15 pts llms.txt bonus + 5 pts AI Discovery bonus, capped at 100):
Bot weights reflect the relative importance assigned by the VerisAI scoring model. Blocking major AI crawlers carries a higher penalty than blocking lower-priority crawlers. Weights should be reviewed periodically as platform behavior changes.
| Bot | Platform | Points | AI chatbot share (US, May 2026) |
|---|---|---|---|
| GPTBot | ChatGPT / OpenAI | 39 | 60.6 % |
| Google-Extended | Gemini / Google AI | 28 | 15.1 % |
| OAI-SearchBot | ChatGPT Search / OpenAI | 15 | — (part of ChatGPT) |
| PerplexityBot | Perplexity AI | 1 | 5.4 % |
| ClaudeBot | Claude / Anthropic | 9 | 5.0 % |
| GrokBot | Grok / xAI | 2 | 0.6 % |
| llms.txt | All platforms | +15 bonus | — |
| AI Discovery endpoints | All platforms | +5 bonus | — |
Weights updated 15 August 2026 to worldwide web-visit share (FirstPageSage/Similarweb, May 2026), replacing the earlier U.S.-only set; the share column below is the older U.S. figure and is retained for context only. Bot weights are not a direct proportional mapping of share — they are the worldwide chatbot user-share allocation mapped onto the 94-point budget. Weighting by observed crawler volume was investigated on 14 August 2026 and deliberately rejected: the value of unblocking a crawler is the audience behind it, not the number of requests it makes. Reviewed quarterly.
- robots.txt policy (per bot): ALLOWED / BLOCKED / NOT_SPECIFIED, evaluated against RFC 9309 — exact user-agent tokens, wildcards, terminal
$, and specificity by literal path length. A bot earns its weight when the site's own declaration does not deny it. - Page retrievability (once, as VerisAI): whether the declared VerisAI crawler received a usable page. Reported alongside the per-bot policy above, never as a statement about a named provider. A rate limit (HTTP 429) or a failure on VerisAI's side is recorded as unmeasured and never scored as a block — a scanner's own crawl rate must not lower a site's score.
- llms.txt bonus (+15 pts): Present at domain root, HTTP 200, non-empty content — the file exists and carries real content. It is a guidance file, not an access-control mechanism — its presence is not permission and its absence is not a block.
- llms-full.txt (detected, not scored here): Full content dump for AI consumption — used as +10 pts signal in Layer 7 Citation Readiness per platform
- AI Discovery endpoints bonus (+5 pts max, added 2026-07-26):
/.well-known/ai.txt(+1),/ai/summary.json(+2),/ai/faq.json(+1),/ai/service.json(+1) — content-validated (not just HTTP 200) to avoid false positives from static-asset hosts that return a fallback page for any unmatched path. A community standard (geo-checklist.dev), not yet a baseline requirement like llms.txt. - Informative bot checks: ChatGPT-User, anthropic-ai, Claude-SearchBot, Claude-Web, Claude-User, Google-CloudVertexBot, Googlebot, BingBot, Perplexity-User, Meta-ExternalAgent, Meta-ExternalFetcher, Amazonbot, Applebot-Extended, Applebot, Bytespider, DuckAssistBot, MistralAI-User, DeepseekBot, and CCBot are tracked for robots.txt diagnostics but do not change L1 score.
Result: BLOCKED gates AI Readiness and the overall score to 0. Diagnostic layers may still be computed and shown — Monitoring uses them to show which individual gaps exist — but they cannot lift the gated result.
Theory: If AI bots cannot access your site, all other optimizations are irrelevant. Bot weights are model assumptions used by VerisAI and should be reviewed periodically as platform behavior changes. llms.txt and llms-full.txt are separate signals: llms.txt is an access/index file (L1), llms-full.txt is a content completeness signal for citation (L7). GPTBot documentation, Google AI crawlers
Layer 2: Server-Side Rendering (Binary Gate)
Quality: GOOD / PARTIAL / FAILED
Checks (4 critical elements):
- Valid
<title>tag (not empty/placeholder) <h1>heading exists- Text content >500 characters
- JSON-LD schema present
Scoring:
- 0 missing = GOOD (SSR Factor = 1.0)
- 1-2 missing = PARTIAL (SSR Factor = 0.7)
- 3+ missing = FAILED → AI Readiness = 0
Theory: AI bots rely on server-rendered HTML. Missing critical elements = empty page for bots. HTML5 spec, Google structured data
Layer 3: Indexability (Penalty System)
Penalties: 0 to -60 points
Canonical Tags (-20 pts per issue, max -60):
- Missing canonical tag: -20
- Multiple conflicting canonicals: -20
- Relative URL (not absolute): -20
Language Declaration (-10 pts): Missing <html lang="xx">: -10
JSON-LD Validation (-20 pts): Broken/invalid JSON-LD: -20
Theory: Technical errors confuse AI crawlers about which version to index. Canonical URLs, JSON-LD spec
Layer 4: Ground Truth Completeness (Max 100)
Layer 4 measures whether AI crawlers can extract canonical company facts from the site. It uses deterministic JSON-LD, HTML, and visible-text extraction; it does not call an LLM.
- Identity core (35 pts): company name (12), meaningful company description of at least 30 characters (12), canonical URL/domain (6), and logo or brand visual identity (5).
- Business facts (25 pts): products/services (12), target market/audience/vertical (5), HQ/country/location (5), and USP or positioning statement (3).
- Trust/entity proof (20 pts): contact email or form/contact-page signal (2-5), phone or physical address (2-4), sameAs/entity anchors (4), Organization/LocalBusiness/Corporation/ProfessionalService schema (3), privacy/terms link (2), and detailed author bio (2).
- Extractable structure (20 pts): meaningful visible text of at least 800 characters (7), H1 (4), valid parseable JSON-LD (4), links to about/contact/services/products pages (3), and key facts front-loaded in the first 30% of text (2).
Critical missing cap: if company name, description, products/services, or meaningful visible text is missing, Layer 4 cannot exceed 69.
Extraction reads Organization, LocalBusiness, Corporation, ProfessionalService, SoftwareApplication, and WebSite JSON-LD, including nodes nested in @graph, together with canonical/meta/heading/logo/contact/navigation HTML signals and visible-text patterns.
Scope note: detectContentType_() still labels VIDEO, AUDIO, ARTICLE, PRODUCT, ORGANIZATION, or GENERIC for UI and L7 compatibility, but it does not determine the Layer 4 score. The former type-specific model is retained read-only as legacy_content_quality and never flows into layer4.score.
Component 2: SEO Foundation (Layer 5-6)
Formula: SEO Score = (Technical SEO + On-Page SEO) / 2
Layer 5: Technical SEO (Max 100)
- Sitemap (30 pts): sitemap.xml exists (15) + Valid XML (15)
- HTTPS (20 pts): HTTPS enabled (20)
- Mobile-Friendly (25 pts): Viewport meta tag (15) + width=device-width (10)
- Performance (25 pts): CSS files <5 (10) + JS files <10 (10) + Images <50 (5)
Theory: Technical foundation enables discovery and indexing. Sitemap protocol, Core Web Vitals
Layer 6: On-Page SEO (Max 100)
- Open Graph (25 pts): og:title (8) + og:description (8) + og:image (9)
- Twitter Cards (15 pts): twitter:card (8) + twitter:title (7)
- Internal Links (20 pts): ≥5 links (20) or 2-4 links (10)
- Images (20 pts): Alt coverage ≥80% (20) or ≥50% (10)
- Headings (20 pts): Valid H1-H2 hierarchy (20) or Single H1 (10)
Theory: Proper meta tags enable social sharing and preview generation. Open Graph Protocol, Twitter Cards
Component 3: Citation Readiness (Layer 7 + L9 adjustment)
Formula: Citation Readiness = max(0, (GPT*0.546 + Gemini*0.284 + Claude*0.093 + Perplexity*0.013 + Grok*0.024 + Meta*0.040) + L9 penalties)
Weighting: L7 is market-share-weighted using VerisAI Intel W15/2026. Weights are revised quarterly and moved to worldwide share on 15 August 2026: GPT 54.6%, Gemini 28.4%, Claude 9.3%, Perplexity 1.3%, Grok 2.4%, Meta 4.0%.
Scope note: Citation Readiness scores measure technical prerequisites and proxy signals. They do not predict exact citation frequency. Actual AI citations also depend on authority, brand mentions, source availability, platform behavior, and ranking systems that AI platforms do not fully disclose.
Shared signal across all platforms: llms-full.txt (+10 pts each)
If /llms-full.txt is present at domain root (HTTP 200), each platform score receives +10 pts. This file provides a structured full-content dump for direct AI corpus ingestion — a distinct signal from llms.txt (which is an access index scored in Layer 1). All platform scores are capped at 100.
Layer 9: Negative Signals (Citation Adjustment: 0 to -18)
- Purpose: a penalty-only, deterministic adjustment to the final Citation Readiness composite. It does not change AI Readiness, SEO, or the individual L7 platform scores.
- Signals: high CTA density (-4), popup/interstitial marker (-5), suspected keyword stuffing (-5), and a missing author on an L4-detected article (-4).
- Scope: popup/interstitial detection reads only server-delivered HTML and inline CSS, so it cannot reliably detect JavaScript-triggered popups; a class or id alone is not proof of an intrusive overlay.
Scope note: L9 is a capped heuristic and not a substitute for manual UX review. Like L7, it measures proxy signals and does not prove actual AI citation behavior.
GPT / ChatGPT (OpenAI) Readiness (Max 100)
- OAI-SearchBot Access (25 pts): Allowed in robots.txt – citation bot for ChatGPT Search
- GPTBot Access (10 pts): Allowed in robots.txt – training bot, affects brand representation
- JSON-LD Schema (25 pts): Valid JSON-LD present
- Content Quality (20 pts): Layer 4 score ≥70
- Citation Structure (20 pts): H2 subheadings (10) + Lists/bullets (10)
- llms-full.txt (+10 pts, capped at 100): Full content dump present
Note: OAI-SearchBot drives ChatGPT Search citations. GPTBot affects training data representation only. GPTBot documentation, OAI-SearchBot documentation
Gemini (Google AI) Readiness (Max 100)
- Google-Extended not denied in robots.txt (20 pts): The site has not opted out of Google AI training data. Until 28 August 2026 this was 10 points for the declaration plus 10 for an HTTP request sent under Google's crawler User-Agent; that request measured VerisAI's access rather than Google's, so it was removed and its points consolidated here. The maximum is unchanged.
- Current Gemini readiness inputs: Google-Extended policy, valid JSON-LD, specific content type, Organization schema, visible dates, visible contact context, and
llms-full.txtwhere available. - Known gap: On-page E-E-A-T proxies are scored — author markup, author-page links, Person/LinkedIn
sameAs, Wikipedia/WikidatasameAs, publication and modification dates. What remains a gap is independently verified authority: backlinks, confirmed author credentials, and externally validated citations. A high proxy score is not evidence of authority. - llms-full.txt (+10 pts, capped at 100): Full content dump present
Note: Gemini cites from Google Search index — Google-Extended is AI training opt-out proxy only, not a direct crawl access bot. VerisAI scores on-page E-E-A-T proxies; it does not measure independently verified authority. E-E-A-T guidelines, Google-Extended
Claude (Anthropic) Readiness (Max 100)
- L1 Gateway PASS (20 pts): Site not blocking bots
- Citation Metadata (20 pts): Canonical URL (10) + og:url (10)
- Trust Signals (10 pts): HTTPS enabled
- Answer Engine Optimization (40 pts): FAQ/Q&A section (15) + Question-format headings (15) + Definition lists (10)
- Rendering Quality (10 pts): SSR quality GOOD
- llms-full.txt (+10 pts, capped at 100): Full content dump for Brave Search ingestion
Note: Claude uses Brave Search index for web citations. FAQ and question-format headings are key extraction signals for answer retrieval. Anthropic docs
Perplexity Readiness (Max 100)
- PerplexityBot Access (25 pts): Allowed in robots.txt (25) – important access signal; blocking it can prevent Perplexity from indexing or citing the content
- Answer Engine Optimization (45 pts): FAQ section (20) + Question-format headings (15) + Definition lists (10)
- Rendering Quality (15 pts): SSR quality GOOD (15)
- Content Quality (15 pts): Layer 4 score ≥70 (15)
- llms-full.txt (+10 pts, capped at 100): Structured full content for Perplexity indexing
Note: PerplexityBot is an important access signal — blocking it can prevent Perplexity from indexing or citing the content. FAQ sections and question-format headings can improve answer extraction and citation readiness. How Perplexity works
Grok (xAI) Readiness (Max 100)
- GrokBot not denied in robots.txt (35 pts): The primary declared access signal for xAI indexing. Was 25 points plus 10 for an HTTP request under GrokBot's User-Agent, consolidated on 28 August 2026 for the reason given under Gemini. The maximum is unchanged, but a denied declaration now caps Grok at 65, below the citable threshold.
- Twitter/X card metadata (15 pts):
twitter:cardmetadata present - X sameAs in Organization schema (10 pts): Organization schema links to twitter.com or x.com profile
- Freshness (15 pts):
dateModifiedpresent in structured data - Content Quality (15 pts): Layer 4 score ≥70
- llms-full.txt (+10 pts, capped at 100): Structured full content for xAI indexing
Note: GrokBot access, X-native metadata, X profile identity, and freshness are the primary Grok readiness signals. xAI
Why Equal Weighting (1/3 + 1/3 + 1/3)?
AI Readiness (33.3%)
Technical capability: is your content open to AI crawlers and readable by them?
SEO Foundation (33.3%)
Discoverability: Can humans and search engines find you?
Citation Readiness (33.3%)
Authority: does the page carry what AI platforms look for in a source they can trust? This measures readiness, not observed citations.
Critical Insight: All three pillars must be present:
- High AI Readiness + Poor SEO = Nobody finds you
- High SEO + Poor Citation = AI ignores you
- High Citation + Blocked Gateway = Score = 0
Score thresholds defined
- PASS / CITABLE (≥70)
- The page meets minimum AI visibility requirements. Signals are sufficient for AI crawlers to access, parse, and potentially cite the content.
- WARN / PARTIAL (40–69)
- Significant gaps exist. The page may be accessible but has missing structured data, weak content signals, or citation barriers that reduce readiness for AI citation likelihood.
- FAIL / NOT CITABLE (<40)
- Critical failures detected — blocked crawling, missing SSR, or insufficient content. AI systems have little they can reliably read or attribute here.
Knowledge Diff Methodology
Purpose: compare AI-generated company narratives against deterministic, crawler-visible website facts.
Ground truth source
AI knowledge gap snapshot uses VerisAI's VCL Layer 4 Ground Truth Completeness output as the source of website facts. The website fact set is derived from fetched HTML, structured data, visible content, and identity signals evaluated by the VCL content layer. It is not generated by asking an LLM to invent or infer the company's official facts.
Pre-comparison gate
The AI comparison runs only when crawler-visible ground truth is strong enough for a reliable diff. The current gate requires sufficient Layer 4 quality, sufficient ground truth completeness, and no critical missing identity facts. If the gate fails, the user receives a "website ground truth needed" result instead of an AI narrative comparison.
Compared fields
- Company name and canonical identity
- Company description and positioning
- Products and services
- Headquarters country and location signals where available
- Target market, employee range, USP, certifications, and contact signals where available
AI narrative providers
When the gate passes, VerisAI takes a same-run snapshot across ChatGPT/OpenAI, Gemini, Claude, Perplexity, and Grok. Each platform answer is compared with the same L4-derived ground truth so results can be read per platform and across the aggregate report.
Diff categories
- Matched: the AI answer aligns with the website fact.
- Discrepancy: the AI answer states a different value or materially different interpretation.
- Missing in AI: the website exposes a fact that the AI answer does not include.
- Hallucinated by AI: the AI answer contains a claim not supported by the L4-derived website facts.
Scope note: AI knowledge gap output is a point-in-time diagnostic snapshot. 100webs benchmark pages track historical rank movement where prior snapshot data is available. Customer monitoring drift view is being prepared as a weekly per-domain view; real-time alerting, competitor comparison, and guaranteed citation detection are not currently offered.
EU AI Act public-evidence check Methodology
Purpose: map public website evidence that may indicate EU AI Act exposure or disclosure work before August 2026.
Evidence source
The EU AI Act public-evidence check reads reachable public website pages and extracts public signals such as chatbot references, AI-generated content claims, automated recommendation language, sensitive-context wording, disclosure text, contact paths, and governance or policy pages. It does not inspect internal systems, contracts, product logs, model inventories, or private documents.
Readiness output
The scan returns an operational exposure tier, public evidence findings, missing disclosure signals, and recommended next actions for internal owners, counsel, risk, or compliance review. The result is a lead triage and evidence-preparation tool, not a legal classification engine.
Boundary
VerisAI does not provide legal advice, legal certification, formal EU AI Act conformity assessment, or a definitive compliance determination. Any flagged item must be confirmed by qualified legal counsel before a compliance claim is made.
Industry Standards & Documentation
Official Standards Bodies:
- Schema.org – Structured data vocabulary
- Google Search Central – SEO best practices, E-E-A-T
- W3C – Web standards (HTML, accessibility)
- IETF – Internet protocols (robots.txt RFC 9309)
AI Platform Documentation:
- OpenAI GPTBot – Bot access guidelines
- Google AI Crawlers – CloudVertexBot, Google-Extended
- Anthropic Claude – AI assistant documentation
- Perplexity AI – Search quality overview
Validation Tools:
Academic & Industry Research:
Version History
v2.7.1 (2026-08-31): Observatory L1 aggregation clarified
- Report 01 v1.2 retains all 78,110 EU-27 company websites in the full sample and reports the 14,708 latest retained L1=0 results as a crawler-access market outcome.
- L2–L8 averages and cross-layer relationships use the 63,402 websites with L1>0 because L1=0 provides no downstream page content to analyse.
- The report aggregation rule is explicitly separated from the scanner persistence rule for a new incomplete observation.
v2.7.0 (2026-08-28): L1 no longer sends providers' crawler User-Agents
- Removed the six per-bot HTTP probes. A User-Agent string does not authenticate a crawler, so those requests measured VerisAI's access while the result was reported as the provider's.
- L1 now scores what robots.txt declares for each named bot. Whether the page is readable is measured once under
VerisAI-IndexScannerand reported as a separate axis. - Added the UNMEASURED status for a robots.txt that could not be read, kept distinct from BLOCKED.
- L7: the Gemini and Grok HTTP-access awards were consolidated onto each provider's robots.txt declaration (10+10→20, 25+10→35). Provider maxima and the composite scale are unchanged.
v2.5.4 (2026-07-27): L4 Ground Truth Completeness correction
- L4: replaced the obsolete type-specific content model with the live deterministic Ground Truth Completeness model: identity core, business facts, trust/entity proof, and extractable structure.
- Documented the critical-missing cap of 69 and clarified that detected content type remains a UI/L7 label only; legacy type-specific scoring does not flow into Layer 4 score.
v2.5.3 (2026-07-27): L9 Negative Signals citation adjustment
- L9 Negative Signals: added a deterministic, penalty-only adjustment (0 to -18) to the final citation composite for high CTA density, popup/interstitial markers, suspected keyword stuffing, and missing article authors.
- L9 does not change AI Readiness, SEO, or individual platform readiness scores. Popup/interstitial detection is limited to server-delivered HTML and inline CSS and cannot reliably detect JavaScript-triggered popups.
v2.5.2 (2026-07-26): L1 AI Discovery endpoints bonus
- L1 Gateway: added a content-validated bonus (max +5 pts) for AI Discovery endpoints —
/.well-known/ai.txt,/ai/summary.json,/ai/faq.json,/ai/service.json(community standard, geo-checklist.dev). - Validation checks JSON structure and minimum field lengths, not just HTTP 200, to avoid false positives from static-asset hosts that serve a fallback page for any unmatched path.
v2.5.0 (2026-06-14): L1 access interpretation and L7 platform weighting update
- Clarified layer interpretation: L1 checked whether AI systems can access a website, L2-L4 whether they can understand it, L7 citation readiness. L1 superseded by v2.7.0: it now reports what robots.txt declares, with readability measured separately.
- Clarified that L1 separates declared crawler policy from observed bot-level HTTP access because robots.txt alone does not prove usable AI crawler access. Superseded by v2.7.0 (28 August 2026): the per-bot HTTP probes were removed and the second axis is now a single request under VerisAI’s own identity.
- L7: added Meta AI proxy readiness and updated market-share weighting: GPT 61.2 % • Gemini 15.5 % • Claude 10.2 % • Perplexity 7.1 % • Grok 2.0 % • Meta 4.0 %. Superseded 15 August 2026 by the worldwide weighting shown in the Citation Readiness section above.
v2.4.0 (2026-05-06): Current services and monitoring status
- Added EU AI Act public-evidence check methodology and legal boundary language.
- Clarified that 100webs historical rank snapshots are implemented where prior benchmark data exists.
- Clarified that weekly per-domain monitoring drift is in progress and real-time alerting is not currently offered.
v2.3.0 (2026-05-02): Knowledge Diff ground truth methodology
- Documented that Knowledge Diff uses VCL Layer 4 Ground Truth Completeness as the deterministic source of website facts.
- Added the pre-comparison gate: weak website ground truth blocks AI narrative comparison and returns a ground-truth-needed result.
- Clarified that AI knowledge gap output is a point-in-time diagnostic snapshot, not customer monitoring, real-time alerting, or guaranteed citation detection.
v2.2.0 (2026-04-06): llms.txt / llms-full.txt signal separation
- L1 Gateway: llms.txt bonus increased from +10 to +15 pts — standard transitioning from optional to baseline as Anthropic, Cursor, Mintlify adopt it; validation tightened (HTTP 200 + non-empty content required)
- L1 Gateway:
llms-full.txtnow detected and stored in score details — separate signal fromllms.txt - L7 Citation Readiness:
llms-full.txtadded as +10 pts to platform scores, capped at 100 — full content dump enables direct AI corpus ingestion independently of platform-specific crawlers - Rationale:
llms.txt= access/index signal (belongs in L1 Gateway);llms-full.txt= content completeness signal (belongs in L7 Citation Readiness). Merging both into one L1 bonus was a design flaw
v2.1.0 (2026-04-06): L1 Gateway — market-share-weighted bot scoring
- Replaced equal-weight formula with market-share-weighted per-bot points; the v2.5.0 weights were GPTBot 36 pts, Google-Extended 20 pts, OAI-SearchBot 14 pts, PerplexityBot 10 pts, ClaudeBot 10 pts, and GrokBot 4 pts.
- Methodological rationale: blocking the dominant AI platform (ChatGPT, 60.2 % market share) cannot carry the same score penalty as blocking a minority platform — equal weighting created false equivalence exploitable by competitors
- Market share data source: VerisAI Market Intelligence Pipeline, Week 15/2026 (AI platform referral traffic analysis)
- Commitment: bot weights will be revised quarterly (Q1 January • Q2 April • Q3 July • Q4 October) based on current market share data from the VerisAI Intel pipeline
v2.4.1 (2026-05-29): Grok added to L7, L1 bot weights rebalanced
- L7: added Grok (xAI) as 5th platform.
- L1: OAI-SearchBot promoted from informative to scored. Weights AT THAT TIME (superseded 15 August 2026 by the worldwide set shown above): GPTBot 36, Google-Extended 20, OAI-SearchBot 14, PerplexityBot 10, ClaudeBot 10, GrokBot 4; total bot weight 94, plus llms.txt bonus capped at 100.
v2.0.0 (2026-02-25): L7 methodology revision based on actual AI bot behavior research
- GPT: replaced GPTBot-only scoring with OAI-SearchBot (25 pts) as primary citation bot + GPTBot (10 pts) for training; rebalanced remaining points
- Gemini: documented the move away from CloudVertexBot-only interpretation and clarified that Gemini cites via Google Search index; E-E-A-T proxy scoring was a known gap at that time; on-page proxies have been scored since L4 v2.0.3 and L7 v2.1.0, and the remaining gap is independently verified authority.
- Claude: replaced Privacy/Terms scoring with FAQ (15), Question headings (15), Definition lists (10); reflects Brave Search index citation behavior
- Perplexity: PerplexityBot elevated to 25 pts as critical gate; FAQ expanded to 20 pts; removed logo/footer credibility signals
- Added informative-only bot list to L1 (OAI-SearchBot, ChatGPT-User, Claude-Web, etc.) — tracked but not scored
v1.0.1 (2026-02-18): Scope clarification
- Added Citation Readiness scope note: scores reflect technical eligibility indicators, not actual citation probability
v1.0.0 (2026-02-15): Initial public release
- Established equal-weight formula: (AI Readiness + SEO + Citation) / 3
- Documented all 8 layers with point breakdowns
- Added platform-specific scoring (GPT, Gemini, Claude, Perplexity)
- Linked official documentation sources
Quarterly Review Commitment: L1 bot weights are revised every quarter (Q1 Jan · Q2 Apr · Q3 Jul · Q4 Oct) based on current AI platform market share data from the VerisAI Intel pipeline. The full methodology is updated whenever AI platforms publish new official crawler documentation, new LLM citation research emerges, or web standards change in ways that affect AI visibility measurement.
Run a free domain check first, then use this methodology to understand the score, the risks, and the fixes behind it.