Auditing your site for the agentic era: beyond traditional SEO
Search engines used to index keywords and count backlinks. Now, models parse your DOM to synthesize answers. Here is how to test what agents actually see when they crawl your pages.
For thirty years, optimizing a website meant optimizing for an indexer. Googlebot would fetch your HTML, parse the keywords, tally incoming backlinks, and rank your page in a list of ten blue links.
If a human clicked, they landed on your page. If they stayed, they were yours.
That model is no longer the whole story. Today, a rapidly growing share of user queries never result in a website visit. Users ask questions directly to Claude, ChatGPT, Gemini, or Google AI Overviews. These systems do not merely index pages; they synthesize answers.
The model acts as an intermediary reader. It queries the web, reads three or four relevant sources in a few hundred milliseconds, extracts facts, reconciles discrepancies, and presents a summarized answer to the user.
A user asks an AI agent which platform supports multi-region CMS sync. The agent crawler (Claude, ChatGPT or Gemini) fetches two sources with a budget of about 150 ms per fetch. Site A is a client-side SPA that returns an empty <div id="root"> shell, so parsing fails and it is skipped. Site B is server-rendered with semantic HTML, JSON-LD and an /llms.txt manifest; it parses in 48 ms with high confidence, and the agent answers: "Site B offers native multi-region sync across 12 regions", citing Site B.
If your website is ambiguous, slow to render, or reliant on client-side scripts to display text, the agent skips you or hallucinates incorrect information about your product.
Here is how modern AI agents actually read your website, and a checklist for conducting a machine-readability audit.
The physics of an agentic crawl
Most web developers assume that because Googlebot can execute JavaScript, modern AI scrapers execute JavaScript just as generously.
They do not.
Running a headless Chromium browser instance with full JavaScript evaluation costs roughly 10x to 30x more compute than performing a simple HTTP GET request. When an AI engine needs to synthesize an answer to an urgent user query, it operates under extreme latency constraints—typically under two seconds for the entire synthesis loop.
Because of this:
- Raw HTTP parsers are the primary pass: AI scrapers often attempt to extract content directly from raw server-rendered HTML.
- Strict timeout budgets: If a scraper does run headless browser evaluation, its timeout budget for hydration is ruthlessly short (often 150ms to 500ms). If your content requires three cascading
fetch()requests before the main text renders, the scraper will capture an empty skeleton. - Aggressive noise stripping: Scrapers strip navigation menus, cookie banners, tracking scripts, and complex SVG trees before passing the text token stream into the model's context window.
If your content only lives in JavaScript state, to an agent, your content does not exist.
The 4-step machine-readability audit
To verify whether your site is legible to AI search engines, run through this four-step audit.
1. The Raw Curl Test (Server-Side Delivery)
Open your terminal and run a raw curl request against your marketing pages:
curl -s -A "Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://docs.perplexity.ai/)" \
https://example.com/pricing/ | grep -i "enterprise"If this command returns empty or only shows an empty <div id="root"></div>, your site is invisible to raw HTTP scrapers.
The fix: Implement Server-Side Rendering (SSR) or Static Site Generation (SSG). Every critical piece of product information, pricing tier, and technical specification must exist in the raw HTML string emitted by the server. Blinkk ships most websites with Root.js, which server-side renders by default.
2. Semantic Document Hierarchy
When an LLM summarizes a webpage, it relies heavily on semantic tags to construct entity relationships. It maps <h1> to the core topic, <h2> to sub-themes, and <table> or <dl> to structured attributes.
Common failure patterns we see in client audits:
- The Div-Only Hierarchy: Using
<div class="title-xl">instead of proper<h1>through<h3>tags. - Orphaned Numbers: Displaying statistics with decorative typography:
<div class="stat-number">99.9%</div>
<div class="stat-caption">Uptime guaranteed</div>An LLM parser frequently fails to associate these two disjointed text nodes, leading to citations like "offers high uptime" rather than "guarantees 99.9% uptime."
- Tabbed Content Hidden in State: Accordions and tabs that do not exist in the DOM until clicked. Use native HTML
<details>and<summary>elements or render tab content into the HTML with CSS toggles rather than conditional DOM mounts.
3. Comprehensive JSON-LD Knowledge Graphs
Schema.org markup via JSON-LD is the native language of information extraction. It allows you to declare facts unambiguously so the model does not have to guess.
A complete machine-readable audit checks for:
OrganizationSchema: Identifying company founders, subsidiaries, official social profiles, and headquarters.Product/SoftwareApplicationSchema: Specifying features, pricing models, supported platforms, and release dates.FAQPageSchema: Providing direct question-and-answer pairs that an answer engine can quote verbatim.BreadcrumbListSchema: Clarifying site taxonomy and parent-child page relationships.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "Does Blinkk support multi-language localization?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Yes. Every project ships with a content engine configured for automated and manual localization pipelines across 40+ locales."
}
}
]
}
</script>When an LLM reads this script tag, its confidence score for your answer increases dramatically compared to parsing conversational marketing copy.
4. The /llms.txt Discovery Standard
In 2024, the web began adopting the /llms.txt standard—a lightweight, markdown-formatted file placed at the root of a domain (similar to robots.txt) designed specifically for AI models.
While robots.txt tells crawlers what not to index, /llms.txt tells models how to understand your company.
A properly structured /llms.txt includes:
- A one-line summary of your company's core mission.
- Direct links to primary canonical pages with clear, factual summaries.
- Pointers to documentation, pricing structures, and technical whitepapers.
At Blinkk, we generate our /llms.txt dynamically from our page registry and content pipeline. When an agent queries our site, it gets an authoritative, noise-free summary of our services and engineering capabilities in less than 50 milliseconds.
Winning the synthetic search era
The brands that win in the generative era will not be the ones that stuff keywords into hidden meta tags. They will be the ones that respect the machine reader.
By engineering websites with pristine server-rendered semantics, explicit structured schemas, and dedicated model manifests, you ensure that when someone asks an AI about your market, your story is told accurately, completely, and authoritatively.
Built-in SEO & AEO insights
To prevent machine readability from becoming a one-time audit that quietly drifts out of date after launch, Blinkk ships websites with an integrated SEO & AEO insights panel built directly into the Root CMS sidebar.
Rather than forcing teams to juggle disparate dashboards across Google Search Console, CrUX reports, and third-party crawlers, the panel gives content strategists and engineering teams agency-grade visibility right where the site is maintained. It executes continuous Generative Engine Optimization (GEO) audits—scoring AI readiness, verifying robots.txt crawler access for AI agents (such as GPTBot, ClaudeBot, PerplexityBot, and Google-Extended), validating JSON-LD knowledge graphs, and flagging missing schema types.
At the same time, it surfaces high-impact striking-distance keywords and question-style queries tailored for answer engines, tracks field Core Web Vitals, and leverages grounded market research to recommend new pages based on actual demand gaps.
Further reading
- The llms.txt specification and the Google Devsite Lighthouse llms.txt reference
- Google Search Central's guide to AI optimization and Googlebot crawling
- schema.org and Google's introduction to structured data
- JSON-LD specification
- Google's JavaScript SEO basics and robots.txt introduction
- Lighthouse overview and the agentic browsing discussion
- GEO: Generative Engine Optimization
- Root.js, the open-source engine we build hybrid sites with