Crawler and robots strategy
Separate training bots from live-citation bots. A common mistake is blocking OAI-SearchBot, PerplexityBot and similar citation crawlers while allowing traffic that will never produce citations.
Technical research · Technical GEO
Technical GEO research is not “write more articles.” It asks whether AI systems can access pages, parse answer units, align brand entities and verify descriptions with external evidence. This page maps repeatable technical levers and priorities. They improve citation conditions; they do not guarantee a one-off or permanent appearance in answers.
Research scope
We define technical work as lowering the cost for machines to understand and verify a brand. Content quality, industry authority and third-party sources still decide whether an answer is worth citing; technical work decides whether systems get a fair chance to read you correctly.
Six technical levers
In practice, fix problems that mean content “cannot be read at all,” then move to structure, extraction and entities. Wrong order means content investment never shows up in AI answers.
Separate training bots from live-citation bots. A common mistake is blocking OAI-SearchBot, PerplexityBot and similar citation crawlers while allowing traffic that will never produce citations.
Critical answers should not exist only in a client-rendered DOM. If first-paint HTML lacks service definitions, FAQs or addresses, AI retrieval may fail to extract them reliably.
Markup such as Organization, WebSite, FAQPage and Article must match visible content. Wrong or empty schema increases misunderstanding; it does not automatically add points.
Lead every section with a stand-alone answer, then add conditions, steps, tables and limits. Align headings to buyer questions; prefer table fields for comparisons.
Keep Chinese and English names, addresses, phone numbers, service scope, authors and official accounts consistent; use sameAs to verifiable external profiles so the brand is not split into multiple entities.
Slow responses, frequent 5xx errors or resources that block rendering make crawling unstable. Core Web Vitals are health metrics—not the only GEO score.
Execution priority
| Priority | Work | Why first | How to verify |
|---|---|---|---|
| P0 | Allow live-citation crawlers; fix mistaken blocks on sitemaps / critical paths | Great content is invisible if it cannot be read | robots tests, crawl logs, public URLs return full HTML |
| P1 | Make core service / about / location pages server-readable answers | Extraction happens at paragraph level, not slogan level | Definitions, service scope and FAQs still visible with JS off |
| P2 | Align Organization / FAQ / Article markup with visible content | Reduces machine confusion about names, services and authorship | Structured-data testing tools show no conflicts; fields match on-page text |
| P3 | Rebuild high-intent pages as answer units + internal topic clusters | Improves extractability and topical coherence | Fixed-query retest: mentions, cited URLs, description accuracy |
| P4 | Entity consistency and external evidence | Supports verification—does not replace the official-site foundation | Cross-platform name/address consistency; fewer sample description errors |
Priorities are for diagnostic sequencing. Actual order depends on stack, CMS, approval workflows and industry compliance. For finance, healthcare and other high-risk categories, accuracy and disclosability come before exposure.
Common failure modes
Animation, modals and client-side routing fill the first paint; answer text loads late or hides inside interactive widgets.
Markup claims FAQs or addresses the page does not show—or Chinese and English names contradict each other.
The homepage has a brand card; service pages are slogans and images. Models cannot find a citable service definition.
After robots and schema fixes, no content or external evidence work and no fixed-query review—so nobody can tell if anything improved.
Measuring technical interventions
After technical work ships, keep at least two fixed-sampling rounds—before and after. Platforms, language, region, login state and record fields must match for comparison.
If technical access is open but answers stay thin, the next priority is usually content and external evidence—not more markup.
Limits
Full sampling definitions: research methodology. Public brand readiness: GEO Guide list.
Next steps
Access, trust, extraction, entities and performance—audit breakdown and a 47-check scope.
How each platform shows sources, plus Hong Kong–first versus overseas/English market configuration differences.
Fixed queries, repeat sampling, mention/citation definitions and publication boundaries.
Build a baseline from your real queries and prioritise technical, content and entity gaps.
Technical work asks whether machines can reach and parse the page; content work asks whether the answer deserves citation. Run both—but unblock retrieval first when pages are unreadable.
No. Markup must match visible content, and the underlying definitions, FAQs, and entity facts still need to be clear and verifiable.
Usually no as a blanket rule. Training bots and retrieval bots differ. Blocking citation crawlers reduces the chance of being fetched and cited.
Yes, but critical answers should appear in server-returned HTML—or you must ensure retrieval systems reliably obtain equivalent content. Client-only rendering is riskier.
Retest mentions, cited URLs, and description accuracy with the same queries, platforms, and regions—and check whether citations land on the repaired answer pages.
Want a baseline of how AI answers describe your brand?
Request a free Audit