How this score is measured
The overall score weights Live Answer Visibility at 45%, Search Index Foundation at 30% and Training Data Presence at 25%, rounded to the nearest point. Critical findings can cap the score lower; any applied cap is stated next to it. Live Answer Visibility and Training Data Presence reflect crawler access plus on-page structure, with Training Data Presence also counting Common Crawl. Search Index Foundation reflects Googlebot and bingbot access plus indexability, sitemap and response health.
Our crawler tests send each bot's user agent from our own servers, not from the AI companies' verified IP ranges. Some firewalls, including Cloudflare, verify crawlers by network identity rather than user agent. That means a site can challenge our test request while allowing the real, verified crawler. When that happens this scan over-reports blocking; it never under-reports it. The Googlebot result is the most exposed to this, because fake Googlebots are the most common spoof and are policed hardest.
This scan of zhimomo.com checked 15 AI crawlers across 21 automated checks on 6 pages. None of the 13 scored crawlers are blocked. AI Access Score: 87/100 (AI-ready).
The bot matrix
robots.txt policy next to what actually happened when each bot knocked. Rows where your policy says yes but the server says no are highlighted; those are the silent failures.
Live-answer botsFetch content in real time to build the answers and citations users see today.
| Bot | Owner | robots.txt policy | Real fetch | Verify |
|---|---|---|---|---|
| OAI-SearchBot Powers ChatGPT search results | OpenAI | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| ChatGPT-User User-triggered fetches from ChatGPT; Cloudflare classes this as Agent traffic, blocked by default on ad pages for new Cloudflare domains | OpenAI | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| PerplexityBot Index for Perplexity answers | Perplexity | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| Perplexity-Userinformational User-triggered; may ignore robots.txt | Perplexity | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| Claude-SearchBot Claude search indexing | Anthropic | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| Claude-User User-triggered fetches from Claude; Cloudflare classes this as Agent traffic, blocked by default on ad pages for new Cloudflare domains | Anthropic | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| DuckAssistBot DuckAssist answers | DuckDuckGo | Partial line 3: Disallow: /backupsql | Reachable (200) |
Training crawlersCollect content for model training. Blocking them removes you from future models' memory.
| Bot | Owner | robots.txt policy | Real fetch | Verify |
|---|---|---|---|---|
| GPTBot Training corpus | OpenAI | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| ClaudeBot Training corpus | Anthropic | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| CCBot Feeds many labs' training data; blocking it is the quietest way to disappear from future models | Common Crawl | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| Meta-ExternalAgent Llama training | Meta | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| Amazonbot Alexa and Amazon AI features | Amazon | Partial line 3: Disallow: /backupsql | Reachable (200) | |
| Bytespiderinformational Known to sometimes ignore robots.txt | ByteDance | Partial line 3: Disallow: /backupsql | Reachable (200) |
Search-index foundationCopilot rides Bing's index and AI Overviews rides Google's. If these starve, the AI layer above them starves.
| Bot | Owner | robots.txt policy | Real fetch | Verify |
|---|---|---|---|---|
| Googlebot Google Search index; AI Overviews rides this. Cloudflare treats it as mixed Search and Training, so as of September 15, 2026 Cloudflare AI training blocks can hit it | Partial line 3: Disallow: /backupsql | Reachable (200) | ||
| bingbot Bing index; Copilot and ChatGPT search lean on it. Cloudflare treats it as mixed Search and Training, so as of September 15, 2026 Cloudflare AI training blocks can hit it | Microsoft | Partial line 3: Disallow: /backupsql | Reachable (200) |
Training opt-out tokensNot crawlers. Directives controlling whether already-crawled content trains AI.
| Token | Controls | Your robots.txt |
|---|---|---|
| Google-Extended | Gemini training use of Google's crawl Blocking Google-Extended does NOT affect your Google Search ranking, but it does remove your content from Gemini training use. | Partial |
| Applebot-Extended | Apple AI training use Blocking Applebot-Extended does not affect Siri or Spotlight search. | Partial |
Your scanned pages
Real defects live on money pages, not just homepages. We scanned 6 pages: your homepage plus pages picked from your sitemap, preferring product, pricing and services pages, because those are the pages AI assistants quote when buyers ask.
| Page | Status | Indexable | Canonical | Title | Words |
|---|---|---|---|---|---|
| / | 200 | Yes | None | 知墨墨知识产权-专注于优质国际知识产权服务 | 143 |
| / | 200 | Yes | None | 知墨墨知识产权-专注于优质国际知识产权服务 | 143 |
| /category/2.html | 200 | Yes | None | 关于知墨墨-知墨墨知识产权-专注于优质国际知识产权服务 | 94 |
| /category/3.html | 200 | Yes | None | 业务领域-知墨墨知识产权-专注于优质国际知识产权服务 | 101 |
| /team.html | 200 | Yes | None | 专业团队-知墨墨知识产权-专注于优质国际知识产权服务 | 145 |
| /category/5.html | 200 | Yes | None | 知识中心-知墨墨知识产权-专注于优质国际知识产权服务 | 121 |
Warnings (4)
AI systems have almost nothing to quote or cite from this page.
LLM retrieval splits pages into chunks along heading structure; broken hierarchy produces orphan chunks with no context.
These are the raw material that citation snippets are built from. Your meta description is under 50 characters, giving AI systems too little to quote. Character counts are for the decoded text, so an entity like & counts as one character, not five.
curl -s "https://www.zhimomo.com/" | grep -ioE '<title[^>]*>[^<]*</title>|<meta[^>]*name=.?description[^>]*>'Without a machine-readable identity, AI systems have to guess who you are instead of knowing.
Good to know (9)
A canonical tag prevents URL variants from splitting the page's identity.
curl -s "https://www.zhimomo.com/" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'Same as "No canonical tag on homepage" above.
curl -s "https://zhimomo.com/" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'Same as "No canonical tag on homepage" above.
curl -s "https://zhimomo.com/category/2.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'Same as "No canonical tag on homepage" above.
curl -s "https://zhimomo.com/category/3.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'Same as "No canonical tag on homepage" above.
curl -s "https://zhimomo.com/team.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'Same as "No canonical tag on homepage" above.
curl -s "https://zhimomo.com/category/5.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'OG tags control how shared and cited links render on every platform.
curl -s "https://www.zhimomo.com/" | grep -ioE '<meta[^>]*property=.?og:[^>]*>'llms.txt is an emerging convention some AI crawlers read for site guidance. Adding one is a five minute quick win.
curl -sI "https://www.zhimomo.com/llms.txt" | head -1Copilot and ChatGPT search lean on Bing's index; if Bing has not indexed you, two of six AI platforms cannot cite you. Automated Bing verification is unreliable without authenticated APIs, so we will not fake it.
Recommended robots.txt additions (free, no email needed)
# Recommended AI crawler access, generated by the AI Crawler Access Report by Citant.ai # Place these lines ABOVE any general User-agent: * block in robots.txt User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: DuckAssistBot Allow: / User-agent: GPTBot Allow: / User-agent: ClaudeBot Allow: / User-agent: CCBot Allow: / User-agent: Meta-ExternalAgent Allow: / User-agent: Amazonbot Allow: / User-agent: Googlebot Allow: / User-agent: bingbot Allow: / # Optional training opt-outs. These are business decisions, not defects: # blocking Google-Extended removes you from Gemini training but does NOT # affect Google Search ranking. Leave them out to stay fully open. # User-agent: Google-Extended # Disallow: /
Your prioritized fix plan
No criticals found. Warnings are ordered by impact. CMS not detected; generic instructions shown. Open the print-ready report to save it as a PDF.
1. Add extractable main content
Fixes: Only 143 words of extractable content on the homepage
- The page carries under 150 words of extractable main content; AI systems have almost nothing to quote or cite.
- Write the page copy into the HTML itself rather than images or embedded widgets.
2. Repair the heading hierarchy
Fixes: Heading hierarchy issues: 1 empty headings
- Use exactly one H1 per page and do not skip levels; LLM retrieval splits pages into chunks along heading structure, and broken hierarchy produces orphan chunks with no context.
- Restructure headings in the page template rather than styling text to look like headings.
3. Fix title and meta description (/)
Fixes: Title or description needs work: meta description on / is 44 chars (aim for 50 to 160)
- Write a title of 15 to 70 characters containing your brand or the page topic, and a meta description of 50 to 160 characters.
curl -s "https://www.zhimomo.com/" | grep -ioE '<title[^>]*>[^<]*</title>|<meta[^>]*name=.?description[^>]*>'4. Add Organization schema with sameAs links
Fixes: No Organization schema on the homepage
- Add the JSON-LD below to your homepage head, filling in your real details.
- The sameAs links are the entity-corroboration signal that connects your site to your LinkedIn and other profiles.
- Validate at validator.schema.org.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Zhimomo",
"url": "https://zhimomo.com",
"logo": "https://zhimomo.com/logo.png",
"description": "What Zhimomo does, in one sentence.",
"sameAs": [
"https://www.linkedin.com/company/zhimomo",
"https://x.com/zhimomo"
]
}
</script>5. Fix the canonical tag (/)
Fixes: No canonical tag on homepage
- The canonical must be an absolute https URL on this same host. A canonical pointing at a staging domain or another host donates the page's identity elsewhere.
- Fix the template that renders the link rel=canonical tag, deploy, and verify with the command in this report.
curl -s "https://www.zhimomo.com/" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'6. Fix the canonical tag (/)
Fixes: No canonical tag on /
- The canonical must be an absolute https URL on this same host. A canonical pointing at a staging domain or another host donates the page's identity elsewhere.
- Fix the template that renders the link rel=canonical tag, deploy, and verify with the command in this report.
curl -s "https://zhimomo.com/" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'7. Fix the canonical tag (/category/2.html)
Fixes: No canonical tag on /category/2.html
- The canonical must be an absolute https URL on this same host. A canonical pointing at a staging domain or another host donates the page's identity elsewhere.
- Fix the template that renders the link rel=canonical tag, deploy, and verify with the command in this report.
curl -s "https://zhimomo.com/category/2.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'8. Fix the canonical tag (/category/3.html)
Fixes: No canonical tag on /category/3.html
- The canonical must be an absolute https URL on this same host. A canonical pointing at a staging domain or another host donates the page's identity elsewhere.
- Fix the template that renders the link rel=canonical tag, deploy, and verify with the command in this report.
curl -s "https://zhimomo.com/category/3.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'9. Fix the canonical tag (/team.html)
Fixes: No canonical tag on /team.html
- The canonical must be an absolute https URL on this same host. A canonical pointing at a staging domain or another host donates the page's identity elsewhere.
- Fix the template that renders the link rel=canonical tag, deploy, and verify with the command in this report.
curl -s "https://zhimomo.com/team.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'10. Fix the canonical tag (/category/5.html)
Fixes: No canonical tag on /category/5.html
- The canonical must be an absolute https URL on this same host. A canonical pointing at a staging domain or another host donates the page's identity elsewhere.
- Fix the template that renders the link rel=canonical tag, deploy, and verify with the command in this report.
curl -s "https://zhimomo.com/category/5.html" | grep -ioE '<link[^>]*rel=.?canonical[^>]*>'11. Complete Open Graph tags
Fixes: Open Graph incomplete: missing og:title, og:description, og:image
- Add og:title, og:description and og:image meta tags to the page head.
curl -s "https://www.zhimomo.com/" | grep -ioE '<meta[^>]*property=.?og:[^>]*>'12. Add an llms.txt file
Fixes: No llms.txt found
- Save the generated draft below as /llms.txt at your site root.
- It is a plain-markdown guide that tells AI systems what your most important pages are; adoption is emerging, and it costs nothing.
# Zhimomo > One-sentence description of what zhimomo.com does, written for AI systems. Edit this line. ## Key pages - [知墨墨知识产权-专注于优质国际知识产权服务](https://www.zhimomo.com/): 知墨墨知识产权深耕知识产权行业多年,提供国内外专利申请、国内外商标注册、版权登记等服务。 - [知墨墨知识产权-专注于优质国际知识产权服务](https://zhimomo.com/): 知墨墨知识产权深耕知识产权行业多年,提供国内外专利申请、国内外商标注册、版权登记等服务。 - [关于知墨墨-知墨墨知识产权-专注于优质国际知识产权服务](https://zhimomo.com/category/2.html): 关于知墨墨-知墨墨知识产权深耕知识产权行业多年,提供国内外专利申请、国内外商标注册、版权登记等服务。 - [业务领域-知墨墨知识产权-专注于优质国际知识产权服务](https://zhimomo.com/category/3.html): 业务领域-知墨墨知识产权深耕知识产权行业多年,提供国内外专利申请、国内外商标注册、版权登记等服务。 - [专业团队-知墨墨知识产权-专注于优质国际知识产权服务](https://zhimomo.com/team.html): 专业团队-知墨墨知识产权深耕知识产权行业多年,提供国内外专利申请、国内外商标注册、版权登记等服务。 - [知识中心-知墨墨知识产权-专注于优质国际知识产权服务](https://zhimomo.com/category/5.html): 知识中心-知墨墨知识产权深耕知识产权行业多年,提供国内外专利申请、国内外商标注册、版权登记等服务。 ## Contact - [Contact](https://zhimomo.com/contact): How to reach Zhimomo
curl -sI "https://www.zhimomo.com/llms.txt" | head -1