Linux 软件免费装
Banner图

Site Answers – llms.txt, AI Crawler Log & Markdown

开发者 lalokupfer
更新时间 2026年10月8日 05:47
PHP版本: 7.4 及以上
WordPress版本: 7.1
版权: GPLv2 or later
版权网址: 版权信息

标签

markdown llm llms.txt aeo ai crawler

下载

1.0.0 1.0.1 1.0.2

详情介绍:

AI assistants are reading your website. Site Answers has one job: make your site readable to them, show you whether they actually came, and let you decide which of them may. These parts do that job, and together they are the whole plugin. It publishes an llms.txt file. llms.txt is an emerging convention — a plain, readable Markdown index at the root of your site that tells AI assistants what you have published and where to find it. Site Answers builds yours from your own posts and pages: your site name, your tagline, then one line per URL with its title and a short description. It is rebuilt automatically whenever you publish, edit or delete something, so it never goes stale. It publishes llms-full.txt too. Where llms.txt is the map, llms-full.txt is the whole territory: the full text of every page llms.txt lists, one after another, in one file, so an AI assistant can read your site in a single visit. Same pages, same exclusions, capped at 5 MB with a note saying where the rest is. You can switch it off in Settings. It logs which AI crawlers visit. GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, CCBot and the others are recognised by name. For each visit you see which crawler came, which page it asked for, and whether it got the page or an error. That answers a question site owners currently cannot answer at all: is AI search actually reading my content, and which pages does it want? It serves a clean Markdown copy of every page. Add .md to the end of any post or page address and you get the article as plain Markdown — headings, paragraphs, lists, links and tables — with the navigation, sidebars, share buttons and scripts left out. That is what an automated reader wants, and it is far less likely to be misquoted than a page it had to dig out of a theme. Each HTML page also advertises its Markdown twin in a Link header, so a crawler can find it without guessing. It tells you when your robots.txt is turning AI crawlers away. An empty crawler log often has a simple cause: a line in robots.txt — added by another plugin, a host, or years ago by hand — asking GPTBot or ClaudeBot to stay away. The Overview reads the robots.txt your site actually serves, the way the crawlers read it, and names every AI crawler it blocks. And it lets you choose, crawler by crawler. In Settings, each AI crawler has three choices: No change, Allow or Block. Block ChatGPT's training crawler and keep the one that fetches your page when someone asks ChatGPT about you, or the other way round. Your choices go into WordPress's own robots.txt. Until you pick something, robots.txt stays exactly as it was. How this is different Plenty of plugins will write an llms.txt for you. Almost all of them stop at that point: the file is written, and you never find out whether anything read it. This one closes the loop — it publishes the file, serves your pages in a form a machine can actually read, and then shows you whether any AI crawler turned up and what it asked for. Publishing without measuring is guesswork, and until now measuring meant reading raw server logs. Three further differences, each of which you can check in the code: What it does not do This plugin does not connect to any external service. It makes no outbound HTTP request of any kind — no API, no phone-home, no telemetry, no analytics, no third-party libraries. Everything it shows you was produced by your own site, on your own server. It stores no IP addresses. The crawler log records the crawler's name, the page path and the time — nothing else. Anything after the ? in an address is discarded before the record is written, because query strings on real sites carry e-mail addresses, order references and reset tokens. Records older than 90 days are deleted automatically. It changes nothing on its own. Your robots.txt is left exactly as it is until you choose Allow or Block for a crawler, and no file is ever written to your server: if you have a robots.txt file of your own, the plugin tells you so and leaves it alone. Why this matters AI assistants increasingly answer questions using content they read from websites, and being readable to them is becoming its own discipline — answer engine optimization, or AEO, alongside the generative engine optimization people now talk about next to ordinary SEO. Both come down to the same two practical questions: can a large language model find and parse what you published, and is it actually doing so? Site Answers gives you a straight answer to the second one, and does the groundwork for the first. Who it is for Anyone who publishes writing and wants to know whether AI search engines are reading it. Bloggers, documentation sites, shops with real content, agencies who need to show a client what is happening. Privacy No personal data of any kind is collected, stored or transmitted. Because no IP address is recorded and nothing leaves your server, this plugin does not put your site into consent-banner territory.

安装:

  1. Install and activate the plugin.
  2. Visit Site Answers → Overview. Your llms.txt link is at the top.
  3. Optionally visit Site Answers → Settings to choose which content is listed and which AI crawlers may read your site.
There is nothing to sign up for and no key to enter.

屏幕截图:

  • The llms.txt file itself, built from your own posts and pages.
  • Settings: which content is listed, the full-text file, and an Allow / Block / No change choice for each AI crawler.

升级注意事项:

1.0.2 Adds a "Get early access" link to the Overview screen. Nothing else changes, and nothing is sent anywhere. 1.0.1 Adds llms-full.txt, a robots.txt check for AI crawlers, and per-crawler Allow / Block choices. Your robots.txt does not change unless you choose. 1.0.0 First release.

常见问题:

What is llms.txt?

It is a proposed convention — like robots.txt or sitemap.xml, but written for large language models. A single Markdown file at /llms.txt that tells an AI assistant what a site contains, in a form it can read without wading through navigation, adverts and scripts. It is young and not every AI engine reads it yet, which is exactly why it costs almost nothing to publish one now.

Where is my llms.txt?

At /llms.txt on your site, and the Overview screen links straight to it.

What is llms-full.txt?

The companion to llms.txt. Instead of a list of links, it holds the full text of every page llms.txt lists, as Markdown, in one file at /llms-full.txt. Each page starts with its title and a Source: line with its address. It leaves out exactly what llms.txt leaves out — drafts, private and password-protected posts, and pages marked "noindex" — and stops at 5 MB. On a large site the first visit may not include every page yet: pages are converted a few at a time so the file never times out, and the file says so when that happens.

My llms.txt says "not found". Why?

Almost always because Settings → Reading → "Discourage search engines from indexing this site" is switched on. Site Answers honours that: if you have asked search engines to stay away, it does not publish an index of your content on your behalf. The Overview screen tells you when this is the reason. Switch that setting off and the file appears immediately. If that setting is already off, re-save Settings → Permalinks once — that rebuilds WordPress's internal address table.

Which AI crawlers are detected?

GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI); ClaudeBot, Claude-User and anthropic-ai (Anthropic); PerplexityBot and Perplexity-User (Perplexity); Google-Extended and GoogleOther (Google); CCBot (Common Crawl); Bytespider (ByteDance); Amazonbot (Amazon); Applebot-Extended (Apple); Meta-ExternalAgent (Meta); cohere-ai (Cohere); DuckAssistBot (DuckDuckGo); YouBot (You.com); and Diffbot. Google-Extended and Applebot-Extended are robots.txt names only: Google and Apple fetch pages with their ordinary crawlers and use these names to let you opt out of AI training. They can be blocked in Settings, but they never appear in the visit log.

How do I get the Markdown version of a page?

Add .md to the end of its address. A post at https://example.com/hello-world/ is also served at https://example.com/hello-world.md. The HTML page carries a Link: <…>; rel="alternate"; type="text/markdown" header pointing at it, which is how an automated reader is meant to discover it. Drafts, private posts, password-protected posts and anything you have marked "noindex" are not served this way, exactly as they are kept out of llms.txt. The Markdown copy also carries X-Robots-Tag: noindex and a canonical link back to the real page, so it cannot compete with your own article in search results.

Does this plugin send anything to anyone?

No. It makes no outbound HTTP request at all. There is no account, no API key, no telemetry and no analytics of any kind. The Overview screen has one ordinary link, "Get early access", to a page about a possible paid add-on. Nothing is sent when the screen loads; the page opens in your browser only if you click, and your e-mail address reaches us only if you type it there.

Is there a paid version?

Not yet. Everything described here is free and stays free. We are considering an optional paid add-on that answers your visitors' questions from your own content; the "Get early access" link on the Overview screen lets you ask to be told if it is built. It is never shown as a notice, and nothing in this plugin is locked behind it.

Does it record IP addresses?

No. The crawler name, the page path and the time — that is the whole record. Query strings are discarded too.

My server log shows more AI crawler visits than this plugin does. Why?

If your site uses full-page caching — nginx FastCGI cache, WP Rocket, Cloudflare, a host level cache — then a cached page is served without WordPress running at all, so the plugin never sees that visit. This is true of every WordPress-level bot log, not just this one. Treat the numbers here as a floor rather than a total.

Does it work with my SEO plugin?

Yes. If a page is marked "noindex" in Yoast SEO, Rank Math, SEOPress or The SEO Framework, Site Answers leaves it out of llms.txt by default. You can turn that off in Settings. For any other rule, the siteanswers_llms_txt_exclude_post filter lets you exclude a page in code.

Will it list my drafts or private posts?

No. Only published content. Drafts, pending, private, trashed and password-protected posts are never listed, whatever you choose in Settings.

What happens when I delete the plugin?

Everything it created is removed: its database table, every option and cached value it wrote, and its scheduled cleanup job. Deactivating the plugin, by contrast, leaves your crawler history intact so you can switch it back on without losing anything.

Can I block AI crawlers like GPTBot?

Yes, if you choose to. Site Answers → Settings has a choice for every crawler it knows: No change, Allow or Block. Block adds a Disallow: / rule for that crawler to your robots.txt; Allow adds a rule inviting it in. Nothing changes until you choose, and each crawler's row says what it is for — some collect training data, others open your page only when someone asks an AI assistant about it, and blocking those means the assistant cannot read you when asked. robots.txt is a request that well-behaved crawlers follow, not a lock. If your site has a robots.txt file of its own on the server, WordPress's robots.txt is never used, so the choices cannot apply — the Settings screen tells you when that is the case.

Why does the Overview say my robots.txt blocks a crawler I never blocked?

Because something else did: another plugin, your host, or a line in a robots.txt file on your server. The Overview reads the robots.txt your visitors actually get and names every AI crawler it turns away, whoever wrote the rule. If you set a crawler to Allow and another rule still blocks it, the Overview says that too.

更新日志:

1.0.2 1.0.1 1.0.0