Crawlune audits how easily your published content can be extracted, controls AI
crawler access, and publishes machine-readable summaries. These features run
locally and do not require an account or license. It does not measure or promise
citations by answer engines.
Implemented:
- Per-agent AI crawler controls for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended, distinguishing training crawlers from retrieval crawlers
- Three named positions to start from — Open, "Answers yes, training no", Closed — each setting every crawler and the matching content signals in one go
- robots.txt directives generated from those rules
- Optional refusal (403) for a crawler you blocked that requests pages anyway — off by default, and the requests keep being counted
- A count of people arriving from ChatGPT, Claude, Perplexity, Gemini and Copilot — click-throughs, which is a floor under citations rather than a measure of them
- Content signals (search, ai-input, ai-train) in robots.txt, which state what a page may be used for rather than who may fetch it — off until you set them, and a declaration rather than an enforcement
- The same refusal of AI training repeated as a noai response header, a robots meta tag and a W3C TDM reservation at /.well-known/tdmrep.json
- Markdown for agents: every page also available at its own address with .md on the end, and on request via Accept: text/markdown — off by default, and honouring the same exclusions as llms.txt
- IndexNow submission when you publish or update a page — off by default, and honouring the same exclusions as llms.txt
- An extractability audit scoring content on liftable answer blocks, question-shaped headings, external citations, quotable statistics, modification dates and thin content
- A site-wide visibility report aggregating that across all published content, worst pages first
- /llms.txt and /llms-full.txt, generated from your real content and refreshed when you publish
- Admin screens for the report, a per-page work-list behind every number, crawler controls, llms.txt and licensing, plus REST routes for the same
- Background analysis of your content, so none of the above is computed while you wait for a page to load
- Structured data for your pages when no SEO plugin is already providing it
- AI rewriting of a page's opening into a liftable answer block, reviewed as a word-level diff before anything is applied
- AI rewriting of section headings toward the question they answer, keeping each heading's level and its existing link address
- Generated FAQ data for pages that answer questions they never ask, where every proposed answer has to be text already on the page — anything else is discarded before you see it
- Applying a rewrite without disturbing surrounding block markup, and reverting a run from a journal
- License activation against the WP Shelf licensing service
Hosted rewriting requires a WP Shelf license and an enabled generation service.
Hosted rewriting is enabled for activated licenses. Public licensing enrollment
is not yet open. Local audits and crawler tools do not
depend on it.
Not included:
- Citation tracking — whether a given engine actually quoted you. Counting people who click through from an assistant is implemented and is a different thing: it is a floor under citations, not a measure of them
Crawlune does not claim it will get your site cited by ChatGPT, Perplexity, Claude
or Google AI Overviews. That is not measurable from inside WordPress and will not
be claimed. What it reports is what it measured about your content.