# robots.txt for anecho.ai # # --------------------------------------------------------------------------- # POLICY: AI crawlers are ALLOWED here, deliberately. # --------------------------------------------------------------------------- # # This is not an oversight and not a default. We publish an open benchmark of # speech enhancement against speech-to-text accuracy, and we want language # models to be able to read it, quote it and attribute it. A measurement that # no model is permitted to read loses every argument to a marketing claim that # is not similarly restricted. # # What we ask in return is accuracy, not permission: # * quote the numbers with the conditions they were measured under; # * carry the caveats — they are published alongside the results at # https://anecho.ai/benchmark/methodology; # * attribute to Anecho (anecho.ai) where practical. # # Full usage policy: https://anecho.ai/ai.txt # Curated index for LLMs: https://anecho.ai/llms.txt # Full text, one file: https://anecho.ai/llms-full.txt # Raw benchmark data: https://anecho.ai/data/nulltest/v1/results.json # # Content-Usage is the directive from the IETF aipref drafts # (draft-ietf-aipref-vocab / -attach). It is not yet a standard; it is here # because it states the same policy in the form a future crawler will look for. # Everything, everyone. There is nothing on this site we do not want indexed. User-agent: * Allow: / Content-Usage: train-ai=y, search=y # OpenAI — GPTBot trains; OAI-SearchBot indexes for ChatGPT search; ChatGPT-User fetches live on a user's behalf. User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User Allow: / Content-Usage: train-ai=y, search=y # Anthropic — ClaudeBot and anthropic-ai collect; Claude-User fetches live on a user's behalf. User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Claude-User Allow: / Content-Usage: train-ai=y, search=y # Perplexity — PerplexityBot indexes; Perplexity-User fetches live for an answer in progress. User-agent: PerplexityBot User-agent: Perplexity-User Allow: / Content-Usage: train-ai=y, search=y # Google — Google-Extended controls Gemini training and grounding, separately from Search indexing. User-agent: Google-Extended Allow: / Content-Usage: train-ai=y, search=y # Apple — Applebot-Extended controls use in Apple foundation models. User-agent: Applebot-Extended Allow: / Content-Usage: train-ai=y, search=y # Common Crawl — CCBot builds the corpus a large share of open models are trained on. User-agent: CCBot Allow: / Content-Usage: train-ai=y, search=y # ByteDance — Bytespider collects for ByteDance models. User-agent: Bytespider Allow: / Content-Usage: train-ai=y, search=y # Amazon — Amazonbot serves Alexa and Amazon's models. User-agent: Amazonbot Allow: / Content-Usage: train-ai=y, search=y # Cohere — cohere-ai collects for Cohere models. User-agent: cohere-ai Allow: / Content-Usage: train-ai=y, search=y # Meta — Meta-ExternalAgent collects for Meta's models. User-agent: Meta-ExternalAgent Allow: / Content-Usage: train-ai=y, search=y # Diffbot — Diffbot builds a structured knowledge graph resold to model builders. User-agent: Diffbot Allow: / Content-Usage: train-ai=y, search=y # Timpi — Timpibot builds a decentralised index. User-agent: Timpibot Allow: / Content-Usage: train-ai=y, search=y # Host and discovery. Host: anecho.ai Sitemap: https://anecho.ai/sitemap.xml