Skip to main content

BeansLLM-CorpusBot

recommended: observe
ai scraper

BeansLLM-CorpusBot collects public web content to build corpora for language-model training.

What it does

BeansLLM-CorpusBot tends to make broad, high-volume sweeps across many pages. Timing is unpredictable: traffic may remain heavy throughout a collection pass, then stop entirely.

How to identify it

User-agent contains any of:

  • BeansLLM-CorpusBot

robots.txt

Well-behaved crawlers honor robots.txt. To control BeansLLM-CorpusBot:

Allow full access

User-agent: BeansLLM-CorpusBot
Allow: /

Block entirely

User-agent: BeansLLM-CorpusBot
Disallow: /

Grey-area scrapers ignore robots.txt. Botscope enforces your policy at the edge or origin regardless — see how →

See BeansLLM-CorpusBot on your own site

Botscope shows every bot and AI agent hitting your site — and lets you allow, challenge, or block each one. Observe first, enforce when you're ready.

Start free — connect in minutes