ClueWeb-Crawler
ClueWeb-Crawler is the academic web crawler operated by Carnegie Mellon University's Language Technologies Institute, run by Professor Chenyan Xiong's research group. It collects pages for the ClueWeb line of research web corpora, which are distributed to researchers for information retrieval and natural language processing work and are used to construct and evaluate LLM pretraining datasets. The operator documents compliance with robots.txt and crawl-delay, conservative per-host rate limiting, automatic back-off on 429 and 503 responses, and an opt-out request form, and publishes a contact address in the user-agent string.
At a glance
- Operator: Carnegie Mellon University
- Type: Research
- RSL category:
ai-train
How Centinel checks it
- User agent: The request calls itself this crawler. Anyone can send the same string.
ClueWeb-Crawler publishes nothing Centinel can check a source against, so a match reports the name and leaves the source unconfirmed. The match tokens, verification domains, and address feeds are not published here.
Allowing or blocking it
The crawler object in the /validate response sets access_allowed to true only for a verified source that your tenant allowlists. A policy rule can allow or block this crawler by its category, Research.