common crawl
title: Common Crawl category: concepts created: 2025-01-04 updated: 2025-01-04 tags: [common-crawl, training-data, web-scraping, mai-thinking-1, microsoft, public-datasets, data-sources, internet-archive] sources: [raw/feeds/2026-06-11--ainews-microsoft-build-mai-thinking-1-and-mai-family-models.md] confidence: high
Common Crawl
Large-scale web crawling project that provides open datasets of web page content, widely used as a foundational training data source for large language models. Notable for being transparently disclosed as a primary data source for mai-thinking-1.
MAI-Thinking-1 Usage
Microsoft explicitly disclosed using Common Crawl data alongside private sources in MAI-Thinking-1 training, representing part of their clean-data-lineage approach with transparent data sourcing.
Data Processing
For MAI-Thinking-1, Common Crawl data underwent:
- Heavy extraction and deduplication work
- Targeted sub-pipelines for different domains
- Quality filtering and curation processes
- Integration with private data sources
Industry Standard
Common Crawl serves as a standard foundational dataset across the AI industry, providing consistent web-scale text data that enables comparative model development and reproducible research.
Transparency Value
Microsoft's explicit disclosure of Common Crawl usage demonstrates commitment to data transparency, contrasting with more secretive training data practices at other frontier AI companies.
See also
- mai-thinking-1
- clean-data-lineage
- Training-Data
- Web-Scraping