The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Overview
FineWeb is a 15-trillion token pretraining dataset derived from 96 Common Crawl snapshots. It is designed to produce better-performing large language models than other open datasets. The dataset and its curation methodology are fully documented and ablated to advance understanding of high-quality data curation.
Best for
Researchers and engineers building or benchmarking open LLMs with high-quality pretraining data.
Use cases
- Pretraining large language models from scratch
- Ablation studies on data curation techniques
- Benchmarking open-source dataset quality for LLM training
Notes
Pros
- Proven to improve LLM performance over other open datasets
- Fully documented and ablated curation process
- Large scale with 15 trillion tokens from diverse web sources
Cons
- Requires significant compute resources to process and use
- Derived only from Common Crawl, limiting domain coverage
- Not a ready-to-use tool; requires integration into training pipelines
Pairs with
Other entries in the index that connect to this one.