# The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

## Overview

FineWeb is a 15-trillion token pretraining dataset derived from 96 Common Crawl snapshots. It is designed to produce better-performing large language models than other open datasets. The dataset and its curation methodology are fully documented and ablated to advance understanding of high-quality data curation.

## Best for

Researchers and engineers building or benchmarking open LLMs with high-quality pretraining data.

## Use cases

- Pretraining large language models from scratch
- Ablation studies on data curation techniques
- Benchmarking open-source dataset quality for LLM training

## Notes

## Pros

- Proven to improve LLM performance over other open datasets
- Fully documented and ablated curation process
- Large scale with 15 trillion tokens from diverse web sources

## Cons

- Requires significant compute resources to process and use
- Derived only from Common Crawl, limiting domain coverage
- Not a ready-to-use tool; requires integration into training pipelines

## Pairs with

Other entries in the index that connect to this one.
