- A Jeu de données Hugging Face est une collection structurée de données hébergée sur le Hub Hugging Face, consultable à l’adresse huggingface.co/datasets — plus de 300 000 jeux de données publics en 2026.
- Chargez n’importe quel jeu de données en une seule ligne :
from datasets import load_dataset; ds = load_dataset('stanfordnlp/imdb') - Chaque jeu de données inclut des sous-ensembles typés (train/validation/test), un schéma de caractéristiques (texte, image, audio, étiquettes) et une option de chargement différé (streaming) pour les fichiers à l’échelle du téraoctet.
- Publiez vos propres données avec
ds.push_to_hub('votre-utilisateur/votre-jeu-de-donnees')après avoir exécutéhuggingface-cli login.
Un jeu de données Hugging Face est une collection structurée et versionnée de données stockée sur le Hub Hugging Face et consommée via la bibliothèque Python datasets . Chaque jeu de données expose un ou plusieurssous-ensembles (généralement train, validation et test), un schéma typé décrivant chaque colonne, ainsi que des métadonnées couvrant la licence, la catégorie de tâche et les étiquettes linguistiques. La bibliothèque gère le téléchargement, la mise en cache et la conversion de format, de sorte que vous n’avez presque jamais à manipuler directement les fichiers bruts.caractéristiques schéma décrivant chaque colonne, ainsi que des métadonnées couvrant la licence, la catégorie de tâche et les étiquettes linguistiques. La bibliothèque gère le téléchargement, la mise en cache et la conversion de format, de sorte que vous n’avez presque jamais à manipuler directement les fichiers bruts.
- Rechercher des jeux de données sur le Hub
- Installation
- Charger un jeu de données
- Structure d’un jeu de données : sous-ensembles et caractéristiques
- Chargement différé (streaming) de jeux de données volumineux
- Filtrage et traitement
- Conversion vers d’autres formats
- Publier votre propre jeu de données sur le Hub
- Remarques relatives à la plateforme
- Utiliser les jeux de données pour l’ajustement fin (fine-tuning) et l’évaluation
- Questions fréquemment posées
Rechercher des jeux de données sur le Hub
L’interface principale de découverte est huggingface.co/datasets. Les filtres disponibles dans l’interface utilisateur sont les suivants :
- Tâche — classification de texte, réponse aux questions, segmentation d’images, traduction, résumé, etc.
- Langue — codes ISO 639-1 (en, zh, fr, de, …)
- Licence — Apache 2.0, MIT, CC-BY, CC0, OpenRAIL, etc.
- Catégorie de taille — moins de 1 000 lignes jusqu’à plus d’un milliard de lignes
- Modalité — texte, image, audio, vidéo, données tabulaires, multimodal
Vous pouvez également effectuer des recherches par programmation à l’aide du client Python du Hub :
from huggingface_hub import list_datasets
results = list_datasets(filter='task_categories:text-classification', limit=20)
for ds in results:
print(ds.id, ds.downloads)
L’API REST du serveur Datasets expose un point de terminaison /valid qui liste tous les jeux de données disposant d’exportations précalculées au format Parquet, permettant ainsi des aperçus rapides et des échantillons de lignes sans télécharger l’intégralité du jeu de données.
Installation
Ledatasets La bibliothèque fonctionne sous Python 3.8+ et est indépendante de la plateforme.
pip install datasets
Pour prendre en charge les images et l’audio, installez les modules complémentaires appropriés :
pip install datasets[vision]# Pillow
pip install datasets # librosa, soundfile
Pour vous authentifier afin d’accéder à des jeux de données privés ou de publier des données sur le Hub :
pip install huggingface_hub
huggingface-cli login # demande votre jeton d’accès HF
Votre jeton est stocké dans ~/.cache/huggingface/token sur macOS/Linux, ou dans %USERPROFILE%.cachehuggingfacetoken sous Windows. Vous pouvez aussi définir directement la variable d’environnement HF_TOKEN .
Charger un jeu de données
Le point d’entrée principal est load_dataset(). En l’absence d’argumentsplit , celle-ci renvoie un objet DatasetDict contenant tous les sous-ensembles disponibles :
from datasets import load_dataset
ds = load_dataset('stanfordnlp/imdb')
print(ds)
# DatasetDict({
# train: Dataset({features: ['text', 'label'], num_rows: 25000})
# test: Dataset({features: ['text', 'label'], num_rows: 25000})
# })
train = ds['train']
print(train.features)
# {'text': Value(dtype='string'), 'label': ClassLabel(names=['neg', 'pos'])}
print(train[0]['text'][:120])
Loading a Single Split
train = load_dataset('stanfordnlp/imdb', split='train')
# Returns a Dataset directly, not a DatasetDict
Loading a Named Configuration
Many datasets define named configs for language variants, domain subsets, or schema versions. Pass the config name as the second positional argument:
ds = load_dataset('Helsinki-NLP/opus_books', 'en-fr')
To list all available configs for a dataset before loading:
from datasets import get_dataset_config_names
print(get_dataset_config_names('Helsinki-NLP/opus_books'))
Loading from Local Files
Pass a format name and file path instead of a Hub dataset ID. Supported formats include CSV, JSON/JSONL, Parquet, Arrow, plain text, and the ImageFolder/AudioFolder conventions.
ds = load_dataset('csv', data_files='my_data.csv')
ds = load_dataset('json', data_files={'train': 'train.jsonl', 'test': 'test.jsonl'})
ds = load_dataset('imagefolder', data_dir='./photos/')
Structure d’un jeu de données : sous-ensembles et caractéristiques
Check which splits a dataset has before loading:
from datasets import get_dataset_split_names
print(get_dataset_split_names('stanfordnlp/imdb'))
# ['train', 'test', 'unsupervised']
Lecaractéristiques dict maps column names to typed descriptors. Common feature types:
| Feature Type | Exemple | Remarques |
|---|---|---|
Valeur |
Value(dtype='string') |
Scalar — string, int32, float32, bool, etc. |
ClassLabel |
ClassLabel(names=['neg','pos']) |
Stored as int; decoded to name on access |
Sequence |
Sequence(Value('int32')) |
Variable-length list of a typed value |
Image |
Image() |
PIL Image; stored as bytes, decoded lazily |
Audio |
Audio(sampling_rate=16000) |
Dict with array et sampling_rate keys |
Translation |
Translation(languages=['en','fr']) |
Dict keyed by language code |
Chargement différé (streaming) de jeux de données volumineux
For datasets too large to download — Common Crawl, The Pile, LAION-5B — pass streaming=True. Data is fetched and decoded on the fly without filling your disk:
ds = load_dataset('allenai/c4', 'en', split='train', streaming=True)
for example in ds.take(1000):
print(example['text'][:80])
Streaming returns anIterableDataset rather than a Dataset. It supports.map(), .filter(), .shuffle(buffer_size=N), et .take(N), but not random indexing or len(). To get a fixed slice without streaming the whole dataset:
ds = load_dataset('allenai/c4', 'en', split='train[:50000]')
Split slicing accepts absolute row counts ([:50000]), percentages ([:10%]), and stepped ranges ([10%:20%]).
Filtrage et traitement
All operations run in Apache Arrow and use multiprocessing by default. The result is cached on disk; re-running the same .map() on the same data returns the cache instantly.
# Filter rows
short = train.filter(lambda x: len(x['text']) < 500)
# Batched map — much faster for tokenization
def tokenize(batch):
return tokenizer(batch['text'], truncation=True, padding='max_length')
tokenized = train.map(tokenize, batched=True, batch_size=256, num_proc=4)
# Column operations
tokenized = tokenized.remove_columns(['text'])
ds = ds.rename_column('label', 'labels')
# Shuffle and select
ds = ds.shuffle(seed=42).select(range(10000))
Conversion vers d’autres formats
| Target Format | Méthode |
|---|---|
| Pandas DataFrame | ds.to_pandas() |
| PyTorch Dataset | ds.with_format('torch') |
| TensorFlow Dataset | ds.to_tf_dataset(columns=[...], batch_size=32) |
| NumPy arrays | ds.with_format('numpy') |
| Parquet file | ds.to_parquet('output.parquet') |
| JSON / JSONL | ds.to_json('output.jsonl') |
| CSV | ds.to_csv('output.csv') |
Publier votre propre jeu de données sur le Hub
After running huggingface-cli login, push anyDataset ou DatasetDict object directly:
from datasets import Dataset, DatasetDict
import pandas as pd
df = pd.read_csv('my_data.csv')
ds = Dataset.from_pandas(df)
ds.push_to_hub('your-username/my-dataset', private=False)
To push train and test splits together:
split = ds.train_test_split(test_size=0.1)
DatasetDict({'train': split['train'], 'test': split['test']}).push_to_hub('your-username/my-dataset')
The Hub stores datasets as sharded Parquet files and auto-generates a dataset preview viewer. Add a README.md (Dataset Card) with YAML front-matter to make your dataset filterable by task, language, and license in the Hub search UI.
Remarques relatives à la plateforme
Cache Paths
| Plateforme | Default Cache Path | Override Env Var |
|---|---|---|
| macOS / Linux | ~/.cache/huggingface/datasets/ |
HF_DATASETS_CACHE |
| Windows | %USERPROFILE%.cachehuggingfacedatasets |
HF_DATASETS_CACHE |
Windows
Windows uses thespawn start method for multiprocessing, which requires your script entry point to be inside a if __name__ == '__main__': guard. Without this,.map(num_proc=4) will either hang or raise a RuntimeError. If you’re running in a Jupyter notebook, either use num_proc=1 or installmultiprocess aux côtés de datasets, which the library will prefer over the standard-library multiprocessing module.
Disk Space
Large datasets (Common Crawl, RedPajama, LAION) consume hundreds of gigabytes when fully cached. Use streaming=True to avoid downloads. To see what’s in your cache, run python -c "from datasets import inspect_dataset; print(inspect_dataset.__doc__)" or browse the cache directory directly. Cached datasets are stored as Arrow files organized by dataset name and hash; delete subfolders manually to reclaim space.
Utiliser les jeux de données pour l’ajustement fin (fine-tuning) et l’évaluation
The standard fine-tuning pipeline is: load dataset → tokenize with.map(batched=True) → set format to 'torch' → pass to aTrainer or a custom training loop. The Hugging Face transformers library’s Trainer class accepts a Dataset object directly for its train_dataset et eval_dataset arguments.
When evaluating models against benchmark datasets, the Classement des grands modèles linguistiques (LLM) provides scores across common evaluation sets — useful context when deciding which dataset to target for your own benchmark. If you’re weighing whether to fine-tune and self-host versus calling an API, the Calculatrice du seuil de rentabilité auto-hébergement vs API can model the cost crossover by request volume. For raw per-token API cost across providers, use the Calculateur de coûts d’API. And if you plan to run a fine-tuned model locally, VRAM is the hard constraint — the Calculatrice VRAM estimates GPU memory requirements from model size and quantization precision.
Questions fréquemment posées
What is the difference between a Dataset and a DatasetDict?
A Dataset is a single split — one table of rows and columns. A DatasetDict is a dict-like container holding multiple splits, and is the default return type of load_dataset() when you omit the split argument. Access individual splits by key: ds['train'], ds['test'], etc. If you passsplit='train', you get a bare Dataset directly.
How do I load a private dataset from the Hub?
Authenticate first with huggingface-cli login, or set the HF_TOKEN environment variable to your access token. Then call load_dataset('org/private-dataset', token=True). Le bloc token=True flag tells the library to use the cached or environment-variable credential. For CI/CD pipelines, set HF_TOKEN as a secret and omit the interactive login step.
Why does load_dataset() take a long time on the first call?
The first call downloads raw data files, converts them to Apache Arrow format, and writes the cache to disk. For large datasets this can take minutes or longer. Subsequent calls on the same machine return the cached Arrow files almost instantly. If you’re on a slow connection or have limited disk space, pass streaming=True to process data on the fly without caching it locally.
Can I use Hugging Face datasets without an internet connection?
Yes. Once a dataset is cached, set the environment variable HF_DATASETS_OFFLINE=1 etload_dataset() reads from the local cache without making any network requests. This is useful for air-gapped servers, reproducible offline runs, or HPC clusters where worker nodes lack internet access but share a network file system with the cache directory.
What file format does the Hub use internally?
Datasets stored on the Hub are served as Apache Parquet files, split into shards. The datasets library downloads these shards and converts them to Apache Arrow (.arrow) files for local caching. You can bypass the library entirely and access the Parquet shards directly via the Hub file browser or by using huggingface_hub.hf_hub_download().
How large can a Hugging Face dataset be?
There is no enforced size limit, and multi-terabyte datasets exist on the Hub (LAION-5B image-text pairs, large Common Crawl snapshots). For datasets above a few gigabytes, the Hub stores data as multiple Parquet shards rather than a single file. Use streaming=True in the datasets library to work with these datasets without downloading them in full, or download specific shards via the data_files argument.
