
When a developer uploads a dataset to the Hugging Face Hub, the result is not a single opaque blob but a versioned repository whose every byte is governed by specific storage technology. Understanding what that repository physically is — a Git tree, a Dataset Card, a set of Parquet shards — helps authors publish more deliberately and reason about storage cost, deduplication behavior, and downstream loading. This deep dive walks through the directory structure that every dataset must follow, the role Xet plays inside the Git layer, and the Google Cloud infrastructure that hosts the Hub at scale. The audience is engineers and dataset authors who already know how to call push_to_hub and want to know what actually happens on the server side.

On the Hugging Face Hub, a "dataset" is not an opaque binary object. It is a Git repository, exactly like a model or Space repo, and it carries the same commit history, branches, and revision tags that any other Git project would (Hugging Face Hub datasets overview). When you call push_to_hub("username/my_dataset") from the 🤗 Datasets library, the client initializes that repo on the Hub's Git backend and pushes a curated tree of files; the same effect can be reproduced by hand with the git CLI (Adding datasets to the Hub).
A dataset repository is a directory whose contents the Hub recognizes in a fixed shape. Concretely, every dataset repo holds:
Because the repo is a Git tree, every change to those files is tracked, and the full revision history is preserved inside the same repository (Adding datasets to the Hub).
The Hub recognizes a defined set of file types, including Parquet (.parquet), CSV/TSV, JSON Lines and JSON, Arrow (.arrow), plain text, common image/audio/video/PDF formats, WebDataset (.tar), and Lance (.lance) (Adding datasets to the Hub). When push_to_hub is called from 🤗 Datasets, the library serializes the in-memory Dataset into Parquet shards before committing them to the repo, taking advantage of the columnar format's compact size and reload speed (Using datasets from the Hub).
The Hub does not store "this file is the train split" anywhere. Instead, the split each file belongs to — train, validation, or test — is inferred from the file or directory name. Files named train-00000-of-00003.parquet are recognized as part of the train split; a sibling directory named validation/ likewise becomes that split (Working with datasets on the Hub). Following this naming convention is exactly what triggers the in-browser Dataset Viewer: the viewer is generated automatically when the repo uses supported formats and a supported directory layout, with no metadata sidecar required (Hugging Face Hub datasets overview).
Publishing a dataset is structurally the same operation as initializing and pushing a Git repository whose working tree happens to contain data files, an optional loader script, and a README card.

On the Hub, the documentation page that appears when you open any dataset repository is rendered from a single file: README.md at the root of the Git tree. Authors create this file either by clicking Create Dataset Card in the repository UI or by pushing the file directly with Git. Because the file lives inside the same Git repository as the data shards, every edit to the card is a normal commit — reviewable, revertable, and versioned like code.
Above the Markdown body of the card sits the Metadata UI, which exposes the most important discovery fields:
Choosing a value in any of these widgets writes a corresponding YAML tag into the top of the README.md. These structured tags are what the Hub's search, filter, and category pages read; without them, a dataset is effectively invisible to discovery even if its Markdown body is excellent. Optional tags — including annotations_creators, size_categories, dataset_creators, and language_creators — extend the schema and are listed in the Dataset Card specification in the huggingface/hub-docs repository under datasetcard.md. That specification is the authoritative reference for which values are accepted (Hugging Face Hub — Adding datasets).
The Metadata UI handles machine-readable metadata; the rest of the README.md is free-form Markdown that documents the corpus itself. The Hub guidance recommends covering four areas: use cases of the dataset, limitations of the dataset, where the data comes from (provenance), and important ethical considerations. The Import dataset card template link in the editor scaffolds these headings automatically, and the CNN DailyMail card at huggingface.co/datasets/cnn_dailymail is cited in the official documentation as a concrete reference for how a complete card reads (Hugging Face Hub — Adding datasets). A well-written card lets downstream users judge relevance, spot biases, and assess risks before they download anything (LLM Course — Chapter 5).
Card quality is not only a documentation concern — it gates the Dataset Viewer. The Viewer renders only when the repository follows the supported directory structure and the card is in place to make the page meaningful. A repo with correctly placed Parquet shards but no card still loses the in-browser preview (Hugging Face Hub — Datasets overview). In practice, the two concerns reinforce each other: a metadata-rich card surfaces the dataset in search, and the supported structure lets the Viewer let visitors sanity-check rows before committing to a download.

Every Git repository on the Hugging Face Hub is built on Xet, the storage backend that replaces the traditional "one blob per file" Git object model for large content. According to the Hub documentation, Xet "intelligently splits files into unique chunks" so that uploads and downloads move only the bytes that have actually changed between commits, rather than retransmitting entire files (Hugging Face Hub docs). For dataset authors, this means that fixing a typo in a README or appending rows to one Parquet shard typically does not cause the other shards to be re-uploaded.
The property that makes this chunked layout cheap is content-addressable deduplication. Each chunk is identified by a hash of its contents, and identical chunks — whether they repeat inside a single file, across sibling files in the same repository, or across unrelated repositories hosted by different users — are stored exactly once in the underlying object store. The same documentation describes this as "content-addressable deduplication for large-scale object storage," which is what allows the Hub to absorb the substantial duplication inherent in public ML corpora (common tokenizer outputs, repeated license headers, shared Parquet row groups) without paying storage cost proportional to that redundancy (Hugging Face Hub docs).
This is a meaningful departure from the Git Large File Storage (LFS) pattern that many engineers have encountered elsewhere. With Git LFS, large files are replaced by small pointer objects, and the real bytes live in a separate, opt-in storage service that requires explicit configuration and client-side setup. Xet, by contrast, is the default and only backend the Hub uses for repository storage: it applies uniformly to model repos, dataset repos, and Spaces, with no per-file opt-in flag and no separate credential flow for chunk retrieval. The deduplication benefits are therefore global across the Hub rather than confined to files an author remembered to track.
A version-sensitive caveat applies. Xet is a relatively recent introduction, and older repositories on the Hub may have been committed before Xet adoption. Their blobs predate the chunked layout, so they will not benefit from cross-file deduplication or incremental re-upload in the same way that newly created or migrated repositories do. For most authors this is invisible — pushing to a fresh repo automatically uses Xet — but anyone inspecting storage accounting or rebuilding historical commits on a long-lived dataset should keep in mind that pre-Xet history can behave differently at the storage layer than recent commits.

A dataset repository on the Hugging Face Hub is, at its core, an ordinary Git repository. This single design choice is what unlocks most of the features authors rely on without ever calling a Hub-specific API. Because every change is a Git commit, the repository carries an immutable commit history, line-level diffs between any two revisions, named branches, tags, and merge commits produced by pull requests (Hugging Face Hub docs).
In practical terms, a single dataset repo can host multiple labeled versions of the data over time. A maintainer can publish v1, v2, and v3-cleaned as parallel branches, push a tag for a frozen release, and rely on the Hub to render diffs in the web UI just like GitHub would (Hugging Face Hub overview). Around that Git core the Hub layers additional collaboration features that are documented as shared by all repositories: pull requests and discussions for code review, webhooks that fire on repository events, notifications, collections to group related repos, and per-repository settings including storage limits and license files (Hugging Face Repositories docs).
git CLI WorksBecause the transport itself is Git, the standard git CLI is a first-class client. Authors can git clone, git add, git commit, and git push against a Hub repo and get full version-control semantics. Large files do not break this model because Xet sits transparently above Git, intelligently splitting files into unique chunks and managing their storage efficiently (Hugging Face Hub overview). To a user, git push looks like any other Git workflow; under the hood, large Parquet shards are deduplicated and chunked by Xet rather than copied verbatim.
For reproducibility, downstream consumers are asked to pin a specific revision — a commit hash, branch name, or tag — when loading a dataset, guaranteeing that the exact bytes used in a past experiment can be retrieved even after the dataset is updated (🤗 Datasets repository).
Git-based versioning is not free: every revision stores its files (with deduplication at the Xet layer), and history accumulates. For data where history is unnecessary, Hugging Face offers Storage Buckets as an alternative — S3-like object storage containers powered by the same Xet backend but without versioning (Hugging Face Hub overview). Buckets are non-versioned, mutable, and rely on content-addressable deduplication rather than commit history.
The Hub documentation recommends buckets for use cases such as training checkpoints, logs, intermediate artifacts, and any large collection of files that does not need version control (Storage Buckets use cases). The 🤗 Datasets library can also read and write directly to Storage Buckets, which lets users keep raw mutable data in buckets while still serving a versioned, Git-based dataset repo downstream (🤗 Datasets repository).
For authored, consumer-facing dataset releases where reproducibility matters, the Git-backed repository model is the right default. For ephemeral training state where every epoch would otherwise bloat history, Storage Buckets are the recommended alternative — same Xet storage layer, no commit graph.

The Hub exposes two distinct storage primitives that share the same underlying Xet chunking layer. Git repositories are version-controlled folders: every push produces a commit, every change is diffable, and branches let you stage work without disturbing main. They are the default home for models, datasets, and Spaces, and they are the only primitive that gives you an immutable history of what was published and when (Hub documentation).
Storage Buckets, by contrast, are S3-like object storage containers. There are no commits, no diffs, and no branches — files are mutable, can be overwritten in place, and the bucket has no notion of "last week's version." The trade-off is simplicity and scale: buckets are designed for large files that change frequently and do not need version control, with content-addressable deduplication applied at the Xet layer (Hub documentation).
The decision rule is straightforward. Publish the dataset itself, the Dataset Card, schema files, evaluation scripts, and any change that benefits from a reviewable history as a Git repository — typically inside the dataset repo on the Hub. Push training checkpoints, training logs, and ephemeral intermediate artifacts (large NumPy buffers, pre-tokenized shards you may regenerate, scratch visualizations) into a Storage Bucket, where you can mutate them freely without inflating the Git history. Typical bucket use cases documented by the Hub include storing training checkpoints, logs, and intermediate artifacts, or any large collection of files that does not need version control (Hub documentation).
A common misconception is that switching primitives changes storage efficiency. It does not. Both Git repositories and Storage Buckets are powered by the Xet backend, which intelligently splits files into unique chunks and applies content-addressable deduplication across the Hub. Whether a 4 GB Parquet shard lives in a dataset repo or in a bucket, identical chunks inside it are stored only once, and partial overlap between files is captured at chunk granularity rather than at the file level (Hub documentation). The practical consequence is that storage cost follows byte content, not the primitive you chose — so the choice between Git and Buckets can be made on versioning semantics alone.

Under every dataset, every Parquet shard, and every Git object on the Hub sits a Google Cloud substrate that the partnership pages describe as the platform's "optimal infrastructure." The scale numbers publicly cited on the Hugging Face–Google Cloud customer page frame the size of that substrate: more than two million open models, around 1.5 million datasets, and a community of 13 million AI builders. Serving that volume of traffic and storage from a single origin would be impractical, so the integration layers Google-managed services directly into the delivery path.
Downloads are accelerated through two Google Cloud entry points:
For training, the partnership unlocks native TPU support across the entire open model library, meaning developers can target Google's Tensor Processing Units directly from Hugging Face training pipelines rather than restricting themselves to GPU-bound code paths.
Safety is treated as a first-class requirement, not an afterthought. The Hub documentation lists malware scanning, GPG commit signing, organization-level access control, and user access tokens as the baseline security features available to every repository. On top of that, the Google Cloud collaboration adds built-in security scanning for models, which integrates with Google Cloud's broader enterprise-grade protection story. Together these layers mean that the bytes uploaded via push_to_hub are inspected before they are served back out.
For dataset authors, the practical consequence is operational rather than visible in any API call. They do not provision buckets, configure CDNs, set up Kubernetes clusters, or wire up scanning pipelines. The Hub, sitting on Google Cloud, exposes a Git-shaped interface on the front and a globally distributed, scanned, enterprise-grade storage and serving fabric on the back. That is why a dataset load_dataset call from a notebook in one region and a fine-tuning job running on Vertex AI in another can both reach the same files quickly, with the same security guarantees applied uniformly.
In short, the Google Cloud backbone is what turns a Git repository full of Xet-backed objects into something that behaves like a globally available content-delivery network, without the author ever needing to engineer one.

Knowing that a dataset on the Hub is a Git repository backed by Xet, with Parquet as the on-disk default and content-addressable chunk dedup underneath, changes how you should approach the very first push_to_hub call. A handful of deliberate decisions made before upload will save bandwidth, storage cost, and reviewer friction later.
Structure splits by name. The Hub recognizes dataset splits from file or directory names that match split identifiers such as train, validation, and test. A repository that lays out files like data/train-00000.parquet, data/validation-00000.parquet, data/test-00000.parquet is parsed automatically, surfaces correctly in the Dataset Viewer, and can be loaded with load_dataset("your/dataset", split="train") without a custom loading script. Departing from this naming convention typically forces you to ship a dataset script, which is an additional maintenance burden.
Prefer Parquet for tabular and text data. push_to_hub already serializes in-memory datasets to Parquet by default, and Parquet is recommended for most datasets because of its efficient compression, rich typing, and broad tooling support. CSV and JSON Lines remain supported and are easier to parse, but they are not recommended for datasets larger than several gigabytes; switching to Parquet early avoids a costly migration later.
Choose WebDataset .tar shards for large image or audio corpora. For image and audio at scale, raw files incur per-file overhead, while WebDataset shards group samples into sequentially readable tar archives that stream efficiently. Parquet is still appropriate for the same data when the use case is analytics, filtering, or metadata parsing. The Hub's supported format table (.parquet, .csv, .tsv, .jsonl, .arrow, .tar, .lance, plus common media extensions) reflects these trade-offs.
Invest in the Dataset Card and metadata fields. Discoverability on the Hub is driven by the README/Dataset Card content and structured metadata, not by the data files themselves. A well-written card — task description, citation, licensing, splits, and example rows — is what makes the dataset findable, citable, and safe to consume.
Plan for Xet's deduplication behavior. Xet splits files into unique chunks and deduplicates them content-addressably. Practically, this means that appending new rows to an existing Parquet file and pushing the new revision is dramatically cheaper than re-uploading the whole file, because unchanged chunks are already stored. Designers of update pipelines should structure datasets as append-friendly shard sets rather than monolithic files, and should chunk uploads via the huggingface_hub library using the documentation guides on uploading a folder by chunks, large-upload tips and tricks, and repository storage limits and recommendations.