> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mirage.strukto.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# HF Datasets

> Mount a Hugging Face dataset repository as a read-only filesystem and browse its shards with shell commands.

The HF Datasets VFS mounts a [Hugging Face Dataset](https://huggingface.co/datasets)
repo at some prefix such as `/ds/`.
All reads are lazy: only the bytes you actually `cat`/`head` get transferred.

For credential setup, see [HF Datasets Setup](/home/setup/hf_datasets).

## Install

```bash theme={null}
# No extra needed: the Hub API is reached over aiohttp, a core dependency.
uv add mirage-ai
```

## Config

```python theme={null}
import os

from mirage import MountMode, Workspace
from mirage.vfs.hf_datasets import HfDatasetsConfig, HfDatasetsVFS

config = HfDatasetsConfig(
    repo_id=os.environ["HF_DATASET_REPO"],   # "namespace/dataset-name"
    token=os.environ.get("HF_TOKEN"),
    # Optional:
    # endpoint="https://huggingface.co",
    # revision="main",
    # key_prefix="train/",
)
vfs = HfDatasetsVFS(config)
ws = Workspace({"/ds": vfs}, mode=MountMode.READ)
```

`HfDatasetsConfig` takes `repo_id` in `namespace/dataset-name` form plus an
optional access token. Public datasets need no token.

## Reading, not writing

This mount is read-only, the way a [`github`](/python/vfs/github) mount is. A Hub write is
a **commit**, and a POSIX write cannot say where a commit ends, so `echo >`,
`rm`, `cp` and `mv` are refused here rather than silently making one commit
per file. The [`hf` CLI](/python/cli/hf) is the write half: `hf download --local-dir`
puts a copy on a ram or disk mount, which is an ordinary writable filesystem,
and `hf upload` sends it back as a single commit.

## Filesystem Layout

Maps dataset repo files to virtual paths under the mount prefix.

For example, if dataset `AlienKevin/SWE-ZERO-12M-trajectories` contains:

```text theme={null}
README.md
data/train-00000-of-01000.parquet
data/train-00001-of-01000.parquet
```

Then mounting at `/ds/` exposes:

```text theme={null}
/ds/
  README.md
  data/
    train-00000-of-01000.parquet
    train-00001-of-01000.parquet
```

## Example

```python theme={null}
import asyncio
import os

from dotenv import load_dotenv

from mirage import MountMode, Workspace
from mirage.vfs.hf_datasets import HfDatasetsConfig, HfDatasetsVFS

load_dotenv(".env.development")

config = HfDatasetsConfig(
    repo_id=os.environ.get("HF_DATASET_REPO",
                           "AlienKevin/SWE-ZERO-12M-trajectories"),
    token=os.environ.get("HF_TOKEN"),
)
vfs = HfDatasetsVFS(config)


async def main() -> None:
    ws = Workspace({"/ds": vfs}, mode=MountMode.READ)

    r = await ws.execute("ls /ds/")
    print(await r.stdout_str())

    r = await ws.execute("cat /ds/README.md | head -n 20")
    print(await r.stdout_str())

    r = await ws.execute("find /ds/ -name '*.parquet' | head -n 5")
    print(await r.stdout_str())


if __name__ == "__main__":
    asyncio.run(main())
```

## Shell Commands

Every read command in [HF Buckets'](/python/vfs/hf_buckets#read-commands) set works here, as
do the text processing and path utilities, which only read. What does not is
the [File Operations](/python/vfs/hf_buckets#file-operations) group: this mount is read-only,
so `rm` and `touch` are refused, as is any command asked to write into it.

## Cache

Uses `IndexCacheStore` with `index_ttl = 600` (10 minutes). Directory
listings are cached and populate file-size/type entries for `stat`'s fast
path, so a `readdir` + per-entry `stat` (which `ls`, FUSE `getattr`, and
most shell commands trigger) costs one HTTP request instead of N.

## Use Cases

* **AI agents inspecting datasets**: Mount, browse the README, read byte
  ranges from large shards without downloading the whole dataset
* **Dataset triage**: `ls`, `stat`, `find` to see what's in a repo before
  committing to a full local copy
* **Sandboxed access**: Pin a `revision` for reproducibility
