> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nanovm.dev.lithosai.cloud/llms.txt
> Use this file to discover all available pages before exploring further.

# Use a dataset

> Mount a read-only dataset granted to your organization and read it from a sandbox.

A dataset is a read-only collection of files that LithosBox grants to your
organization. Name it when you create a sandbox, and the sandbox sees its files at
`/datasets/<name>`. The data is attached to the sandbox rather than copied into
its writable disk, so a large dataset does not use your 16 GiB (or larger) disk.

LithosBox grants datasets to organizations. Organizations enabled for upload can
publish their own (below); otherwise, to request a dataset,
[contact support](/launch#support). The SDK's `dataset` argument requires version
0.7.1 or later; upload requires 0.7.2.

```python theme={null}
from lithosbox import LithosBox

box = LithosBox()
sandbox = box.sandboxes.create(image="python:3.12-slim", dataset="YOUR_DATASET")
print(sandbox.info.dataset)
print(sandbox.run("ls -la /datasets/YOUR_DATASET").check().stdout)
```

`sandbox.info.dataset` names the attached dataset, or is an empty string when the
sandbox has none. Omit `image` to use the default environment with the dataset.

## Upload your own

For organizations enabled for upload (others get `FeatureNotEnabledError`, `403`).
Files go straight from your machine to storage; the platform then builds the
dataset, stages it on every host, and grants it to your organization.

```python theme={null}
ds = box.datasets.create("parquet-2026-10-01")   # names: lowercase, digits, . _ -
ds.upload("/data/parquet")                        # parallel; re-run to resume
ds.publish()
ds.wait()                                         # uploading > building > pinning > ready
sandbox = box.sandboxes.create(dataset=ds.name)
```

Or from a shell: `lithosbox datasets upload /data/parquet --name parquet-2026-10-01`
(then `lithosbox datasets status NAME` / `list`). Only regular files are uploaded
(symlinks are refused) and the folder layout is kept under `/datasets/<name>/`.
A published dataset does not change: new data is a new name.

## Create over HTTP

Send `dataset` with `POST /vms`:

```sh theme={null}
curl -fsS https://api.sandbox.lithosai.cloud/vms \
  -H "Authorization: Bearer $LITHOSBOX_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"image": "python:3.12-slim", "dataset": "YOUR_DATASET"}'
```

The sandbox object includes a `dataset` field when a dataset is attached; the
field is absent otherwise. See [Create a sandbox](/reference/http-api#create-a-sandbox).

## Read dataset files

Process dataset files inside the sandbox with `run()` and return only the
results. The tools you use, such as DuckDB or pandas, come from your image:

```python theme={null}
result = sandbox.run(
    "df -h /datasets/YOUR_DATASET && ls /datasets/YOUR_DATASET | head"
)
print(result.check().stdout)
```

`sandbox.files.read()` runs inside the sandbox too, but returns at most about
16 MiB of text (12 MiB with `read_bytes()`) per call. It suits small files such
as a README or manifest, not dataset contents. Write outputs to a writable path
such as `/work`; see [Work with files](/guides/files) to download them.

The mount is read-only. Writing, deleting, or renaming anything under
`/datasets` fails with `Read-only file system`.

## Forks, snapshots, and archives

A dataset stays with the sandbox and everything made from it:

* **Forks** inherit the parent's dataset. A fork request cannot name a dataset;
  over HTTP, sending `dataset` to `/branch` returns `400`.
* **Snapshots** record the dataset, and a restore mounts it again. Only your
  organization can restore such a snapshot, and only while the dataset is still
  granted to your organization at the same version; otherwise the restore
  returns `403`.
* **Archive and unarchive** keep the dataset attached.

A snapshot does not include the dataset's files, so the dataset adds nothing to
your snapshot storage.

## Limits and errors

* A sandbox can have one dataset, chosen at creation. It cannot be added to or
  removed from an existing sandbox.
* Dataset names use lowercase letters, digits, `.`, `_`, and `-`, start with a
  letter or digit, and are at most 64 characters.
* `dataset` cannot be combined with `template` or `snapshot_id`, and templates
  cannot include a dataset. Datasets work with `image` or the default image.
* A create with a dataset always starts its image fresh rather than from a
  prepared template, so it takes longer than other creates.

| Error                                               | Cause                                                                                                          | What to do                                                 |
| --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| `ValueError` (SDK), `InvalidRequestError`, or `400` | `dataset` combined with `template` or `snapshot_id`, an invalid name, or `dataset` sent to a fork or template. | Remove the conflicting field.                              |
| `AuthError` / `403`                                 | The dataset is not granted to your organization, or a snapshot's dataset is no longer granted at that version. | Check the name, or contact support.                        |
| `ServerError` / `503` naming the dataset            | No host that can run the sandbox currently holds the dataset.                                                  | Retry after a short wait. If it persists, contact support. |

## Usage

Your organization's usage reports each granted dataset's size as
`dataset_storage_byte_seconds`, for as long as the grant exists, whether or not a
sandbox is using it. The console's usage view shows it as **Dataset storage**.
See [Usage](/reference/http-api#usage) for the fields.

When finished:

```python theme={null}
sandbox.delete()
box.close()
```

Deleting the sandbox does not remove the dataset or its grant.
