Skip to main content
A dataset is a read-only collection of files that LithosBox grants to your organization. Name it when you create a sandbox, and the sandbox sees its files at /datasets/<name>. The data is attached to the sandbox rather than copied into its writable disk, so a large dataset does not use your 16 GiB (or larger) disk. LithosBox grants datasets to organizations. Organizations enabled for upload can publish their own (below); otherwise, to request a dataset, contact support. The SDK’s dataset argument requires version 0.7.1 or later; upload requires 0.7.2.
sandbox.info.dataset names the attached dataset, or is an empty string when the sandbox has none. Omit image to use the default environment with the dataset.

Upload your own

For organizations enabled for upload (others get FeatureNotEnabledError, 403). Files go straight from your machine to storage; the platform then builds the dataset, stages it on every host, and grants it to your organization.
Or from a shell: lithosbox datasets upload /data/parquet --name parquet-2026-10-01 (then lithosbox datasets status NAME / list). Only regular files are uploaded (symlinks are refused) and the folder layout is kept under /datasets/<name>/. A published dataset does not change: new data is a new name.

Create over HTTP

Send dataset with POST /vms:
The sandbox object includes a dataset field when a dataset is attached; the field is absent otherwise. See Create a sandbox.

Read dataset files

Process dataset files inside the sandbox with run() and return only the results. The tools you use, such as DuckDB or pandas, come from your image:
sandbox.files.read() runs inside the sandbox too, but returns at most about 16 MiB of text (12 MiB with read_bytes()) per call. It suits small files such as a README or manifest, not dataset contents. Write outputs to a writable path such as /work; see Work with files to download them. The mount is read-only. Writing, deleting, or renaming anything under /datasets fails with Read-only file system.

Forks, snapshots, and archives

A dataset stays with the sandbox and everything made from it:
  • Forks inherit the parent’s dataset. A fork request cannot name a dataset; over HTTP, sending dataset to /branch returns 400.
  • Snapshots record the dataset, and a restore mounts it again. Only your organization can restore such a snapshot, and only while the dataset is still granted to your organization at the same version; otherwise the restore returns 403.
  • Archive and unarchive keep the dataset attached.
A snapshot does not include the dataset’s files, so the dataset adds nothing to your snapshot storage.

Limits and errors

  • A sandbox can have one dataset, chosen at creation. It cannot be added to or removed from an existing sandbox.
  • Dataset names use lowercase letters, digits, ., _, and -, start with a letter or digit, and are at most 64 characters.
  • dataset cannot be combined with template or snapshot_id, and templates cannot include a dataset. Datasets work with image or the default image.
  • A create with a dataset always starts its image fresh rather than from a prepared template, so it takes longer than other creates.

Usage

Your organization’s usage reports each granted dataset’s size as dataset_storage_byte_seconds, for as long as the grant exists, whether or not a sandbox is using it. The console’s usage view shows it as Dataset storage. See Usage for the fields. When finished:
Deleting the sandbox does not remove the dataset or its grant.