Skip to content
Repo

HuggingFaceDataset

from modal_dojo import HuggingFaceDataset

A dataset loaded from a Hugging Face datasets repository.

Attributes

hf_repo str

Hugging Face dataset repository ID.

hf_split str

Source dataset split. Default: "train"

hf_config str | None

Source dataset configuration name.

input_column str

Source prompt column.

output_column str | None

Source answer column.

input_format Literal['text', 'messages', 'raw']

Shape of the source prompt: plain text to convert into messages, preformatted messages, or raw model input. Default: "text"

system_prompt str

System message added to formatted examples. Default: ""

prompt_template str

Template applied to each source prompt. Default: "{input}"

always_download bool

When training, always download the dataset from Hugging Face instead of caching it. Default: False

HuggingFaceDataset(hf_repo: str, *, hf_split: str = 'train', hf_config: str | None = None, input_column: str, output_column: str | None = None, input_format: Literal['text', 'messages', 'raw'] = 'text', system_prompt: str = '', prompt_template: str = '{input}', always_download: bool = False)
apply_chat_template() -> bool

Whether to apply the model’s chat template to the input.

cache_key() -> str | None
input_key() -> str

Prompt column name.

label_key() -> str | None

Ground-truth column name, or None when rows carry no label.

output_format() -> str

The on-disk format written by write(), either parquet or jsonl.

rows() -> Iterable[DatasetRow]

Load raw examples.

Returns

An iterable collection of raw examples.

validate_written(path: str) -> None

Validate the materialized file format and required columns.

write(path: str) -> None

Materialize training data at path.