HuggingFaceDataset
from modal_dojo import HuggingFaceDatasetA dataset loaded from a Hugging Face datasets repository.
Attributes
hf_repo str
Hugging Face dataset repository ID.
hf_split str
Source dataset split. Default: "train"
hf_config str | None
Source dataset configuration name.
input_column str
Source prompt column.
output_column str | None
Source answer column.
input_format Literal['text', 'messages', 'raw']
Shape of the source prompt: plain text to convert into messages, preformatted messages, or raw model input. Default: "text"
system_prompt str
System message added to formatted examples. Default: ""
prompt_template str
Template applied to each source prompt. Default: "{input}"
always_download bool
When training, always download the dataset from Hugging Face instead of caching it. Default: False
Constructor
Section titled “Constructor”HuggingFaceDataset(hf_repo: str, *, hf_split: str = 'train', hf_config: str | None = None, input_column: str, output_column: str | None = None, input_format: Literal['text', 'messages', 'raw'] = 'text', system_prompt: str = '', prompt_template: str = '{input}', always_download: bool = False)apply_chat_template
Section titled “apply_chat_template”apply_chat_template() -> boolWhether to apply the model’s chat template to the input.
cache_key
Section titled “cache_key”cache_key() -> str | Noneinput_key
Section titled “input_key”input_key() -> strPrompt column name.
label_key
Section titled “label_key”label_key() -> str | NoneGround-truth column name, or None when rows carry no label.
output_format
Section titled “output_format”output_format() -> strThe on-disk format written by write(), either parquet or jsonl.
rows() -> Iterable[DatasetRow]Load raw examples.
Returns
An iterable collection of raw examples.
validate_written
Section titled “validate_written”validate_written(path: str) -> NoneValidate the materialized file format and required columns.
write(path: str) -> NoneMaterialize training data at path.