Skip to content
Repo

HarborDataset

from modal_dojo import HarborDataset

A dataset loaded from Harbor tasks.

HarborDataset(*, split: Literal['all', 'train', 'eval'] = 'train', dataset_name: str = '', path: str | None = None, task_root: str = '', task_glob: str = '*', task_names: list[str] | None = None, instruction_path: str = 'instruction.md', label_metadata_path: str | None = None, test_data_dir: str | None = None, prompt_template: str = '{instruction}', system_prompt: str = '', train_size: int | None = None, eval_size: int | None = None, train_repeats: int = 1, eval_repeats: int = 1, shuffle_tasks: bool = False, shuffle_seed: int = 0, always_download: bool = False) -> None
apply_chat_template() -> bool

Whether to apply the model’s chat template to the input.

cache_key() -> str | None
input_key() -> str

Prompt column name.

label_key() -> str

Ground-truth column name, or None when rows carry no label.

load(split: Literal['all', 'train', 'eval'] = 'all') -> Any
output_format() -> str

The on-disk format written by write(), either parquet or jsonl.

rows() -> Iterable[DatasetRow]

Load raw examples.

Returns

An iterable collection of raw examples.

validate_written(path: str) -> None

Validate the materialized file format and required columns.

write(path: str) -> None

Materialize training data at path.