Skip to content
Repo

MultimodalDataset

from modal_dojo import MultimodalDataset

Dataset of text prompts paired with image, audio, or video data.

MultimodalDataset(rows: Iterable[dict[str, Any]] | None = None, *, modality: Literal['image', 'audio', 'video'] = 'audio', media_column: str | None = None) -> None
apply_chat_template() -> bool

Whether to apply the model’s chat template to the input.

cache_key() -> str | None
input_key() -> str

Prompt column name.

label_key() -> str

Ground-truth column name, or None when rows carry no label.

output_format() -> str

The on-disk format written by write(), either parquet or jsonl.

rows() -> Iterable[DatasetRow]

Load raw examples.

Returns

An iterable collection of raw examples.

source_rows() -> Iterable[dict[str, Any]]
validate_written(path: str) -> None

Validate the materialized file format and required columns.

write(path: str) -> None

Materialize training data at path.