Skip to content
Repo

TrainConfig

from modal_dojo import TrainConfig

A dataset, model, and recipe for one training run.

Attributes

dataset DatasetConfig

Training dataset materialized in the framework's /data volume.

model ModelConfig

Model identity and weight download behavior.

recipe SlimeRecipe | MilesRecipe

Training framework, Modal resources, and framework arguments.

eval_dataset DatasetConfig | None

DatasetConfig | None Optional dataset used by the framework's internal evaluation loop. It is materialized independently from the training dataset.

resume_from_checkpoint Checkpoint | None

Start a new run from this Megatron checkpoint's weights with a fresh optimizer and LR schedule; recipe.num_rollout counts from zero. To continue the source run in place with its Adam state and iteration count, leave resume_from_checkpoint unset and point recipe.load at the checkpoint directory.

detach bool

Keep training on Modal if the local train() wait is interrupted. False stops the app. Default: True

group_id str | None

Sweep ID assigned by TrainingGroup.

group_overrides dict[str, Any] | None

Variant overrides keyed by dotted field path.

group_axes list[str] | None

Swept field paths that group dashboard variants.

context_plan_line() -> str | None

Summarize effective context length and parallelism on one line.

framework: Framework
launch(*, show_output: bool = True) -> TrainingRun

Start one training run in a detached Modal app.

Returns

The launched training run.

train(*, show_output: bool = True) -> TrainingRun

Run one training configuration.

Returns

The completed training run.