TrainingRun
from modal_dojo import TrainingRunA launched training run that can be inspected, awaited, or loaded by ID.
Attributes
training_run_id str
Stable id for this run in the metadata volume.
modal_app_id str
Modal app id after the run is spawned. Default: ""
modal_app_url str
Dashboard URL for that Modal app. Default: ""
framework Framework
Training framework that executed the run.
config Any
Serialized train config captured at launch.
dataset_id str
Dataset id materialized for this run. Default: ""
deployment_id str
Linked deployment id, when the run served a model. Default: ""
status TrainingRunStatus
Modal Dojo-level run status. Default: running
framework_status SlimeStatus | MilesStatus | None
Latest phase reported by the framework.
created_at int
Unix time the record was created. Default: 0
started_at int
Unix time training started. Default: 0
ended_at int | None
Unix time the process exited.
completed_at int | None
Unix time the run reached a terminal success.
updated_at int
Unix time the record was last written. Default: 0
duration_seconds int | None
Wall time from start to end, when known.
step_times dict[str, dict[str, int | None]] | None
Per-step timings, when present.
substep_times dict[str, dict[str, dict[str, float | int | bool | None]]] | None
Per-substep timings, when present.
error_message str | None
Failure text. Empty while running or after success.
metadata dict[str, Any] | None
Extra keys such as sweep group_id.
function_call_id str
Modal FunctionCall id used to wait on the spawned run. Default: ""
app_name str
Modal app name used as the default checkpoint volume prefix. Default: ""
source_model Any
Model used for training.
metrics dict[str, Any]
Final scalar metrics written at the end of the run. Default: {}
add_dashboard_component
Section titled “add_dashboard_component”add_dashboard_component(*, name: str, component_type: 'DashboardComponent | str', from_path: str | os.PathLike[str], replace: bool = False) -> dict[str, Any]Attach a local dashboard component to this run.
The source is uploaded to the shared dashboard-overlay Volume and the resulting immutable manifest is associated with this run’s metadata. This operation is independent of the training job and may be called after a run has started or completed.
apply_framework_status
Section titled “apply_framework_status”apply_framework_status(update: FrameworkStatusUpdate) -> FrameworkStatus | NoneApply a framework status update without saving the run.
Returns
The resolved status, or None for an invalid phase.
checkpoint_dir
Section titled “checkpoint_dir”checkpoint_dir: strcheckpoints
Section titled “checkpoints”checkpoints() -> list[Checkpoint]Snapshot of committed megatron iter_* directories.
Miles LoRA adapters are not listed. Serving needs
convert_megatron_checkpoint_to_hf.
close() -> NoneStop the detached Modal app. Safe to call more than once.
done() -> boolTrue if status is not RUNNING or the FunctionCall has finished.
Does not stop the Modal app.
error: str | Nonefrom_id
Section titled “from_id”from_id(run_id: str, *, is_async: bool = False) -> TrainingRun | Awaitable[TrainingRun]from_stored_data
Section titled “from_stored_data”from_stored_data(data: object) -> TrainingRunBuild a run from persisted data and snapshot its metadata keys.
function_call
Section titled “function_call”function_call: Anygroup_id
Section titled “group_id”group_id: str | NoneThe sweep ID stored in metadata.
latest_checkpoint
Section titled “latest_checkpoint”latest_checkpoint() -> 'Checkpoint | None'model: ModelConfigBuild a ModelConfig whose path targets the latest megatron checkpoint.
record_latest_rollout
Section titled “record_latest_rollout”record_latest_rollout(rollout: TrainingRolloutResult) -> NoneStore the saved rollout summary in this run’s metadata.
result
Section titled “result”result(*, timeout: float | None = None, stop_app_on_success: bool = True) -> TrainingRunsave(*, is_async: bool = False) -> None | Awaitable[None]wait(*, timeout: float | None = None) -> TrainingRunBlock until this run is done. Does not stop the Modal app.
wait_all
Section titled “wait_all”wait_all(runs: Sequence[TrainingRun], *, poll_interval: float = 30) -> list[TrainingRun]Wait for every run. Close each run when that run is done.