Skip to content
Repo

TrainingRun

from modal_dojo import TrainingRun

A launched training run that can be inspected, awaited, or loaded by ID.

Attributes

training_run_id str

Stable id for this run in the metadata volume.

modal_app_id str

Modal app id after the run is spawned. Default: ""

modal_app_url str

Dashboard URL for that Modal app. Default: ""

framework Framework

Training framework that executed the run.

config Any

Serialized train config captured at launch.

dataset_id str

Dataset id materialized for this run. Default: ""

deployment_id str

Linked deployment id, when the run served a model. Default: ""

status TrainingRunStatus

Modal Dojo-level run status. Default: running

framework_status SlimeStatus | MilesStatus | None

Latest phase reported by the framework.

created_at int

Unix time the record was created. Default: 0

started_at int

Unix time training started. Default: 0

ended_at int | None

Unix time the process exited.

completed_at int | None

Unix time the run reached a terminal success.

updated_at int

Unix time the record was last written. Default: 0

duration_seconds int | None

Wall time from start to end, when known.

step_times dict[str, dict[str, int | None]] | None

Per-step timings, when present.

substep_times dict[str, dict[str, dict[str, float | int | bool | None]]] | None

Per-substep timings, when present.

error_message str | None

Failure text. Empty while running or after success.

metadata dict[str, Any] | None

Extra keys such as sweep group_id.

function_call_id str

Modal FunctionCall id used to wait on the spawned run. Default: ""

app_name str

Modal app name used as the default checkpoint volume prefix. Default: ""

source_model Any

Model used for training.

metrics dict[str, Any]

Final scalar metrics written at the end of the run. Default: {}

add_dashboard_component(*, name: str, component_type: 'DashboardComponent | str', from_path: str | os.PathLike[str], replace: bool = False) -> dict[str, Any]

Attach a local dashboard component to this run.

The source is uploaded to the shared dashboard-overlay Volume and the resulting immutable manifest is associated with this run’s metadata. This operation is independent of the training job and may be called after a run has started or completed.

apply_framework_status(update: FrameworkStatusUpdate) -> FrameworkStatus | None

Apply a framework status update without saving the run.

Returns

The resolved status, or None for an invalid phase.

checkpoint_dir: str
checkpoints() -> list[Checkpoint]

Snapshot of committed megatron iter_* directories.

Miles LoRA adapters are not listed. Serving needs convert_megatron_checkpoint_to_hf.

close() -> None

Stop the detached Modal app. Safe to call more than once.

done() -> bool

True if status is not RUNNING or the FunctionCall has finished.

Does not stop the Modal app.

error: str | None
from_id(run_id: str, *, is_async: bool = False) -> TrainingRun | Awaitable[TrainingRun]
from_stored_data(data: object) -> TrainingRun

Build a run from persisted data and snapshot its metadata keys.

function_call: Any
group_id: str | None

The sweep ID stored in metadata.

latest_checkpoint() -> 'Checkpoint | None'
model: ModelConfig

Build a ModelConfig whose path targets the latest megatron checkpoint.

record_latest_rollout(rollout: TrainingRolloutResult) -> None

Store the saved rollout summary in this run’s metadata.

result(*, timeout: float | None = None, stop_app_on_success: bool = True) -> TrainingRun
save(*, is_async: bool = False) -> None | Awaitable[None]
wait(*, timeout: float | None = None) -> TrainingRun

Block until this run is done. Does not stop the Modal app.

wait_all(runs: Sequence[TrainingRun], *, poll_interval: float = 30) -> list[TrainingRun]

Wait for every run. Close each run when that run is done.