Getting your data
Updating code that uses the previous dataset API? See the dataset migration guide for before-and-after examples and changes to evaluation and caching.
Models need data to train on. Data can be either curated using proprietary and/or readily available sources, or synthetically generated by the environment a model trains in, potentially with ground truth data to verify against.
Here, we’ll focus primarily on the former, and in the next guide, we’ll show how you can do the latter.
Hugging Face
Section titled “Hugging Face”The HuggingFaceDataset class handles all the subtle ways in which datasets differ.
For example, statworx/haiku contains a keywords column with only a single word as input, and a text column that contains a ground-truth label. Since the input isn’t in the OpenAI chat completions API format, we need to set input_format to text so that the Modal Dojo can format each prompt in the dataset as a single user message.
from modal_dojo import HuggingFaceDataset
haiku_dataset = HuggingFaceDataset( "statworx/haiku", input_column="keywords", output_column="text", input_format="text",)You can customize this by adding a system prompt, or by providing a prompt template to add additional text to the user message:
haiku_dataset = HuggingFaceDataset( # ... input_format="text", system_prompt="You are an expert poet.", prompt_template="Write a haiku about {input}.",)Note that only input is allowed.
Other datasets like zhuzilin/dapo-math-17k are already formatted in the OpenAI chat completions format. For these datasets, set input_format to messages:
math_dataset = HuggingFaceDataset( "zhuzilin/dapo-math-17k", input_column="prompt", output_column="label", input_format="messages",)As any seasoned ML veteran will tell you, we need separate datasets for training and validation/evaluation to properly train a model. This is made easy with HF’s slicing syntax:
train_dataset = HuggingFaceDataset( ..., hf_split="train[:1000]",)eval_dataset = HuggingFaceDataset( ..., hf_split="train[1000:]",)If a dataset is gated, you can create a Modal Secret named huggingface-secret that the Modal Dojo will auto-detect:
modal secret create huggingface-secret HF_TOKEN=hf_...Harbor
Section titled “Harbor”A Harbor dataset provides a series of tasks, each containing instructions, files for the coding environment, and tests.
from modal_dojo import HarborDataset
hello_dataset = HarborDataset(dataset_name="harbor/hello-world")If you have the files installed locally, you can instead use the filepath:
hello_dataset = HarborDataset(path="/path/to/task_root")Some tasks provide additional metadata:
hello_dataset = HarborDataset( # ... label_metadata_path="task.toml", test_data_dir="tests",)During training, each task will be converted into a single user prompt. Like HuggingFaceDataset, you can optionally provide a system prompt or override the template for the user message. However, if the metadata contains variables, you can also use them in prompt_template!
hello_dataset = HarborDataset( # ... system_prompt="You are an expert Python programmer.", prompt_template="Perform task {task_name} stored at {task_path}: {instruction}",)Creating a custom dataset
Section titled “Creating a custom dataset”To use your own data, likely stored in an external source or a Modal Volume, you simply subclass DatasetConfig:
from modal_dojo import DatasetConfig
prompts = [(prompt, label), ...] # external source
class MyCustomDataset(DatasetConfig): def input_key(self) -> str: return "messages"
def label_key(self) -> str: return "label"
def rows(self): for prompt, label in prompts: yield { self.input_key(): [{"role": "user", "content": prompt}], self.label_key(): label, }
dataset = MyCustomDataset()When the environment generates the prompts (i.e., no initial dataset), you must use OnlineRollout:
from modal_dojo import OnlineRollout
recipe = Qwen3_5_4B_Recipe( # ... custom_generate_function=fn, rollout_batch_size=4,)
dataset = OnlineRollout(n_rows=recipe.rollout_batch_size)