Task Index
The parquet table of pointers a training run samples from
The trainer never reads task content. It samples rows from a parquet table in which every row is a pointer.
- Input — the task directories.
- Output — a parquet table, one pointer row per task.
# one index row (built by utils/create_task_index.py)
{
"prompt": [{"role": "user", "content": "/data/openswe_filtered/astropy__astropy-15883"}],
"reward_model": {"style": "rule", "ground_truth": None}, # the verifier decides
"extra_info": {
"harbor_task_path": "/data/openswe_filtered/astropy__astropy-15883",
"instance_id": "astropy__astropy-15883",
"data_source": "harbor",
},
}python utils/create_task_index.py \
--tasks_dir /data/openswe_filtered \
--output /data/harbor_indexes/train_openswe_filtered.parquet \
--split trainThe split pays off in the stage below. An index costs kilobytes per thousand tasks, so filtering to a difficulty band, dropping contaminated instances, and unioning subsets across sources are all plain dataframe operations. No task content is copied or rewritten.