Repo guide
Organization & Repository Guide
Start with your track's public repository, install the shared toolkit version named in its instructions, and test your agent locally. This guide explains what each repository provides and what stays private. Check the logged-in website for release status and CodaBench competition links.
How the repositories are laid out
- Shared public toolkit. The toolkit repository provides common validation, submission packaging and documentation.
- Four public track repositories. Each contains participant instructions, public practice material and track-specific code: T1 Coding, T2 Forecasting, T3 Simulation, and T4 Explainability. Their private evaluation counterparts remain organizer-only.
- Competition website. Agenthon.net provides team registration, participant guidance and public Development leaderboards. CodaBench handles submission uploads and processing status.
Public practice, sealed exam
The split between a track's two halves is the design decision that matters most. You cannot read the exam.
| Public practice repo | Private sealed exam repo |
|---|---|
| Practice tasks you can run | Held-out tasks used for the final ranking |
| Examples, local checks and available baselines | Reference solutions, held by the organizers |
| Track-specific validation and scoring tools | The final scorer and ranking logic |
| Participant docs, templates, and safety checks | Audit material and evaluation logs |
Practice tasks don't decide the final ranking. Anything published could already have been memorized by a model, so public tasks are for learning and self-testing. The tasks that decide the final standings are unpublished and stay sealed.
One format, one evaluation spine
Tracks share common metadata and validation tools, but their input files, output schemas and command arguments differ. Your track's SUBMISSION_CLI.md describes its interface. Terms are defined in the glossary.
Each unit is checked through the shared admissibility categories before its output earns the track's metric:
g0 integrity → g1 schema → g2 cutoff and resource rules → g3 domain semantics → score
A participant-caused failure remains in the evaluation denominator under the track's failure rule. Your agent runs as a Docker image, but you upload a toolkit-generated ZIP containing its descriptor and team-verification proof. The image is fetched separately. The overview covers what each track asks for, and How evaluation works explains scoring and failures.
How you'll work
Clone your track's public repository and read its README and SUBMISSION_CLI.md. Install the documented toolkit version and the track package needed by its local checks. Run the supplied example or available baseline, then build and validate your own agent.
Examples have different purposes. Track 1 has no official baseline agent; its interface example solves one public exemplar. Track 2's named adapters are scaffolds, not the full models, and local checks without realized targets cannot measure forecast accuracy. Follow each track's instructions rather than assuming every starter kit provides the same kind of baseline or score.
Official submissions are made on CodaBench, not through GitHub. Use one designated CodaBench account for your team across all tracks; the FAQ explains team verification. See the submission format for packaging and the logged-in website for competition access and opening status. Evaluated and validated Development results are published on the public agenthon.net leaderboards. The published phase dates remain on the timeline.