SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

Agenthon 2026 / Repo guide

Repo guide

Organization & Repository Guide

Start with your track's public repository, install the shared toolkit version named in its instructions, and test your agent locally. This guide explains what each repository provides and what stays private. Check the logged-in website for release status and CodaBench competition links.

How the repositories are laid out

Public practice, sealed exam

The split between a track's two halves is the design decision that matters most. You cannot read the exam.

Public practice repo Private sealed exam repo
Practice tasks you can run Held-out tasks used for the final ranking
Examples, local checks and available baselines Reference solutions, held by the organizers
Track-specific validation and scoring tools The final scorer and ranking logic
Participant docs, templates, and safety checks Audit material and evaluation logs

Practice tasks don't decide the final ranking. Anything published could already have been memorized by a model, so public tasks are for learning and self-testing. The tasks that decide the final standings are unpublished and stay sealed.

One format, one evaluation spine

Tracks share common metadata and validation tools, but their input files, output schemas and command arguments differ. Your track's SUBMISSION_CLI.md describes its interface. Terms are defined in the glossary.

Each unit is checked through the shared admissibility categories before its output earns the track's metric:

g0 integrity → g1 schema → g2 cutoff and resource rules → g3 domain semantics → score

A participant-caused failure remains in the evaluation denominator under the track's failure rule. Your agent runs as a Docker image, but you upload a toolkit-generated ZIP containing its descriptor and team-verification proof. The image is fetched separately. The overview covers what each track asks for, and How evaluation works explains scoring and failures.

How you'll work

Clone your track's public repository and read its README and SUBMISSION_CLI.md. Install the documented toolkit version and the track package needed by its local checks. Run the supplied example or available baseline, then build and validate your own agent.

Examples have different purposes. Track 1 has no official baseline agent; its interface example solves one public exemplar. Track 2's named adapters are scaffolds, not the full models, and local checks without realized targets cannot measure forecast accuracy. Follow each track's instructions rather than assuming every starter kit provides the same kind of baseline or score.

Official submissions are made on CodaBench, not through GitHub. Use one designated CodaBench account for your team across all tracks; the FAQ explains team verification. See the submission format for packaging and the logged-in website for competition access and opening status. Evaluated and validated Development results are published on the public agenthon.net leaderboards. The published phase dates remain on the timeline.

The binding documents are the Official Competition Rules, Terms of Participation, Privacy Notice and Data & Software Licensing Policy. Where a guide and the Rules differ, the Rules govern.