SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

Agenthon 2026 / How evaluation works

Evaluation

How Evaluation Works

Agenthon 2026 uses shared run checks and track-specific scoring. This guide describes Development; its resources and results do not certify Final resource grants or Final standings. Use the current track instructions for details and the logged-in website for submission opening status. The glossary explains the terms.

What a submission is

Your agent runs as a container image: a self-contained package holding your code and everything it needs to run. Each image implements a single command, and which command depends on the track — solve for T1 Coding, forecast for T2 Forecasting, simulate for T3 Simulation, and analyze for T4 Explainability.

All four tracks run without general internet access. Supported Coding, Forecasting and Explainability submissions can call the provided House Nemotron model through a restricted connection, within the published limits. Simulation has no network access. Include all required dependencies and permitted artifacts in your container image; downloads and external API calls are unavailable during evaluation.

Upload the toolkit-generated submission ZIP, containing the descriptor and a verification proof linking it to your agenthon.net team. The descriptor identifies an immutable image digest; the image is pulled separately. The Team Key itself is not included in the ZIP. See the submission format and use the CodaBench competition links provided on the logged-in website.

Admissibility comes first

Each evaluation unit passes through four categories of checks before its output can earn the track's metric. These checks prevent invalid formats, rule violations or incorrect semantics from earning credit.

A participant-caused failure or missing output remains in the evaluation denominator and contributes the track's specified worst value. A failed unit is not silently excluded to improve the submission's average. Failure outcomes use published codes. An organizer or infrastructure fault is handled separately, not automatically counted as a participant failure.

Scores and Development results

Use CodaBench to submit and monitor processing. Evaluated and validated Development results appear on the public leaderboards at agenthon.net, visible to everyone. The four tracks use different metrics and scales, so their numeric scores cannot be compared or added.

Local smoke checks help identify interface and output problems. They do not prove official acceptance, accuracy without reference targets, or performance on private evaluation tasks.

Keeping results honest

Leakage and cheating are handled by design, not by trust. Tasks carry data cutoffs, so that a submission is judged on what it could have known at the time rather than on the answer it is asked to predict. Held-out material stays sealed. Competition material can carry hidden markers, so material that was copied rather than solved becomes detectable. Leading submissions may be rerun before results are final.

Private evaluation tasks, reference answers and organizer evaluation artifacts are not participant downloads. Agents receive only the task inputs required for execution; publication of a team score does not release the private evaluation package. The organizers do not publish every check in detail.

Where to go next. The overview covers the competition's structure and its public and private repositories, and the repo guide walks through what you get to work with.

The binding documents are the Official Competition Rules, Terms of Participation, Privacy Notice and Data & Software Licensing Policy. Where a guide and the Rules differ, the Rules govern.