Evaluation
How Evaluation Works
Agenthon 2026 uses shared run checks and track-specific scoring. This guide describes Development; its resources and results do not certify Final resource grants or Final standings. Use the current track instructions for details and the logged-in website for submission opening status. The glossary explains the terms.
What a submission is
Your agent runs as a container image: a self-contained package holding your code and everything it needs to run. Each image implements a single command, and which command depends on the track — solve for T1 Coding, forecast for T2 Forecasting, simulate for T3 Simulation, and analyze for T4 Explainability.
All four tracks run without general internet access. Supported Coding, Forecasting and Explainability submissions can call the provided House Nemotron model through a restricted connection, within the published limits. Simulation has no network access. Include all required dependencies and permitted artifacts in your container image; downloads and external API calls are unavailable during evaluation.
Upload the toolkit-generated submission ZIP, containing the descriptor and a verification proof linking it to your agenthon.net team. The descriptor identifies an immutable image digest; the image is pulled separately. The Team Key itself is not included in the ZIP. See the submission format and use the CodaBench competition links provided on the logged-in website.
Admissibility comes first
Each evaluation unit passes through four categories of checks before its output can earn the track's metric. These checks prevent invalid formats, rule violations or incorrect semantics from earning credit.
- g0 integrity — what ran is what you submitted, on the data the task intended.
- g1 schema — the output has the form the track requires. Malformed output follows the track's participant-failure rule.
- g2 cutoff and resources — the run respected the task's data cutoff and stayed inside its resource limits.
- g3 domain semantics — the output makes sense as finance, not merely as data of the right shape.
A participant-caused failure or missing output remains in the evaluation denominator and contributes the track's specified worst value. A failed unit is not silently excluded to improve the submission's average. Failure outcomes use published codes. An organizer or infrastructure fault is handled separately, not automatically counted as a participant failure.
Scores and Development results
Use CodaBench to submit and monitor processing. Evaluated and validated Development results appear on the public leaderboards at agenthon.net, visible to everyone. The four tracks use different metrics and scales, so their numeric scores cannot be compared or added.
- Coding: official pass@1 is the share of tasks solved in one execution per task, with a fixed denominator and no confidence interval. Optional local pass@3 reports are development diagnostics.
- Forecasting: the raw normalized forecast loss is lower-is-better. The website's displayed leaderboard score is converted so higher is better. There is no separately scored text-ablation or information-uplift component.
- Simulation: Development throughput is practice feedback, based on submitted timing that is checked for consistency. It does not establish controlled Final timing or Final standings.
- Explainability: follow the current Track 4 guide for prediction, evidence and reasoning evaluation. Do not confuse prediction intervals in an answer with confidence intervals for an aggregate leaderboard score.
Local smoke checks help identify interface and output problems. They do not prove official acceptance, accuracy without reference targets, or performance on private evaluation tasks.
Keeping results honest
Leakage and cheating are handled by design, not by trust. Tasks carry data cutoffs, so that a submission is judged on what it could have known at the time rather than on the answer it is asked to predict. Held-out material stays sealed. Competition material can carry hidden markers, so material that was copied rather than solved becomes detectable. Leading submissions may be rerun before results are final.
Private evaluation tasks, reference answers and organizer evaluation artifacts are not participant downloads. Agents receive only the task inputs required for execution; publication of a team score does not release the private evaluation package. The organizers do not publish every check in detail.
Where to go next. The overview covers the competition's structure and its public and private repositories, and the repo guide walks through what you get to work with.