Executive summary
Competition Overview
Agenthon 2026 is a NeurIPS 2026 Competition Track on verifiable AI for quantitative finance. It asks one question: can an AI agent do real finance work, and can an automated evaluation confirm the answer is right rather than merely plausible? This guide describes Development. Read the current track instructions for its detailed requirements, and check the logged-in website for release status and CodaBench competition links.
What Agenthon 2026 is
Agenthon extends the Alphathon program into its first NeurIPS edition. It keeps Alphathon's four-track structure and its hard finance setting, and adds what a research competition needs to be trusted: sealed held-out data, automated leakage controls, reproducible reruns, and public leaderboards. Leakage means answer material reaching a submission it shouldn't reach, through the data it is given or the data it was trained on.
Every submission makes a claim: this code is correct, this forecast is honest about its uncertainty, this market simulation behaves like the real thing, this prediction rests on the evidence it cites. Agenthon tests the claim before it scores it.
The four tracks
The four tracks run in parallel and share one evaluation spine. Each has its own verb, headline metric, and domain check.
| Track | Verb | Metric | Gate |
|---|---|---|---|
| T1 Coding | solve | pass@1 | pytest + financial invariants |
| T2 Forecasting | forecast | CRPS composite | as-of cutoff + calibration |
| T3 Simulation | simulate | events/sec | semantic regression |
| T4 Explainability | analyze | Track 4 composite | faithfulness + embargo |
The glossary defines the metric and gate terms used here.
T1 Coding You build an agent that solves quantitative-finance coding tasks. Its work is checked by ordinary unit tests and by financial invariants, the identities a correct answer has to satisfy however the code was written. Official scoring uses one execution per task and reports the share solved (pass@1). Optional local pass@3 reports are not the official leaderboard metric.
T2 Forecasting You build an agent that forecasts financial time series from numeric data plus a time-stamped text corpus. The score evaluates the forecast distribution. Understanding whether text improves forecasts is a research aim; there is no separately scored information-uplift or text-ablation component.
T3 Simulation You submit an ABIDES-compatible market simulator that runs faster. Speed counts only if the simulator preserves matching-engine semantics and still reproduces the stylized facts of real markets, the statistical signatures real market data reliably shows. Development results are practice feedback and do not establish controlled Final timing or Final standings.
T4 Explainability You build an agent that predicts labels, values, or rankings for rows of a table, and supports each prediction with citations from a frozen evidence corpus. Follow the current Track 4 guide for prediction, evidence and reasoning evaluation. Prediction intervals in your answer are distinct from uncertainty measures attached to an aggregate leaderboard score.
How a submission is judged
Your agent runs as a Docker container image, a self-contained package holding your program and its dependencies. It implements the command-line verb for its track. You upload a toolkit-generated ZIP identifying that image and proving your team's registration; the ZIP does not contain the image itself. See the submission format for the details.
All four tracks run without general internet access. Supported Coding, Forecasting and Explainability submissions can call the provided House Nemotron model through a restricted connection. Simulation has no network access. Include required dependencies and permitted artifacts in your image before evaluation.
Each unit is checked for integrity, output schema, cutoff and resource rules, and track semantics: the g0 to g3 admissibility gates. Valid outputs receive the track's metric. Participant failures remain in the evaluation denominator at the track's specified worst value; they are not silently dropped. Organizer-side faults are handled separately.
Use CodaBench to upload and monitor processing. Validated Development results are published on the public leaderboards at agenthon.net. Each track has its own score and reporting rules; Track 1 does not publish a confidence interval.
Public practice, sealed exam
Each track ships as a pair of repositories. The answers stay sealed.
| Public practice | Private exam |
|---|---|
| Public practice units | Private-test held-out units |
| Examples, local checks and available baselines | Oracle solutions and final scorer |
| Manifest and canary safety checks | Canary registry and audit logs |
A unit is one item of evaluation: a coding task, a forecast, a simulation scenario, an explainability question. A canary is a marker planted in sealed material so that leaks become detectable. The smoke scorer in the public repo lets you check your own runs against public material before you send a submission in.
The two phases
The competition has two phases: Development and a joint Final + Verification phase. Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE. The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens. Other dates are on the timeline on the main page.
- Development. Build and test against the public starter packages. CodaBench handles submissions and processing; agenthon.net publishes the public Development leaderboards. Follow the logged-in website for submission opening status.
- Final + Verification. One final submission per entered track is evaluated on sealed held-out units. Within this same phase, organizers rerun leading submissions and review reproducibility. Verification does not require a second participant submission.
How evaluation works goes a level deeper on submissions and the gates.