SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

Agenthon 2026 / Competition overview

Executive summary

Competition Overview

Agenthon 2026 is a NeurIPS 2026 Competition Track on verifiable AI for quantitative finance. It asks one question: can an AI agent do real finance work, and can an automated evaluation confirm the answer is right rather than merely plausible? This guide describes Development. Read the current track instructions for its detailed requirements, and check the logged-in website for release status and CodaBench competition links.

What Agenthon 2026 is

Agenthon extends the Alphathon program into its first NeurIPS edition. It keeps Alphathon's four-track structure and its hard finance setting, and adds what a research competition needs to be trusted: sealed held-out data, automated leakage controls, reproducible reruns, and public leaderboards. Leakage means answer material reaching a submission it shouldn't reach, through the data it is given or the data it was trained on.

Every submission makes a claim: this code is correct, this forecast is honest about its uncertainty, this market simulation behaves like the real thing, this prediction rests on the evidence it cites. Agenthon tests the claim before it scores it.

The four tracks

The four tracks run in parallel and share one evaluation spine. Each has its own verb, headline metric, and domain check.

Track Verb Metric Gate
T1 Coding solve pass@1 pytest + financial invariants
T2 Forecasting forecast CRPS composite as-of cutoff + calibration
T3 Simulation simulate events/sec semantic regression
T4 Explainability analyze Track 4 composite faithfulness + embargo

The glossary defines the metric and gate terms used here.

T1 Coding You build an agent that solves quantitative-finance coding tasks. Its work is checked by ordinary unit tests and by financial invariants, the identities a correct answer has to satisfy however the code was written. Official scoring uses one execution per task and reports the share solved (pass@1). Optional local pass@3 reports are not the official leaderboard metric.

T2 Forecasting You build an agent that forecasts financial time series from numeric data plus a time-stamped text corpus. The score evaluates the forecast distribution. Understanding whether text improves forecasts is a research aim; there is no separately scored information-uplift or text-ablation component.

T3 Simulation You submit an ABIDES-compatible market simulator that runs faster. Speed counts only if the simulator preserves matching-engine semantics and still reproduces the stylized facts of real markets, the statistical signatures real market data reliably shows. Development results are practice feedback and do not establish controlled Final timing or Final standings.

T4 Explainability You build an agent that predicts labels, values, or rankings for rows of a table, and supports each prediction with citations from a frozen evidence corpus. Follow the current Track 4 guide for prediction, evidence and reasoning evaluation. Prediction intervals in your answer are distinct from uncertainty measures attached to an aggregate leaderboard score.

How a submission is judged

Your agent runs as a Docker container image, a self-contained package holding your program and its dependencies. It implements the command-line verb for its track. You upload a toolkit-generated ZIP identifying that image and proving your team's registration; the ZIP does not contain the image itself. See the submission format for the details.

All four tracks run without general internet access. Supported Coding, Forecasting and Explainability submissions can call the provided House Nemotron model through a restricted connection. Simulation has no network access. Include required dependencies and permitted artifacts in your image before evaluation.

Each unit is checked for integrity, output schema, cutoff and resource rules, and track semantics: the g0 to g3 admissibility gates. Valid outputs receive the track's metric. Participant failures remain in the evaluation denominator at the track's specified worst value; they are not silently dropped. Organizer-side faults are handled separately.

Use CodaBench to upload and monitor processing. Validated Development results are published on the public leaderboards at agenthon.net. Each track has its own score and reporting rules; Track 1 does not publish a confidence interval.

Public practice, sealed exam

Each track ships as a pair of repositories. The answers stay sealed.

Public practice Private exam
Public practice units Private-test held-out units
Examples, local checks and available baselines Oracle solutions and final scorer
Manifest and canary safety checks Canary registry and audit logs

A unit is one item of evaluation: a coding task, a forecast, a simulation scenario, an explainability question. A canary is a marker planted in sealed material so that leaks become detectable. The smoke scorer in the public repo lets you check your own runs against public material before you send a submission in.

The two phases

The competition has two phases: Development and a joint Final + Verification phase. Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE. The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens. Other dates are on the timeline on the main page.

How evaluation works goes a level deeper on submissions and the gates.

The binding documents are the Official Competition Rules, Terms of Participation, Privacy Notice and Data & Software Licensing Policy. Where a guide and the Rules differ, the Rules govern.