Awards
$1,500, $1,000 and $500 in every track.
The first, second and third teams in each of the four tracks win cash awards: $12,000 in total.
Agenthon
NeurIPS 2026 Competition Track
Verifiable AI for quantitative finance.
A four-track competition testing whether AI agents can produce finance outputs that survive automated, leakage-controlled, cheat-resistant evaluation.
Awards
The first, second and third teams in each of the four tracks win cash awards: $12,000 in total.
We're delighted to welcome NVIDIA, Bloomberg and AllianceBernstein as sponsors of Agenthon 2026. Thank you for your support, for helping us organize the competition, and for collaborating with us on its science and task development. Meet our sponsors
Key dates
Registration runs from 17 August to 12 October 2026. Registration and Development close together at 23:59 Anywhere on Earth (AoE, UTC−12) on 12 October. The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens.
Sign up and form your team.
Public practice repositories, iteration, and the validation leaderboard.
Papers for the NeurIPS workshop were due at 23:59 AoE.
One submission per entered track is evaluated on sealed held-out units; organizers rerun leading submissions and review reproducibility within the same phase.
The Agenthon workshop, with final presentations by the winners, in Atlanta, Georgia.
The competition closes with a workshop and final presentations from the winning teams at NeurIPS in Atlanta, Georgia, on Saturday 12 December. All dates are given in the official rules, which govern if anything here is out of step.
Teams
You can compete alone or with up to two teammates. You join one team and stay on it, and that one team can enter as many of the four tracks as it wants.
Registration is per team, and only registered teams can submit solutions or appear on the leaderboard. Your team uses one account on the competition platform, chosen by the team — the FAQ explains what that means for everyone else on the team.
Have a question about entering? Start with the FAQ.
Core question
Agenthon extends the Alphathon program into its first NeurIPS edition. The competition keeps the four-track structure and hard finance setting, while adding sealed held-out data, automated leakage controls, reproducible reruns, and public leaderboards.
Agents in all four tracks run in sandboxed Docker containers without general internet access. Supported Coding, Forecasting and Explainability submissions can call House Nemotron through a restricted connection. Simulation has no network access. Build required dependencies and permitted artifacts into your image before evaluation.
Four tracks
Build a Docker agent that solves quantitative finance coding tasks under pytest and financial-invariant checks.
Forecast future panels using time-series data plus a time-stamped text corpus. The score evaluates the forecast distribution; information uplift is a research aim.
Submit an ABIDES-compatible simulator that is faster while preserving matching-engine semantics and market stylized facts. Development results are practice feedback, not controlled Final timing.
Predict labels, values, or rankings over tabular entities, supported by a frozen evidence corpus. Follow the track guide for prediction, evidence and reasoning evaluation.
Protocol
Each unit is checked for admissibility and evaluated using its track's scoring rules. Participant failures remain in the evaluation denominator; organizer faults are handled separately. If two Final submissions finish a track with the same ranking score, the tie is broken in favour of the one uploaded earlier.
Upload the toolkit-generated ZIP identifying your agent image and proving team registration.
g0-g3 verify integrity, schema, cutoff/resource rules, and domain semantics.
Track-specific scoring produces Development feedback. Track 1 uses pass@1 without a confidence interval.
Public/private firewall
Each track has a public practice repo and a private sealed exam repo. Public repos include practice tasks, examples, local checks, available baselines and participant docs. Private repos hold evaluation tasks, reference answers, scoring material and audit records.
Leaderboard
Development leaderboards on agenthon.net are public and visible to everyone. Each track uses its own scale. Higher is better for the displayed leaderboard scores; the raw Forecasting loss is converted for this display. Use CodaBench for uploads and processing status. Follow the logged-in website for submission opening status. From 5 October 2026, 00:00 AoE (12:00 UTC), the Track 1 board ranks only runs uploaded after 25 September 2026 at 05:00 UTC, for every team (Track 1 README, rules 8 and 9).
The leaderboard is updated periodically from the competition platform.
T1 Coding: teams 1–25 of 94
T1 Coding: teams 26–50 of 94
T1 Coding: teams 51–75 of 94
T1 Coding: teams 76–94 of 94
No team matches that search.
T2 Forecasting: teams 1–25 of 116
T2 Forecasting: teams 26–50 of 116
T2 Forecasting: teams 51–75 of 116
T2 Forecasting: teams 76–100 of 116
T2 Forecasting: teams 101–116 of 116
No team matches that search.
T3 Simulation: teams 1–25 of 41
T3 Simulation: teams 26–41 of 41
No team matches that search.
No team matches that search.
2026 phases
Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE.
Upcoming
Build and test with the public starter packages. Check the logged-in website for submission opening status.
Teams practice and iterate with the public starter packages. Submission opening is announced separately.
One submission per entered track is evaluated on sealed private-test units. Organizers rerun leading submissions and review reproducibility in the same phase. There is no separate verification submission. Run outputs, logs, and score files stay hidden.
Awards
The top three teams in each track win a cash award, decided by the final standings after verification. That is $3,000 per track and $12,000 in total.
In each of the four tracks.
In each of the four tracks.
In each of the four tracks.
Amounts are in US dollars. A team's award is split equally among its registered members. Eligibility, verification, tax and payment requirements are set by the official rules, and a confirmed winner who accepts an award publishes the code needed to reproduce the method.
Scientific outputs
Gate failures are labeled and aggregated to show where finance agents break down.
T2 investigates whether text and reasoning improve forecasts; this is not a separate scored component.
T3 measures throughput only after semantic fidelity and stylized facts survive checks.
T4 requires evidence-backed predictions that do not cite future or unsupported facts.
Call for papers
Accepted papers will be presented as posters at the Agenthon workshop at NeurIPS in Atlanta on Saturday 12 December. The program will follow.
Our Supporters
A number of sponsorship tiers and opportunities are available, including monetary and awards sponsorship, event space, data, infrastructure, and compute.
Agenthon Questions and Tasks are of real relevance to hedge funds and asset managers, spanning alpha forecasting, portfolio optimization, computational statistics, machine learning, and AI. Questions are provided jointly by the SQA, CEWIT and our Question Partners.
See Questions 2025 for last year's questions and a feel for Agenthon priorities.
Organizing Committee
Assistant Professor, Department of Applied Mathematics and Statistics, Stony Brook University
Vice President, Society of Quantitative Analysts
Chief Investment Officer, Atlas Ridge Capital
Adjunct Professor, NYU Courant
Executive Advisory Board, Columbia Business School, Program for Financial Studies
Head of Machine Learning Strategy, CTO Office
Bloomberg, Toronto, Canada
Head of Quant Technology Strategy, Office of the CTO
Bloomberg, New York, USA
Global Head of Capital Markets Strategy
NVIDIA Corporation, USA
T1 · Coding
Zhikang Dong
Track Lead T1
Independent Researcher
T2 · Forecasting
Ruolan Sun
Track Lead T2
Ph.D. Student, Stony Brook University
T3 · Simulation
Haohan Xu
Track Lead T3
Ph.D. Student, Stony Brook University
T4 · Explainability
Mathew Thiel
Track Lead T4
Quant Research Analyst, validityBase