SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

NeurIPS 2026 Competition Track

Agenthon 2026

Verifiable AI for quantitative finance.

A four-track competition testing whether AI agents can produce finance outputs that survive automated, leakage-controlled, cheat-resistant evaluation.

Agenthon 2026 logo
Submission Docker agent
Admissibility g0-g3 gates
Ranking Track metrics

Awards

$1,500, $1,000 and $500 in every track.

The first, second and third teams in each of the four tracks win cash awards: $12,000 in total.

We're delighted to welcome NVIDIA, Bloomberg and AllianceBernstein as sponsors of Agenthon 2026. Thank you for your support, for helping us organize the competition, and for collaborating with us on its science and task development. Meet our sponsors

Key dates

Key dates.

Registration runs from 17 August to 12 October 2026. Registration and Development close together at 23:59 Anywhere on Earth (AoE, UTC−12) on 12 October. The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens.

  1. 17 Aug – 12 Oct

    Registration

    Sign up and form your team.

  2. 28 Aug – 12 Oct

    Development

    Public practice repositories, iteration, and the validation leaderboard.

  3. 30 Sep

    Call for papers closed

    Papers for the NeurIPS workshop were due at 23:59 AoE.

  4. 13 – 25 Oct

    Final + Verification

    One submission per entered track is evaluated on sealed held-out units; organizers rerun leading submissions and review reproducibility within the same phase.

  5. Sat 12 Dec

    Workshop at NeurIPS, Atlanta

    The Agenthon workshop, with final presentations by the winners, in Atlanta, Georgia.

The competition closes with a workshop and final presentations from the winning teams at NeurIPS in Atlanta, Georgia, on Saturday 12 December. All dates are given in the official rules, which govern if anything here is out of step.

Teams

Enter as a team of one to three.

You can compete alone or with up to two teammates. You join one team and stay on it, and that one team can enter as many of the four tracks as it wants.

Registration is per team, and only registered teams can submit solutions or appear on the leaderboard. Your team uses one account on the competition platform, chosen by the team — the FAQ explains what that means for everyone else on the team.

Have a question about entering? Start with the FAQ.

Core question

Can AI agents produce finance answers that can be checked by machine?

Agenthon extends the Alphathon program into its first NeurIPS edition. The competition keeps the four-track structure and hard finance setting, while adding sealed held-out data, automated leakage controls, reproducible reruns, and public leaderboards.

Agents in all four tracks run in sandboxed Docker containers without general internet access. Supported Coding, Forecasting and Explainability submissions can call House Nemotron through a restricted connection. Simulation has no network access. Build required dependencies and permitted artifacts into your image before evaluation.

Four tracks

Same competition spine, four finance problems.

T1 Coding

Quant-finance coding agents

Build a Docker agent that solves quantitative finance coding tasks under pytest and financial-invariant checks.

Verb
solve
Metric
pass@1
Gate
pytest + invariants
T2 Forecasting

Reasoning-augmented time series

Forecast future panels using time-series data plus a time-stamped text corpus. The score evaluates the forecast distribution; information uplift is a research aim.

Verb
forecast
Metric
CRPS composite
Gate
as-of cutoff + calibration
T3 Simulation

Accelerated market simulation

Submit an ABIDES-compatible simulator that is faster while preserving matching-engine semantics and market stylized facts. Development results are practice feedback, not controlled Final timing.

Verb
simulate
Metric
events/sec
Gate
semantic regression
T4 Explainability

Evidence-grounded prediction

Predict labels, values, or rankings over tabular entities, supported by a frozen evidence corpus. Follow the track guide for prediction, evidence and reasoning evaluation.

Verb
analyze
Metric
Track 4 composite
Gate
faithfulness + embargo

Protocol

Submit, check, score.

Each unit is checked for admissibility and evaluated using its track's scoring rules. Participant failures remain in the evaluation denominator; organizer faults are handled separately. If two Final submissions finish a track with the same ranking score, the tie is broken in favour of the one uploaded earlier.

01

Submit

Upload the toolkit-generated ZIP identifying your agent image and proving team registration.

02

Check

g0-g3 verify integrity, schema, cutoff/resource rules, and domain semantics.

03

Score

Track-specific scoring produces Development feedback. Track 1 uses pass@1 without a confidence interval.

Public/private firewall

Transparent grading, sealed answers.

Each track has a public practice repo and a private sealed exam repo. Public repos include practice tasks, examples, local checks, available baselines and participant docs. Private repos hold evaluation tasks, reference answers, scoring material and audit records.

Public practice Private exam
Public practice units Private-test held-out units
Examples, local checks and available baselines Oracle solutions and final scorer
Manifest and canary safety checks Canary registry and audit logs

Leaderboard

Competition scores, one track at a time.

Development leaderboards on agenthon.net are public and visible to everyone. Each track uses its own scale. Higher is better for the displayed leaderboard scores; the raw Forecasting loss is converted for this display. Use CodaBench for uploads and processing status. Follow the logged-in website for submission opening status. From 5 October 2026, 00:00 AoE (12:00 UTC), the Track 1 board ranks only runs uploaded after 25 September 2026 at 05:00 UTC, for every team (Track 1 README, rules 8 and 9).

The leaderboard is updated periodically from the competition platform.

T1 Coding T2 Forecasting T3 Simulation T4 Explainability

T1 Coding

T1 Coding: teams 1–25 of 94

Rank Team Score
1 Kaile 0.7674
2 Saifuddin 0.7209
3 Pluvia 0.5349
4 Proxima 0.5000
5 Arsen Ibragimov 0.4535
6 Chirayu 0.3837
7 Yan Su 0.3605
8 RandomGuyHavingFun 0.3488
9 MindAlchemist 0.2907
10 RizYL 0.2791
11 abc123 0.2674
11 Nana Champions 0.2674
13 dbuddha 0.2558
14 DeepSucc 0.2442
15 Cornfield Chase 0.2326
15 Formaggio 0.2326
17 Autonomous Alpha 0.2209
17 fiftyfifty 0.2209
19 Liqourena 0.2093
19 Tengen Toppa 0.2093
19 Pluto 0.2093
22 Anonymous 0.1977
22 in3lab 0.1977
22 APTAPT 0.1977
22 Quiet Signal 0.1977

T1 Coding: teams 26–50 of 94

Rank Team Score
22 Husky River Trading 0.1977
27 eagle 0.1860
27 Submartingale 0.1860
27 CMCCGDYDDICT 0.1860
27 eoo0m 0.1860
31 Just-in-Time 0.1744
31 VerityLovityFalsityCruelty 0.1744
31 BoGo 0.1744
31 Paragon 0.1744
31 Make It Make Cents 🙏 0.1744
36 Undefined 0.1628
36 TappuKiMKC 0.1628
36 Proof of Alpha 0.1628
36 AbangFrabi 0.1628
36 verifiquant 0.1628
36 FedorAzarov 0.1628
36 Powerhouse 0.1628
36 Eyyyyyy 0.1628
44 Sean 0.1512
44 Four Sigma 0.1512
44 Money Miner 0.1512
44 1991 0.1512
44 Give me a job 0.1512
49 gheware 0.1395
49 Warkop PuteraPrakoso 0.1395

T1 Coding: teams 51–75 of 94

Rank Team Score
49 Masato 0.1395
52 iluss 0.1279
52 Duo_tech 0.1279
52 Jade 0.1279
55 MicroCat 0.1163
55 hnhparitosh 0.1163
55 Agent_007 0.1163
55 always win 0.1163
55 Nietzsche 0.1163
60 TEst 0.1047
61 UNIVERSEPLAYER 0.0930
61 Sir Thaddeus 0.0930
61 HTHY 0.0930
61 Banana Kingdom 0.0930
65 Agent Alpha 0.0814
66 The Unscheduled Penguins 0.0698
67 Apex 0.0581
68 GAOKUO 0.0465
69 attention is qkv 0.0349
70 volatility_labs_stony_brook 0.0233
71 Urithiru 0.0116
— Aamir Abdul Azeez — not scored
— Jin & Pei — not scored
— AgentReady — not scored
— 3226 — not scored

T1 Coding: teams 76–94 of 94

Rank Team Score
— AlbertQuant — not scored
— Verifiable Capital — not scored
— a-team — not scored
— FAWDA — not scored
— 9:00pm — not scored
— C_cake — not scored
— huhudawang — not scored
— Quant Science Flow — not scored
— phtree — not scored
— DaoyuShu4094 — not scored
— Future Ability — not scored
— Echoing — not scored
— Luca — not scored
— Artifical General Invesment — not scored
— Flowfront — not scored
— nogugu — not scored
— Outsiders — not scored
— Wealth-Building Lab — not scored
— IMAnonymous — not scored

No team matches that search.

T2 Forecasting

T2 Forecasting: teams 1–25 of 116

Rank Team Score
1 Yan Su -1.3403
2 DuML -1.4388
3 Ywin -1.5483
4 Flowfront -1.5548
5 huhudawang -1.5682
6 alpha-extractor -1.5725
7 Made in Heaven -1.6351
8 Samoyed -1.6775
9 Chirayu -1.7656
10 Not Today -1.7697
11 Quiet Signal -1.7917
12 HBIT -1.8358
13 F1_1 -1.8367
14 Jin & Pei -1.8424
15 Proof of Alpha -1.8515
16 labubu666 -1.8693
17 CMCCGDYDDICT -1.8758
18 Pawsibble -1.9221
19 Apex -1.9618
20 HackStreet Boys -1.9722
21 Nexus of f(x) -2.0134
22 Warkop PuteraPrakoso -2.0303
23 mikelou1 -2.0579
24 Just-in-Time -2.0898
25 VerityLovityFalsityCruelty -2.0903

T2 Forecasting: teams 26–50 of 116

Rank Team Score
26 DKYnumber1 -2.1028
27 Autonomous Alpha -2.1097
28 Jigglyduck -2.1111
29 Paragon -2.1128
30 Super Debugging -2.1651
31 Saifuddin -2.1769
32 spidy -2.1999
33 Sapien -2.2129
34 urayaha -2.2155
35 lifeisluck -2.2172
36 TradeFlare -2.2310
37 LastDigitsOfPi -2.2338
38 Masato -2.2448
39 HG Market System -2.2556
40 Cornfield Chase -2.2726
41 quack -2.2893
42 Sitadel Insecurities -2.2962
43 JJ LAB -2.3146
44 Formaggio -2.3310
45 iluss -2.3557
46 Money Miner -2.3569
47 abc123 -2.3662
48 5 o'clock -2.3690
49 time -2.3694
50 fiftyfifty -2.3772

T2 Forecasting: teams 51–75 of 116

Rank Team Score
51 pranshu rastogi -2.3774
52 1991 -2.3901
53 AsOf Quant -2.3990
54 Tengen Toppa -2.4107
55 Bayes Street -2.4108
56 Auror -2.4301
57 QIE DOG -2.4329
58 Javier Amo -2.4429
59 Dog parents -2.4467
60 lingsio -2.4480
61 hnhparitosh -2.4484
62 C_cake -2.4513
63 in3lab -2.4570
64 randomwalk -2.4666
65 LiLaiLai -2.4670
66 DoublePai -2.4681
67 Aamir Abdul Azeez -2.4722
68 ZZKK -2.4795
69 In Sync -2.4848
70 Natpaphon -2.4891
71 Verifiable Capital -2.4918
72 GAOKUO -2.4932
73 SMTM -2.4938
74 krnb -2.4961
75 LIZARD -2.5086

T2 Forecasting: teams 76–100 of 116

Rank Team Score
76 1111 -2.5090
77 PSS_ -2.5096
78 Garros Solo Lab -2.5139
79 GXR123 -2.5169
80 Nietzsche -2.5272
81 Yongchang -2.5273
82 Try hard -2.5347
83 AbangFrabi -2.5381
84 Banana Kingdom -2.5388
85 Jade -2.5458
86 JiaxinLi -2.5482
87 pluuuuuus -2.5521
88 Husky River Trading -2.5567
89 Probably Right -2.5643
90 Liqourena -2.5769
91 Launchlane Forecast -2.5799
92 iKun-XinglinForrest -2.5994
93 TheAxiom -2.5999
94 YIC Quant -2.6067
95 Luca -2.6137
96 lwz888 -2.6141
97 Undefined -2.6151
98 ARKANE -2.6357
99 OOBs -2.6373
— Reference forecaster M0 reference, not ranked -2.6412
100 SoloTraveler -2.6765

T2 Forecasting: teams 101–116 of 116

Rank Team Score
101 ah_a1g -2.6829
101 Give me a job -2.6829
103 OffTheTape -2.7369
104 BigWorld -2.7476
105 katharsis -2.7795
106 K-uant -2.7985
107 PA-Agent -2.8354
— 9:00pm — not scored
— ZOLDIK — not scored
— Urithiru — not scored
— Quant Science Flow — not scored
— phtree — not scored
— Hatenabase — not scored
— Edge-X — not scored
— Artifical General Invesment — not scored
— Outsiders — not scored

No team matches that search.

T3 Simulation

T3 Simulation: teams 1–25 of 41

Rank Team Score
1 fuzz 231024.2407
2 abc123 221088.2954
3 IMAnonymous 215963.1445
4 Proof of Alpha 211712.3699
5 Cornfield Chase 210201.3605
6 Yan Su 210016.3629
7 Warkop PuteraPrakoso 201094.4945
8 VerityLovityFalsityCruelty 198767.4670
9 DKYnumber1 196323.2757
10 iMak AI Lab 190224.1037
11 S2WISH 189138.3673
12 Autonomous Alpha 185936.6020
13 Lumia 185889.4276
14 Paragon 184046.5937
15 Made in Heaven 181871.6467
16 Quantumonster 179955.5150
17 Money Miner 179846.1354
18 Verifiable Capital 178103.6278
19 PA-Agent 177343.4124
20 Jin & Pei 176009.6963
21 Yongchang 161575.8191
22 Probably Right 157361.2474
23 HackStreet Boys 156661.8792
24 fiftyfifty 152138.2119
25 huhudawang 134182.0441

T3 Simulation: teams 26–41 of 41

Rank Team Score
26 chilli 110245.5255
27 Saifuddin 84836.4432
28 PumpkinChicken 20625.2674
29 1991 17167.8369
30 Apex 16467.6082
31 CMCCGDYDDICT 15559.8991
32 Banana Kingdom 14180.6623
33 Just-in-Time 13813.8474
34 Quant Science Flow 13797.7076
35 Auror 13708.7408
36 Outsiders 12116.2977
— TEst — not scored
— thalachira — not scored
— Hertz — not scored
— Artifical General Invesment — not scored
— Sequoia — not scored

No team matches that search.

T4 Explainability

Rank Team Score
1 Genshin, launch! 0.8401
2 Just-in-Time 0.5787
3 Saifuddin 0.5234
4 abc123 0.5058
5 MALIU 0.5025
6 Paragon 0.4920
7 stone stone 0.4902
8 Sitadel Insecurities 0.4797
9 Quiet Signal 0.4683
10 huhudawang 0.4512
11 OmniSync 0.4365
12 ccczy 0.4029
13 FAWDA 0.4001
14 SignalCraft 0.3993
15 Apex 0.3954
16 Flowfront 0.3400
17 LOOOONG 0.1475
— yohoho — not scored
— Chaewon-Research — not scored
— Northline — not scored

No team matches that search.

2026 phases

Development, then Final + Verification.

Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE.

Upcoming

Development

Build and test with the public starter packages. Check the logged-in website for submission opening status.

0% complete
  1. Upcoming Aug 28 - Oct 12

    Development

    Teams practice and iterate with the public starter packages. Submission opening is announced separately.

  2. Upcoming Oct 13 - Oct 25

    Final + Verification

    One submission per entered track is evaluated on sealed private-test units. Organizers rerun leading submissions and review reproducibility in the same phase. There is no separate verification submission. Run outputs, logs, and score files stay hidden.

Awards

Cash awards in all four tracks.

The top three teams in each track win a cash award, decided by the final standings after verification. That is $3,000 per track and $12,000 in total.

  1. 1st place

    $1,500

    In each of the four tracks.

  2. 2nd place

    $1,000

    In each of the four tracks.

  3. 3rd place

    $500

    In each of the four tracks.

Amounts are in US dollars. A team's award is split equally among its registered members. Eligibility, verification, tax and payment requirements are set by the official rules, and a confirmed winner who accepts an award publishes the code needed to reproduce the method.

Scientific outputs

Agenthon is designed to explain failures, not just rank winners.

Cross-track failure map

Gate failures are labeled and aggregated to show where finance agents break down.

Information uplift

T2 investigates whether text and reasoning improve forecasts; this is not a separate scored component.

Speed-realism frontier

T3 measures throughput only after semantic fidelity and stylized facts survive checks.

Faithfulness under embargo

T4 requires evidence-backed predictions that do not cite future or unsupported facts.

Call for papers

The call for papers is closed.

Accepted papers will be presented as posters at the Agenthon workshop at NeurIPS in Atlanta on Saturday 12 December. The program will follow.

Our Supporters

Sponsor Agenthon 2026.

A number of sponsorship tiers and opportunities are available, including monetary and awards sponsorship, event space, data, infrastructure, and compute.

Agenthon Questions and Tasks are of real relevance to hedge funds and asset managers, spanning alpha forecasting, portfolio optimization, computational statistics, machine learning, and AI. Questions are provided jointly by the SQA, CEWIT and our Question Partners.

See Questions 2025 for last year's questions and a feel for Agenthon priorities.

Organizing Committee

Thank you from the organizers.

Lead Organizers

Pawel Polak

Assistant Professor, Department of Applied Mathematics and Statistics, Stony Brook University
Vice President, Society of Quantitative Analysts

Website · LinkedIn

Christos Koutsoyannis

Chief Investment Officer, Atlas Ridge Capital
Adjunct Professor, NYU Courant
Executive Advisory Board, Columbia Business School, Program for Financial Studies

Website · LinkedIn

Industry Co-Organizers

David Rosenberg

Head of Machine Learning Strategy, CTO Office
Bloomberg, Toronto, Canada

LinkedIn

Gary Kazantsev

Head of Quant Technology Strategy, Office of the CTO
Bloomberg, New York, USA

LinkedIn

Ioana Boier

Global Head of Capital Markets Strategy
NVIDIA Corporation, USA

LinkedIn

Track Leads

T1 · Coding

Quant-finance coding agents

Zhikang Dong
Track Lead T1
Independent Researcher

LinkedIn · GitHub

T2 · Forecasting

Reasoning-augmented time series

Ruolan Sun
Track Lead T2
Ph.D. Student, Stony Brook University

LinkedIn · GitHub

T3 · Simulation

Accelerated market simulation

Haohan Xu
Track Lead T3
Ph.D. Student, Stony Brook University

LinkedIn · GitHub

T4 · Explainability

Evidence-grounded prediction

Mathew Thiel
Track Lead T4
Quant Research Analyst, validityBase

LinkedIn · GitHub