Short definitions of the words used across this site and in the competition materials.
These definitions explain the Development workflow; the track instructions and governing
competition policies provide the precise requirements. Submission availability is
announced separately in your signed-in agenthon.net account.
The competition and its pieces
Competition and benchmark
A benchmark is a set of tasks with agreed rules for scoring them. A competition
adds teams, deadlines, phases, and a ranking. Agenthon 2026 is both: a benchmark for
verifiable AI in quantitative finance, run as a NeurIPS 2026 competition.
Track
One of the four contests inside Agenthon 2026: coding, forecasting, simulation, and
explainability. Each track has its own tasks, its own headline metric, and its own
pair of repositories.
Phase
The competition has two phases: Development and a joint Final + Verification phase.
Development supports local practice and hosted evaluation with provisional results.
The joint phase evaluates one final submission per entered track. Within the joint
phase, organizers rerun leading submissions and review reproducibility before results
are final. There is no separate verification submission.
Development runs through 12 October 2026. The joint Final + Verification phase runs
from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE. Other dates live on the
timeline on the main page.
Unit
One item that gets scored on its own: a single coding task, a single forecast, a single
simulation scenario, a single analysis question.
Task card
The short description file that travels with every unit. It tells you what the task
asks for. You read task cards; you never edit them.
Split (public-dev, validation, private-test)
A label for a dataset's role in development or evaluation. Public practice material
lets you build and test locally. Hosted evaluation uses organizer-controlled packages;
a split label does not mean that its tasks or answers are public. Private evaluation
tasks and reference answers are not published with the results.
Baseline
An example method used as a point of comparison. Starter packages differ: some
include runnable baselines; others provide interface examples and checks. A local
baseline score is not an official evaluation score.
Repos, submissions, and the sandbox
Public repo and private repo
The public repository is your track's practice kit: its interface, examples, local checks
and available baselines. Organizer repositories hold private evaluation material and reference
answers. Participants do not need access to those repositories to submit.
Firewall
The separation between public practice material and private evaluation tasks,
reference answers and audit data. Review and automated checks help prevent private
material from being published with participant packages or results.
Docker image
A packaged program with its dependencies and permitted artifacts. Official submissions
reference a Linux/amd64 image by immutable digest. The backend pulls and runs those
image bytes under the track's resource and network rules; identical image bytes do not
imply identical hardware or timing on every machine.
Submission
The ZIP uploaded to CodaBench, containing a descriptor and generated team proof. It references the container image that will run; it does not contain the container layers or Team Key.
Descriptor
submission.json: structured information identifying the team, track, phase, image and declared models. The toolkit validates and seals it when packaging.
Image digest
An immutable identifier for specific container image bytes. Rebuilding an image requires updating its reference and repacking the submission.
Team Number
The numeric team identifier issued on agenthon.net. It is different from the descriptor’s derived Team ID.
Team Key
A confidential team credential used locally by the toolkit to create verification proofs. It is entered at a hidden prompt and is not put in the submission ZIP. Keep it private.
Team ID
The identifier the toolkit derives from your Team Number and Team Key for the submission descriptor. Use the alias command to obtain it; do not leave an example ID in your descriptor.
Team proof
Verification data bound to a particular descriptor that demonstrates possession of the Team Key without including the key. The first valid proof links the submitting CodaBench account to the registered team; keep the packaged proof private.
CodaBench
The platform used to request entry to a track, upload packaged submissions and follow processing status. Public Development leaderboard results appear on agenthon.net.
House model
The organizer-provided Nemotron service that supported Coding, Forecasting and Explainability agents can call through the restricted evaluation connection. Its request budget is separate from container runtime limits.
Submission verb
The single command word the organizers use to start your image: solve for
coding, forecast for forecasting, simulate for simulation,
analyze for explainability. The verb is stable.
Sandbox
The controlled environment in which a submission runs. All four tracks have no
general internet access. Supported Coding, Forecasting and Explainability submissions
can call the House Nemotron service through a restricted connection. Simulation has no
network access. Dependencies and permitted artifacts must be included before the run.
Smoke scorer
A scorer shipped in the public repo so you can grade your own runs before you submit.
It grades against public material, so treat it as a check on whether your output is well
formed and roughly on track, not as a preview of your final standing.
Gates, scoring, and staying honest
Admissibility gate
A check on integrity, output format, resource use or task semantics. Participant
failures follow the track's scoring rules: a failed unit contributes the track's
precommitted worst value rather than disappearing from the denominator. Organizer
faults are a separate case and require review rather than a participant score.
The gates g0 to g3
The four gates run in order. g0 covers integrity: the files and the image are what they
claim to be. g1 covers schema: the output has the shape the track asked for. g2 covers
cutoff and resource rules: no future data, and the run stayed inside the limits set for
everyone. g3 covers domain semantics: the output makes sense for the subject matter,
which is specific to each track.
Organizer fault
An infrastructure or evaluation failure requiring organizer review rather than a
participant performance score. An absent score does not automatically mean zero.
Development leaderboard
The public boards on agenthon.net showing
eligible team scores for each track. They update periodically and are visible to
everyone. Development results are provisional and are separate from Final rankings.
Ranking score
The value used to order a track's leaderboard. It may be direction-adjusted from a
raw metric such as a lower-is-better forecasting loss. Scores from different tracks are
not directly comparable.
Confidence interval
An uncertainty estimate accompanying a score where the track's scoring contract
provides one. It may be computed by resampling units, called the bootstrap. Confidence
intervals are not promised for every track or displayed universally on the website.
Reproducibility
Run the same submission again and the result should hold up. During the joint
Final + Verification phase, organizers rerun leading submissions to check that their
scores hold up.
Data cutoff (embargo)
The information boundary specified by a task or track. Participant data and derived
artifacts must respect the applicable cutoff and published exceptions. For example,
using the future outcome itself is not a valid forecast.
Leakage
Using information that was not available at the cutoff, by accident or on purpose. A
common example is citing a document that was published after the question's cutoff
date.
Canary
A hidden marker placed in competition material. If it later turns up in a submission's
output, that points to material being copied rather than solved. It is how
contamination gets detected after the fact.
Headline metrics
pass@1 and pass@k
pass@1 measures the fraction of tasks passed with one execution per task and is the
hosted Coding Development score. A local pass@k experiment measures success across
multiple attempts; it must not be confused with the hosted one-execution score.
CRPS
The continuous ranked probability score grades a probabilistic forecast, meaning a
forecast that gives a range of outcomes with probabilities attached rather than a single
number. It rewards landing close to what happened and being honest about how uncertain
you were. Lower is better, and the forecasting track's headline metric is built on
it.
Events per second
A simulation performance measurement: how many events a simulator processes per
second. Performance is meaningful only alongside correctness and semantic checks.
Development feedback is provisional; shared Development conditions do not certify
comparable Final timing.
Coverage
How much of a required target is covered, as defined by the metric. Forecast interval
coverage asks how often an outcome lies within a predicted range. Evidence or reasoning
coverage concerns required support or arguments. These are different measures; use the
definition in the relevant track's scoring guide.
Faithfulness
Whether the claims in an answer are supported by the sources it cites. This is
distinct from whether the answer is correct or its reasoning addresses the question;
follow the Explainability track's published scoring guidance for those components.
Information uplift
A comparison between a system using additional information, such as text, and a
comparable system without it. It asks whether that information improves results; it is
not permission to use data outside the track's cutoff or artifact rules.
Market and document vocabulary
Limit order book
The live list of every outstanding buy and sell order in a market, organized by price.
The highest price a buyer will pay is the best bid, and the lowest price a seller will
accept is the best ask.
Matching engine
The part of an exchange that decides which buy orders meet which sell orders, and at
what price. A simulator that matches orders differently from the reference produces
different trades and, from there, different prices, no matter how fast it runs.
ABIDES
An open-source, agent-based market simulator. It models individual traders sending
orders into an order book and records the results. Follow the Simulation track's
published interface for the compatibility your submission must provide.
Stylized facts
Statistical patterns that real markets show over and over: large moves arriving in
bursts, extreme moves being more common than a bell curve would suggest, and similar
regularities. A simulator can be very fast and still produce output with none of them,
which is why they get checked.
Evidence corpus
The collection of source material provided for evidence-based questions. Use it and
any permitted additional artifacts according to the Explainability track's information
cutoff, disclosure and artifact rules.
Citation
A pointer from a claim in an answer back to the document and passage that supports it,
like a footnote. Citations make a submission show its work.