SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

Agenthon 2026 / Glossary

Glossary

Glossary

Short definitions of the words used across this site and in the competition materials. These definitions explain the Development workflow; the track instructions and governing competition policies provide the precise requirements. Submission availability is announced separately in your signed-in agenthon.net account.

The competition and its pieces

Competition and benchmark

A benchmark is a set of tasks with agreed rules for scoring them. A competition adds teams, deadlines, phases, and a ranking. Agenthon 2026 is both: a benchmark for verifiable AI in quantitative finance, run as a NeurIPS 2026 competition.

Track

One of the four contests inside Agenthon 2026: coding, forecasting, simulation, and explainability. Each track has its own tasks, its own headline metric, and its own pair of repositories.

Phase

The competition has two phases: Development and a joint Final + Verification phase. Development supports local practice and hosted evaluation with provisional results. The joint phase evaluates one final submission per entered track. Within the joint phase, organizers rerun leading submissions and review reproducibility before results are final. There is no separate verification submission.

Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE. Other dates live on the timeline on the main page.

Unit

One item that gets scored on its own: a single coding task, a single forecast, a single simulation scenario, a single analysis question.

Task card

The short description file that travels with every unit. It tells you what the task asks for. You read task cards; you never edit them.

Split (public-dev, validation, private-test)

A label for a dataset's role in development or evaluation. Public practice material lets you build and test locally. Hosted evaluation uses organizer-controlled packages; a split label does not mean that its tasks or answers are public. Private evaluation tasks and reference answers are not published with the results.

Baseline

An example method used as a point of comparison. Starter packages differ: some include runnable baselines; others provide interface examples and checks. A local baseline score is not an official evaluation score.

Repos, submissions, and the sandbox

Public repo and private repo

The public repository is your track's practice kit: its interface, examples, local checks and available baselines. Organizer repositories hold private evaluation material and reference answers. Participants do not need access to those repositories to submit.

Firewall

The separation between public practice material and private evaluation tasks, reference answers and audit data. Review and automated checks help prevent private material from being published with participant packages or results.

Docker image

A packaged program with its dependencies and permitted artifacts. Official submissions reference a Linux/amd64 image by immutable digest. The backend pulls and runs those image bytes under the track's resource and network rules; identical image bytes do not imply identical hardware or timing on every machine.

Submission

The ZIP uploaded to CodaBench, containing a descriptor and generated team proof. It references the container image that will run; it does not contain the container layers or Team Key.

Descriptor

submission.json: structured information identifying the team, track, phase, image and declared models. The toolkit validates and seals it when packaging.

Image digest

An immutable identifier for specific container image bytes. Rebuilding an image requires updating its reference and repacking the submission.

Team Number

The numeric team identifier issued on agenthon.net. It is different from the descriptor’s derived Team ID.

Team Key

A confidential team credential used locally by the toolkit to create verification proofs. It is entered at a hidden prompt and is not put in the submission ZIP. Keep it private.

Team ID

The identifier the toolkit derives from your Team Number and Team Key for the submission descriptor. Use the alias command to obtain it; do not leave an example ID in your descriptor.

Team proof

Verification data bound to a particular descriptor that demonstrates possession of the Team Key without including the key. The first valid proof links the submitting CodaBench account to the registered team; keep the packaged proof private.

CodaBench

The platform used to request entry to a track, upload packaged submissions and follow processing status. Public Development leaderboard results appear on agenthon.net.

House model

The organizer-provided Nemotron service that supported Coding, Forecasting and Explainability agents can call through the restricted evaluation connection. Its request budget is separate from container runtime limits.

Submission verb

The single command word the organizers use to start your image: solve for coding, forecast for forecasting, simulate for simulation, analyze for explainability. The verb is stable.

Sandbox

The controlled environment in which a submission runs. All four tracks have no general internet access. Supported Coding, Forecasting and Explainability submissions can call the House Nemotron service through a restricted connection. Simulation has no network access. Dependencies and permitted artifacts must be included before the run.

Smoke scorer

A scorer shipped in the public repo so you can grade your own runs before you submit. It grades against public material, so treat it as a check on whether your output is well formed and roughly on track, not as a preview of your final standing.

Gates, scoring, and staying honest

Admissibility gate

A check on integrity, output format, resource use or task semantics. Participant failures follow the track's scoring rules: a failed unit contributes the track's precommitted worst value rather than disappearing from the denominator. Organizer faults are a separate case and require review rather than a participant score.

The gates g0 to g3

The four gates run in order. g0 covers integrity: the files and the image are what they claim to be. g1 covers schema: the output has the shape the track asked for. g2 covers cutoff and resource rules: no future data, and the run stayed inside the limits set for everyone. g3 covers domain semantics: the output makes sense for the subject matter, which is specific to each track.

Organizer fault

An infrastructure or evaluation failure requiring organizer review rather than a participant performance score. An absent score does not automatically mean zero.

Development leaderboard

The public boards on agenthon.net showing eligible team scores for each track. They update periodically and are visible to everyone. Development results are provisional and are separate from Final rankings.

Ranking score

The value used to order a track's leaderboard. It may be direction-adjusted from a raw metric such as a lower-is-better forecasting loss. Scores from different tracks are not directly comparable.

Confidence interval

An uncertainty estimate accompanying a score where the track's scoring contract provides one. It may be computed by resampling units, called the bootstrap. Confidence intervals are not promised for every track or displayed universally on the website.

Reproducibility

Run the same submission again and the result should hold up. During the joint Final + Verification phase, organizers rerun leading submissions to check that their scores hold up.

Data cutoff (embargo)

The information boundary specified by a task or track. Participant data and derived artifacts must respect the applicable cutoff and published exceptions. For example, using the future outcome itself is not a valid forecast.

Leakage

Using information that was not available at the cutoff, by accident or on purpose. A common example is citing a document that was published after the question's cutoff date.

Canary

A hidden marker placed in competition material. If it later turns up in a submission's output, that points to material being copied rather than solved. It is how contamination gets detected after the fact.

Headline metrics

pass@1 and pass@k

pass@1 measures the fraction of tasks passed with one execution per task and is the hosted Coding Development score. A local pass@k experiment measures success across multiple attempts; it must not be confused with the hosted one-execution score.

CRPS

The continuous ranked probability score grades a probabilistic forecast, meaning a forecast that gives a range of outcomes with probabilities attached rather than a single number. It rewards landing close to what happened and being honest about how uncertain you were. Lower is better, and the forecasting track's headline metric is built on it.

Events per second

A simulation performance measurement: how many events a simulator processes per second. Performance is meaningful only alongside correctness and semantic checks. Development feedback is provisional; shared Development conditions do not certify comparable Final timing.

Coverage

How much of a required target is covered, as defined by the metric. Forecast interval coverage asks how often an outcome lies within a predicted range. Evidence or reasoning coverage concerns required support or arguments. These are different measures; use the definition in the relevant track's scoring guide.

Faithfulness

Whether the claims in an answer are supported by the sources it cites. This is distinct from whether the answer is correct or its reasoning addresses the question; follow the Explainability track's published scoring guidance for those components.

Information uplift

A comparison between a system using additional information, such as text, and a comparable system without it. It asks whether that information improves results; it is not permission to use data outside the track's cutoff or artifact rules.

Market and document vocabulary

Limit order book

The live list of every outstanding buy and sell order in a market, organized by price. The highest price a buyer will pay is the best bid, and the lowest price a seller will accept is the best ask.

Matching engine

The part of an exchange that decides which buy orders meet which sell orders, and at what price. A simulator that matches orders differently from the reference produces different trades and, from there, different prices, no matter how fast it runs.

ABIDES

An open-source, agent-based market simulator. It models individual traders sending orders into an order book and records the results. Follow the Simulation track's published interface for the compatibility your submission must provide.

Stylized facts

Statistical patterns that real markets show over and over: large moves arriving in bursts, extreme moves being more common than a bell curve would suggest, and similar regularities. A simulator can be very fast and still produce output with none of them, which is why they get checked.

Evidence corpus

The collection of source material provided for evidence-based questions. Use it and any permitted additional artifacts according to the Explainability track's information cutoff, disclosure and artifact rules.

Citation

A pointer from a claim in an answer back to the document and passage that supports it, like a footnote. Citations make a submission show its work.

The binding documents are the Official Competition Rules, Terms of Participation, Privacy Notice and Data & Software Licensing Policy. Where a guide and the Rules differ, the Rules govern.