Agenthon 2026 / FAQ
FAQ
Frequently Asked Questions
The questions people ask most often about entering. For anything more precise — eligibility,
licensing, privacy, and the full competition rules — see the
Rules, Terms,
Privacy Notice, and Licensing Policy.
Submission availability is announced separately in your signed-in agenthon.net account.
This guide explains the Development workflow; it does not announce that uploads are open.
Entering
Who can take part?
Anyone 18 or over for whom taking part is lawful. If you need your employer's or
university's approval to enter or to submit work you make on their time, get it first.
Organizers, judges, their households, and anyone with access to the held-out test
material are not eligible for prizes — if that might include you, ask us before
entering.
How big can a team be?
One to three people. Entering alone is fine.
You join one team and stay on it — one person, one competition identity, one team. If
you want to take on several tracks, your team enters them together rather than you
joining a different team for each.
Each team picks a captain who handles official messages and decides the final
submission. See Who can submit? below.
Tell us before the joint Final + Verification phase opens if your roster changes.
Who can submit?
Choose one CodaBench account for all your team's uploads across every track you enter.
The first valid team-verification proof links that account to your agenthon.net team.
Other members build, test and are credited, but every upload comes from the linked
account. A different account cannot claim an already-linked team.
When Development submissions open, the daily upload limit per team is 1 for Track 1
(Coding) and 5 for each of Tracks 2–4. The total limit is 20 uploads per team per track, and
23 on Track 1 (raised from 20 on 25 September 2026 at 05:00 UTC).
Held or cancelled uploads still count, even without a score. A replacement upload uses
another attempt; local checking and packing do not.
In the joint Final + Verification phase your team makes one submission for each
track it entered, and there is no resubmission. Organizers perform verification within
that same phase; your team does not make a second verification submission.
What are the competition dates?
Development runs through 12 October 2026. The joint Final + Verification phase runs
from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE. Registration runs from 17 August to 12 October 2026.
Registration and Development close together; the
other announced dates are unchanged.
The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens.
How is our registration checked?
The toolkit uses your agenthon.net Team Number and Team Key to create a proof bound
to your submission descriptor. Organizers check that proof against your registered
team. Your Team Key is not included in the ZIP. Matching an email address alone does
not verify a submission.
Website registration, requesting entry on CodaBench and verification of the uploaded
proof are separate steps. Keep your Team Key and packaged proof private. If your key
changes or a teammate has already linked an account, contact the organizers before
uploading again.
Can we enter more than one track?
Yes. One team can register for as many of the four tracks as it wants, and it is the
same team throughout — you don't form a separate team per track. You pick one final
submission for each track you enter.
Building your submission
What do we actually submit?
Prepare a Linux/amd64 container image and upload a small ZIP that identifies it. The
ZIP contains submission.json and the toolkit-generated
team-claim.json, not container layers, your Team Key or registry credentials.
The descriptor names your image by its immutable digest.
Build the image for linux/amd64: an image without a linux/amd64
build will not run on the evaluation hosts. On a Mac with Apple silicon or another ARM
machine, docker build builds for arm64 unless you ask for amd64:
docker build --platform linux/amd64 -t YOUR_IMAGE .
Then push it:
docker push YOUR_IMAGE
Before you put the digest in submission.json, check that exact reference.
It should print linux/amd64:
docker buildx imagetools inspect --format '{{.Image.OS}}/{{.Image.Architecture}}' YOUR_IMAGE@sha256:YOUR_DIGEST
If you build on a Mac, keep macOS file attributes out of your layers: a layer in which a file carries com.apple.provenance cannot be unpacked on the evaluation hosts, and the run fails before your code starts. A normal docker build with COPY keeps them out with BuildKit 0.22 or later (docker buildx inspect shows the version). If you add a layer from a tarball made on the Mac, create it with tar --no-xattrs; xattr -cr often leaves the attribute in place.
Your image implements solve for
T1 Coding, forecast for
T2 Forecasting, simulate for
T3 Simulation, or analyze for
T4 Explainability.
Use the submission guide and your
track's public repository to build, test and package.
Install the toolkit version identified by the current
shared toolkit README.
It has to run start to finish without anyone stepping in, and stay inside the size,
runtime, memory, and dependency limits your track publishes.
How do we package and upload?
Start from your track's Development descriptor and replace its example image with
your own digest-pinned image. Obtain your descriptor's team ID with:
qfbench2 submission alias --team-number YOUR_TEAM_NUMBER
Enter the Team Key at the hidden prompt. Put the resulting team ID into
submission.json, replacing the example ID. Run the track's local checks,
then package:
qfbench2 submission pack --descriptor submission.json --team-number YOUR_TEAM_NUMBER --out submission.zip
The toolkit prompts for the Team Key, seals the descriptor and creates the proof.
Repack after changing the descriptor or image digest. A copied example's team ID must
not be left in the descriptor.
When uploads are available, use the competition links in your signed-in agenthon.net
account. Request entry on CodaBench, accept the terms, follow its approval instructions,
and upload submission.zip under My Submissions → Development.
Will my submission have internet access while it runs?
All four tracks run without general internet access. Supported Coding, Forecasting
and Explainability submissions can call the provided House Nemotron model through a
restricted connection. Simulation has no network access.
Include required dependencies and permitted artifacts in your image before submitting.
Downloads, package installation from the internet and arbitrary external API calls
are unavailable during evaluation.
What is the House allowance?
For supported House submissions, each evaluation unit allows up to 25 admitted
generation requests, with at most 4,000 output tokens per request. Both are counted
by the House route, and together they are the whole model budget: there is no
per-unit token allowance. The earlier figure of 1,000,000 input plus 100,000 output
tokens per unit is withdrawn and nothing replaces it. The model's context window is a
separate limit on a single request: a request whose prompt, together with its output
limit (4,000 tokens unless you ask for fewer), does not fit in the window is refused
before admission and uses none of the unit's requests.
An admitted request uses budget even if its response is lost or the upstream model
service returns an error. A retry can consume another request. Use your track's
published runtime instructions for timing and request-accounting details.
Can I use open-source libraries and pretrained models?
Yes, where your track allows it and the licence permits. Following the licence terms is
your responsibility.
Watch one thing in particular: your container must not redistribute a component
in a way its licence forbids. If you can use something but not ship it, ask us how to
handle it before you build it in.
Can I use my own or other external data?
Only if the track allows it, you have lawful access, and it respects the track's cutoff
date. Employer, sponsor, embargoed, and privileged-access data are out unless we
authorise them in writing.
The cutoff applies to derived information too: features, retrieval indexes, caches,
fine-tuned weights and synthetic data. Follow the published track policy and any
expressly approved model exception; do not infer permission from schema acceptance.
Declare every model you actually use in the descriptor's required models
array. Use the approved disclosure for House access rather than guessing its metadata.
A genuinely model-free submission uses models: []; a model-free Simulation
submission keeps category: "simulator". Do not invent a model entry. On
Track 1 a model-free submission validates but earns no credit: from 5 October 2026,
00:00 AoE (12:00 UTC), a Track 1 task counts as passed only if the agent used the House model to
solve it at run time (Track 1 README, rule 9).
Scoring and results
What happens if a run fails an admissibility check?
Participant failures are handled under the track's published scoring rules. A failed
unit is not silently dropped: it contributes the track's precommitted worst value.
An organizer or infrastructure fault is held for review rather than recorded as a
participant's zero. A missing score is not automatically a measured zero.
The submission guide explains the
failure-code vocabulary. Do not assume every internal diagnostic or private evaluation
log is available to participants.
If an upload is held or appears stuck, contact the organizers with its submission ID
and visible status before uploading another copy. The upload has already used an
attempt. Never include your Team Key, proof, password or token in a support message.
Where do we see results, and can we iterate against them?
CodaBench receives uploads and shows processing status. Validated Development scores
appear on the public leaderboards on agenthon.net,
visible to everyone. Results update periodically; finishing on CodaBench and appearing
on the website need not happen at the same moment.
Development scores are provisional and help you iterate. Do not try to infer hidden
labels, extract private evaluation data or exploit scoring defects; report suspected
problems. Public practice examples are available in the track repositories. Private
evaluation tasks and reference answers are not published with the results.
From 5 October 2026, 00:00 AoE (12:00 UTC), the Track 1 board ranks only runs uploaded after
25 September 2026 at 05:00 UTC, for every team (Track 1 README, rules 8 and 9).
Do I have to open-source my submission?
A confirmed winner who accepts a prize publishes the code needed to reproduce the
method under the competition's licensing rules. Other entrants are not required to
open-source their work. The Licensing Policy gives the detail,
and the Privacy Notice governs submission handling.
Image visibility is a separate practical choice. The ordinary supported route uses an
anonymously pullable public container, which other people can download, including the
files bundled inside it. A successful pull while logged in does not prove anonymous
access.
If your image must remain confidential, contact the organizers privately before
uploading and obtain confirmation of the supported handoff and exact usable image
reference. Selecting organizer_mirror alone does not transfer an image or
provide registry access. Do not publish confidential material to work around a pull
failure, and never place credentials in the ZIP or image.
Who do I ask?
Email admin@agenthon.net for private matters,
or use your track's public GitHub repository for technical
questions that contain no secrets or confidential artifacts.
For a stuck upload, include its submission ID and visible status. Keep an eye on the
address you registered with and announcements in your agenthon.net account.