SQA Agenthon
Agenthon Submissions Deadline: October 12th, 23:59 AoE
(October 13th at 7:59 AM New York, October 13th at 12:59 PM London, October 13th at 7:59 PM Singapore)
Announcing Competition Awards!

Agenthon 2026 / FAQ

FAQ

Frequently Asked Questions

The questions people ask most often about entering. For anything more precise — eligibility, licensing, privacy, and the full competition rules — see the Rules, Terms, Privacy Notice, and Licensing Policy.

Submission availability is announced separately in your signed-in agenthon.net account. This guide explains the Development workflow; it does not announce that uploads are open.

Entering

Who can take part?

Anyone 18 or over for whom taking part is lawful. If you need your employer's or university's approval to enter or to submit work you make on their time, get it first.

Organizers, judges, their households, and anyone with access to the held-out test material are not eligible for prizes — if that might include you, ask us before entering.

How big can a team be?

One to three people. Entering alone is fine.

You join one team and stay on it — one person, one competition identity, one team. If you want to take on several tracks, your team enters them together rather than you joining a different team for each.

Each team picks a captain who handles official messages and decides the final submission. See Who can submit? below.

Tell us before the joint Final + Verification phase opens if your roster changes.

Who can submit?

Choose one CodaBench account for all your team's uploads across every track you enter. The first valid team-verification proof links that account to your agenthon.net team. Other members build, test and are credited, but every upload comes from the linked account. A different account cannot claim an already-linked team.

When Development submissions open, the daily upload limit per team is 1 for Track 1 (Coding) and 5 for each of Tracks 2–4. The total limit is 20 uploads per team per track, and 23 on Track 1 (raised from 20 on 25 September 2026 at 05:00 UTC). Held or cancelled uploads still count, even without a score. A replacement upload uses another attempt; local checking and packing do not.

In the joint Final + Verification phase your team makes one submission for each track it entered, and there is no resubmission. Organizers perform verification within that same phase; your team does not make a second verification submission.

What are the competition dates?

Development runs through 12 October 2026. The joint Final + Verification phase runs from 13 to 25 October 2026. Registration and Development close on 12 October at 23:59 Anywhere on Earth (AoE, UTC−12). Final + Verification closes on 25 October at 23:59 AoE. Registration runs from 17 August to 12 October 2026. Registration and Development close together; the other announced dates are unchanged.

The last Development runs start by 20:00 UTC on 12 October: an upload that has not started by then is not run. The evaluation fleet is in scheduled maintenance on 13 October from 08:00 to 12:00 UTC, when the Final + Verification phase opens.

How is our registration checked?

The toolkit uses your agenthon.net Team Number and Team Key to create a proof bound to your submission descriptor. Organizers check that proof against your registered team. Your Team Key is not included in the ZIP. Matching an email address alone does not verify a submission.

Website registration, requesting entry on CodaBench and verification of the uploaded proof are separate steps. Keep your Team Key and packaged proof private. If your key changes or a teammate has already linked an account, contact the organizers before uploading again.

Can we enter more than one track?

Yes. One team can register for as many of the four tracks as it wants, and it is the same team throughout — you don't form a separate team per track. You pick one final submission for each track you enter.

Building your submission

What do we actually submit?

Prepare a Linux/amd64 container image and upload a small ZIP that identifies it. The ZIP contains submission.json and the toolkit-generated team-claim.json, not container layers, your Team Key or registry credentials. The descriptor names your image by its immutable digest.

Build the image for linux/amd64: an image without a linux/amd64 build will not run on the evaluation hosts. On a Mac with Apple silicon or another ARM machine, docker build builds for arm64 unless you ask for amd64:

docker build --platform linux/amd64 -t YOUR_IMAGE .

Then push it:

docker push YOUR_IMAGE

Before you put the digest in submission.json, check that exact reference. It should print linux/amd64:

docker buildx imagetools inspect --format '{{.Image.OS}}/{{.Image.Architecture}}' YOUR_IMAGE@sha256:YOUR_DIGEST

If you build on a Mac, keep macOS file attributes out of your layers: a layer in which a file carries com.apple.provenance cannot be unpacked on the evaluation hosts, and the run fails before your code starts. A normal docker build with COPY keeps them out with BuildKit 0.22 or later (docker buildx inspect shows the version). If you add a layer from a tarball made on the Mac, create it with tar --no-xattrs; xattr -cr often leaves the attribute in place.

Your image implements solve for T1 Coding, forecast for T2 Forecasting, simulate for T3 Simulation, or analyze for T4 Explainability.

Use the submission guide and your track's public repository to build, test and package. Install the toolkit version identified by the current shared toolkit README.

It has to run start to finish without anyone stepping in, and stay inside the size, runtime, memory, and dependency limits your track publishes.

How do we package and upload?

Start from your track's Development descriptor and replace its example image with your own digest-pinned image. Obtain your descriptor's team ID with:

qfbench2 submission alias --team-number YOUR_TEAM_NUMBER

Enter the Team Key at the hidden prompt. Put the resulting team ID into submission.json, replacing the example ID. Run the track's local checks, then package:

qfbench2 submission pack --descriptor submission.json --team-number YOUR_TEAM_NUMBER --out submission.zip

The toolkit prompts for the Team Key, seals the descriptor and creates the proof. Repack after changing the descriptor or image digest. A copied example's team ID must not be left in the descriptor.

When uploads are available, use the competition links in your signed-in agenthon.net account. Request entry on CodaBench, accept the terms, follow its approval instructions, and upload submission.zip under My Submissions → Development.

Will my submission have internet access while it runs?

All four tracks run without general internet access. Supported Coding, Forecasting and Explainability submissions can call the provided House Nemotron model through a restricted connection. Simulation has no network access.

Include required dependencies and permitted artifacts in your image before submitting. Downloads, package installation from the internet and arbitrary external API calls are unavailable during evaluation.

What is the House allowance?

For supported House submissions, each evaluation unit allows up to 25 admitted generation requests, with at most 4,000 output tokens per request. Both are counted by the House route, and together they are the whole model budget: there is no per-unit token allowance. The earlier figure of 1,000,000 input plus 100,000 output tokens per unit is withdrawn and nothing replaces it. The model's context window is a separate limit on a single request: a request whose prompt, together with its output limit (4,000 tokens unless you ask for fewer), does not fit in the window is refused before admission and uses none of the unit's requests.

An admitted request uses budget even if its response is lost or the upstream model service returns an error. A retry can consume another request. Use your track's published runtime instructions for timing and request-accounting details.

Can I use open-source libraries and pretrained models?

Yes, where your track allows it and the licence permits. Following the licence terms is your responsibility.

Watch one thing in particular: your container must not redistribute a component in a way its licence forbids. If you can use something but not ship it, ask us how to handle it before you build it in.

Can I use my own or other external data?

Only if the track allows it, you have lawful access, and it respects the track's cutoff date. Employer, sponsor, embargoed, and privileged-access data are out unless we authorise them in writing.

The cutoff applies to derived information too: features, retrieval indexes, caches, fine-tuned weights and synthetic data. Follow the published track policy and any expressly approved model exception; do not infer permission from schema acceptance.

Declare every model you actually use in the descriptor's required models array. Use the approved disclosure for House access rather than guessing its metadata. A genuinely model-free submission uses models: []; a model-free Simulation submission keeps category: "simulator". Do not invent a model entry. On Track 1 a model-free submission validates but earns no credit: from 5 October 2026, 00:00 AoE (12:00 UTC), a Track 1 task counts as passed only if the agent used the House model to solve it at run time (Track 1 README, rule 9).

Scoring and results

What happens if a run fails an admissibility check?

Participant failures are handled under the track's published scoring rules. A failed unit is not silently dropped: it contributes the track's precommitted worst value. An organizer or infrastructure fault is held for review rather than recorded as a participant's zero. A missing score is not automatically a measured zero.

The submission guide explains the failure-code vocabulary. Do not assume every internal diagnostic or private evaluation log is available to participants.

If an upload is held or appears stuck, contact the organizers with its submission ID and visible status before uploading another copy. The upload has already used an attempt. Never include your Team Key, proof, password or token in a support message.

Where do we see results, and can we iterate against them?

CodaBench receives uploads and shows processing status. Validated Development scores appear on the public leaderboards on agenthon.net, visible to everyone. Results update periodically; finishing on CodaBench and appearing on the website need not happen at the same moment.

Development scores are provisional and help you iterate. Do not try to infer hidden labels, extract private evaluation data or exploit scoring defects; report suspected problems. Public practice examples are available in the track repositories. Private evaluation tasks and reference answers are not published with the results.

From 5 October 2026, 00:00 AoE (12:00 UTC), the Track 1 board ranks only runs uploaded after 25 September 2026 at 05:00 UTC, for every team (Track 1 README, rules 8 and 9).

Do I have to open-source my submission?

A confirmed winner who accepts a prize publishes the code needed to reproduce the method under the competition's licensing rules. Other entrants are not required to open-source their work. The Licensing Policy gives the detail, and the Privacy Notice governs submission handling.

Image visibility is a separate practical choice. The ordinary supported route uses an anonymously pullable public container, which other people can download, including the files bundled inside it. A successful pull while logged in does not prove anonymous access.

If your image must remain confidential, contact the organizers privately before uploading and obtain confirmation of the supported handoff and exact usable image reference. Selecting organizer_mirror alone does not transfer an image or provide registry access. Do not publish confidential material to work around a pull failure, and never place credentials in the ZIP or image.

Who do I ask?

Email admin@agenthon.net for private matters, or use your track's public GitHub repository for technical questions that contain no secrets or confidential artifacts.

For a stuck upload, include its submission ID and visible status. Keep an eye on the address you registered with and announcements in your agenthon.net account.

The binding documents are the Official Competition Rules, Terms of Participation, Privacy Notice and Data & Software Licensing Policy. Where a guide and the Rules differ, the Rules govern.