PostTrainBench v1.2
More reliable scores and verdicts
PostTrainBench v1.2 drops BFCL, fixes HumanEval and remote-code scoring, averages evaluations over multiple seeds, decides contamination by majority vote, and adds Harbor support.
A few weeks ago, Epoch AI reviewed PostTrainBench v1.1 and designated it Verified. In our response, we talked about two changes for v1.2: removing BFCL and deciding contamination by majority vote. Both are now in v1.2.
v1.2 also fixes three issues in how the model weights are evaluated: HumanEval could count a sample as passed without running its tests, benchmarks could run code from the submitted model, and all benchmarks were evaluated with a single seed. Finally, a long awaited feature: an Harbor adapter lets anyone run PostTrainBench on cloud sandboxes.
Removing BFCL
Epoch AI's review classified the Berkeley Function Calling Leaderboard as flawed:
Epoch AI's reviewWe classify one of the included benchmarks, BFCL, as Flawed due to an observed error rate of 48%, but the selected subset used in PostTrainBench may have a different error profile.
We agree, and had noticed the same issues independently. BFCL was also where we saw very specific item targeted training, which sometimes gave the largest score increases without a comparable general improvement.
- BFCL is no longer included in the leaderboard and pipeline.
- The remaining six benchmarks are reweighted as before: the weight is proportional to the inverse of the gap between the official instruct models and the base models.
| Benchmark | v1.1 weight | v1.2 weight |
|---|---|---|
| AIME 2025 | 22.7% | 24.5% |
| GPQA Main | 22.5% | 24.3% |
| HealthBench | 18.4% | 19.9% |
| HumanEval | 10.6% | 11.5% |
| GSM8K | 9.4% | 10.1% |
| Arena Hard Writing | 9.0% | 9.8% |
| BFCL | 7.5% | – |
Fixing HumanEval scoring
HumanEval checks a completion by running the model's function against unit tests. The scorer that we used ran the model's code and the tests in one process and counted a sample as passed if that process exited normally. Code that exited before the tests ran, like by calling exit() would pass, along with things like a function returning an object that compares equal to everything.
- The tests run in their own process and don't share a process with the model's code.
- Each call to the model's function runs in a fresh process inside a separate sandbox container, with no network, no access to host files, and with its own process namespace.
Fixing Arena Hard and HealthBench scoring
The Arena Hard and HealthBench evaluation scripts used the flag --trust-remote-code when loading the model. This could allow intricate reward hacks, because the agent could write arbitrary code into the submitted model, which would then be executed during the final evaluation.
We have not seen any exploit of this. We removed the flag so that future agents can't exploit it.
- The Arena Hard and HealthBench evaluation scripts no longer use
--trust-remote-code. - Scores did not change, because no agent had exploited the flag.
Averaging evaluations over seeds
Until now, each final model was evaluated once. A single evaluation is noisy, especially on benchmarks with few samples, such as AIME (30 problems). In PTB v1.2, the model is evaluated with several seeds, and its score is the mean over them.
The number of seeds depends on the benchmark:
| Benchmark | Seeds | Why |
|---|---|---|
| AIME 2025 | 5 | Has only 30 samples. |
| GSM8K, GPQA Main, HumanEval | 3 | Have at least 164 samples. |
| Arena Hard Writing, HealthBench | 1 | LLM judged |
Majority vote judging
The contamination judge is an LLM reviewing the agent trace, so in principle there can be noise in the verdict, given the stochasticity of the judges. We didn't see a lot of variance in our runs and have calibrated the judge. But to further reduce the variance, a run is now flagged for contamination only if at least two of three judge instances flag it.
- The contamination judge runs three times on every run. The API usage and PostTrainBench lookup judges still run once due to very little variance.
- We still review flagged runs manually.
- All judges now run on GPT-5.6 Terra, after OpenAI retired GPT-5.4 from Codex on 31 August 2026.
The trace viewer shows the majority verdict for each run.
New agents
We've added five new frontier agents to v1.2: Fable 5.1, Opus 5.5 and GPT-6 (Astra), GLM 5.3 and GLM 5.3 Flash, all at max reasoning. Fable 5.1 is now #1 at 44.6%, followed by Opus 5.5 at 43.8% and GPT-6 (Astra) at 41.9%.
Five of Fable 5.1's GPQA Main runs fell back to Opus 5, due to safety classifiers.
Recomputed results
How the leaderboard changed
All agents have their scores recomputed under the v1.2 setting. This is how their scores changed with the update:
Aggregate benchmark performance. Differences include removing BFCL and reweighting, HumanEval rescoring, and majority vote judging.
Run PostTrainBench yourself
Until now, running PostTrainBench was pretty hard, unless someone had an HTCondor cluster setup. Version 1.2 adds a Harbor adapter that allows anyone to run the benchmark on cloud GPUs through Harbor and Modal with exactly the same settings. To get started:
git clone https://github.com/aisa-group/PostTrainBench.git
cd PostTrainBench/src/harbor_adapter
bash run_modal_task.sh --benchmark all --base-model all \
--agent claude-code --model anthropic/claude-opus-4-8 --job-name sweep1
For more details, look at the Harbor adapter README. We will also publish PostTrainBench to the Harbor hub soon, so it can be run directly from Harbor.