Harness optimizers decide how to update an agent's harness. ActiveSaddler decides which training scenarios should teach it next: a curriculum that co-evolves with the harness.
vs. AutoSaddler · same optimizer, same rollout budget
Abstract
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.
Method Overview
Existing harness optimizers train on scenarios in an order fixed up front. ActiveSaddler makes this curriculum active: it proactively picks which scenarios the optimizer sees next as the harness evolves, leaving the optimizer itself unchanged.
We model this choice as a non-stationary multi-armed bandit whose arms are failure patterns: harness weaknesses shared by several failed runs. Pulling an arm aims the next round of optimization at that weakness.
Three Key Components
This framing raises two challenges. First, the arms are not given in advance. Weaknesses surface only when failures are diagnosed, and new ones keep appearing as the harness changes. Second, each iteration runs only a few scenarios. Scenarios the optimizer has not run yet may hide weaknesses no one has seen, so the curriculum must weigh fixing the weaknesses it already knows (exploitation) against running unseen scenarios to find new ones (exploration). The Failure-Pattern Extractor addresses the first challenge; the Arm Prioritizer and the Exploration Controller address the second.
Where do the arms come from?
Groups failures caused by the same weakness into one arm, registering a new arm when a new weakness appears. The pool starts empty and grows online.
Which known weakness next?
An LLM scores each arm’s learning progress (severity, fixability, breadth, side-effect risk) and samples one by score. Scores are refreshed at every pull.
Fix known weaknesses, or look for new ones?
Each iteration, it either Pulls a known arm (exploit) or Draws unseen scenarios that may reveal new weaknesses (explore).
One iteration of ActiveSaddler, step by step
- Explore or exploit. The Exploration Controller looks at the current harness, the known arms, the scenarios not yet seen and the optimization history, and chooses to Draw or Pull.
- Draw. Pick a small batch of scenarios the optimizer has not seen yet.
- Pull. Otherwise, score every arm, sample one by priority, and build the batch from the scenarios where that weakness showed up.
- Optimize. Run the batch with the current harness and let the optimizer improve it (here AutoSaddler: diagnose, patch, re-run, validate on the dev set, reflect, evolve).
- Extract. Describe the failures seen before and after the patch, and match each to an existing arm, a combination of arms, or a new arm.
- Update. Record what happened and repeat until the rollout budget runs out, then return the harness with the best dev score.
Watch the Curriculum Co-evolve
Here are those three components at work on the actual GAIA2 run behind the paper’s results, replayed iteration by iteration. Each iteration unfolds in three steps:
- The Exploration Controller decides whether to Draw or Pull, and you can read its own reasoning.
- On a pull, the Arm Prioritizer re-scores every arm, and the needle samples one by priority.
- The harness optimizer patches on that batch.
Watch probability mass shift as weaknesses get fixed and new ones appear. Hover an arm to see its scores, or drag the timeline to scrub.
gpt-5.5 (reasoning: medium) · Optimizer: gpt-5.5 (reasoning: xhigh) · Cost covers bothPull probability across the arm pool
no arm scored yetThe arm pool starts empty. The controller must draw unseen scenarios to discover failure patterns.
Exploration Controller
Forced draw: the arm pool is still empty, so there is nothing to pull yet.
Arm Prioritizer
Not invoked on a draw. The batch is sampled from the unseen pool without replacement:
Harness optimizer outcome
Inside the Arm Pool
Every row is one failure-pattern arm, and every column is one arm-selection iteration. Some arms are resolved: their priority collapses after an accepted patch, as with P3. Others keep recurring: they gain supporting scenarios and are pulled again and again, as with P8. Click any arm to trace its lifecycle, and jump into the replay at any of its pulls.
Incomplete source enumeration before cross-app aggregation produces wrong numeric answers
Key Results
ActiveSaddler finds stronger harnesses than every baseline on both benchmarks, with the same gpt-5.5 models and the same rollout budget (1,400 rollouts on GAIA2, 490 on Terminal-Bench 2.0).
Test Results
| Harness | Type | GAIA2 Test (300) | Terminal-Bench 2.0 Test (40) |
|---|---|---|---|
| Default Agent / Terminus 2 | Manual | 53.6 ± 1.1 | 64.2 ± 2.9 |
| Terminus-KIRA | Manual | – | 69.2 ± 3.8 |
| GEPA | Optimizer | 54.2 ± 2.2 | 65.8 ± 5.2 |
| Meta-Harness | Optimizer | 54.2 ± 1.2 | 66.7 ± 5.2 |
| AutoSaddler | Optimizer | 55.4 ± 1.2 | 72.5 ± 0.0 |
| AutoSaddler w/ Category Acc. Order | Fixed curriculum | 55.9 ± 1.3 | 70.8 ± 1.4 |
| AutoSaddler w/ Scenario Acc. Order | Fixed curriculum | 55.7 ± 1.2 | 73.3 ± 1.4 |
| ActiveSaddler | Adaptive curriculum | 59.8 ± 1.0 | 80.0 ± 2.5 |
Test Pass@1 (mean ± std. over three test-time executions). Fixed easy-to-hard curricula barely help AutoSaddler, and can even hurt it; the adaptive curriculum adds +4.4 and +7.5 pp.
Ablation Studies
Every component matters. Removing any one of them erases most of the gain.
| Variant | GAIA2 | Δ | TB2 | Δ |
|---|---|---|---|---|
| ActiveSaddler | 59.8 ± 1.0 | 80.0 ± 2.5 | ||
| w/o failure-pattern arms (category arms) | 56.8 ± 0.8 | −3.0 | 69.2 ± 2.9 | −10.8 |
| w/o failure-pattern arms (scenario arms) | 56.2 ± 0.2 | −3.6 | 73.3 ± 3.8 | −6.7 |
| w/o Arm Prioritizer | 55.3 ± 2.1 | −4.5 | 72.5 ± 2.5 | −7.5 |
| w/o Exploration Controller (explore every 5 iters) | 55.3 ± 0.9 | −4.5 | 71.7 ± 1.4 | −8.3 |
RQ1: Do Failure-Pattern Arms Matter?
Replacing failure-pattern arms with category or scenario arms drops Pass@1 by 3.0–3.6 pp on GAIA2 and 6.7–10.8 pp on TB2. A failure-pattern arm keeps an unresolved weakness as a target, so the curriculum can revisit it and refine patches that did not work. Category arms are too coarse: successes on some scenarios mask weaknesses left unresolved in others. Scenario arms are too fine: one weakness is scattered across many arms and is rarely revisited under a limited budget. As a result, failure-pattern arms hit still-failing scenarios far more often, spending less of the budget on scenarios the harness already passes.
RQ2: Does Adaptive Arm Prioritization Matter?
Giving every arm the same priority drops Pass@1 by 4.5 pp on GAIA2 and 7.5 pp on TB2. With adaptive scoring, priority keeps moving across arms rather than settling on a fixed few (see P3 and P8 in the explorer above). Because the budget follows each weakness’s changing learning potential, more of the observed failures end up fixed than under uniform priority.
RQ3: Does Adaptive Exploration Matter?
Replacing the Exploration Controller with a fixed schedule (explore every 5 iterations) drops Pass@1 by 4.5 pp on GAIA2 and 8.3 pp on TB2: when to explore matters. The controller pulls while worthwhile unresolved weaknesses remain and draws once the known arms have little left to teach. At its pulls, 2.0× (GAIA2) and 6.9× (TB2) as many arms are unresolved as at its draws, while a fixed schedule is blind to this (0.8× and 1.3×).
Conclusion
We introduced ActiveSaddler, an automated curriculum-learning approach for harness optimization that adapts which training scenarios generate the execution feedback used for each harness update. Rather than fixing scenario selection before optimization, ActiveSaddler formulates curriculum construction as a non-stationary bandit that dynamically instantiates failure-pattern arms, prioritizes them by their evolving learning potential, and balances revisiting known weaknesses with exploring unseen scenarios. Across GAIA2 and Terminal-Bench 2.0, ActiveSaddler improves test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer with a fixed scenario order, respectively, while also outperforming existing harness optimizers and fixed curricula. Our ablations further show that these gains rely on failure-pattern arms, adaptive arm prioritization, and adaptive exploration. Overall, these results establish automated curriculum learning as a complementary dimension of harness optimization: performance depends not only on how feedback is used for updates, but also on which training scenarios generate that feedback.