ActiveSaddler Automated Curriculum Learning for Agent Harness Optimization

*Work done during an internship at Microsoft., †Corresponding authors.

Harness optimizers decide how to update an agent's harness. ActiveSaddler decides which training scenarios should teach it next: a curriculum that co-evolves with the harness.

vs. AutoSaddler · same optimizer, same rollout budget

GAIA2
+4.4 pp
higher test Pass@1
AutoSaddler 55.4% ActiveSaddler 59.8%
4.6 ×
cheaper to 58.5% dev
AutoSaddler $1,360 ActiveSaddler $298
Terminal-Bench 2.0
+7.5 pp
higher test Pass@1
AutoSaddler 72.5% ActiveSaddler 80.0%
1.7 ×
cheaper to 78.9% dev
AutoSaddler $220 ActiveSaddler $128

Abstract

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.

Method Overview

Existing harness optimizers train on scenarios in an order fixed up front. ActiveSaddler makes this curriculum active: it proactively picks which scenarios the optimizer sees next as the harness evolves, leaving the optimizer itself unchanged.

We model this choice as a non-stationary multi-armed bandit whose arms are failure patterns: harness weaknesses shared by several failed runs. Pulling an arm aims the next round of optimization at that weakness.

Overview of ActiveSaddler: Exploration Controller, Arm Prioritizer with stochastic selection, harness optimization, and Failure-Pattern Extractor forming a loop
Overview of ActiveSaddler. At each iteration, the curriculum chooses between exploring unseen scenarios and revisiting an instantiated arm, optimizes the harness on the selected scenario batch, and updates its state: the failure-pattern arms, their priorities, and the unseen pool.

Three Key Components

This framing raises two challenges. First, the arms are not given in advance. Weaknesses surface only when failures are diagnosed, and new ones keep appearing as the harness changes. Second, each iteration runs only a few scenarios. Scenarios the optimizer has not run yet may hide weaknesses no one has seen, so the curriculum must weigh fixing the weaknesses it already knows (exploitation) against running unseen scenarios to find new ones (exploration). The Failure-Pattern Extractor addresses the first challenge; the Arm Prioritizer and the Exploration Controller address the second.

Failure-Pattern Extractor

Where do the arms come from?

Groups failures caused by the same weakness into one arm, registering a new arm when a new weakness appears. The pool starts empty and grows online.

Arm Prioritizer

Which known weakness next?

An LLM scores each arm’s learning progress (severity, fixability, breadth, side-effect risk) and samples one by score. Scores are refreshed at every pull.

Exploration Controller

Fix known weaknesses, or look for new ones?

Each iteration, it either Pulls a known arm (exploit) or Draws unseen scenarios that may reveal new weaknesses (explore).

One iteration of ActiveSaddler, step by step
  1. Explore or exploit. The Exploration Controller looks at the current harness, the known arms, the scenarios not yet seen and the optimization history, and chooses to Draw or Pull.
  2. Draw. Pick a small batch of scenarios the optimizer has not seen yet.
  3. Pull. Otherwise, score every arm, sample one by priority, and build the batch from the scenarios where that weakness showed up.
  4. Optimize. Run the batch with the current harness and let the optimizer improve it (here AutoSaddler: diagnose, patch, re-run, validate on the dev set, reflect, evolve).
  5. Extract. Describe the failures seen before and after the patch, and match each to an existing arm, a combination of arms, or a new arm.
  6. Update. Record what happened and repeat until the rollout budget runs out, then return the harness with the best dev score.

Watch the Curriculum Co-evolve

Here are those three components at work on the actual GAIA2 run behind the paper’s results, replayed iteration by iteration. Each iteration unfolds in three steps:

  1. The Exploration Controller decides whether to Draw or Pull, and you can read its own reasoning.
  2. On a pull, the Arm Prioritizer re-scores every arm, and the needle samples one by priority.
  3. The harness optimizer patches on that batch.

Watch probability mass shift as weaknesses get fixed and new ones appear. Hover an arm to see its scores, or drag the timeline to scrub.

Curriculum Replay
The real GAIA2 · Run 1 optimization run · 51 iterations · 1,393 rollouts
Task agent: gpt-5.5 (reasoning: medium) · Optimizer: gpt-5.5 (reasoning: xhigh) · Cost covers both
Iteration1/ 51
Arm pool0failure patterns
Unseen scenarios75/ 75
Rollouts used6· $12
Best dev accuracy55.4%
Best dev accuracy so farCurriculum decisionDraw share: 44% of iterations 2–26 → 20% of 27–5150%55%60%65%56.9% · draw58.5% · P4 patch61.5% · draw63.1% · P8 patchdrawsonlyArm focus · pull probability per weakness15101520253035404550iter
Pull a failure-pattern armDraw unseen scenariosNew best dev accuracyPatch acceptedDev evaluation

Pull probability across the arm pool

no arm scored yet

The arm pool starts empty. The controller must draw unseen scenarios to discover failure patterns.

1

Exploration Controller

DRAW unseen scenarios0 known arms · 75 unseen scenarios

Forced draw: the arm pool is still empty, so there is nothing to pull yet.

2

Arm Prioritizer

Not invoked on a draw. The batch is sampled from the unseen pool without replacement:

0wlkxc7v5wh08hgfug
3

Harness optimizer outcome

Batch0wlkxc7v5wh08hgfug

Inside the Arm Pool

Every row is one failure-pattern arm, and every column is one arm-selection iteration. Some arms are resolved: their priority collapses after an accepted patch, as with P3. Others keep recurring: they gain supporting scenarios and are pulled again and again, as with P8. Click any arm to trace its lifecycle, and jump into the replay at any of its pulls.

Arm Lifecycle Explorer
All 35 failure-pattern arms × 34 arm-selection iterations · click a row
5202530354550P1P2P3P4P5P6P7P8P9P10P11P12P13P14P15P16P17P18P19P20P21P22P23P24P25P26P27P28P29P30P31P32P33P34P35pull probability0 → 60%sampled armnot yet in pool
P8resolved by own patch at iteration 30

Incomplete source enumeration before cross-app aggregation produces wrong numeric answers

Pull probability0%20%40%60%51020304050
Hover the chart for per-iteration values · 4 supporting scenarios at the end of the run · patch accepted, rejected, batch already passing, new supporting scenario joined
Replay its pulls:

Key Results

ActiveSaddler finds stronger harnesses than every baseline on both benchmarks, with the same gpt-5.5 models and the same rollout budget (1,400 rollouts on GAIA2, 490 on Terminal-Bench 2.0).

GAIA2: best-so-far dev accuracy versus cumulative task-agent rollouts for six methods
GAIA2. Best-so-far dev accuracy over the rollout budget: ActiveSaddler reaches 63.1%, while the best baseline stops at 58.5%.
Terminal-Bench 2.0: best-so-far dev accuracy versus cumulative task-agent rollouts
Terminal-Bench 2.0. ActiveSaddler reaches 89.5%, while the best baseline stops at 78.9%.

Test Results

HarnessTypeGAIA2 Test (300)Terminal-Bench 2.0 Test (40)
Default Agent / Terminus 2Manual53.6 ± 1.164.2 ± 2.9
Terminus-KIRAManual–69.2 ± 3.8
GEPAOptimizer54.2 ± 2.265.8 ± 5.2
Meta-HarnessOptimizer54.2 ± 1.266.7 ± 5.2
AutoSaddlerOptimizer55.4 ± 1.272.5 ± 0.0
AutoSaddler w/ Category Acc. OrderFixed curriculum55.9 ± 1.370.8 ± 1.4
AutoSaddler w/ Scenario Acc. OrderFixed curriculum55.7 ± 1.273.3 ± 1.4
ActiveSaddlerAdaptive curriculum59.8 ± 1.080.0 ± 2.5

Test Pass@1 (mean ± std. over three test-time executions). Fixed easy-to-hard curricula barely help AutoSaddler, and can even hurt it; the adaptive curriculum adds +4.4 and +7.5 pp.

Ablation Studies

Every component matters. Removing any one of them erases most of the gain.

VariantGAIA2ΔTB2Δ
ActiveSaddler59.8 ± 1.080.0 ± 2.5
w/o failure-pattern arms (category arms)56.8 ± 0.8−3.069.2 ± 2.9−10.8
w/o failure-pattern arms (scenario arms)56.2 ± 0.2−3.673.3 ± 3.8−6.7
w/o Arm Prioritizer55.3 ± 2.1−4.572.5 ± 2.5−7.5
w/o Exploration Controller (explore every 5 iters)55.3 ± 0.9−4.571.7 ± 1.4−8.3

RQ1: Do Failure-Pattern Arms Matter?

Replacing failure-pattern arms with category or scenario arms drops Pass@1 by 3.0–3.6 pp on GAIA2 and 6.7–10.8 pp on TB2. A failure-pattern arm keeps an unresolved weakness as a target, so the curriculum can revisit it and refine patches that did not work. Category arms are too coarse: successes on some scenarios mask weaknesses left unresolved in others. Scenario arms are too fine: one weakness is scattered across many arms and is rarely revisited under a limited budget. As a result, failure-pattern arms hit still-failing scenarios far more often, spending less of the budget on scenarios the harness already passes.

Failure hit rate by arm definition: failure pattern, category, scenario
Failure hit rate: the share of iterations whose batch includes a scenario that is still failing.

RQ2: Does Adaptive Arm Prioritization Matter?

Giving every arm the same priority drops Pass@1 by 4.5 pp on GAIA2 and 7.5 pp on TB2. With adaptive scoring, priority keeps moving across arms rather than settling on a fixed few (see P3 and P8 in the explorer above). Because the budget follows each weakness’s changing learning potential, more of the observed failures end up fixed than under uniform priority.

Failure-to-success conversion rate for ActiveSaddler, without Arm Prioritizer, and AutoSaddler
Failure-to-success conversion rate: the share of observed failures that the optimized harness later fixes.

RQ3: Does Adaptive Exploration Matter?

Replacing the Exploration Controller with a fixed schedule (explore every 5 iterations) drops Pass@1 by 4.5 pp on GAIA2 and 8.3 pp on TB2: when to explore matters. The controller pulls while worthwhile unresolved weaknesses remain and draws once the known arms have little left to teach. At its pulls, 2.0× (GAIA2) and 6.9× (TB2) as many arms are unresolved as at its draws, while a fixed schedule is blind to this (0.8× and 1.3×).

Distinct failure patterns discovered over iterations, with annotated Draw and Pull rationales at iterations 2, 7, 24 and 30
Exploration over the run: failure patterns discovered over time, with the controller’s own reasoning at four Draw/Pull decisions.

Conclusion

We introduced ActiveSaddler, an automated curriculum-learning approach for harness optimization that adapts which training scenarios generate the execution feedback used for each harness update. Rather than fixing scenario selection before optimization, ActiveSaddler formulates curriculum construction as a non-stationary bandit that dynamically instantiates failure-pattern arms, prioritizes them by their evolving learning potential, and balances revisiting known weaknesses with exploring unseen scenarios. Across GAIA2 and Terminal-Bench 2.0, ActiveSaddler improves test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer with a fixed scenario order, respectively, while also outperforming existing harness optimizers and fixed curricula. Our ablations further show that these gains rely on failure-pattern arms, adaptive arm prioritization, and adaptive exploration. Overall, these results establish automated curriculum learning as a complementary dimension of harness optimization: performance depends not only on how feedback is used for updates, but also on which training scenarios generate that feedback.