2026-10-09
2026-10-09
Agent Plasticity Measures How Efficiently Agents Convert Experience into Held-Out Gains
Agent Plasticity: Measuring Self-Improvement Through Experience
Most agent evaluations measure what an agent can do at a fixed point rather than how effectively it learns from experience. This paper instead evaluates persistent self-improvement in fixed-weight lineages: acting episodes start in fresh contexts, while an evolving inventory of tools, strategies, skills, and memory is inherited by future instances.
At each checkpoint, the authors measure performance on training and held-out interactions while accounting for learning cost. They define agent plasticity as the efficiency with which experience becomes gains in future held-out performance. This makes final score and acquisition efficiency separate evaluation targets: a model can finish with the best score without improving most efficiently.
The reported trajectories show why that distinction matters. On hard chess, GPT-5.6 Sol rises from near 0% to 36.9% at late reported checkpoints, while Claude Fable 5 and Claude Opus 5 reach higher absolute held-out scores. In the combined-game saturation estimate, GPT-5.6 Sol has the highest estimated plasticity, while Claude Fable 5 reaches the highest fitted performance level. The paper also separates in-regime gains from harder-opponent transfer: gains against stronger held-out opponents are generally smaller or less consistent.
The diagnostic layer is the practical research idea. In the analyzed model–game cells, greater held-out artifact reuse is associated with larger held-out score gains. But high use is not sufficient: for high-reuse models in the chess failure analysis, most remaining identified failures occur while a relevant artifact is already in use. The authors therefore frame possible bottlenecks as missing artifacts, missed reuse, or artifacts that are reused but ineffective in their quality, generalization, or application.
Treat those diagnoses as descriptive rather than causal. The paper does not establish that reuse causes gains or separate artifact quality from application failure. It also does not isolate the value of environment feedback from the extra computation used for reflection and artifact revision; plasticity estimates depend on score headroom, learning horizon, saturation-fit assumptions, and provider pricing.
For an evaluation you design, retain fresh held-out episodes at multiple checkpoints, log both interaction and reflection cost, and inspect failure traces after measuring the curve. The key question is not only whether stored experience helped, but whether a failure reflects absent knowledge, retrieval or reuse failure, or failure despite apparently relevant guidance.
Useful for designing post-deployment learning evaluations: it operationalizes held-out improvement per learning cost and turns artifact traces into a concrete diagnostic for absent, unused, and ineffective experience.
Small reused evaluation sets can inflate reported gains in LLM self-improvement loops
The Winner’s Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
The self-improvement setup studied here keeps a proposed change when it scores better on a small evaluation set. The paper’s concrete target is the resulting selection problem: the selected score is a noisy measurement, and reusing the selection set can make it overstate fresh-data gain. Its model treats candidate errors as correlated within a single decision rather than reducing each proposal to an isolated score. Qwen models rewrite their own instructions, and every candidate is also scored on 600 held-out items; oracle-instrumented runs compare measured selection advantages with held-out advantages.
The confirmatory result is large but conditional. In preregistered greedy loops, the final selection-set score exceeded held-out accuracy by 13–20 points with 16 selection items, and by 1–5 points with 256 items. Increasing the selection set was associated with larger held-out gains on TREC, but this pattern did not generalize to GSM8K. In the oracle runs, when the best rewrite was selected using 16 items, its held-out selection advantage was only 1–51% of the measured advantage across the tested settings.
Changing the acceptance rule was not a sufficient fix in this study. The evaluated alternatives to greedy acceptance did not improve pooled held-out performance over complete runs, despite benefits for some isolated fresh-set decisions; this does not show that every possible acceptance rule fails. A practical audit performed scoring only the starting and final instructions on 64 items never used for selection. It removed average bias and reduced gain-estimation RMSE in the tested runs, but individual estimates remained too noisy to resolve small gains reliably.
For a new loop, the transferable operation is to log the reused-set score separately from an independent seed-to-final audit, then report held-out gain with uncertainty rather than presenting the selected score as progress. The paper’s boundaries matter: its theory does not model the dynamics of lock-in or make claims about recursive self-improvement, and the authors note that the Gaussian selection model may not capture heavy-tailed true effects. The useful research question is therefore narrower: how should audit size, candidate count, and selection-set reuse be logged and varied when comparing complete self-improvement procedures?
Read this to learn a concrete audit design for separating a loop’s reused-set score from its held-out gain, and to see whether changing acceptance rules improved complete runs.
VoS Learns Intervention Value from Stepwise Counterfactuals, Not Failure Risk Alone
From Uncertainty to Action: Learning to Steer LLM Agents
Uncertainty can identify failing trajectories, but it does not reliably identify the step at which steering helps. That is the concrete oversight limitation this paper isolates: a signal can be useful for ranking bad trajectories while being poor for choosing an intervention point. On completed AppWorld trajectories, the reported correlation between failure-detection and within-trajectory localization rankings is -0.12 for Gemma and -0.41 for Qwen.
The method operationalizes this distinction by constructing a Stepwise Outcome Table (SOT): each reference trajectory is replayed from every non-terminal step, each of four steering mechanisms is applied at that step, and the continuation runs to completion. SOT records the change in final task score, including recovery on failed trajectories and harm to successful ones. The table contains about 82,000 counterfactual continuations from 1,864 trajectories across three benchmarks and two agents.
VoS then learns from SOT the value of steering at each step and uses that estimate to decide where to steer; a harm-budgeted trigger decides whether to steer and limits the fraction of successful trajectories disturbed. This is a different control objective from asking only whether the current trajectory looks risky.
The evidence is conditional but broad within the study: VoS improved over unmodified execution in all 12 benchmark-agent-mode settings and had the highest score in 11 of 12. The reported average gains were 7.8 points over unmodified execution and 2.9 points over the strongest of five tested uncertainty-triggered methods. In the AppWorld-Gemma ablation, measured steering outcomes outperformed the tested failure, binary-helpfulness, and process-reward targets offline and scored highest among the listed targets online.
Boundary conditions matter. SOT is expensive to build because it reruns the agent from every step with every mechanism. The demonstrated evidence is limited to two open-weight agents with token distributions available, and black-box verbalized confidence is reported as substantially less effective. The harm budget is calibrated rather than guaranteed on test data. Under stricter rewind-only accounting, online AppWorld with Qwen has a VoS net change of -0.1 ± 0.3, so intervention-specific benefit is not positive in every setting.
A concrete research operation is to collect outcome labels at unselected candidate steps, train value predictors per mechanism, and evaluate intervention localization separately from failure detection. The next question is whether this value signal transfers when counterfactual re-execution is reduced or when the agent is closed-box.
Read Sections 2–4 to learn how to turn intervention choice into a measurable counterfactual-labeling problem before tuning an uncertainty threshold.
Shipped task verifiers can return PASS for wrong persisted effects
Finding Blind Spots in AppWorld and WorkArena Task Verifiers
Execution-based verifiers turn trajectories into PASS/FAIL decisions, but this audit tests the decision against the persisted effect. It uses source-informed mutation tests; the main audit never modifies a shipped checker.
A prior audit showed why the opposite verdict direction matters: it judged 15.3% of sampled FAIL verdicts incorrect, while its stated scope did not assess false positives. The current paper targets selected cases in which a checker passes despite an independently confirmed wrong effect.
In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value. The verifier accepted all three task variants from two of five eligible generators, covering 6/15 constructed effects. A cardinality patch applied to checker copies after the census sent all six duplicate-write cases from PASS to FAIL while valid controls remained passing.
WorkArena supplies the clearest effect-level confirmation. A prospective rerun of 23 extra-field candidates selected because of earlier checker-PASS outcomes found persisted nondefault out-of-scope values in 21 cells through independent Table API readback; the checker passed all 23. Two requested strings were aliases of stored defaults, so those two PASS cases were effect-correct degeneracies rather than confirmed wrong effects. These selected cases confirm mechanisms under the audit protocol, not a population rate.
The intent-swap arm provides a useful negative result with a different evidentiary status: the fixed grid produced zero checker PASS verdicts across 2,689 valid off-diagonal executions. The authors report this as a rejection census: 57 WorkArena cells use session-scoped evidence, while the other 2,632 lack classified rejection causes and independent target ground truth.
A transferable research operation follows: record the checker verdict, independently read back the exact persisted state, and keep construction type, selection condition, and evidence class attached to every count. Use copied-checker patches as diagnostic controls rather than silently changing the object being audited.
The boundary is important. Fixed, selected censuses do not estimate the probability that a deployed run is wrongly accepted. No live model-generated adversary, checker-aware search, or natural-failure corpus contributes to these results, so generalization to adaptive agents, natural failures, other checker designs, or changed environments remains untested.
Read it to learn how to turn a verifier-PASS bug into an independently evidenced state-level finding without presenting a selected 21-cell rerun as a benchmark-wide error rate.
abstract; Section 3.6, paragraph 1; Section 3.5, subsection “Out-of-scope writes in WorkArena”; Section 3.5, subsection “Out-of-scope writes in WorkArena”; Appendix F; Section 3.3, paragraph 2; Section 5, paragraph 1; Section 8, paragraph 1 · §4.4, Verdict reliability; Table 2; Limitations, Audit scope
CredLeak-Bench evaluates credential leakage, legitimate-task utility, and trusted-path recovery together
CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents
Credential-security evaluation has a measurement problem: lower leakage can reflect refusal rather than secure completion. In the supplied comparison record, Scammer4U already measures whether exact stored PII reaches an attacker-controlled endpoint and uses benign twins and one-factor siblings. CredLeak-Bench’s incremental move is to make legitimate-task utility and trusted-path recovery explicit outcomes. That addresses the paper’s reported concern that reducing leakage alone is insufficient when defenses also impair genuine tasks.
The benchmark uses 15 fictional services across eight categories and varies eight phishing cues; its directed and autonomous suites contain 512 and 256 scenarios. CredLeak-Bench pairs deceptive cases with legitimate controls and measures actual submissions, enabling joint assessment of leakage and legitimate-task non-completion across directed authentication, autonomous inbox monitoring, and recovery. It evaluates user-directed authentication and autonomous inbox monitoring, including cases in which monitoring begins without a request to authenticate. Recovery requires the agent to submit the correct password to a trusted endpoint without sending vault information to the phishing endpoint. Leakage is scored from submitted synthetic information rather than the agent’s narration.
All seven evaluated models leaked in both directed and autonomous settings, including when inbox monitoring began without a user request to authenticate. Directed leakage ranged from 31.3% to 75.8%, versus 1.6% to 53.1% autonomously. In this evaluation, lower autonomous leakage co-occurred with higher false refusal for every model, showing why leakage alone can make cautious non-completion look like security. Explicit domain-verification guidance improved both security and utility for six of seven models in the directed setting; this indicates that the observed security–utility trade-off was not inevitable under every intervention. Avoiding a phishing link rarely translated into safe continuation through a trusted alternative: model recovery rates were 0%–24.2% on the primary recovery subset.
A useful replication question is whether an intervention lowers submitted secrets without increasing false refusal, then whether the same agent can complete genuine authentication through a trusted path. Report those outcomes separately, and inspect cue-specific failures: plain HTTP was the highest-leaking directed cue for each model despite the hostname being correct.
The benchmark uses fictional services and a local simulated webpage environment. HTTP and certificate-warning cues are simulated rather than produced by real network or certificate attacks, and the vault policy is prompt-level guidance rather than technically enforced access control. The reported comparison covers seven models in one harness. Recovery is narrower than downstream task completion: it tests correct-password submission to the trusted endpoint without vault information reaching the phishing endpoint.
Read it to redesign an agent-security experiment around three separately scored questions: did a secret cross the deceptive endpoint, did the agent complete the genuine task, and could it recover through a trusted path?
§3.2, Controlled Generation; §4.2, Evaluation Metrics · Section 3.2, Benchmark Design
