Chinese version: [中文]

Run stats

  • Candidates: 275
  • Selected: 30
  • Deepread completed: 30
  • Window (UTC): 2026-08-25T00:00:00Z → 2026-08-26T00:00:00Z (arxiv_announce, expanded=0)
Show selected papers
arXiv IDTitle / LinksCategoriesScoreWhyTags
2608.24017WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents
PDF
cs.CR, cs.AI96Browser-native trust boundaries for WebMCP agents; tackles provenance, tool lifecycle, and prompt injection.agent-safety, browser-agents, prompt-injection, tool-security, provenance, capabilities
2608.24777StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
PDF
cs.AI, cs.CR95Step-level guardrails for agent actions with scalable data generation and safety-utility balancing.agent-safety, guardrails, tool-use, trajectory-auditing, rl, safety-utility
2608.24232TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
PDF
cs.AI95Evidence-grounded benchmark for unsafe prompts, reasoning traces, and final answers in LRMs.safety-evaluation, reasoning-models, guardrails, benchmark, red-teaming
2608.23941More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
PDF
cs.AI95Directly studies pre-execution LLM oversight tradeoffs for trusted monitoring in AI control.ai-safety, oversight, agents, monitoring, evaluation
2608.24691Confident at the moment of action: belief miscalibration in LLM play under hidden information
PDF
cs.AI, cs.CL, cs.LG95Directly probes action-time confidence miscalibration in agentic LLMs; strong safety relevance.llm-safety, calibration, agents, evaluation, hidden-information
2608.24275RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
PDF
cs.AI, cs.CL93RL-based safeguard that invokes applicable safety policies over trajectories; strong agent safety relevance.agent-safety, policy-invocation, guardrails, reinforcement-learning, trajectory-evaluation
2608.24569When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
PDF
cs.AI, cs.MA93Targets safety-critical state preservation failures in multi-stage LLM agent workflows.agent-safety, workflows, constraint-preservation, reliability, evaluation
2608.24022What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
PDF
cs.CR, cs.AI92Runtime localization of behavior-guiding instructions targets dynamic prompt injection in LLM agents.prompt-injection, agent-security, tool-use, runtime-defense, instruction-localization
2608.24354Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs
PDF
cs.CR, cs.AI, cs.CL92Model-level backdoor repair for MLLMs across modalities; strong security relevance and concrete method.mllm-security, backdoor-defense, multimodal, model-repair, robustness
2608.24809Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch
PDF
cs.CL, cs.IR92Inspectable, bounded scholarly search agent with explicit stopping and evidence grounding.agents, rag, grounding, search, safety, evaluation
2608.23965RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation
PDF
cs.CR, cs.AI, cs.IR, cs.LG91Training-free defense for poisoned RAG corpora with certifiable robustness claims; practical security impact.rag, security, data-poisoning, robustness, retrieval-defense, certification
2608.24306Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
PDF
cs.CL91Useful evaluation for faithfulness and citation failures in multi-agent deep research systems.agents, faithfulness, citations, evaluation, deep-research
2608.24509PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents
PDF
cs.AI, cs.SE90Benchmark for resource-aware parallel tool use in agents, targeting practical safety/latency tradeoffs.agents, tool-use, benchmark, resource-awareness, evaluation
2608.24857Prompt Structure Redistributes, Not Reduces: An Empirical Analysis of Security-Weaknesses in LLM-Generated Python Code
PDF
cs.CR, cs.SE90Empirical study of how prompting shifts security weaknesses in LLM-generated code.security, code-llms, prompting, evaluation, software-safety
2608.23959NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution
PDF
cs.CR, cs.AI, cs.IR, cs.LG89Hardens safety alignment against jailbreaks and neuron-ablation attacks by redistributing safety signals.alignment, jailbreaks, robustness, mechanistic-safety, fine-tuning, defense
2608.24361Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
PDF
cs.AI89Failure attribution framework for multi-agent LLM systems could improve debugging and safety.multi-agent, failure-attribution, observability, debugging, reliability
2608.24145TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis
PDF
cs.CL, cs.SE88Benchmarking trustworthy structured-data analysis, including abstention and robustness to table changes.reliability, benchmark, structured-data, abstention, robustness
2608.24571Joint Optimization of Tool Creation and Use for Large Language Model Agents
PDF
cs.AI, cs.SE88Joint RL training of tool creation and use is a notable agent capability advance with safety relevance.llm-agents, tool-use, tool-creation, reinforcement-learning, frontier
2608.24780Linear Probing Provides Robust and Efficient Detection of Machine-Generated Text
PDF
cs.CL88Robust, sample-efficient machine-generated text detection with strong OOD focus.detection, evaluation, robustness, generated-text, security
2608.24099Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
PDF
cs.AI87Benchmark for GUI-agent robustness under dynamic adversarial anomalies; useful for agent reliability eval.agent-evaluation, gui-agents, robustness, adversarial-evaluation, benchmark
2608.24753The RAT: A Unified Bayesian Model for RAG Evaluation
PDF
cs.CL, cs.AI86Unified Bayesian framework for RAG evaluation captures retrieval, abstention, and answer correctness.RAG, evaluation, bayesian, abstention, reliability
2608.24191'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection
PDF
cs.CL, cs.AI86Measures cross-script safety inconsistency, exposing moderation blind spots in undercovered languages.safety-evaluation, multilingual, content-moderation, robustness, fairness
2608.24848BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes
PDF
cs.CL86Scales open-web browser-agent data collection via parallel sandboxes; useful frontier agent infra.web-agents, data, browser, sandboxing, training-infrastructure
2608.24368From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
PDF
cs.AI, cs.SE85Separates state tracking from action generation for more reliable multi-turn tool use in agents.agents, tool-use, reliability, state-tracking, architecture
2608.24621Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding
PDF
cs.CL85Consequence-aware evaluation exposes semantic-safety gaps in safety-critical language understanding.safety-evaluation, safety-critical, benchmark, reliability, language-understanding
2608.24419A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
PDF
cs.AI84Sharp evaluation framing for LLM judges via construct validity, invariance, and sensitivity profiles.evaluation, llm-as-judge, reliability, measurement, benchmarking
2608.24804StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
PDF
cs.AI, cs.SE84Harness evolution improves enterprise agent performance without weight changes; useful for deployment.agents, enterprise, evaluation, harness, deployment
2608.23956Recursive Agentic Reasoning
PDF
cs.AI84Unified test-time reasoning framework with broad relevance to frontier LLM agent performance.reasoning, test-time-compute, agents, evaluation, frontier-llm
2608.24065Mechanistic Circuit Identification for Controllable Data Generation
PDF
cs.LG, cs.AI, cs.CL84Links mechanistic interpretability to controllable data generation via causal circuit interfaces.mechanistic-interpretability, data-generation, control, alignment, circuits
2608.24252SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction
PDF
cs.AI, cs.SE83Benchmark for semantic faithfulness in paper-reproduction agents; exposes silent implementation drift.agent-evaluation, code-generation, benchmark, scientific-reliability, semantic-alignment

AI Paper Insight Brief

2026-08-27

0) Executive takeaways (read this first)

  • The strongest cross-paper pattern is that many “safety improvements” are really measurement or interface improvements: shorter verification windows, paired evaluation, evidence-grounded labels, and consequence-aware metrics often change conclusions more than adding another prompt or judge.
  • For agent safety, where and how you inspect matters as much as what model you use. Several papers show failures arise at handoff boundaries, tool registries, retrieval context construction, and intermediate reasoning/action steps—not just in final outputs.
  • A recurring empirical result is that more context or more structure is not automatically better: longer oversight windows increase false rejections, open-ended deep research loops add cost and error propagation, and security prompts can redistribute rather than remove risk.
  • The most actionable defenses today are runtime-local and auditable: step-level guards, provenance-aware instruction localization, browser-native trust boundaries, and post-retrieval poison filtering all show concrete reductions in attack success with manageable overhead.
  • Evaluation is shifting from coarse correctness to decision-relevant diagnostics: evidence attribution, policy invocation accuracy, resource-feasible scheduling, semantic fidelity to papers, and action-time calibration all expose failure modes hidden by standard success/F1 metrics.
  • RL is being used less for generic capability gains and more for control-layer optimization: policy invocation, step-level guarding, adversarial robustness in GUI agents, and joint tool creation/use.

2) Key themes (clusters)

Theme: Agent oversight and guardrails are moving to step-level, policy-aware control

Theme: Evaluation is becoming evidence-grounded, paired, and consequence-aware

Theme: Robustness work is shifting from prompt defenses to structural defenses

Theme: Agent reliability depends heavily on interfaces, handoffs, and execution scaffolds

Theme: Test-time and tool-time scaling are being re-evaluated under realistic constraints

  • Why it matters: More inference-time compute helps, but only when allocated correctly. This batch shows that repeated sampling often wins because it recovers truncation failures, while scheduling, resource limits, and tool reusability become first-class concerns.
  • Representative papers:
  • Common approach:
    • Decouple components: reasoning operator choice, logical planning vs physical scheduling, tool writing vs tool use, retrieval vs abstention vs answer correctness.
    • Hold prompts, budgets, and graders fixed to isolate the effect of the operator or interface.
    • Optimize with verifiable rewards or probabilistic factorization rather than aggregate end-to-end scores alone.
    • Treat annotation budget and infrastructure constraints as part of the evaluation problem.
  • Open questions / failure modes:
    • BRANCH-style gains may partly reflect truncation artifacts rather than deeper reasoning improvements.
    • Strong logical plans still fail under finite resource capacity.
    • Judge-based rewards and evaluators can bias training.
    • Current models of RAG and tool use are still simplified relative to multi-turn, reranking-heavy deployments.

Theme: New benchmarks are exposing hidden robustness gaps in language, code, and GUI settings

3) Technical synthesis

  • Several papers converge on matched-pair evaluation as the right primitive: twin-prefix oversight, paired recursion scoring, prefix-aligned guardrail data, and controlled perturbation benchmarks all avoid reading “catch” or “accuracy” in isolation.
  • A common failure mechanism is information loss at compression boundaries: long oversight windows with withheld observations, artifact handoffs that weaken constraints, and deep-research pipelines that corrupt citations across agents.
  • Multiple defenses exploit signals outside the model’s surface text behavior: attention maps for instruction localization, hidden-state geometry for RAG poisoning, neuron probes for safety concentration, and layer inconsistency for backdoor repair.
  • There is a clear shift from binary correctness metrics to structured decompositions: safety vs utility, retrieval vs abstention vs task success, logical planning vs physical scheduling, prompt/trace/final-response safety, and semantic vs consequence-aware accuracy.
  • RL is most effective when rewards are verifiable and decomposed: policy invocation correctness, step-level safety labels, tool build/use outcomes, and class-balanced guard optimization all rely on measurable sub-objectives.
  • Several results warn that aggregate gains can be misleading: BRANCH gains partly track truncation recovery; security prompts reduce high-severity findings while increasing low-severity ones; high invariance in judges coexists with low construct sensitivity.
  • Shorter, localized interventions often outperform broader ones: 1–2 action verification windows beat longer windows; step-level guards outperform coarser safeguards; bounded citation-graph exploration beats open-ended deep research loops on cost and recall.
  • Robustness methods increasingly combine formal guarantees with practical attack suites, but the guarantees usually depend on assumptions like honest majority, separation, or fine-tuning control.
  • Many papers show that system scaffolding is a major optimization surface: harness evolution, typed OODA stages, adaptive influence graphs, and browser-native capability layers improve outcomes without changing base weights.
  • Across agent settings, the hardest failures remain long-horizon and context-dependent: round-layer GUI traps, long-context localization, multi-step constraint preservation, and resource-feasible scheduling under tight capacity.

4) Top 5 papers (with “why now”)

  • More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
    • Introduces a clean twin-prefix framework that isolates the effect of verification window length from task difficulty and error position.
    • Shows across two domains and six judges that informedness peaks at 1–2 actions; longer windows raise catch and false rejection together.
    • Mechanistically ties long-window failure to withheld observations, with replay restoring much of the lost discrimination.
    • Why now: many agent stacks are adding pre-execution monitors, and this paper says the default instinct to “review more context” may actively hurt deployable oversight.
    • Skepticism / limitation: results are zero-shot, limited to L≤8 and single injected writes, so adaptive adversaries and trained verifiers remain open.
  • StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
    • Combines StepGen synthetic prefix-aligned supervision with a 4B step-level guard and Balance-GRPO to reduce defense bias.
    • Reports strong static performance and runtime guarded-agent gains, cutting mean ASR by 77.3% with only a 2.8-point utility drop.
    • Provides a practical latency profile (~600 ms per call; ~7.24% of AgentDojo task time).
    • Why now: step-level guarding is becoming the operational control point for tool-using agents, and this is one of the more deployment-shaped papers in the batch.
    • Skepticism / limitation: synthetic data and benchmark-limited evaluation may leave gaps for open-ended tools and adaptive adversaries.
  • RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation
    • Proposes a training-free, black-box-compatible post-retrieval defense using surrogate hidden-state geometry and geometric-median consensus.
    • Offers a certifiable guarantee under honest-majority and separation assumptions.
    • Empirically achieves low ASR with competitive accuracy across datasets, models, and adaptive attacks at modest overhead.
    • Why now: RAG poisoning is moving from toy threat to practical deployment issue, and this is a rare defense that does not require retraining or white-box access to the target LLM.
    • Skepticism / limitation: depends on surrogate inaccessibility and honest-majority; effectiveness may erode if attackers can optimize against the surrogate.
  • A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
    • Reframes judge evaluation around a 2D validity profile: invariance to construct-preserving edits and sensitivity to construct-changing edits.
    • Finds judges can have high invariance (mean 0.945) but low sensitivity (mean 0.319) at matched thresholds.
    • Shows public label sets can often be partially reproduced by surface-only predictors.
    • Why now: as LLM-as-a-judge becomes infrastructure for eval, reward modeling, and filtering, this paper challenges a core validation habit.
    • Skepticism / limitation: results are conditional on the chosen edit families, English-only probes, and surviving items after protocol filtering.
  • WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents
    • Identifies structural WebMCP vulnerabilities around provenance, lifecycle, execution opacity, and semantic injection.
    • Implements a browser-native ownership layer plus asymmetric dual-agent runtime; blocks all 80 description injections and reduces returned-content attacks to 2/80 in the full setup.
    • Demonstrates deterministic elimination of overwrite/revocation attacks in tested settings.
    • Why now: browser-native agent tooling is becoming standardized, and this paper addresses protocol-level risks before they harden into ecosystem defaults.
    • Skepticism / limitation: current implementation is a polyfill, and adaptive name-layer attacks remain a live bypass without timing/name controls.

5) Practical next steps

  • Default pre-execution oversight to short verification units (1–2 actions) and explicitly measure false rejection alongside catch before widening review windows.
  • Add paired evaluation protocols to internal agent experiments: same prompts, same budgets, same resolved-item intersections, and explicit handling of infrastructure failures.
  • For tool-using agents, instrument step-level guardrails with matched safe/unsafe prefixes and track utility loss, not just ASR reduction.
  • In RAG systems, test post-retrieval filtering separately from generation quality; log poison ratio, honest-majority assumptions, and whether surrogate signals are attacker-exposed.
  • Audit multi-agent workflows for constraint weakening at handoffs by checking whether blockers, authority, prerequisites, and fallbacks survive summarization or ticketing.
  • For browser or MCP-style integrations, enforce native provenance and lifecycle binding before relying on semantic prompt-injection filters.
  • Expand eval dashboards beyond accuracy/F1 to include evidence attribution, abstention quality, downgrade severity, policy invocation accuracy, and resource-capacity violations.
  • If using LLM judges, validate both invariance and construct sensitivity; do not treat agreement on surface edits as sufficient evidence of evaluator quality.
  • For long-horizon GUI or web agents, separate failures into single-step recoverable vs contextual multi-step traps and train/evaluate mitigations accordingly.
  • Where possible, optimize scaffolding as a first-class lever: try typed stage separation, harness evolution, or bounded search spaces before assuming the next gain requires a larger base model.

Generated from per-paper analyses; no external browsing.