中文版:/paper-news/2026-07-28-zh/

Run stats

  • Candidate papers: 219
  • Selected papers: 30
  • Deepreads completed: 10
  • Time window (UTC): 2026-07-27T00:00:00Z → 2026-07-28T00:00:00Z (arxiv_announce, expanded=0)
  • Evidence basis: This issue is synthesized from the full selected-paper set, candidate abstracts, and 10 completed deepreads available locally. Where deeper reads are missing, claims below are framed as selected-paper synthesis rather than full-paper consensus.
Expand to view the selected-paper list used for this synthesis
arXiv IDTitle / LinksCategoryScoreSelection reasonTags
2607.24625Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
PDF
cs.CR, cs.AI95Formal taint-confinement framework for LLM agents with concrete prompt-injection mitigation design.agent-safety, prompt-injection, information-flow-control, permissions, llm-agents, security
2607.24392When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
PDF
cs.CR, cs.LG95Systematic jailbreak-defense tradeoff study on safety, over-refusal, and inference cost.llm-safety, jailbreaks, defenses, evaluation, robustness
2607.23929MemTX: Transactional Belief Commit for Stateful Agent Memory
PDF
cs.AI94Transactional memory with provenance/permissions directly targets unsafe agent actions from polluted state.agents, agent-safety, memory, tool-use, provenance, permissions
2607.23999ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents
PDF
cs.CR93Strong agent-security benchmark measuring post-injection traces, containment, recovery, and utility tradeoffs.agent-safety, benchmark, prompt-injection, tool-use, containment, evaluation
2607.24343Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
PDF
cs.LG, cs.AI, cs.CL93Per-field conformal risk control for high-risk LLM tool-call arguments; strong agent safety relevance.agents, tool-use, conformal, risk-control, calibration, safety
2607.24054Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
PDF
cs.AI92Introduces success provenance auditing so agent evals distinguish reasoning from answer leakage.agents, evaluation, auditing, benchmark, information-leakage
2607.23982Moral Hazard in Multi-Agent Language Models
PDF
cs.MA, cs.AI92Introduces a controlled multi-agent safety game for hidden-action failures and evaluates open LMs.multi-agent, safety, evaluation, social-simulation, coordination, benchmark
2607.24300Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
PDF
cs.CL, cs.MA91Targets self-improving agent verification failure, a core alignment risk with deployment-relevant framing.alignment, agents, self-improvement, verification, evaluation, reliability
2607.24653Kimi K3: Open Frontier Intelligence
PDF
cs.CL, cs.LG91Open frontier MoE with 1M context, vision, RL post-training, and scaling-efficiency claims.frontier-llm, moe, long-context, multimodal, reasoning, scaling
2607.24484What do Reward Models Memorize?
PDF
cs.LG, cs.CL91Important alignment result: reward models memorize shortcuts and misgeneralize on preference data.alignment, reward-models, preference-learning, memorization, generalization, RLHF
2607.23933SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
PDF
cs.DC, cs.AI, cs.LG, cs.PF91Agent sandbox scheduling for MCP tool use; strong systems relevance to safe, efficient agent deployment.agents, sandboxing, MCP, systems, serving, tool-use
2607.24604Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
PDF
cs.CL, cs.AI90Shows revision loops can reduce code-agent reliability; proposes evidence-bound repair contracts.agents, code-generation, reliability, evaluation, self-correction
2607.24720The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
PDF
cs.CL, cs.AI, cs.LG90Controlled study of long-horizon planning from pretraining to agentic distillation; high frontier relevance.llm, agents, planning, distillation, pretraining, post-training
2607.24010When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost
PDF
cs.LG89Useful budget-aware Active RAG evaluation separating utility, calibration, and cost under retrieval decisions.rag, evaluation, calibration, efficiency, retrieval, llm-systems
2607.24645Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
PDF
cs.LG, cs.AI, cs.CL89Interpretability for SAE features via causal logit-effect geometry; useful for LLM understanding and steering.interpretability, SAE, mechanistic-interpretability, LLMs, steering
2607.24562Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models
PDF
cs.AI88Hierarchical group-conditional conformal control targets subgroup risk, not just average LLM risk.calibration, fairness, selective-prediction, conformal, reliability
2607.24717DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
PDF
cs.CL, cs.AI88Per-example pretraining data orchestration could materially improve LLM training quality and reuse.llm, pretraining, data-curation, data-quality, orchestration
2607.24112Scaling GUI Agents with Visual State Transitions
PDF
cs.AI88New pretraining axis for GUI agents using state transitions; clear frontier agent capability advance.GUI-agents, multimodal, pretraining, world-models, agents
2607.24063The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards
PDF
cs.AI87Reframes hallucination benchmarking around factuality-cost tradeoffs, improving deployment-relevant evaluation.hallucination, factuality, evaluation, efficiency, benchmarking, deployment
2607.24167Falsifiable Commitment Planning for Self-Correcting Web Agents
PDF
cs.AI87Falsifiable commitment planning adds explicit evidence checks for robust long-horizon web agents.web-agents, planning, self-correction, reliability, agent-safety
2607.24647Efficiency Matters in Autonomous Research
PDF
cs.AI, cs.LG87Pushes autonomous research evaluation beyond final quality to efficiency under budget, a key agent metric.agents, autonomous-research, evaluation, efficiency, benchmarking
2607.24174Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection
PDF
cs.CR86Directly studies prompt injection against LLM log analysis in SOC workflows using realistic attack traces.prompt-injection, security, llm-applications, red-teaming, soc, evaluation
2607.24368Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
PDF
cs.CL85Benchmark exposes memory retrieval blind spots from implicit associations in long-term agent memory.agent-memory, benchmark, retrieval, long-term-memory, evaluation
2607.23955EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
PDF
cs.AI85RL for search agents with evidence-constrained teacher backoff; relevant to reliable agentic RAG.agents, rag, reinforcement-learning, search, evidence, reliability
2607.24339Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
PDF
cs.AI, cs.CL84Runtime control layer for affect-regulated LLM agents; interesting architecture with injection-isolated controller.llm-agents, runtime-control, alignment, robustness, architecture, monitoring
2607.24667Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating
PDF
cs.AI84New fixed-lag smoothing view of test-time memory offers practical long-context efficiency gains.llm, long-context, memory, inference, efficiency
2607.24354Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
PDF
cs.AI84Improves multimodal prompt optimization with visual failure feedback; practical VLM adaptation method.VLMs, prompt-optimization, multimodal, evaluation, adaptation
2607.23970Understanding Machine Unlearning Through the Lens of Mode Connectivity
PDF
cs.LG, cs.AI, cs.CL83Analyzes machine unlearning via mode connectivity; useful for privacy/safety understanding.unlearning, privacy, optimization, theory, safety
2607.24651Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
PDF
cs.CV, cs.CL, cs.IR83Studies attribution hallucination in document VLMs and tests text-based evidence interfaces.multimodal, hallucination, attribution, evaluation, documents, grounding
2607.24268Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
PDF
cs.CL82Good reliability eval: separates execution failure states from correctness, exposing hidden benchmark artifacts.evaluation, reliability, reasoning, benchmarking, failure-analysis, language-models

AI Paper Insight Brief

2026-07-28

0) Executive takeaways (read this first)

  • The strongest shift today is that agent safety is becoming a state-management problem. Papers on memory commit, taint confinement, and per-field tool-call control all move safety into runtime state rather than treating it as a prompt wrapper.
  • Endpoint metrics keep losing explanatory power. A correct answer can come from leakage, a zero-violation run can still propagate taint widely, and a defense with strong jailbreak numbers can quietly destroy usability or latency budgets.
  • The practical design pattern is granular control, not monolithic guardrails: separate high-risk roles, isolate contaminated branches, preserve parent state, and demand explicit evidence before irreversible action.
  • A second systems thread today is that efficiency and safety are converging. Sandbox scheduling, budget-aware retrieval, and evidence-bound code repair all assume the real unit of improvement is a controlled workflow, not a single response.
  • Because only 10 deepreads are locally completed for this date, the broad framing below is deliberately anchored to the selected-paper set plus those completed analyses, not to a full 30-paper deepread pass.

2) Key themes (clusters)

Theme: Stateful safety

  • Why it matters: The day’s strongest safety papers stop treating agent context like a flat prompt history. Instead they model belief state, permissions, taint propagation, and action eligibility as explicit runtime objects.
  • Representative papers:
  • Common approach:
    • Add explicit lifecycle state, commit checks, rollback paths, or permission rules around agent memory and tool calls.
    • Distinguish safe influence from unsafe influence at the level of roles, branches, or records.
    • Treat irreversible actions as governed transitions, not just another generation step.
  • Open questions / failure modes:
    • Most systems still rely on trusted sanitizers, ledgers, or tool metadata.
    • Bounded verification does not remove covert channels or undeclared side effects.
    • More control logic can trade safety gains for operator complexity or calibration burden.

Theme: Trace-first evaluation

  • Why it matters: Several papers today make the same argument from different angles: final outcomes are insufficient statistics. To judge a system, you have to inspect how success or failure happened.
  • Representative papers:
  • Common approach:
    • Compare matched traces, not just matched endpoint labels.
    • Separate correctness from provenance, utility from containment, and safety from cost.
    • Measure where a policy fails: admission, preservation, retrieval trigger, revision loop, or action execution.
  • Open questions / failure modes:
    • Trace-rich evaluations are harder to standardize and more expensive to run.
    • Synthetic or single-model studies may not transfer cleanly to production traffic.
    • Better instrumentation does not automatically imply better interventions.

Theme: Granular control

  • Why it matters: The papers worth watching do not promise universal guards. They propose narrower, testable contracts: per-role risk budgets, typed revision receipts, budget-aware retrieval thresholds, and sandbox prewarming tied to predicted tool use.
  • Representative papers:
  • Common approach:
    • Replace global safety claims with local contracts tied to fields, states, budgets, or exact code snapshots.
    • Optimize for recoverability and auditability alongside task success.
    • Make deployment choices explicit: what gets cached, what gets revised, what gets certified, what gets deferred.
  • Open questions / failure modes:
    • Narrow contracts can leave gaps between controlled subsystems.
    • Calibration and threshold transfer remain hard under drift.
    • Many promising methods still lack broad workload validation.

3) Technical synthesis

This issue is intentionally explicit about its evidence basis. The local archive contains the full selected-paper list and abstracts for the date, but only 10 completed deepreads. So the framing below is a synthesis from those completed analyses plus the broader selected-paper metadata, not a claim of equal reading depth across all 30 selected papers.

The clearest research movement is importing systems discipline into agent state. MemTX treats memory writes as tentative beliefs that need commit checks, action gating, and rollback. APPA treats unsafe reads as taint events that should be confined to disposable child branches unless a checked derivative comes back. Beyond Aggregate Risk makes the same point at the tool-call level: rare but high-risk argument roles deserve their own guarantees. These papers do not try to make agents vaguely safer; they redefine the unit of control.

The second movement is refusing to trust endpoint metrics. ContainmentBench shows that matched zero-harm endpoints can still conceal sharply different propagation and utility profiles. Success Is Not Self-Explanatory argues that correct answers are not enough without success provenance. When LLM Defenses Backfire shows that defense choice is a three-way deployment trade-off among safety, over-refusal, and cost. Together they make a strong case that the next generation of agent evaluation has to be trace-aware and side-effect-aware.

The third movement is narrow contracts over universal guardrails. Budget-aware Active RAG, evidence-bound code repair, falsifiable commitment planning, and speculative sandbox scheduling all work by defining exactly what gets measured and certified at each step. That is a quieter but more believable path to reliable autonomy than promising a single defense layer that solves everything.

4) Top papers

  1. MemTX: Transactional Belief Commit for Stateful Agent Memory
    Best first read if you care about what happens after an agent writes something wrong to shared memory. The paper’s contribution is practical and legible: belief commit, action gating, and typed rollback become explicit runtime mechanics.
    Caveat: the verification is bounded and only repairs what the system has actually recorded in provenance.

  2. Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
    Strong companion paper because it solves a real deployment tension: how do you inspect untrusted content without permanently poisoning the parent context? The branch-and-sanitize design is more concrete than most prompt-injection defenses.
    Caveat: the threat model still relies on trusted sanitizers and excludes covert channels.

  3. When LLM Defenses Backfire
    Worth reading if you ever have to choose a defense under latency, refusal, or budget constraints. Its main value is operational: it shows that “best defense” depends on what collateral damage you can tolerate.
    Caveat: the benchmark scope is still limited to current defense families and open-source mid-sized models.

  4. ContainmentBench
    Important because it makes post-injection containment measurable as a trajectory object rather than just a pass/fail label. That is exactly the level at which many real agent incidents differ.
    Caveat: evidence comes from a synthetic single-model study with a fixed authorization-ledger assumption.

  5. Beyond Aggregate Risk
    A smart tool-safety paper: it argues that the dangerous parts of a tool call should not hide behind whole-action averages. That is a crisp, reusable idea for structured-action agents.
    Caveat: rare roles still depend on pooled fallback if calibration data are thin.