Chinese version: [中文]

Run stats

  • Candidates: 252
  • Selected: 30
  • Deepread completed: 30
  • Window (UTC): 2026-07-28T00:00:00Z → 2026-07-29T00:00:00Z (arxiv_announce, expanded=0)
Show selected papers
arXiv IDTitle / LinksCategoriesScoreWhyTags
2607.25255SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
PDF
cs.MA, cs.CR95Semantic info-flow defense for malicious cross-agent propagation; highly relevant MAS safety.multi-agent, security, information-flow, agent-safety, defense
2607.25560Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
PDF
cs.AI94Black-box method exposes proprietary agent skills from trajectories; strong agent-security relevance.agents, security, privacy, model-extraction, black-box-eval
2607.25987\textsc{IH-Benchmark}: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
PDF
cs.CR, cs.SE93Strong benchmark for instruction-hierarchy conflicts across system/user/tool settings.benchmark, instruction-following, tool-use, robustness, evaluation
2607.26041Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
PDF
cs.AI, cs.CV93Step-level benchmark for GUI agents' transition understanding; directly relevant to agent reliability.agents, benchmark, GUI, evaluation, reliability
2607.26034Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
PDF
cs.AI, cs.CY, cs.GT, econ.GN93Behavioral evidence on AI race dynamics and safety tradeoffs; directly relevant to governance and incentives.ai-safety, governance, race-dynamics, behavioral-experiment, incentives
2607.25297Hybrid Analysis for Secure MCP Tool Use in LLM Agents
PDF
cs.CR, cs.AI92MCP tool-use defense with hybrid analysis targets a key real-world agent attack surface.MCP, tool-use, security, agents, defense
2607.25400COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
PDF
cs.AI91Compiles workflows into constrained execution, directly addressing agent workflow misalignment.agents, alignment, workflow, tool-use, control
2607.25914Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
PDF
cs.AI, cs.CR, cs.NI91Cross-vendor tool trust model for autonomous agents; concrete safety mechanism for tool-use systems.agent-safety, tool-use, trust-management, autonomous-systems, security
2607.25451Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
PDF
cs.LG, cs.CR91Measures verbatim extraction under quantization directly; strong privacy relevance for deployed LLMs.llm-privacy, memorization, quantization, data-extraction, deployment
2607.25953Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
PDF
cs.CL, cs.CY91Benchmark for responsible LLM mediation in elections; evaluates epistemic modesty under imperfect info.evaluation, politics, reliability, benchmark, epistemics
2607.25364Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
PDF
cs.AI, cs.SE90Server-verified action claims for tool execution offer practical governance without trusting rationales.tool-use, verification, governance, agents, security
2607.25227Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
PDF
cs.CR, cs.LG90Shows bit-flip attacks can induce targeted cognitive bias in LLM decisions without obvious breakage.LLM-security, model-integrity, bit-flip, adversarial-attacks, decision-making
2607.25398HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
PDF
cs.AI, cs.CL89Benchmark for long-context policy adherence in agentic settings with MCP tools; very reusable.benchmark, long-context, agents, instruction-following, MCP
2607.25294CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
PDF
cs.CV, cs.AI, cs.CL, cs.LG89New benchmark for multimodal context learning; useful for evaluating grounding and knowledge acquisition.benchmark, multimodal, evaluation, context-learning, grounding
2607.25880Stemma: Induced Decision Regions Reveal LLM Provenance
PDF
cs.CR, cs.AI, cs.CL89Provenance testing via induced decision regions may strengthen black-box lineage and misuse auditing.security, provenance, auditing, black-box, llm
2607.25907Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
PDF
cs.LG, cs.AI, cs.CL88Targets evaluation-awareness latents, highlighting a threat to validity of LLM safety evaluations.alignment, evaluation, interpretability, latent-control, safety
2607.26057Pass the Baton: Trajectory-Relayed On-Policy Distillation
PDF
cs.CL, cs.AI88Improves on-policy distillation by fixing failed reasoning prefixes; promising for efficient reasoning training.llm-training, distillation, reasoning, on-policy, post-training
2607.25857Shieldstral
PDF
cs.CL, cs.CV87Small multimodal safety classifier with strong results and large-scale data recipe; deployable impact.safety, multimodal, classifier, moderation, efficiency
2607.25816Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
PDF
cs.AI87Agent/tool-call efficiency advance with joint agent-speculator RL; relevant to practical agent systems.agents, tool-use, reinforcement-learning, efficiency, inference
2607.25308CAST: Game Solvers as Turn-Level Teachers for LLM Agents
PDF
cs.CL, cs.AI87Turn-level credit from solver teachers addresses sparse rewards in long-horizon LLM agent training.agents, rlvr, credit-assignment, reasoning, games
2607.25659CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
PDF
cs.AI87Token-level credit assignment for rubric-guided RL could improve post-training reliability and alignment.alignment, rlhf, post-training, credit-assignment, llm-training
2607.25619SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
PDF
cs.SE, cs.CR86Targets malicious skill files in coding agents, a timely supply-chain risk with runtime detection.coding-agents, supply-chain, security, runtime-detection, malware
2607.25479Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
PDF
cs.CR, cs.AI, cs.LG85Architectural backdoors in VLM supply chains are novel and security-critical for model deployment.VLM, backdoors, supply-chain, security, representation-steering
2607.25225SecDrift: Measuring Sector-Conditioned Security Drift in AI-Generated Code
PDF
cs.CR, cs.LG, cs.SE85Benchmark on security drift in AI-generated code across critical sectors; useful deployment evaluation.code-generation, security, benchmark, evaluation, critical-infrastructure
2607.25634AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations
PDF
cs.AI, cs.CL85Audits pedagogical risks with rationales and evidence spans; concrete safety evaluation for education AI.ai-safety, evaluation, auditing, education, bias, factuality
2607.25995Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
PDF
cs.CR, cs.AI85Tests whether live runtime context improves LLM-generated Kubernetes security patch correctness.security, agents, kubernetes, patching, deployment
2607.25600Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
PDF
cs.IR, cs.AI, cs.CL84Uses verbalized uncertainty to route retrieval in RAG QA; useful for reliability and efficiency.RAG, uncertainty, retrieval, QA, reliability
2607.25502Anti-Backdoor Coreset Selection via Cumulative Entropy
PDF
cs.LG, cs.CR84Training-time backdoor defense via coreset selection; practical security angle with concrete mechanism.security, backdoor-defense, data-poisoning, coreset, robustness
2607.25970Reinforcement Learning for Code Optimization
PDF
cs.LG, cs.AI84Concrete RL pipeline for code optimization with sandboxing and reward design; notable frontier capability work.code, reinforcement-learning, efficiency, sandbox, llm
2607.25485PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
PDF
cs.AI, cs.CL83Patient-facing health agent benchmark emphasizes safety-critical evaluation beyond QA tasks.healthcare, agents, benchmark, safety, evaluation

AI Paper Insight Brief

2026-07-30

0) Executive takeaways (read this first)

  • Agent safety work is shifting from prompt-local defenses to runtime and workflow control: multiple papers show that preserving provenance, constraining execution, or compiling workflows beats relying on the model to “remember policy.”
  • Several results argue that context alone is not a reliable safety lever. Sector framing did not reliably improve code security, long handbook policies were often ignored, and suppressing an “evaluation-awareness” latent did not reliably change behavior.
  • The strongest practical wins came from structured control surfaces: taint propagation in multi-agent systems, workflow compilation/interpreters, server-verified action claims, and runtime tool monitoring.
  • Benchmarks are getting more diagnostic and less forgiving: new suites isolate instruction hierarchy conflicts, long-context policy adherence, desktop transition understanding, multimodal context learning, and patient-facing agent failures rather than just end-task success.
  • On the frontier-progress side, several papers show that better credit assignment and execution-aware RL matter: solver-derived turn-level credit, token-level rubric credit, relay-style on-policy distillation, and timing-aware code RL all improve learning efficiency or capability.
  • Security threats are broadening beyond prompt injection to supply-chain, memory-integrity, provenance, and IP leakage: bit-flip stance hijacking, architectural VLM backdoors, skill-file malware, trajectory-based skill extraction, and black-box provenance testing all look increasingly operational.

2) Key themes (clusters)

Theme: Runtime governance for agents and tools

Theme: Policy-following and hierarchy robustness are still weak

Theme: Supply-chain and post-deployment attacks are becoming more realistic

Theme: Better credit assignment is driving agent/RL progress

  • Why it matters: Several papers attack the same bottleneck: sparse or misallocated learning signal in long-horizon reasoning and tool use. The pattern is to inject finer-grained supervision without fully changing the training stack.
  • Representative papers:
  • Common approach:
    • Replace coarse trajectory-level reward with turn-, token-, or prefix-local signals.
    • Use existing structure as teacher signal: solvers, rubric-conditioned counterfactuals, teacher handoffs, or calibrated execution timing.
    • Keep compatibility with GRPO/DAPO-style pipelines rather than introducing heavy new models.
    • Add stabilization tricks—normalization, ramps, larger rollout groups, or bounded interventions—to make noisy signals trainable.
  • Open questions / failure modes:
    • Many methods depend on privileged structure: exact solvers, criteria-free prompts, strong teachers, or calibrated execution services.
    • Gains are often domain-bounded so far: games, math, or competitive programming.
    • More granular credit can destabilize training without careful scheduling and normalization.
    • Transfer to open-world agent tasks remains mostly unproven.

Theme: Diagnostic benchmarks are exposing hidden capability gaps

3) Technical synthesis

  • A recurring design pattern is “compile or normalize natural language into a smaller trusted object”: WCFGs in COVENANT, typed explanation packets in EBTE, taint labels in SafeFlow, and yes/no policy queries in Shieldstral.
  • Several papers separate semantic intent from execution evidence: MTGuard compares declared tool intent to eBPF-observed behavior; EBTE checks model claims against authoritative facts; KuTIE checks scanner-clearing patches against runtime dependency preservation.
  • The strongest evaluations increasingly use negative controls to isolate causal effects: SecDrift’s matched baseline and placebo sectors, KuTIE’s topology-independent controls, BeyondUncertainty’s route-count-matched random routing, and latent-suppression placebo directions.
  • Across security papers, stealth preservation is central: CogBias keeps perplexity and MMLU nearly unchanged, VLM architectural backdoors preserve clean accuracy, and skill-file screening focuses on low-FPR deployability.
  • Multiple works show that model choice matters more than prompt framing when the intervention is weak or implicit: SecDrift found model differences more reliable than sector wording; HANDBOOK.md shows long policy context alone is insufficient.
  • There is a broad shift from single-turn prompt defense to graph/state-based defense: SafeFlow, COVENANT, MTGuard, and AgentToolMO all reason over trajectories, dependencies, or workflow state.
  • RL/optimization papers converge on localized credit with lightweight integration: solver advantages, token replay weights, relay handoffs, and ranked timing rewards all preserve existing training backbones while sharpening signal.
  • Several benchmarks reveal dissociations between adjacent capabilities: desktop action-family recognition exceeds payload recovery; multimodal grounding differs from knowledge induction; task completion in health agents is near-ceiling while triage remains weak.
  • Practical deployment trade-offs are explicit: BeyondUncertainty saves retrievals but increases total tokens; MTGuard improves detection but adds ~12.39s per tool call when both audits run; COVENANT improves success but raises latency and model calls.
  • Provenance and integrity are becoming measurable at inference time: Stemma uses induced decision regions for black-box lineage, while runtime anomaly detection and snippet screening aim to catch compromised artifacts without full retraining.

4) Top 5 papers (with “why now”)

SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

  • Cuts average ASR from 69.3% to 12.7% by preserving taints and validating forbidden source–sink paths across the whole agent workflow.
  • Important because it targets a real blind spot in multi-agent systems: harmful intent can be split into locally benign subtasks.
  • The design is operationally concrete: staged hard sinks, deterministic rule application, attribution paths, and a closed label schema.
  • Useful now for teams moving from single-agent demos to delegated multi-agent workflows with sensitive tools.
  • Skepticism: effectiveness depends heavily on instrumentation quality, provenance completeness, and trusted wrappers.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

  • Shows that even strong frontier agents often fail to treat long policy documents as binding authority; best strict pass@1 is only 36.2%.
  • The benchmark is unusually decision-useful: 65 realistic containerized tasks, 20–124 page handbooks, and 824 deterministic verifier criteria including forbidden side effects.
  • Why now: many enterprise deployments assume “put the SOP in context” is enough; this paper says it usually is not.
  • Useful as a regression suite for policy adherence, especially for MCP/tool-heavy enterprise agents.
  • Skepticism: the paper summary does not provide a strong limitations section beyond benchmark design notes.

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

  • Introduces a post-deployment threat where as few as ~12 bit flips can shift model stance on targeted topics with ASRs up to 84.6% while preserving general capability.
  • The attack is notable because it is trigger-free, persistent, and aimed at downstream decision bias rather than obvious model breakage.
  • Why now: open-weight deployment, quantization, and edge inference make memory-integrity attacks more relevant than purely training-time threats.
  • Useful for red-teaming model integrity assumptions and motivating ECC/hash verification plus semantic monitoring.
  • Skepticism: assumes offline white-box localization and practical bit-flip capability on target hardware.

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

  • Raises success from 43.2% to 73.1% across 3,000 paired executions by compiling workflow prose into a control-flow graph and enforcing node-level checks.
  • Strong evidence that controller-owned traversal is a bigger lever than hoping the model self-enforces procedural text.
  • Why now: organizations already have SOPs and workflows in prose, not formal languages; this offers a path to operationalize them.
  • Useful for high-stakes workflows where trace completeness and argument correctness both matter.
  • Skepticism: frontend compilation is not end-to-end sound, and runtime overhead is substantial.

Reinforcement Learning for Code Optimization

  • Shows that timing-aware RL can work if measurement, reward design, and optimizer stability are engineered together; e.g. p50 pass@1 rises from 18.0% to 31.3% on Qwen 2.5 7B and 30.7% to 50.4% on CWM 32B.
  • The contribution is less “new RL algorithm” than a full stack for making noisy execution-time rewards usable.
  • Why now: code agents are moving from correctness to efficiency, and naive timing rewards are too noisy to train on.
  • Useful for teams building execution-grounded code optimization or runtime-aware coding agents.
  • Skepticism: scope is narrow—single-file Python competitive programming with expensive infrastructure.

5) Practical next steps

  • Add workflow-level controls before expanding agent autonomy: provenance logging, staged sinks, and deterministic release rules are repeatedly higher-leverage than prompt tweaks.
  • Treat long policies and system prompts as advisory unless externally enforced; compile them into executable guards, node contracts, or tool-call policies where possible.
  • Instrument tool use with runtime observability: sandboxing, process/file/network traces, and side-effect verification should be standard for MCP or similar tool protocols.
  • Build evaluation suites that separate task completion from policy compliance; include forbidden-side-effect checks, hierarchy conflicts, and near-miss metrics.
  • For RAG systems, test selective retrieval controllers against both quality and token cost; retrieval savings alone may hide higher total spend.
  • Add integrity controls for deployed/open-weight models: weight hashing, ECC where available, artifact provenance checks, and semantic drift monitoring for stance or moderation shifts.
  • Screen third-party agent assets—skills, tools, model code—with hybrid filters that combine cheap static prefilters and targeted LLM review to keep FPR and latency deployable.
  • When training agents, prioritize finer-grained credit assignment: turn-level solver signals, token-level rubric weighting, or prefix-local teacher interventions appear more sample-efficient than pure terminal rewards.

Generated from per-paper analyses; no external browsing.