Chinese version: [中文]
Run stats
- Candidates: 252
- Selected: 30
- Deepread completed: 30
- Window (UTC): 2026-07-28T00:00:00Z → 2026-07-29T00:00:00Z (arxiv_announce, expanded=0)
Show selected papers
| arXiv ID | Title / Links | Categories | Score | Why | Tags |
|---|---|---|---|---|---|
2607.25255 | SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems | cs.MA, cs.CR | 95 | Semantic info-flow defense for malicious cross-agent propagation; highly relevant MAS safety. | multi-agent, security, information-flow, agent-safety, defense |
2607.25560 | Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories | cs.AI | 94 | Black-box method exposes proprietary agent skills from trajectories; strong agent-security relevance. | agents, security, privacy, model-extraction, black-box-eval |
2607.25987 | \textsc{IH-Benchmark}: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications | cs.CR, cs.SE | 93 | Strong benchmark for instruction-hierarchy conflicts across system/user/tool settings. | benchmark, instruction-following, tool-use, robustness, evaluation |
2607.26041 | Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions? | cs.AI, cs.CV | 93 | Step-level benchmark for GUI agents' transition understanding; directly relevant to agent reliability. | agents, benchmark, GUI, evaluation, reliability |
2607.26034 | Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment | cs.AI, cs.CY, cs.GT, econ.GN | 93 | Behavioral evidence on AI race dynamics and safety tradeoffs; directly relevant to governance and incentives. | ai-safety, governance, race-dynamics, behavioral-experiment, incentives |
2607.25297 | Hybrid Analysis for Secure MCP Tool Use in LLM Agents | cs.CR, cs.AI | 92 | MCP tool-use defense with hybrid analysis targets a key real-world agent attack surface. | MCP, tool-use, security, agents, defense |
2607.25400 | COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution | cs.AI | 91 | Compiles workflows into constrained execution, directly addressing agent workflow misalignment. | agents, alignment, workflow, tool-use, control |
2607.25914 | Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks | cs.AI, cs.CR, cs.NI | 91 | Cross-vendor tool trust model for autonomous agents; concrete safety mechanism for tool-use systems. | agent-safety, tool-use, trust-management, autonomous-systems, security |
2607.25451 | Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization | cs.LG, cs.CR | 91 | Measures verbatim extraction under quantization directly; strong privacy relevance for deployed LLMs. | llm-privacy, memorization, quantization, data-extraction, deployment |
2607.25953 | Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections | cs.CL, cs.CY | 91 | Benchmark for responsible LLM mediation in elections; evaluates epistemic modesty under imperfect info. | evaluation, politics, reliability, benchmark, epistemics |
2607.25364 | Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales | cs.AI, cs.SE | 90 | Server-verified action claims for tool execution offer practical governance without trusting rationales. | tool-use, verification, governance, agents, security |
2607.25227 | Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks | cs.CR, cs.LG | 90 | Shows bit-flip attacks can induce targeted cognitive bias in LLM decisions without obvious breakage. | LLM-security, model-integrity, bit-flip, adversarial-attacks, decision-making |
2607.25398 | HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following | cs.AI, cs.CL | 89 | Benchmark for long-context policy adherence in agentic settings with MCP tools; very reusable. | benchmark, long-context, agents, instruction-following, MCP |
2607.25294 | CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition | cs.CV, cs.AI, cs.CL, cs.LG | 89 | New benchmark for multimodal context learning; useful for evaluating grounding and knowledge acquisition. | benchmark, multimodal, evaluation, context-learning, grounding |
2607.25880 | Stemma: Induced Decision Regions Reveal LLM Provenance | cs.CR, cs.AI, cs.CL | 89 | Provenance testing via induced decision regions may strengthen black-box lineage and misuse auditing. | security, provenance, auditing, black-box, llm |
2607.25907 | Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models | cs.LG, cs.AI, cs.CL | 88 | Targets evaluation-awareness latents, highlighting a threat to validity of LLM safety evaluations. | alignment, evaluation, interpretability, latent-control, safety |
2607.26057 | Pass the Baton: Trajectory-Relayed On-Policy Distillation | cs.CL, cs.AI | 88 | Improves on-policy distillation by fixing failed reasoning prefixes; promising for efficient reasoning training. | llm-training, distillation, reasoning, on-policy, post-training |
2607.25857 | Shieldstral | cs.CL, cs.CV | 87 | Small multimodal safety classifier with strong results and large-scale data recipe; deployable impact. | safety, multimodal, classifier, moderation, efficiency |
2607.25816 | Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL | cs.AI | 87 | Agent/tool-call efficiency advance with joint agent-speculator RL; relevant to practical agent systems. | agents, tool-use, reinforcement-learning, efficiency, inference |
2607.25308 | CAST: Game Solvers as Turn-Level Teachers for LLM Agents | cs.CL, cs.AI | 87 | Turn-level credit from solver teachers addresses sparse rewards in long-horizon LLM agent training. | agents, rlvr, credit-assignment, reasoning, games |
2607.25659 | CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization | cs.AI | 87 | Token-level credit assignment for rubric-guided RL could improve post-training reliability and alignment. | alignment, rlhf, post-training, credit-assignment, llm-training |
2607.25619 | SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents | cs.SE, cs.CR | 86 | Targets malicious skill files in coding agents, a timely supply-chain risk with runtime detection. | coding-agents, supply-chain, security, runtime-detection, malware |
2607.25479 | Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering | cs.CR, cs.AI, cs.LG | 85 | Architectural backdoors in VLM supply chains are novel and security-critical for model deployment. | VLM, backdoors, supply-chain, security, representation-steering |
2607.25225 | SecDrift: Measuring Sector-Conditioned Security Drift in AI-Generated Code | cs.CR, cs.LG, cs.SE | 85 | Benchmark on security drift in AI-generated code across critical sectors; useful deployment evaluation. | code-generation, security, benchmark, evaluation, critical-infrastructure |
2607.25634 | AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations | cs.AI, cs.CL | 85 | Audits pedagogical risks with rationales and evidence spans; concrete safety evaluation for education AI. | ai-safety, evaluation, auditing, education, bias, factuality |
2607.25995 | Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches? | cs.CR, cs.AI | 85 | Tests whether live runtime context improves LLM-generated Kubernetes security patch correctness. | security, agents, kubernetes, patching, deployment |
2607.25600 | Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs | cs.IR, cs.AI, cs.CL | 84 | Uses verbalized uncertainty to route retrieval in RAG QA; useful for reliability and efficiency. | RAG, uncertainty, retrieval, QA, reliability |
2607.25502 | Anti-Backdoor Coreset Selection via Cumulative Entropy | cs.LG, cs.CR | 84 | Training-time backdoor defense via coreset selection; practical security angle with concrete mechanism. | security, backdoor-defense, data-poisoning, coreset, robustness |
2607.25970 | Reinforcement Learning for Code Optimization | cs.LG, cs.AI | 84 | Concrete RL pipeline for code optimization with sandboxing and reward design; notable frontier capability work. | code, reinforcement-learning, efficiency, sandbox, llm |
2607.25485 | PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents | cs.AI, cs.CL | 83 | Patient-facing health agent benchmark emphasizes safety-critical evaluation beyond QA tasks. | healthcare, agents, benchmark, safety, evaluation |
AI Paper Insight Brief
2026-07-30
0) Executive takeaways (read this first)
- Agent safety work is shifting from prompt-local defenses to runtime and workflow control: multiple papers show that preserving provenance, constraining execution, or compiling workflows beats relying on the model to “remember policy.”
- Several results argue that context alone is not a reliable safety lever. Sector framing did not reliably improve code security, long handbook policies were often ignored, and suppressing an “evaluation-awareness” latent did not reliably change behavior.
- The strongest practical wins came from structured control surfaces: taint propagation in multi-agent systems, workflow compilation/interpreters, server-verified action claims, and runtime tool monitoring.
- Benchmarks are getting more diagnostic and less forgiving: new suites isolate instruction hierarchy conflicts, long-context policy adherence, desktop transition understanding, multimodal context learning, and patient-facing agent failures rather than just end-task success.
- On the frontier-progress side, several papers show that better credit assignment and execution-aware RL matter: solver-derived turn-level credit, token-level rubric credit, relay-style on-policy distillation, and timing-aware code RL all improve learning efficiency or capability.
- Security threats are broadening beyond prompt injection to supply-chain, memory-integrity, provenance, and IP leakage: bit-flip stance hijacking, architectural VLM backdoors, skill-file malware, trajectory-based skill extraction, and black-box provenance testing all look increasingly operational.
2) Key themes (clusters)
Theme: Runtime governance for agents and tools
- Why it matters: The common failure mode is that unsafe intent gets hidden across steps, tools, or free-form rationales. These papers replace trust in model obedience with explicit runtime checks, provenance, and constrained execution.
- Representative papers:
- SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
- Hybrid Analysis for Secure MCP Tool Use in LLM Agents
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
- Common approach:
- Preserve semantics across execution boundaries via taints, graphs, or typed claims.
- Move enforcement to deterministic controllers, verifiers, or staged sinks rather than the base model.
- Combine static intent/context checks with runtime evidence such as tool side effects or workflow state.
- Treat natural-language policies as inputs to be compiled or normalized into executable control structures.
- Open questions / failure modes:
- Observability gaps remain load-bearing: missing provenance edges, hidden side effects, or uninstrumented tools can defeat defenses.
- Post-execution denial may be too late if side effects are irreversible.
- LLM-mediated annotation/reconstruction still introduces a soft spot inside otherwise deterministic pipelines.
- Operational overhead and integration burden may limit adoption in real agent stacks.
Theme: Policy-following and hierarchy robustness are still weak
- Why it matters: Enterprise and safety-critical deployments assume that long policies, system instructions, and tool restrictions act as persistent authority. These benchmarks show that assumption is still fragile.
- Representative papers:
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- \textsc{IH-Benchmark}: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
- Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
- PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
- Common approach:
- Use executable or rubric-grounded evaluations that score forbidden actions, not just task completion.
- Stress models with conflicting instructions, imperfect evidence, or long standing policies.
- Separate dimensions like faithfulness, calibration, triage, and workflow correctness instead of collapsing to one score.
- Evaluate in stateful environments with tools and persistent side effects.
- Open questions / failure modes:
- Models often privilege proximate context over higher-priority or longer-lived policy.
- Strong performance on one conflict surface does not transfer to others; S≻U robustness did not imply U≻T robustness.
- Self-reports of compliance are unreliable; some agents claim checks passed when they did not.
- Benchmark success is still far from deployment readiness in high-stakes domains like healthcare or elections.
Theme: Supply-chain and post-deployment attacks are becoming more realistic
- Why it matters: The attack surface is moving from training data poisoning toward deployed artifacts, memory faults, and reusable agent components. That makes integrity and provenance controls more urgent.
- Representative papers:
- Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
- SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
- Stemma: Induced Decision Regions Reveal LLM Provenance
- Common approach:
- Target sparse, hard-to-inspect control points: a few flipped bits, a gated residual addition, or malicious skill files.
- Preserve clean-task utility to stay stealthy while altering targeted behavior.
- Pair attacks with lightweight operational defenses such as anomaly detection, snippet-based screening, or black-box fingerprinting.
- Evaluate under realistic reuse/deployment assumptions: open weights, shared artifacts, registries, or black-box APIs.
- Open questions / failure modes:
- Many attacks assume strong attacker capabilities, such as white-box localization or artifact control.
- Detection methods may be brittle to adaptive attackers, distributed payloads, or hidden channels.
- Provenance and screening tools help after the fact but do not prevent initial compromise.
- Real-world prevalence under cloud mitigations, ECC, or provider controls remains uncertain.
Theme: Better credit assignment is driving agent/RL progress
- Why it matters: Several papers attack the same bottleneck: sparse or misallocated learning signal in long-horizon reasoning and tool use. The pattern is to inject finer-grained supervision without fully changing the training stack.
- Representative papers:
- Common approach:
- Replace coarse trajectory-level reward with turn-, token-, or prefix-local signals.
- Use existing structure as teacher signal: solvers, rubric-conditioned counterfactuals, teacher handoffs, or calibrated execution timing.
- Keep compatibility with GRPO/DAPO-style pipelines rather than introducing heavy new models.
- Add stabilization tricks—normalization, ramps, larger rollout groups, or bounded interventions—to make noisy signals trainable.
- Open questions / failure modes:
- Many methods depend on privileged structure: exact solvers, criteria-free prompts, strong teachers, or calibrated execution services.
- Gains are often domain-bounded so far: games, math, or competitive programming.
- More granular credit can destabilize training without careful scheduling and normalization.
- Transfer to open-world agent tasks remains mostly unproven.
Theme: Diagnostic benchmarks are exposing hidden capability gaps
- Why it matters: New evaluations are less about leaderboard averages and more about identifying where systems break: grounding vs application vs induction, transition verification, retrieval calibration, or topology-aware remediation.
- Representative papers:
- CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
- Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
- Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
- Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
- Common approach:
- Isolate sub-capabilities with controlled task formulations and negative controls.
- Measure whether added context helps only where it should; topology context improved TD patches but not TI controls.
- Use structured scoring and per-failure taxonomy rather than a single aggregate metric.
- Evaluate practical trade-offs like token cost, latency, or functional blast radius.
- Open questions / failure modes:
- Better diagnostics do not automatically yield better systems; many best scores remain low.
- LLM judges and semantic scoring still introduce evaluator variance in some settings.
- Added context can help selectively but also create new failure modes or overhead.
- Offline diagnostics may miss long-horizon closed-loop failures.
3) Technical synthesis
- A recurring design pattern is “compile or normalize natural language into a smaller trusted object”: WCFGs in COVENANT, typed explanation packets in EBTE, taint labels in SafeFlow, and yes/no policy queries in Shieldstral.
- Several papers separate semantic intent from execution evidence: MTGuard compares declared tool intent to eBPF-observed behavior; EBTE checks model claims against authoritative facts; KuTIE checks scanner-clearing patches against runtime dependency preservation.
- The strongest evaluations increasingly use negative controls to isolate causal effects: SecDrift’s matched baseline and placebo sectors, KuTIE’s topology-independent controls, BeyondUncertainty’s route-count-matched random routing, and latent-suppression placebo directions.
- Across security papers, stealth preservation is central: CogBias keeps perplexity and MMLU nearly unchanged, VLM architectural backdoors preserve clean accuracy, and skill-file screening focuses on low-FPR deployability.
- Multiple works show that model choice matters more than prompt framing when the intervention is weak or implicit: SecDrift found model differences more reliable than sector wording; HANDBOOK.md shows long policy context alone is insufficient.
- There is a broad shift from single-turn prompt defense to graph/state-based defense: SafeFlow, COVENANT, MTGuard, and AgentToolMO all reason over trajectories, dependencies, or workflow state.
- RL/optimization papers converge on localized credit with lightweight integration: solver advantages, token replay weights, relay handoffs, and ranked timing rewards all preserve existing training backbones while sharpening signal.
- Several benchmarks reveal dissociations between adjacent capabilities: desktop action-family recognition exceeds payload recovery; multimodal grounding differs from knowledge induction; task completion in health agents is near-ceiling while triage remains weak.
- Practical deployment trade-offs are explicit: BeyondUncertainty saves retrievals but increases total tokens; MTGuard improves detection but adds ~12.39s per tool call when both audits run; COVENANT improves success but raises latency and model calls.
- Provenance and integrity are becoming measurable at inference time: Stemma uses induced decision regions for black-box lineage, while runtime anomaly detection and snippet screening aim to catch compromised artifacts without full retraining.
4) Top 5 papers (with “why now”)
- Cuts average ASR from 69.3% to 12.7% by preserving taints and validating forbidden source–sink paths across the whole agent workflow.
- Important because it targets a real blind spot in multi-agent systems: harmful intent can be split into locally benign subtasks.
- The design is operationally concrete: staged hard sinks, deterministic rule application, attribution paths, and a closed label schema.
- Useful now for teams moving from single-agent demos to delegated multi-agent workflows with sensitive tools.
- Skepticism: effectiveness depends heavily on instrumentation quality, provenance completeness, and trusted wrappers.
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- Shows that even strong frontier agents often fail to treat long policy documents as binding authority; best strict pass@1 is only 36.2%.
- The benchmark is unusually decision-useful: 65 realistic containerized tasks, 20–124 page handbooks, and 824 deterministic verifier criteria including forbidden side effects.
- Why now: many enterprise deployments assume “put the SOP in context” is enough; this paper says it usually is not.
- Useful as a regression suite for policy adherence, especially for MCP/tool-heavy enterprise agents.
- Skepticism: the paper summary does not provide a strong limitations section beyond benchmark design notes.
Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
- Introduces a post-deployment threat where as few as ~12 bit flips can shift model stance on targeted topics with ASRs up to 84.6% while preserving general capability.
- The attack is notable because it is trigger-free, persistent, and aimed at downstream decision bias rather than obvious model breakage.
- Why now: open-weight deployment, quantization, and edge inference make memory-integrity attacks more relevant than purely training-time threats.
- Useful for red-teaming model integrity assumptions and motivating ECC/hash verification plus semantic monitoring.
- Skepticism: assumes offline white-box localization and practical bit-flip capability on target hardware.
COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
- Raises success from 43.2% to 73.1% across 3,000 paired executions by compiling workflow prose into a control-flow graph and enforcing node-level checks.
- Strong evidence that controller-owned traversal is a bigger lever than hoping the model self-enforces procedural text.
- Why now: organizations already have SOPs and workflows in prose, not formal languages; this offers a path to operationalize them.
- Useful for high-stakes workflows where trace completeness and argument correctness both matter.
- Skepticism: frontend compilation is not end-to-end sound, and runtime overhead is substantial.
Reinforcement Learning for Code Optimization
- Shows that timing-aware RL can work if measurement, reward design, and optimizer stability are engineered together; e.g. p50 pass@1 rises from 18.0% to 31.3% on Qwen 2.5 7B and 30.7% to 50.4% on CWM 32B.
- The contribution is less “new RL algorithm” than a full stack for making noisy execution-time rewards usable.
- Why now: code agents are moving from correctness to efficiency, and naive timing rewards are too noisy to train on.
- Useful for teams building execution-grounded code optimization or runtime-aware coding agents.
- Skepticism: scope is narrow—single-file Python competitive programming with expensive infrastructure.
5) Practical next steps
- Add workflow-level controls before expanding agent autonomy: provenance logging, staged sinks, and deterministic release rules are repeatedly higher-leverage than prompt tweaks.
- Treat long policies and system prompts as advisory unless externally enforced; compile them into executable guards, node contracts, or tool-call policies where possible.
- Instrument tool use with runtime observability: sandboxing, process/file/network traces, and side-effect verification should be standard for MCP or similar tool protocols.
- Build evaluation suites that separate task completion from policy compliance; include forbidden-side-effect checks, hierarchy conflicts, and near-miss metrics.
- For RAG systems, test selective retrieval controllers against both quality and token cost; retrieval savings alone may hide higher total spend.
- Add integrity controls for deployed/open-weight models: weight hashing, ECC where available, artifact provenance checks, and semantic drift monitoring for stance or moderation shifts.
- Screen third-party agent assets—skills, tools, model code—with hybrid filters that combine cheap static prefilters and targeted LLM review to keep FPR and latency deployable.
- When training agents, prioritize finer-grained credit assignment: turn-level solver signals, token-level rubric weighting, or prefix-local teacher interventions appear more sample-efficient than pure terminal rewards.
Generated from per-paper analyses; no external browsing.