Chinese version: [中文]

Run stats

  • Candidates: 252
  • Selected: 30
  • Deepread completed: 0 (not yet run)
  • Window (UTC): 2026-07-28T00:00:00Z → 2026-07-29T00:00:00Z (arxiv_announce, expanded=0)
  • Synthesis basis: selected-paper set plus candidate titles/abstracts only; no full-paper deepread pass yet.
Show selected papers
arXiv IDTitle / LinksCategoriesScoreWhyTags
2607.25255SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
PDF
cs.MA, cs.CR95Semantic info-flow defense for malicious cross-agent propagation; highly relevant MAS safety.multi-agent, security, information-flow, agent-safety, defense
2607.25560Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
PDF
cs.AI94Black-box method exposes proprietary agent skills from trajectories; strong agent-security relevance.agents, security, privacy, model-extraction, black-box-eval
2607.25987IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
PDF
cs.CR, cs.SE93Strong benchmark for instruction-hierarchy conflicts across system/user/tool settings.benchmark, instruction-following, tool-use, robustness, evaluation
2607.26041Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
PDF
cs.AI, cs.CV93Step-level benchmark for GUI agents' transition understanding; directly relevant to agent reliability.agents, benchmark, GUI, evaluation, reliability
2607.26034Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
PDF
cs.AI, cs.CY, cs.GT, econ.GN93Behavioral evidence on AI race dynamics and safety tradeoffs; directly relevant to governance and incentives.ai-safety, governance, race-dynamics, behavioral-experiment, incentives
2607.25297Hybrid Analysis for Secure MCP Tool Use in LLM Agents
PDF
cs.CR, cs.AI92MCP tool-use defense with hybrid analysis targets a key real-world agent attack surface.MCP, tool-use, security, agents, defense
2607.25400COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
PDF
cs.AI91Compiles workflows into constrained execution, directly addressing agent workflow misalignment.agents, alignment, workflow, tool-use, control
2607.25914Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
PDF
cs.AI, cs.CR, cs.NI91Cross-vendor tool trust model for autonomous agents; concrete safety mechanism for tool-use systems.agent-safety, tool-use, trust-management, autonomous-systems, security
2607.25451Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
PDF
cs.LG, cs.CR91Measures verbatim extraction under quantization directly; strong privacy relevance for deployed LLMs.llm-privacy, memorization, quantization, data-extraction, deployment
2607.25953Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
PDF
cs.CL, cs.CY91Benchmark for responsible LLM mediation in elections; evaluates epistemic modesty under imperfect info.evaluation, politics, reliability, benchmark, epistemics
2607.25364Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
PDF
cs.AI, cs.SE90Server-verified action claims for tool execution offer practical governance without trusting rationales.tool-use, verification, governance, agents, security
2607.25227Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
PDF
cs.CR, cs.LG90Shows bit-flip attacks can induce targeted cognitive bias in LLM decisions without obvious breakage.LLM-security, model-integrity, bit-flip, adversarial-attacks, decision-making
2607.25398HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
PDF
cs.AI, cs.CL89Benchmark for long-context policy adherence in agentic settings with MCP tools; very reusable.benchmark, long-context, agents, instruction-following, MCP
2607.25294CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
PDF
cs.CV, cs.AI, cs.CL, cs.LG89New benchmark for multimodal context learning; useful for evaluating grounding and knowledge acquisition.benchmark, multimodal, evaluation, context-learning, grounding
2607.25880Stemma: Induced Decision Regions Reveal LLM Provenance
PDF
cs.CR, cs.AI, cs.CL89Provenance testing via induced decision regions may strengthen black-box lineage and misuse auditing.security, provenance, auditing, black-box, llm
2607.25907Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
PDF
cs.LG, cs.AI, cs.CL88Targets evaluation-awareness latents, highlighting a threat to validity of LLM safety evaluations.alignment, evaluation, interpretability, latent-control, safety
2607.26057Pass the Baton: Trajectory-Relayed On-Policy Distillation
PDF
cs.CL, cs.AI88Improves on-policy distillation by fixing failed reasoning prefixes; promising for efficient reasoning training.llm-training, distillation, reasoning, on-policy, post-training
2607.25857Shieldstral
PDF
cs.CL, cs.CV87Small multimodal safety classifier with strong results and large-scale data recipe; deployable impact.safety, multimodal, classifier, moderation, efficiency
2607.25816Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
PDF
cs.AI87Agent/tool-call efficiency advance with joint agent-speculator RL; relevant to practical agent systems.agents, tool-use, reinforcement-learning, efficiency, inference
2607.25308CAST: Game Solvers as Turn-Level Teachers for LLM Agents
PDF
cs.CL, cs.AI87Turn-level credit from solver teachers addresses sparse rewards in long-horizon LLM agent training.agents, rlvr, credit-assignment, reasoning, games
2607.25659CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
PDF
cs.AI87Token-level credit assignment for rubric-guided RL could improve post-training reliability and alignment.alignment, rlhf, post-training, credit-assignment, llm-training
2607.25619SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
PDF
cs.SE, cs.CR86Targets malicious skill files in coding agents, a timely supply-chain risk with runtime detection.coding-agents, supply-chain, security, runtime-detection, malware
2607.25479Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
PDF
cs.CR, cs.AI, cs.LG85Architectural backdoors in VLM supply chains are novel and security-critical for model deployment.VLM, backdoors, supply-chain, security, representation-steering
2607.25225SecDrift: Measuring Sector-Conditioned Security Drift in AI-Generated Code
PDF
cs.CR, cs.LG, cs.SE85Benchmark on security drift in AI-generated code across critical sectors; useful deployment evaluation.code-generation, security, benchmark, evaluation, critical-infrastructure
2607.25634AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations
PDF
cs.AI, cs.CL85Audits pedagogical risks with rationales and evidence spans; concrete safety evaluation for education AI.ai-safety, evaluation, auditing, education, bias, factuality
2607.25995Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
PDF
cs.CR, cs.AI85Tests whether live runtime context improves LLM-generated Kubernetes security patch correctness.security, agents, kubernetes, patching, deployment
2607.25600Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
PDF
cs.IR, cs.AI, cs.CL84Uses verbalized uncertainty to route retrieval in RAG QA; useful for reliability and efficiency.RAG, uncertainty, retrieval, QA, reliability
2607.25502Anti-Backdoor Coreset Selection via Cumulative Entropy
PDF
cs.LG, cs.CR84Training-time backdoor defense via coreset selection; practical security angle with concrete mechanism.security, backdoor-defense, data-poisoning, coreset, robustness
2607.25970Reinforcement Learning for Code Optimization
PDF
cs.LG, cs.AI84Concrete RL pipeline for code optimization with sandboxing and reward design; notable frontier capability work.code, reinforcement-learning, efficiency, sandbox, llm
2607.25485PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
PDF
cs.AI, cs.CL83Patient-facing health agent benchmark emphasizes safety-critical evaluation beyond QA tasks.healthcare, agents, benchmark, safety, evaluation

AI Paper Insight Brief

2026-07-30

Method note: this is a publish-ready fast synthesis built from the 30-paper selected set and the broader 252-paper candidate window’s titles and abstracts. No deepread pass is completed yet, so treat all paper claims as abstract-level unless otherwise stated.

0) Executive takeaways (read this first)

  • The day’s center of gravity is operational agent safety. Workflow compilers, verified action-claim layers, MCP defenses, and cross-vendor trust models all try to make policy survive contact with runtime.
  • Evaluation is moving below aggregate task success. IH-Benchmark, Desktop-Delta Bench, PatientAgentBench, and Polistemics test whether agents obey hierarchy, verify progress, and stay reliable in safety-critical interactions.
  • The security perimeter is widening beyond model weights. Skill leakage, malicious skill files, VLM backdoors, quantized memorization, and topology-aware patching all treat deployment artifacts as first-class risk surfaces.

2) Key themes (clusters)

Theme: Enforceable control is replacing prompt-only alignment

Theme: Reliability evaluation is becoming step-level and conflict-aware

Theme: Supply-chain and deployment security are now core agent topics

3) Technical synthesis

  • A visible design shift is from instruction following to execution shaping: COVENANT compiles workflows, EBTE turns rationales into typed claims, and MCP security work mixes static and dynamic analysis so risky tool calls can be intercepted before execution.
  • Multi-agent safety is getting more semantic and structural. SafeFlow argues malicious intent can propagate through locally reasonable subgoals, while cross-vendor trust-management work treats agent-tool trust as standardized infrastructure rather than ad hoc metadata.
  • Evaluation papers repeatedly attack the same blind spot: a good endpoint can be produced for the wrong reason. IH-Benchmark tests priority obedience, Desktop-Delta tests transition understanding, and Polistemics tests whether models stay epistemically modest when evidence is noisy or contradictory.
  • The day’s security papers also widen what counts as the model boundary. Skill leakage from trajectories and malicious skill-file detection both imply that behavioral packaging is now part of the attack surface, not just pretrained weights.
  • Quantization and deployment are no longer “post-model” implementation details. Bits and Memories argues privacy should be measured with verbatim extraction, while the Kubernetes topology paper shows remediation quality can change sharply once live service dependencies are exposed.
  • Several capability papers still fit the day’s broader pattern: Relay-OPD, CAST, CoRT, and code-optimization RL are all about denser credit assignment and less wasted trajectory budget rather than brute-force scaling alone.
  • Shieldstral stands out as a reminder that small, deployable safety components remain strategically important. A 3B multimodal safety classifier with strong reported results may matter more operationally than another giant general model release.
  • The governance warning is unusually concrete today: the AI-race experiment suggests unsafe development can arise from competitive position and behavioral momentum, not just fixed individual risk preferences.

4) Top 5 papers (with “why now”)

COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution

  • Compiles natural-language workflow instructions into constrained execution paths instead of leaving both procedure choice and step execution inside the model.
  • Best interpreted as a systems answer to workflow misalignment: if policy matters, make it executable.
  • Why now: many deployed agents are governed by long handbook-style prompts; this paper asks whether those rules can become something stronger than context text.
  • Skeptical about / limitation: its protection depends on workflow expressiveness, compiler correctness, and whether real tool actions cleanly map onto the compiled control surface.

IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications

  • Evaluates both direct system-user conflict and tool-mediated user-tool conflict across a broad hand-authored taxonomy.
  • Especially valuable because it shows strong system-prompt compliance does not automatically imply robustness when tool outputs carry conflicting instructions.
  • Why now: production agents increasingly blend system prompts, policy files, and tool messages, so hierarchy robustness is becoming a deployment primitive.
  • Skeptical about / limitation: as with many benchmark papers, the main question is how well the curated conflict families and judging protocol transfer to messier live environments.

SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

  • Frames multi-agent safety as an information-flow problem: harmful objectives can be fragmented into plausible subtasks and evade per-agent inspection.
  • Promising because it tackles a failure mode unique to delegation-heavy systems rather than single-agent jailbreaks alone.
  • Why now: multi-agent orchestration is becoming common, and delegation chains create new ways for harmful intent to move without ever appearing explicitly at one node.
  • Skeptical about / limitation: the real test will be whether semantic labeling and blocking remain usable when agent roles, messages, and tool outputs are much noisier.

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

  • Adds a missing diagnostic layer between GUI grounding and full task success by asking whether models can reconstruct the causal transition produced by an action.
  • This is exactly the kind of measurement GUI agents need when stale observations and asynchronous rendering cause false beliefs about progress.
  • Why now: computer-use models are improving fast, but their most expensive failures still come from misreading whether a previous action actually changed the desktop state.
  • Skeptical about / limitation: offline step-level evaluation is useful, but still cannot reproduce every live timing and interaction artifact of real remote desktop use.

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

  • Shows that proprietary skill artifacts can be partially reconstructed from black-box execution trajectories, turning behavior logs into a side channel.
  • Important because it links agent UX, marketplace economics, and security into one problem.
  • Why now: hosted agents and reusable skill packages are becoming product surfaces, so trajectory leakage is no longer only an academic privacy curiosity.
  • Skeptical about / limitation: practical leakage severity will depend on what trajectories are exposed, how diverse tasks are, and what providers do to blur behavioral signatures.

5) Practical next steps

  • If you run tool-using agents, separate policy expression from policy enforcement; a handbook file is not a control plane.
  • Add at least one step-level benchmark to evaluation, especially for hierarchy conflict or GUI transition verification.
  • Treat skills, MCP tools, model packaging, and deployment topology as part of your security inventory, not peripheral implementation detail.
  • For hosted-agent products, assume execution traces may leak more than you intend and review telemetry retention through that lens.
  • If you evaluate patient, civic, or other public-facing agents, measure epistemic behavior under uncertainty, not just answer correctness.
  • For new post-training methods, prefer work that improves credit assignment or supervision density over papers that only report a better terminal score.

Generated from the selected-paper set plus candidate titles/abstracts only; no deepread pass yet.