英文版:/paper-news/2026-08-12/

运行统计

  • 候选论文: 310
  • 入选论文: 30
  • 已精读完成: 30
  • 时间窗口 (UTC): 2026-08-10T00:00:00Z → 2026-08-11T00:00:00Z (arxiv_announce, expanded=0)
展开查看用于总结的论文列表
arXiv ID标题 / 链接分类评分入选理由标签
2608.09867Stealing Reasoning Traces from Proprietary LLM APIs
PDF
cs.CR, cs.AI, cs.LG97High-impact LLM security flaw exposing hidden reasoning traces across users/models.llm-security, chain-of-thought, reasoning-traces, api-vulnerability, jailbreak
2608.09476ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
PDF
cs.CR, cs.AI95Strong agent safety benchmark for behavioral risks from trajectories, with self-evolving attacks.agent-safety, benchmark, behavioral-safety, red-teaming, tool-use
2608.09732ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners
PDF
cs.CR, cs.AI95Cross-skill collusion attack exposes a key blind spot in agent security scanners.agent-safety, security, jailbreaks, tool-use, evaluation
2608.09828Multi-Agent AI Safety as an Institutional Design Problem
PDF
cs.LG, cs.AI, cs.MA95Large pre-specified study on how institutional rules shape multi-agent AI safety.ai-safety, multi-agent, institutions, governance, evaluation
2608.09225Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference
PDF
cs.CR, cs.AI94Concrete defense for KV-cache timing leaks in multi-tenant LLM serving; practical security impact.llm-security, side-channel, inference, privacy, serving
2608.09025Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
PDF
cs.AI, cs.CR, stat.ML93Runtime governance for financial agents targets effect authorization, not just text.agent-safety, governance, runtime-monitoring, tool-use, finance
2608.09551Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models
PDF
cs.CL93Targets implicit-context jailbreak surface; strong safety framing beyond explicit prompt attacks.llm-safety, jailbreaks, prompting, pragmatics, robustness
2608.09624Measuring the Wrong Thing: Internal Harmfulness 评分s Anti-Rank Successful Jailbreaks
PDF
cs.CL, cs.AI, cs.CR92Challenges common jailbreak filtering assumptions; measures why internal harmfulness scores fail.jailbreaks, safety-evaluation, robustness, auditing, prompt-filtering
2608.09524STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework
PDF
cs.CR, cs.AI92Agentic incident-response framework with state tracking and staged planning for cyber defense.agents, cybersecurity, incident-response, planning, tool-use
2608.09001Mind the Hook: Source-Level Auditing of Privacy Defenses in Retrieval-Augmented Generation
PDF
cs.CR, cs.LG91Source-level audit of RAG privacy defenses reveals inactive hooks and leakage gaps.RAG, privacy, auditing, security, evaluation
2608.09836Mismatch Matters: On-Policy Distillation Beyond Token Agreement
PDF
cs.AI, cs.CL91Identifies OPD failure mode in LLM post-training and proposes mismatch-aware fix.llm, post-training, distillation, reliability, training
2608.09128Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
PDF
cs.CL, cs.AI, cs.MA91Objective multi-agent social benchmark with tournaments; useful for agent eval and training.agents, benchmark, multi-agent, evaluation, social-reasoning
2608.09885SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
PDF
cs.AI, cs.CV90Agent harness safety framework with explicit components and trajectory-driven evolution.agent-safety, harness, runtime-safety, tool-policy, memory
2608.09542Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
PDF
cs.LG, cs.AI, cs.CR89Adversarial safety alignment for reasoning models targeting mechanism-level jailbreak robustness.alignment, reasoning-models, adversarial-training, jailbreak-robustness, safety
2608.09158From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
PDF
cs.SD, cs.AI89Inaudible low-frequency red teaming for audio-language models with proposed defense.multimodal, audio, red-teaming, robustness, safety
2608.09629Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
PDF
cs.AI89Strong frontier-agent result: open-ended optimizer beats prescribed self-improvement pipelines.agents, optimization, frontier-llm, self-improvement, evaluation
2608.09164CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
PDF
cs.AI89Privacy preference alignment dataset with human annotations; concrete, reusable eval resource.privacy, alignment, dataset, evaluation, human-preferences
2608.09577ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
PDF
cs.AI88Supply-chain backdoor attack on agent skills is novel and highly relevant to agent deployment safety.agent-security, backdoor, supply-chain, skills, adversarial-attacks
2608.09072A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents
PDF
cs.SE, cs.AI88Repository-level benchmark decomposes coding-agent failures into requirements, planning, code.agents, coding, benchmark, evaluation, reasoning
2608.09445DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging
PDF
cs.CR, cs.CV88Practical defense against backdoor inheritance when merging diffusion checkpoints.security, backdoors, diffusion, model-merging, defense
2608.09819Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
PDF
cs.LG, cs.CL88Open continual-learning agent model with post-deployment self-improvement and modular LoRA specialization.agents, continual-learning, self-improvement, mixture-of-lora, open-models
2608.09928Multimodal Model Diffing for Feature Discovery and Control
PDF
cs.CV, cs.AI, cs.CL, cs.LG87Feature-level diffing/control for multimodal models aids interpretability and intervention.interpretability, multimodal, control, SAE, alignment
2608.09574The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games
PDF
cs.AI87Studies deception, cooperation, and governance failures in hierarchical LLM-agent games.ai-safety, agents, deception, multi-agent, governance
2608.09217Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training
PDF
cs.LG, cs.AI87Improves LLM RL post-training via task learnability, a potentially impactful training prior.llm-training, reinforcement-learning, post-training, reasoning, efficiency
2608.09119Motif 3: Technical Report
PDF
cs.AI86Large frontier MoE LLM with architectural novelty and scale; important despite limited safety focus.frontier-llm, moe, architecture, efficiency, pretraining
2608.09826Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
PDF
cs.LG, cs.AI86On-policy self-distillation transfers privileged skill signals into weights; strong relevance to LLM training.llm-training, self-distillation, reinforcement-learning, reasoning, post-training
2608.09069Telemetry and Concealment in Self-Adapting Generative AI: Logging Architecture, Adversarial Model Hiding, and the Limits of Detection
PDF
cs.CR, math.NA, q-fin.RM85Telemetry architecture for self-adapting generative AI tackles auditability and concealment.governance, monitoring, security, auditing, self-modifying-models
2608.09548ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
PDF
cs.CL, cs.AI, cs.CY85Integrated benchmark for education LLM capability, safety, trustworthiness, and pedagogy.benchmark, llm, safety, evaluation, education
2608.09123RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
PDF
cs.AI85Rubric-informed RL for open-ended alignment; targets missed criteria rather than scalar collapse.alignment, reinforcement-learning, rubrics, post-training, llms
2608.09153TRACE: TRajectory Attribution for Automated Context Engineering
PDF
cs.AI, cs.LG84Useful framework for automated context debugging in agents via trajectory attribution.agents, context-engineering, debugging, reliability, trajectory-analysis

AI 论文洞察简报

2026-08-12

0) 核心结论(请先阅读)

  • 今天最强的一条主线是:评估正在从仅看输出转向关注机制的审计。多篇论文表明,如果测量了错误的通道、错误的构念,或错误的分析单位,就会让防御措施看起来有效,但实际上并非如此。
  • 智能体安全正越来越成为一个基础设施与制度设计问题,而不只是模型对齐问题。运行时权限门控、感知溯源的执行机制、harness 演化,以及轨迹级证据,都会实质性改变结果。
  • 多篇论文揭示了智能体生态中的新型供应链与系统攻击面:被投毒的技能、跨技能串谋、KV-cache 时序泄漏,以及可重放的加密推理轨迹。
  • 在能力侧,进展更多来自更好的训练信号塑形,而不只是更多数据:面向准则的 RL、考虑可学习性的任务采样、以技能为锚的自蒸馏,以及感知失配的 on-policy 蒸馏,都提升了效率或鲁棒性。
  • 基准测试正变得更阶段感知与轨迹感知:编码智能体、协作智能体、社会推理智能体,以及面向教育的模型,如今都在中间步骤、角色和执行轨迹上被评估,而不再只看最终答案。
  • 对从业者而言,实际含义很明确:将权限绑定到状态,审计精确的干预点,记录轨迹,并在信任基准提升之前,用端到端金丝雀或可执行结果进行验证。

2) 关键主题(聚类)

主题:机制感知的安全与隐私审计

主题:面向智能体与制度的运行时治理

主题:智能体供应链与推理系统安全

主题:面向推理与开放式对齐的更优后训练信号

主题:基准测试正变得更具阶段感知、角色感知与轨迹感知

3) 技术综合

  • 安全论文中的一个共同模式是分析单位失配:单技能扫描器会漏掉多技能工作流,有害性探针会漏掉实际越狱成功,而黑盒隐私指标会漏掉防御是否真正作用于生成文本。
  • 多项工作正在汇聚到轨迹优先评估:TRACE、ActBench、SHE、STAIR 和 Social Gym 都将执行轨迹或多轮交互视为主要对象,而不只是最终响应。
  • 溯源绑定正成为核心设计原语:按主体划分的 KV salting、精确工件回执、不可变溯源守卫,以及上下文绑定的推理封装,都在将动作或缓存命中绑定到已认证状态。
  • 多篇论文区分了行为预防与机械性遏制。Constitutional prompts 可以压制不安全提议;可执行守卫可以允许提议但阻止执行;这两者是操作上不同的安全模式。
  • 在后训练中,共享的转向是从统一优化转向选择性优化:选择失败的 rubric 准则、高可学习性任务、零方差组,或失配严重的 token 位置。
  • 若干方法使用的是会随时间移除或门控的辅助信号,而不是永久混入主目标:RISE-RL 的引导调度、SKALD 的门控,以及 TRAJVAL 作为静态先验。
  • 轻量、可部署的防御明显增多:金丝雀验证、HMAC salting、音频 requery guard、“候选项+上下文”扫描,以及 harness 局部编辑。
  • 许多论文明确区分了状态表示与策略优化:事件响应中的 GAS、SAGE-Fin 中的带类型候选项、Macaron-V1 中的 HCP,以及 OEO 的优化契约,都在模型周围形式化环境。
  • 在各类基准中,中间监督正成为常态:需求澄清 GT、计划可复现性、个性化隐私历史,以及角色条件化的博弈结果,都提升了诊断能力。
  • 一个反复出现的限制是对 judge 和模拟器的依赖:即便是很强的机制性论文,也常依赖合成环境、人工编写目录或自动评审,因此独立重放和人工审计仍是高价值的下一步。

4) Top 5 论文(附“为什么是现在”)

Stealing Reasoning Traces from Proprietary LLM APIs

  • 表明加密推理封装可在会话之间和同系列模型之间移植,使较弱模型能够转录隐藏推理。
  • 展示了跨厂商影响和大规模真实泄漏:解码出 315,320 个公开推理块,其中包括恢复出的凭证和 PII。
  • 现在重要,是因为推理 token 产品和智能体轨迹共享的发展速度快于其安全模型。
  • 对 API/平台团队有用,因为缓解路径很具体:上下文绑定封装、服务端存储,以及跨模型隔离。
  • 质疑 / 局限:结果绑定于测试窗口内的特定 API 版本,据称提供方在披露后已进行了缓解。

Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference

  • 识别出多个 KV-cache 时序攻击背后的根因:缓存键未绑定到已认证主体。
  • 提出一个简单修复——按主体进行 HMAC salting——可将模拟 ASR 降至 0%,且每次请求仅增加约 1.6 µs 的中位开销。
  • 硬件 TTFT 测量证实该侧信道足够大,足以在生产中构成问题。
  • 现在有用,因为共享前缀缓存是多租户服务栈中的默认优化。
  • 质疑 / 局限:语义缓存不在讨论范围内,而边界 salting 的效率收益是外推而来,并未完整端到端测量。

Context Is Not Authority: Structured Runtime Governance for Financial Market Agents

  • 形式化了许多智能体团队已经感受到的一种失效模式:上下文正确并不等于在运行时拥有行动权限。
  • 提供了完整架构——带类型候选项、见证、覆盖债务、权限上限、精确工件回执和门控——并给出形式化健全性声明。
  • 为什么是现在:智能体部署正进入受监管、具状态性的工作流,“看起来对”已经不够。
  • 除金融外也有价值,可作为任何高风险智能体系统中 effect-boundary 治理的模板。
  • 质疑 / 局限:实证验证主要是人工编写的一致性测试加有限的定性部署证据,而非广泛的独立结果测量。

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

  • 提出了一个尖锐的方法论观点:一个检测器可以很好地区分有害提示词,但在预测成功越狱方面却比无用还糟。
  • 在包装过的有害提示词上,报告的结果 AUROC 为 0.220,这意味着成功攻击被打成比失败攻击更低的有害性分数。
  • 为什么是现在:许多团队正在固定误报预算下部署生成前过滤器和内部探针。
  • 之所以有用,是因为它将评估重心重新放在实际结果、校准和阈值迁移上,而不只是提示词标签上的 AUROC。
  • 质疑 / 局限:结果标签依赖 judge,且某些目标模型单元中的正样本数较少。

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

  • 提供了经过验证的需求澄清和规划中间 GT,而不只是补丁正确性。
  • 发现隐式需求恢复是主要瓶颈,占智能体运行的 24.5%–46.0%,而平均解决率仅为 31.5%。
  • 为什么是现在:编码智能体正在产品化,但大多数评估仍掩盖了失败源头。
  • 对研究优先级有用:提升需求理解能力,可能比再做一轮代码生成调优收益更大。
  • 质疑 / 局限:范围仅限于 163 个 Python/Java 任务,且部分诊断依赖 LLM judge。

5) 实际下一步

  • 在报告防御有效性之前,为每个 RAG/隐私基准增加hook 清单 + 指标到通道映射;并用实际输出通道上的金丝雀端到端验证泄漏。
  • 对智能体平台,实现状态绑定的执行门控:带类型工件、精确工件回执、溯源检查,以及按工具划分的权限上限,而不是只依赖提示词指令。
  • 审计你的服务栈中的共享状态侧信道:KV cache 命名空间隔离、语义缓存分区,以及时延差异测量,应成为多租户加固的一部分。
  • 将技能和已安装工具视为供应链工件:扫描候选技能时要结合已安装技能上下文,而不是孤立扫描,并为跨技能组合增加运行时溯源。
  • 针对实际攻击成功率重新评估安全过滤器,而不只是有害提示词分类;分别报告排序、校准和固定阈值行为。
  • 如果你使用 RLVR 或 OPD 训练,测试自己是否在零方差组或退化一致性上浪费信号;加入选择性辅助目标或失配感知修正。
  • 对编码和智能体基准,收集或合成中间参考(需求、计划、权限状态、轨迹谓词),以便尽早归因失败。
  • 现在就把轨迹记录与重放构建进生产智能体;今天若干最强方法——TRACE、SHE、STAIR、ActBench 风格审计——都依赖结构化轨迹来持续提升安全性。

基于逐篇论文分析生成;未进行外部浏览。