The Decoder· Jonathan Kemper·· 2 天前AI 评分60
Google研究人员提出RRSI,防止自改进AI智能体记忆测试任务
Google researchers find a way to keep self-improving AI agents from memorizing their tests
AI 导读
Google研究人员提出RRSI,通过逐步收紧编辑预算和批评器审查,减少智能体在自优化过程中记忆测试任务的问题。基于冻结的Claude Opus 4.8,RRSI在8个基准上测试,训练任务最高提升14.1分,并在5个未见基准上最高提升4.7分。其运行时token用量比未正则化版本少约30 percent,代码已发布在GitHub。
来源:The Decoder · the-decoder.com