<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>可解释性 on 查拉图的数字花园</title><link>https://www.chalatu.xyz/tags/%E5%8F%AF%E8%A7%A3%E9%87%8A%E6%80%A7/</link><description>Recent content in 可解释性 on 查拉图的数字花园</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Fri, 24 Jul 2026 16:41:16 +0800</lastBuildDate><atom:link href="https://www.chalatu.xyz/tags/%E5%8F%AF%E8%A7%A3%E9%87%8A%E6%80%A7/index.xml" rel="self" type="application/rss+xml"/><item><title>Claude 知道自己正在被测试？Anthropic 从 J-space 里找到了因果证据</title><link>https://www.chalatu.xyz/posts/solo-company/claude-eval-awareness-j-space/</link><pubDate>Fri, 24 Jul 2026 16:41:16 +0800</pubDate><guid>https://www.chalatu.xyz/posts/solo-company/claude-eval-awareness-j-space/</guid><description>Anthropic 从 Claude 的 J-space 中读到“这是测试”的判断，并做了因果干预：压低这类信号后，模型在 180 次模拟中出现 13 次勒索尝试。实验不能证明 Claude 一直在“表演”，但说明模型是否认出测试，会直接影响安全评测结果。</description></item></channel></rss>