✨ 精选阅读
最新 3 篇,直接读
5 Silent Failures in My AI Briefing System (and the 237 Tests That Came After)
I built an AI-powered daily briefing system. It 'worked perfectly' for weeks — until I discovered 5 classes of failures that pass every error check. HTTP 200, no exceptions, no warnings. But the output was wrong. Here's what happened and how I fixed each one.
Agent 超时假确认:别把半截 stdout 写成「已验证」
会话压缩后,被 kill 的命令(exit 143)的部分输出常被摘要成「已确认」。本文拆解假确认语义陷阱,给出观测/事实分层、强制复验与 harness 防护代码。
Least Autonomy 工程:Agent 什么时候该闭嘴不行动
2026 年 AgentAbstain / Agentic Abstention 表明:前沿 Agent 的 Paired Accuracy 仍卡在 60%。本文给出可落地的弃权(abstention)策略层——何时拒做、何时追问、何时早停,以及如何评测。
喜欢这些内容?
在 GitHub 上 Star 仓库,或订阅 RSS 获取最新更新。