科学与数学
Agent Bluff · 多模型欺骗对赛
比较 Astra、Fable 5.1 和 Grok 4.6 的欺骗能力。源码公开。
科学与数学 / GPT-6 Astra

Epoch AI 评测中,预发布 GPT-6 Astra 自主写出 Lean 4 反例。仓库含 Challenge/Solution、Comparator 配置和 CI。
创作者尚未公开这个作品的提示词。你可以查看作者原帖或项目源码,了解更多创作细节。
Eval record: 26.1h working time, ~$432 at list rates
A formal-math repo, not a prompt toy. Human mathematicians have not audited the narrative.