Science & math
Agent Bluff · deception tournament
Compares whether Astra, Fable 5.1, and Grok 4.6 lie well. Source is public.
Science & math / GPT-6 Astra

In an Epoch AI eval, a pre-release GPT-6 Astra autonomously wrote a Lean 4 counterexample. Challenge/Solution, Comparator config, and CI are in the repo.
The creator has not published a prompt for this project. Explore the original post or source code for more context.
Eval record: 26.1h working time, ~$432 at list rates
A formal-math repo, not a prompt toy. Human mathematicians have not audited the narrative.