Back to projects

Science & math / GPT-6 Astra

Agent Bluff · deception tournament

Agent Bluff · deception tournament

About this project

Compares whether Astra, Fable 5.1, and Grok 4.6 lie well. Source is public.

Eval harness

Published prompt

The creator has not published a prompt for this project. Explore the original post or source code for more context.

Creator & source notes

Scores belong in the repo; this catalog does not repeat unaudited numbers.

Science & math

Lean 4 disproof of the Köthe conjecture

In an Epoch AI eval, a pre-release GPT-6 Astra autonomously wrote a Lean 4 counterexample. Challenge/Solution, Comparator config, and CI are in the repo.