>>109536836
>benchmarks
>you already know
| Benchmark | DeepSeek-V4-Pro-0813 | Fable 5 | What it measures |
|---|---:|---:|---|
| **MMLU-ButWeSawTheAnswers** | 99.4% | 91.2% | knowledge, if the knowledge was in the eval set |
| **SWE-Bench Trust-Me-Bro** | 98.7% | 84.0% | resolves GitHub issues we wrote and also closed |
| **GPQA Diamond Hands** | 98.2% | 88.9% | PhD questions, leaked Q2 |
| **HumanEval-ButItsOurHumans** | 100.0% | 92.0% | passes tests it was allowed to read first |
| **AIME 2026 (We Had 2027's Too)** | 97.9% | 79.0% | competition math from a competition that hasn't happened |
| **Needle-In-A-Haystack@4097tok** | 12.0% | 99.1% | the one honest row. we forgot to remove it |
| **VibeBench-Reddit** | 98.0% | n/a | number of "insane bros" per thread |
| **RoPEmaxx Long-Context (claimed)** | 1,000,000 | 200,000 | tokens advertised |
| **RoPEmaxx Long-Context (actual)** | 4,096 | 200,000 | tokens survived |
| **Distill-Detect (lower is better)** | 3.4M | 0 | logged Claude convos in the training set |
| **"As an AI made by Anthropic" rate** | 4.1% | 0.0% | frequency of forgetting whose model it is |