🔬
Research Digest
6 articles
Claude Formalized Fermat’s Last Theorem. It Did Not Discover a New Proof.
Claude generated a 13-million-line Lean formalization of Fermat’s Last Theorem in 11 days. The real advance is not a new proof—it is verification throughput, agent scaffolding, and a public artifact that exposes exactly what the kernel checked.
Ibrahim Denis Fofanah·Sep 8, 2026·9 min
Same Model, Same Benchmark: 54.8% or 99.9% Depending on the Harness
GPT-6 Astra scored 54.82% or 99.95% on ARC-AGI-3 at the same reasoning level. The model did not change; the evaluation harness did. That gap is a lesson for anyone measuring agents.
Ibrahim Denis Fofanah·Sep 6, 2026·9 min
Brain Waves to Words: What Brain2Qwerty Actually Does, and What It Doesn't
Meta's Brain2Qwerty decodes typed sentences from brain activity with no surgery, at 61% word accuracy. The catch: the scanner is a room, and the participants could type.
Ibrahim Denis Fofanah·Jul 13, 2026·6 min
RAG vs. Fine-Tuning: A 2026 Decision Framework for Practitioners
Stop arguing. Here's a decision tree grounded in cost, latency, and drift.
Ibrahim Denis Fofanah·Feb 23, 2026·8 min
5 Papers That Explain How LLM Alignment Actually Works
RLHF, Constitutional AI, DPO, and Anthropic's Sleeper Agents result showing safety training can teach a model to hide rather than behave.
Ibrahim Denis Fofanah·Feb 16, 2026·6 min
State of Coding Agents: Who Actually Wins on Real-World Tasks?
Agents score 90%+ on SWE-bench. A controlled trial found developers were 19% slower with AI, and thought they were 20% faster. Why both are true.
Guest Contributor·Feb 14, 2026·7 min