AI
AI Twitter Daily Digest
Sources Aaron Levie / @levie Andrej Karpathy / @karpathy Andrew Ng / @AndrewYNg Aravind Srinivas / @AravSrinivas Awni Hannun / @awnihannun Fei-Fei Li / @drfeifei Gary Marcus …
Sources
Highlights
Today’s discourse is sharply divided between breakthrough capabilities in formal verification and glaring limitations in open-ended reasoning. OpenAI’s Astra model showcased significant leaps in solving major theoretical math problems, igniting heated debates over whether this signals impending AGI or simply massive compute applied to verifiable search spaces. Simultaneously, harsh economic realities are surfacing, from macro warnings about the tech market’s over-reliance on OpenAI to rapid price-to-performance disruption by DeepSeek’s new V4-Flash model.
Top Stories
- OpenAI’s Astra Solves 10 Major Math Problems: An internal version of OpenAI’s upcoming Astra model (referred to by some as GPT-5.6) has reportedly solved 10 open problems in mathematics, theoretical computer science, and quantum complexity. The newly discovered proofs, which come with Lean certificates and chain-of-thought walkthroughs, tackle complex areas like the existence of nonsofic groups and Connes’ Rigidity Conjecture.
- DeepSeek V4-Flash Delivers a “2.0 Moment”: The newly released DeepSeek V4-Flash is causing a massive stir by successfully completing Fable 5 benchmark tasks at a reported 105x lower total cost. Industry observers note that two orders of magnitude improvements in price-performance are exceedingly rare and highly disruptive to the current AI stack.
- OpenAI Termed a “Load-Bearing” Market Risk: A stark macroeconomic warning circulated characterizing OpenAI as the single massive failure point for the current hyperscaler market structure. The analysis argues that Microsoft Azure’s growth and hundreds of billions in backlogs at both Microsoft and Oracle are precariously tied to OpenAI’s compute expenditures, warning of cascading market effects if OpenAI fails to go public within the next 8 months.
- The AI Harness Becomes the Ultimate Differentiator: Box CEO Aaron Levie argues that as AI tasks scale to tens or hundreds of millions of tokens, the “harness”—the system routing sub-tasks to the right models at the right time—will match raw model capability in importance for reducing cost and maximizing accuracy. He predicts a widening divergence where deep domain AI (math, science, coding) goes vertical while standard consumer productivity levels off.
- Claude Code Proves 3.7x Pricier Than Alternatives: A cost analysis of AI coding agents by Composio found Hermes and Pi Agent leading the pack on average task cost ($0.39 and $0.40 respectively), while Claude Code was the most expensive at an average of $1.47 per task.
Articles Worth Reading
The “All Cognition is Alike” Fallacy Gary Marcus penned a sharp critique of the “AGI-is-near” community’s jubilant reaction to the new Astra math proofs. He argues that success in formal, verifiable domains like discrete math does not automatically translate to domains requiring open-world reasoning, pointing out that human expertise in one cognitive domain never guarantees universal competence. The debate escalated significantly, with tech commentators like Matt Shumer directly accusing Marcus of constantly moving the goalposts over the last decade.
The Solemn Gravity of Client Capital Following the meltdown of the SALP fund, Jeff Park dissects an interview with Leopold Aschenbrenner, author of “Situational Awareness,” pointing out severe red flags in his fundamental approach to institutional risk. Park highlights that for fiduciaries and stewards of client capital, “not blowing up” isn’t a secondary objective, but the only task, issuing a stark warning against a tech subculture that increasingly celebrates and justifies reckless financial misadventures.
LLMs Still Can’t Write a YouTube Script Physicist and creator Sabine Hossenfelder offers a grounding reality check on current model capabilities in open-ended creative workflows. Despite attempting to use ChatGPT, Claude, Grok, and Gemini, she found them fundamentally incapable of generating interesting, coherent video scripts or reliably fact-checking her work. She notes that over the past two years, their utility for these specific tasks has arguably regressed, leaving them useful only for basic grammar fixes and gathering references.