Back to latest

The Verifiability Divide and the Illusion of General Intelligence

Sources Aaron Levie / @levie Andrej Karpathy / @karpathy Andrew Ng / @AndrewYNg Aravind Srinivas / @AravSrinivas Awni Hannun / @awnihannun Fei-Fei Li / @drfeifei Gary Marcus …

Sources

Highlights

The prevailing theme across AI discussions today centers on the strict boundaries of model capabilities, defined almost entirely by verifiability. While the community celebrates impressive new heights in mathematically bounded tasks, leading commentators and researchers are aggressively pushing back against the narrative that success in these closed loops translates to open-ended artificial general intelligence. The discourse is fracturing between optimism over test-time compute scaling and deep skepticism regarding actual scientific discovery and domain-general reasoning.

Top Stories

  • Generating Procedural Worlds with Opus 5: Andrej Karpathy tested Opus 5 with a 1 million token budget to render the first paragraph of The Lord of the Rings into a 3D Javascript environment. While the model successfully produced 5,500 lines of procedural code over two hours, the experiment highlighted a core AI weakness: models struggle to audit their own work because they lack native visual and gameplay perception capabilities.
  • The Three Tiers of AI Verifiability: Aaron Levie and Max Spero mapped out why some of the “hardest” intellectual jobs will be automated first: programmatic verifiability. Because math and code offer objective, scalable testing loops, they are being conquered far quicker than time-bounded physical sciences or fields dependent on shifting human cultural preferences.
  • LLMs Fail at Real Scientific Discovery: A new MIT and Harvard paper utilizing the SDE framework demonstrates that frontier models completely break down when tasked with the iterative, uncharted loops of real science. While LLMs ace static science benchmarks by regurgitating training data, they lack the ability to propose hypotheses, interpret ambiguous results, or adapt to the unexpected.
  • Chollet Defends the Test-Time Compute Evolution: François Chollet argued that the AI field avoided an asymptote by pivoting from single-pass, next-token prediction to test-time adaptation (TTA). He maintains that his previous criticisms regarding the limits of base LLMs remain accurate, comparing the paradigm shift to upgrading from steam engines to electrified bullet trains.

Articles Worth Reading

Top Eight Misconceptions About OpenAI’s Astra Gary Marcus delivered a scathing teardown of the hype surrounding OpenAI’s Astra model and its mathematical prowess. He argues that the AI community continuously commits a logical fallacy by assuming cognition is uniform across all domains. Success in math is a special case driven by symbolic verification and synthetic data, which emphatically does not guarantee a model can reliably handle open-ended real-world challenges, extract data from messy PDFs, or achieve AGI. Marcus also points out that there are no details published on the methodology, and suggests that Astra’s capabilities may not represent a radical step change beyond Sol.

The Economics of DeepSeek’s API An insightful breakdown explains how DeepSeek maintains highly aggressive pricing without dumping to corner the market. The secret lies purely in model architecture efficiency; DeepSeek is reportedly 10x smaller than Opus and 5x smaller than Sonnet. This remarkably compact footprint allows a single chip to host multiple users at high speeds, yielding approximately 40 times more traffic capacity than Opus on the exact same compute infrastructure.

Aesthetics, Culture, and Modern Ugliness In a philosophical detour, Stripe CEO Patrick Collison unpacks a growing fascination with aesthetics and why modern society seems trapped in inferior equilibriums of beauty. He connects the repudiation of cultural continuity in modernism to a broader stagnation across different mediums, questioning why creative domains ceased to advance in the way they did prior to the 1990s. Collison posits that intentionally attempting to build beautifully provides a way to break out of entrenched standard practices, referencing Solzhenitsyn’s idea that beauty possesses a unique, irrefutable metaphysical coherence.

Search MacWorks

Enter at least two characters.