NEWS
The Brief
Gemini 3.7 Flash lands at $0.75 per million input tokens and $3.75 per million output, half of what 3.6 Flash cost at launch, while beating it across the coding, document, and …
Top Story
Google has cut the effective price of its agent workhorse model in half three weeks after shipping the previous one, and the cadence — not the benchmark deltas — is the news. Gemini 3.7 Flash lands at $0.75 per million input tokens and $3.75 per million output, half of what 3.6 Flash cost at launch, while beating it across the coding, document, and workflow evals that Google chose to publish. The model itself is a modest step; the pricing and the shipping rhythm are the move.
The numbers are real but unsurprising for a point-release refresh. On FrontierCode 1.1 Main, 3.7 Flash hits 43.6% against 3.6 Flash’s 34.4%; on the long-horizon software-engineering eval DeepSWE v1.1 it scores 65.3% versus 49.0%. WebDev Arena Elo moves from 1538 to 1588. The bigger gaps are in knowledge work: GDP.pdf, a document-comprehension benchmark, jumps from 22.0% to 34.0%, and AutomationBench — completing real business workflows — from 17.0% to 30.4%. Those last two are where the product intent is clearest: this is a model aimed at agents that read contracts, reconcile spreadsheets, and drive multi-step tool calls, not at chat.
What matters is what Google chose not to do. It did not hold 3.7 Flash for a next “major” release. It shipped 3.5 Flash Cyber, 3.6 Flash, and now 3.7 Flash in roughly a quarter — a release cadence that reads like a continuous-deployment pipeline for the model layer. And it cut the price at the moment of release, through the end of the year, rather than after. That combination is the economic signal: Google is pricing agents as a volume business and using speed to deny competitors a settled price-performance reference point. Every autonomy product that prices off a per-token budget — retry loops, sub-agent orchestration, 24/7 “personal agents” like Gemini Spark, which picks up 3.7 Flash today — just got its arithmetic rewritten. Half the token cost is, in effect, double the retry budget at the same spend, and agents burn retries faster than any chat workload.
The halved price also does what halved prices usually do: it moves the cost question off the table and onto reliability. At $3.75 per million output tokens, the binding constraint on an autonomous workflow stops being unit economics and becomes first-pass success rate — every failed tool call and every hallucinated step still costs wall-clock time and supervision, and those don’t get 50% cheaper because the tokens do. The evals Google highlighted are precisely the ones that reward getting it right once: higher first-pass code accuracy, fewer retries, disciplined execution. That is the honest read of why those particular benchmarks lead the announcement.
There are two things worth flagging as unconfirmed. Google’s own post ties 3.7 Flash to its Frontier Safety safeguards and bioresilience posture, a framing worth watching given the CBRN and cyber-offense language in the release. And the timing is pointed: DeepSeek shipped its first agent product the same day, so 3.7 Flash is arriving with a competitor’s agent debut already priced into the room.
What this changes is specific and measurable, not abstract. Every developer who priced an agent on 3.6 Flash tokens should re-run that math today, because the input cost just halved with a model that clears the old one on the evals that count. The model layer is now moving on a three-week cycle at collapsing prices, and the teams that win the next six months are the ones that stop treating model choice as a quarterly decision and start treating it as a rolling one. The party whose next move actually decides the shape of this — whether the three-week cadence is a sprint or the new normal — is DeepSeek, whose Harness preview is the first real counterweight to the Flash pipeline, and whose pricing answer will tell us whether Google just reset the floor or merely joined a race to it. Gemini 3.7 Flash
Also Today
DeepSeek Harness developer preview · Source DeepSeek’s first agent product, Harness, is out in developer preview, and its most consequential decision is architectural: every capability — models, tools, sessions, sandboxes, even the UI — is a plugin on top of Cordis, so teams swap or extend components in configuration without forking DeepSeek’s source. Every run records an append-only session log of system prompts, reasoning, tool calls and context injections, viewable in a Trajectory panel that resume, fork and replay all draw from. Four runtime modes (standard, code, minimal, creator) frame Harness as an evaluator’s bench as much as an agent shell. DeepSeek is betting the winning position is owning the harness, not the model.
Spaghettifying DRAM · Source Chris Domas’s latest is a one-line exploit with outsized consequences: a single bit-flip in the DRAM controller’s bank-swizzle register rewires the lowest stage of the address translation pipeline, and every security primitive above it — PSP, microcode, SMM carveouts, the ones invisible even to the kernel — is built on physical addresses the controller can now silently rearrange. Because the controller’s transform is a GF(2) linear map, the scrambled memory can be reconstructed with linear algebra. It targets AMD Family 16h, the last generation whose datasheets document the translation registers as unlockable; newer parts just leave that out. This is the clearest argument yet that the fence at the memory controller, not the TLB, is the real last line.
Accelerating GPT-5.6 Sol Ultrafast · Source Cerebras and OpenAI are previewing Ultrafast, a tier that runs GPT-5.6 Sol at up to 750 output tokens per second with, they claim, no quality compromise — 11x faster than Fable 5 and 5x faster than Opus 4.8 on fast mode, and on Humanity’s Last Exam the entire 2,500-question set finished in 11 hours 11 minutes versus Fable’s 78 hours 27 minutes. The mechanism is the familiar Cerebras thesis: 44 GB of on-chip SRAM keeps weights resident, so inference becomes a data-movement problem the wafer-scale architecture doesn’t have. As agents move onto critical paths, the argument that latency is a feature is getting harder to argue with, but the speed-vs-intelligence tradeoff is only postponed, not abolished.
Claude users are mad that Anthropic’s new watermarks will catch them using it · Source Anthropic has started watermarking Claude’s editorial output, inserting invisible markers to satisfy the EU AI Act’s Transparency Code, and the predictable Reddit blowback has arrived: students caught reusing AI-organized paragraphs, writers objecting to the ‘digital tattoo,’ and one thread pointing out the irony of watermarking work built on scraped training data. The counterargument — that the watermark only catches verbatim copy-paste, which is the one case where detection is clearly legitimate — is sound, and so is the observation that watermarking deters the careless while the careful paraphrase or resynthesize around it. It is an honesty tax that falls almost entirely on people who were never going to get away with it anyway.
Anthropic: Introducing The Conceptual Reasoning Index · Source Anthropic and Redwood Research have published the Conceptual Reasoning Index, three benchmarks — LMCA for argumentation judgment, ACCoRD for logical consistency, DTBench for decision-theoretic reasoning — that measure models on the unverifiable, no-feedback-loop questions where AI risk work actually happens. The headline number: Opus 5 sits at 73.6 against an estimated ceiling around 91, and scores have climbed roughly linearly since late 2024 with no flattening. DTBench is already near saturation, LMCA is a year out from it. Reading the trend lines is reading a claim about how long the reasoning gap lasts, and Anthropic is making its bet explicit rather than letting it stay implicit in a benchmark leaderboard.
In Brief
- Codex in the ChatGPT desktop app for Linux has moved to preview, available as an official x64 RPM. (Source)
- Apple is in talks to pay publishers for news content to give Siri AI access to current events, per The Wall Street Journal. (Source)
- Someone is running mass vulnerability scans that spoof known AI crawler user-agents like ClaudeBot to hide in analytics, according to KnownAgents. (Source)
- Microsoft is merging its consumer Copilot and Microsoft 365 Copilot apps into one unified client ahead of a broader ‘super app’ later this year. (Source)
- The tiny-JPEGs-look-different-in-Chrome mystery turns out to be a deliberate JPEG decoding optimization, not a rendering bug. (Source)
- A single multi-line log entry can translate into 49KB+ of disk writes on ext4 and 110KB+ on btrfs through systemd-journald, per a Debian bug report. (Source)
- Bloomberg reports Anthropic is in talks for a roughly $6 billion deal to acquire AI-infrastructure startup Decart, its biggest acquisition yet. (Source)
- A blog post written days after OpenAI claimed to have solved ten major math problems asks what sort of maths LLMs are actually good at. (Source)
- A write-up argues HTML over WebSockets lets you build real-time SPAs with barely any JavaScript. (Source)
- Wednesday’s total solar eclipse produced measurable internet traffic dips across Iceland, Spain and Portugal as people looked up. (Source)
- A review of 50 open source projects finds AI is accelerating development while adding new attack surfaces from unfamiliar contributions. (Source)
- Apple has submitted a court-ordered proposal to charge developers up to 15 percent for linking to purchase options outside the App Store in the US. (Source)
One Line
Physical addresses are really more of a suggestion.
— xoreaxeaxeax, in the skitter-creek-bath-salts README