THE BRIEF
DeepSeek's real announcement today was a cache reduction that happens to ship with a …
V4.1-Flash needs a quarter of the HBM and an eighth of the SSD storage for its KV cache compared with the previous generation — and because cache-hit charges are the line item …
Top Story
DeepSeek’s real announcement today was a cache reduction that happens to ship with a model. V4.1-Flash needs a quarter of the HBM and an eighth of the SSD storage for its KV cache compared with the previous generation — and because cache-hit charges are the line item that dominates long-running agent bills (DeepSeek says so explicitly in the thread), that is a change to what agents can afford to keep resident, not just what they pay per call. Plenty of labs shipped better models today.
The model: 552B-parameter mixture of experts, with a new causal encoder–decoder architecture activating 8B parameters on input and 16B on output. That asymmetry is the more interesting number than the 552B. Input context is where agent workloads actually spend tokens — retrieved documents, tool results, prior turns — so an architecture that activates 8B for input is a bet that the scaling pressure sits in the context window rather than in reasoning depth. DeepSeek calls V4.1-Flash the smallest model in a new architecture family, which is an announcement of larger siblings without saying so.
It is live on the DeepSeek API as deepseek-flash with native multimodal support; V4-Flash and V4-Flash-Vision-Exp are retired, with the old model IDs routing to V4.1-Flash for compatibility. No end date is given for those redirects, which matters if you are pinned to the old IDs in production — “temporarily route” is doing unlabelled work in that sentence.
On price: DeepSeek says lower API prices, keeps the peak/off-peak split, and off-peak sits at 50% of peak. No per-million-token figures appear in the thread as published. The direction is clear and the magnitude is unverified until someone diffs the pricing page against last generation’s — which is the first thing worth doing, because “we’re passing the savings on” is unfalsifiable as written.
The deployment offer is the part that doesn’t fit the usual open-weights script. DeepSeek is publishing weights to Hugging Face and working with the community on inference support — and in the same breath soliciting “a large-scale deployment with 2,000 GPUs + a storage cluster.” That is a vendor offering to help you build the alternative to its own API. It’s a coherent position only if you read the cache compression as the real product: a cache that is an eighth the size makes self-hosting plausible at exactly the scale where the API becomes a recurring line item.
Set against the rest of the day, the shape is clear. OpenAI’s Agents API and Cognition’s SWE-2 are platform plays — orchestration, tooling, evaluation harnesses. DeepSeek is commoditizing the layer beneath all of them. Every agent platform pricing page written today is a markup on tokens whose cost just moved, and the platforms that look healthiest on unit economics are the ones with the least lock-in to a single inference supplier right now.
What I don’t know, and would want before drawing a firm conclusion: there are no independent evals in hand. The thread’s benchmark post is truncated at “ahead of” — ahead of what, and on which suite, is not stated in the material available. The 196B Engram parameter count is not a third-party rumor: it is in DeepSeek’s own technical report for V4.1-Flash, which describes a 552B backbone plus 196B Engram conditional-memory parameters. And the cache claims are ratios against V4-Flash, not absolute numbers, so a team already on a different serving stack should not assume a direct 4× transfer.
The concrete thing to watch: whether the off-peak discount survives contact with demand. DeepSeek is using price to flatten load, which works when idle capacity exists and stops working the moment V4.1-Flash becomes the default for batch agent work. If off-peak quietly narrows within a quarter, the headline savings were a launch promotion. If it holds, the cache math is real and the rest of the field has a serving-economics problem it will have to answer with weights, not with API features. DeepSeek v4.1 Flash
Also Today
Anthropic Says It Blocked Possible Efforts to Build Biological Weapons · Source Anthropic’s threat-intelligence report says it disrupted several cases this year of scientists using Claude for biological work that could aid bioweapons development, and that it could not tell legitimate inquiry from misuse, so it acted anyway. The same report adds a genuinely new category: three cases in China, two in Russia and one in Yemen (contextually the Houthis) of using Claude to develop software for conventional weapons — firearms, missiles, armed drones, bombs — plus Russian state media generating fabricated Moldova election coverage. The bioweapons framing got the headline, but the dual-use ambiguity it admits is the finding: no classifier can separate vaccine research from pathogen engineering, so Anthropic is publishing its own judgment calls as the safeguard. This is a vendor grading its own homework, and the counts deserve that skepticism.
Shopify moves back to Native from React Native · Source Shopify is moving its mobile apps off React Native and back to Swift and Kotlin, the reversal of its 2020 all-in bet, because coding agents changed the cost of the thing React Native existed to avoid: writing the same feature twice. The Shop app went from proof of concept to a shipped native rebuild in 12 weeks; the 300-plus-screen Shopify app follows this year. The transferable machinery is Helix, which slices a migration into checkpoints that each need passing tests, a visual match, two adversarial reviewers and a human sign-off. The real prerequisite wasn’t model quality but architecture: business logic decoupled from UI so agents can exercise it through a CLI in milliseconds instead of driving a simulator. Shopify rebuilt its apps for agents first, then let agents rebuild them.
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra · Source Cognition’s SWE-2 scores 50.0% on FrontierCode 1.1 Main, one point behind Fable 5.1 at 64% less cost, post-trained from Kimi K3 (2.8T parameters) with RL scaled to the multi-trillion-parameter regime and all effort levels trained in a single run via a cost penalty λ tuned to the local slope of the base model’s Pareto frontier. The efficiency numbers are the interesting ones: SWE-2 medium beats SWE-1.7 while taking 58% fewer turns and making its first real edit at a median 18 steps versus 48. But Terminal-Bench 4 tells the other half — 27.3% against Fable 5.1’s 55.8% and Astra’s 57.9% — so this is a model tuned hard for repository work, not long-horizon terminal autonomy. The Pareto-slope penalty is the piece other labs will copy.
OpenAI Agents API · Source OpenAI’s Agents API exposes the Codex harness as a managed service: it runs sessions, orchestration, context compaction and recovery, while you supply tools and pick an environment, with sandboxed code execution, file edits, MCP connections and subagents. Billing is the model rate plus tool and container rates. The load-bearing line is in the data controls: sessions are US-residency only, Zero Data Retention is not supported, and choosing a self-hosted sandbox explicitly does not make the endpoint ZDR-eligible. Managed agent runtimes are now table stakes; what differentiates them is whether your prompts and code get retained, and OpenAI is telling enterprise buyers the answer is yes. For regulated workloads, that exclusion matters more than the feature list.
TSMC Reported to Speed Up 1.4nm Mass Production, Trial Production Starting as Early as April 2027 · Source Reports say TSMC’s Taichung A14 (1.4nm) site is running ahead of plan: the first of four fabs has finished steelwork, with its first manufacturing phase completed around April 2027, trial production possible as early as Q3 2027, and volume production around mid-2028 — against the 2028 date TSMC reaffirmed at its Q1 earnings call. Planned spend is roughly $49B. A14 promises up to 15% more performance at equal power, or 30% less power at equal performance, with over 20% more logic density, on second-generation nanosheet transistors with NanoFlex Pro. Intel targets 14A risk production in H2 2027 and Samsung’s SF1.4 around 2029. The schedule is press-reported, not a TSMC commitment, and the binding constraint from here is yield, not construction — but with A14 volume production reported around mid-2028, TSMC’s date is roughly level with Intel’s 2028 14A high-volume ramp rather than a year ahead of it.
In Brief
- Apple took the wraps off the iPhone 18 Pro and Pro Max, shown in a burgundy finish, with the usual newsroom rollout of specs and imagery. (Source)
- A Hacker News user reports that OpenAI has repeatedly re-enabled the ‘allow training’ setting after it was switched off, an unverified single-account claim that lands the same day OpenAI confirmed its new Agents API endpoint is not Zero Data Retention eligible. (Source)
- OpenAI launched ChatGPT for Financial Services, pairing built-in financial data with GPT-6 Astra for research, modeling and client-ready materials. (Source)
- Consumer-rights researchers have collected Sony’s own website language about players ‘owning’ digital games, as four California PlayStation buyers who spent hundreds of dollars on digital goods say they were left with only a license. (Source)
- Google shipped the Gemini app for Windows, its first dedicated desktop client for the assistant. (Source)
- A piece argues the safety case for autonomous vehicles is now overwhelming, citing an estimate that self-driving tech could prevent 580,000 deaths a year. (Source)
- A researcher claims on X that OpenAI may have taken another major mathematical proof without credit — an unconfirmed accusation, and one of several circulating against the same lab this week. (Source)
- PlanetScale opened Neki in platform preview, a new product built from eight years of running large-scale database infrastructure. (Source)
- Researchers released a demo of WeWorm, described as the first zero-click worm spreading through WeChat calls on both iOS and Android, in which the victim need not answer the call or interact with the device. (Source)
- The DOJ won a temporary pause on an order requiring documents from 14 agencies in its antitrust case against Apple. (Source)
- Norwegian researchers spoofed a helicopter’s GPS in 40 minutes during the 2024 Jammertest campaign at Andøya, demonstrating that satellite navigation can be quietly steered to a false position. (Source)
- A community mod got native DLSS frame generation running on Nvidia’s RTX 20-series Turing cards, undercutting Nvidia’s position that the feature requires the Ada Lovelace optical flow accelerator. (Source)
One Line
We don’t hold on to a decision just because it was successful at the time. When a core assumption changes, we’re willing to go back and ask whether it’s still the right call.
— Shopify Engineering, ‘Back to Native’ — the post announcing the move from React Native back to Swift and Kotlin