Back to latest

The Brief

Qwen 3.8 27B, released on Hugging Face today in FP8, scores 61.7 on SWE-bench Pro and 42.2 on DeepSWE 1.1 — against Claude Opus 4.6 Max's 53.4 on the former and no posted …

Top Story

An open-weights model just beat the closed frontier on the agentic coding benchmarks the field actually argues about. Qwen 3.8 27B, released on Hugging Face today in FP8, scores 61.7 on SWE-bench Pro and 42.2 on DeepSWE 1.1 — against Claude Opus 4.6 Max’s 53.4 on the former and no posted number at all on the latter. It holds 90.3 on LiveCodeBench v6 to Opus’s 88.8, 79.0 on Qwen’s own software-engineering benchmark to Opus’s 63.8, and 84.3 on OSWorld-Verified computer use to Opus’s 72.7. The number that should stop you is the one at the top of the card: 27B parameters, dense, a native vision-language model with 262K context extensible to a million tokens. This is no longer open weights trailing closed models by a generation. At the top of the agentic stack the gap has closed, and the model that closed it runs on hardware a company already owns.

Read the card before you trust the table, because Qwen’s own reporting demands it. SWE-bench Pro and DeepSWE 1.1 were evaluated with the Claude Code harness at temp 1.0 and a 256K window — real conditions, but conditions Qwen chose, not ones you will necessarily deploy under. The multimodal table corrects “a small number of incorrect ground-truth annotations” in MathVision and CharXiv and recomputes every affected score from the fixed versions, which flatters the Vision numbers in ways that deserve a skeptical eye. And Qwen still loses the raw terminal race (73.0 to Opus’s 78.2) and trails on HLE (30.8 to 40.0) and GPQA (89.2 to 91.3). So this is not “open beats closed everywhere.” It’s that the skills that carry real work — software engineering, computer use, long-horizon office tasks, where Qwen hits 70.7 on CoWorkBench to Opus’s 68.2 — are where the open model now leads.

The engineering is what makes the claim stick. The weights ship FP8, fine-grained quantization at block size 128, with performance Qwen calls nearly identical to the unquantized model — so the headline numbers are roughly what you get when you download, not what a frontier lab reports from an unreachable cluster. The architecture is a hybrid: sixteen gated DeltaNet blocks of sparse linear attention sitting under gated attention layers, which is how a 27B dense model sustains a 262K native context at all. It serves on vLLM, SGLang, and TokenSpeed out of the box, thinking mode on by default with per-request disable and a tunable reasoning_effort. Nothing here needs an API key, a data-use agreement, or a trust decision.

That last point is what ties today’s other news into one story. The same edition has Google making homomorphic encryption practical, Anthropic explaining its text watermarking, and Google letting users strip visible watermarks from AI output — every one of them straining at the same question: how much of what the model does can you verify? Open weights is the bluntest answer available, because verification stops being a provenance problem the moment you can run the thing and inspect it yourself. The price war now forcing OpenAI and Anthropic to cut rates as Chinese rivals undercut them makes that answer cheaper by the week, and when the open model beats the closed one on the benchmarks your engineers will copy into a sprint, the premium for keeping your code out of someone else’s inference rack starts to look like a tax on nothing.

The move that decides where this goes is Qwen’s own: the hosted 3.8-27B API on Qwen Cloud, coming soon, with built-in tools and a 1M default context. It’s the first clean shot at how Alibaba prices a model that beats the US closed labs on their own turf. If it launches near the current price war’s numbers, the closed labs’ agentic margins aren’t just under pressure — they’ve got a reference price set by the model that’s beating them. Qwen 3.8 27B is out: open weights, best local dense model yet

Also Today

Google Is Making Private AI Practical with Homomorphic Encryption · Source Google’s HEIR compiler is out, and it moves homomorphic encryption from a research niche to a tool a non-expert could plausibly ship. HEIR takes a pre-trained model that runs on plaintext and recompiles it to operate on ciphertext, so a server can run recommendation or fraud detection without ever seeing the input. The demos — a recommendation model with Belfort Labs and LG, a Kitsune intrusion detector with Niobium — run single-threaded on CPU, which tells you the latency story honestly. The catch is the one HEIR can’t compile away: a cryptographer still built these. Making the compiler approachable compresses the cost of private inference; it does not yet erase the expertise needed to reach the first usable result.

How Claude’s text watermarking works · Source Anthropic’s watermarking announcement reads as an unusually candid accounting of what provenance tools can and cannot do. Claude’s output will carry a SynthID-Text pattern only in the low-stakes word choices where either option is fine, which means exact text like code and math gets almost no watermark, and lightly-edited human writing gets little to none. Detection stays probabilistic, needs a key, and can’t say who prompted the model. The honest limits matter because the EU AI Act pushed this out globally before a regional switch existed. What Anthropic is describing, in effect, is a marking system whose real signal is how much of the text Claude actually chose.

Firefox is now the last major browser that still supports uBlock Origin · Source With Microsoft Edge following Chrome to Manifest V3, Firefox is now the only major browser where uBlock Origin still runs at full strength, and Mozilla’s response was a one-line promise that support “isn’t going anywhere.” The significance is mostly structural: uBlock Origin’s MV2-era API access is exactly what lets it block requests that Chromium’s replacement extension system can’t intercept as precisely. For the ad-blocking faithful that makes Firefox the last uncompromised option, and Mozilla has decided that’s a feature to lean into rather than a legacy to phase out. It is, in a genuinely strange way, the clearest retention strategy a small browser has had in years — a reason to switch that no rival can copy.

OpenAI and Anthropic in price war as Chinese AI rivals gain ground · Source OpenAI cutting GPT-5.6 Luna’s price 80 percent and Anthropic launching Opus 5 at “half the price” of Fable 5 is the first time the two labs have publicly competed on cost rather than benchmark, and Silicon Data’s index says token prices from leading US labs have fallen almost a quarter since mid-July. The pressure is upstream: capable open-weights Chinese models let DoorDash and Airbnb shop around, and both labs need to show revenue growth ahead of trillion-dollar IPOs even as they push enterprises toward usage-based billing. Cheap has become a performance claim in its own right, and it reorders the competitive frame around who can sustain the margin.

Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes · Source A reopened systemd issue claims journald writes tens of kilobytes of disk per log line — 49KB or more on ext4, 110KB+ on btrfs — such that a VM pushing two lines a second sustains ~50 IOPS. The reporter is explicit that this is a repeat of #15292, closed years ago, and that it survives kernel write coalescing. The concrete number is the story: journald’s binary format stores metadata far out of proportion to the 150-byte haproxy lines that generated it, which is why a nominally trivial log stream can grind a disk-backed VM. Whether the diagnosis or the blame is fully fair, a metric this far from “within an order of magnitude of syslog” is a real operational cost hiding in a default.

In Brief

  • Reuters reports Apple trained its own large language model for the China market with help from Alibaba, rather than relying on a US-built model. (Source)
  • A federal judge ordered Google to strip out the “anticompetitive friction” that makes rival Android app stores harder to find and install, in remedies stemming from the Epic case. (Source)
  • Apple sent a fresh wave of mercenary spyware threat notifications to users in 110 countries it suspects may have been targeted. (Source)
  • President Trump imposed tariffs of up to 100 percent on heavier and “sensitive” drones, including models over 55 pounds or equipped with docking stations or thermal imaging. (Source)
  • Google will now let users strip visible watermarks from AI-generated content while keeping the invisible SynthID marking in place. (Source)
  • Geoffrey Litt argues that understanding, not generation, is the new bottleneck for AI systems, citing a script Claude wrote during a website migration as the example. (Source)
  • RustDesk now supports true unattended remote access on Wayland, closing one of the harder gaps in Linux remote desktop. (Source)
  • Cloudflare detailed how it detects Model Context Protocol traffic and helps secure it, noting most permission systems were designed around a single human user rather than agents. (Source)
  • Grok 4.6 is now available in GitHub Copilot across the CLI, IDE, and cloud products. (Source)
  • OpenAI offered a hands-on look at its take on agent memory, drawing on computer history to explain the design. (Source)
  • An FT-sourced graph undercuts Dwarkesh Patel’s prediction that Anthropic would end the year near a $100–150B revenue run rate, since roughly $11.5B in Q2 would require more than doubling by Q4 even as prices fall and competition rises. (Source)
  • After a few days at Usenix Security, a researcher reflects on “Going Dark” and the era of law enforcement hacking. (Source)

One Line

This is officially Firefox’s killer app. I will never go back to Chrome so long as uBlock Origin is supported on Firefox.

— @hispanicat7hedisco.bsky.social, in a Bluesky post quoted by PCWorld

Search MacWorks

Enter at least two characters.