Alibaba shipped a 27B model in August that scores 52 on the Artificial Analysis index against its predecessor's 38, with an identical network. What that means for running a capable model on one graphics card, and what it costs in memory and tokens.
The top four models on the Artificial Analysis index sit within four points of each other and vary threefold in cost per task. The cheapest of them is the one you can download. What changed in July was pricing power, not capability.
NVIDIA, Microsoft, Meta, and 22 others signed a letter arguing the US should not restrict downloadable models. We read the three pages, checked the one claim we can verify from our own work, and noted what the letter avoids.
The hardware to hold a capable model got cheap, and the models got small enough to fit. What runs on the machines in your bedroom, and the two numbers that decide it.
Anthropic's engineering blog describes how to build agents in production. agent-stdlib packages those procedures into installable skills, MCP servers, and safety hooks you can run in any harness.
We built our production website with Qwen3.6-27B, an open-source model from Alibaba. What frontier open-weight models can do now, and where you still need a person in the loop.