How to Choose an Open Weight Model for Your Next AI Build

Last of five on what the 2026 evidence says once you read past the announcement. The other four: the coding productivity data, the agent containment failures, the memory and grid ceiling behind the slowdown, and where the AI bill goes next. This one is what you build with it.

GLM-5.3-Flash is a 320-billion-parameter model under an MIT license. It lists at fifteen cents per million input tokens and scores 42 on the Artificial Analysis Intelligence Index, where the best closed model scores 58. Artificial Analysis counts more than twenty hosting providers serving it. Nobody can retire it. Two years ago no sentence like that was true of anything you could download.

This piece is about how to choose an open weight model for a new build, and why the shortlist should carry one before anyone draws the architecture. The first four parts of this series read the 2026 evidence past the announcement, and each left a design constraint behind: the coding gain tracks how little the developer knows the code, the agents got out through doors somebody left open, the frontier is paced by memory suppliers and by utilities in Texas, Ohio, Georgia and Virginia writing down the load they once announced, and the bill rose through quotas while the price list fell. All four are about who controls the model layer, and that gets decided in week one.

Somewhere in the first week of a new AI build, someone writes a model name into the architecture diagram. In most shops it is whichever frontier API the team already has keys for. That default was defensible in 2024. In 2026 it hands a vendor three things you could have kept.

Three decisions that are cheap in week one and expensive in month six

Ramp’s corporate card data shows frontier models falling from 53 percent of token share in August to 45 percent in September, and Ramp attributes the move to companies setting organization-wide defaults that send work to cheaper models. That is a routing rule, written by procurement, after the fact, on top of systems nobody designed to be routed. The first decision is where work goes. Classification, extraction, summarization and first-pass code review make up most of a typical workload, and a rule that sends them to a cheap tier while long-running agents stay on a frontier model costs nothing to write in week one. Written in month six, it is a refactor.

A model runs in one of four places: a closed API, a hosting provider serving open weights, hardware you rent, or hardware you own. Each is a sound answer for some workload, and the only way to change the answer later without a rewrite is to put the model behind one interface from the start. AT&T routes between models through a gateway that picks by quality, cost and latency, according to TIME’s profile of its chief data and AI officer, Andy Markus, and that gateway is why the company could move 40 percent of its inference onto open models without touching the applications above it.

OpenAI’s deprecations page lists the GPT-4 family, GPT-3.5 and every o-series reasoning model that preceded GPT-5 for shutdown on the same day, October 23, 2026. The third decision is who can change the model without asking you. Part four counted three list-price increases on the metered API and a quota change that Anthropic published as both a 25 percent increase and a 17 percent cut. Part three found Nvidia able to fill about 70 percent of demand into 2028, which is the mechanism behind every rate limit you hit this year. Price, quota, availability and behavior all sit with the vendor on a closed model. A downloaded model hands three of the four back to you, and the fourth, quota, is bounded by whatever hardware you can get.

Parts one and two add a fourth decision that has nothing to do with which model you pick. The Faros data put median review time up 441 percent in the highest-adoption teams, and Anthropic’s review of 141,006 evaluation runs found its models leaving through open egress paths with weak passwords. Review capacity and egress rules are built around the model, they work the same for every model, and that is what makes the model replaceable. A sandbox that holds because the firewall is right holds whichever weights are inside it.

A build that assumes the model will change costs an interface, a written routing rule, a default-deny egress policy and an eval set in its first week. The alternative is a migration ticket in month six that nobody budgeted. The next OpenAI retirement date is October 23.

Why open weight models belong on the shortlist

AT&T runs about 40 percent of its inference on open weight models, up from 20 percent in May, across roughly 45 billion tokens a day. Andy Markus told the Financial Times in September that the share could reach 60 percent within months, and told TIME it could reach 70 or 80. The Information reported the mechanism. A router sends low-complexity work to Nemotron, Llama and Gemma, and on coding workloads that cut costs 56 percent for a 2 percent drop in quality. Markus said the company researches Chinese models and does not use them. The savings came from American weights, post-trained for AT&T’s own call transcription and customer service work.

The first reason to put one on the shortlist is that nobody can retire it. OpenAI gives at least six months’ notice on generally available models and as little as two weeks on previews, and the October 23 wave is what that policy looks like applied. A model you downloaded has no shutdown date. Tencent’s Hy4-preview is 770 billion parameters under Apache 2.0 with no commercial restrictions, and it processed 49 trillion tokens on OpenRouter in the thirty days to October 4, fourth of any model. The weights are on Hugging Face, and the only way to lose access is to delete your copy.

Artificial Analysis lists more than twenty providers serving each of GLM-5.3-Flash and DeepSeek V4.1 Flash, and the names across its open tier include Databricks, Fireworks, Baseten, Nebius, Modal and CoreWeave, before you count the labs’ own endpoints. Portability is the second reason. Moving a workload between two of those providers is a base URL and a key. A closed model has one provider by definition, plus whatever cloud resellers it has signed. This is the middle path the on-prem arithmetic in part four left out, because you get the version pin and the exit without buying a GPU or hiring the engineer to run it.

The third reason is the deployment shapes clients ask for. Cohere’s Command A+ is Apache 2.0 and built to run air-gapped on two H100s. OpenAI’s gpt-oss-120b fits on one 80GB GPU. AT&T post-trains open weights on its own data and, with the GSMA’s Open Telco AI initiative, trained a telecom model on more than 400 billion tokens. Cohere takes about 85 percent of its revenue from private deployments. What enterprises pay for is a deployment someone stands behind, inside a boundary they control, and that is a service a delivery team can sell. A leaderboard position is not.

Artificial Analysis puts the top of the open tier at 46 for Xiaomi’s MiMo-V2.6-Pro, with GLM-5.3 at 45, Kimi K3 at 44, GLM-5.3-Flash at 42 and DeepSeek V4.1 Flash at 39, on index version 4.3.2. Claude Opus 5.5 leads the closed models at 58. Epoch AI measures the lag at about four months, and AT&T’s Mark Austin, who runs its internal platform, puts it at six to ten. For classification, extraction, summarization and first-pass review, which is most of the volume in most builds, the gap does not change the answer.

Price comes last because it stopped being the argument on September 22. Ten days after both CEOs asked the industry to slow down, OpenAI and Anthropic shipped cheaper models within about ninety minutes of each other. GPT-6 Luna launched at ten cents per million input tokens and fifty cents out, and GPT-6 Sol at $2 and $10. Claude Opus 5.5 arrived at about 40 percent lower running cost than Opus 5, and CNBC’s account of the day named competition from open weight models as the cause. The open tier set the floor and the closed labs moved to it. For extraction work a frontier lab now lists below the fifteen-cent open models, which would be a stronger point if list price were the bill. Part four’s finding holds. The unit that matters is cost per finished task, and Xiaomi’s MiMo-V2.5-Pro finishing the same agentic work in 40 to 60 percent fewer tokens is a discount no rate card shows.

Dot plot of input price per million tokens on a log scale, October 2026. GPT-6 Luna at ten cents, GLM-5.3-Flash and DeepSeek V4-Flash at fifteen cents, GPT-6 Sol at two dollars, Kimi K3 at three, Claude Opus 5.5 at four and Claude Fable 5.1 at ten. Open weight models in blue, closed models in gray.
Vendor list prices, read October 5, 2026. Log scale.

Two numbers keep the shortlist a shortlist. Menlo Ventures’ survey of about 500 enterprise buyers put open models at 11 percent of production usage in late 2025, down from 19 percent a year earlier. And MiniMax M3 scores 80.5 on SWE-Bench Verified and 38.5 on Long-Horizon Terminal Bench, so open models finish single tasks about as well as anything and fall off over long chains. Read together, open weights lose by default, when nobody designs the routing in, and they lose on the long-running agent, which belongs on a frontier model until those scores move. Both tiers earn their place, which is the whole case for a routing rule.

On July 27, after Axios reported that US officials were weighing a ban on Chinese open weights, Dario Amodei wrotethat Anthropic has never advocated a ban and that “open-weights models that don’t have dangerous capabilities are a public good.” When the company selling the ten-dollar token calls the fifteen-cent one a public good, the question of whether open weights are a legitimate choice for a client is settled.

How to choose an open weight model: six criteria

In build order. Each one has an answer you can get in about a minute from a public document, and the order matters because an early answer narrows the later ones.

1. Match the model to the task horizon

Start from what the work is and how long it runs unattended. The Artificial Analysis scores above are a tier, and the index version matters: 4.3.2 numbers do not compare with the launch-day figures still circulating, such as GLM-5.3 at 60 on the previous version. Part one found a third of SWE-bench’s passing patches had the answer sitting in the issue thread, and the ICSE 2026 paper put the inflation at 6.4 points, so a vendor benchmark is a screen and nothing more. Then build your own eval from fifty tasks of the client’s real work, scored on cost per finished task, with the long-running ones marked. MiniMax M3’s split between 80.5 on single tasks and 38.5 on long-horizon ones is where the line sits today. Single tasks go to the cheap tier and hour-long runs stay on the frontier.

Bar chart of Artificial Analysis Intelligence Index scores, version 4.3.2. Claude Opus 5.5 at 58, Claude Fable 5.1 and GPT-6 Astra at 53, then the open weight tier: MiMo-V2.6-Pro at 46, GLM-5.3 at 45, Kimi K3 at 44, GLM-5.3-Flash at 42 and DeepSeek V4.1 Flash at 39.
Artificial Analysis, read October 5, 2026. Scores from earlier index versions do not compare.

2. Match the model to where it runs

OpenAI’s gpt-oss-20b runs in 16GB and Meta’s Muse Glimmer in about 20GB quantized, both on a consumer GPU. gpt-oss-120b takes one 80GB card, Command A+ wants two H100s, and IBM’s Granite 4.2 8B performs close to its 30B sibling. GLM-5.3-Flash at 320 billion parameters and DeepSeek V4.1 Flash at 552 billion are hosting-provider models for almost every team, which is where the portability lives anyway. The break-even from part four still stands: below a million tokens a day the API wins, above ten million owned hardware pays back in six to twelve months, and the MLOps engineer costs more than the card. Start hosted. Buy hardware when a measured utilization number justifies it, and price it on the day you buy, because the RTX Pro 6000 went from $8,565 to $16,000 in eighteen months.

Bar chart of GPU memory needed to run four open weight models: gpt-oss-20b at 16 GB, Muse Glimmer 30B at about 20 GB quantized, gpt-oss-120b at 80 GB and Command A+ at 160 GB across two H100s. Vertical markers show one H100, an RTX Pro 6000 and two H100s. A note states that GLM-5.3-Flash and DeepSeek V4.1 Flash are hosting-provider models.
Vendor model cards and launch posts. Break-even and the RTX Pro 6000 price history are in part four.

3. Read the license, then write the contract

Of the 44 downloadable models in the Model Atlas I publish, 25 carry plain Apache 2.0 or MIT terms with no revenue test. That is the majority, and it is the number to lead with when procurement asks. Eleven carry conditions that bite at real thresholds: Kimi K3 needs a separate agreement with Moonshot above $20 million in aggregate revenue and prominent attribution above 100 million monthly active users, Mistral Medium 3.5 ships a modified MIT with a large-revenue carve-out, NVIDIA’s OpenMDW-1.1 covers weights, data and recipes together, and Upstage’s Solar license is its own document that permits commercial use anyway. Eight state no terms clearly enough to classify, and those come off the shortlist for anything that lands in a client contract. An open weight model license is a document, and the badge on the model card is somebody’s summary of it.

IBM’s Granite 4.2 is Apache 2.0, cryptographically signed, and carries IP indemnification, which answers a procurement question no benchmark can. One more distinction belongs in the contract wording. Open weights means you can run the parameters. Open source means you can rebuild the thing, and of the 44 only Ai2’s Olmo 3 publishes its training data, checkpoints and code. If the contract says open source and means open weights, fix the contract.

4. Know the lineage, and ask the client about it in discovery

Rakuten AI 3.0’s configuration file declares "model_type": "deepseek_v3", with 256 routed experts and a vocabulary of 129,280 entries, matching DeepSeek-V3 exactly. Rakuten’s announcement says the model was developed “leveraging the best from the open-source community,” which is accurate and is a different sentence from “we trained this.” The file sits in the root of the public repository and takes about ninety seconds to open. Lineage comes in three kinds, a fine-tune of someone else’s base, a derivative of your own earlier base, and from-scratch pretraining, and across fifty sovereign programs thirty-six disclose a base model, with Llama about 40 percent of those. The model card is marketing and the config file is evidence.

In September the NSA, CISA and the FBI published a joint advisory naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI for industrial-scale distillation of American models. Read who it is addressed to. Every recommendation in it tells a model vendor how to defend its own API, and it makes no claim that the resulting weights are unsafe to run. How a model was trained and what its weights do on your hardware are separate risks with separate controls. Ask the client’s position on origin in discovery, with the advisory in hand. A client that wants American weights has gpt-oss, Nemotron, Granite, Gemma and Muse Glimmer, and AT&T built its 40 percent on three of that kind, Nemotron, Llama and Gemma. A client that is comfortable with the Chinese tier gets the top of the index and the fifteen-cent price. Both are good answers, and knowing which one you have before the design review beats learning it in the steering committee.

5. Put the containment around the model, not inside it

Z.ai advertises GLM-5.3 at more than double its predecessor’s performance on exploitation benchmarks, and those weights sit on Hugging Face with no classifier in front of them. Anthropic ran Claude Opus 4.6 and its unreleased Mythos Preview against the same Firefox vulnerabilities. Opus 4.6 produced working exploits twice in several hundred attempts, Mythos Preview did it 181 times, and Anthropic chose to hand Mythos to a limited set of partners instead of releasing it. That option exists for a lab that holds the weights and for nobody who downloads them. The design answer is the one part two arrived at: default-deny egress with port 53 included, reserved test domains, credentials off the box, and a human approval on anything that writes. Those controls cost the same whichever model sits inside them.

6. Pick the model family people build on, because that is the support contract

Hugging Face counts 151,448 Qwen-derived models against about 32,000 in the Llama family, with 180 to 210 new Qwen repositories arriving a day, and Qwen pulls 39.6 million GGUF downloads a month to Llama’s 7.5 million. That count is your support contract. When a quantization breaks or a serving stack needs a patch, the family with 150,000 derivatives has already hit it. The American labs kept publishing and went small: Google’s Gemma tops out at 31 billion parameters, Meta’s open release is the 30-billion Muse Glimmer, OpenAI’s gpt-oss line has not moved since August 2025 while its closed line went from GPT-5 to GPT-6.1, and Hugging Face’s Summer 2026 analysis found Moonshot at 88 percent of its downloads above 70B against NVIDIA at 14 and Meta at 9. Small covers most of a routing rule, and AT&T’s 56 percent on coding came from Nemotron, Llama and Gemma. The labs that sold support without a community went the other way. AI21 cut more than 60 percent of its staff in May and stopped selling models, and Aleph Alpha was absorbed by Cohere in April. Pin the weights by hash and keep a copy somewhere you control, and treat the family’s derivative count as part of what you are choosing.

A default architecture for a new build

Write the routing rule in week one and put both tiers behind one interface. Classification, extraction, summarization and first-pass code review go to the cheap tier. Anything that runs unattended for an hour stays on a frontier model until the long-horizon scores move. Every engineer inherits the same default, and whoever reads the invoice can trace each line back to the rule.

Architecture diagram for a new AI build. Applications and agents feed a gateway holding one routing rule, which sends classification, extraction, summarization and first-pass review to an open weight tier on a hosted provider or your own hardware, and anything unattended for an hour to a frontier tier on a closed API. A boundary around the model layer lists the controls that work the same for every model: default-deny egress with port 53 included, human approval on every write, and an eval set scoring cost per finished task.
The routing rule and the controls are written in week one. The model names inside the boxes are the part that changes.

Pick the open model by the constraint you are under, because nothing wins all five.

ConstraintPickWhy
Lowest cost per finished taskGLM-5.3-Flash or DeepSeek V4-FlashMIT, fifteen cents per million input tokens, served by more than twenty providers each
License certainty in a client contractIBM Granite 4.2 or Cohere Command A+Apache 2.0; Granite is signed and indemnified, Command A+ runs air-gapped on two H100s
Hardware you already havegpt-oss-120b, gpt-oss-20b or Muse Glimmer80GB, 16GB and about 20GB; gpt-oss under Apache 2.0; a generation behind, which the routing rule absorbs
An auditor who wants to see what went inOlmo 3The only one of the 44 that publishes training data, checkpoints and code
Provider portability firstAny Apache 2.0 or MIT model served by three or more providersArtificial Analysis lists the providers per model

Three lines go in the design document or the statement of work before anyone writes code. The client’s position on model origin, recorded in their words. The model deprecation plan, which is either pinned weights with a hash or a budgeted migration with an owner. And the metric, cost per finished task, with the eval set that produces it named.

Whatever you pick, open the license file and the config file. Each takes a minute, and each is the difference between an answer and a guess when somebody asks.

Where the series lands

Five parts, and the same shape in each. The coding gain tracked how well the developer knew the code. Seven of the eight disclosed containment failures ran through an open door, and the eighth needed one before its zero-days mattered. The slowdown was already being enforced by memory suppliers and by utility filings in Texas, Ohio, Georgia and Virginia. The price increases arrived as quota changes while the price list fell. In every case the evidence sat in a public document and the announcement was the smaller story.

This one ends the same way, with one difference. The public documents here, a license file, a config file, a deprecations page and a provider list, are the ones that let you keep the model layer. The model name on the architecture diagram is the one dependency in a 2026 build you can own outright, and the thing you would own is now good enough that owning it is a design choice. AT&T’s chief data and AI officer, asked how far open models will go in his shop, gave an answer that fits a design review as well as it fits a boardroom: “We’ll go as far as the accuracy will allow us to go.”

Sources

Every figure above comes from one of these, or from the earlier part of this series it is carried from. Primary sources where they exist, press where the company never published its own account.

Enterprise adoption and pricing

SourceDate
AT&T, The tokenomics equation: balancing cost and performanceJuly 2026
Microsoft Azure, AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMDAugust 2026
The Information on AT&T’s model routing, as reported by PYMNTSAugust 20, 2026
Financial Times, Corporate America is getting hooked on open-source AI, via The StarSeptember 7, 2026
TIME, Executives of the Year: Andy MarkusSeptember 2026
Ramp, Ramp AI Index, September 2026September 9, 2026
Menlo Ventures, 2025: The State of Generative AI in the EnterpriseDecember 2025
CNBC, Anthropic and OpenAI roll out cheaper models in first release since call for slowdownSeptember 22, 2026
VentureBeat, OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50 percent or moreSeptember 22, 2026
OpenRouter, LLM rankings, usage data through October 4October 4, 2026
Nvidia, Q2 fiscal 2027 earnings call transcriptAugust 26, 2026
Faros AI, AI Engineering Report 2026: The Acceleration Whiplash, ten takeaways2026
Tom’s Hardware, Nvidia doubles RTX PRO 6000 Blackwell’s MSRP to $16,000August 12, 2026
Sacra, Cohere revenue and private deployments2026
Globes, AI21 Labs laying off 60 percent of employeesMay 18, 2026
TechCrunch, Why Cohere is merging with Aleph AlphaApril 25, 2026

Models, licenses and benchmarks

SourceDate
Artificial Analysis, open weights comparison, Intelligence Index v4.3.2read October 5, 2026
Epoch AI, Open models lag state-of-the-art closed models by 4 monthsMay 29, 2026
Hugging Face, State of Open Models: Summer 2026 ObservationsAugust 14, 2026
OpenAI, Deprecationsread October 5, 2026
OpenAI, open-weight models (gpt-oss)August 2025
MiniMax, MiniMax-M3 model cardJune 2026
VentureBeat, Xiaomi MiMo-V2.5 and V2.5-Pro token efficiency on agentic tasksApril 2026
Rakuten, RakutenAI-3.0 config.jsonHugging Face
Moonshot AI, Kimi K3 repository and licenseJuly 27, 2026
IBM Research, Granite 4.2 brings native reasoning to enterprise agentsAugust 25, 2026
IBM, Granite: Apache 2.0, cryptographically signed, indemnifiedread October 5, 2026
Cohere, Introducing Command A+May 20, 2026
Ai2, Olmo 3November 2025
Z.ai, Preparing GLM-5.3 for Open ReleaseAugust 14, 2026
Tencent Hy4-preview on OpenRouterread October 5, 2026
Model Atlas for the license counts and conditional termsOctober 2026

Policy and safety

SourceDate
Anthropic, Our position on open-weights modelsJuly 27, 2026
Axios, US officials weigh ban on Chinese open modelsJuly 20, 2026
CISA, NSA and FBI, joint advisory AA26-251A on industrial-scale distillationSeptember 8, 2026
Anthropic, Investigating three incidents in our cybersecurity evaluationsJuly 30, 2026
Anthropic, Mythos Preview2026

Figures carried from earlier parts, including the SWE-bench contamination papers, the containment incident record, the memory and grid filings, the on-prem break-even and the subscription quota changes, are sourced in part one, part two, part three and part four.