News/Guide
Guide · Aug 16, 2026

Four models in nine days. You need a policy, not an opinion on each one.

At this cadence, evaluating every new model is a full-time job nobody assigned. The way out is a standing rule that answers the question before it is asked.

361361 NetworkEditorial team4 min read

Kimi K3 on 6 August. MAI-Code-1.1-Flash on 11 August. Gemini 3.7 Flash on 13 August. Grok 4.6 on 14 August. Four models in nine days, in one product, each arriving with a policy for an administrator to enable.

Most teams are handling this by treating each release as a decision. That does not scale at this cadence, and it produces the worst outcome available: a backlog of unanswered admin toggles, developers who cannot get a model they read about, and no actual policy.

Separate enabling from adopting

These are two different decisions and conflating them is why the backlog exists.

Enabling means a model appears in the picker. The cost is a policy toggle and an entry on a list; the risk is almost entirely about cost and data handling, not quality. Adopting means the model becomes what your team reaches for by default, and that genuinely deserves evaluation.

Most model releases warrant a fast yes or no on enabling and no opinion at all on adopting. Answering the first question quickly is what stops the second one from piling up.

The three-tier structure

Write these down once and most new releases classify themselves.

  • Default — what runs when nobody chooses. One model, changed deliberately and rarely, ideally a cheap fast one because most work is routine
  • Available — enabled and selectable by anyone who wants it. Should be generous; the cost of an extra entry in a dropdown is close to zero
  • Blocked — not enabled, for a stated reason: a data-handling concern, a pricing model you cannot forecast, or a vendor your organisation has a position on. The reason matters more than the decision

The default deserves the evaluation, and nothing else does

Everything you would do for a serious model evaluation — the three saved tasks, the refactor, the bug with a reproducible test, the explanation of code you already understand — applies to the default and to nothing else.

That is the whole efficiency gain. One careful evaluation per quarter beats four rushed ones per fortnight, and the model in the default slot is the only one whose quality affects work nobody consciously chose.

A useful rule for changing the default: a new model has to beat the incumbent on your own tasks, not on a benchmark, and it has to do it twice. Agent runs are not deterministic, and a win that a second run erases is not a win.

What actually deserves your attention

Not every release is equal. Four signals mean a release is worth reading properly rather than classifying:

  • A deprecation attached — MAI-Code-1-Flash stops existing on 10 September, and a model disappearing is an incident if something automated names it
  • A new pricing mechanism — Kimi K3 arriving billed by token rather than by premium request changed what a heavy user costs, which is a budget question, not a model question
  • A "Preview" in the policy name — Gemini 3.7 Flash is gated behind one, and preview status means it can change under you
  • A new capability class rather than a better score — native vision in a cheap small model changes what work you can route to it

Write the standing answer down

Four sentences in a place your team can find, which converts every future release from a discussion into a lookup.

  • Our default model is X, and we review that in [month]
  • New models are enabled unless they fail [your stated criteria] — data handling, pricing you cannot forecast, or a vendor position
  • Anyone may select any enabled model; nobody needs permission to try one
  • Deprecations and pricing-mechanism changes get escalated; everything else gets classified

Then check quarterly, not per release

Once a quarter, look at three things: which models your team actually used according to the per-model token report GitHub shipped on 11 August, whether your default is still the right one, and whether anything on the blocked list has changed enough to revisit.

That is fifteen minutes, four times a year, and it replaces a decision every two days. The point of a policy is not to be right about every model — it is to stop the question consuming attention that belongs somewhere else.

More news