Every week our pipeline scrapes the model catalogs and vendor blogs, then judges each item against one question: would this actually improve something we run in production? Verdicts below. Watch = interesting but unproven claims · Adopt = earned a place · Ignore = noise · Product Input = a feature idea, not a factory change.
30 items judged. 4 kept, all from model-watch (Tier 1, direct registry evidence): two INCUMBENT:YES alias re-points (haiku-4.5 affecting quincy's locked-consumer launch path, and sonnet-4.6 affecting my-day/lengua cost-bound incumbents) escalate to Immediate Factory Upgrade per the outage-prevention exception, with quincy ranked highest given launch proximity. One INCUMBENT (UNVERIFIED) alias re-point (opus-4.7) lands as Factory Candidate — first step is a config grep across 5 apps, resolving to a migrate-to-Opus-5 cleanup per standing Opus-tier deprecation policy rather than a re-eval of a dead-end snapshot. One new model (deepseek-v4-flash-0731) is aggressively priced but has no benchmark yet, so it's held at Watch pending its already-queued auto-eval scorecard. The remaining 26 items are all Tier 3 practitioner-blog essays (chip-huyen, eugene-yan, hamel-husain, one HF/IBM vendor post) covering evals, agents, and prompting — generic, mostly dated 2023-2025, none naming a Factory product decision or artifact gap; 7 were already-backlogged duplicates of prior triage. All dropped per the promotion floor: reading, not compounding capability.
New model deepseek/deepseek-v4-flash-0731 posts an aggressive price point ($0.09/$0.18 per M prompt/completion tokens, $0.018/M cached-read) with a 1M-token context window — cheap enough to plausibly undercut the current cost-bound fallback tier (claude-haiku-4.5) if quality holds, but no benchmark evidence exists yet.
Why this verdict: Tier 1 source (direct provider/registry pricing data). Novelty Medium — the price point is aggressive enough to be plausibly frontier-relevant, unlike routine catalog churn, but quality is entirely unverified pending the automated eval already in motion. Expected gain Medium if quality holds; effort Low since evaluation is already automated. Held at Watch, not Candidate, because no benchmark evidence exists yet.