Every week our pipeline scrapes the model catalogs and vendor blogs, then judges each item against one question: would this actually improve something we run in production? Verdicts below. Watch = interesting but unproven claims · Adopt = earned a place · Ignore = noise · Product Input = a feature idea, not a factory change.
30 items judged, 26 dropped, 4 kept. Bulk of intake was tier-3 blog/newsletter synthesis or already-backlogged canonical reading (hamel, karpathy, lilian-weng) with no named product decision or artifact this week — dropped per the default-discard mandate. Of five model-watch entries, two were variant/codename SKUs of already-covered families (gemini-3.5-flash-lite, gpt-5.6-luna) and dropped; claude-opus-5 and kimi-k3 clear Factory Candidate — opus-5 as a deprecated-4.6/4.7/4.8-pin cleanup across nine named apps, kimi-k3 as a doctrine-confirmed frontier-adjacent cheap-swap candidate, both pending their auto-eval scorecards. gemini-3.6-flash held at Watch pending its own scorecard. Two Supabase changelog items: the remote MCP server claim was dropped as already-implemented (this session's own MCP auth already uses OAuth, not legacy PAT), while the Feb 2026 Postgres agent-rules update is held at Watch pending a gap-check against the existing best-practices skill to avoid duplicating it.
Version-bumped Gemini Flash-tier model detected via automated catalog watch; auto-eval scorecard is queued but not yet produced, so its price/quality position relative to the current Flash-tier rotation pin is unverified.
Why this verdict: source tier 1 (direct catalog detection); novelty Medium — a real version bump, not a bare variant suffix, so it clears the SKU-drop rule; materiality is plausible but unconfirmed without a scorecard, so expected gain is held Low until verified against effort Low (fully automated). Watch, not Candidate, until the eval lands.