This website uses cookies

Read our Privacy policy and Terms of use for more information.

In partnership with

Happy Tuesday, friends.

The most interesting AI release this week can't write you a sentence. It picks one of your options in under 40 milliseconds, and that is exactly why it matters. I went looking for the same pattern outside the labs and found it in a Bucharest accounting startup and on Google's own plan page.

The Cloudflare and Amazon piece is the one to read first if you run support or sales operations. The DeepSeek story shows what happens to the expensive tier once the cheap one is good enough, and if anyone on your team uses free Gemini, the last story has a date that lands on Friday.

Cloudflare and Amazon shipped models that only choose, and they cost almost nothing

Cloudflare released Clef and Clef-flash on 1 October, and Amazon shipped Strands Decider 2B the same day. These are "decision models": you hand one an input and a set of typed questions, and it returns an answer with a confidence score instead of text. Clef-flash costs $0.09 per million input tokens, charges nothing for output and answers in a median 38.8 milliseconds. Both weights are Apache 2.0, and Amazon's 2-billion-parameter version runs on a MacBook at around 150 milliseconds.

They are chasing TypeSafe's Jev, which launched in mid-September at $0.042 per million input tokens and was, according to Clouded Judgement, used by about 13% of paid teams on Vercel's AI Gateway within 24 hours. Treat the "100x cheaper" headline carefully: the same analysis puts it closer to 4x against a mainstream reasoning model, and the big models still win on accuracy. Its author's estimate is that 10 to 15% of today's tokens could move to decision models. Ticket routing, urgency scoring, intent detection and escalation are the obvious jobs, and as eesel's comparison puts it, "the sorting is the easy half."

Pull 500 support tickets your team already labelled this quarter, run them through Clef-flash and look at where it disagrees with your people. If it matches on most of them, put your expensive model behind it and let that one see only the cases the small model flags as low confidence.

Blu Dot surpasses 2,000% ROAS with self-serve CTV ads

Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:

After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.

The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.

“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”

Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.

A Romanian accounting startup uses AI for one step and people for the rest

Declaro launched in May with about €1 million from its founders and private investors, and it reached roughly 200 clients in five months. The target is 1,000 within a year. It handles about 25,000 transactions a month, and around half its clients moved over from another provider. The product connects to e-Factura, Romania's national e-invoicing system, and checks the tax authority's SPV messages every 15 minutes.

The AI part is deliberately narrow. The company says it uses AI selectively for document extraction, runs its own software for verification and keeps accountants for processing, advice and support. That is the decision-model shape from the first story: a cheap machine step at the front and a person at the point of judgement. CEO Alexandru Puiu was COO of Keez until March 2025, and these are company figures from a single interview, so read them as an early signal.

The trigger is the part worth copying. A state mandate left thousands of small companies needing compliant paperwork, and nobody wants to do it by hand. Look for a rule that took effect in your market in the last 12 months and left a boring, repetitive task behind.

DeepSeek retired its flagship because the cheap model was good enough

DeepSeek released V4.1-Flash on 10 September and began redirecting all V4-Pro requests to it from 14 September, at lower prices, until a V4.1 Pro exists. Off-peak output costs $0.60 per million tokens, and TNW cites Bloomberg Intelligence putting the overall price cut at about 32%. The model has 552 billion parameters but activates only 8 billion to read and 16 billion to write, which is how it stays cheap.

The news this week is the gap. Bloomberg Intelligence's Robert Lea, as relayed by AI Weekly, put V4.1 Flash at 81.1 on LiveBench against 83.4 for the best US model on the board, a gap of about 3% that stood near 15% earlier this year. On agentic coding it scores 77.3 against 66.1. DeepSeek's own testing shows it matching a leading US model on coding but trailing badly on reasoning exams (36.8 against 56.3 on Humanity's Last Exam), and the leaderboard numbers are one analyst's read.

Run your five hardest real tasks on the Flash tier before your next model contract renews, and let the pricier tier earn its place by failing something.

Google is rationing the better Gemini from 9 October

From 9 October, free users of the Gemini app lose Flash and Pro and keep only Flash-Lite. Reports citing 9to5Google trace the change to Google's updated support pages. AI Plus at $4.99 a month loses Pro as well. AI Pro at $19.99 keeps Flash-Lite, Flash, Pro and Deep Think, and Ultra at $99.99 keeps everything.

The change applies to the consumer app, not the developer API, and none of the coverage I found carries a reason from Google. Compute-based usage limits have applied since May, so this is a second step in the same direction. Put it next to the decision models and DeepSeek's retired Pro and the logic is the same everywhere: the cheap tier is the default, and the best model is a paid escalation.

Friday is the date to plan around. Find out who on your team does work prompts on a free personal Gemini account, and decide this week who gets a paid seat and who moves to the API.

Short Signals

Productivity: Agent-Reach. One install gives your coding agent read and search access to web pages, YouTube transcripts, RSS and GitHub, plus Reddit, X and LinkedIn with login. Tell your agent to install it from the repo's install doc (MIT). Logged-in platforms carry account-ban risk, so use a throwaway account. Try it on a Monday competitor sweep.

Design: OpenPencil. An open-source Figma alternative that opens .fig files, has a built-in AI chat and ships an MCP server with 100+ tools for Claude Code, Cursor and Windsurf. Install with brew install --cask openpencil or use the web app (MIT). Ask your coding agent to inspect and export one real frame from a file you already have.

Marketing: OpenMontage. Turns Claude Code, Cursor, Copilot, Windsurf or Codex into a video production studio, from script to edited cut. git clone, then make setup (Python 3.10+, FFmpeg, Node 18+). The licence is AGPL-3.0, so check it before building it into a client service. Make one 30-second explainer from your landing page and see how far it gets.

Sales: OpenOutreach. A self-hosted agent: describe your product and market, and it finds prospects, explains why each one qualifies and emails from your own mailbox. Install with uv tool install openoutreach (GPLv3); it needs an LLM key, a BetterContact key (40 free credits) and an email app password. EU senders should check consent rules and deliverability before pointing it at a real list.

Dev: DwarfStar (ds4). A local inference engine for a short list of open models (DeepSeek V4 family, GLM 5.x, Qwen 3.8 Flash Next) on Metal, CUDA and ROCm. git clone and make (MIT); Apple Silicon wants 96GB+ RAM, with SSD streaming on smaller machines. Worth a weekend if your data can't leave the building.

Next edition soon,
Çelik