Something broke in the price of intelligence this week, and four stories fell out of it. Companies stopped buying the expensive model by default. Anthropic cut its own price before anyone made it. A young insurer in Serbia proved you can beat centuries-old incumbents with a technique from 2016 and a payroll of twenty-five-year-olds. And Washington started arguing about whether the cheap option should stay legal.
The through line is that the winners this week were not the ones with the best model. Read the Ominimo round and the Sent raise back to back and you get the same shape twice: small teams attacking large categories by owning the layer underneath instead of renting it. Then read the open-weights letter and notice who refused to sign.
The expensive model stopped being the default
On July 26 the Wall Street Journal reported that corporate America has flipped from "tokenmaxxing," burning tokens to look serious about AI, to "thrift-maxxing," shopping for the cheapest model that finishes the job. Fortune published the ladder the same weekend: $50 per million output tokens on Anthropic's Fable 5, $15 on Moonshot's Kimi K3, $4.40 on Z.AI's GLM-5.2, and roughly 87 cents on DeepSeek V4-Pro. In one week this month, Chinese models took 57% of the tokens US firms consumed on OpenRouter.
Anthropic saw it coming. Two days before the Journal ran, it shipped Claude Opus 5 at $25 per million output tokens, half of Fable 5, which it beats. Both labs filed confidentially for IPOs in June, and gross margin is the one number no public investor has seen yet. Anthropic is tracking near 44%. PitchBook's read is that below 35% the fair value compresses by 70% or more, and it now has OpenAI leaning toward 2027 rather than this autumn.
The number worth stealing is Cursor's. Its field CTO priced building a browser from scratch at over $10,000 on GPT-5.5 alone, against $1,339 for the same job split between Cursor's own Composer and Opus 4.8. That is the measurement to run this month: cost per finished task, not cost per token. And when a vendor starts handing you six figures in free credits, and several are, read it as a statement about their prospectus rather than their generosity.
The New Rules of Online Visibility
Your customers are searching in places your strategy doesn’t reach.
So before your business is buried and left behind, you need to understand the new rules of SEO.
BELAY's SEO in the Age of AI report explains how search is changing, what AI means for your visibility, and the practical steps small businesses like yours can take to stay visible.
BELAY’s U.S.-based Marketing Assistants turn strategy into execution, helping your business stay visible, credible, and competitive in every search.
Novi Sad prices risk better than Zurich does
Ominimo took the first tranche of its Series B Monday, €20.1 million led by EBRD Venture Capital, at €1.48 billion, or $1.6 billion. The Recursive called it Serbia's first unicorn, though it registers in Hungary and engineers in Belgrade and Novi Sad, so Serbian-founded is the honest phrasing. Three ex-McKinsey insurance consultants, Dušan Komar, Laslo Horvath and Dennis Weinbender, started it in 2024. Zurich led the Series A fifteen months ago at €212.75 million, a 7x.
They did not win with a frontier model. Incumbents sort ten million drivers into ten to a hundred thousand risk buckets, so careful drivers subsidise reckless ones. Ominimo prices each driver individually with XGBoost, a technique data scientists have had since 2016. The staffing is the moat: engineers and data scientists are over two thirds of a hundred-person company against under 5% at a traditional insurer, average age 25, with eight international maths and physics Olympiad medalists on the payroll.
Gross written premium went from €26.3 million in 2024 to €157.8 million in 2025 to roughly €307 million annualised, at a combined ratio between 92% and 100% where competitors growing this fast usually run above 115%. They grew without buying growth: one in five policies sells through their own site, the rest through comparison engines and brokers where drivers already shop. If you are attacking an incumbent, the question is not whether your model is smarter. It is whether you can ship a pricing change in a day where they need six months.
A Kosovar founder just became a phone company
Sent raised $12 million on Tuesday, an oversubscribed Series A led by Companyon Ventures with Bessemer, Urban Innovation Fund and CP Overture. Co-founders Daniel Vataj and Betim Drenica built one API that sends over SMS, WhatsApp or RCS and picks whichever channel will reach the person, across 75-plus carrier relationships and more than a billion phone numbers. Twenty thousand developers route through it. Fourteen months ago the same company raised $3.55 million.
The line that matters is not the money. Sent is now a licensed telecommunications carrier in the United States, approved by the FCC, which moves it from a vendor sitting on top of carrier infrastructure to a participant inside it, with the right to allocate numbers directly. That is the Stripe move and the Plaid move: stop being a nicer interface and become the thing underneath. The category is worth about $77 billion, and while Twilio, Sinch, Infobip, Vonage and Route Mobile carry more than 60% of the world's paid messages, none holds 20% alone.
Notice who Infobip is: a Croatian company out of Vodnjan, in Istria. The incumbent under attack and the team attacking it are both regional stories, one built at home and one built in New York. Anyone sending verification codes or order confirmations should pull last month's delivery rates by channel before the next renewal comes up. Only 13% of consumers treat SMS as their main channel, and most companies are still paying about eleven cents a message to reach them there.
Fifty companies signed a letter. Anthropic did not.
On July 24 Nvidia published Open Weights and American AI Leadership, a three-page letter against US restrictions on Chinese open-weight models. It opened with 25 names including Microsoft, Meta, IBM, Palantir, Hugging Face, Mistral, a16z and Y Combinator, then doubled to 50 within a day as Google, OpenAI and Cloudflare joined. Anthropic, Amazon and xAI stayed off it. Three days later Dario Amodei answered in public: "Anthropic has never advocated for a ban on open-weights models."
There is no instrument yet. No bill, no executive order, no Commerce rule, just four mechanisms held in reserve, Treasury Secretary Bessent floating sanctions and Entity List designations, and a doctrine separating ordinary distillation from "large-scale, covert industrial distillation." Beijing is running the mirror image: its Ministry of Commerce has been consulting on export controls over whether foreigners should keep downloading Chinese model weights. The cheap tier could close from both ends inside one quarter.
The evidence is awkward for both capitals. The joint UK and US AI Safety Institute assessment on July 23 found Kimi K3 achieved code execution on zero of 41 exploit samples, where the most cyber-capable frontier models average twenty. When Hugging Face was breached this month, its commercial models refused the forensic queries, so the team self-hosted GLM 5.2 to read the attacker's log. Should the cheap tier be load-bearing in your cost model, the restriction reaches you through your US cloud provider's terms rather than through Brussels. Write down this week which workloads you would have to reprice.
Short Signals
Five tools to install or test this week, tagged by the seat they help.
Productivity: Meta AI will now run jobs on a schedule. Meta added Scheduled Tasks and Artefacts to its iOS and Android apps on July 24: recurring actions on a timer plus a library of everything the assistant has produced. Calendar-based morning briefings are the obvious first use. Limited markets at launch, with WhatsApp promised. Set one recurring job you currently do by hand on Monday mornings and see whether the output is worth opening.
Design: Figma's auto layout now behaves like CSS. Figma changed how auto layout computes on July 24, closing the small mismatches that made engineers add spacing corrections at handoff. New frames get it automatically, existing frames stay legacy until you opt in, and the legacy fallback disappears in January 2027. Flip one live handoff file and ask the engineer building it whether their usual corrections vanish.
Marketing: Braze shipped a hosted MCP server. Braze's July 23 release puts a remote MCP endpoint in early access, so Claude, ChatGPT or Cursor can query campaign, Canvas and segment analytics and write email templates and Content Blocks, with OAuth sign-in and no user PII exposed. There is a separate EU endpoint. Start read-only: ask which of your last ten campaigns underperformed and why, before letting it touch a template.
Sales: Clay turned market sizing into a plain-English query. Clay opened Search in beta on July 28, combining company, people and job data in one natural-language question, alongside Reporting that shows which sourced prospects became actual pipeline and a Lookalikes tool. Write your ICP as one sentence, then run Reporting against a list you sourced last quarter. If none of it converted, the definition is the problem.
Dev: Claude's API stops hard-failing on refusals. Anthropic added server-side fallbacks on July 24, routing flagged requests to another model instead of returning an error, plus the ability to swap an agent's available tools mid-conversation without invalidating the prompt cache. Both are in beta. Turn fallbacks on for any agent running unattended, where a refusal currently means a silent dead end.
Next edition soon,
Çelik


