TOGETHER WITH TODAY'S PARTNER

Good morning, {{first_name|there}}. The cheapest frontier model on earth introduces surge pricing tomorrow — and almost nobody has repriced their stack for it.

Read time: 3 minutes. Same time, every weekday — rate today's issue at the bottom.

🚀 The Big Story: DeepSeek ends the price war it started

Starting tomorrow, August 16, DeepSeek splits its API into peak and off-peak billing — and peak output on V4-Pro jumps from $0.87 to $3.96 per million tokens. That's 4.6x, on the model that made $0.14 inputs normal.

  • The new rates: V4-Pro cache-miss input goes $0.435 to $1.32 at peak; V4-Flash output goes $0.28 to $1.32. Off-peak is exactly half of peak, so nothing goes back to yesterday's price.

  • The clock that matters: peak is 01:00-04:00 and 06:00-10:00 UTC — roughly 9pm-midnight and 2am-6am US Eastern. If your traffic is US business hours, you land in off-peak and eat about 2.3x, not 4.6x.

  • The reason given: DeepSeek says it's adjusting "to allocate resources more reasonably." Translation: inference demand outran the subsidy. It warned developers on August 6 and confirmed rates on the 13th.

Jason's take: The floor was never a technology curve — it was a subsidy, and subsidies end. Anyone who priced a product off DeepSeek's rock-bottom tokens just found out their gross margin was rented, not owned. The lesson isn't "leave DeepSeek," it's that single-provider inference is now a business risk, not an architecture preference.

⚡ Quick Hits

  • Gemini 3.7 Flash lands cheap and good. Google's Wednesday drop jumped DeepSWE from 49% to 65.3% and ships at $0.75 in / $3.75 out per million — with 50% off through year-end.

  • OpenAI's Sol hits 750 tokens per second. The Cerebras-powered GPT-5.6 Sol runs up to 14x faster than standard, which quietly makes real-time agent loops viable.

  • Anthropic cut Opus 5 to roughly half of Fable 5. The frontier labs are discounting the same week the discount house raises prices — read that twice.

  • Anthropic signed a $9.1B, 20-year compute lease. Riot Platforms' 191MW Rockdale, Texas campus delivers 96MW by December 2027 — capacity is being booked out to 2048.

  • Apple built a China-only model with Alibaba. It has registered on-device generative AI with Chinese regulators to fight Huawei on home turf.

  • Encrypted reasoning blocks turned out to be portable. Researchers recovered 315,320 reasoning blocks, 367 PII artifacts, and 182 credentials across Anthropic, OpenAI, and Google systems.

📡 Trending on X

  • "Rented margin" is the phrase of the week. Builders are posting their DeepSeek invoices and doing the 4.6x math in public — the replies are half panic, half I-told-you-so.

  • The open-weights crowd is split. V4-Pro is still MIT-licensed at 1.6T parameters, so one camp says "just self-host" while the other points out GPU rental costs more than the new API price for most teams.

  • Meta's Glimmer bet reopened the open-vs-closed fight. A 30B Apache 2.0 model that runs on a laptop landed the same week the cheap API stopped being cheap — the timing is doing a lot of arguing for Meta.

  • Multi-provider routing is having a moment. Gateway and router repos are trending as teams discover their "vendor-agnostic" architecture had exactly one vendor in it.

📺 Trending on YouTube

🛠 The Workflow: Price-proof your AI stack before Sunday

You have one day. This is the 30-minute version that protects your margin.

  1. Pull your last 30 days of token usage and write down one number: monthly spend.

  2. Multiply it by 4.6 for peak and 2.3 for off-peak. If either number breaks your P&L, you have a pricing problem, not an AI problem.

  3. Check when your jobs actually run. Move every batch, cron, and backfill outside 01:00-04:00 and 06:00-10:00 UTC — that alone halves the hit.

  4. Turn on prompt caching. Cache hits stay dramatically cheaper than cache misses, and most apps re-send the same system prompt thousands of times a day.

  5. Put a router in front of your calls today — Gemini 3.7 Flash and Opus 5 both got cheaper this week. One config change beats one code rewrite.

Reply with the word "ROUTER" and I'll send you the one-page provider-cost comparison I built this morning.

🧰 Trending Tools

  • Rork Max — builds real iPhone apps from plain language. For operators who want a paid micro-app without hiring a dev.

  • Google Stitch 2.0 — turns a text description into a usable UI design. For founders who keep losing weeks to "what should this look like."

  • AdAnt AI — generates ad concepts and full campaign angles. For marketers who bill for creative volume.

  • Fathom 3.0 — records and summarizes every call automatically. For consultants who charge for follow-through, not note-taking.

  • Muse Code — writes, tests, and validates software autonomously. For small teams trying to ship like a big one.

📣 Put your brand here. The AI Innovator reaches AI-first operators, creators, and marketers every weekday. Primary sponsorships are now booking.

That's a wrap

Monday: the first real invoices from DeepSeek's new pricing land — we'll show what teams actually paid versus what they budgeted.

Enjoying the new format? Share your link — every referral keeps this free:

How was today’s email?

(Tell us what you liked or what could be better)

Login or Subscribe to participate

You're reading the 3-minute AI Innovator — same time, every weekday. Hit reply and tell me what you want more of. I read every one.