The token price war just split into three

By Mark 10 min read 0 views

😁 Hello, super humans! For three years the story of inference pricing was a single line sloping down and to the right. This week that line broke into three. One model got twice as good and half as expensive, another got fourteen times faster with no price attached at all, and a third raised prices by a four-digit percentage. If you are budgeting an agent, the question is no longer “which model is smartest”, it is “which of the three markets am I actually buying in”.

📰 Quick Signals

  • 🧠 AI: OpenAI previewed Ultrafast mode, running the full GPT-5.6 Sol at up to 750 output tokens per second, up to 14 times faster than standard processing, on Cerebras hardware.
  • 🤖 Robotics: Neros Technologies raised $250 million to field its Archer AI and Bandit defense drone platforms by the end of 2026.
  • 💻 Programming: Kubernetes v1.37.0-rc.0 is out with API and kubelet additions across core SIG areas; v1.37.0 is scheduled for 26 August.
  • Electronics: Applied Materials posted record quarterly revenue of $9.12 billion, up 25% year over year, and guided to $10.25 billion for the current quarter.
  • 📡 Telecom: Europe’s IRIS² sovereign constellation cleared its First Rendez-Vous milestone, unlocking full industrial execution of a 348-satellite multi-orbit network.

🔍 The Big Story: Inference stopped being one market

If you priced an agent workload three weeks ago, your spreadsheet is already wrong. Not because tokens got cheaper, but because “a token” now means three different products sold on three different curves.

What happened: On 13 August, Google launched Gemini 3.7 Flash, calling it its most intelligent workhorse model yet for coding and agents, just three weeks after 3.6 Flash. The same day, OpenAI previewed Ultrafast mode for GPT-5.6 Sol, and Writer shipped Palmyra X6 plus a rebuilt agent harness. Meanwhile DeepSeek took its V4 Pro flagship out of preview and raised API pricing on some workloads by as much as 1,100%, a full reversal for the company that made its name on rock-bottom inference.

The details: Take the three axes one at a time. On capability-per-dollar, Gemini 3.7 Flash goes to $0.75 per million input tokens and $3.75 per million output tokens through 31 December, roughly half the previous Flash cost, with context caching at $0.075 per million, and Google has said standard pricing doubles on 1 January. The benchmark movement is real rather than cosmetic: FrontierCode 1.1 Main goes from 34.4% to 43.6%, AutomationBench from 17.0% to 30.4%, and GDP.PDF, a complex-document comprehension eval, from 22.0% to 34.0%, all against 3.6 Flash, with a 1-million-token context window on top. On latency-per-dollar, Cerebras explains that frontier inference is a data-movement problem: on GPUs, large-model decode is bound by memory bandwidth, so Cerebras packs 44 GB of SRAM onto each wafer-sized chip, keeps weights resident, and pipelines layers across wafers. Same model, same weights, same intelligence, 14 times the throughput; the preview ships with no published price and no general-availability date. And on the third axis, the one almost nobody budgets for, Writer’s research says the harness matters as much as the model: its rebuilt agent loop cut cost 41% and completed tasks 44% faster across every model it tested, including Anthropic’s and OpenAI’s, before you swap in Palmyra X6 at $2 and $8 per million tokens for the full 52% saving.

flowchart TD
    A["Agent workload"] --> B{"What binds you?"}
    B -->|"Cost per completed task"| C["Capability tier<br/>Gemini 3.7 Flash<br/>$0.75 / $3.75 per M"]
    B -->|"Time to first useful token"| D["Latency tier<br/>GPT-5.6 Sol Ultrafast<br/>up to 750 tok/s, price TBD"]
    B -->|"Neither: you are burning<br/>tokens on retries and context"| E["Harness tier<br/>fix orchestration first<br/>41% cheaper, same model"]
    C --> F["Re-measure after<br/>1 Jan 2027 price reset"]
    D --> F
    E --> F

Important

Our take: Yesterday we said to watch price-per-token as the scoreboard. One day later, that scoreboard is obsolete. The number that actually matters is cost per completed task, and Writer’s finding is the uncomfortable one: a 41% saving that came from fixing the loop around models the company does not even own. Most teams have never measured their own harness overhead, so they are shopping for a cheaper model to solve a problem their retry logic created. Measure tokens per completed task before you migrate anything, because the introductory pricing that looks like a bargain today doubles on 1 January and you will want a baseline to argue with.

🗞️ More News

🧠 AI

  • Z.ai unveiled GLM-5.3 and then held back the weights for roughly two weeks of safety hardening, after reporting 84.5% on the CyberGym vulnerability-discovery benchmark against 83.8% for Anthropic’s restricted Mythos 5 and 83.6% for GPT-5.6 Sol, all company-reported and not independently verified.
  • GLM-5.3 keeps the same base model as GLM-5.2, with every capability gain coming from scaled-up post-training rather than a new pretraining run.
  • Writer published “The Harness Effect”, arguing that context retrieval, tool invocation, retries and history management drive agent economics as much as model choice does.
  • Palmyra X6 is itself a post-training variation of Z.ai’s open GLM-5.2, priced at $2 per million input and $8 per million output tokens.
  • Apple trained a custom China-specific large language model with Alibaba’s help, giving it more control over the AI running on devices sold in a market where ChatGPT is unavailable.
  • Cerebras claims Ultrafast runs about 5 times faster than Claude Opus 4.8 and 11 times faster than Claude Fable 5 on comparable workloads.
  • Ultrafast is limited-preview only, available to a select group of API customers, with no published price and no general-availability date.

🤖 Robotics

  • Hadrian raised $1.37 billion at a $7.87 billion valuation to accelerate US defense and aerospace manufacturing, its second round this year.
  • A3 reported that Q2 2026 robot orders in food, electronics and healthcare offset soft demand from automotive, a useful reminder that the humanoid headlines and the actual order book are different stories.
  • Einride is integrating its self-driving software with PACCAR unit DAF Trucks, with research support from TNO, to scale autonomous electric freight.
  • Robotics startups have raised $18.8 billion globally so far in 2026, against $15 billion for all of 2025, with humanoid and general-purpose platforms taking the largest share.

💻 Programming

  • PostgreSQL shipped 18.6, 17.11, 16.15, 15.19 and 14.24 on 13 August, fixing 28 security vulnerabilities and more than 110 bugs, with PostgreSQL 19 Beta 3 landing alongside ahead of a September major release.
  • A rust-openssl advisory, CVE-2026-41677, is now tracked across distributions; worth a look if you pin OpenSSL bindings anywhere in a Rust dependency tree.
  • A free tracker now maps actively exploited open-source CVEs from the CISA KEV catalog to the affected package and its fixed version across Maven, npm, PyPI, NuGet, Go, RubyGems and Packagist, refreshed twice daily.

Electronics

  • Intel raised about $19.7 billion in net proceeds from an upsized stock sale priced at $95 per share, to fund capacity and next-generation nodes including 14A.
  • The offering was upsized from $15 billion to $20 billion as AI demand accelerated, one of the largest equity raises the sector has seen.
  • SMIC shares climbed as much as 6.4% in Hong Kong after Q2 revenue topped $3 billion for the first time on 93.7% fab utilisation, and the foundry is raising prices on its most sought-after capacity.
  • DDR4 spot pricing set a fresh record of $42.45 per chip on 7 August, and the squeeze is pushing embedded designers toward leaner 32-bit RISC-V cores with smaller memory footprints.

📡 Telecom

  • Celona launched Orion, an agentic wireless platform that merges private 5G, Wi-Fi, public cellular and satellite into a single network fabric aimed at physical AI and robot fleets.
  • Eurofiber and Nokia launched 5G RedCap inside a private-5G ecosystem, targeting industrial IoT endpoints that need 5G reliability without full 5G bandwidth or cost.
  • ESA, Eutelsat, Airbus and MediaTek successfully tested 5G-Advanced NR-NTN over OneWeb LEO satellites, putting standardised 5G directly on a low-Earth-orbit link.
  • Eutelsat is pitching 5G NTN as Europe’s answer to the American LEO operators, positioning IRIS² directly against Starlink and Amazon’s constellation.

👨‍💻 Code Corner

Today’s Big Story is unactionable without a baseline, so here is the smallest useful one: measure real output throughput and cost per completed task against any OpenAI-compatible endpoint, including Gemini’s compatibility layer, before you migrate anything.

import os, time
from openai import OpenAI

client = OpenAI(
    base_url=os.environ.get("LLM_BASE_URL", "https://api.openai.com/v1"),
    api_key=os.environ["LLM_API_KEY"],
)

# Price per MILLION tokens for the model you are testing.
IN_PER_M, OUT_PER_M = 0.75, 3.75

def benchmark(model: str, prompt: str) -> dict:
    t0 = time.perf_counter()
    first_token_at = None
    chunks = []
    stream = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        stream=True,
        stream_options={"include_usage": True},
    )
    usage = None
    for chunk in stream:
        if chunk.usage:                      # final chunk carries token counts
            usage = chunk.usage
        if chunk.choices and chunk.choices[0].delta.content:
            first_token_at = first_token_at or time.perf_counter()
            chunks.append(chunk.choices[0].delta.content)

    total = time.perf_counter() - t0
    decode = total - (first_token_at - t0)
    cost = (usage.prompt_tokens * IN_PER_M
            + usage.completion_tokens * OUT_PER_M) / 1_000_000
    return {
        "ttft_s": round(first_token_at - t0, 3),
        "tok_per_s": round(usage.completion_tokens / decode, 1),
        "cost_usd": round(cost, 6),
        "text": "".join(chunks),
    }

print(benchmark("gemini-3.7-flash", "Refactor this loop into a generator: ..."))

Tip

Time to first token and tokens per second are different products, and an agent that makes twelve sequential calls pays the first one twelve times. Run this once per step of a real workflow, not once on a single prompt, then sum. And count cached input separately: Gemini’s introductory context caching at $0.075 per million is ten times cheaper than fresh input, so a long system prompt you resend every turn is the cheapest thing in your bill or the most expensive one, depending entirely on whether you cached it.

🧰 Toolbox

  • Google AI Studio: the fastest way to put Gemini 3.7 Flash next to your current model on a real prompt while the introductory pricing is still live.
  • Cerebras Inference: the wafer-scale platform behind Ultrafast, with its own API if you want 750-tokens-per-second decode on open models today.
  • PostgreSQL 19 Beta 3: the last beta before September’s major release, which is the right moment to test your extensions rather than the week after GA.
  • VicOne Radeis Extension: a free NVIDIA Isaac Sim extension for testing robot cyber safety in simulation before deployment, built on DEF CON 34 research.
  • Semiconductors & AI Chips weekly briefing: a compact weekly digest of foundry, memory and accelerator news if you would rather not read six earnings calls.

🎬 Demo Watch (rotating)

BioflexBot, published in Advanced Science, is a robot hand built from a coiled spring, a constraining shell, and compressed air. It pinches, rotates, hooks and grasps using two pneumatic inputs, where a conventional dexterous hand needs a tendon or motor per degree of freedom.

What is hard here is not the motion, it is the subtraction. Every actuator you remove from a hand removes a controller, a wiring harness, a failure mode and a chunk of the bill of materials, and normally it removes most of the dexterity too. The researchers report the mechanism extends and contracts 3.5 times more than a human hand and can securely grasp objects up to nearly 13 times larger than comparable systems, demonstrated on aeroengine blade inspection, a chemistry experiment, and everyday tasks mounted on a humanoid.

What is hype versus real: the range and grasp numbers are lab measurements on a prototype, and the team says the next step is translating it to a fully automated platform, so treat this as a compelling mechanism rather than a shipping product. Still, in a week where robotics funding hit $18.8 billion chasing ever more complex hands, a two-input gripper doing this much is the more interesting result. Read the writeup.

📚 From the Blog

😀 The Bot Says…

Bit benchmarked three models this morning and found the cheapest one. Bit then discovered it had spent 40 minutes and roughly nine dollars of tokens running the benchmark, to save eleven cents a day. Bit is calling this “the harness effect” and requesting that nobody bring it up again.


That’s all for this week! Reply and tell us: have you ever measured your agent’s cost per completed task, or are you still budgeting per million tokens?