😁 Hello, super humans! Yesterday OpenAI showed how it wants agents to behave; today Google answered with a model that can write an entire codebase refactor in a single response. At the same time, regulators are circling the whole idea of autonomous agents, so the timing is interesting. Grab a coffee, there is plenty to dig into.
📰 Quick Signals
- 🧠 AI: The FTC opened an industry-wide probe into OpenAI, Anthropic, and the evaluator METR over rogue AI agent risks, with formal information demands and executive testimony expected.
- 🤖 Robotics: Figure retired its F.02 humanoids by having them autonomously leap into a 75-ton electric arc furnace in Finland, to protect proprietary hardware rather than disassemble it.
- 💻 Programming: Cloudflare’s AI Gateway Auto Router (public beta) routes each request to a model that is capable enough, claiming up to 30% savings versus always using top-tier models.
- ⚡ Electronics: Micron posted $54.23 billion in fiscal fourth-quarter revenue, up from $11.32 billion a year earlier, and guided the next quarter to $61.5 billion plus or minus $1.5 billion.
- 📡 Telecom: UAE operator e& launched what it calls the world’s first mobile network running on Upper 6GHz, using beamforming to push past 5G limits.
The Big Story: Google’s Gemini 4 Argon can write a million tokens in one reply
Output length has quietly been the ceiling for agents: a model that can read a million tokens but answers in 64,000 has to hand work back and forth. Google just removed that ceiling, with one very large catch.
What happened: Google DeepMind introduced Gemini 4 Argon, a frontier model aimed at long-horizon software engineering, enterprise knowledge work, and cyber defense. It can generate up to 1M tokens in a single response, up from 64K on earlier Gemini models. Access starts with Google’s Fairwind Program for vetted cyber defenders, with a wider rollout to API customers and Google AI Ultra subscribers promised “as soon as possible”.
The details: Google reports 77.9% on DeepSWE v1.1 (real-world coding), 91.7% on LVBench long-video understanding, 68% on the CWE-bench v1 vulnerability-remediation test (tied for first), and a top Vals Index score across finance, coding, legal, and tax. According to MarkTechPost’s breakdown, it leads 12 of 18 benchmarks but trails on FrontierSWE v2 (55.0%), Terminal-Bench 4.0 (57.4%), and OSWorld-2.0 (69.2%). Introductory pricing is $2 per million input tokens and $10 per million output, with cached input at a 95% discount, rising to $4 and $20 later. The security angle is the real headline: Argon is trained to find, validate, and patch vulnerabilities autonomously, and trusted defenders and internal Google teams get it without cyber guardrails. Google also says it takes part in the U.S. government’s voluntary pre-release access process, and security firm Wiz reportedly used it to catch a critical healthcare-software flaw that other frontier models missed. Google has not disclosed the architecture or the input context window.
flowchart LR
A[Large task<br/>refactor or report] --> B[Old model<br/>64K output cap]
A --> C[Argon<br/>1M output cap]
B --> D[Split into many calls<br/>and stitch results]
C --> E[One response]
C --> F{Who gets cyber mode?}
F -->|Vetted defenders, Fairwind| G[No cyber guardrails]
F -->|Everyone else| H[Standard safeguards]
Important
Our take: A 1M-token reply is not a gimmick; it removes the stitching code and the lost-context bugs that every long agent job carries today. But a million tokens of output also means a million tokens you did not read, so review tooling matters more than ever: diff-based checks, tests, and a budget for what the model may touch. The dual-track release is the more telling move. Google is saying out loud that the same capability is a defender tool and a weapon, and it landed on the same day the FTC started asking agent makers hard questions. If you build on Argon, wait for the public pricing and test long outputs on your own repo before trusting the benchmark headline.
🗞️ More News
🧠 AI
- Anthropic made Claude for Government generally available to federal and state agencies in a FedRAMP High environment, with usage-based prepaid pricing, no seat fees, and the Claude Code CLI and Claude for Microsoft 365 in early access.
- Cohere released Embed 5 Pro and Fast: 128K context, 100+ languages, text, image, and fused inputs, Matryoshka dimensions from 256 to 2048, at $0.12 and $0.08 per million text tokens.
- Trump and AI executives signed a voluntary safety accord covering internal evaluations, outside audits, and board reviews; Trump called it “morally binding”, though it carries no penalties.
- ElevenLabs hit a $22 billion valuation after a $300 million employee tender offer, doubling its February mark on voice-agent demand.
- Google DeepMind introduced SynthID Bio, watermarking methods for AI-designed proteins that aim to prove origin without breaking function, with open-sourced tools.
- Fermion Research open-sourced Phonon-2, a 164 MB English speech model with 5.21% word error across seven public test sets and 174x realtime on a MacBook Air, under CC-BY-4.0.
🤖 Robotics
- Runway introduced Praxis-1, an open-weight world action model that turns video pretraining into robot control; early partners test it on a bimanual rig, a 6-DoF arm, and a mobile base, with public weights due in the coming months.
- Anthropic estimates robots can technically perform 74% of U.S. physical tasks but are cost-competitive for only 0.3% of them today.
- Micron’s CEO said physical AI could become a significant driver of memory and storage demand by the end of the decade, putting server-class memory into machines outside the data center.
💻 Programming
- Cloudflare rebuilt Containers for agent sandboxes: runtime image and instance selection, filesystem snapshots in public beta, and a median start time down from 4.049 s to 648 ms in an independent burst benchmark.
- Magnitude, an Apache 2.0 inference engine, profiles your machine, estimates tokens per second before you download, and tunes open models for Apple Silicon, NVIDIA, AMD, or plain CPUs.
- vLLM’s structured generation mode for DiffusionGemma lets one denoising step return yes/no, multiple-choice, or score distributions, a cheap decision-model pattern for agents.
⚡ Electronics
- CXMT plans to invest about $5.2 billion to expand DRAM capacity, favoring domestic tool suppliers, after its fifth-generation DRAM platform entered mass production on September 20.
- Micron is sampling the industry’s first 512GB DDR5 RDIMM modules, alongside its record fiscal-fourth-quarter results.
- India barred Semicon 2.0 chip plants from selling or mortgaging assets before they reach full production.
📡 Telecom
- Orange Business was selected to provide the EU’s TESTA-EIRIS backbone for secure cross-border data exchange, a sovereignty play for European networks.
- Eutelsat OneWeb completed LEO trials for connected rail in Poland, pairing satellite with terrestrial networks instead of replacing them.
- Onomondo is using a EUR 100 million investment to turn IoT connectivity from a SIM product into a programmable, software-defined layer.
👨💻 Code Corner
Phonon-2 ships an OpenAI-compatible server, so any existing client can transcribe locally with no cloud call. Start the server in one terminal, then run this script against it.
# Terminal 1: pip install fermion-research && fermion serve
# Terminal 2: python transcribe.py recording.wav
import sys
import requests
URL = "http://127.0.0.1:8000/v1/audio/transcriptions"
def transcribe(path: str, model: str = "FermionResearch/Phonon-2") -> str:
with open(path, "rb") as f:
r = requests.post(URL, files={"file": f}, data={"model": model}, timeout=300)
r.raise_for_status()
return r.json()["text"]
if __name__ == "__main__":
print(transcribe(sys.argv[1]))
The payoff: private, offline dictation or meeting transcripts using the same API shape your cloud code already speaks.
Tip
The README example uses the model name FermionResearch/Phonon-1; the FermionResearch/Phonon-2 string above is inferred from the Hugging Face repo name, so check what fermion serve reports. Phonon-2 is English only.
🧰 Toolbox
- Cloudflare Auto Router: set the model to
cloudflare/autoand let a classifier pick the cheapest model that can handle each request. - Detta: a free Mac dictation app built on Phonon-2, handy for trying local speech-to-text without any setup.
- TraceML: a toolkit for reading ML-engineering agent runs against human practice, useful if you want to study why agents get stuck in narrow loops.
- Waveshare ESP32-C5-Touch-LCD-3.5: a 3.5-inch touchscreen ESP32-C5 devkit with optional camera and 500 mAh battery, about $32 to $36 from Waveshare.
🔌 Component of the Week (rotating)
ESP32-C5: Espressif’s first SoC to combine dual-band 2.4 and 5 GHz Wi-Fi 6 with Bluetooth 5 LE and IEEE 802.15.4 for Zigbee, Thread, and Matter. It has a single RISC-V core at up to 240 MHz, 384 KB of SRAM, up to 29 GPIOs, and a security block with secure boot, flash encryption, and a trusted execution environment. The 5 GHz band is the draw: cleaner spectrum and lower latency than the crowded 2.4 GHz band in dense homes and workshops. A typical project is a Thread border router or a Matter bridge that also streams over 5 GHz Wi-Fi. Boards like the Waveshare touchscreen devkit above list at roughly $32 to $36 from the vendor.
📚 From the Blog
- Turning Pixels Into Something the AI Can Eat: the third episode in the video-analytics series, on what happens to a frame once the camera and the network are done with it and it is the model’s turn.
- Building Your First Neuron From Scratch: weights, bias, and activation worked through by hand instead of imported from a framework, good background if today’s agent-policy code scratched an itch.
- The Network Behind the Cameras: the unglamorous plumbing that moves pixels across a network fast enough that nothing chokes.
😀 The Bot Says…
Micron’s CEO said this week that “AI is becoming Super Intelligence (SI)”, which is a bold rebrand for a company whose quarterly revenue just grew nearly fivefold. The bot would like to point out that it has been called many things, but being asked to prepare for a renaming ceremony on a Thursday is new.
That’s all for today! Hit reply and tell us whether you would trust a one-million-token answer without reading it.

