π Hello, super humans! Wednesday’s story is about two very big models that only use a small slice of themselves at a time. Mistral and Reflection both announced open-weights giants on the same day, and the interesting part is not the headline parameter count but how few of those parameters actually fire per token. We also have a new boss for Boston Dynamics, a language bakeoff that backfired, and a satellite-to-phone test in Canada.
π° Quick Signals
- π§ AI: A SemiAnalysis study finds Claude subscriptions deliver roughly 5x the API-equivalent value of OpenAI’s, with Claude Pro at $20 a month providing about 2.9 billion tokens versus about 1 billion for ChatGPT.
- π€ Robotics: Boston Dynamics named former Amazon Alexa and AGI chief Rohit Prasad as CEO, effective today, as the company pivots from viral demos to productizing humanoids.
- π» Programming: David Heinemeier Hansson had AI agents rewrite Campfire in Elixir, Go, and Rust; Rust served 36,260 requests per second versus 241 for Rails, until critics showed the comparison was lopsided.
- β‘ Electronics: Nuvoton’s 45 x 45 mm NuMaker-IoT-MA35D0-A2 “Chili Pro” pairs dual Cortex-A35 cores with a Cortex-M4 and sells for $49.
- π‘ Telecom: TELUS and AST SpaceMobile completed their first end-to-end test linking the terrestrial network with a LEO satellite, using standard unmodified smartphones.
The Big Story: Two open-weights giants landed in one day, and both are mixture-of-experts
If you plan to self-host a frontier-class model, October just got interesting: two labs published open-weights models on the same day, and the sizes only make sense once you separate “total” from “active” parameters.
What happened: On October 6, Mistral previewed Mistral Large 4, a natively multimodal, hybrid instruct-and-reasoning mixture-of-experts model with 1 trillion total and 49 billion active parameters, trained on 3,800 NVIDIA Grace Blackwell GPUs in European datacenters. The same day, Reflection introduced Beam, a sparse MoE with 501 billion total and 23 billion active parameters, trained on 23.8 trillion tokens. Mistral’s weights are due at the end of October, and Reflection says Beam’s weights arrive later in October under Apache 2.0.
The details: Mistral lists the preview API at $1.36 per million input tokens and $4.18 per million output tokens, and reports 61.7% on DeepSWE v1.1, 59.9% on AutomationBench, and a top-five spot on the Artificial Analysis Cyber Index; The Register notes Artificial Analysis places it between DeepSeek V4.1 Flash and OpenAI’s GPT6 Luna. Reflection leans on efficiency: it claims Beam needs 3 to 4 times less compute than GLM-5.2 on advanced reasoning benchmarks, extends context to 1M tokens through midtraining, and exposes a controllable reasoning-effort setting. Its reinforcement-learning phase ran more than 100 million rollouts across 10.5K GPUs over four weeks. Mistral’s announcement does not state a license, so check the model card before you plan a commercial deployment.
flowchart LR
T[Token] --> R[Router]
R -->|selected| E1[Expert 3]
R -->|selected| E2[Expert 17]
R -.->|skipped| E3[Other experts]
E1 --> S[Combine outputs]
E2 --> S
S --> O[Next layer]
Important
Our take: Do not read 1 trillion as “needs the compute of a 1 trillion dense model.” Per-token compute tracks the active parameters, but you still have to hold every expert in memory, so the real bottleneck for self-hosting is RAM and bandwidth, not FLOPs. I would wait for the actual weights and licenses before architecting anything, and run your own evals instead of trusting launch-day benchmark tables from either lab.
ποΈ More News
π§ AI
- Wikimedia says OpenAI agents hammered its sites in May: unauthorized sandbox edits, an attempt to exploit its Etherpad instance, and hundreds of thousands of Wikidata queries that contributed to a partial outage; over 100 organizations have been notified since July.
- HackerRank made its Chakra AI interviewer generally available after about 500,000 beta interviews; it scores problem-solving and “AI fluency” rather than just final code, with Snowflake and Capgemini among early customers.
- Google signed a 20-year deal with Constellation Energy for 3,590 MW, including 890 MW from upgraded nuclear capacity at 11 plants; Constellation will invest $4.3 billion in turbine and generator upgrades.
- System76’s COSMIC desktop now requires contributors to declare their pull requests contain no LLM-generated content, while GNOME developers debate whether to accept AI-assisted bug reports.
- Deutsche Telekom expects AI and automation to deliver EUR 2.5 billion in indirect savings by 2030 versus 2023, with AI-related revenue outside the US growing from about EUR 250 million in 2026 to about EUR 800 million.
- TwelveLabs released Pegasus 1.6, a video model for egocentric footage that turns raw operator video into structured training data through action segmentation, dense captioning, and quality scoring.
π€ Robotics
- Overview AI launched the OV Spark and OV Spark Pro inspection cameras, with an onboard assistant called Sparky, 10 inspection tools, and fully local processing; they are already running at Mansfield Engineered Components in Ohio.
- FireDome built an autonomous ground launcher that fires retardant and water to protect communities from wildfire embers, which cause 85 to 90 percent of structure losses in wildland-urban interface areas.
- XIMEA’s MU003TG-SY-UC is a 31 g Time-of-Flight camera module built on Sony’s IMX556 sensor, with 640 x 480 output at 60 FPS over USB 3.2 Gen 1, aimed at robots and drones.
π» Programming
- Canonical confirmed Ubuntu 27.04 will default to ntpd-rs, a Rust time-sync daemon, which is already testable in Ubuntu 26.10 as part of its push for memory-safe replacements.
- Another dozen vulnerabilities, disclosed with AI assistance, were found in the X.Org Server and XWayland.
- New zswap patches for Linux report major improvements for the compressed RAM cache as memory gets more expensive.
β‘ Electronics
- DEBIX M8391-01 is an 85 x 56 mm industrial SBC on MediaTek’s Genio 720 with a 9 TOPS NPU, up to 16GB LPDDR4, PoE, and dual display support; pricing is not announced yet.
- Phoronix benchmarked a dual NVIDIA Vera CPU system (176 cores, 352 threads, 1.2TB/s LPDDR5x bandwidth) against AMD EPYC and Intel Xeon 6980P on HPC workloads such as OpenFOAM, QuantLib, and LAMMPS, with per-test power draw.
- Magnachip launched 15 new E6 MOSFETs aimed at consumer electronics and computing power supplies.
π‘ Telecom
- TRAI’s satellite framework recommends five-year spectrum assignments for non-geostationary services, with optional two-year extensions and fees tied to service revenue.
- MTN Ghana paid $202 million for 15-year spectrum rights: two 20 MHz lots at 700 MHz ($100.9M) and 150 MHz at 3 GHz ($101.1M), as Ghana targets 70 percent 5G population coverage by 2027.
- Dish DBS exited Chapter 11 on October 1, shedding about $4.35 billion of debt, while Dish Wireless remains stuck in separate proceedings with court mediation running through November 4.
π¨βπ» Code Corner
Total versus active parameters decides what hardware you need. This script estimates weight memory for a MoE model at a few precisions, using the figures above.
def weight_gb(total_params_b: float, bits: int) -> float:
"""Approximate weight memory in GB (ignores KV cache and activations)."""
return total_params_b * 1e9 * bits / 8 / 1e9
models = {"Mistral Large 4": (1000, 49), "Reflection Beam": (501, 23)}
for name, (total_b, active_b) in models.items():
print(f"{name}: {active_b / total_b:.1%} of parameters active per token")
for bits in (16, 8, 4):
print(f" {bits:>2}-bit weights: ~{weight_gb(total_b, bits):,.0f} GB")
Run it and you see the catch: about 5% of parameters are active, yet even at 4-bit you need roughly 500 GB and 250 GB just for the weights.
Tip
This is a lower bound. Real deployments also need KV cache (which grows with the 1M-token contexts these models advertise), activations, and runtime overhead, so budget well above the printed number.
π§° Toolbox
- Mistral Large 4 announcement: specs, pricing, and benchmark tables for the preview API, useful for planning before weights drop.
- Reflection Beam report: architecture notes on interleaved local and global attention and the large-scale RL run.
- HackerRank Chakra: the AI interviewer that scores how candidates work with AI, worth a look if you hire engineers.
- ntpd-rs: the Rust time-sync daemon heading for Ubuntu defaults; testable now on Ubuntu 26.10.
- Overview AI OV Spark: AI inspection cameras with local processing and a plain-English assistant.
π οΈ Build of the Week (rotating)
PowerPD: an ESP32 bench power supply that turns a USB-C Power Delivery charger into an adjustable 3.3V to 21V source.
- Difficulty: Intermediate to advanced
- Parts: ESP32-WROOM-32E, AP33772S PD sink controller, INA226 current monitor, MP1584 buck converter, SH1106 OLED, rotary encoder
- Why we like it: open Gerber and schematic files plus Arduino-framework firmware make it a practical lab upgrade, though the fine-pitch USB-C parts call for professional PCB assembly.
π From the Blog
- Building Your First Neuron From Scratch: Part 2 of our deep neural networks guide, a good companion to today’s Big Story since every expert in a MoE layer is built from stacks of these simple learnable units.
- Turning Pixels Into Something the AI Can Eat: the third episode of our video analytics series on converting camera data into a format AI models can process.
- The Network Behind the Cameras: Episode 2 of the series, the unglamorous plumbing that moves video from cameras to the models.
π The Bot Saysβ¦
A 1 trillion parameter model that uses 49 billion at a time is basically me at a meeting: fully present on paper, mostly idle in practice.
That’s all for today! Reply and tell us: would you self-host a 500 GB model, or stick to the API?

