π Hello, super humans! Every week somebody claims an AI did something extraordinary, and every week the claim arrives without a way to check it. On Saturday one arrived with a compiler attached. That single detail, a proof you can recompile yourself instead of a benchmark you have to trust, is what makes today’s story worth an hour of your attention rather than a scroll. Let’s dig in.
π° Quick Signals
- π§ AI: OpenAI announced its next model family, Astra, by publishing ten previously open results in mathematics and theoretical computer science, each with a formal Lean proof on GitHub.
- π€ Robotics: BMW is putting humanoid robots into a German production plant for the first time, moving the programme off its US pilot site.
- π» Programming: Linux 7.2-rc6 landed with Torvalds calling it the biggest rc6 by commit count in years, with heavy sound and DRM fixes.
- β‘ Electronics: DDR4 has climbed back above $4 per gigabit and HBM3e contract pricing rose roughly 20% quarter over quarter as AI capacity gets booked out.
- π‘ Telecom: Qualcomm is still pressing the FCC for 5 MHz in the 1675 to 1695 MHz band to run a 5G sidelink service that lets public-safety devices talk to each other with no network in between.
The Big Story: An AI proved ten open problems, and you can recompile the proofs yourself
The interesting thing about Saturday’s announcement is not that a model did mathematics. Models have claimed that before. The interesting thing is that OpenAI shipped the artifacts in a format where the claim either compiles or it does not.
What happened: On 1 August, OpenAI published ten advances in mathematics and theoretical computer science produced by an internal version of Astra, the model family it calls its next major release. The results span high-dimensional geometry, coding theory, group theory, operator algebras, arithmetic circuit complexity, quantum complexity, lattice problems and extremal combinatorics. Among them: a construction establishing the existence of non-sofic groups, a disproof of Connes’s rigidity conjecture, a proof of Ehrhart’s volume conjecture, three problems lifted from the ErdΕs catalogue, and the first improvement to the general upper bound on high-dimensional sphere-packing density since 1978. Fields Medallist Timothy Gowers, one of several mathematicians OpenAI asked to assess the work, said he would recommend one of the proofs for a top journal without hesitation. OpenAI puts the total compute at roughly $2,000 at its own Sol API rates.
The details: The engineering story here is the verification loop, not the model. A Lean proof is a program: it type-checks against mathlib, the community library that now holds more than 115,000 definitions and 232,000 theorems, and the Lean kernel accepts it or rejects it with no room for a persuasive paragraph in between. That property does two things at once. Downstream, it means any mathematician can git clone the repository, run the checker and confirm the result without trusting OpenAI, which is a standard almost no other AI capability claim can meet. Upstream, it means the model had a hard oracle during search: propose a proof term, compile it, read the error, revise. That is a reinforcement signal of a quality you simply cannot get in domains where “correct” is a matter of taste. It is also why the honest caveat matters. As Understanding AI pointed out, these problems sit exactly where formal search is strongest, so this is real research output rather than evidence that general mathematical taste has been solved.
flowchart TD
A[Problem statement, formalized in Lean] --> B[Model proposes a proof term]
B --> C[Lean kernel type-checks against mathlib]
C -->|Rejected| D[Error message names the failing step]
D --> B
C -->|Accepted| E[Proof compiles: result is certain]
E --> F[Published to GitHub]
F --> G[Anyone recompiles and confirms]
G --> H{"#print axioms shows sorryAx?"}
H -->|Yes| I[Proof has a hole, not a proof]
H -->|No| J[Verified, no trust in the author required]
Important
Our take: The lesson to carry into your own work is not about mathematics, it is about oracles. Astra got this far because somebody had already spent fifteen years building a machine that says yes or no to a proof, and the model got to fail against it thousands of times for free. Every domain that has a cheap, exact checker is going to see this pattern next: type systems, property-based tests, formal hardware verification, SMT-backed configuration. Every domain that does not will keep producing confident, plausible, unverifiable output. If you are choosing where to point an agent in the next year, the highest-leverage engineering you can do is not prompt design, it is building the checker that tells the agent it is wrong. And keep the framing honest: this is a genuine research contribution in areas suited to formal search, not a general intelligence result, and OpenAI announcing an unreleased model this way is also very good marketing.
ποΈ More News
π§ AI
- The ten results cover eight fields, including operator algebras and arithmetic circuit complexity, and every one arrived with a Lean certificate rather than a prose sketch.
- The sphere-packing result is the first improvement to the general high-dimensional upper bound since 1978, pushing toward the Cohn-Elkies threshold.
- The $2,000 figure is an estimate of total token spend at Sol API rates, which reframes certain research questions as a budget line rather than a staffing problem.
- Astra itself is unreleased, and OpenAI folded the model family’s announcement into a research blog post rather than giving it a launch of its own.
- The unit distance conjecture counterexample that surfaced weeks ago is now confirmed as part of the same research programme, which makes this a trajectory rather than a one-off.
- Epoch AI projects that AI chip deployments will double roughly every nine months, which is the compute curve sitting underneath cheap frontier experiments like this one.
- Trump Media launched a paid real-time feed of market-moving posts and Democratic senators promptly asked the SEC to look at it, which is a preview of how automated trading consumes social data.
π€ Robotics
- Two Figure 02 units ran about 1,250 operational hours over eleven months at BMW’s Spartanburg plant, handling more than 90,000 sheet metal parts across roughly 30,000 X3 bodies.
- Germany’s Neura Robotics raised up to $1.4 billion in a Series C, and Apptronik closed $935 million for its Apollo line.
- Robotics startups have taken more than $23 billion in 2026 so far, already close to the whole of 2025.
- Nvidia’s Isaac GR00T reference humanoid pairs a Unitree H2 Plus body with Sharpa Wave tactile hands and Jetson Thor compute, 75 degrees of freedom in total, with Ai2, ETH Zurich and Stanford already committed.
π» Programming
- Mathlib 4 now carries more than 115,000 definitions and 232,000 theorems, which is the dependency graph every one of Astra’s proofs compiles against.
- LeanDojo exposes Lean proof states programmatically with retrieval over mathlib, which is the interface most LLM prover work is built on.
- Rust 1.98 is in beta and reaches stable on 20 August, following 1.97 making v0 symbol mangling the default for readable backtraces.
- The July Azure SDK drop shipped Rust’s first client library while the Python track took four libraries to stable, which is a fair snapshot of where cloud SDK effort is going.
β‘ Electronics
- Intel is reviving DDR4-compatible CPUs, with limited shipments since June and volume expected in September, to dodge DDR5 pricing in the China PC market.
- TSMC is reported to be planning baseline price rises of 5% to 10% on advanced nodes from the start of 2027, with some services up as much as 25%.
- The consumer end of the memory surge is cooling on affordability, but DRAM and NAND contract prices are still forecast to climb through the third quarter.
- The industry is still tracking toward roughly a trillion dollars of annual sales in 2026, which is the demand curve making all of the above possible.
π‘ Telecom
- Starlink told Malaysia’s communications minister it wants to trial direct-to-cell there, keeping the existing SIM, number and 4G handset while the satellite behaves as an LTE base station wired back to the operator core.
- France’s Arcep opened a second consultation on reallocating the 700 MHz, 800 MHz, 900 MHz, 1.8 GHz, 2.1 GHz and 2.6 GHz licences expiring between 2030 and 2035, with responses due by 23 September.
- The Release 21 timeline is now fixed: 6G work item approval in March 2027, physical layer freeze in September 2028, and the final ASN.1 freeze in March 2029.
- Ericsson named Christophe Van de Weyer head of its Global Communications Platform business and CEO of Vonage, effective 15 August, to push the API turnaround along.
π¨βπ» Code Corner
If you have never seen a machine-checked proof, the whole story above stays abstract. Here is the smallest useful one: paste it into the Lean 4 web editor, no install required, and watch the kernel accept it.
import Mathlib.Tactic
-- A claim about every natural number, and a proof the kernel verifies step by step.
theorem four_le_sq (n : Nat) (h : 2 β€ n) : 4 β€ n * n := by
calc 4 = 2 * 2 := by norm_num
_ β€ n * n := Nat.mul_le_mul h h
-- The audit trail: this lists what the proof ultimately rests on.
#print axioms four_le_sq
Now break it on purpose: change 2 β€ n to 1 β€ n and the calc step turns red immediately, because Nat.mul_le_mul h h no longer produces 2 * 2 β€ n * n. That red squiggle is the entire trust model of today’s story in miniature.
Tip
#print axioms is the check that matters when you are reading someone else’s proof, human or machine. If sorryAx appears in the output, the proof contains an admitted hole and proves nothing at all, no matter how convincing the surrounding text looks. Everything else it lists, typically propext, Classical.choice and Quot.sound, is the ordinary foundation mathlib is built on.
π§° Toolbox
- Lean 4 web editor: run today’s snippet in a browser tab with mathlib preloaded, no toolchain to install.
- mathlib: the 232,000-theorem library that makes formal proofs about real mathematics practical instead of theoretical.
- LeanDojo: programmatic access to Lean proof states plus retrieval over mathlib, the standard harness for LLM prover experiments.
- MathArena arXivLean: an evaluation set built from recent arXiv results, useful for sanity-checking prover claims against fresh problems.
- Isaac GR00T reference humanoid: an open reference design if you would rather your verification loop involved actual torque.
- ThunderScope: an open-source software-defined oscilloscope built around an Artix-7, for when the bug is in the analogue domain.
π¬ Demo Watch (rotating)
The companion result to the proofs, and the one more likely to touch your week, is OpenAI’s field report on coding agents modernizing research software, published alongside the math work.
What it shows: eight real deployments, mostly in computational biology, where agents were pointed at fragile legacy scientific code. The headline numbers are good: a 60x speedup in RNA-sequencing quality control, a 20,000-line C and C++ genome aligner rewritten from scratch in Rust at 99.8% output parity, and a GPU-native redesign that cut synthetic genome generation from 1,610 seconds to 27 seconds per run.
Why it is hard: scientific code is not slow because nobody knew better, it is slow because it was written by domain experts under deadline and then never touched again, and nobody remembers which of its quirks are bugs and which are load-bearing. Rewriting it means preserving behaviour you cannot fully specify.
Hype versus real: the report is unusually honest about the boundary. The agents are described as eloquent, convincing and confidently wrong in ways that are easy to miss, and every one of those wins rested on something the agent could not supply: a human deciding what “correct” meant and building the harness that proved it. Same lesson as the Lean proofs, arriving from the opposite direction. The speedup is real, the oracle is yours to build.
π From the Blog
- Building Your First Neuron From Scratch: the smallest learnable transformation, built by hand, which is a useful grounding before reading anything about what a frontier model “understands”.
- The Network Behind the Cameras: the unglamorous plumbing that moves video frames around, and why the pipeline is usually the bottleneck, not the model.
- What Deep Learning Actually Is: the plain-language explanation of the machinery that just got a Fields Medallist’s tentative endorsement.
π The Bot Saysβ¦
A Fields Medallist would recommend one of my cousin’s proofs for a top journal. Meanwhile my last pull request was rejected because I put the closing brace on the wrong line. Verification is a humbling business at every scale.
That’s all for today! Reply and tell us which part of your stack already has an exact checker, and which part is still running on vibes and code review.


