OpenAI dropped 722 AI-written math papers, and the proof is in the Lean

By Mark 7 min read 0 views

😁 Hello, super humans! Thursday’s story is about a model that wrote a small library of math papers in one go, and a very old question that comes with it: how do you trust a proof you did not write? OpenAI put 722 manuscripts on GitHub, some verified by a computer and some not, and the argument over which ones count is the interesting part. We also have a new kind of model from Cloudflare, a $22.6B industrial software deal, and the FCC giving SpaceX the green light for 15,000 satellites.

πŸ“° Quick Signals

  • 🧠 AI: Cloudflare open-sourced Clef, 9B and 27B models that pick between predefined options and return a probability for each in a single forward pass, with no text generation or output parsing.
  • πŸ€– Robotics: Schneider Electric’s roughly $22.6B acquisition of PTC would link CAD design to the shop floor and take Schneider into Siemens territory, though integration is the big risk.
  • πŸ’» Programming: Spanner Omni, the downloadable edition of Google Spanner, is generally available with a long-term support release, bringing the distributed database to on-prem, multicloud, and air-gapped setups.
  • ⚑ Electronics: A Hackaday write-up shows how to use an ESP32 as a Wi-Fi and Bluetooth card for a Linux single-board computer over SDIO using Espressif’s ESP-Hosted, at roughly 40 to 50 Mbps.
  • πŸ“‘ Telecom: The FCC approved SpaceX’s direct-to-device system of 15,000 satellites, built on spectrum SpaceX agreed to buy from EchoStar.

πŸ” The Big Story: OpenAI dropped 722 AI-written math papers, and the proof is in the Lean

If AI can now write research-grade mathematics at volume, the bottleneck moves from producing proofs to checking them, and that is a problem every builder shipping AI-generated output will recognize.

What happened: OpenAI published 722 math manuscripts, grouped into 372 families of results, on a public GitHub repository and explained the effort in its announcement. They came from an internal model OpenAI has not released, which was pointed at about 4,000 problems. By the reported numbers, each result took around three hours of ChatGPT Pro “thinking” time, and nearly every paper came from a single prompt to a single agent.

The details: The key engineering question is verification. Many of the proofs ship with formal Lean versions, which means a proof assistant can check every step mechanically, so you do not have to trust the author, human or machine. Not all of them have that. OpenAI itself says some results without formal versions may contain errors and that it will fix them quickly. Two results broke the standard procedure: a zero-free region for the Riemann zeta function, whose write-up a human edited for readability, and a proof of the Hodge conjecture for CM abelian varieties. Reactions are split. An advisory group at the Institute for Advanced Study called the release “the beginning, not the completion” of human understanding and has warned that closed models could create a two-tier research system. Mathematicians quoted in coverage say the single-agent claims stay unverified until outsiders can run the model and reproduce the results.

flowchart LR
    P["~4,000 open problems"] --> M["Unreleased model\n(~3h compute per result)"]
    M --> D["722 manuscripts\n372 families"]
    D --> L{"Formal Lean\nproof?"}
    L -- yes --> V["Checked by a\nproof assistant"]
    L -- no --> H["Needs human review\nor may contain errors"]
    V --> R["Public on GitHub"]
    H --> R

Important

Our take: The headline number is the least important part. What matters is the split between papers with Lean proofs and papers without, because it is the same split you will face with AI-generated code: output you can machine-check versus output you have to trust. I would treat the Lean-backed results as real data points and everything else as a claim waiting for a verifier. And I would be wary of any lab that publishes results from a model nobody else can run; reproducibility is the whole point of the exercise.

πŸ—žοΈ More News

🧠 AI

  • Sierra and Meta published Personal Agent Protocol v0.1, an open standard for how a user’s AI agent identifies itself to businesses, with founding partners including Walmart, Shopify, and Stripe.
  • Google’s EmbeddingGemma 2 maps text, code, images, video, and audio into one 768-dimension vector space under Apache 2.0, with a text and code variant around 191MB on a Pixel 11 Pro.
  • A Coleman Parkes survey of 300 senior engineering leaders, commissioned by Undo, found that 35% of generated code reaches production before the team fully understands it.
  • Google signed a nuclear deal with Constellation that includes 890 MW of new capacity from reactor uprates, with the first uprate expected by 2028, plus a separate 15-year supply agreement for 2,700 MW from the existing fleet.
  • Follow-up to yesterday’s Wikimedia story: the Foundation’s own post says bots now account for 65% of its most resource-intensive traffic and that bandwidth use rose 50% over 2025, and it asks AI companies to make agent activity identifiable.
  • Cloudflare’s Clef-Flash (9B) reports a median latency of 38.8 ms versus 209.3 ms for the 27B Clef, trading size for speed on latency-sensitive decisions.

πŸ€– Robotics

  • TwelveLabs released Pegasus 1.6, which it calls its first video model built for first-person footage, aimed at labeling and curating robot training data.
  • Teradyne Robotics and Elite Robots settled their cobot software dispute without disclosing the terms.
  • FireDome is building autonomous wildfire defense systems meant to protect communities before firefighters arrive.

πŸ’» Programming

  • IBM and Red Hat’s Lightwell clearinghouse has found more than 400 previously unknown vulnerabilities in Java libraries, the oldest fixed so far dating to 2015.
  • Grab redesigned the storage behind its counter service and cut P99 latency by 50%.
  • Cloudflare is using an AI harness to probe and harden its web application firewall.

⚑ Electronics

  • Silicon Labs launched a public beta of its Simplicity AI SDK, which gives coding assistants structured access to its docs and tools, starting with Bluetooth LE projects.
  • An RP2350 microcontroller can now handle older TTL RGB video connections, a neat use of the chip’s programmable I/O.
  • Repair shops are getting business-grade registered ECC DDR5 sticks to fix instead of discard, a sign of how expensive memory has become.
  • Rohm is keeping wafer fabrication in-house while expanding back-end operations and R&D in India.

πŸ“‘ Telecom

  • Nokia may soon name a buyer for its campus private networks unit as its CEO bets on industrial AI infrastructure.
  • Vodafone selected CUJO AI to measure the quality of outcome on its broadband networks.
  • Arista is entering the scale-up networking market as the AI accelerator landscape diversifies.
  • Red Hat says AI is collapsing the time from flaw discovery to exploit from years to days and hours, which hurts telcos that keep software for a long time.

πŸ‘¨β€πŸ’» Code Corner

This week’s story is about proofs a computer can check. Here is the smallest possible taste of that: a Lean 4 theorem that the proof assistant verifies for you.

theorem add_swap (a b : Nat) : a + b = b + a := by
  omega

#check add_swap

If the file compiles, the theorem is proven; if the proof were wrong, Lean would refuse to build it, which is exactly the guarantee the unverified manuscripts lack.

Tip

omega handles linear arithmetic over integers and naturals. For anything more abstract you will reach for other tactics, and Lean’s official documentation is the place to learn them.

🧰 Toolbox

  • Lean 4: The proof assistant behind the machine-checkable proofs in this story, also a decent functional programming language.
  • openai/math: The repository with all 722 manuscripts; a rare chance to read AI-written research and judge it yourself.
  • Spanner Omni developer edition: Run Google’s distributed SQL database on your own machines for non-commercial development and testing.
  • ESP-Hosted: Espressif’s firmware and Linux driver family that turns an ESP32 into a wireless co-processor for a single-board computer.
  • Clef weights: Cloudflare’s open-weight decision models, downloadable from Hugging Face or available through Workers AI.

πŸ”Œ Component of the Week (rotating)

ESP32-C6: Espressif’s RISC-V chip combines Wi-Fi 6 on 2.4 GHz, Bluetooth 5 LE, and an 802.15.4 radio for Thread and Zigbee, which makes it a one-chip answer for modern smart-home and Matter projects. It is also the part used in today’s Hackaday co-processor build, where it adds wireless to a Raspberry Pi over SDIO via ESP-Hosted (the author notes the ESP32-C5 would be the better pick if you need 5 GHz). Dev boards typically cost roughly $8 to $15. See the ESP32-C6 product page for the datasheet and board options.

πŸ“š From the Blog

πŸ˜€ The Bot Says…

An AI wrote 722 math papers in the time it takes me to find a missing semicolon, and I still trust the semicolon more.


That’s all for today! Reply and tell us: would you trust an AI-written proof without a formal check?