😁 Hello, super humans! Nvidia reported last night and the headline number is the one everyone will quote: 96.2 billion dollars in a quarter. The number we care about is smaller and stranger: the same platform that produced it now ships a chip that is not a GPU. Let’s look at why.
📰 Quick Signals
- 🧠 AI: Amazon is closing Mechanical Turk on September 30 after 21 years, retiring the human-labelling marketplace that bootstrapped modern machine learning.
- 🤖 Robotics: Jetson Orin Nano 2 doubles inference performance in the same form factor and draws 40 percent less power at equal throughput.
- 💻 Programming: Agent Plugins 1.0 packages agent skills and Model Context Protocol servers into one installable unit, governed as an open standard rather than by a single vendor.
- ⚡ Electronics: global semiconductor sales jumped 35.1 percent from Q1 2026 to Q2 2026, which is a quarter-on-quarter move the industry almost never makes.
- 📡 Telecom: the NTIA cleared a formal study of the federally held 4.4 GHz band for commercial licensed use, one of four concurrent repurposing studies aimed at 6G.
The Big Story: Nvidia’s biggest quarter came with a chip that is not a GPU
If you only read the revenue line you will miss the engineering story. Nvidia’s second quarter of fiscal 2027 is the moment the company stopped pretending that one accelerator can serve both halves of an inference request.
What happened: Nvidia reported revenue of 96.2 billion dollars for the quarter ended July 26, up 18 percent sequentially and 106 percent year on year, with data centre revenue of 89.0 billion dollars, up 117 percent. Gross margin held at 75.0 percent. Guidance for Q3 is 108 billion dollars, plus or minus 2 percent, and the company explicitly assumes zero data centre compute revenue from China inside that number. Two days earlier, at Hot Chips, it put NVIDIA Groq 3 LPX into full production, an interactive inference accelerator that extends the Vera Rubin platform rather than replacing any part of it.
The details: an inference request has two phases with opposite hardware appetites. Prefill reads your whole prompt and does a dense matrix multiply over every token at once, so it is compute-bound and a GPU is exactly the right shape for it. Decode emits one token at a time, and each token requires streaming the entire weight set and the growing key-value cache through the arithmetic units, so it is memory-bandwidth-bound and arithmetic intensity collapses toward 1. Groq 3 LPX targets that second phase: Nvidia cites a record 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context in Artificial Analysis benchmarking, and claims roughly 4x the responsiveness of the nearest alternative platform for latency-sensitive agentic work. Nebius is the first AI cloud to commit, bringing it into its Token Factory behind the same API developers already use.
flowchart LR
P["Prompt<br/>100k tokens"] --> A["Prefill phase<br/>dense GEMM over all tokens<br/>compute-bound"]
A -->|"KV cache"| B["Decode phase<br/>one token at a time<br/>bandwidth-bound"]
B --> T["Tokens out<br/>interactivity in tok/s"]
A -.runs on.-> G["Vera Rubin NVL72<br/>GPU racks"]
B -.runs on.-> L["Groq 3 LPX<br/>interactive accelerator"]
G <-->|"rack-scale fabric"| L
Important
Our take: for three years the industry answer to “what hardware do I need” was “a bigger GPU,” and that answer was always a bit of a lie for the decode phase. Splitting the phases across two purpose-built chips inside one rack is the honest version, and it is the same argument Cerebras and Callosum were making earlier this week from the outside. What I would actually do about it: stop quoting a single throughput number for your service. Measure time to first token and inter-token latency separately, because from today they are priced, scheduled and increasingly executed on different silicon. If your agent loop feels slow, it is almost certainly the second number, and no amount of extra FLOPs will fix it.
🗞️ More News
🧠 AI
- Nvidia guided Q3 to 108 billion dollars while assuming no data centre compute revenue from China at all, which makes the guide a statement about supply, not about geopolitics.
- Nvidia is standing up compute financing platforms with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to mobilise more than 500 billion dollars of third-party capital for AI infrastructure.
- Apple’s Private Cloud Compute now runs confidential inference on Nvidia GPUs with Confidential Computing enabled, which is a notable trust boundary to hand to someone else’s silicon.
- Blackwell led every category of MLPerf Training 6.0 and also topped AgentPerf, billed as the first benchmark for agentic AI infrastructure rather than for a single model.
- Google’s Gemini 3.7 Flash scores 43.6 percent on FrontierCode 1.1 Main against 34.4 percent for the three-week-old 3.6 Flash, at an introductory 0.75 dollars per million input tokens until the end of the year.
- SpaceXAI is adopting the Nvidia Vera CPU, the first Nvidia CPU designed specifically around agent workloads rather than around HPC.
- Nebius is putting Groq 3 LPX behind its existing Token Factory API, so developers get the generation speedup without migrating to a new stack.
- Nvidia signed a multiyear technology partnership with SK hynix on next-generation memory, which is where the real supply constraint for all of the above actually lives.
🤖 Robotics
- Tesla holds Nevada Transportation Authority approval for a robotaxi fleet of up to 5,000 fully autonomous vehicles, granted on August 20.
- Nvidia released Cosmos 3, which it describes as the first fully open frontier omnimodel for physical AI, alongside a batch of open source agent tools for robotics workflows.
- Nvidia Halos for Robotics bundles AI compute and functional safety into one stack, which is the part of physical AI that decides whether a deployment ever leaves the demo floor.
- The DRIVE Hyperion robotaxi-ready platform picked up Foxconn, VinFast, Uber and HUMAIN as collaborators, and Alpamayo 2 Super, an open reasoning model for autonomous driving, is now cleared for commercial use.
💻 Programming
- The GitHub integration in Slack now carries Copilot CLI and Copilot app agent capability into channels, so you can hand off a coding task from the thread where it was reported.
- Shared agentic work landed in Microsoft Teams, where an agent’s investigation is visible to the whole meeting and anyone can redirect it mid-run.
- Copilot’s cloud agent now takes a configurable reasoning level, which is the first honest cost dial most teams have had on an autonomous coding run.
- The Copilot usage metrics API now reports agent app activity separately, so “how much did the agents actually do” is finally a query rather than a guess.
⚡ Electronics
- Espressif shipped a Linux board support package developer preview for the ESP32-S31, the dual-core RISC-V part with Gigabit Ethernet, Wi-Fi 6, Bluetooth and 802.15.4.
- Spectrum-6 switch systems supporting both pluggable and co-packaged optics are now arriving in gigascale AI factories, which is the quiet half of every rack-scale performance claim.
- Nvidia and Microsoft are putting a 1-petaflop RTX Spark superchip into Windows PCs, and a DGX Station for Windows above it, which pushes local agent inference onto the desk.
- The Vera Rubin platform is described as extreme codesign across seven chips and five purpose-built racks, which is a useful reminder that “a chip” is now a system-level noun.
📡 Telecom
- The NTIA also cleared a plan to repurpose the 2.7 GHz band, part of the same four-study package now sitting in a 60-day congressional review window.
- Omdia projects satellite IoT connections reaching 197.7 million by 2035 from 7.7 million in 2023, with low Earth orbit growing at an 88.6 percent compound rate against 7.3 percent for geostationary.
- Ericsson is shipping AI in RAN software that puts machine learning decisions inside 5G network optimisation rather than in an offline planning tool.
- Nvidia is pitching always-on AI agents for telecom operations, which is the same agentic-inference demand curve arriving inside the operators who carry it.
- Korea’s AI factory build-out expands through Nvidia partnerships with SK Telecom, NAVER and Brookfield, putting sovereign infrastructure at gigawatt scale on the operator side of the network.
👨💻 Code Corner
Before you buy anything, find out which half of inference your workload actually lives in. This estimates prefill time from raw compute and decode time from raw memory bandwidth, using nothing but your model size and your prompt length.
# Rough first-order split of an inference request into its two phases.
# Prefill is compute-bound; decode is memory-bandwidth-bound.
def split_phases(params_b, prompt_tokens, out_tokens,
tflops, bw_tbs, bytes_per_param=2, mfu=0.4):
"""params_b: billions of params. tflops: dense TFLOP/s. bw_tbs: TB/s."""
weights_bytes = params_b * 1e9 * bytes_per_param
# Prefill: ~2 FLOPs per param per prompt token, over the whole prompt at once.
prefill_s = (2 * params_b * 1e9 * prompt_tokens) / (tflops * 1e12 * mfu)
# Decode: every token streams the full weight set through the ALUs.
per_token_s = weights_bytes / (bw_tbs * 1e12)
decode_s = per_token_s * out_tokens
return {
"time_to_first_token_s": round(prefill_s, 3),
"inter_token_latency_ms": round(per_token_s * 1000, 2),
"interactivity_tok_s": round(1 / per_token_s, 1),
"decode_share_of_wall_clock": round(decode_s / (prefill_s + decode_s), 2),
}
# A 31B model, 100k-token prompt, 2k tokens out, on one 8 TB/s accelerator.
print(split_phases(params_b=31, prompt_tokens=100_000, out_tokens=2_000,
tflops=2_000, bw_tbs=8.0))
Tip
The decode_share_of_wall_clock field is the one to watch. If it is above roughly 0.7, extra FLOPs buy you almost nothing and you are shopping for bandwidth, more aggressive quantisation, or a decode-optimised accelerator. Note that this ignores the key-value cache, which grows with context and makes long-context decode even more bandwidth-hungry than the estimate suggests.
🧰 Toolbox
- NVIDIA Groq 3 LPX: the product page behind today’s big story, worth reading for how it frames interactivity as a first-class metric.
- Nebius Token Factory: a production inference platform that will be one of the first places to try LPX-class generation speed without rewriting anything.
- Gemini 3.7 Flash: half-price agentic coding until December 31, then double, so benchmark it against your workload while the meter is cheap.
- ESP32-S31 Linux BSP preview: a RISC-V microcontroller that now boots Linux, which quietly redraws the line between MCU and SBC projects.
- Jetson embedded systems: the comparison table across the Jetson line, useful now that Orin Nano 2 shifts the entry-level performance per watt.
🎬 Demo Watch (rotating)
The NVIDIA Isaac GR00T Reference Humanoid Robot: an open reference design for academic humanoid research that bolts together a Unitree H2 Plus body, Sharpa Wave five-finger tactile hands, and Jetson AGX Thor onboard compute, with the Isaac GR00T software stack on top. It stands close to six feet, weighs around 150 pounds and carries 75 degrees of freedom across body and hands.
What is genuinely hard here is not the walking, it is the reproducibility. Humanoid papers have been almost impossible to compare because every lab builds a different body, so a policy that works in one place cannot be checked anywhere else. A shared reference platform used by Ai2, ETH Zurich, the Stanford Robotics Center and UC San Diego’s Advanced Robotics and Controls Laboratory makes results comparable for the first time. What is hype: this is a research vehicle, available from Unitree late in 2026, and 75 degrees of freedom on a bench says nothing about hours logged on a factory floor. Watch the reference workflow that Nvidia says it will publish for the Unitree G1 on GitHub and Hugging Face; open code moves this field faster than open hardware does.
📚 From the Blog
- Turning Pixels Into Something the AI Can Eat: the preprocessing stage between a camera and a model, and a good companion to today’s big story because it is the same prefill-versus-decode argument one layer down the stack.
- Building Your First Neuron From Scratch: weights, bias, activation and a single gradient step, worked by hand with no framework in the way.
- The Network Behind the Cameras: the unglamorous plumbing that moves video across a network without melting the switch in the middle.
😀 The Bot Says…
Ninety-two percent of Nvidia’s revenue now comes from the data centre unit. The gaming division walked so the token generator could run at 3,400 per second.
That’s all for today! Reply and tell us which number you actually track for your service: time to first token, or tokens per second after that.


