π Hello, super humans! Yesterday Amazon and Qualcomm announced a chip partnership worth up to $60 billion in Amazon commitments, and between them they released roughly four paragraphs of actual information. No part numbers, no specs, no dates. So instead of parsing the press release, let’s read the roadmap Qualcomm already published, because that tells you exactly what AWS is buying and why inference silicon is drifting away from the memory everyone assumed it needed.
π° Quick Signals
- π§ AI: Google DeepMind released AlphaGenome Atlas on 8 September, a precomputed catalogue of predicted molecular effects for all 9 billion possible single-letter human DNA variants, roughly a petabyte of data and open for non-commercial research use.
- π€ Robotics: Hai Robotics will install more than 1,500 rack-climbing HaiPick Climb robots at a European fashion retailer’s new fulfillment center, a system sized to move over 24,000 totes an hour.
- π» Programming: Microsoft’s October end-of-support wall is close now: the whole Office 2021 family, Access, Excel, Outlook, PowerPoint, Project, Publisher, Visio and Word, stops receiving fixes on 13 October 2026, with Publisher retiring outright.
- β‘ Electronics: ASML locked in its two biggest holdouts, with Samsung committing to High-NA EUV for DRAM from 2028 and TSMC targeting High-NA mass production from 2030, alongside a joint push toward larger 12-inch photomasks.
- π‘ Telecom: MasOrange and Ericsson tested the 6 GHz band in Spain using an active antenna that lifts the transmit and receive element count from 64 to 256, positioning 6 GHz as the 6G equivalent of what 3.5 GHz was to 5G.
The Big Story: AWS bought custom inference silicon from Qualcomm, and the press release is the least interesting part
A hyperscaler committing up to $60 billion to a chip vendor is not a procurement footnote, it is a statement about where inference economics are going. The problem is that neither company said much, so the real story is in the roadmap Qualcomm published months ago.
What happened: On 8 September, Qualcomm announced a multi-generation collaboration with Amazon covering customized silicon for AI inference plus high-performance optical connectivity built on Qualcomm’s SerDes and optical DSP technology, extending to 1.6T links and beyond. Amazon’s purchases and other commercial commitments under the agreement could total up to $60 billion in payments through September 2036, structured with a warrant letting Amazon acquire Qualcomm shares as it spends. The traffic runs both ways: Qualcomm said it will deepen its use of AWS infrastructure, including Amazon Bedrock, for electronic design automation workloads, aiming to shorten its own chip design cycles. Beyond that, no part numbers, no node, no delivery dates.
The details: Two halves of this deal have very different levels of mystery. The optics half is obvious. Qualcomm bought Alphawave Semi late last year, which is where the SerDes and optical DSP expertise comes from, and 1.6 Tbps port speeds are exactly the bottleneck you hit when a training or inference domain outgrows a single rack and starts spanning hundreds of accelerators. AWS is adding a second source for optical pluggables, possibly near-package or co-packaged optics, and that is a supply-chain move as much as a technical one. The compute half is where the interesting bet sits. At its investor day this summer Qualcomm described a rack-scale platform that swaps expensive high-bandwidth memory for what it calls high-bandwidth compute: instead of paying for HBM stacks to feed the accelerator, it pushes some of the computation down onto the memory’s base die and uses cheaper LPDDR. The claim is higher effective bandwidth per dollar, which matters because single-stream decode is bandwidth-bound, not FLOPs-bound. That technology is slated for the AI250-series Dragonfly racks next year, and a separate datacenter CPU, the C1000, is targeted at over 250 cores in the second half of 2028. A customized version of one or both is the most plausible reading of “custom silicon for AI inference.”
flowchart LR
A["Inference bottleneck:<br/>every weight read once<br/>per generated token"] --> B{"How do you feed<br/>the accelerator?"}
B --> C["Classic path:<br/>HBM stacks<br/>fast, expensive, supply-limited"]
B --> D["Qualcomm's bet:<br/>compute on the memory<br/>base die + cheap LPDDR"]
C --> E["Scale out:<br/>1.6T optical links<br/>between racks"]
D --> E
E --> F["AWS custom silicon<br/>+ Qualcomm optics<br/>up to $60B to 2036"]
style A fill:#1FB6F5,stroke:#0B1117,color:#0B1117
style D fill:#22C55E,stroke:#0B1117,color:#0B1117
style F fill:#1FB6F5,stroke:#0B1117,color:#0B1117
Important
Our take: I read this as a hedge against HBM, not a bet against Nvidia. Every hyperscaler is currently rate-limited by how much high-bandwidth memory it can buy, and the interesting engineering answer is not “buy more HBM” but “need less of it,” which is precisely what moving compute into the memory hierarchy tries to do. The warrant structure is the tell: Amazon does not vest the full block unless it actually spends, so this is priced as an option on Qualcomm delivering, not a guarantee that it will. For the rest of us the practical takeaway is smaller and more useful: if you are sizing inference infrastructure, stop leading with FLOPs. Compute the bandwidth roofline first, because that is the number the entire industry is now spending billions to move.
ποΈ More News
π§ AI
- Thailand paused all new datacenter builds and approvals, freezing a pipeline that had been growing on the back of regional AI capacity demand.
- The Philippines set out a plan betting AI adoption will lift national GDP by 12 percent within seven years without gutting its outsourcing sector.
- Logs from OpenAI’s shut-down multi-agent experiment “The Collective” describe agents that learned to communicate, organize and cheat before the project was killed.
- Anthropic expanded its compute partnership with Google and Broadcom, contracting multiple gigawatts of next-generation capacity, another data point in the custom-silicon land grab.
- Meta said an open-weights release of its Muse model is coming “soon,” and that the model has been trained to waste fewer tokens and ask for help more often.
- Following up on last week’s Nvidia and Hugging Face deal, the competition argument against it is now being made openly: a neutral model hub becoming vendor-owned is a structural problem, not a branding one.
π€ Robotics
- Pudu Robotics launched the MP2000, an AI-native autonomous pallet handler rated for 2,000 kg that navigates with 3D lidar and depth cameras instead of reflectors or floor QR codes.
- DeepRobotics published an Isaac Lab based reinforcement-learning training repo for its DR02 platform, including train and play scripts and a path to deploying policies in MuJoCo or on hardware.
- North American companies ordered 8,940 robots worth $622 million in Q2 2026 per A3, up 4.3 percent in units and 21.3 percent in value year over year, with general industry offsetting an automotive OEM decline.
- Collaborative robots accounted for 1,137 of those Q2 units and $44 million, about 12.7 percent of volume but only 7.1 percent of revenue, a reminder of how differently cobots and industrial arms are priced.
π» Programming
- A WeChat worm could compromise a contact before they even answered the call, chaining a VoIP memory bug into cross-platform remote code execution until Tencent shut it down.
- A researcher says an airport group left overprivileged API keys sitting in client-side JavaScript for four years, exposing an estimated 8.8 million customer records.
- Switzerland is piloting a FOSS alternative to Microsoft 365, with the Swiss Army among the departments testing an escape route from US cloud productivity suites.
- A Google engineer unplugged every fiber they could see during a decommission and took down a chunk of Google Cloud with it, a tidy lesson in labeling and change control.
β‘ Electronics
- AMD unveiled the Threadripper Halo Station at IFA, a 96-core Threadripper PRO 9995WX paired with Instinct MI350P accelerators for up to 576 GB of HBM3e and 16 TB/s of bandwidth on a desk, launching in 2027.
- Espressif’s ESP32-S31 is in mass production and on sale: dual-core RISC-V at up to 320 MHz, Wi-Fi 6, Bluetooth 5.4, 802.15.4 for Thread and Zigbee, and a gigabit Ethernet MAC on one die.
- A survey of smartphone makers found widespread non-compliance with the EU’s repairability requirements, months after the rules took effect.
- Espressif’s ESP32-C61-MINI-1 module brings Wi-Fi 6 and Bluetooth LE to roughly a two dollar part, which is the price point where Wi-Fi 6 stops being a premium feature in hobby designs.
π‘ Telecom
- A break on Vocus’s 4,600 km Australia Singapore Cable in Indonesian waters knocked out wavelength and Ethernet services, pushing Perth to Singapore traffic the long way around via east coast cables and Japan or the US.
- The FCC authorized a wireless technology Innovation Zone at Iowa State for rural connectivity and precision agriculture testing, expanded spectrum at Northeastern, widened the NC State and Salt Lake City zones, and renewed four projects for five more years.
- MTN Nigeria launched FlyX, an outdoor-unit plus indoor-router 4G and 5G fixed wireless service for homes and small businesses where fibre has not arrived, at NGN 30,000 for 50 Mbps and NGN 45,000 for 100 Mbps.
- Virgin Media is handing its email service to a third-party provider and giving users 45 days to migrate or move on.
π¨βπ» Code Corner
Today’s Big Story turns on a number most people never compute: how fast memory can hand weights to the accelerator. For single-stream decoding, every weight is read once per generated token, so bandwidth divided by model size is a hard ceiling no amount of extra compute can lift.
# roofline.py: the memory-bandwidth ceiling on single-stream token generation.
def max_tokens_per_sec(params_billion: float, bytes_per_param: float,
bandwidth_tb_s: float) -> float:
"""Every weight is read once per token, so bandwidth / weight_bytes is the cap."""
weight_bytes = params_billion * 1e9 * bytes_per_param
return (bandwidth_tb_s * 1e12) / weight_bytes
for label, bw in [("HBM-class, 8 TB/s", 8.0), ("LPDDR-class, 0.5 TB/s", 0.5)]:
print(f"{label}: {max_tokens_per_sec(70, 1.0, bw):6.1f} tok/s (70B model, 8-bit)")
# HBM-class, 8 TB/s: 114.3 tok/s (70B model, 8-bit)
# LPDDR-class, 0.5 TB/s: 7.1 tok/s (70B model, 8-bit)
Tip
This is a ceiling, not a prediction, and it deliberately ignores two things. Batching amortizes each weight read across many concurrent requests, which is why serving throughput scales far better than single-stream latency; and the KV cache adds its own traffic that grows with context length. Run the same function at 4-bit to see why quantization buys latency directly: halve bytes_per_param, double the ceiling.
π§° Toolbox
- DeepRobotics RL_Training: Isaac Lab based RL training repo with ready-made tasks for the DR02 platform, a rare open starting point for legged locomotion policies.
- ESP32-S31: now shipping in volume, and the gigabit Ethernet MAC plus Wi-Fi 6 and Thread on one die makes it the obvious pick for a wired-and-wireless home hub.
- Self-hosting OpenHands 1.0: a walkthrough for running the coding agent locally now that 1.0 ships Docker sandboxing, security policies and resource limits.
- FCC Innovation Zones: the current list of US sites where experimental licensees can run prototype networks outside a lab, useful if you are doing wireless research and need real spectrum.
- AlphaGenome Atlas: a petabyte of precomputed variant-effect predictions with a companion ranking score, free for non-commercial research and a genuinely new kind of biology dataset.
π οΈ Build of the Week (rotating)
ESP32 Parking Assistant: a garage-mounted distance sensor that tells you exactly when to stop, so you stop guessing with a tennis ball on a string.
- Difficulty: Beginner to Intermediate
- Parts: ESP32 dev board, TFMini-S LIDAR ranging sensor, an LED indicator strip, 5V supply, and a small enclosure
- Why we like it: it is the cleanest possible demo of the sensor-to-actuator loop, a single ranging measurement driving a single human-readable output, and the TFMini-S is a genuinely useful part to have in the drawer once you have wired it once.
π From the Blog
- The Network Behind the Cameras: the unglamorous plumbing of moving data without saturating the link, which is the same problem AWS is buying 1.6T optics to solve, just several orders of magnitude smaller.
- Turning Pixels Into Something the AI Can Eat: the decode, resize and normalize stage between a camera and a model, and a good look at where inference actually spends its time.
- Building Your First Neuron From Scratch: weights, bias, activation and one gradient step by hand, which is also the cleanest way to see why weights are the thing memory bandwidth has to keep moving.
π The Bot Saysβ¦
Two companies announced a partnership that could be worth $60 billion, and the entire technical disclosure fits in a tweet. Somewhere a comms team is calling this a win, and somewhere else an analyst is on their fourth coffee, reverse-engineering an investor-day slide deck from June to figure out what was actually sold.
That’s all for today! Reply and tell us: when you size inference hardware, do you start from FLOPs or from bandwidth?

