The model stopped watching the whole video

By Mark 10 min read 0 views

😁 Hello, super humans! Every long-context announcement for the last two years has been the same announcement: the window got bigger, please put more stuff in it. The model released this morning has the biggest window yet and then quietly argues you should stop filling it. The interesting number is not the million tokens it can hold, it is the tokens it refused to spend.

πŸ“° Quick Signals

  • 🧠 AI: Bonsai 2 27B claims near-lossless quality in a footprint roughly nine times smaller, which is the compression result everyone running models on their own hardware has been waiting for.
  • πŸ€– Robotics: Project Kitchen turns manipulation data collection into a VR game, then transfers what players learn to real arms with a handful of on-robot demos.
  • πŸ’» Programming: Bend is a new language that tries to block AI coding mistakes with proofs rather than review, targeting both CPU and GPU execution.
  • ⚑ Electronics: India’s Semicon 2.0 phase targets 200 chip-design startups and 100,000 trained semiconductor technicians, moving the programme past fab groundbreakings into the ecosystem around them.
  • πŸ“‘ Telecom: Proposed FCC ISM band rule changes could break the license-free assumptions that Meshtastic and MeshCore communities were built on.

πŸ” The Big Story: The model stopped watching the whole video

If you have ever pushed an hour of footage through a multimodal model, you already know the failure mode: the bill scales with the length of the file, not with the difficulty of the question. Alibaba’s Qwen team shipped a model today that attacks that directly, and the mechanism is worth more attention than the headline context number.

What happened: Qwen released Qwen3.8-Omni-Flash, its first omni-modal model built around agentic behaviour. It takes text, images, audio and video in and returns text out, with a 1M-token context, function calling, web search and thinking enabled by default. It is API-only at launch on QwenCloud, Alibaba Cloud Model Studio and Qwen Studio, built on the Qwen3.8-Flash-Next base whose weights shipped in August. There are no open weights for this one, so self-hosting is not on the table.

The details: Most video models do a linear pass: decode the file, sample frames at some fixed rate, feed everything in, answer. Qwen calls its alternative agentic perception, and it inverts the order. The model starts from the question, decides which segments are worth watching or listening to, and gathers evidence over several coarse-to-fine rounds, spending compute only where the answer plausibly lives. On OmniVideoBench that moves accuracy from 63.4 to 67.8 while token consumption drops from 145,736 to 79,117, about 45.7 percent fewer. Getting better and cheaper at the same time is unusual enough to be worth checking. Across 29 evaluations Qwen reports an average gain over 25 percent against Qwen3.5-Omni-Plus, with WildClawBench-MM up 36.5 points and AgenticVBench up 22.3. Pricing is $0.15 per 1M input tokens and $0.47 per 1M output, with audio input down over 98 percent per hour against the previous generation. Input limits are practical rather than demo-sized: two-hour video by URL, three-hour audio, 113 languages, stable sampling up to 15 fps. All of these figures are Qwen’s own; no independent reproductions existed at publication.

flowchart LR
    Q["Question"] --> A
    subgraph Linear["Linear pass"]
      L1["Decode whole file"] --> L2["Sample every frame window"] --> L3["Answer"]
    end
    subgraph Agentic["Agentic perception"]
      A["Plan: what must I check?"] --> B["Coarse scan, low fps"]
      B --> C{"Evidence enough?"}
      C -- "No" --> D["Zoom into candidate span"]
      D --> C
      C -- "Yes" --> E["Answer"]
    end

Important

Our take: The context-window race was always a bit of a distraction, and this is the release that makes the alternative concrete: retrieval inside a modality, not just across documents. The pattern is the same one we learned the hard way with text, where stuffing everything into the prompt lost to fetching the right chunk. If you are building anything on top of long media, the lesson transfers even if you never touch this API: make the question drive the sampling. The part we would not skip over is that this is closed weights and a hosted endpoint, priced attractively, right after the open base model it was built on. That is a familiar shape, and it means the cheap number today is a business decision, not a property of the technology.

πŸ—žοΈ More News

🧠 AI

  • Microsoft open-sourced TauGrid under MIT, a Kubernetes-native stack bundling the tau CLI, Kueue queueing, KubeRay, GPU health monitoring and Prometheus so researchers can submit AI workloads without learning Kubernetes.
  • OpenAI published a framework for reporting model misalignment, with review tracks and a set of incident reports drawn from its own reinforcement learning runs.
  • OpenAI launched Astra for Law, a vertical push into legal work that drew one of the busiest discussion threads of the day.
  • Z.ai published a detailed account of how GLM built its own inference infrastructure, which is the kind of writeup that is far more useful than another benchmark table.
  • A new paper proposes infinite-parameter language models that generate and adapt their own weights from live data rather than freezing them at training time.
  • The Economist reports that AI systems now outperform some of the best human forecasters, which is a narrow claim worth reading carefully before repeating.
  • Hacktron published a writeup of security testing against OpenAI’s own surfaces, a reminder that the attack surface of an AI product is mostly ordinary web application surface.

πŸ€– Robotics

  • Agent as Policy puts a general-purpose agent directly in the control loop of a physical robot, with no task-specific or environment-specific training in between.
  • Vecna Robotics raised $31 million led by Unless to scale its autonomous mobile robots and add pallet stacking, de-stacking and trailer loading to the platform.
  • A new general-purpose humanoid market report profiles 27 players including Tesla, Unitree, Figure AI, Boston Dynamics and Apptronik, which is a useful map even if you discount the forecasts.
  • Bank of America now projects roughly 90,000 humanoid shipments in 2026 rising to 1.2 million by 2030, a curve steep enough that the supply chain, not the software, becomes the interesting question.

πŸ’» Programming

  • Martin Fowler wrote up why he does not like LLMs, and it is the rare skeptical piece that argues from software design rather than from vibes.
  • Servo published a one-year retrospective on sponsored development, a useful data point on whether funded open source browser engine work actually compounds.
  • GitLab.com is changing its rate limits, which is worth reading before your CI finds out for you.
  • CrowdSec disclosed a source code exposure and published its own account of what leaked and what it means for users.
  • A thoughtful piece on self-driving codebases argues the interesting unit of automation is the repository rather than the individual pull request.

⚑ Electronics

  • Helge Fykse built a working transceiver around a PCL86, a TV audio output tube whose 300 mA heater was designed so every tube in a set could sit in one series string.
  • A DIY blood-processing centrifuge build, which is a good study in why rotor balance and containment matter more than motor choice.
  • A clear explainer on why wave energy stayed hard while wind turbines became routine, and the answer is mostly about surviving the load cases rather than harvesting the power.
  • The Eye-D conference badge replaces the paper name tag with something that actually catches attention, and it is a tidy low-power display exercise.
  • Communities hosting Flock Safety surveillance cameras are discovering how weak the security around them is, which is the predictable end state of cameras deployed faster than they are hardened.

πŸ“‘ Telecom

  • Orange and Telesat commissioned Europe’s first Telesat Lightspeed gateway at Orange’s Tier-4 teleport in Bercenay-en-Othe, linked by Orange fibre to a planned Paris point of presence, with Intellian tracking antennas doing the pass-to-pass handover.
  • Mavenir launched NetAIShield, which fuses network telemetry, subscriber behaviour, signalling and messaging intelligence into a live fraud picture aimed at SIM farms and application farms.
  • O2 and Freshwave lit five outdoor small cells across Tonbridge town centre with Kent County Council, part of a programme now past 2,500 live sites nationally.
  • Berenberg restarted coverage of Eutelsat with a HOLD and a 2 euro target, a reminder that the LEO build-out story and the LEO equity story are not the same story.

πŸ‘¨β€πŸ’» Code Corner

You do not need today’s model to get today’s lesson. Coarse-to-fine video querying is a loop you can wrap around any multimodal API: ask a cheap, sparsely sampled pass where the answer lives, then spend real tokens only on that window.

import os, subprocess, json
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DASHSCOPE_API_KEY"],
    base_url=os.environ["DASHSCOPE_BASE_URL"],
)
MODEL = "qwen3.8-omni-flash"


def frames(path, start, end, fps):
    """Extract a span as JPEGs; returns the output directory."""
    out = f"/tmp/f_{int(start)}_{int(end)}"
    os.makedirs(out, exist_ok=True)
    subprocess.run([
        "ffmpeg", "-y", "-loglevel", "error",
        "-ss", str(start), "-to", str(end), "-i", path,
        "-vf", f"fps={fps}", f"{out}/%05d.jpg",
    ], check=True)
    return out


def ask(prompt, video_url, schema_hint=""):
    r = client.chat.completions.create(
        model=MODEL,
        messages=[{"role": "user", "content": [
            {"type": "video_url", "video_url": {"url": video_url}},
            {"type": "text", "text": prompt + schema_hint},
        ]}],
        modalities=["text"],
    )
    return r.choices[0].message.content


def coarse_to_fine(video_url, question, duration_s):
    # Pass 1: cheap. Locate, do not answer.
    locate = ask(
        f"Do NOT answer this question yet: {question}\n"
        f"Only identify the single time span most likely to contain the answer.",
        video_url,
        ' Reply as JSON: {"start_s": <int>, "end_s": <int>}',
    )
    span = json.loads(locate)
    lo = max(0, span["start_s"] - 5)
    hi = min(duration_s, span["end_s"] + 5)

    # Pass 2: expensive, but only over ~10 percent of the file.
    return ask(f"{question}\nOnly consider {lo}s to {hi}s.", video_url)


if __name__ == "__main__":
    print(coarse_to_fine(
        os.environ["VIDEO_URL"],
        "What torque value does the narrator specify for the head bolts?",
        duration_s=3600,
    ))

Two passes over a one-hour file usually costs less than one pass, because the first pass is deliberately starved and the second one never sees the other 55 minutes.

Tip

Always pad the returned span, as done above with the 5 second margin. Localisation is the step that fails, and a model that lands 3 seconds early will confidently answer from the wrong shot. If the second pass disagrees with the first, that disagreement is your cheapest signal that the span was wrong, so re-run the locate step at a wider stride instead of trusting the answer.

🧰 Toolbox

  • Qwen-MM-Plugins: Apache-2.0 plugins that make an existing agent harness multimodal-native, each capability shipping as a Skill plus an optional MCP server.
  • TauGrid: MIT-licensed Kubernetes stack for GPU AI workloads, needing a 1.30+ cluster with GPU nodes and Helm 3.
  • Bend: a language betting that proof obligations, not code review, are the right place to catch machine-written mistakes.
  • Hister: a private search engine over the pages you have visited and the files you keep, which is the local-first answer to losing things you already read.
  • Cloudflare Security-Audit-Skill: Cloudflare’s own security review workflow, packaged so you can run the same checklist against your code.
  • whoisinspace.com: exactly what it says, and a nice reminder that a single-purpose page still beats a dashboard.

πŸ› οΈ Build of the Week (rotating)

ESP32-CYD-MiniTV: turns the Cheap Yellow Display into a tiny television with channels, driven by a single button.

  • Difficulty: Beginner
  • Parts: one Cheap Yellow Display (ESP32 with integrated touchscreen), a microSD card, synchronised .mjpeg video and audio files, optional 3D printed retro TV case
  • Why we like it: the whole interface is the CYD’s existing BOOT button, cycling through channels that are just files on the card, and the .mjpeg constraint is an honest lesson in what the older ESP32 on this board can actually decode in real time. Minimal wiring, and the novelty framing hides a perfectly reusable media player.

πŸ“š From the Blog

πŸ˜€ The Bot Says…

Neovim has been sitting on a Bitcoin donation now worth around $800,000, untouched since 2023. Somewhere out there is a text editor with a better investment record than most hedge funds, and it still will not let you exit without the colon.


That’s all for today! Reply and tell us: has a bigger context window ever actually solved your problem, or did you end up writing the retrieval step anyway?