ai-powered-markdown-translatorArticle translated from fr to en with gpt-5.6-sol.
Monday, September 21 belongs to xAI: Grok 4.7 launches and lands in Cursor and GitHub Copilot the same day, with no delay between the model’s announcement and its arrival in the tools. OpenAI chooses the same day to reveal that an internal model trained since August 28 has solved more than one hundred open mathematics problems, and to establish an advisory group of mathematicians following a statement signed by 27 Fields Medalists. Google opens Googlebook preorders, ElevenLabs puts an editing agent at the heart of its video editor, and Hugging Face rereleases its tokenization library with speeds 3 to 30 times faster.
Grok 4.7 launches at xAI and arrives in Cursor and GitHub Copilot the same day
September 21 — xAI releases Grok 4.7, its most capable model for coding and knowledge work. The base model is larger than Grok 4.6’s, and its reinforcement learning phase was extended using a task mix weighted toward problems that take several hours. xAI expects this to produce a model that checks its own work more effectively and handles long contexts better.
The figures paint a mixed picture rather than showing outright dominance. The price, however, remains unchanged: 6 per million output tokens, as with Grok 4.6—half the input price of GPT-5.6 Sol and one-fifth that of Fable 5.1. The argument rests on this price-performance ratio.
| Published benchmark | Grok 4.7 (xHigh) | Grok 4.6 (High) | GPT-5.6 Sol (Max) | Fable 5.1 (Max) |
|---|---|---|---|---|
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| EEBench (electrical engineering) | 64.0% | 53.0% | 39.4% | 56.4% |
| Input price (dollars / M tokens) | 2 | 2 | 4 | 10 |
| Output price (dollars / M tokens) | 6 | 6 | 20 | 50 |
The API exposes grok-4.7 with a 500,000-token context and four reasoning-effort levels, from low to xhigh; pricing doubles beyond 200,000 prompt tokens. xAI also highlights an entirely new guardrail stack: first place on LatchBio’s biosecurity benchmark with 62.4%, and 3.3% of risky dual-use cyber prompts allowed through on HackerBench v0.3.
Deployment is immediate at two integrators. Cursor serves the model from day one—a consequence of SpaceX’s acquisition of Anysphere, completed on August 14. The model remains xAI’s: the post published the same day on Cursor’s blog consists of one sentence and links to the full announcement on x.ai. It is xAI that highlights CursorBench 4.0, Cursor’s benchmark for long-running tasks, where Grok 4.7 scores 46.3% at the same price. GitHub Copilot, meanwhile, is deploying it across all its surfaces, from VS Code to Xcode, for plans ranging from Pro to Enterprise. Its billing deserves attention: Grok 4.7 falls outside the request-multiplier system and is billed at the provider’s public usage rate. Since a new model is enabled by default, controlling spending requires action in the model policy, not merely doing nothing.
🔗 Grok 4.7 — xAI 🔗 Grok 4.7 is now available in GitHub Copilot
OpenAI establishes an advisory group of mathematicians and reveals more than one hundred solved open problems
September 21 — OpenAI announces that it is working with an independent advisory group of nine mathematicians, formed by the members themselves and hosted at the Institute for Advanced Study in Princeton. The context carries more weight than the announcement.
On August 28, we began training a new internal model. In addition to resolving the Navier–Stokes Millennium Prize problem, this model has now resolved more than 100 long-standing open problems across most areas of mathematics. — OpenAI, Advisory Group on Mathematics and Artificial Intelligence
The model has not been publicly named, and OpenAI writes that the pace of this progress surprised its own mathematicians. This acceleration prompted a direct response: on September 11, a statement titled “A Severe Misalignment of AI in Mathematics” was published with a Zenodo DOI and the signatures of 27 Fields Medalists, including Terence Tao, Peter Scholze, and Cédric Villani. Their argument is not that the models are wrong. Solving a major problem has always signaled new understanding, subsequently developed by the community until it becomes teachable; the signatories fear that the mass production of true or false statements will destroy this fertile ground instead of nurturing it, and that these rushed announcements will leave no time for proper exposition.
The arrangement rests firmly on independence: the group may issue unsolicited advice, comment publicly on OpenAI’s impact on mathematics, and change its membership without the company’s approval, while its members are not paid by OpenAI. What remains is the sentence OpenAI places at the end of the same paragraph: the group will not advise the company on the pace of its internal progress. Its remit therefore covers the dissemination of results, not their production—everything except what triggered the protest. Martin Hairer, meanwhile, appears on both lists, as a signatory of the statement and a member of the group.
🔗 A Severe Misalignment of AI in Mathematics
Googlebook, Google’s laptop lineup designed around Gemini
September 21 — Google opens preorders for Googlebook, which it presents as a category of its own: the Android technology stack, desktop foundations inherited from ChromeOS, all designed around Gemini. Five models launch simultaneously, manufactured by Acer, ASUS, Dell, HP, and Lenovo, starting at $899. Deliveries begin on October 4 in the United States and on October 5 in Canada, the United Kingdom, Ireland, France, Germany, and Australia.
| Announced specification | Value |
|---|---|
| Starting price | $899 |
| Processors | Intel Core Ultra Series 3 or Snapdragon X Elite |
| NPU | more than 45 TOPS |
| Memory | 16 GB standard, up to 32 GB |
| Battery life | up to 14 hours of video, 16 hours of web browsing |
| Software support | updates for up to 10 years |
Three Gemini features are unique to the machine. Magic Pointer, triggered by shaking the cursor, brings Gemini to the item displayed on the screen—analyzing a suspicious email or combining several images—and Google specifies that it activates only on request and can be disabled. Rambler turns a rambling dictation into structured text with headings and task lists, including when the speaker switches languages mid-sentence. Create My Widget generates desktop widgets from a natural-language description.
For developers, every machine ships with Antigravity, Google’s development-agent platform, and an isolated Linux environment with a full terminal—the documentation explicitly mentions the ability to run Claude Code or Antigravity CLI there. This isolation relies on a Level 5-certified pKVM hypervisor, which Google presents as a first in the category, complemented by the Titan hardware root of trust. Continuity with the Android phone does the rest: resuming a task started on mobile and streaming mobile applications into a desktop window. Every Googlebook includes 12 months of Google AI Pro with 5 TB of storage, 3 months of YouTube Premium and Photoshop, and one year of GeForce NOW.
ElevenLabs adds an editing agent to Studio 4.0
September 21 — ElevenLabs releases Studio 4.0 in ElevenCreative, an update to its video editor. The central idea can be summed up in one sentence: everything happens in one place. Video, image, voice, music, and sound-effect generation is launched from within the project itself, every ElevenCreative model can be called without leaving the editor, and editing, subtitling, and export all take place in the same interface. A shot can be reworked without restarting the entire generation process, and subtitles can be manipulated like any other clip on the timeline.
The most structurally significant new feature is Studio Agent, a co-editor present in every project.
Studio Agent is your AI co-editor, built into every project. Describe the change you want and it makes the edit. Step in and make cuts whenever you like, then hand the timeline back to the agent. — @ElevenLabs on X
This back-and-forth is what distinguishes the tool from opaque generation: editors can take control whenever they want, then hand the timeline back. The timeline itself has also been rebuilt for multi-scene, multi-track projects—advertising campaigns, tutorials, or short-form series—with frame-level zoom, clip snapping, and the ability to move a group of clips without throwing the rest out of sync. Team comments attach to a specific clip or asset rather than becoming scattered across separate threads.
A note for tracking purposes: at the time of the scan, the announcement existed only on X, in a six-post thread. The ElevenLabs blog ended with its September 10 post, while the documentation changelog ended with the September 14 entry.
🔗 Studio 4.0 in ElevenCreative
Hugging Face releases tokenizers v1, 3 to 30 times faster with identical IDs
September 21 — Hugging Face releases the candidate version of tokenizers v1, the Rust library that turns text into sequences of integers before a model reads it. The initial observation is simple: tokenization has never been the bottleneck in machine-learning pipelines, but the balance is shifting as models become faster, and a GPU waiting for its CPU is an idle GPU.
| Metric | Measured value |
|---|---|
| Encoding speedup over v0.23, single thread | 3 to 30 times |
| Benchmark machine | Apple M4 Max |
| Low and high ends of the range | t5-base and gpt2 |
| Scaling across eight workers | 76% of linear scaling |
| Token IDs produced | identical to v0.23 |
| Candidate version installation | cargo add tokenizers --pre |
The constraint the team imposed on itself is more interesting than the gains themselves: v1 produces exactly the same token IDs as v0.23; the API, vocabulary, and merge ranks do not change; and the library remains general-purpose rather than specializing in BPE. The update therefore requires nothing more than changing the installed version.
Six changes account for most of the work, and the most unusual is called bitcannon. The regular expression that splits text into pre-tokens is a fixed parameter shipped with the tokenizer: there is no reason to have a general-purpose engine interpret it for every encoding operation. Hugging Face therefore writes the splitting function by hand and has it operate on bit streams using SIMD, deciding 64 bytes per register operation. The tradeoff is explicit: only recognized patterns (GPT-2, cl100k, o200k, Tekken, DeepSeek) benefit, while the others retain the regex path—which explains the range from 3 to 30. These changes are joined by an allocation-free, branchless merge loop, a thread-local word cache, and native parallelism. Inference-only C and C++ bindings, intended for ExecuTorch and llama.cpp, are announced for release before 1.0.0.
🔗 tokenizers v1: encode, decode and scaling, measured
Devin Cloud comes to the terminal, with SSH access to the agent’s machine
September 21 — Cognition brings Devin Cloud to the terminal: devin --cloud, or /cloud from an active session, creates, controls, and resumes a cloud session without leaving the CLI. The mechanism relies on a symmetric command: /handoff transfers the current task to Devin’s virtual machine—Mac, Linux, or Windows—and, when run from the cloud, performs the reverse journey by retrieving the pull request branch. Cloud sessions survive the terminal that launched them.
And for the first time, SSH into Devin’s dedicated VM, then /handoff the work back to your own device. — @cognition on X
devin ssh provides full access to the environment the agent has built for itself: editing code, running a development server, forwarding ports, and retrieving files with scp. The runtime environment is no longer a black box. SWE-2 sessions are free in Devin Cloud until October 8.
🔗 Devin Cloud in your terminal
Qwen Code v0.24.3 instruments its Web Shell and persists its sessions via JDBC
September 21 — Qwen Code moves to v0.24.3 the day after v0.24.2: thirty-one pull requests and no breaking changes. Most of the visible work concerns the Web Shell, whose execution results are no longer raw text but a structure separating the command, its output, execution details, and elapsed time. A trajectory tab, disabled by default, shows per-turn durations, time to first token, token count, and the duration of each tool call.
The least visible part is probably the most foundational. The runtime gains JDBC persistence for its bindings and sessions, allowing state to be shared across multiple broker processes and session ownership to be retained after a restart. Runtime tools are also routed through bwrap with an immutable workspace policy—the direct continuation of the sandbox introduced the previous day, whose setting remains unavailable to users. On the same day, the TypeScript SDK v0.1.14 bundles this CLI, followed by Qwen Code Desktop v0.24.3.
🔗 Qwen Code v0.24.3 release notes
Hugging Face releases relore, repository memory for coding agents
September 21 — Hugging Face releases relore under the Apache-2.0 license, a tool that gives a coding agent the memory of a repository. It began with an incident: on Transformers issue #48630, opened on September 8, two fixes had already been proposed when the company’s continuous integration agent was about to write a third. All the information was on GitHub, but scattered across issues, pull requests, and comments that were invisible from the original issue.
relore indexes this history while recording who said what: a maintainer’s decision does not carry the same weight as a contributor’s hypothesis, and machine-generated content is excluded by default. A working clone is maintained alongside the index, served through the same HTTP API, and queries are processed locally—only ingestion contacts GitHub. The commands follow: inflight lists the threads that claim to close an issue, search --trust authoritative surfaced PR #39847 as the source of the regression, and why starts from a line of code to retrieve the review comments surrounding it. During a field test, the agent reported that defs gave it a better overview of a file than reading it in full, using 5% of the tokens.
🔗 relore, repository memory for coding agents
Perplexity post-trains Computer on its users’ errors and adds video
September 21 — Perplexity Research explains how the company post-trains the model powering its Computer agent—GLM 5.2, as the article reveals in passing—using real user sessions. Synthetic environments produce controlled tasks that are easy to verify; real sessions provide two signals they do not: corrections made by the user and errors returned by tools. Sessions containing personal data and those belonging to users who opted out of training are excluded before any sampling.
Each turn then receives one of three treatments: cross-entropy imitation for error-free turns from successful sessions, Kullback-Leibler divergence correction for faulty turns with a valid hint, or simply remaining in context. The same set of weights acts as both teacher and student, but only the teacher sees the hint. Three automated judges identify the potentially responsible turns, and at least two must agree, because the last turn before the complaint is the actual cause only about half the time.
| Tool error rate | Measured value |
|---|---|
| Offline, original GLM 5.2 | 2.79% |
| Offline, rejection sampling only | 1.35% |
| Offline, with self-distillation | 0.87% |
| Online, between two checkpoints | 2.24%, then 1.77%, a relative reduction of 21.2% |
The authors are explicit about what these figures do not demonstrate: the checkpoints compared offline were not trained on the same data, task-level results remain mixed, and user dissatisfaction does not change significantly—2.58% versus 2.54%—in A/B tests involving roughly one hundred thousand users per arm.
On the same day, Computer gains video generation, handled by two third-party models: MiniMax H3 and ByteDance Seedance 2.5. Whether it is a campaign clip, product demonstration, or social media visual, the finished video appears in the same thread as the text and images produced for the same request. It is available to Pro and Max subscribers, with no mention of free access or availability through the API.
🔗 Learning from Real-World Experience 🔗 Perplexity Computer and video models
Gemini Notebook makes Live Chat and Interactive Learning Overviews widely available
September 21 — Gemini Notebook reached two deployment milestones on the same day. Live Chat reaches 100% of Ultra subscribers on mobile: real-time voice conversation with a notebook in roughly 100 languages, powered by Google’s latest audio models, with nearly entirely hands-free step-by-step guidance. A few hours later, Interactive Learning Overviews leave their gradual rollout and become available to all users, with no subscription required; located under Reports, they bring source summaries and studio artifacts together in a single interactive space.
Both features had been introduced on September 15 in the announcement of the new study tools. The September 21 posts add no features: they confirm actual availability, complete for the first and open to everyone for the second.
🔗 Live Chat rolled out to 100% 🔗 Interactive Learning Overviews open to everyone
Google.org commits $4 million to AI training for educators
September 21 — Google.org commits $4 million to Digital Promise, announced in New York on the sidelines of the United Nations General Assembly. The funding expands the Google AI Educator Series, a free training program announced earlier in the year for the 6 million primary, secondary, and higher-education teachers in the United States. Designed with ISTE around a train-the-trainer approach, it focuses on skills transferable from one tool to another rather than on a specific product and consists of self-paced modules validated by digital badges.
Digital Promise is involved in three areas: adapting the training for state education agencies and school districts, partnering with community college networks for undergraduate educators, and conducting a two-year study of what actually helps higher-education instructors, whose findings will be published in a free public guide.
🔗 Google.org and Digital Promise
Runway expands its Workflows to compositing, transparency, and depth maps
September 21 — Runway expands its Workflows—the sequences of operations assembled in its application—with five post-production-oriented components: Compositing, Alpha, HDR, Depth Map, and RGB Depth. The announcement’s argument fits in a single sentence: build more of the production pipeline in one place, without switching to another tool between steps. Compositing and alpha-channel handling are intended for assembling multiple image layers; depth maps, in both forms, provide a geometric signal for controlling effects that depend on distance from the camera. The announcement was made on X with a 31-second demonstration, with no corresponding post on Runway’s news page at the time of the scan.
🔗 Expanded Workflows in Runway
NVIDIA launches DSX Ready, a qualification program for powering and cooling AI factories
September 21 — NVIDIA launches DSX Ready, a program that qualifies partner products against the requirements of its DSX reference architecture for AI factories. The starting observation is practical: power, cooling, water, the site, and the electrical grid now determine what a builder can actually deploy, and equipment poorly matched to the rest of the design does not turn compute capacity into useful production.
| Qualified category | Launch partners | Qualification process |
|---|---|---|
| Battery energy storage | Hitachi Energy, LG Energy Solution, Tesla | Partner testing, NVIDIA review and approval |
| Cooling distribution units | LG Electronics, LiquidStack, Vertiv | Self-qualification suite |
NVIDIA sets an explicit limit on what the label means: passing qualification does not replace site-level engineering and provides no guarantee of site stability. Other categories covering infrastructure and software are expected to be added gradually.
NVIDIA treats agent security as an engineering problem, layer by layer
September 21 — NVIDIA publishes its position on the security of agentic systems: it is an engineering problem, requiring defined requirements, enforceable controls, named owners, and evidence that the safeguards hold. The stack is divided into three levels—models provide capabilities, harnesses organize context, tools, and workflows, and the runtime environment supplies the infrastructure where actions occur—each carrying its own controls.
The chosen example is telling: an agent updates a customer record, encounters malicious instructions in an attached document, and attempts to export the data to an unauthorized destination. NVIDIA then expects a network policy to block the transfer and protected logs to retain the attempted tool call, the authorization decision, and the result. The right to update a record does not extend to the right to export it, and an agent may request access without ever being able to grant it to itself. On the tooling side, OpenShell enforces policies beyond the agent’s reach; Cisco layers DefenseClaw on top, and JFrog verifies the skills an agent can access.
On the same day, a second post applies the same reasoning to physical AI and connects it to the Halos foundation; it appears in the briefs below, along with its two market projections.
Earth-2 assimilates local observations to refresh weather forecasts
September 21 — NVIDIA publishes a tutorial on Earth-2’s AI data assimilation tools, intended for sectors that already have their own measurements: wind and solar farms, local radars and sensors, and satellite observations. The first technique, SDA, applies to diffusion models such as CorrDiff and StormCast: at each denoising step, it compares the intermediate result with observations and guides the next step, without retraining.
| Measured technique | Scope and resolution | Gain reported in the example |
|---|---|---|
| CorrDiff-SDA | Europe, from 0.25 degrees to 2.2 km | Wind RMSE reduced by 54% at held-out stations |
| StormCast-SDA | Continental United States, 3 km, initialized HRRR | Wind RMSE reduced by 7.2% across six time steps |
| HealDA | Global, 1-degree HEALPix grid | Global atmospheric state estimated in a few seconds |
The benefit is primarily operational: numerical analyses require substantial compute time and are produced on fixed schedules, whereas SDA continuously incorporates observations and closes the gap with the current time. HealDA, meanwhile, treats each data stream as a point cloud in space and time, aggregates it onto the target grid, and then passes it through a vision transformer-style backbone network.
🔗 Data assimilation in Earth-2
Multiverse Computing prunes blocks like an Ising glass and gains 23 MMLU points
September 21 — Removing entire transformer blocks is one of the least expensive ways to accelerate a large language model, and one of the most brutal: the model literally becomes shorter, but removing the wrong blocks causes it to collapse. Multiverse Computing argues that the question is framed incorrectly. Existing methods score each block in isolation—magnitude, sensitivity, block influence—which the authors characterize as mean-field approaches, even though the effect of removing one block depends on which others are removed at the same time.
Selection therefore becomes a constrained binary optimization problem that maps onto an Ising glass: a disordered spin system with all-to-all interactions and a fixed number of upward-pointing spins. The practical benefit can be summed up in one sentence: the energy of this system is an inexpensive proxy for the score the pruned model will achieve, making it possible to rank a very large number of candidate configurations without evaluating any of them on benchmarks. At 50% compression of Llama-3.3-70B-Instruct, the authors claim a gain of nearly 23 percentage points on MMLU over the best competing method.
🔗 Pruning LLMs like a physicist
Briefs
NVIDIA published seven articles and posts on September 21: three are covered above, while the other four appear here, with no technical announcements.
- NVIDIA calls for security at every layer of physical AI — a position paper backed by two market projections: 49 million Level 3 to 5 autonomous vehicles installed by 2035 according to ABI Research, and approximately 60 million industrial robots deployed between 2026 and 2035 according to Omdia. Four security requirements are tied to the NVIDIA Halos foundation. 🔗 source
- NVIDIA highlights the scaling of Egypt’s AI ecosystem — a reception at the Grand Egyptian Museum featuring opening remarks by Paolo Guglielmini, EMEA vice president. The introduction mainly notes that AI factories are coming online in South Africa, Morocco, and Nigeria. 🔗 source
- NVIDIA lists five companies applying AI to low-carbon energy — an overview published for New York Climate Week, spanning fusion acceleration and the reuse of recycled electric-vehicle batteries. ThinkLabs AI is building digital twins and agents on CUDA to shorten grid-connection timelines. 🔗 source
- NVIDIA announces the Golden Ticket winners for GTC Berlin — six winners selected by a jury for the October 20–22, 2026 conference. No technical content. 🔗 source
- Qwen-Image-2.1 supported on day one by vLLM, SGLang, ComfyUI, and Hugging Face — SGLang-Diffusion runs the model on a single 24 GB RTX 4090 with CPU offloading, generating a 1024-by-1024 image in 18.7 seconds while using 22.7 GiB. vLLM describes it as a 7.1-billion-parameter diffusion transformer paired with Qwen3-VL-8B. 🔗 source
- OpenAI Academy adds four learning paths — Build with AI, Lead AI Adoption, AI for Educators, and AI for College Students join Apply AI at Work. Each course ends with an assessment that awards a badge; the developer path separates operating Codex from building on the API. 🔗 source
- Runway hosts a hackathon in San Francisco centered on its API — a one-day event on September 30 at the Masonic, alongside the Runway AI Summit, to build an agent, application, or creative workflow, with no pitch deck or approval process. 🔗 source
- HeyGen highlights Code2Video, a benchmark built with Google DeepMind and Kaggle — it measures a question distinct from writing code: whether that code produces a video worth watching. Published as a comment on a post about the agentic video stack based on code generation. 🔗 source
- Suno demonstrates sound effects applied through simple descriptions — a one-minute demonstration ranging from warm saturation to a more experimental result. No feature name, eligible plan, or accompanying blog post: a product showcase rather than a launch. 🔗 source
- Live Gemini shopping demonstration on Discord — a session announced for Tuesday, September 22 at 11 a.m. PT on the official Discord server, focused on planning a purchase, comparing options, and searching from a photo. No new feature. 🔗 source
- Together AI and Ollama announce a session on open coding agents — Parth Sareen from Ollama and Zain Hasan from Together AI will host a session on deploying coding agents based on open models. No date or location was included in the post. 🔗 source
- Weights released for two abandoned Boris-2 training runs — the first plateaued at a loss of 3,100 after 30 billion tokens because of an incompatibility between Muon and AdamW gradients; the second was stopped at around 7 billion tokens. A story of failure more instructive than the models themselves. 🔗 source
- A preregistered protocol addressing directional bias in judge models — the hypothesis is that, for narrative text, errors lean toward explicitly stating a sentiment rather than encoding it. The protocol was published before any data collection, together with the result that would lead to abandoning the construct and a conflict of interest declared by the author. 🔗 source
- An open-stack reproduction of Jev falls short by 0.6 points — a user who says they cannot program delegates a day of experiments to Claude Code to rebuild Jev using open, CPU-only stacks. Eight failed approaches, which they consider more instructive than the one that worked. 🔗 source
- Jevini separates decision design from execution — a large reasoning model designs a versioned graph of typed judgments, which a small decision model executes for each new situation. Promotional content from the vendor, to be treated accordingly. 🔗 source
- Cohere maintains its brand communications, with no product announcement — a September 20 post continues the AI for Empowerment campaign by restating the intent attributed to cofounder Aidan Gomez, while two visuals from September 21 recount the company’s journey to the student hackathon Hack The North in Waterloo. No model, feature, or pricing. 🔗 campaign · 🔗 hackathon
What It Means
Day zero is becoming the distribution norm. Grok 4.7 does not wait: it is available in Cursor and GitHub Copilot on the day it is announced, while Qwen-Image-2.1 was supported by vLLM, SGLang, ComfyUI, and Hugging Face as soon as it was released. The delay between a model release and its availability in everyday tools, once measured in weeks, is trending toward zero—in Cursor’s case because the editor and the lab have belonged to the same group since August 14, and elsewhere because integration is prepared before the announcement. The interesting question is no longer when a model will arrive, but what it will cost once it does.
That is precisely what Copilot’s billing policy changes. Grok 4.7 leaves the request-multiplier system and is instead billed at xAI’s public usage rates. Combined with the automatic enablement of new models, this shifts a cost decision to platform administrators: doing nothing amounts to opening a line item indexed to a third-party provider’s pricing. From xAI’s perspective, the argument is symmetrical—the same $2 input price as Grok 4.6, real gains on Terminal-Bench, but one place behind Fable 5.1 on long coding tasks: what is being sold here is price-performance, not first place.
For agents, three distant announcements tell the same story: the execution environment is no longer a black box. devin ssh opens the agent’s virtual machine to an interactive session, relore gives the agent a repository’s history of past decisions rather than only the code’s current state, Perplexity post-trains its model on the exact turns where it failed for users, and Qwen Code routes its tools through bwrap. NVIDIA provides the theoretical framework for this shift by arguing that agent security must be handled layer by layer, beyond the agent’s reach, with logs that serve as evidence. The inspectable agent and the constrained agent are two sides of the same need.
What remains is the question that OpenAI’s announcement raises without addressing. An internal model that solved more than a hundred open problems in under a month prompted an institutional response—an independent, unpaid advisory group free to express its views—whose mandate explicitly excludes the pace of internal progress, the very subject of the protest by 27 Fields medalists. The contrast with the rest of the day is instructive: while the pace of spectacular results is being debated, a tokenization library achieves a 3-to-30-fold speedup while guaranteeing identical IDs, and a compression method borrows from spin glasses to gain 23 MMLU points at 50% compression. That work does not make headlines, but it determines what tomorrow’s inference will cost.
Sources
- Grok 4.7 — xAI
- @SpaceXAI on X, Grok 4.7 announcement
- Grok 4.7 is now available in GitHub Copilot
- Advisory Group on Mathematics and Artificial Intelligence
- A Severe Misalignment of AI in Mathematics
- Googlebook preorders
- ElevenLabs on X, Studio 4.0
- ElevenLabs on X, Studio Agent
- tokenizers v1: encode, decode and scaling, measured
- Devin Cloud in your terminal
- Cognition on X, devin ssh
- Qwen Code v0.24.3 release notes
- relore, repository memory for coding agents
- Learning from Real-World Experience
- Perplexity on X, video models in Computer
- Gemini Notebook on X, Live Chat
- Gemini Notebook on X, Interactive Learning Overviews
- Google.org and Digital Promise
- Runway on X, expanded Workflows
- NVIDIA DSX Ready
- Securing the agent stack
- NVIDIA, physical AI security and the Halos foundation
- Data assimilation in Earth-2
- Pruning LLMs like a physicist
- NVIDIA, Egypt’s AI ecosystem
- NVIDIA, five companies and low-carbon energy
- NVIDIA AI on X, GTC Berlin Golden Tickets
- Alibaba Qwen on X, Qwen-Image-2.1 support
- Expanding OpenAI Academy with new learning paths
- Runway on X, API hackathon
- HeyGen on X, Code2Video benchmark
- Suno on X, sound effects described in natural language
- Gemini App on X, shopping demonstration
- Together AI on X, session with Ollama
- Boris-2-0907 and Boris-2-0917
- Summarization Bias: A Pre-Registered Test for a Directional Failure in LLM Judges
- I benchmarked Jev against open CPU-only stacks
- Decision Circuit Engineering
- Cohere on X, AI for Empowerment campaign
- Cohere on X, Hack The North