Search

Claude Unifies Chat and Cowork, Zed Opens Delta to Everyone, and Vera Rubin Debuts in MLPerf

ai-powered-markdown-translator

Article translated from French to English with gpt-5.6-sol.

View project on GitHub ↗

This Wednesday, Anthropic is merging chat and Cowork into a single Claude that chooses its own tools, and launching Claude Docs and Claude Slides in beta. Zed is opening Delta to everyone by replacing the pull request with an in-thread review, while NVIDIA is debuting Vera Rubin NVL72 in MLPerf Inference v6.1 as electricity emerges as the key constraint for AI factories. Branded agents are also entering advertising, both in ChatGPT and on YouTube, while Cohere signs the definitive agreement with Aleph Alpha and Mistral arrives in Firefox.


One Claude: Chat and Cowork Merge, Claude Docs and Claude Slides Arrive in Beta

September 16 — Anthropic is merging chat and Claude Cowork: there is no longer a mode to choose, as Claude decides for itself whether to search the web, read files, or execute code, and delivers finished files. Long-running tasks continue in the cloud even when the computer is shut down, but local files, the integrated browser, and computer control (computer use, in beta on Pro and Max) require Claude Desktop to remain open.

Today, chat and Cowork start merging into one Claude. The direction: one Claude that carries context across everything you’re working on, wherever you are. — Boris Cherny (Anthropic) on X

The new Claude Docs and Claude Slides create documents and presentations, while Claude Design, launched in April, now works within conversations. All three are in beta on paid plans, subject to Enterprise administrators’ choice. Each creation can be shared via a link, with presentations exportable to PowerPoint or PDF. In Manual mode, the default setting, Claude asks before every action; in Auto mode, automated security checks precede every action.

Plan or audience concernedAnnounced schedule
Pro and Max (web, desktop, mobile)“Over the next few weeks,” with nothing to enable
Team and Free“Soon,” with no date
EnterpriseAt least 30 days’ notice before any change

The switch is permanent, but Cowork tasks, projects, connectors, skills, and artifacts are preserved. At launch, the help center notes several limitations: no importing from GitHub or conversation branching, incognito conversations opening in the old experience, older Cowork tasks missing from search, and Dispatch closed to new users. Chat and Cowork had already shared projects and artifacts since July 7, and memory since August 25.

🔗 Anthropic blog post 🔗 Help center

Design, Slides, and Docs Also Work in Claude Code

On the same day, @ClaudeDevs said that Claude Design, Claude Slides, and Claude Docs also work in Claude Code: users can request a design review deck or an interface mockup by pointing Claude to the repository’s actual files and RFCs, edit the result themselves or within the conversation, and then share the link. For Design, this is not entirely new, since the /design command has been available in research preview since August 17: Slides and Docs are the new additions. The tweet does not specify a minimum version, plans, or documentation, and the notes for the latest published release do not mention them.

🔗 Announcement from @ClaudeDevs on X


Zed Opens Delta in Public Beta and Replaces the Pull Request with an In-Thread Review

September 16 — Zed is opening Delta to everyone. Launched on August 12 by registration, then on Windows on September 4 for registered users, the tool can now be freely downloaded on macOS, Linux, and Windows, used on the web without installation, and its threads can be followed from a mobile device.

The central idea is to collaborate without committing or pushing. A teammate invited into the conversation with the agent sees the same worktrees, can resume them on their own machine, ask the agent about its choices, or continue if the author disconnects. Code review, previewed on September 10, takes the place of the pull request: a subthread guides the reviewer through the branch’s changes, with the original agent’s context, on an isolated copy of the worktrees. The /land command then has the agent squash the changes into a single commit, pass CI, merge after team approval, and delete the branch.

On Delta’s repository, pull requests have been disabled since the week before the announcement, and 33 people have merged 570 changes there since then; the Zed editor repository remains on GitHub “for now.” Under the hood, DeltaDB complements Git versioning with incremental deltas that preserve changes and messages from both humans and agents between two commits; the commit remains the checkpoint, so a teammate who never opens Delta sees an ordinary Git repository.

Zed says it wants to replace GitHub.com and calls this way of working continuous engineering; its undated roadmap includes Git storage in DeltaDB and eventually envisions CI-like checks within the thread. Delta is free during the beta, paid plans are announced as coming “soon,” with no price or date, and Zed promises to retain a free version.

🔗 Zed’s post about Delta’s public beta 🔗 Zed’s thread about the /land command


Vera Rubin NVL72 Debuts in MLPerf Inference v6.1 with Up to 3.7 Times the Throughput of GB300 NVL72

September 16 — NVIDIA is debuting Vera Rubin NVL72, its next-generation rack-scale platform, in MLPerf Inference v6.1, for which MLCommons has published the results, with preview submissions. According to NVIDIA’s comparisons between its own submissions, the rack achieves up to 3.7 times the throughput of GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios, using vLLM and NVIDIA Dynamo, and up to 2.5 times the throughput on DeepSeek-R1 using TensorRT-LLM. Nebius also submitted preview results for Vera Rubin NVL72.

NVIDIA attributes these gains to joint hardware and software design: NVFP4 precision for weights, attention, and the KV cache; disaggregated serving that separates prefill and decoding; large-scale expert parallelism; and a sixth-generation NVLink domain, for which NVIDIA claims 10 times the packet throughput of standard Ethernet. The current generation also published its figures:

MLPerf Inference v6.1 testSubmitted systemResult published by NVIDIA
DeepSeek-R1 offline, from 72 to 288 GPUsGB300 NVL72, from 1 to 4 racks99% scaling efficiency
WAN 2.2, text to videoGB300 NVL720.65 720p videos per second, 5.7 s per video
Qwen3-VL, v6.1 versus v6.0GB300 NVL72Up to 1.6 times, thanks to software
Edge-Agentic, new test (Qwen3.6-27B)Jetson AGX Thor, TensorRT Edge-LLMSubmission with no figure in the post

NVIDIA also reports gains achieved after the deadline on GPT-OSS-120B and DLRMv3, which MLCommons has not yet verified. It claims 19 partners, including eight using multi-node Blackwell NVL72 systems, and mentions MLCommons’ future MLPerf Endpoints benchmark, which is intended to standardize the measurement of agentic inference.

🔗 NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut


Electricity at the Heart of AI Factories: DSX MaxLPS at Lambda and the AEMA Alliance

September 15 and 16 — Two NVIDIA announcements put the power grid at the center of AI deployment, with Emerald AI involved in both: an AI factory controlled by its software curtailed its consumption at the request of a power provider, then Emerald AI, Google, and NVIDIA launched an alliance for flexible data centers.

AI Infra Summit: Lambda Validates DSX MaxLPS, an AI Factory Curtails Consumption at the Grid’s Request

September 15 — At the AI Infra Summit, Ian Buck’s NVIDIA keynote focused on the tokens produced per megawatt in AI factories. Lambda validated DSX MaxLPS, NVIDIA software that monitors GPU and rack power consumption and reallocates power according to workloads: 19 HGX B200 nodes operated within the power budget of 16 nodes at full power, increasing throughput from approximately 4 million to 5 million tokens per second (+24%) and improving performance per watt by 23%. NVIDIA projects up to 40% more GPU capacity for Vera Rubin NVL72 within the same power budget, “in suitable environments.” One evening in August, a signal from Silicon Valley Power automatically reduced a Santa Clara AI factory managed by Emerald AI Conductor from 4 MW to 3 MW without interrupting priority inference; more than 200 signals have since been honored, with a response time of under one minute. DSX Flex, which handles this type of signal, will receive its first dedicated commercial deployment in a 96 MW Vera Rubin factory in Manassas, Virginia.

🔗 NVIDIA live blog

AEMA, an Alliance for Data Centers That Adapt to the Grid

September 16 — Emerald AI, Google, and NVIDIA are launching the AI Energy Management Alliance (AEMA), a coalition for data centers capable of adjusting their consumption according to grid conditions: shifting workloads, drawing from storage, relying on associated generation, or responding to incidents. Its principles are to define ride-through, curtailment, and incident-response obligations before grid connection; standardize metrics; create faster connection pathways for verifiable flexibility commitments; and allocate costs according to actual impacts. The alliance is focused on the United States, and its launch partners have not been named. On its website, it cites two estimates: 100 GW that could be unlocked on the existing grid with moderately flexible data centers, and $733 million in avoided costs per GW of new flexible data centers. NVIDIA and Emerald AI had already presented flexible AI factories at CERAWeek on March 31.

🔗 NVIDIA post about AEMA 🔗 AEMA website


Branded Agents Enter Advertising, in ChatGPT and on YouTube

September 16 — On the same day, OpenAI and Google are placing brands’ conversational agents within ad formats in the United States: in testing with selected advertisers at OpenAI and in an opt-in beta for merchants at Google.

ChatGPT Ads Tests Sponsored Agents and Connects to HubSpot and Shopify

OpenAI is adding AI to both sides of its ChatGPT Ads platform. With Sponsored Agents, clicking an ad in ChatGPT can open a clearly labeled conversation with an agent sponsored by the brand, before offering a link to its website. OpenAI specifies that this interaction is distinct from ChatGPT’s independent responses and separate from the original conversation.

ChatGPT Ads updateStatus as of September 16
Sponsored AgentsIn testing with selected advertisers in the United States
Ads Manager Plugin, prompt-driven campaignsRolling out
Suggested copy and visuals in Ads ManagerRolling out
Copy customization and translationLaunched, available on an opt-in basis
HubSpot integrationAvailable
ChatGPT Ads app for ShopifyUS merchants, with other ChatGPT Ads markets from September 23

The post gives neither pricing nor the number of advertisers. Tested in ChatGPT in February, advertising gained a beta, self-service Ads Manager in May, and on August 31 OpenAI claimed $1 billion in annualized revenue as it expanded self-service access.

🔗 Reimagining advertising with AI

Google Places Business Agent in YouTube Ads and Expands Agentic Commerce

Google is publishing its agentic commerce updates for the holiday season. Business Agent, which arrived in Search “earlier this year,” according to Google, is entering beta in YouTube ads in the United States on an opt-in basis for merchants: viewers can ask the agent about products without leaving YouTube. AI performance insights, which compares a brand’s share of voice in AI Mode and AI Overviews, is becoming available in Merchant Center in Australia, Canada, India, New Zealand, and the United States. In Merchant Center, the UCP (Universal Commerce Protocol) integration hub—the protocol introduced by Google on January 11—is adding cart handoff to the merchant’s website and checkout-flow testing, with analytics coming “soon”; these features are rolling out gradually in the United States, followed by Australia and Canada “early next year.” Tapestry (Coach, Kate Spade) is already selling through UCP in Search and the Gemini app. According to Google, product-feed best practices deliver an average of 5% more conversions the following month, and in tests with lululemon, the brand’s conversational attributes appeared in 50% of relevant AI Mode recommendations.

🔗 Google post about agentic commerce


Cohere Signs with Aleph Alpha and Launches Confidential Computing in Model Vault

September 16 — Cohere is making a series of announcements, including two aimed at customers that want to retain control over their data and technology.

Aleph Alpha: definitive agreement and dual headquarters in Berlin and Toronto

Cohere and Aleph Alpha have signed the definitive business combination agreement, whose proposed terms were announced on April 24. The combined company will operate under the Cohere name, with more than 1,000 employees, planned dual headquarters in Berlin and Toronto, and a research center in Heidelberg at Aleph Alpha’s site. The transaction, for which no amount has been disclosed, remains subject to final regulatory approvals; it is expected to close “later this year.” Ilhan Scheer, co-CEO of Aleph Alpha, will then become Cohere’s chief operating officer, and co-founder Samuel Weinbach its chief research officer. The partnership with Schwarz Group continues on STACKIT, its sovereign cloud.

No government or enterprise should have to choose between capable AI and control over their tech. That belief is exactly why we’re joining forces with Aleph Alpha. Together, we’ll meet the rising global demand for frontier AI that’s both powerful and secure. — Aidan Gomez (Cohere), quoted by @cohere on X

🔗 Cohere’s post about the agreement

Confidential Computing in Model Vault, in early access

Cohere is launching Confidential Computing in Model Vault, its dedicated, fully managed inference platform, to protect data while it is being processed, rather than only at rest and in transit. Three mechanisms are described: end-to-end encrypted inference, hardware-enforced isolation extending to the GPU with NVIDIA Confidential Computing, and a signed attestation token available on demand and verifiable with the hardware manufacturer’s public keys. Cohere says that no one, “not even us,” can access customer workloads, without mentioning any third-party audit. Early access is limited to a small group of beta customers, with no pricing, general availability date, or list of models, and the announcement is limited to tweets and a product page. The feature had been announced as coming soon on September 9.

🔗 Cohere’s announcement on X 🔗 Model Vault product page


Mistral powers Smart Window, Firefox’s AI assistant

September 16 — In a joint post, Mistral AI and Mozilla announced that Smart Window, Firefox’s AI browsing assistant still in beta, is now powered by Mistral models, without specifying which ones. The assistant helps untangle complex research and gather sources from open tabs. Mistral is initially powering it for users in France and North America, with the United Kingdom and Germany expected “later this year.” By default, conversations are not retained on Mozilla’s servers, and partners such as Mistral commit to zero data retention. For Mistral, which says it “traditionally” serves businesses, the agreement opens up access to the general public. Arthur Mensch, co-founder and CEO of Mistral, describes two open-source advocates working together, while Anthony Enzor-DeMeo, CEO of Mozilla Corporation, believes a browser should not be a one-way funnel.

🔗 Joint post by Mistral and Mozilla


Grok Build retains a project’s conventions and decisions across sessions

September 16 — SpaceXAI is giving Grok Build, its coding agent, a memory. After each completed turn, Grok reviews it in the background without blocking the session and records the lasting information in a Markdown note: team conventions, decisions and their rationale, and project facts such as the command used to run tests. Task status, tentative conclusions, secrets, and anything already covered by the repository or its documentation are excluded. Notes are organized by project, with a global set for preferences that apply everywhere. The /dream command groups them into thematic files and also runs automatically at intervals, while /memory opens a read-only browser. When returning to a project, Grok reads the topics for the area where it is about to work, but the current conversation takes precedence over the notes. Memory applies only to new sessions. The May 25 beta already listed persistence of decisions across sessions; the post details how it works.

🔗 Memory in Grok Build


Qwen Code v0.24.0 confines the agent at the kernel level on Linux and changes bash hooks

September 16 — Two days after v0.23.4, Qwen Code v0.24.0 adds a kernel-based sandbox on Linux via bubblewrap (bwrap), without containers, root privileges, a daemon, or an image. When enabled, it relaunches the CLI behind a read-only view of the file system, where only the workspace and a few directories (temporary, cache, Qwen, Git) remain writable; qwen sandbox displays the configuration, and --verify tests the confinement. The PR author lists its limitations: no separate PID namespace, the host’s Unix sockets remain reachable, and Git and Qwen settings can be modified. Breaking change: in hooks executed under bash, $QWEN_PROJECT_DIR and its Claude and Gemini equivalents, which are read from the environment rather than substituted, must be enclosed in double quotation marks. web_search is capped per session (200 by default, 10,000 at most), a budget such as +500k in the message sets the size of an entire turn, and cross-session messaging is enabled by default.

🔗 Release notes


Devin launches Code Scans, goal-driven audits of the entire codebase

September 16 — Cognition is launching Code Scans in Devin: users start with a broad objective rather than a specific change, optionally supplying their own criteria (team standards, an accessibility rubric, migration requirements), and Devin investigates the codebase, prioritizes its findings, discusses them, and then opens the requested pull requests. Launched using /scan in the web application and already mentioned in the September 2 release notes, the scan uses the Agentic MapReduce architecture of Devin Security Swarm, introduced in July, in four phases: planning, batching, parallel exploration, and synthesis. Cognition published two trials: on Dioxus, a full debug build of 22 crates fell from 58.6 to 21.0 seconds, a 64% reduction; an SEO scan of devin.ai and cognition.com identified 44 findings, after which devin.ai’s Ahrefs score rose from 87 to 92 and its slow pages decreased by 73%. No pricing was specified.

🔗 Introducing Code Scans


Claude for Small Business expands to 43 workflows and 27 new integrations

September 15 — In a post that our September 15 roundup had missed, Anthropic expanded Claude for Small Business, its small-business plugin launched in May with 15 workflows. It now includes 43, with 27 new integrations including Shopify, Salesforce, TikTok, Atlassian, Zoom, Xero, Gusto, Square, Stripe, and Zapier, and claims more than 900,000 installations. Each workflow starts in approval mode, and autonomy can be configured reversibly on a workflow-by-workflow basis. Anthropic is also relaunching training: a fall tour with Tenex across 10 US cities, more than 150 approved training organizations offering over 750 workshops, and 14 partner webinars. The offering is available on all paid plans; the spring tour recap was published on September 10.

🔗 Anthropic’s post about Claude for Small Business


ElevenLabs launches Reception, an AI receptionist for small businesses

September 16 — ElevenLabs is launching Reception, an AI receptionist for small businesses built on its ElevenAgents platform. It answers questions about services, opening hours, or the address around the clock, in the caller’s language, then books the appointment in the calendar, either as overflow when the team does not answer or after closing time. The service has its own website, reception.ai, and can be tried free for 14 days with 30 minutes of calls.

Reception planMonthly price excluding taxesIncluded call minutes
Basic29 dollars75
Plus79 dollars275
Premium199 dollars1,000
EnterpriseContact for pricingCustom

All plans include more than 70 languages, Google Calendar, Calendly, Zapier, webhooks, and an MCP server.

🔗 Introducing Reception, by ElevenAgents 🔗 Reception pricing


Synthesia opens its Interactive Avatar API, offering real-time avatars at €0.10 per minute

September 16 — Synthesia is making its Interactive Avatar API generally available, enabling products to integrate avatars capable of listening, speaking, and responding in real time. Customers provide their LLM, knowledge base, and logic; Synthesia handles the face, voice, and lip synchronization. The avatar joins a LiveKit room as a video participant through a plugin that, according to the FAQ, currently supports only LiveKit agents written in Python. The API is billed according to usage at €0.10 per minute, with no commitment and 500 free minutes. A Pipecat plugin and an end-to-end offering with managed LLM, knowledge, and voice are announced as “coming soon,” without a date. In July, Synthesia published the research framework for its interactive avatar models.

🔗 Synthesia’s announcement on X 🔗 Synthesia’s Interactive Avatars page


Koa, Salesforce’s CRM model built on Nemotron 3 Super

September 15 — At Dreamforce, Salesforce unveiled Koa, its first reasoning model dedicated to customer relationship management (CRM), created by post-training NVIDIA Nemotron 3 Super with NeMo RL, NeMo Gym, and NeMo AutoModel on a synthetic dataset. Salesforce controls the weights and runs the model on its own infrastructure; according to Salesforce, no customer data was used for training or is used for inference. On CRM Bench, its own benchmark (updating an opportunity, routing a case, scheduling a follow-up), Koa matches or outperforms leading models with three times fewer errors, though the models used for comparison are not named. Already used internally in Slack, it will enter a pilot program in October as a selectable model in Agentforce, including at Formula 1 and Xero, before an expected general release in winter 2026 in US regions.

🔗 NVIDIA’s post about Jensen Huang at Dreamforce


Runway details real-time moderation for its upcoming video models

September 16 — Runway plans to launch real-time video-generation models “in the coming months” and explains how it will moderate them. With continuous streaming, users see the video before any final review. Runway therefore designed and tested a faster system: accelerated input screening, synchronous moderation that will analyze frames from the stream and stop playback as soon as problematic content appears, followed by a review of the complete video before any download or sharing. The classifier, Zentropi’s CoPE-B, is a LoRA adapter on Gemma-4-26B-A4B-it, a mixture-of-experts model with 25.2 billion parameters, of which 3.8 billion are active; screening takes less than 0.5 seconds on average. In testing, borderline inputs briefly allowed prohibited content to appear before playback was stopped. The system will supplement Runway’s existing safeguards, and real-time videos will carry C2PA provenance signals like its other creations.

🔗 Moderation in Real Time


Briefs

  • Zed 1.20.1 adds Gemini 3.8 Flash — the editor’s weekly stable release also brings an optional cursor animation, opens Markdown files in rendered preview mode, and provides macOS and Linux binaries that are around 25% smaller; it fixes a flaw that could expose private files to collaborators through project search. 🔗 source
  • Replit opens custom connectors in beta (Custom Connectors) — workspace administrators can define an HTTPS REST API authenticated with an API key for the Agent to use beyond the official connectors; the documentation mentions Pro and Enterprise customers, but a callout reserves the beta for Enterprise. 🔗 source
  • NVIDIA entrusts agents with porting TileGym to cuTile Rust — eight days after the launch of CUDA Rust, a multi-agent skill included with TileGym ported all 24 public operators in the library, retaining an average of 99.5% of cuTile Python’s performance across 347 configurations measured on DGX B200. 🔗 source
  • OpenAI publishes a guide to Admin Console analytics — the Usage, Insights, and Outcomes views, the last of which measures the share of merged commits with Codex contributions, are intended to connect the use of ChatGPT Work and Codex with results; the cited 245% return on investment is explicitly hypothetical, and the linked documentation remained unavailable when we checked. 🔗 source
  • The OpenAI API introduces controls for key creation — added to the changelog on September 15, five days after expiration dates arrived, the API Key Governance section lets administrators allow only service account keys, only user-owned project keys, or block all new keys; existing keys remain unchanged. 🔗 source
  • Copilot: budget increase requests reach general availability — with usage-based billing for Copilot Business and Enterprise, a member who has exhausted their AI credits can request an additional budget; owners and billing managers set and approve the amount, immediately restoring access. 🔗 source
  • AI Scan for pull requests no longer requires CodeQL’s default configuration — AI-powered vulnerability detection on pull requests is expanding to eligible repositories without that configuration, in public preview for GitHub Advanced Security customers on github.com; GitHub Enterprise Server is not supported. 🔗 source
  • GitHub Advanced Security configurations can be enforced across organizations — since September 15, enterprise administrators can prevent organization administrators, and not only repository owners, from modifying security configurations set at the enterprise level. 🔗 source
  • GitHub disables SHA-1 over HTTPS — as announced on April 20, the September 15 changelog confirms the end of SHA-1 in HTTPS connections to github.com and its partner CDNs, including GitHub Enterprise Cloud; browsers, API clients, and Git over HTTPS are affected, while GitHub Enterprise Server is not. 🔗 source
  • MiniMax H3 becomes available serverlessly (serverless) on Together AI — announced in a September 15 tweet at 4:58 p.m. Pacific Time, the open 33-billion-parameter video model, which generates clips lasting 4 to 15 seconds at up to 2K with native stereo sound, is listed there at $0.1391 per video. 🔗 source
  • MiniMax unveils H3 IP Edition — in a recap of its Tokyo conference with KAGAMI AI, MiniMax says it unveiled an edition of H3 associated with officially licensed Japanese intellectual properties, without providing a license catalog, date, or pricing. 🔗 source
  • Runway adds five third-party models with its Fall 2026 collection — Fish Audio S2.1 Pro, Cartesia Sonic 3.6, MiniMax H3 Max, FLUX Video Upscale, and FLUX Video Edit are joining the platform, with no pricing announced; several were already available elsewhere, including FLUX Video Upscale from Black Forest Labs since August 20 and H3 Max since August 27. 🔗 source
  • Luma launches Camera Angles — released on September 15 at 4:26 p.m. Pacific Time, the Luma Agents feature regenerates the subject of a single photo from selected angles while preserving details, with no named model or figures provided. 🔗 source
  • Grok Voice arrives on fal — the inference platform offers SpaceXAI’s voice model for Speech to Speech and real-time use; according to fal, it responds in 0.70 seconds and costs $0.00083 per second, and a live workshop with SpaceXAI is scheduled for September 22. 🔗 source
  • OpenText and Cohere partner for regulated industries — announced at the ALL IN AI conference in Montreal, the partnership will integrate North and Cohere models into OpenText Aviator agents, deployable on-premises or in private, public, or sovereign clouds; the joint offering is expected to reach customers in early 2027. 🔗 source
  • Manus forms training partnerships in Singapore — the agent is included in SkillsFuture AI Subscription, which offers six months of free access to eligible Singaporeans enrolled in certain AI courses, and in the Singtel AI Pass; an inter-campus challenge drew more than 1,100 registrations, and Manus is launching a global call for learning partners. 🔗 source
  • The University of Manchester forecasts air pollution with Earth-2 — retrained in two days on an 8-GPU node of the Isambard-AI supercomputer, Earth-2 CorrDiff covers the entire United Kingdom at a resolution of 2 to 3 km², and the team promises to release the data and workflows as open source. 🔗 source
  • Children’s Hospital of Philadelphia models hearts in seconds — built on the open-source MONAI framework, CHOP’s service produces cardiac models that previously required around four hours of work, with approximately 200 cases planned this year, while tools built on Newton and currently being integrated can bring device-placement simulations close to real time. 🔗 source
  • Google publishes a study on American teenagers and AI — conducted with RXN among more than 1,000 young people aged 13 to 17, the U.S. Future Report study indicates that 94% of them used AI in the past year and that 55% cross-check answers against other sources; no product was announced. 🔗 source
  • Perplexity publishes a guide to AI text detectors — its example, presented as hypothetical, shows that a detector that misclassifies 2% of human-written texts could flag 64 out of 1,000 articles, 19 of which were written by humans, resulting in an error rate of nearly 30%; the guide advises treating a flag as the beginning of a review. 🔗 source
  • GiniGEN AI ranks 425 OpenRouter models — a community post published on the Hugging Face blog: the ranking adds time-to-first-token latency measured across 329 models and a Korean-language quality score; only 8.5% of scored models receive an A for honorific forms and 9.4% for Korean institutions. 🔗 source

What It Means

The agent is becoming the interface, and the choice of mode is disappearing. In the new Claude, users no longer decide between a conversation and a workspace: Claude chooses among web search, files, and code, then produces a document or presentation. Delta applies the same logic to code delivery, with review taking place in the thread and a /land command that handles commit, CI, and merge in sequence instead of using a pull request. Devin Code Scans starts from an objective rather than a modification, and Grok Build remembers a project’s conventions and decisions from one session to the next. Human control is shifting without disappearing: Manual mode is the default in Claude, team approval is required before merging in Zed, and findings are discussed before any pull request in Devin.

Agents are also entering the commercial process. OpenAI is testing Sponsored Agents that take over from an advertisement, Google is placing Business Agent in YouTube ads and enabling payments with UCP, and both separate the branded agent from everything else: a labeled conversation distinct from ChatGPT’s independent responses on one side, and an opt-in beta for merchants on the other. For small businesses, Claude for Small Business claims 43 workflows and more than 900,000 installations, while ElevenLabs sells a receptionist that books appointments starting at $29 per month. Synthesia charges €0.10 per minute for an avatar that speaks inside its customers’ products, and Salesforce makes Koa a selectable model in Agentforce: the agent no longer merely assists the company; it addresses customers on the company’s behalf.

Trust is becoming a selling point, and vendors are trying to make it verifiable. Cohere and Aleph Alpha emphasize sovereignty, with dual headquarters in Berlin and Toronto, while Model Vault issues an attestation token that the customer controls using the hardware manufacturer’s public keys. Mistral gives Firefox a zero-retention commitment, Salesforce keeps Koa’s weights on its own infrastructure, Runway plans to cut off a video stream as soon as problematic content appears, OpenAI lets administrators restrict API key creation, and Qwen Code confines its agent at the kernel level while publishing a list of its limitations. The strongest guarantees are those customers can verify themselves; an in-house benchmark such as CRM Bench or a “not even us” claim without a cited third-party audit remains merely an assertion.

Finally, computing is running up against electricity constraints. Vera Rubin NVL72 delivers up to 3.7 times the throughput of GB300 NVL72 according to NVIDIA, but the most concrete announcements from these two days concern megawatts: Lambda fits 19 nodes within the power budget of 16, an AI factory in Santa Clara is scaling down from 4 to 3 MW at its provider’s request, and AEMA is calling for faster grid connections in exchange for verifiable flexibility. Throughput per rack remains the technical selling point, but access to the grid determines how quickly those racks can enter service.


Sources