AI Weekly: Four Frontier Models in Four Days
Week of August 11 to 18, 2026 Four labs shipped frontier models within four days of each other this week, and every one of them was tuned for the same thing: agents that stay on task. SpaceXAI released Grok 4.6 and closed its Cursor acquisition, Google shipped Gemini 3.7 Flash at half price, DeepSeek took V4 Pro to general availability and then raised its prices, and Z.ai announced GLM-5.3 with cybersecurity claims that real CVE databases partially back up. Below the model layer, the MCP stateless spec entered its adoption window, and the memory market quietly delivered the most consequential news of all: 2027 DRAM and HBM capacity is reportedly already sold out. As always, the order is models first, then tooling, then standards, then infrastructure. Models set what is possible, tooling determines who can use it, standards decide whether the pieces connect, and infrastructure sets the cost. Models: Grok 4.6, Gemini 3.7 Flash, DeepSeek V4 Pro GA, and GLM-5.3 Grok 4.6 bets everything on long-horizon agents SpaceXAI released Grok 4.6 on August 12, and the release notes read like a thesis statement about where frontier labs think the value is. This is a post-training upgrade over Grok 4.5 rather than a larger base model. The lab held the foundation constant and spent the improvement budget on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning inside agentic environments. The goal is agents that stay on a task across many steps without drifting. The specs: a 500,000-token context window, a new xhigh reasoning-effort level above the existing ladder, and tiered pricing at $2 per million input tokens, $0.50 for cached input, and $6 per million output tokens below 200K prompt tokens. Above that threshold, prices double to $4, $1, and $12. The model is generally available through the xAI API as grok-4.6, is the default model in Grok Build, and ships in Cursor with doubled included usage for the first week. The independent numbers are genuinely interesting. Artificial Analysis scores Grok 4.6 at 61 on its Intelligence Index, up five points from Grok 4.5 and tied with GPT-5.6 Sol Max for third place overall. On AA-Briefcase, a long-horizon professional work benchmark, it posts an Elo of 1,577, narrowly above Claude Fable 5 Max at 1,574. The efficiency story stands out even more: Artificial Analysis reports Grok 4.6 completed its AA-Briefcase workloads in roughly 53 turns and about 0.5 billion input tokens on average, against roughly 103 turns and 2 billion input tokens for Claude Opus 5 Max. Fewer turns means less re-read context on every step, which compounds into real cost savings for production agents. Now the honest caveats. The bolded wins on GDPval-AA v2 and AA-Briefcase sit inside published confidence intervals, so they are statistical ties rather than leads. The comparison set in SpaceXAI's own table excludes Claude Opus 5, which currently tops the Artificial Analysis index at 63. And on the coding rows engineering teams care about most, Grok 4.6 still trails: 65.9% on DeepSWE v1.1 against 73% for GPT-5.6 Sol Max, and 26% on Terminal-Bench v3.0, nearly double its predecessor and still last among the listed frontier models. Artificial Analysis also places it at $0.84 per completed task, less economical than GPT-5.6 Luna and GLM-5.2. Grok 4.6 is a real step forward for long-running agent work and an incomplete one for coding. Gemini 3.7 Flash: coding gains at half price, for now Google released Gemini 3.7 Flash on August 13, 23 days after Gemini 3.6 Flash, and priced it to move. The introductory rate is $0.75 per million input tokens and $3.75 per million output tokens, with output charges including thinking tokens. That pricing expires on December 31, 2026, after which the rate doubles to $1.50 and $7.50, exactly what 3.6 Flash cost at launch. Google also applied the promotional rate to 3.6 Flash, so through year-end the migration decision is about capability, not list price. The specs are unchanged from 3.6 Flash: a 1,048,576-token input context window, a 65,536-token output limit, a March 2026 knowledge cutoff, and multimodal input across text, image, video, audio, and PDF with text output. The API exposes tunable thinking levels of low, medium, and high, and returns an error on the unsupported minimal setting. Availability spans the Gemini API, Google AI Studio, Antigravity, Android Studio, Gemini Enterprise, and Gemini Spark. The benchmark story is all software engineering, and every headline number Google published is a coding or automation test. The flagship result is DeepSWE v1.1 at 65.3%, against 49.0% for Gemini 3.6 Flash, a 16-point generational jump on a long-horizon software engineering eval. Google-reported numbers also show GDM-MRCR v2 long-context retrieval improving from 91.8% to 97.0% at 128K, and OSWorld-2.0 computer use rising from 33.8% to 47.9%. Those figures are vendor-reported, so treat them as release evidence rather than independent results. On the independent side, Artificial Analysis scores the model 56 on its Intelligence Index against 52 for 3.6 Flash, and ranks it first of 186 models on output speed at 340.1 tokens per second. GPT-5.6 Terra still leads on DeepSWE, Terminal-Bench, and OSWorld in cross-vendor comparisons. The practitioner takeaway: a model that resolves an agentic task in fewer intermediate steps saves both the output tokens on those steps and the input overhead of re-reading a growing conversation on every call. At $0.75 input with a 1M window, high-volume document extraction, agentic search, and classification workloads are exactly where this price cut compounds. Test it before January, because the price doubles after that. DeepSeek V4 Pro goes GA, then raises prices DeepSeek moved V4 Pro to general availability this week with the 0813 checkpoint, ending a preview that began with the April 24 launch. The company updated its API pricing page on August 12 to map the deepseek-v4-pro endpoint to DeepSeek-V4-Pro-0813, and OpenRouter listed the model the same day. There was no blog post and no press release, just a changed model table. Existing integrations keep the same model name and base URL, and the release retains the 1-million-token context window, 384K maximum output, thinking and non-thinking modes, tool calls, and native Responses and Anthropic API compatibility. The benchmark claims are large and unverified. DeepSeek's own table shows broad agent and coding gains over the Pro Preview build, with reported improvements of up to 49.9 percentage points on individual tests, and a Humanity's Last Exam with tools score rising from 48.2 to 60.0. No third-party evaluator has replicated the headline numbers yet. Where independent measurement exists, the picture is more modest: Artificial Analysis scores V4 Pro at 53 on its Intelligence Index, one point above DeepSeek's own near-free V4 Flash at 52 and ten points below Claude Opus 5 at 63. One neutral harness places it second on SWE-bench Verified at 96.40%, behind only Claude Opus 5, while LiveBench ranks it last of seven frontier peers on agentic coding. Strong patch-style coder, weak long-horizon agent. The bigger story is the price reset. Since May, V4 Pro has cost $0.435 per million input tokens on a cache miss, $0.003625 on a cache hit, and $0.87 per million output. This week DeepSeek moved both V4 models to peak and off-peak billing, with V4 Pro at $0.66 input and $1.98 output off-peak and $1.32 and $3.96 at peak. Cache-hit input rises up to 12-fold at peak hours. Even after the increase, per unit of work DeepSeek stays cheap, at roughly $0.06 per completed benchmark task against $2.34 for Claude Opus 5 on the Artificial Analysis measure. But the direction matters: the era of DeepSeek pricing as a loss-leader appears to be ending, and the company itself warns of further increases with no timeline disclosed. Teams that built cost models on DeepSeek's flat rates should rerun the math on their actual traffic hours. GLM-5.3 arrives with CVE receipts Z.ai announced GLM-5.3 on August 14, its new flagship for complex software engineering, long-horizon agentic tasks, and cybersecurity work. Architecturally it follows the same playbook as Grok 4.6: the GLM-5.2 base model, roughly 750 billion parameters, held constant, with all the claimed gains coming from expanded post-training. It supports Low, High, and Max thinking effort and a 1-million-token context window. The distinctive claim is security research capability. Z.ai says the GLM-5 line found 2,436 real vulnerabilities, and unlike most vendor claims, this one has partial external validation: FreeBSD and Red Hat CVE entries credit the model line. That is a new kind of benchmark, one where the scoreboard is public vulnerability databases rather than a lab-controlled harness. The access story is the catch. There are no open weights at launch, a break from Z.ai's history, and no public API for roughly two weeks. Availability starts with GLM Coding Plan subscribers, whose tiers run $18 Lite, $80 Pro, and $168 Max per month, now on a credit system. For a lab that built its reputation on open weights, shipping a closed flagship behind a subscription is a strategic tell worth watching. The rest of the week's releases Four smaller releases filled out the window. Alibaba's Qwen team shipped Qwen3.8-27B on August 14, continuing its fast open-weights cadence. NVIDIA released Nemotron 3.5 Lightning 30B A3B in NVFP4, notable for shipping natively in the 4-bit format its Blackwell hardware accelerates. Dots Studio put out dots3-note Preview, and Mixedbread released Toast 1, a new embedding model. None of these moves the frontier, and all of them widen the menu of small models cheap enough to run everywhere. Tooling: Grok Bot, the Cursor Acquisition, and Public Agent Evals Grok Bot gives agents your logins SpaceXAI opened early beta access to Grok Bot on August 11, one day before Grok 4.6, and the pairing is deliberate. Grok Bot is the product and Grok 4.6 is the engine. The pitch, in the launch post's words, is AI teammates that sign in to your tools, use them like you do, and come back with finished work. Each bot gets its own persistent cloud computer, so jobs keep running when you step away. The beta launched on Mac and iOS first, with Windows and Linux desktop builds available and Android to follow. Access is gated to SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers. The differentiator is the absence of integrations. Most agents, including Claude Code and OpenAI Codex, reach external services through APIs or MCP connectors that someone had to build. Grok Bot drives the browser and desktop directly, so it works on software with no API at all. Inside SpaceXAI, the reported internal uses include a sales bot updating a CRM from call transcripts, an ops bot processing invoices from Gmail, and an engineering bot reproducing a bug, filing the ticket, and handing off the fix. Security teams should read the documentation before anyone expenses this. Every bot a user creates shares one cloud computer, one set of browser sessions, and one credential pool, and SpaceXAI's own docs warn against treating separate bots as a security boundary. There is no published architecture or safety documentation yet. Early user reports flag slowness and missed steps on simple tasks like newsletter unsubscribes. The product category is compelling and the credential model deserves a hard look from every IT department whose power users hold a qualifying subscription. SpaceX closes the Cursor acquisition The corporate story behind those bundled subscriptions resolved this week: SpaceX completed its acquisition of Cursor on August 14. Cursor announced it will join the SpaceXAI team to work on Grok, Grok Build, Grok Bot, the Grok API, and Cursor itself. The deal traces back to the April partnership that gave Cursor access to the Colossus training supercomputer, and it lands two months after SpaceX went public. The practical effects are already visible: Grok Build has defaulted to grok-4.6 since August 12 with up to 8 parallel subagents, and Grok 4.6 shipped day-one in Cursor with doubled usage for the first week. The most popular AI IDE is now a division of a rocket company, and its model roadmap is now Grok's roadmap. Teams standardized on Cursor with non-Grok models should watch how model routing and pricing evolve over the next quarter. Rails publishes its agent eval raw data The most useful tooling artifact of the week came from an unexpected publisher. The Ruby on Rails team added Grok 4.6, GLM-5.3, Gemini 3.7 Flash, and Claude Opus 4.8 to its agent benchmark and published all 792 raw run directories, every command, diff, and verdict included. Grok 4.6 was the best of the newcomers, completing 52 of 63 runs and landing just behind GPT-5.6 Sol, with frontier-tier Rails API recall at a $49 total campaign cost. The failure-mode analysis is the part worth internalizing. When Claude Fable 5 failed, only a quarter of its failed runs touched the files where the fix lives, so it failed by looking in the wrong place. When the GPT-5.6 models failed, nearly 80% of the time they found the right files and fixed them incorrectly. The Rails team flags the sample as small, but if the pattern holds, it changes how you review each model family's output: audit Claude's navigation, audit GPT's edits. Framework maintainers publishing reproducible agent evals with raw trajectories is exactly the norm this industry needs, and it puts vendor benchmark tables in their proper place. Standards: MCP Goes Stateless and the Clock Starts The Model Context Protocol's 2026-07-28 specification shipped three weeks ago, and this was the week adoption work got real. The headline change is that MCP is now stateless at the protocol layer. The initialize handshake is gone, the Mcp-Session-Id header is gone, and any server instance behind ordinary HTTP infrastructure can answer any request. The release also brings Multi Round-Trip Requests, header-based routing, cacheable list results, authorization hardening around OAuth and OpenID Connect, and a formal extensions framework covering MCP Apps and the Tasks extension for long-running work. Two adoption signals landed inside this window. Google Cloud published engineering guidance on scaling agent infrastructure on the stateless spec, walking through why the session-oriented design hit a hard wall in cloud-native deployments and how the removed handshake changes load balancing. And the Enterprise-Managed Authorization extension reached stable status, with Anthropic, Microsoft, and Okta adopting it so organizations can centrally manage authorization and end users can reach every connected MCP server through a single login. Repeated consent prompts have been the loudest enterprise complaint about MCP, so EMA adoption is the item to track. The scale numbers explain the urgency. The project reports close to half a billion SDK downloads a month across Tier 1 SDKs, with the TypeScript and Python SDKs each past one billion total downloads. Tier 1 SDK maintainers are expected to ship stateless support within the validation window, so if you operate MCP servers, your dependency updates over the next month carry breaking changes. Deprecated features from the old spec, including the legacy session model, need migration plans now rather than at the deadline. One adjacent standards note from the data world: Apache communities spent this week drafting the other kind of AI standard. Arrow proposed concurrent PR limits for non-committers after an uptick in unresponsive AI-generated contributions, DataFusion opened a policy discussion on LLM-generated PRs, and Iceberg debated norms for AI-generated review comments. Contribution governance for AI-assisted work is becoming a standard in its own right, written one dev list at a time. Infrastructure: The 2027 Memory Wall DRAM and HBM for 2027 are already gone The most consequential infrastructure news of the week fits in one sentence: 2027 DRAM and HBM capacity is reportedly fully allocated, a year and a half before that supply exists. Buyers are receiving only 60% to 70% of requested volumes and often paying deposits upfront. Adata's chairman estimates HBM and AI servers will consume nearly 70% of total DRAM capacity, and SK Group's chairman expects 2027 AI chip demand to rise 60% to 100%. Memory, not GPUs, is the binding constraint on the AI buildout, and every model provider's 2027 pricing already has this baked in whether they say so or not. The knock-on effects reach consumer hardware too, with AI server demand squeezing DRAM and NAND supply and pushing PC component prices up. China scales out with supernodes DIGITIMES' week-of-August-10 roundup puts numbers on China's alternative path. China's intelligent-computing capacity hit 2,185 EFLOPS in the first half of 2026, up 177% year over year, with 15th Five-Year Plan investment in the computing network potentially reaching CNY4 trillion, about $593 billion. The architecture bet is the supernode: combine more domestic accelerators with high-speed interconnects and system-level optimization to offset weaker single-chip performance. Guohai Securities forecasts the domestic supernode market growing from CNY88.9 billion this year to CNY1.109 trillion in 2028. Huawei's 18-tier pagoda system is the flagship example of the thesis that system architecture, not transistor shrink, drives the next phase of performance. In the same vein, Washington is reportedly preparing restrictions on Chinese-made optical transceivers used in AI data centers, extending export controls from compute into networking. The buildout leaves the ground SpaceX and Nvidia's Starmind program kept generating consequences this week after the August 4 announcement. Each Starmind satellite carries Nvidia Rubin GPUs and Vera CPUs, with peak power raised 67% to roughly 250 kW, enough for a full Vera Rubin NVL72 rack in orbit, cooled by 160 square meters of deployable liquid radiators. Prototypes target early 2027. Elon Musk declared SpaceX exclusive to Nvidia, and the more interesting detail for terrestrial buyers is his statement that the simplified NVL72 design built for orbit will deploy on the ground as well, because SpaceX considers it a radical simplification of the standard rack. The company is also building Terafab, a chip fab budgeted at $20 billion to $25 billion, to address its own compute shortages. Adjacent to all of this, Nvidia is reportedly closing in on a $100 billion credit guarantee deal supporting OpenAI's Ohio data center. The capital structures underneath the AI buildout keep getting stranger, and the week's OCP APAC summit in Taipei delivered the fitting summary: the bottleneck is moving beyond the GPU to power distribution, cooling, fiber density, and memory. What to Watch Next Week Watch for independent evaluations of DeepSeek's 0813 benchmark claims, the GLM-5.3 public API and whether open weights follow, the first enterprise security reviews of Grok Bot's shared-credential model, and Tier 1 MCP SDK releases landing stateless support. And keep an eye on memory pricing announcements, because the 2027 sellout will start showing up in 2026 contracts. If you want to go deeper on AI agents, data infrastructure, and how they fit together, from agentic analytics to lakehouse architecture, browse my full catalog of books at books.alexmerced.com.
This is a summary aggregated from Dev.to. Read the complete article on the original site:
Read full article at Dev.to