Podcast Episode
Gemini 3.7 Flash Halves the Price, Cerebras Hits 750 Tokens/Sec, and the Agent Harness Land Grab Begins
August 14, 2026
0:00
11:48
Google ships Gemini 3.7 Flash just three weeks after 3.6 Flash with big coding gains and a 50% introductory price cut, while OpenAI and Cerebras preview an Ultrafast mode running at up to 750 tokens/sec. DeepSeek and Arcee both open-source their internal agent harnesses, MiniMax drops an open-weights music model that runs on consumer GPUs, and a $40M raise plus a new custom-eval platform signal that benchmarking is becoming its own industry.
Google's Three-Week Turnaround
Gemini 3.7 Flash arrived barely three weeks after 3.6 Flash, and Google is positioning it as the default workhorse for coding, web development, knowledge work and agentic tasks. The reported jumps are substantial: DeepSWE from 49.0% to 65.3%, FrontierCode from 34.4% to 43.6%, and AutomationBench from 17.0% to 30.4%. There's an introductory 50% price cut through the end of the year at $0.75/$3.75 per 1M input/output tokens, doubling later. Independent testing put it roughly 4 points higher on aggregate intelligence indices, at around 340 output tokens/sec with a 1M context window, and it rolled out across the Gemini API, AI Studio, Android Studio, Gemini Enterprise and third-party coding tools within hours.The Speed Race Moves to Silicon
OpenAI previewed an Ultrafast mode for GPT-5.6 Sol powered by Cerebras, hitting up to 750 tokens/sec and around 14x standard speed for select API customers, aimed at voice, support, commerce and security workloads. It prompted an immediate observation that tool latency, not model latency, is about to become the real bottleneck in agentic systems. On the open side, Red Hat AI released DSpark, a speculator for Kimi-K3 delivering roughly 4x faster decoding, from about 110 to 435 tokens/sec per user. Prime Intellect shipped Prime Flash MoE, Blackwell-optimised CUDA kernels fusing routing-aware GEMMs, SwiGLU and quantisation.Harnesses Go Open Source
The day's most discussed release was DeepSeek Harness, open-sourced under MIT as a developer preview. The interest was architectural rather than benchmark-driven: composable plugins, multiple modes, visible trajectories, and KV-cache-aware append-only history semantics. Arcee open-sourced NAC under Apache 2.0, an internal harness for long-running asynchronous work that has already produced a significant share of code in its own training and data pipelines. Nous expanded Hermes into Bot Mode, where agent profiles become persistent named bots with routines, memory and bot-to-bot messaging. OpenAI added opt-in desktop activity as context, with a timeline view and user controls.Music You Can Run at Home
MiniMax released MiniMax-Music3 as an open-weights music model, an 8B language model paired with a 2.7B diffusion transformer that turns a prompt plus lyrics into full songs and runs on consumer hardware. The same company also took the top spot on the Video Edit Arena with a reported 32-point lead over the next best systems.Who Grades the Graders
Artificial Analysis launched Optima, a platform for building custom benchmarks on internal workloads from uploaded datasets, agent traces or plain-language descriptions, tracking quality, cost per task and time per task. Vals raised a $40M Series A at a $400M valuation and expanded into repo-derived coding benchmarks, an R&D capability index and a cyber evaluation suite.Uncomfortable Research
Three papers cut against the grain: skill libraries can actively hurt agents, with 307 failures attributed to loaded skills; context compactors retain only about 17% of persistent session constraints without a dedicated extractor; and leaderboard variance is dominated by agent-task interaction, with the agent main effect under 3% in several benchmarks.Six Labs Sign Up on Provenance
Anthropic, OpenAI, Google, Meta, Microsoft and Mistral all signed the EU Code of Practice on transparency of AI-generated content, with OpenAI stating it wants to extend provenance signals to all modalities including text. Sceptics argue text watermarks could be learned and stripped, or would interfere with code generation.Published August 14, 2026 at 1:56am