You're offline - Playing from downloaded podcasts
Back to All Episodes
Podcast Episode

Inkling-Small Goes Open, Gemini Robotics 2, and OpenAI's 80% Price Cut

July 31, 2026

0:00
13:04
Podcast Thumbnail

Thinking Machines ships Inkling-Small, a tiny-active open multimodal model, while Google DeepMind's Gemini Robotics 2 gives one AI brain full-body control over many robots. OpenAI slashes GPT-5.6 prices by up to 80%, Cursor reveals cloud agents now write 56% of its merged code, and a new study finds every major model quietly drifts toward Japan.

Thinking Machines Ships Inkling-Small

Thinking Machines released Inkling-Small, an open-weights, natively multimodal Mixture-of-Experts model with 276B total parameters but only 12B active at any moment. It claims performance close to the far larger original Inkling at roughly a quarter of the size, supports a 1M-token context in deployment stacks, and can run Python to inspect images mid-reasoning. The open inference ecosystem picked it up within hours, with day-0 support, single-server deployment, and local GGUF builds. Independent benchmarks put it around 40 on the Intelligence Index, with particular strength in coding and science.

Google DeepMind's Gemini Robotics 2

Gemini Robotics 2 is pitched as "one brain for any robot," moving beyond tabletop manipulation to whole-body humanoid control, advanced dexterity, and multi-robot collaboration. A companion model, Gemini Robotics ER 2, handles high-level embodied reasoning: observing, planning, tracking progress, and recovering from failed steps across multi-minute tasks. The same checkpoint reportedly controlled multiple hardware types and adapted to a new two-arm robot with fewer than 200 examples.

OpenAI Slashes GPT-5.6 Prices

OpenAI cut GPT-5.6 pricing aggressively, up to 80% on one tier and 20% on another, and added a Sol Fast option at up to 2.5x lower latency. The company attributes the savings to systems-level efficiency across the model, inference stack, and agentic harness, with automatic code review moving to the cheaper tier at roughly 10x lower cost.

Cursor's Cloud Agents Write 56% of Merged PRs

Cursor shared a striking adoption stat: cloud agents went from 1 in 10 merged pull requests in December to 56% today, achieved by giving agents their own cloud computers and letting them improve their environments. Similar reports describe engineers working entirely through cloud agents, including one building and testing a native iOS app inside a full macOS environment.

ARC-AGI-3 and the "Model Isn't the System" Debate

Much of the technical discussion focused on evaluation. ARC creator François Chollet stressed that long-horizon benchmarks now measure the whole agent system, memory retention, truncation policy, compaction, and tool orchestration, not just base weights. The same model scored wildly differently depending on its harness, in one case jumping from 7.8% to 38.3%.

The Productisation of AI Memory

Perplexity launched Projects, turning workspaces into persistent hubs with shared files and a memory layer nicknamed Brain. Infrastructure is maturing too, with one provider migrating 400M+ agent memories between databases for faster recall. But a cautionary study found filesystem-style memory can halve retrieval cost at scale without improving final answer quality.

Kernel Gains and a Curious Japan Bias

Under the headlines, systems work delivered real wins: a team more than doubled end-to-end performance on AMD's MI355X accelerators, and a custom attention kernel jumped past 2x over the standard approach by swapping in a faster approximate instruction. Finally, a study across 31,680 prompts and 24 languages found major models disproportionately drift toward Japan after fine-tuning, most likely a reflection of lopsided training data.

Published July 31, 2026 at 2:23am

More Recent Episodes