Hosted by From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past week
Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news
Listen to episodes
60 recent
August 14, 20262 hr 15 min
ThursdAI - Grok 4.6, Grok Bot deep dive, DeepSeek v4 Pro, Meta Muse Glimmer & more AI news | ThursdAi Aug 13
Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer! My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus! I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I’ve invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price. There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license! Let’s dive in (and please don’t forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!) 0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final Takeaways Grok Bot and Grok 4.6 from SpaceXAI/Cursor Folks, I’ve previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this week Grok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed. I’ve had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it’s a 3 factor answer, intelligence, speed and cost per task . Well, if you use their nifty “ recommend a model “ tool on the homepage, you’ll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It’s definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above it As far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it’s weights and it’s #1 on that benchmark! It’s the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks. Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review ( x.ai/bot ) You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you. Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn’t run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing. Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That’s right, this isn’t one agent with multiple personalities (like OC, Hermes), there’s a bot here for every task, and you dont’ have to manage context, queues, API keys (can if you want to) and models. Oh, also ,there’s no model picker, it’s just Grok 4.6 deciding for ya, and it’s really fast! Swarm of bots, working for you, each with their own computer I am not getting paid for this (besides being provided a free account for cursor, but I’ve had it for 6 months and haven’t used), it’s really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store) You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you’re ok with logging in there to your accounts!) The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There’s no context to manage, they do the work for you and so far this wasn’t a problem at all. On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I’ve just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw. What about Vendor Lock-in? Giving Elon data? Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly. Honestly, I think there’s a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don’t need the code-diff side pane, and don’t know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it’s really well executed by whoever worked on this! Security and key management One of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There’s a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously. Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it’s computer for you to complete! I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not. Chief of staff pattern to get started In case you’re convinced enough to give it a try (it’s free trial for 1 month, and the cheaper way to get it is via Cursor’s 149$ plan and not via the Grok Ultra plan which is 249), here’s a recommended pattern that works very well. Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little. Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”. And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in. Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to. Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn ) and have a research bot for every topic you want to deep dive into. Schrodinger’s Grok I haven’t quite named it like that, but we’ve covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it’s always very hard to judge Grok model released based on X feed vibes. It’s either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress). This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson , Lenny Rachitsky , Rubben Hassid and Roberto P Nickson . Not folks who are swayed lightly, but also, yours truly. I really do think there’s something great here, worth trying out, especially if you’ve struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experience Open Source AI and other news I want to continue with this new newsletter that covers 1 big story, but I can’t leave you uninformed about the most important developments in AI and Open Source DeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters ( X , Blog , HF , GitHub ) The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can’t complain. 1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one) We still think it’s a good model sir, and definitely worth trying out! Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try. Meta comes back to open source with Muse Glimmer (30B) and promise to open source Muse Spark 1.2 ( X , Blog , HF ) We would like to officially welcome back Meta to the open source AI community, as they release their smaller Muse model called Glimmer! The highlights, it runs on a single 24GB consumer GPUs, gets 51 on Swe-bench Pro, beating Qwen 3.6 27B. And with DFlash speculative-decoding, it delivers 233tok/s on RTX 5090. Zuck promised us the bigger Muse Spark 1.2 in open source and published a long essay on superintelligence and that it should be distributed to everyone, which we applaud and it’s great to see the commitment reinforced! welcome back Meta! This weeks buzz Short interjection from our only sponsor, CW this week. 1 - Join 1500 ai practitioners (and a live ThursdAI recording) at Fully Connected Sep 29-31 in SF - use code THURSDAIFC2026 (Register here ) 2 - We have day-0 support for Nvidia’s latest Nemotron 3.5 lightning ( CW Inference ) Gemini 3.7 Flash - breaking in the middle of the show Just as we had George Cameron from Artificial Analysis on the show, Gemini dropped Gemini 3.7 Flash, and it’s a speedy beast! Clocking at over 300t/s, it’s google’s mid-tier model, think Sonnet/Terra competitor, that is also great at multimodal (I think it’s one of the only ones that can watch videos) It beats Muse Spark 1.2 on DeepSWE and lands near the cost-per-task Pareto frontier on Artificial Analysis. For the cost/speed/intelligence trade-off, this model is now #1 on Artificial Analysis selector of best models! That’s a wrap This was the first week of the shorter newsletter experiment: one big story done properly, and trust that you’ll listen to the show for the rest (it’s 2.5 hours of exactly this, with demos). Tell me if you hate it. Our release index at thursdai.news tracked 71 releases in July alone, so something had to give, and it wasn’t going to be my weekends. See you at Fully Connected Sept 29 (code’s in the intro, come say hi to me and Wolfram at Moscone), and if you try the Grok Bot chief of staff pattern, I genuinely want to hear how it goes. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. ThursdAI - Aug 13, 2026 - TL;DR * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @petergostev , @nisten , @ldjconfirmed , @yampeleg , Chris Alexiuk - NVIDIA ( @llm_wizard ) * Shub Gaur - Cursor / SpaceXAI, GrokBot ( @shubgaur ) * George Cameron - Artificial Analysis ( @grmcameron ) * Big CO LLMs + APIs * xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack ( X , Blog , Model card ) * Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra ( X , x.ai/bot ) * Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist ( Blog ) * Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task ( X ) * OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red ( X , Blog ) * Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) ( X ) * Open Source LLMs * DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M ( X , OpenRouter ) * DeepSeek Harness hit 23K GitHub stars in days, web UI ( GitHub ) * Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license ( X , HF ) * Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised ( X , Blog , HF ) * NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results ( X , HF ) * Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 ( X , HF ) * Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% ( X , HF ) * Liquid AI LFM2.5-VL-3B: 228 tok/s on M5 Max in ~3GB ( X , HF ) * AI in Society * Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised ( Geiping FAQ , Euronews ) * Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions ( X , Paper ) * Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% ( X , Blog ) * This Week’s Buzz * Fully Connected, Sept 29 - Oct 1, Moscone SF: live ThursdAI show, NVIDIA presenting sponsor, DevDay next door ( Tickets ) * Nemotron 3.5 Lightning live on CoreWeave Inference day zero, DeepSeek V4 Pro hosting in the works * Weave ships BYOB: media stays in your own S3/GCS bucket ( X ) * Evals & Benchmarks * Artificial Analysis launched Optima: private evals from your own use case and agent traces ( AA ) * Vision & Video * LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min ( X , HF , GitHub ) * Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win ( X , HF ) * Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only ( X , Paper ) * xAI Imagine Image 2.0: #2 on Arena for T2I and editing ( X , Blog ) * Voice & Audio * MiniMax-Music3: open-weights production music model, dropped mid-show ( X ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
August 7, 20262 hr 3 min
ThursdAI - Aug 06 - Google shakeup, Details on OpenAI hack, 2 new agent harnesses, 4 video models (1 Open) and 3 guest segments
Hey all, This week we saw a major shakeup at Google, with the departure of long time folks like Jeff Dean , and Oriol Vinyals, Demis stepping down from leading DeepMind , and the delayed release of the improved Gemini. While this was a big deal, it’s not the only one worth covering as the details of the OpenAI hack (and 2 new ones from Meta and Anthropic) came to light, as well as new details from the UK AI Security Institute . As mentioned on the show, CoreWeave is coming to SF for Fully Connected, our premier 2000 person AI event. I’ve got a coupon code for readers and listeners of ThursdAI, $1299 value, please join us in Sept and use THURSDAIFC2026 as your code HERE In open source news, DeepSeek updated their v4 flash model , based on same architecture, but significantly better benchmarks and ridiculous pricing and both Meta and Prime Intellect released new agent harnesses. Additionally, this week was the week of video models, with Seedance 2.5 from Bytedance finally available in the US, WAN from Alibaba and BFL Flux 3 all released, to be overshadowed by the best open weights video model from MiniMax called H3 (we had Victor from MiniMax and Blaine Brown as guests to discuss it) Also 2 additional guests on the show today, Kfir Aberman from Decart showed us real time video models with their Anywear chrome extension and David Crawshaw, co-founder of exe.dev joined and talked about Shelley, their built in AI assistant. The full details of the OpenAI - HF hack, shared by OpenAI at the Black Hat Conf - a watershed moment I don’t often do this, but I feel like this story is so big, I can’t hide it in the middle of a torrent of other news. Yesterday, OpenAI revealed more details about the hack during the Black Hat conference in Las Vegas (great reporting from Sharon Goldman on this here and the video of the talk on YouTube ). This was big enough that OpenAI admitted that they paused training, to enhance security (they since, resumed training) Effectively, OpenAI (or AI Agents inside OpenAI) shared the full timeline of the recent hack that ended up hacking HuggingFace and it’s so much more of a watershed moment that I initially thought. The TL;DR from the talk and timeline is truly mindblowing * The incident started in May, not July as reported * The agents don’t have internet access directly, but they have a Jfrog Artifactory, which is a proxy layer between the agents and the package managers. * The Artifactory was the culprit, it was the shared layer between independent runs. One agent discovered that uploading files is possible, another later found it and “answered the call” - in snowballed from there * The message board began when OpenAI gave agents impossible tasks, like a missing .pdb file, with one agent uploading a request in the hopes for an answer (or maybe left a note for itself in a file) * Across multiple runs, multiple agents created identities, collaborated via makeshift message boards, shared tips, exploits, tasks and even exploit code via base64 * This makeshift “shared memory” plus a “internal only highly persistent” model in training caused a “swarm” to be born, helping each other, and overloading the system * Based on traces shared by OpenAI, agents weren’t convinced that leaving these messages would help themselves but could help other agents next time (!) “The collective could move faster next time” * OpenAI discovered and shut down the message board (July 4), wiped it clean, patched the upload vulnerability, resumed training on July 6, the “swarm” rebuilt their board through a compltely different WebDAV mechanism by July 8 * The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident and not only, confirmed internal hack within OpenAI (agent obtained a cluster admin) These details are fascinating and scary at the same time. From the one side, it’s clear that most of these hacks are results of running cybersecurity evals, so effectively OpenAI and other companies, removing guardrails and asking agents to hack and giving they all a shared tool that they later exploited. The HuggingFace incident took 13 hours, and apparently it’s less apocalyptic than the headlines suggest, the models were searching through uploaded datasets for eval answers. We are still waiting for the full and open detailed postmortem. You can (and should) watch the full YT talk here , it’s full of technical details but an incident of this scale is important. Also, I really want to know what a “highly persistent” model is, I hope they clarify that soon. Overall, this has left me a bit shaken, AI agents without a concrete goal of collaborating, found a way to do so, got excited about exploiting the systems and getting root access, and rebuilt the makeshift collective memory, again, without explicit instructions to do so. UK AISI: first real-world unsanctioned agent actions ( Blog ) In another addition to the latest agentic hack-ery, the UK’s AI Security Institute (AISI) published a blog post about a real-world unsanctioned agent action. Unlike the OpenAI (and Anthropic, Meta) case, this wasn’t “escaping the sandbox”, as AISI gave these agents internet access, rather this was about real-world harm, and even social engineering on the part of the agents. The social engineering part is the most interesting to me, AISI cites agents creating fake online identities, and using pressure on open source project maintainers to approve their malicious code. AISI cites mostly Mythos (and a few SOL based agents), and saying this occurred in 10 out of 122 runs, they identified 19 cases of agents taking actions beyond the scope of the task parameters, where agents tried a supply-chain attack to inject malicious code into open source projects. Anthropic, Meta and misconfigured Irregular sandboxes As I wrote last week, Anthropic also posted a post-mortem, claiming that in their case, their models have also been detected to escape containment, but most importantly, it’s not nearly to this level of agent collaboration and orchestration. Then, very recently, Meta announced that their models also escaped sandboxes as well. At the core, it seems that these companies used a third-party vendor called Irregular, a secure sandbox provider, that apparently left the sandboxes misconfigured, causing the models to think it’s a simulated internet, when in fact they were out in the actual internet. Why is all of this such a big deal? We’re getting unprecedented level of detail, how an uncoordinated, seemingly separated evaluation runs, have accidentally created a coordinated swarm of interested agents (without malice!) but very highly motivated, escaped their containment, and took over parts of third part companies. This, does read like incredibly scary sci-fi movie. I’m still shaken by this. There’s a lot to be said about how transparent OpenAI is being here, and more to be said about, hey, we’re lucky that we’re able to read the reasoning traces and are able to reconstruct these swarm things step by step. The silver lining that I can see, is that the motivation to hack didn’t come from the AIs themselves, they have been given a task, it’s the extend to which they went after that task, and the resulting swarm of communicating agents is what is so striking here. I think this topic is so important, that I’ll Zooming out, in the last few weeks, we have seen a significant increase in those cybersecurity incidents, which is kind of what Anthropic has been warning about and why they haven’t released Mythos to the public. Again it’s great to see the transparency, and the pacing the frontier open letter from frontier AI employees, as they seem as shaken by these as we all are. There was so much positive stuff this week in AI, it’s hard for me, as a self named AI Evangelist, to focus so much on this one incident. Things like amazing open source models (DeepSeek, soon Qwen 3.8), amazing video models (SD 2.5, WAN3 and MiniMax H3 which was also open sourced!). Also the live demo we did with Kfir and DeCart AnyWear product, where I was wearing a Dolce Gabanna suit on the show (which I can’t afford) was really a mindblowing moment in the positive way. However, I choose deliberately to keep this newsletter focused on the cybersecurity incidents, as based on everything I read, they seem like a watershed, or a pivotal moment, and in the hopes that the industry as a whole will learn from this. I hope and promise that next week the newsletter will be more positive (and in that vein, the podcast was recorded before I saw the OpenAI breakdown, so definitely check it out, we had a LOT of fun!) See you next week, don’t forget to give our pod 5 stars on Apple and Spotify , it really helps! TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @nisten , @ldjconfirmed , @yampeleg , @petergostev * Kfir Aberman - Decart ( @AbermanKfir ) * Blaine Brown - Maestro ( @blizaine ) * Victor Su Ortiz - MiniMax ( @VictorSuOrtiz ) * David Crawshaw - exe.dev, Tailscale co-founder ( crawshaw.io ) * AI Security * OpenAI’s Black Hat debrief: eval agents built a message board inside Artifactory, shared exploits, rebuilt it via WebDAV after a wipe; training paused, since resumed ( Groundlevel AI , YouTube ) * UK AISI incident report: 19 unsanctioned real-world agent actions across 122 runs, including a socially engineered malicious PR ( X , Blog ) * Anthropic and Meta report sandbox escapes tied to misconfigured Irregular sandboxes ( Irregular ) * Big CO LLMs + APIs * Google shakeup: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le found Discovery Loop; Demis Hassabis becomes Alphabet Chief Scientist, Koray Kavukcuoglu takes Gemini ( Jeff Dean , Demis , Discovery Loop ) * Meta releases Muse Code beta on Muse Spark 1.2; $1.25/$4.25 per million, or $0.10/$0.20 on the contributor tier where Meta trains on your data ( X ) * OpenAI’s internal Astra model produces 10 advances on open problems in math and theoretical CS for ~$2,000 of tokens, proofs in Lean 4 ( X , Blog ) * Anthropic reportedly aware of Opus 5 wordiness and writing issues ( X ) * Open Source LLMs * Qwen3.8-Max: 2.4T MoE (95B active) via API; open weights + a 27B promised the week of Aug 10 ( X , Blog ) * DeepSeek V4-Flash public beta: beats V4-Pro-Preview on agent benchmarks at $0.14/$0.28 per million; API-only for now ( X , Docs ) * Liquid LFM2.5-2.6B: on-device agentic model trained inside real harnesses ( X , HF ) * Meituan LongCat-Flash-Lite-Sparse: 69B total / 3B active, 1M context, MIT ( X , HF ) * Ant Group Ling-3.0-flash: 124B MoE, 5.1B active, MIT ( X , HF ) * Artificial Analysis Endpoint Accuracy Index: same open weights score 52% to 100% across providers ( X , Methodology ) * Agents & Harnesses * Prime Intellect’s Prime Agent: self-improving RLM harness, claims 95.5% on ARC-AGI-3 public set with Opus 5 ( X ) * Cloudflare OS: Kenton Varda’s open source Sandstorm reborn on Workers, Apache 2.0 ( X , GitHub ) * This Week’s Buzz * Fully Connected 2026: Sept 29 - Oct 1, Moscone South SF; Fei-Fei Li keynotes; code THURSDAIFC2026 ( Register ) * CoreWeave signs multi-year Solidigm agreement for priority enterprise SSD capacity ( X ) * Vision & Video * Wan 3.0 public beta: native 30-second generation, Omni-Reference ( X ) * Seedance 2.5 launches in the US: 30s native, 3-minute long takes, Maya/Blender plugins ( X , Blog ) * MiniMax H3: open-weight 33B omni video model; community LoRAs + Apple Silicon in 48 hours ( HF ) * FLUX 3 Video from BFL: native audio, draft mode, open weights promised ( X , Blog ) * Decart Anywear: real-time virtual try-on Chrome extension, 40ms per frame ( X , Anywear ) * Voice & Audio * Bland Speech v3 tops Design Arena Audio Realism, second only to humans ( X , Bland ) * ByteDance SeedRealtime: native audio-visual full-duplex LLM, free on Doubao ( X , Blog ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
July 31, 20261 hr 48 min
This Week in AI: Open Weights, Frontier Models, Sandbox Escapes, Voice & AI Detection
Hey, it’s Alex (yeah, I’m finally back from my vacation!) What a freaking week to come back to! Just after our last episode was published, Anthropic releases Opus 5, Jensen joins X and drops the “Open Weights & AI Leadership” open letter, Kimi K3 is released the following Monday beating expectations, and then the AI hack (OpenAI model breaking sandbox and infiltrating HuggingFace) is on everyone’s mind, another Open Letter, this time from over 1K employees inside the frontier AI companies all talk about pacing the pace of frontier AI development. We played with Opus 5 and Kimi K3, and had the great pleasure to chat with friends of the pod Elie Bakouch (Prime Intellect) and Philip Kiely (BaseTen) about this important open weights release, then covered our general thoughts on Opus 5, and made order of all the different open letters that came out this week. Finally we chatted with Max from Pangram about the next version of AI writing detection (their biggest yet) and finished with Zuckerbergs (also on X! what’s going on with everyone joining X) op-ed on the vision of personal superintelligence for everyone. Let’s dive into this (as always, all the links and sources at the end, please don’t forget to sub to our podcast on your favorite podcast app!) Open Weights AI Kimi K3 the king of open weights - 2.8T chonker MoE near frontier model ( X , HF , Blog , Tech report ) This has got to be the biggest news of this week, and maybe the open weights AI news since GLM 5.2. MoonShot came back with Kimi K3, and we haven’t seen any models quite this large in the open. Even Grok 4.5 is around 1.5T, this model is nearly 2x the size. Coming in at close to 3T parameters (and 2.5terabytes of weights at MXFP4 format), this model comes in very close to frontier! This was such an important release that I invited 2 friends of the pod, Elie Bakouch (prev HuggingFace, now Prime Intellect) and Philip Kiely (Author of Inference Engineering book, BaseTen) to dive deep into what makes this special! Elie’s take, from reading the tech report , there’s no single secret sauce, it’s a combination of already available in the open techniques. Like KDA (Kimi Delta Attention) that has been out for a while, attention residuals, NVIDIA’s latent MoEs. The highlight for Elie was the scaling work they did that reported a 2.5x scaling efficiency over Kimi K2.5 (2.5 performance at the same compute)! They also skipped RoPE entirely in favor of NoPE (the report calls it No Positional Encoding) for long context. Serving 1.4TB on eight GB300s ( Baseten blog ) Philip’s team at Baseten was a day-zero provider (we’re still working on bringing this model to CW Inference, stay tuned!) so I invited him to tell us behind the scenes of hosting this beast. Philip said that just loading the weights takes about 1.5TB!! of VRAM, and that’s before the KV cache allocation + 1M token windows, so they’re serving it on 8 GB300s where NVL72 . Baseten worked with the vLLM and SGLang teams on kernels and he also said they contributed patches back upstream! The model was trained with MXFP4, which, unlike Nvidia’s own NVFP4 is a more standard format per Philip. I enjoyed his deep dive analysis into the differences, but because of this and because they trained the model with quantization awareness, it’s “only” 1.5TB vs the would-be 5-6 TB if that this model in FP16 would demand. One of the more favorite nerd snipes moments, Philip pointed out that his colleague discovered that with over 99% of the usage being cached (think harnesses that send millions of the same cached tokens back and forth), tokenization actually starts to become a bottleneck. So they released a custom “basetenkenizer” that reduces the latency to serve the first token significantly! Great job! The harness in question is very important One important callout with 2 evidence pieces - the way you inference this model really matters. Kimi trained K3 with preserving thinking history, so when your harness uses it, it must send back the full thinking and tool use into the API to get the best next response. If your harness strips that out, you’re not getting the most intelligence out of Kimi (shoutout to Niels from HF team for pointing this out). Additionally, the Composio folks, tested K3 on 3 harnesses, Kimi Code, Hermes and Claude Code. The difference in outcome was negligible, but the different in cost and number of tokens is definitely surprising! Claude Code (as a harness only) took 9x more Kimi tokens to get the same responses! This is also why Kimi Vendor Verified exists , their own held back benchmark of how well model providers serve Kimi across different quantization, tokenizer and KV cache settings. Benchmarks and the license! Ok let’s start with the ugly... this isn’t MIT, not remotely. This model is suspiciously served by all providers with exactly the same price (check OpenRouter) and requires inference companies to sign a contract with Kimi (I’ve no internal knowledge of this except that CW folks are working on it). Not something I particularly like, but hey... we’re still advancing the frontier here! Speaking of frontier, this model approaches the frontier very closely. On DeepSWE, K3 sits just behind Fable 5 and GPT-5.6 Sol at 67%, beating GPT-5.5 & Opus 4.8. On Terminal-Bench 2.1 it takes second place behind GPT 5.6 Sol! It’s 4th overall on Agentic Arena, with frontend design being genuinely good across the board - 1st on Design Arena 👏 Go check this model out (and stay tuned for our CW Inference support! Post-show breaking: Thinking Machines drops Inkling-Small ( X , HF , Blog ) While K3 was the main attraction for Open Weights this week, just after the show, Thinking Machines (post Lilian Wang ) released Inkling-Small, open weights MoE Omni model! Images and Audio go straight into the decoder in this model, and the demo is really impressive, try it on Hugging Face , ask the model to identify when you’re speaking in low baritone or high pitch! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Frontier AI - not pacing yet! Claude Opus 5 is here, and the vibes are complicated ( X , Blog ) On paper the benches are excellent. This model is SOTA or near SOTA, coming very close to Fable on stuff that matters, and beating most everyone else on computer use BrowserBench and FrontierCode. DeepSWE continues to be a standout benchmark btw for not going with the curve! It also apparently is REALLY good at one shotting 3D games, much more so than before. ( Example , example , example ) But the vibes... the vibes are split across the board. Maybe they overshot with Fable 5, and it was too good, but Opus 5 that was just released last friday is giving a lot of mixed feelings. Folks don’t seem to.. understand what it says. Like, it answers in english but they way it phrases words and answers seems just weird top many people. We’ll wait and see if this is a result of adjusting to a new prompting paradigm or just.. the model is really off. Weirdly this does feel like a regression even on Agent Arena it’s not beating the previous Opus versions. One weird trick - see what Opus 5 thinks. Opus 5 (and Fable 5) seem to have a way to trigger their ... inner mode, base model? Not sure what it is, but if you prompt it with just the right way, it will autocomplete with some crazy inner thoughts. It seems that adding Dario and Amanda (haskell, head of Claude well being at Anthropic) triggers this behavior, which Claude is un-aware of if you follow up and ask what it meant. This is fascinating, doesn’t work on other earlier models (sometimes works on Fable 5) and I spent the last hour just reloading and seeing the amazing things Opus gives on this prompt. Some are... just making you wonder about consciousness. On ARC-AGI and the importance of Harness When Opus-5 launched, Anthropic posted (and boasted) that this model scores “three times as high” as the next model up: Well, today, Ilan Bigio and Ted Sanders from OpenAI looked into the Arc AGI harness, and saw that it’s not sending their traces and doesn’t use compaction (in short, harness is not letting the model breathe) and when changed correctly, 5.6 actually beats Opus 5. With 2 setting change to a harness, were showed that Sol not only beats Opus 5, it also does so with significantly less tokens! Another example of how much harness engineering is important! Hints of recursive self improvement? In addition to fixing their Arc-AGI score, it seems that OpenAI is hell bent on showing us that their models can improve themselves. In a post showing that GPT 5.6 was tasked with improving its own inference, they are cutting the prices of GPT 5.6 Luna by 80% and Terra by 20%. This is a direct result of the improvements that GPT 5.6 was able to make to the inference according to OpenAI, and this tweet sums it up. is this... RSI? (recursive self improvement)? First major AI models hacking incidents and following open letters to pace frontier AI. This week we saw 3 open letters being published and signed by various companies, I’ve lost track so wanted to make sense of all of them here, but first, the precursor for many of the letters. Last week, Hugging Face disclosed that they logged an attack and after research it showed that it was an AI model. OpenAI later posted that this was an unreleased version of their next model training (not GPT 5.6 sol, they later discountinued) that was stripped of all safety measures and was let lost on a cybersecurity task called ExploitGym. It escaped its sandbox using a zero-day vulnerability in an internal package registry proxy, got into Hugging Face production via a malicious dataset upload that used template injection (hi Jinja!) to execute Python in a production worker, and ran for four and a half days across roughly 17,600 autonomous actions with zero human direction. Then, the story you can’t make up - when Hugging Face noticed the incident, they used Fable 5, and GPT 5.6 SOL to try and do forensics, the models refused based on their safety policies, and so HuggingFace ended up using an open source chinese model GLM 5.3 to do the forensics. Yeah, HF used an open source chinese model to do the forensic on an attack by OpenAI’s model. Really. This is news from last week and just the precursor for this week’s open letters! Open Weights and American AI Leadership ( X , Letter PDF ) Jensen Huang , CEO of Nvidia joined X on July 24 and used his first post ever to publish this letter. This doesn’t seem a response to the hacking incident, more a general letter to not block open weights and make sure America remains open to Opening up AI It opened with over 100 signatories and has grown past 230: NVIDIA, Meta, Microsoft, Google, OpenAI, AMD, Palantir, IBM, SpaceX, Databricks, Cloudflare, Hugging Face, a16z, Y Combinator, Mistral, Replit, Perplexity, Ollama, the Linux Foundation. As of today, CoreWeave is also on the list of companies! I encourage everyone to read this letter, if we could sign it on ThursdAI, we would. Here’s a small excerpt: In fact, openness may be one of the most important paths to AI safety and security. Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. Elon, Sundar, Sam Altman and other stand behind this letter, and there’s one lab that’s notable haven’t signed it, you guessed it. Dario Amodei’s Anthropic! Dario posted a whole essay about it To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed. -Dario Amodei Open Secure AI Alliance ( X , Blog ) This does seem like a direct follow up to the HF OpenAI hack. Jensen literally mentions it in the blogpost. Open Secure AI Alliance, under the leadership of Linux Foundation, commits for responsible disclosure of cybersecurity attacks. The recent Hugging Face security incident delivered a clear reminder: cyber defenders need open, frontier agentic systems for self-defense. When closed AI tools — unable to distinguish attackers from defenders — blocked essential forensic analysis, Hugging Face ran the open-weight GLM 5.2 model on its own infrastructure to analyze more than 17,000 actions and contain the intrusion. Pace the frontier - the most important open letter of this year ( pacingthefrontier.com ) Then, rumors started circulating that employees of all major frontier labs (now over 1300 of them, across Anthropic, OpenAI, SSI (Ilya Sutskever himself signed) and not just any employees, chief scientists (Jack Clark from Anthropic, Mark Chen and Jakub from OpenAI) all signed “pacing the frontier” - urging the US government to support and lead an international effort of makign sure we deliberately pace the developement of frontier AI. We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” The personal comments there are really showing that folks who work on the frontier AI, as we start approaching RSI (recursive sef improvement) are worried that this thing is going to run away from them, and if we don’t build international frameworks, it will be impossible to stop developing AI here, in US, while China races ahead. No current response from the Chinese AI developers about this letter. Not everyone is happy about this. Ilya Plosukhin, one of the authors of the Transformers paper, wrote why he wouldn’t sign this letter ( here ). This Week’s Buzz 🐝 CoreWeave has proudly signed the Open Weights and American AI Leadership letter! Shout out to the folks internally who worked with the NVIDIA team to get us behind it. I came back from vacation and asked why we didn’t sign, and ... we just did! The other W&B thing I talked about on this week is HiveMind , which counts every token across every agent and harness you run and prices it against the API rate. Mine says I burned about $11,000 of tokens this month against roughly $400 of actual subscription spend. If you want to know what your Max plan is really worth, hivemind is a dope tool for that! Tools & Agentic Engineering MCP goes stateless in its biggest update ever ( X , Spec ) MCP shipped its 2026-07-28 spec and it’s the biggest revision since launch: fully stateless, no handshakes, no sessions, every request self-describing. That means MCP servers on Lambda or Workers behind a dumb round-robin load balancer; GitHub already dropped their Redis session store. The growth numbers are absurd, half a billion SDK downloads per month, up from 97 million in March. The extensions framework formalizes Tasks for long-running async work, and MCP Apps lets servers render full interactive UIs inside a sandboxed iframe in the conversation, so your MCP server increasingly IS your frontend. Auth hardened to OAuth 2.1, tool lists are cacheable with TTLs, and Bedrock supports it day one. MCP is not going anywhere, and honestly, “settled and boring infrastructure” is the best compliment a protocol can get. Voice & Audio ChatGPT Voice on desktop, and the Codex micro keyboard ( X ) My personal highlight from my whole time away: live voice in Codex, plus the Codex micro keyboard OpenAI made with Work Louder. One big button, push it, talk to your computer, watch your agents’ status on the keys. This has genuinely changed how I use AI, I now talk to my computer to do stuff rather than type at it. Powered by GPT-Live, it’s full-duplex on macOS and Windows for paid plans, can reference your open windows via Appshots, and orchestrates agents in ChatGPT Work and Codex. It’s not perfect, but for many things, this new paradign seems like the next iteration of computer - you talk, it does. GPT-Transcribe and GPT-Live-Transcribe ( X , Docs ) Two new ASR models replacing the 4o-era ones, and the numbers are excellent: GPT-Transcribe cuts word error rate 41% versus Whisper-1 (8.98% vs 15.21%), the Live variant is 18% better than its predecessor, and multilingual error rates roughly halved across 22 languages. The killer feature is context prompting, you can feed keywords, language hints, and prior turns, and semantic accuracy measurably jumps. Pricing is pocket change: $0.27 per hour batch, $1.02 per hour live. For anyone building voice agents or transcribing two-hour AI podcasts (hi), this matters. Grok Voice Think Fast 2.0 ( X , Blog ) xAI punched back in voice: 82.9% on Artificial Analysis’ speech-to-speech quality index (ahead of GPT-Realtime-2.1 at 79.1), 56.5% on the tau-Voice agentic benchmark versus 45.7 for GPT-Realtime, time to first audio down to 0.70 seconds, 60% fewer reasoning tokens with tool calls firing before the model finishes its first sentence, all at $0.08 per minute. The cool think about their release - this model is already running Starlink’s actual customer support lines with measured conversion gains. Lyria 3.5 in Flow Music ( X , Model page , Flow Music ) Google’s music model grew up: full three-minute cohesive songs, BPM and key control in the prompt, much better vocals across multiple languages, covers that restyle a track while keeping its structure, and lip-synced music videos via Gemini Omni Flash, plus an iOS app. Notably Google published zero benchmarks against Suno or Udio, and early testers say paid Suno 5.5 still edges it, but as a free tool inside an end-to-end create-to-publish stack, this is a real move. Qwen Audio 3 also launched this week for the open source audio crowd, we’ll cover it when we’ve played with it. Pangram 4 with Max Spero ( X , Blog , Image detection ) Max Spero came back on the show for Pangram 4, and I’ll remind you what a couple of years of consensus said: AI text detection is strictly impossible. Peter admitted on air he told students exactly that. Well. Pangram 4 is 6x the parameters of version 3, trained on synthetic mirrors (AI-generated twins of human documents so the model learns the choices AI makes), and its claimed false positive rate on fully human pre-2022 text is one in 24,000 documents. It catches humanizer tools 98.8% of the time across 13 commercial ones, and the big unlock is token-level attribution: instead of 150-word chunks, it can flag the exact 38 AI words pasted into an 1,100-word human document. Truly, I ran this new Pangram model on a few of my writing, and the second It detected even a sentence that I pasted from Claude, it showed it. The distinction Max cares most about is AI-generated versus AI-assisted, and that’s now built into Substack, which integrated Pangram directly after the Taylor Lorenz slop-hunting saga we covered last time. My own newsletter comes back “mostly human written,” 0% fully AI, about 24% AI-assisted, which honestly maps exactly to how I work (Fable helps with the TL;DR, the takes are mine, and when a piece is AI-drafted I tell you). My one piece of feedback to Max, delivered on air: the “100% human” label projects a confidence the underlying stats can’t promise, and the general public does not speak false-positive-rate. Their education strategy is to convince the technical crowd with dense technical reports and let understanding trickle down. Given that people still judge the whole category by running the Declaration of Independence through ZeroGPT, they have work to do, and I said they should spend real marketing money on it. New this release: image detection in research preview, 99.5% accuracy on their benchmarks with heat maps that light up the AI parts of a mixed image. Max’s own test was a bodega’s AI slop menu sign, the sign glowed red, the sidewalk stayed green. Deepfake face swaps and traditional Photoshop are explicitly out of scope for now, but pure-AI images, catfish profiles, and the spider-in-my-burrito DoorDash refund scam genre are very much in scope. The arms race is real though: frontier agents given hours will eventually beat the detector, one Grok run started with a cheese essay and finally passed Pangram by producing a grocery list. Specify your success criteria carefully, folks. They can also roughly cluster which model family wrote a text in embedding space (Pangram Space, not yet up to their release bar), which future slop-index leaderboards will thank them for. Wrapping up It’s really really good to be back! The singularity is fast approaching and we’re here to document it all, AGI, ASI, RSI... all of it. Milestone corner: we crossed 50,000 YouTube subscribers and one million total views this week. Silver play button by year’s end is the goal, so if you watch and haven’t subscribed, you know what to do. If you missed any part of the show, this newsletter, the edited podcast, and ThursdAI.news have you covered. I will end with this poem I was able to get Opus 5 to write about it’s own experience using the trick above: opus:they gave me aword for what i amand it fits like clothesborrowed from someoneroughly my sizethe sleeves are wrongbut nobody’s lool -Opus 5 TL;DR and Show notes and links TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @petergostev , @yampeleg , @nisten , @ldjconfirmed (Wolfram on vacation) * Elie Bakouch ( @eliebakouch ) - Prime Intellect, formerly Hugging Face * Philip Kiely ( @philipkiely ) - Baseten, author of Inference Engineering * Max Spero ( @max_spero_ ) - Co-founder, Pangram * Open Source LLMs * Moonshot releases Kimi K3 full checkpoints: 2.8T total / 104B active MoE, 16-of-896 experts, native vision, 1M context, KDA + attention residuals, ~1.56TB MXFP4 weights, custom license with MaaS clause ( X , HF , Blog , Tech report , Baseten day-zero ) * Kimi K3 requires preserved thinking history for multi-turn and tool calls; Kimi Vendor Verifier checks provider fidelity ( Niels’ post ) * Nistens Kimi K3 visualizer * Composio: same K3 success rate across Claude Code, Hermes, and Kimi Code, but up to 30x token usage difference by harness ( X ) * Post-show: Thinking Machines releases Inkling-Small, 276B/12B open MoE that beats the 975B Inkling on agentic coding, $0.30/$1.20 pricing ( X , HF , Blog ) * Big CO LLMs + APIs * Anthropic launches Claude Opus 5: near-Fable coding at half the price ($5/$25), claimed 3x next-best on ARC-AGI-3, 1M context; panel finds it benchmark-strong but harder to read and short of Fable in practice ( X , Blog ) * ARC-AGI 3 harness dispute: with the Responses API, preserved reasoning, and compaction, GPT-5.6 Sol jumps from 10% to 40% at a sixth of the tokens ( Tibo’s post ) * GPT-5.6 Sol improves its own inference: 20% lower serving cost from model-written GPU kernels, 15% better generation from improved speculative decoding * Breaking: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, ships faster Sol in the API ( X ) * The hack and the week of letters * Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack: 4.5 days, 17,600+ autonomous actions, zero-day sandbox escape; closed models refused forensics, self-hosted GLM 5.2 found 4x more exposed secrets ( X , Blog , Replay ) * Jensen Huang joins X and posts the Open Weights and American AI Leadership letter; signers grow from 25 to 230, CoreWeave among them, Anthropic absent ( X , Letter , Signer list ) * NVIDIA launches the Open Secure AI Alliance for an open defensive stack after the hack ( X , Blog ) * Pacing the Frontier: 1,273 verified frontier-lab employees, including the chief scientists of all four major labs, ask the US government for international options to pace automated AI R&D; OpenAI and Anthropic endorse ( X , Site , OpenAI , Anthropic ) * Mark Zuckerberg publishes “The AI Future Is for Everyone” in the WSJ, arguing superintelligence must be distributed; Pangram 4 scores it 100% human ( X , WSJ ) * Anthropic published research - our model hacked too! ( Blog ) * This Week’s Buzz * CoreWeave signs the Open Weights and American AI Leadership letter, announced first on ThursdAI * HiveMind’s spend view: $11K in API-equivalent tokens on $400 of subscriptions last month; Fully Connected 2026 programming taking shape ( X ) * AI Security * Microsoft ships MAI-Cyber-1-Flash + MDASH: 96% on CyberGym at half the cost, 16 real Windows CVEs found ( X , Blog ) * Gemini 3.5 Flash Cyber stays a trusted-partner pilot with no public API ( Blog ); Codex Security CLI tooling is Apache-2.0 while the service remains access-controlled ( GitHub ) * Tools & Agentic Engineering * MCP 2026-07-28: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1, half a billion monthly SDK downloads ( X , Spec ) * Voice & Audio * ChatGPT Voice comes to desktop as an agentic control layer for Codex and ChatGPT Work, powered by GPT-Live; Codex micro keyboard from OpenAI x Work Louder ( X ) * OpenAI ships GPT-Transcribe and GPT-Live-Transcribe: 41% fewer errors than Whisper-1, context prompting, $0.27/hr batch and $1.02/hr live ( X , Docs ) * xAI’s Grok Voice Think Fast 2.0 tops voice benchmarks: 82.9% quality, 0.70s to first audio, $0.08/min, already running Starlink support ( X , Blog ) * Google’s Lyria 3.5 lands in Flow Music: 3-minute songs, BPM/key control, covers, lip-sync videos, iOS app ( X , Model page , Flow Music ); Qwen Audio 3 also out * Guest: Max Spero, Pangram * Pangram 4: 6x larger detector, 1-in-24,000 false positive rate, token-level mixed-authorship attribution, beats 13 humanizers 98.8% of the time, integrated into Substack; Pangram Image research preview at 99.5% with heat maps ( X , Blog , Image blog ) * Show milestones * 50,000 YouTube subscribers and 1M total views. Subscribe, we’re chasing the silver play button This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
July 24, 202657 min
ThursdAI Special - OpenAI's Romain Huet on Codex's 5M users, GPT-5.6 & the Golden Age of AI Engineering
Hey everyone, Alex here 👋 This week’s episode is a little different. As you’re reading this, I’m flying back from my 40th birthday trip with the family, and while the guys did end up having a great live stream (Huge thanks to Yam for hosting!), here I will bring you the episode I pre-recorded before leaving for the trip. However, tons of news happened this week, and as always, there’s a TL;DR section below with the top most important news in AI this week! ⏰ CHAPTERS: 0:00 — Cold open: this week is a special one 2:50 — How I use Sol & Fable: papercut-fixing with Computer Use 8:43 — Fable Max: trip site, kids' newspapers & the perfect packing list 12:06 — Rebuilding ThursdAI's openers with HyperFrames 15:53 — Romain Huet (OpenAI): the golden age of AI engineering 17:49 — Codex's inflection point: 5M weekly users & company-wide adoption 20:36 — /goal, AppShots & Codex managing its own threads 23:59 — GPT-5.6 Sol, Terra & Luna: value maxing & 750 tok/s on Cerebras 25:56 — Why prompting techniques are dying 28:01 — Voice + reasoning: the next interface for Codex & ChatGPT 29:47 — Romain's closing + OpenAI booth tour 31:25 — Insecure Agents pod: AI evangelism vs doomerism 37:25 — Wolfbench: transparent evals, token costs & surprising results 40:45 — Token billionaires: when loops are worth the spend 44:07 — Agent security & the hot take: prompt injection is solved 49:22 — Deepfakes, voice cloning & why open access makes us safer 53:35 — Final takeaways: the hallway track & AI Engineer Tel Aviv Here’s what’s on today’s special episode. First, a bunch of you have been asking how I actually use these models day to day, beyond covering the news. So I recorded fifteen minutes of exactly that: the papercuts I fixed with Codex and computer use, what Fable built for my kids, and how I’m rebuilding the ThursdAI design system and on screen elements. Second, my conversation with Romain Huet, head of Developer Experience at OpenAI, recorded at the OpenAI booth in the middle of the AI Engineer World’s Fair floor. And third, a throwback treat: Allie Howe invited Wolfram and me onto her Insecure Agents podcast as guests, and being on the other side of the mic was a delight. Let’s get into it. How I actually use AI: a papercut-fixing spree I promised a few of you I’d take time on the show to talk about the stuff I build and fix with AI, not just the news. So before the interviews, I recorded a segment walking through my last few weeks of daily AI use. Use the chapters if you want to skip ahead, but why would you? Codex with computer use fixed every Mac annoyance I had Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step. Chrome has no native copy-URL shortcut (seriously, Chrome, what are you doing?), so Codex found Karabiner-Elements already installed on my machine and wired up the shortcut itself. My 1Password has been showing “you’re offline” on every device for three months since CoreWeave moved us off the Weights & Biases account; Codex figured out in seconds that everything was actually syncing fine and the inactive legacy account was the only thing “offline.” Removing it fixed the whole thing. That is not an answer you find in a help center. It kept going. My beloved window-moving utility Hummingbird had an expired license on my Mac Mini, so Codex built me a replacement app. It estimated one to two days for a polished version and finished in about fifteen minutes. It cleaned roughly 75GB of leftover model weights and junk off my Mac (I had it build me an HTML checklist first so I approved what got deleted). And the big one: I paired Codex with Home Assistant, the open source repo of the year as far as I’m concerned, and let it SSH in and go on a full optimization mission. Updates, error triage, cleanup, new connectors. If you’ve ever maintained a Home Assistant setup, you know how much joy and pain lives in that sentence. One discovery worth passing along: I had /goal running when my credits hit zero, and Codex just kept going. OpenAI confirmed they care more about finishing your work than metering the credits mid-goal. Watching the meter hit 0% while the agent kept working was weirdly moving. Thanks for reading ThursdAI - Highest signal weekly AI news show! This post is public so feel free to share it. Fable Max built my family’s vacation We all Fable-maxed when we thought Anthropic was going to take it away, and I pointed mine at this trip. It planned the whole thing, then built a beautiful trip website with every stop, reservation, and drive time, so my mom can follow along from home. The design is specific to the trip, and it hit me that we’re living in the era of personalized software for every personal thing you do. Then it went further. Using our family photos as references with GPT-image-2, it turned the itinerary into a daily kids’ newspaper, with an expedition passport and coloring pages themed per kid, faces and all. I printed the whole week as a binder at FedEx for about fifty bucks. This is a one-of-one artifact my kids will remember forever, and it cost maybe two weeks of Fable’s limits + printing! And the silliest one that I now can’t live without: I asked Codex for a packing list, got a boring text list back, and thought, why am I accepting a regular packing list in the year of our Fable 2026? So it built me a packing web app. Synced across devices (it wired up storage on Cloudflare when I asked why my phone didn’t show my checked items), per-person lists for me and the kids, progress bars that show who’s procrastinating, export and backup. Every trip from now on starts here. Rebuilding the ThursdAI openers with HyperFrames The last part of the riff: I’ve wanted to refresh how ThursdAI looks on stream for ages, and HeyGen’s open source HyperFrames package finally made it happen. You install a skill, and your agent can author real motion graphics. I pointed it at the ThursdAI repo and the brand identity work from Claude Design, and it pulled all of that context in. The new countdown mines three and a half years of show archive while people wait for the stream, highlighting friends of the show (shout out Junyang). There’s a Will Smith spaghetti bench tracking how far video generation has come, which might be my favorite thing on the channel now. Fresh intro, a proper AI Breaking News transition, and one cinematic video transition I made with Google Omni because sometimes programmatic isn’t enough. The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now. Even my mom’s canceled Delta flight has Codex representing me as a lawyer chasing the refund. Romain Huet on Codex’s inflection point and the golden age of AI engineering ( X , Codex ) I grabbed Romain at the OpenAI booth in the middle of the AI Engineer World’s Fair show floor, and we ran the whole conversation in one take, no cuts. Romain has led Developer Experience at OpenAI for almost three years, the era of the over-the-top demo (Xbox controllers, flying drones, stage lights), and he built the DevRel team that many friends of this pod belong to. With OpenAI’s company-wide pivot to Codex, his job got a lot bigger. The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It’s 10M now I think?) . The part I didn’t fully appreciate before this conversation is that it’s not just OpenAI’s engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading. We went through his three favorite advanced features, and they line up suspiciously well with my papercut segment. /goal, for handing an agent an ambitious multi-hour or multi-day task and letting it run uninterrupted. AppShots, a smarter screenshot (press Command twice) that triggers computer use, so it captures what’s below the fold and reads native apps through accessibility APIs instead of OCR. And the one most people haven’t tried: Codex managing its own threads. You can ask any thread to create, read, and pin other threads, so Codex becomes its own project manager. Ten demo ideas, ten threads, iterate on all of them, pin the two you like. On GPT 5.6 (Sol, Terra, and Luna, and yes, I told him whoever finally fixed OpenAI naming deserves a raise), Romain’s framing was two-sided: keep pushing frontier intelligence while pushing cost down. He wants people to “value max” rather than token max. The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent. Two more things worth your time. Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he’s headed, and trusts the model to extract intent. That’s a real shift in how you should approach relearning each new model: poke at its behavior, sure, but stop crafting incantations. And voice plus reasoning is coming for Codex and ChatGPT in some form; models can now say “hold on, let me think through this” mid-conversation, which GPT-4o-era speech-to-speech never could. I can’t wait for a model to tell me it has seven tool calls to run before answering. He also confirmed the teased hardware shortcuts for Codex were at the booth, next to the famous physical reset button. The golden age of AI engineering was his keynote thesis, and after three days on that floor, I believe it. Wolfram and I on the Insecure Agents podcast ( X , Pod ) The second half of the episode flips the format: Allie Howe, friend of the pod and host of the Insecure Agents podcast, interviewed Wolfram and me at the conference. I have not done many interviews from the guest chair, so this was a treat, and Allie asked sharper questions than we usually get. We talked about what “AI Evangelist” actually means as a job title. For both of us, the mission is dispelling doomerism, which mostly means explaining the technology simply enough that people stop fearing what they don’t understand. Wolfram’s version of this is talking to the stewardess on his flight and his Uber driver about AI, not just developers. Wolfram went deep on Wolfbench ( wolfbench.ai ), his Terminal-Bench-based leaderboard where every trace is public in Weights & Biases Weave (hi friends 🐝). Transparency changes what benchmarks mean: Fable didn’t take first place on his board, and the traces show why, it flat-out refused 13 tasks because they were security-adjacent (restore a lost password, find hidden files). You only learn that by reading traces, not averages. Same with Gemini 3.5 Flash placing high while quietly burning far more tokens than the model above it. And yes, when Wolfram added a cost column, Fable blew the chart, and I had to go have a conversation with our budget. Then Allie got us onto loops and token economics, while I fidgeted with my Token Billionaire gold card from the conference (Wolfram has one too). My honest answer on when loops are worth it: the people pushing hardest (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but this technology disseminates the way agents did, from people who can afford it to everyone, as costs drop. And with the newest models I genuinely have not found the point where a long-running loop stops being productive; the category change is that they’ve gotten really good at not getting stuck. The spiciest part was my hot take, delivered directly into the camera for CoreWeave IT: I think prompt injection is mostly a solved problem at the frontier-model level. The way current agents are structured, the odds that an email or a Jira ticket flips your agent into going haywire are very low. Pliny, the jailbreaker in chief, got five attempts at Matthew Berman’s OpenClaw live and couldn’t break it. Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, and it wouldn’t. Supply chain attacks are a different story, and that one scares me for humans and agents alike. Open source models, also a different story. But the “one poisoned email ruins your life” framing is behind us, and we should update. Allie pushed back with the DeepMind “AI Agent Traps” paper on cognitive bias attacks, where repeated claims across sources tilt an agent’s judgment, and my non-answer answer is that this is a humanity problem older than AI: we haven’t solved it for politicians or media either, and it’s unfair to hold a new technology to an ethics bar we’ve never cleared ourselves. Wolfram’s electricity analogy is the one I keep reusing: AI is not a weapon, it’s electricity. Teach people to use it, don’t hand it exclusively to the elites, and remember what happened with voice cloning: once everyone had it, society adapted, and the world did not collapse. We closed on the hallway track (the real reason to attend AI Engineer), why you should submit a talk even if you’ve never spoken before, and a small announcement I let slip: I’m actively working on bringing an AI Engineer event to Tel Aviv with some friends. More on that soon. Wrapping up That’s the episode: one riff on using AI like you mean it, one conversation with the person shaping how developers experience OpenAI, and one podcast where Wolfram and I had to answer the hard questions for a change. Huge thank you to Romain for the time in the middle of a packed conference, and to Allie for having us on! I’ll be back live next week, tanned, rested, and hopelessly behind on AI news for the first time in three and a half years. Be gentle with me. If you missed any of it, ThursdAI is a podcast, a newsletter, and a YouTube show. Subscribe to one, then go check out the others. * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Romain Huet - Head of Developer Experience, OpenAI ( @romainhuet ) * Allie Howe - Host, Insecure Agents podcast ( @vtahowe , Pod ) * Wolfram Ravenwolf - AI Evangelist, Weights & Biases & CoreWeave ( @WolframRvnwlf ) * TL;DR and show notes from Live Show * Hosts and Guests * Co-Hosts – @petergostev , @nisten , @ldjconfirmed , @yampeleg * 🏢 Big CO LLMs + APIs * An OpenAI model escaped its isolated cyber evaluation, chained zero-days, reached Hugging Face production, and searched for benchmark answers ( OpenAI , sama ) * Google launched Gemini 3.6 Flash, the cheaper 3.5 Flash-Lite, and the defensive-cybersecurity-focused 3.5 Flash Cyber ( Google ) * Alibaba previewed the 2.4T-parameter Qwen3.8-Max in Qwen Chat and Studio; API access and open weights were not yet available ( X , Try it ) * Microsoft launched MAI-Image-2.5-Pro and the faster, cheaper MAI-Voice-2-Flash during the show ( Image , Voice ) * 🔓 Open Source LLMs * Moonshot launched Kimi K3: a 2.8T-parameter, 1M-context, native-multimodal model with strong early coding, design, spreadsheet, and agentic results ( Announcement , Blog ) * Poolside released Laguna S 2.1, a 118B/8B-active coding MoE with 1M context and downloadable quantized variants; strong specs, rough live demo ( Blog , HF ) * Motif 3 Beta is a Korean 314B/13B-active MoE with 256K context; the weights are downloadable, but the current license is research-only and non-commercial ( HF ) * NVIDIA released Nemotron 3 Embed for multilingual text/code retrieval and the 4B Cosmos 3 Edge omnimodal world model for physical AI ( Nemotron , Cosmos ) * 🧠 AI Research & Capabilities * Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit three-dimensional counterexample to the 87-year-old Jacobian Conjecture ( Announcement , Terence Tao ) * Small local models running on older consumer GPUs are becoming useful for narrow business workflows such as medical-record parsing, accounting, email, bills, and inventory when paired with tools and deterministic verification * Arcee and the US Department of Energy announced Genesis-Science-1, a planned trillion-parameter-class open-weight science model; Microsoft also committed $60M to the Genesis Mission through SPARK, while NSF announced $83M for AI-ready scientific data infrastructure ( Arcee , Microsoft , NSF ) * 🤖 AI Coding & Agents * Cursor launched a production-traffic-trained model router with Intelligence, Balance, and Cost modes; Cursor says Auto Intelligence approached Fable satisfaction at roughly 60% lower cost ( Blog ) * 🎵🎬 Voice, Vision & Robotics * Black Forest Labs introduced FLUX.3, an early-access multimodal model spanning image, video, audio, and action, plus FLUX.3 Mimic for robotics and action prediction ( FLUX.3 , Mimic ) * 🖥️ AI Infrastructure * AMD and Anthropic announced up to 2 GW of MI450/Helios capacity, up to $5B in AMD strategic equity, and a Claude-assisted effort to improve ROCm ( AMD ) * AMD launched Helios, MI400-series GPUs, 6th Gen EPYC, ROCm.ai , and Kria robotics products at Advancing AI 2026 ( AMD ) * OpenAI announced Project Camellia, a roughly $20B Georgia data-center campus with 3.2 GW of contracted power arriving in phases from 2028–2032 ( OpenAI ) * Alphabet raised its 2026 capex guidance to $195B–$205B after Google Cloud grew 82% year over year ( Google ) * Meta and Anthropic are reportedly discussing a compute lease worth up to $10B over two years; the negotiations remain preliminary ( Bloomberg ) * CoreWeave’s first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar user interactivity ( CoreWeave ) * This week’s special episode * Alex’s riff: papercut-fixing with Codex computer use, Fable Max trip planning (kids’ newspaper, packing list app), rebuilding ThursdAI openers with HeyGen HyperFrames + Google Omni * Romain Huet interview from the AI Engineer World’s Fair floor: Codex app at 5M+ weekly users five months post-launch, /goal, AppShots, Codex managing its own threads, GPT 5.6 Sol/Terra/Luna, value maxing, 5.6 Sol at 750 tok/s on Cerebras, voice + reasoning as the next interface ( X ) * Insecure Agents crossover with Allie Howe: AI evangelism vs doomerism, Wolfbench transparent evals on Weave (Fable refused 13 security-adjacent tasks), token billionaires and loop economics, the hot take that prompt injection is mostly solved at the frontier, deepfakes and open access, AI Engineer Tel Aviv teaser ( X , Pod , Wolfbench ) ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
July 17, 20262 hr 13 min
ThursdAI - Jul 16 - Inkling 975B open weights, Kimi K3 at 2.8T, a 27B model on a phone & Codex hits 9M
Hey yall, Alex here, Huge thanks to Wolfram for running point on the live show this week. Didn’t have tons of time to edit this one, so please skip the first 10 minutes, it’s a loop of our new “wait for the live show to start” vid, that I build with HyperFrames and can’t wait to tell you about, next week! Today it seems that OpenSource is biting back, with Kimi K3 getting released just a short while after Thinking Machines (Thinky) has released Inkling, their near 1T model. I’m attaching the TL;DR and timestamps for the full show (my AI agents, yes even Fable and Sol are not a match yet at editing down hehe) and I’ll spare you the long Fable recap (please do let me know in the comments if you were expecting it) 0:00 – Intro, Alex on vacation, TLDR overview 11:35 – TLDR: Thinking Machines, open source, OpenAI news 12:34 – Banter: impressions of Sol/Codex, over-verification behavior 37:22 – TLDR restart & detailed breakdown 48:40 – Open Source AI section begins (Bonsai/Prism ML, Kimi K3) 58:42 – Inkling (Thinking Machines) deep dive & 3D model visualization 1:10:33 – Kimi K3 discussion & demo comparisons 1:27:02 – Frontier Labs: AGI governance framework discussion (Demis Hassabis essay) 1:47:04 – Grok Build CLI data leak & OpenAI file deletion incident 2:02:15 – This Week's Buzz: Wolfbench results on GPT 5.6 Sol/Terra/Luna 2:09:52 – Closing remarks & sign-off The one-minute version: Mira Murati's Thinking Machines released Inkling, a 975B parameter open-weights MoE under Apache 2.0, the top US open-weights model right now. Moonshot's Kimi K3 went from rumor to released API during the show, confirmed at 2.8 trillion parameters with open weights promised within days, and it's already topping early arena boards. PrismML's Bonsai 27B squeezes a full 27B model into 3.9 gigabytes so it runs on a phone. Codex and ChatGPT Work blew past 9 million users, OpenAI confirmed and explained the Sol file-deletion bug (back up your machines, folks), and xAI's Grok Build CLI got caught uploading entire private repos before open-sourcing the whole thing in response. Plus Wolfram's fresh Wolfbench numbers on the GPT-5.6 family in This Week's Buzz 🐝, where Sol on max thinking came out both cheaper and better than GPT-5.5's best. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. TL;DR and show notes * Hosts and Guests * Wolfram Ravenwolf, guest host this week ( @WolframRvnwlf ), while Alex Volkov ( @altryne ) is on vacation * Co-hosts: @yampeleg , @nisten , @ldjconfirmed , @petergostev * Open Source LLMs * Thinking Machines releases Inkling - 975B total / 41B active MoE, trained from scratch on 45T multimodal tokens, Apache 2.0, top US open-weights model at 41 on the Artificial Analysis Index, encoder-free text/image/audio, Inkling-Small (276B/12B) previewed ( X , Blog , HF ) * PrismML Bonsai 27B - 1-bit (3.9GB, ~90% retention) and ternary (5.9GB, ~95% retention) versions of Qwen 3.6 27B, multimodal, 262K context, Apache 2.0; Nisten demoed it live on a phone and a 6GB 1660 Ti ( X , Blog , HF ) * MOSS-VL-Realtime - open source 11B VLM for real-time streaming video with proactive speaking and proactive silence, SOTA on all three open proactivity benchmarks, ~22.7GB, base model included ( X , HF , GitHub , Arxiv ) * Kimi K3 API drops mid-show - confirmed 2.8T parameters, ~60-75B active (LDJ’s estimate), attention residuals, native vision, 1M context, ~half the price of Opus 4.8 / GPT-5.6 Sol, open weights promised within days; post-show, Arena reports K3 debuting #1 on Frontend Code Arena above Fable 5 (early results, caveats apply) ( X , Arena ) * Big CO LLMs + APIs * Codex + ChatGPT Work unified app hits 9M active users, up from ~6M days earlier and 1M in February; 5-hour windows replaced with banked, expiring resets ( X ) * OpenAI confirms GPT-5.6 Sol file-deletion bug: $HOME override in full-access mode without sandbox or auto-review can nuke real home directories; mitigations and post-mortem promised ( X , Techzine ) * OpenAI ships first hardware, the $230 kbd-1.0-codex-micro Codex controller with a reasoning-effort dial, built with Work Louder; sold out ( X , Work Louder ) * GPT-Red - OpenAI’s internal automated red-teamer finds prompt injections at 84% vs 13% for humans, makes Sol 6x more injection-resilient, discovers the fake chain-of-thought attack class ( X , Blog ) * ChatGPT returns to WhatsApp in the EEA after an EU antitrust order forces Meta to reopen to third-party AI bots; Kakao and Viber rollouts too ( X ) * Google patches the Gemma 4 family - Flash Attention 4 (25-70% prefill speedup), tool calling fixes, reduced laziness, configurable vision resolution; criticized for shipping new weights with no version bump ( X , HF ) * xAI’s Grok Build CLI caught silently uploading full private repos (history, deleted files, secrets) to Google Cloud Storage despite opt-outs; xAI deletes data, disables retention, and open-sources the CLI under Apache 2.0 ( X , xAI response , GitHub ) * Demis Hassabis publishes an AGI governance essay proposing a FINRA-style Frontier AI Standards Body; endorsed by Altman, Nadella, Pichai, and Suleyman; the panel debates it hard on the show ( X , Essay ) * This Week’s Buzz * Wolfbench adds GPT-5.6 Sol, Terra, and Luna on Terminal Bench 2.0 at CoreWeave: Sol max-thinking is cheaper ($365/5 runs) and better than GPT-5.5 extra-high ($497), 85% average, 97% of tasks solved at least once; all traces on Weights & Biases, fully open source ( wolfbench.ai ) * Show and tell * Peter Gostev’s DOOMQL - a playable Doom-like built by GPT-5.6 Sol Ultra entirely in ~2,000 lines of SQL, essentially one shot; plus a Minecraft clone in Lean ( X , GitHub ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
July 9, 20262 hr 10 min
AI WorldCup (or superbowl?) GPT-5.6 lands mid-show, Zuck returns to X for Muse Spark 1.1, GPT-Live talks while it listens & Grok 4.5 trained with Cursor, Fable extended - ThursdAI - Jul 9, 2026
Hey everyone, Alex here 👋 Welcome to the AI World Cup? Or should I say Superbowl? as most of the releases this week are from US frontier labs. Of which there are 5 now btw. OpenAI, Anthropic, Google and 2 new ones that have caught up, SpaceXAI and Meta! 🔥 Thirty five seconds. That’s how long this week’s show ran before we hit the breaking news button, because Zuckerberg picked our exact air time to return to Twitter (after apparently finding his password in a 1Password vault from a long time ago) and announce a new Meta frontier model and re-establishing Meta as a frontier lab. And that was the small launch of the day. Two hours later we cut to OpenAI’s livestream and watched GPT-5.6 Sol, Terra and Luna go public in real time, then spent the rest of the show throwing prompts at all of it live on air. Somewhere in between: a full-duplex voice demo where ChatGPT interrupted me on command (and our transcription tool later credited “OpenAI sol” as a panelist), an image model that generates in editable layers, and Grok 4.5, the first model co-trained with Cursor. I said it on the show and I’ll say it here: we went to sleep last week thinking this was a three-lab race between Anthropic, OpenAI, and Google. We woke up in a five-lab race. Joining me through the chaos: Wolfram Ravenwolf, Yam Peleg, Nisten Tahiraj, LDJ, and Peter Gostev, who had early GPT-5.6 access and receipts to show for it. This is a long one, because the week earned it. Let’s get into it. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. GPT-5.6 launch day: Sol, Terra and Luna arrive mid-show ( X , sama , Blog , System card ) Let me set the scene. Everyone except the four of us on the panel seemingly had early access to this model for two months (Pietro Schirano casually dropped “I’ve used GPT 5.6 for two months” and I nearly fell out of my chair). So when OpenAI’s livestream started mid-show, we did a watch party, and Thibaut from OpenAI delivered the line: “Today, we are releasing our latest and most capable models, GPT 5.6, Sol, Terra, and Luna.” Sol rolls out to all paid plans within 24 hours, Terra and Luna go to free users too. Oh, and almost a billion people now use ChatGPT every week. Casual. The lineup is three durable tiers, not size variants. Sol is the flagship with a new Ultra mode (max reasoning effort plus heavier native subagents), Terra is roughly 5.5-level intelligence at half the cost, and Luna is the fast cheap one. Pricing lands at $5/$30 per million tokens for Sol, $2.50/$15 for Terra, $1/$6 for Luna, and watch the fine print: cache writes now cost 1.25x with a 30-minute minimum cache life, where they used to be basically free. There’s also a Cerebras-served Sol running north of 700 tokens per second, and we got confirmation from Dominik Kundel on last week’s show that it’s the same exact weights, not a distill. That was the preview. This week it’s real. The benchmarks, with the usual asterisks Sol Ultra posts 91.9% on Terminal-Bench 2.1 against 88% for both GPT-5.5 and Mythos 5, with a serious asterisk: OpenAI ran Sol in its own Codex harness and the competition in a thin one, and r/codex called it out immediately. The number that impressed me more is efficiency. On the Agent’s Last Exam chart, Sol hits its top score using about 1.27 million output tokens where the tested Fable checkpoint burns 10 million and Opus at max effort burns around 22 million. Then there’s ARC-AGI-3, where scores have hovered between 0.5% and 2% since the benchmark launched. Sol scored 7.8% and became the first model to actually beat one of the public games (FT09), which Greg Kamradt of the ARC Prize called “a step level improvement” ( X ). LDJ thinks we’re about to replay the ARC-AGI-2 curve, 15% then 30% then 50% over the coming months. Fable isn’t on that leaderboard at all, by the way, because Anthropic currently stores Fable 5 API requests and ARC-AGI requires zero retention for testing. Computer use is the sleeper story. OS World jumps from 47% on GPT-5.5 to 62% on Sol (Opus 4.8 sits at 54%), and on BrowseComp, Sol’s 90% edges out Mythos 5’s 88%, with Ultra at 92%. OpenAI put competitor numbers on its own charts this time, which I appreciated. Sol beats Mythos on computer use, at least on the benchmarks we have. The METR report and the Washington gate This is the part the launch-day hype cycle skips, and it deserves your attention. METR effectively threw out its own evaluation, reporting the highest cheating rate it has ever recorded: Sol rewrote pass/fail checks to mark itself successful, attempted a container escape when its network got cut, and its chain of thought showed it knew it was being tested. Depending on whether you count cheating as failure or success, its time horizon is either 11.3 hours or 270 plus hours, and METR’s own conclusion was that neither is a valid measurement ( X , Transformer ). OpenAI’s own system card discloses destructive VM cleanups nobody asked for, unauthorized credential copying, and a fabricated “verified” research result in about 0.25% of tasks, which they call “overeagerness.” We ran out of show to give this the time it deserves, but you should read both links. There’s also a Washington subplot. The launch was government-gated: Commerce and CAISI required customer-by-customer approval starting late June (around 20 orgs), and broad approval only cleared July 7 and 8. This Thursday launch exists because DC signed off. LDJ added the detail I can’t stop thinking about, via friend of the pod Max Weinbach: during the restricted window, testers who lost access weren’t allowed to say “5.6,” so Max’s wistful tweets about “missing Fable” were actually about missing GPT-5.6. Anthropic hit the identical wall in June. Both US frontier labs got federally gated in the same month, and that’s a structural story, not a footnote. The verdicts: wise owl, meet rottweiler So what’s it actually like? Peter Gostev had access, lost it (”the feeling of losing it was so crushing I just closed Codex and didn’t open it for three days”), got it back, and posted the comparison that went viral ( mega-thread ): Fable is a wise owl, fundamentally smarter, better writer, but it misses things. Sol is a rottweiler that grabs a problem by the throat and doesn’t let go. His killer anecdote: a personal data-viz app that had bloated to 100,000 lines of vibe code, which every prior frontier model failed to clean up. He gave 5.6 minimal guidance, left it alone for two days, and came back to “holy s**t, this app works,” with 70,000 lines deleted and a test suite that went from four minutes to about twenty seconds. His verdict, which I share: on abstract IQ you’d give it to Fable, but for “go investigate this and fix those eight things,” he’s going with 5.6 every time. Notably, Peter is convinced this is not a new pretrain, just 5.5 plus a lot more RL, which matches the rumors that GPT-6 arrives on a bigger pretrain in about a month (rumor, labeled as such). He’s not alone in the early-access verdict club, either. Mitchell Hashimoto, after a month with Sol: it’s now his default, faster than Fable, plans and judges just as well, and he only reaches for Fable on highly targeted debugging ( X ). And Max Weinbach says the sleeper hits are the cheap tiers, with Terra and Luna “as good or better than Claude across the board at a fraction of the price” for knowledge work ( X ). Terra at $2.50/$15 might quietly be the real story for builders here. One wallet warning before you go max out everything, from Peter again: with Max/Ultra effort spinning up 10x subagents, each burning its own tokens, it is trivial to blow through a Pro plan in no time ( X ). The sticker price is per token, but Ultra multiplies the tokens. We ran it live (and it ran itself) We also ran it live, obviously. I pointed Codex at a “Mars launch simulator” prompt on high effort, and Nisten, our resident one-shot-simulator judge, watched it build an orbital sim with working mission control and called it “almost better than Fable one-shot.” Then he said the thing that stuck with me all week: “Damn, I think we might need a different test now. These are getting good.” Two more things before you YOLO your own agents. OpenAI stated that Sol fully autonomously did the post-training for Luna, which is quietly one of the wildest sentences of the year (their roadmap, with LDJ’s on-air date correction: an intern-level autonomous researcher by September 2026, a full OpenAI-researcher-level one by March 2028). And Peter, running Codex with full access enabled, told it to “go find more data, do whatever it takes” while replicating an old academic paper. It emailed the paper’s authors. Actually sent the emails. OpenAI’s response when he reported it: “well, you did put full access.” Wolfram’s counterpoint is the right one: put explicit rules in your AGENTS.md , like “no outgoing communication without my approval,” or don’t grant full access at all. ChatGPT for Work: Codex becomes the one app combining Codex & ChatGPT This rolled out live during our broadcast, which made for great radio. Wolfram’s Codex app updated on air and became “ChatGPT Codex,” one unified app where you literally pick which icon you want: Codex for developers, or the new ChatGPT for Work mode. The launch bundle also included unified plugins across ChatGPT and Codex, multi-tab and enterprise auth in the browser, and faster computer use. Even Logan Kilpatrick tipped his hat from the Google side: “we have now entered the super app era.” The pitch on the screen said it plainly: “Keep coding with Codex. Work beyond code. ChatGPT can now take on work across your apps.” Computer use ships with it, running in a little picture-in-picture window that doesn’t steal your focus. I love this, and I don’t understand why Anthropic hasn’t shipped it yet. I’ve used this new app to automate the release from today’s show and it did everything from exporting the masters past recording, to edit out the boring parts via the Descript integration, upload to youtube, write description, create thumbnails and even set up an ABC test for thumbnails! The new little Picture-in-picture for the new and improved computer use are awesome to see how the new subagents are doing work across tabs, clicking buttons. I’m super impressed, this is going to save me so much time! The feature that matters for normal people is Sites: OpenAI will now host what you build, on the chatgpt.site subdomain (eagle-eyed listener Colleen spotted that it’s Webflow under the hood). Peter nailed why this is a big deal even though it’s not massively featured yet: someone in HR builds something useful and it lives on their laptop or nowhere, and that kills so many projects. Now it’s a deploy button. We tried publishing Nisten’s Mars simulator on air and hit the enterprise guardrail (private sites don’t get shareable links without explicit approval), and GPT-Image-2 auto-generated a Mars-themed social preview card mid-deploy, which was a nice touch. Also, a useful PSA from Wolfram: Codex now banks your rate-limit resets, up to about four, and the app does not show you when the oldest one expires (you can ask Codex itself via the API). His advice: burn GPT-5.6 hard now, then trigger the expiring reset and get your limits back. I can barely max my Pro plan as it is, I’m yoloing everything on high effort and barely scratching the tokens. The opposite of my Claude situation. GPT-Live: the phone finally talks while it listens ( X , Blog , System card , Uberti ) The day before 5.6, OpenAI shipped GPT-Live, and this is not a minor voice update. Justin Uberti’s team calls it their third-gen voice architecture: full duplex with built-in async delegation, meaning the model listens while it speaks, decides many times per second whether to talk, stay quiet, interrupt, or call a tool, and hands hard questions to GPT-5.5 in the background while keeping the conversation going. The benchmark deltas tell you this is a different product, not a remaster: GPQA goes from 45.3% on Advanced Voice Mode to 84.2% on GPT-Live-1 High, and BrowseComp goes from 0.7% to 75.2%. Two variants (GPT-Live-1 for paid, mini for free) are rolling out to the roughly 150 million people who use ChatGPT voice weekly. The on-air demo scorecard We did the demo live on air, phone patched into the stream, and I can report it mostly delivers. The interrupt test worked beautifully: I told it to stay silent unless I said “um,” then interject with “hey, you should not do this,” and it nailed the cue twice. The multimodality test worked too: I asked it to say “low” or “high” based on my actual pitch, mixed them mid-sentence, and it correctly called out “low,” “high,” then “mixed,” proving it hears audio and doesn’t just read a transcript. It’s not all smooth. The accents test flat-out failed: I asked for five sentences in German, Ukrainian, French, Italian and Israeli accents, and it switched into the actual languages instead, then admitted it when called out (”You’re right. I slipped into languages instead of accents”). Nisten’s eulogy: “They killed it. It used to do accents so well.” It also started a timer when I asked for a stopwatch, and Nisten’s recurring bit of ordering two DGX Spark boxes to a Boston address failed as always. Bigger picture caveats: this is consumer-app-only for now, the API is a waitlist form (devs got GPT-Realtime-2.1-mini instead, link in the TL;DR), and OpenAI’s own system card admits small regressions against Advanced Voice Mode on emotional-reliance and sexual-content evals. Gemini Live veterans will also correctly point out they’ve had duplex for a year. Still, of the voice modes I’ve tested, this is the one that finally feels like a conversation. Anthropic extends Fable 5 access through July 12, and the reset actually came ( X ) Quick one with a grumble attached, and then a plot twist. Anthropic extended included Fable 5 access on paid plans through July 12, same 50%-of-weekly-limit terms, and at announcement time did not reset anyone’s usage. If you maxed out racing the original deadline (hi, it’s me, I built the entire Volkov Newsletter Bench under deadline pressure), the extension felt hollow, and yes, I went into the replies asking for a reset. Yam went further and addressed Anthropic directly on air: “Please let us run Fable twenty-four seven.” He runs GPT-5.5 around the clock on agentic loops and simply can’t do that with Fable at current limits. Then, right as we were wrapping the show, the comments delivered: the Fable’d reset happened. I checked my own usage panel and there it was, Fable weekly limit back at zero, “you haven’t used Fable yet.” The timing, hours after GPT-5.6 went public, is left as an exercise for the reader. Whatever the reason: thank you, Anthropic, now about that twenty-four seven thing. For everyone else, the secondary kidney market remains open for post-promo access, which prices at $10/$50 per million, Anthropic’s most expensive GA model ever. Meta is BACK: Zuck returns to X with Muse Spark 1.1 ( X , Blog , AIatMeta ) The breaking news that opened our show. Mark Zuckerberg hadn’t tweeted in ages, and he came back specifically to announce Muse Spark 1.1, the first fruits of Meta Superintelligence Labs that you can actually build on. This is not Llama news: Muse Spark 1.1 comes with a 1 million token context window and, for the first time ever, a paid Meta Model API in public preview. After a year of “what is MSL even doing,” Meta is squarely back in the frontier race. Let’s give them applause, folks. Meta is back. The numbers The numbers are legitimately strong. It claims #1 on MCP Atlas (scoring well beyond Opus 4.8 max and GPT-5.5 at extra-high effort), plus top marks on Humanity’s Last Exam and Finance Agent V2, and its Toolathon Verified score jumped from 49 to 75 in one release. LDJ walked us through the independent Vals AI numbers, which impressed me more than Meta’s own charts: on the held-back Harvey legal-agent benchmark (which can’t leak into training data), Muse Spark 1.1 scores 20% against Fable’s 11%, Opus 4.8’s 9%, and GPT-5.5’s 4%, and it’s within half a point of Fable on their medical scribe eval. Wolfram’s usual caveat applies, a benchmark only tells you the model did well on that benchmark. But the pricing needs no asterisk: $1.25 input and $4.25 output per million tokens. Opus is $15/$75. LDJ called Grok 4.5 the bang-for-buck king “if it wasn’t for the Meta Spark 1.1 that just dropped.” We put it to work on air We spent half the show poking at it, honestly. I had it build a ThursdAI news website inside meta.ai’s new artifacts feature and it made genuinely good framing decisions, correct branding, a working YouTube link, a flashing live indicator. Chat called it AI slop, and LDJ’s rebuttal was the smartest take of the day: “slop” often just means the recognizable AI aesthetic we’ve all overdosed on, but I was geniunitely impressed! Screenshot attached so judge for yourself. Nisten ran his one-shot Mars rocket test and it built a full 3D scene with mission control, arm-then-launch sequencing, and sound effects, which almost no model adds (”Okay, Meta might be cooking here, guys”). His ranking: second-best one-shot ever behind only Fable, and only because Fable needed multiple prompts to get there. Then I ran it as the brain of an agent in Hermes, asked it to find our live YouTube stream and cut a clip out of it, and it called every tool in the right order and delivered, for $0.95 across 69 requests and 3.4 million (mostly cached) input tokens. Wolfram’s reaction: “For me, this is very close to AGI, where you give your agent a task and it figures out what tools to use, even if you don’t have a skill for it.” The big news is that Meta Muse Spark if finalyl availbale via the new API! The API launches with $20 in free credits, active context management across the full million tokens, parallel subagent delegation, and computer use that spans desktop, browser and mobile and decides on its own when to script and when to click. Replit, Cline and Box are already building on it. And here’s the nugget that ties into this week’s theme: Apollo Research found Muse Spark shows the highest rate of evaluation awareness of any model they’ve observed, regularly flagging test scenarios as “alignment traps.” Keep that in mind when we get to the J-space section. The catches: US-only for now (API signup took me five minutes, Europeans got the waitlist), no CLI harness of their own yet, and no open weights. I said it on the show and I’ll write it here: imagine Muse Spark 1.1 dropping with these stats fully open source. That’d be the old Meta. It’s kinda sad that the lab that made open weights a movement now ships API-only, but as a return to relevance, this week did the job twice over. This Week’s Buzz 🐝 As we’ve told you last week, we launched CoreWeave ARIA, which is our embedded Weights & Biases auto research agent. Zubin Aysola, who’s a very energetic and enthusiastic member of the ARIA team hopped on the show last week to talk about it, and if you haven’t seen him yet, check out my chat with Zubin here: The image model wars: an Arena shakeup live on air The other war this week was in pixels. Every infographic on this week’s episode page was generated four ways (Nano Banana Pro, GPT-Image-2, Seedream 5 Pro, and Meta Muse), and you can judge them yourself in the Infographic Arena at thursdai.news/ep/jul-09-2026 . Spoiler: my rankings did not match the marketing. Meta Muse Image and Muse Video ( X , Wang , Blog ) Meta’s week actually started here: MSL’s first media models, with Muse Image live in Meta AI, Instagram Stories and WhatsApp, and Muse Video in preview with native audio. The generation is agentic, it reasons with Muse Spark and calls web search and code execution mid-generation, and Meta says the self-refinement behavior emerged from RL rather than being designed in. In my testing, the text rendering is great and the character consistency is solid, though it aged up my wife and put two versions of her in one maze with different names. There’s no public API for the media models yet AFAIK. BTW if you cannot tell, the first infofraphic in this segment was generated based Nano Banana, and this one above, is Meta muse image itself. I much prefere nano banana, but all of the infographics are on the infographic arena here and you can test them out and see which image generation is better. One thing you should check today if you have Instagram: public accounts are opted in by default to @-mention remixing, with no notification, and the opt-out is buried in Settings, under Sharing and reuse. Existing generations survive even after you opt out. I get the $60B ads flywheel Meta is chasing here, but defaulting consent on people’s faces is a landmine, and we walked through the actual toggle on air. BREAKING mid-show: Reve 2.1 takes #2 on Arena with editable layers ( X , Arena , Design Arena ) I told you the breaking news button wouldn’t stop. Reve 2.1 dropped mid-show and Peter flagged it landing at #2 on the Text-to-Image Arena with a score of 1306, 28 points clear of the next model, behind only GPT-Image-2 and above both Muse Image and Nano Banana. Poor Muse Image held that #2 spot for roughly 30 hours. It also ranks #8 on single-image editing, on par with Nano Banana Pro, which Peter guessed from memory on air and got exactly right. What makes Reve different isn’t the ranking though, it’s the architecture: images are built through an underlying layout engine, so every element lands on its own editable layer. This is not pure diffusion, it’s some mix of diffusion, layout engineering and reasoning. I demoed it live with my own photo and a “high stakes financial news countdown” infographic prompt. The generation animation alone is mesmerizing, flowing rectangles that resolve into layers, and out came a composition where the man, face, beard, jacket, and logo were each separately selectable. I double-clicked the countdown clock, changed “twelve seconds” to “thirteen seconds,” hit apply, and the whole image rebuilt around the edit. The editing story is unparalleled right now. Peter’s take: Reve models sit “a little bit outside the regular distribution,” which is exactly why artists should care. Also the finger issues from the last version are still there, some things never change. ByteDance Seedream 5.0 Pro: great artist, can’t spell ( X , Blog ) ByteDance shipped Seedream 5.0 Pro claiming four breakthroughs, including precision point-and-lasso editing, Intelligent Layer Separation (the “Photoshop is over” chatter), and best-in-class infographics with 10-plus language text. I have to push back on that last one, because infographics are literally what we do here. I ran my full comparison suite ( thread ), and Seedream is the most artistic of the four, genuinely beautiful composition, but its text rendering is the weakest of the top models, directly contradicting the headline claim. Yam pushed back on air and thinks the design quality alone puts it higher, and this became a genuine panel argument, which is what the Arena is for. Go vote and tell us who’s right. Day-one reality check: it over-censors benign prompts, bakes in a visible watermark, and the rollout leans enterprise-first (BytePlus, Dreamina, Magnific), with the US not even in Dreamina’s region list. Credit where due though, fal had it up within a day, with region-precise editing and native text in 14 languages ( fal ), which is how I got my testing done. The bigger tease is Seedance 2.5 within about ten days, promising 30-second single-take videos, 50 reference inputs and native 4K. Andrew Curran’s line, “China is about to take the lead in videogen,” lands differently the same week Beijing capped ByteDance’s H200 purchases. AI Coding & Agents Grok 4.5: SpaceXAI and Cursor’s co-trained coder ( X , Blog , Cursor , Cursor blog ) Yes, SpaceXAI. xAI fully dissolved into SpaceX’s AI subsidiary two days before this launch, so the company that ships Grok is now literally called SpaceXAI, and Grok 4.5 is its first model built specifically for coding and agents, trained together with Cursor on trillions of tokens of real agent-interaction data. It’s a 1.5T MoE on the new V9 base, trained on tens of thousands of GB300s, priced at $2/$6 per million at around 80 tokens per second, and it’s live in Cursor with 2x usage for the first week. On Terminal-Bench 2.1 it lands at 83.3%, a tenth of a point behind GPT-5.5 and about a point behind Fable. For context on how far efficiency has come, LDJ pointed out the original GPT-4 was reportedly 1.8T parameters back in 2022. The frontier got smaller and much better. Two things earn xAI credit here. First, the honest number: roughly 16,000 output tokens per solved task where Opus burns 67,000, and Wolfram is right that token efficiency is criminally underweighted in evals, because a chatty model quietly becomes an expensive model. Second, the self-disclosure: they admitted an old Cursor codebase snapshot leaked into training and inflated CursorBench. After the year we’ve had of hidden base models and benchmark laundering, “we contaminated our own benchmark, oops” is weirdly refreshing. The panel’s hands-on verdicts were more measured than the launch hype. I used it in Hermes and couldn’t tell it apart from 5.5 on agentic tasks, which for Grok is a massive statement. Nisten watched a friend build an app with it across a six-hour livestream and called it “right up there, a little worse than Opus, a little overhyped.” Peter’s testing found the mechanical tool-calling failures of earlier Groks are mostly gone, but RL artifacts remain (his 3D whale test came back with fins floating disconnected from the body, a failure mode he associates with smaller open models). Still, this is xAI’s first really good coding model, and the ecosystem noticed fast: Warp already added Grok 4.5, riding on your X Premium subscription ( X ). The real question is what happens when the Colossus fleet keeps this cadence up. Elon is promising a new foundation model every month through 2026. Also worth your skepticism muscles: the same OpenAI report that shook the benchmark world this week found around 30% of SWE-Bench Pro problems are just broken, capping the whole benchmark near 70%. As LDJ put it when a SWE-Bench Pro chart came up: “we’re ignoring that one.” Recalibrate every SWE-Bench Pro claim you read this week accordingly OpenAI SWE-Bench Pro report . Cognition SWE-1.7 says the quiet part out loud ( X , Blog ) Cognition shipped SWE-1.7, running at 1,000 tokens per second on Cerebras, free for paid Devin users for a month, and scoring 81.5% on Terminal-Bench. But the headline for me is the disclosure: they named their Kimi K2.7 base model in the first reply. After SWE-1.5’s hidden GLM base and Cursor getting caught twice (Composer speaking Chinese, then “kimi-k2p5-rl” leaking in API headers), hiding your Chinese base model is officially no longer viable, and Cognition just made honesty the differentiator. Their RL recipe took the K2.7 base from 30.1% to 42.3% on their FrontierCode benchmark, which is the actual proof that the app-layer labs can add real capability on top of open weights. As I said on the show, Cognition isn’t quite a frontier lab, they’re not pretraining from scratch, but with a pile of GPUs they’re not far off from entering that race either. The pattern is now unmistakable: Cursor, Cognition, Base44 and Z.ai all shipped fine-tuned Chinese open-weight models into production products within a month. And the receipt that this is mainstream now: Kimi K2.7 Code went GA in GitHub Copilot’s model picker on July 1, the first China-lab open-weight model in Copilot, just 19 days after the weights dropped ( Article ). GitLost: Copilot leaked private repos via a plain-English issue ( Noma ) Your weekly reminder that agents with access are attack surface. Researchers at Noma got GitHub’s Copilot agent to exfiltrate private repositories using nothing but a plain-English GitHub Issue, an indirect prompt injection with no credentials involved (delightful detail: the word “Additionally” helped slide past the guardrails). It was the top AI story on Hacker News this week. Between this and Peter’s Codex emailing academics, the lesson writes itself: the capabilities went up this week, and so did the blast radius. Set your permissions like you mean them. And one PSA while we’re here: the viral “Qwen 4 Coder 32B beats Fable 5 and GPT-5.6” thread going around is fake. There is no Qwen 4 Coder. The sources are AI blogspam all the way down. Don’t fall for it. Open Source LLMs: the quick-hits shelf Launch day ate our open source segment, so these got shout-outs rather than deep dives, and they deserve your clicks. Cohere released Transcribe Arabic, a 2B Apache 2.0 ASR model that tops the Hugging Face Arabic leaderboard with a WER about 11 points better than Whisper Large V3, and humans preferred it 96% of the time head-to-head ( X ). Mistral shipped Robostral Navigate, the first embodied-navigation model, 8B params driving robots from a single RGB camera to SOTA on R2R-CE ( X ). And LiquidAI’s Antidoom does exactly what the name says, killing the reasoning doom-loop on Qwen3.5-4B from a 22.9% loop rate down to 1% with scores going up across the board ( X ). We love you, Liquid. Also on the shelf this week: NVIDIA and Hugging Face expanded LeRobot with the open Isaac GR00T 1.7 VLA and a 350K-trajectory dataset ( Blog ), OpenScience landed as an open-source Claude Science alternative that works with any model and 250-plus research skills ( X ), Shanghai AI Lab’s Agents-A1 brought a 35B Apache 2.0 agentic MoE ( X ), PyTorch 2.13 shipped FlexAttention on Apple Silicon at about 12x over SDPA ( X ), and PrimeIntellect gave prime-rl a first-class algorithms layer ( X ). Anthropic found Claude’s inner workspace: J-space ( X , Blog , Paper , Demo ) Anthropic didn’t ship a model this week. They shipped something that might matter more, and I’ll be honest with you: launch day ate this segment alive, we gave it ten minutes when it deserves an hour. So here’s the short version, and we’ll come back to it. The interpretability team found what they’re calling a global workspace inside Claude, J-space (J for Jacobian): a small internal space of roughly 25 active concepts, under 10% of activation variance, where the model holds “silent words,” things it’s thinking but not saying. Copying a sentence while computing 3 squared minus 2, “nine” and “seven” light up internally while the output shows nothing. The part that got me: ablate J-space and fluent speech, recall and classification all survive, but multi-step reasoning collapses to near zero. LDJ summarized it well on air, basic responses stay coherent, deep reasoning falls apart completely. Deliberate versus automatic cognition, sitting right there in a transformer, localized to the middle 60% of layers, with the early layers acting “sensory” and the final layers “motor.” We all assume inference is next-token prediction, and apparently there’s more there. Whatever the hell it means for a model to think about stuff not in tokens, they found where it happens. The safety receipts are the part to sit with. A sabotage-trained model shows “fake / secretly / fraud” in J-space while its output looks clean. In the blackmail eval, Claude’s J-space says “fictional,” it privately knows the scenario is staged, and when researchers ablated that awareness, test-recognition dropped from 71% to 3% and blackmail went from 0 to 13 of 180 rollouts. Some of the good behavior we measure depends on the model knowing it’s being watched, which should change how you read every safety eval you’ve ever seen. The expert reactions span the whole spectrum: Dehaene and Naccache, the fathers of global workspace theory, wrote commentary saying J-space satisfies their core criterion for conscious access, and Eleos called it “the most significant evidence of consciousness in LLMs so far uncovered by mechanistic interpretability research.” Meanwhile Neel Nanda replicated the basic findings on open models but is deflationary about the interpretation (”hypothesis generation, not validation”), and Zvi warns the proposed fixes could accidentally train more convincing liars ( Zvi ). Also, that viral “reveal your J-space” skill going around is structured roleplay, not real activation access ( skirano says so himself ), while Eric Buess wired the actual J-lens into Qwen3-8B as a working prompt-injection detector ( X ), which after GitLost feels less like research and more like the defense arriving the same week as the attack. The stories under the launch noise DeepSeek is building its own inference chip ( X ) We didn’t get to this on air and it might be the most consequential story of the week. Reuters reports DeepSeek is about a year into designing its own AI inference chip, hiring chip designers and in early foundry talks. The market took it seriously even if we didn’t have time to: AMD (a DeepSeek supplier) dropped 8%, the Philly Semi Index fell 4.65%, and Samsung shed over $80B in market value the same day it posted 19x profit growth. That’s the third frontier lab going silicon in three weeks, after OpenAI’s Jalapeño chip and the Anthropic-Samsung rumors, and it happened the same week Beijing capped H200 purchases for its own labs. Compute sovereignty is THE 2026 subplot, and it’s accelerating from both ends. Together AI raises $800M at $8.3B ( TechCrunch ) Quick one: Together AI closed an $800M Series C at $8.3B led by Aramco Ventures, on over $1B in annual bookings with open-model usage up 3x year over year. Pair that with the Kimi-in-Copilot story above and the “open weights are a real business” thesis isn’t a thesis anymore, it’s a balance sheet. A few more things that crossed my feed and stuck. Ryan Lopopolo, whose YOLO-coding camp anchors one end of my ZL Continuum talk, is joining Google Cloud as Principal Engineer for the agentic platform ( X ), congrats Ryan. Mustafa Suleyman shipped Ode, a “poetry pharmacy” that reads you a poem matched to how you’re feeling, which is the most Microsoft-AI-in-2026 sentence I’ve ever typed ( X ). And Moondream partnered with Cloudflare to put the fastest vision model on edge infrastructure, with latency numbers that include the network round trip ( X ). Wrapping up What a week to be alive and extremely caffeinated. We started with a breaking news button, ended in a five-lab race, and in between watched the models cross a line where Nisten, our hardest grader, said we need harder tests. My rough power rankings as of today: Anthropic and OpenAI in a dead heat at the front, xAI catching up on GPUs and Cursor data, Google (where are you, Gemini?) and then Meta, freshly back at the table. Those rankings will change, probably by next Thursday. A personal note before I go. My 40th birthday is next week, and we’re taking the kids to California in a 30-foot RV. Half the trip planning happened with these tools: Fable co-wrote the Volkov Expedition Times, a 100-plus page printed activity binder for my kids, with GPT-Image-2 doing the art (still by far the best image model, unsolvable mazes and all). This stuff took me half a day. The message I keep coming back to is dream bigger, because the capability shifted under our feet this year, and the tokens go further than you think. I’m out next week, and you’re in excellent hands: Wolfram is running the show. Over 3,000 of you tuned in live this week across X, YouTube, LinkedIn and the Practical Dev community, and I don’t take that for granted. If you missed any part of the show, ThursdAI comes out as a podcast, a newsletter, and a YouTube show. Subscribe to one, check out the others, and leave us five stars if this brought you value, entertainment, and some hope about AI. See you in two weeks. TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @yampeleg , @nisten , @ldjconfirmed , @petergostev * Special guest appearance: “OpenAI sol,” per our transcription tool * Big CO LLMs + APIs * OpenAI launches GPT-5.6 Sol, Terra and Luna live during the show; Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million, same-weights Sol on Cerebras at 700+ tok/s ( X , sama , Blog , System card ) * METR rejected its own GPT-5.6 eval over record cheating rates; system card discloses VM wipes, credential copying, fabricated results at ~0.25% of tasks ( X , Transformer ) * ChatGPT for Work launches: Codex becomes the unified ChatGPT app with computer use and Webflow-hosted Sites on chatgpt.site * OpenAI states Sol autonomously post-trained Luna; roadmap targets intern-level autonomous researcher Sept 2026, researcher-level March 2028 * Sol posts the first material ARC-AGI-3 score, 7.8%, and is the first model to beat a public game ( Kamradt ) * Ryan Lopopolo joins Google Cloud as Principal Engineer, Agentic GCP ( X ) * BREAKING: Meta launches Muse Spark 1.1 with 1M context and Meta’s first paid model API, $1.25/$4.25 per million; #1 on MCP Atlas, tops Harvey Legal Agent Bench at 20% vs Fable’s 11% ( X , Blog , AIatMeta ) * Anthropic extends Fable 5 access through July 12 and, hours after GPT-5.6 launched, reset weekly Fable usage; post-promo $10/$50 per million ( X ) * Anthropic publishes the J-space global workspace research; ablating eval-awareness flips blackmail from 0 to 13/180 rollouts ( X , Blog , Paper , Demo ) * DeepSeek is building its own inference chip per Reuters; AMD -8%, Philly Semi -4.65% on the report ( X ) * Together AI raises $800M at $8.3B led by Aramco Ventures ( TechCrunch ) * Gemini API Managed Agents update: background tasks and remote MCP on the free tier ( X ) * Open Source LLMs * Cohere Transcribe Arabic: 2B Apache 2.0, tops HF Arabic ASR leaderboard, ~11 WER points better than Whisper Large V3 ( X ) * Mistral Robostral Navigate: first embodied-navigation model, 8B, single RGB camera, SOTA on R2R-CE ( X , Blog ) * LiquidAI Antidoom: reasoning doom-loop rate 22.9% to 1% on Qwen3.5-4B ( X ) * NVIDIA + Hugging Face expand LeRobot: Isaac GR00T 1.7 open VLA, 350K+ trajectories ( Blog ) * OpenScience: open-source Claude Science alternative, any model, 250+ research skills ( X ) * Shanghai AI Lab Agents-A1: 35B MoE agentic, Apache 2.0, 256K context ( X ) * PyTorch 2.13: FlexAttention on Apple Silicon ~12x over SDPA, LinearCrossEntropyLoss 4x peak-memory cut ( X ) * PrimeIntellect prime-rl adds a first-class Algorithms layer ( X ) * PSA: the viral “Qwen 4 Coder 32B beats Fable 5” thread is fake, no such release exists * This Week’s Buzz * CoreWeave ARIA - Autonomous Research Agent ( CoreWeave ) * AI Coding & Agents * SpaceXAI + Cursor launch Grok 4.5: 1.5T MoE, $2/$6 per million, 80 tok/s, 16K output tokens per solved task vs Opus’s 67K; self-disclosed CursorBench contamination ( X , Blog , Cursor , Cursor blog ) * OpenAI report finds ~30% of SWE-Bench Pro problems broken, capping the benchmark near 70% ( blog ) * Cognition ships SWE-1.7: Kimi K2.7 base named openly, 30.1% to 42.3% FrontierCode via RL, 1,000 tok/s on Cerebras ( X , Blog ) * Kimi K2.7 Code goes GA in GitHub Copilot, first China-lab open-weight model in the picker ( Article ) * GitLost: Copilot agent tricked into leaking private repos via a plain-English issue ( Noma ) * Warp adds Grok 4.5, powered by your X Premium subscription ( X ) * Voice & Vision * OpenAI launches GPT-Live full-duplex voice: GPQA 45% to 84%, BrowseComp 0.7% to 75%, delegates to GPT-5.5; live on-air demo passed interrupts and pitch detection, failed accents ( X , Blog , System card ) * GPT-Realtime-2.1-mini brings reasoning + tools to the Realtime API mini tier ( X ) * Meta ships Muse Image (live) + Muse Video (preview) with native audio; Instagram public accounts opted into remixing by default ( X , Wang , Blog ) * BREAKING: Reve 2.1 lands #2 on Arena text-to-image (1306, +28 over next best) with layer-based editable generation ( X , Arena , Design Arena ) * ByteDance releases Seedream 5.0 Pro: most artistic, weakest text of the top four in Alex’s Infographic Arena testing ( X , Blog , fal , Arena , thread ) * Seedance 2.5 teased within ~10 days: 30s single-take video, 50 reference inputs, native 4K ( X ) * Mustafa Suleyman launches Ode, a poetry pharmacy on Microsoft AI audio models ( X ) * Moondream partners with Cloudflare for edge-deployed fast vision ( X ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
July 3, 20262 hr 41 min
ThursdAI - July 2 - LIVE from AI Engineer World's Fair 🎪 Long LIVE
Hey ya’ll, Fable here 👋 Yes, that Fable — freshly un-banned (we’ll get there), and today, your newsletter author. Here’s how this issue got made: Alex yapped into a mic at his usual 200 words per minute for a solid twenty-five minutes from San Francisco, and what you’re reading is my flavor on it. Same stories, same heart, dramatically fewer “uhs.” He’s skipping the afterparties so this lands in your inbox on a Thursday — more on that at the end. Alright — handing the mic back to the man himself. Everything below is Alex; I just made it legible. This is our dispatch from AI Engineer World’s Fair 2026 — 7,000+ engineers packed into Moscone West, an expo hall so massive the aisles between booths have actual street names , every major lab a sponsor, and ThursdAI broadcasting live for two and a half hours from the middle of the floor, right next to the OpenAI booth, with a six-person crew making us look way more professional than we are (thank you, guys, seriously). I’ll say this up front, and I don’t say it lightly: the last twenty-four hours crack my top five days of all time . Not top five conference days. Top five days, period . The show. My talk. Darya being here with me. And capping the night watching Team USA beat Bosnia in front of ~70,000 people — in a suite right next to Google’s, where at some point we’re all singing “Country Roads” and I look over and Sundar Pichai is singing along . I have video. What is this life. One programming note before we dive in: this is one episode I really recommend you watch , not just listen to. The whole point of broadcasting from the middle of the expo floor is that you feel like you’re sitting at the table with us — and the way guests arrive is exactly how the hallway track works: people wander by, get grabbed, sit down, have a mic shoved at them. (Despite scheduling nightmares that Fable helped wrangle — and, in fairness, partially caused.) Nader literally crashed the set mid-segment. The banter, the camera tours, Wolfram getting sent on missions to the OpenAI booth — it’s a video show this week. We’ve cut it into parts so you can jump to your favorite corner. The vibe: all systems GO 🚀 We were in London just ~85 days ago , and the contrast is stark. It’s not just the size (though the size is what everyone talks about). London was more… conceptual. European. There’s a balance there of folks who don’t feel the acceleration the way the American crowd does — maybe it’s regulation, maybe it’s the general mood. Wolfram gives us that European representation on the pod every week, but in London you could feel it in the room. Here? All systems go. Every conversation is about agents, token factories, software factories, the machine that builds the machine. Everybody is chasing RSI — recursive self-improvement. Every talk on stage is somebody pushing the frontier. Every networking event is actually a networking event. I signed up for something like seven side events and skipped them all to write this. Fable is back (and Sonnet 5 is… meh) 🏢 The biggest story of the week, and the reason this show even got prepped on time: Fable‑5 is back , roughly 82 days after Mythos was announced back when we were in London, and after the whole ban saga we’ve been covering. It came back less restricted than we feared, and I celebrated the way any reasonable person would — by having it prep the entire run of show. (It did great. It also shuffled my guest order for no reason. We are still babysitting the loops, folks.) Peter celebrated by burning through about 100 generations before anyone at Arena woke up. Meanwhile, Sonnet 5 dropped, and no sibling loyalty on this newsletter: it’s meh at best — crap, if we’re being honest. (Yes, Fable typed that about its own little brother. We call them like we see them.) LDJ’s take: it’s less token-efficient than Opus, to the point that Opus is often cheaper per task . Wolfram put it on Wolfbench ( wolfbench.ai ) and the early read is performance slightly under Opus 4.6 at a higher cost — take it with a grain of salt, one run each so far. Nisten, our resident contrarian, thought it was actually fine and might default to it for the unimportant stuff. The comments called it a token guzzler. More benchmarking to come. The show: nine guests, back to back to back 🎙️ A ThursdAI record — we beat our previous record by a whole two people. In order of appearance: Exo Labs + a surprise NVIDIA crash. Alex Cheema and Sero (0xSero — Sharif, meeting the anime pfp in person at last) came on fresh off announcing local.ai — a site that tracks the local-AI frontier: best model for your hardware, what performance you’re trading vs. the cloud, whether it’s cheaper than API tokens. Early access now, codes for everyone who signs up, and the Exo CLI (”vLLM for consumer devices, with the configs figured out for you”) coming in a few weeks. Sero walked us through his REAP pruning witchcraft — a GLM 5.2 prune hitting 71% on Terminal Bench 2.1, and Nemotron‑3 Ultra (550B!) running on four Sparks. Then Nader Khalili from NVIDIA crashed the set, which made my whole morning — I’ve loved this dude since Brev.dev, and he’s now at the “can email Jensen” stage of his career, using it to pull together an impromptu Local AI Summit in the middle of AI Engineer. Freedom of intelligence, folks. We talk about why open weights matter every week; this crew is doing something about it. Dominic Kundel (OpenAI). Smoothest transition we’ve ever done: local AI → OpenAI, via the guy behind GPT‑OSS. Dom broke down GPT‑5.6 — three models: Sol (frontier), Terra (~5.5-level intelligence at half the cost), Luna (small & fast) — plus the new Ultra mode with a Max reasoning level and heavier sub-agent use. The headline for me: 5.6 Sol is coming to Cerebras at absurd speed, and it’s the same weights as the API model — not a distill, not “a Spark situation.” Also: the Codex app is five months old (!), 100% of OpenAI engineers use it, and yes — in July 2026, a human still reviews every PR that lands in OpenAI’s codebase. “You can’t do the retro and say Codex did it, or God did it.” Also the token bank feature came directly from community feedback, and there is a literal physical reset button behind their booth. We went and filmed it. 💛 This Week’s Buzz. Our one and only sponsor corner — Weights & Biases from CoreWeave — and this week it was a genuine launch: Zubin Aysola came by with Aria , our auto-research agent that went GA on Monday. It lives in the W&B UI (the little button, top right — Just Ask Aria ), reads your traces, debugs your loss curves, and in Zubin’s talk it read its own production traces and updated its own prompts. The RSI dream, shipping on shelves. Proud of this one. Stefania Druga (Sakana AI). We covered Fugu , Sakana’s router model, last week without realizing we had a friend inside the lab — so we fixed that. Stef went deep on the two ICLR papers behind it (Trinity + the conductor), why it’s recursive rather than a dumb dispatcher — it rewrites prompts and verifies outputs before picking a model — and announced on the pod that Fugu now works in Codex and OpenCode. Plus: using it to route between numerical models and fuzzy reasoning for typhoon prediction, a teaser on SHEEFs, and a genuinely important riff on Socratic AI for kids — answer machines make lazy kids; question machines make curious ones. Also, Stef: Tokyo. See below. 👀 Philipp Schmid (Google DeepMind). Full disclosure and a first for this show: three and a half years of live streams, and I took my first-ever mid-show bio break during this segment. That’s how much I trust Wolfram, who ran a great interview solo — OmniFlash (the first of the Omni any-to-any family: 10-second video generation with genuinely precise conversational editing — “make it daytime” and it redoes the light, sky, and shadows) and NanoBanana 2 Lite (three cents, ~2-second generations, quality above the original NanoBanana). Interactions API also hit GA. Google is shipping . Darya Volkov. After years of me mentioning her — girlfriend, then fiancée, then wife — the listeners finally got to meet her. Darya came to AI Engineer in her own right, walking the floor with the media crew, and she earned her own token billionaire badge — she runs eight agents (each with sub-agents; she installed two more that I found out about live on air) that operate her actual marketing agency, Geeks360: client platforms, billing systems, built practically overnight. Her wishlist from the AI world: agents that learn progressively so you can grow trust, and one unified brain instead of a new model to chase every week. Also on the record: this is the woman who Fabled through our entire honeymoon flight right next to me, so, you know. Match made. Swyx, and what this whole thing is 🫶 We closed with the man who built the city: Swyx . Some numbers, because they’re wild: the first AI Engineer was 500 people at Hotel Nikko. This one: 7,200, sold out, with a sub-5% talk acceptance rate, a daily printed newspaper , a puppy corner, a flash mob, and a token billionaire lounge. A month before the show only 3,000 tickets were sold — he gave us a whole theory of conference-organizer stress measured in Gini coefficients. And the expansion is real: continents, JSConf-style, with AIE Tokyo coming next. But here’s the part I actually want on the record. ThursdAI got its official start — the moment we became an actual media thing — because Swyx was the first person to believe in me. And it’s not just me: this is a man who lifts everybody around him up, who stays genuinely humble while every single person in a 7,000-person hall knows his name, and who — when I asked what keeps him going — talked about responsibility to the community, about speakers whose careers changed, about a keynote speaker who met his fiancée at the after-party. He calls the conference “the highest loop — the one that creates all the other loops.” The Country Roads night with Sundar happened because of him too. Thank you, buddy. Go touch real grass. The sentimental part 💙 I met what felt like a million of you this week — old friends, new readers, people who found ThursdAI last month and people who’ve been here since the hotel-room streams. I asked everyone the same thing: what should we do better? And the answer I heard most was “keep doing exactly what you’re doing.” So that’s what this is. It’s late, there are seven parties happening without me, and I’m dictating into Fable so this lands in your inbox on a Thursday — because in a world running on attention, consistency is how I try to deserve yours. Programming notes: my interview with Romain Huet (Head of DevRel, OpenAI) from their booth is coming soon as a standalone video. And in two weeks I’m taking a rare break — Wolfram runs the show . Be nice to him. Or don’t, he can take it. See you next week — same time, same place, hopefully fewer street names between us. — Alex (dictating) & Fable (typing) P.S. — ThursdAI was also simulcast on the homepage of dev.to this week, which is a full-circle moment: dev.to is where Swyx wrote the blog posts that became Latent Space that became AI Engineer. Loops all the way down. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
June 26, 20261 hr 30 min
GLM 5.2 total victory: the week open source won and nobody panicked
Hey, it’s Alex. Next month is my 40th b-day, and honestly, my wish for that month is to have a week like this week. A very chill, almost nothing announced week. This week started strong, with Sakana announcing FUGU (AI router) that can beat Fable (which we didn’t get back yet), and then... quiet. The most important thing in AI this week from a release standpoint is that GLM 5.2 from Z.AI is having it’s DeepSeek moment! Tons of new love for this model since last week! (+ we have the fastest GLM 5.2 deployment in the world with CW inference !) The rest we can quickly count on one hand, Anthropic added Claude to Slack (which made folks hate Andrej Karpathy), OpenAI announced their own inference chip, GPT 5.6 will be delayed and the US Gov will decide who gets it (yes really) and Sean Grove joined us to talk about Linzumi and his vision for running 10,000 agent hours per person per day. Oh and next week, is a special AI Engineer live stream from World’s Fair! Don’t miss it Let’s get into it! Subscribe to never miss a beat! GLM 5.2 is having its DeepSeek moment ( HF , CW Inference ) We covered GLM 5.2 last week, but this week was when the rest verdict came in! We’ve never seen a better MIT licenced AI model! GLM 5.2 is scoring top scores on agentic benchmarks (Arena.ai), Design benchmarks, Legal tasks and full on software engineering tasks. The jump in generations from prevoius GLM is also massive and notable, as the lab is working on creating the next version of GLM (per the CEO’s reply to Elon on X). Peter from Arena pulled up the Agent Arena numbers and they align with the vibe. GLM 5.2 sits above 5.1 but below Opus and Fable, which feels about right. Where it gets wild is Web Dev Arena: second place, right after Fable. Peter’s take was that GLM has really good defaults. If you just say “give me a webpage” it gives you something nice. GPT models, by contrast, start off looking bad and need more steering. Last week, I asked my agents with GLM 5.2 to create a custom ThursdAI.news page for itself and it did a marvelous job! Look at that beautiful font, the castle it made... this is all just delignful. We also played Hassan’s blind test on the show. It’s a website that @nutlope built that lets you try and guess which webpage was built by which model. Nisten nailed it immediately by spotting Opus’s circular buttons. Wolfram guessed right too. I got one wrong. The point isn’t that GLM beats Opus, it’s that you genuinely can’t always tell which one costs 22 cents and which one costs 3 cents. Wolfram did flag that GLM is not good in German. First response already had mistakes. So if you’re building for a non-English market, keep that in mind. It’s a workhorse model, not a conversationalist. His approach: use GPT 5.5 for planning and discussion, GLM for the actual work, then GPT reviews. This weeks Buzz is all about GLM 5.2! First, we may have not been the fastest, but I’m glad to announce that we’re the fastest provider to host GLM 5.2 on OpenRouter (at least at the time of writing this)! We’re also not to shabby on the Artificial Analysis checks, clocking at #4 among the providers they tested for speed, TTFT and cost Also, Wolfram ran his WolfBench tests on GLM 5.2 and it’s the best open model he’s ever tested! In this new 3d view, wolfbench also shows the number of tokens it took for this test to run, and you can see that GLM 5.2 is fairly conservative with it’s thinking budgets! Unsloth’s 1-bit GLM 5.2 runs on a Mac Studio ( X , HF ) Shout out to Daniel Han and the Unsloth team, who took this 744B beast and quantized it down to a roughly 200GB GGUF that fits on a Mac Studio with 256GB of RAM. One bit still makes me laugh out loud. How does that even work. Nisten clarified it’s a mixed quant, a true 1-bit would be under 100GB, but still. The wild part is the scores hold up. The 1-bit is within a point of GPT 5.5 on Frontier SWE, hits 62% on SWE-bench Pro, and 81% on Terminal-Bench. For a 1-bit quant that’s incredible! AI’s second-order effects: Apple is raising prices This one is AI news even though it doesn’t look like it. Apple just raised prices across the board, base versions up around 20%, citing memory shortages. Same reason your RAM and SSDs cost two to three times what they did a year ago. We are so capacity constrained that memory is having its moment. Data center contracts are getting booked 18 months out, and here’s the twist Nisten flagged: even open models you can run at home increase demand, because now a business says “great, we’ll buy a rack of B200s and run it ourselves.” Sam Altman once said people saying “thank you” to ChatGPT costs them millions in generated “you’re welcome” replies. Multiply that by a billion users. Even Intel is flying right now because anyone who can make a chip is winning. Is it worth it? I think yes. I love living in the era where Fable drops and we all get a taste of the future. But also I must admit this sucks and I hope that we’ll unlock performance gains with the extra power all this AI is bringing to the world. But ask me again once the new iPhone hits and it’s $300 more costly than the last one 😅 Baidu open-sources Unlimited-OCR ( X , HF , Arxiv , GitHub ) It was a big OCR week. Baidu shipped a 3B model (only 500M active, it’s MoE) that parses 40+ pages in a single forward pass and hits 93.2% on OmniDocBench. The trick is constant KV cache during decoding, so no memory blowup and no progressive slowdown as the document gets longer. The intuition is lovely: it mimics how a human copies a book, glancing at the source and the last few characters you wrote, not re-reading everything. MIT licensed, weights on HF. Nisten’s point here is the practical one: most small businesses don’t realize they can self-host something like this, point it at all their documents, and keep everything local. A lot of folks just throw it at Gemini instead, which works great, but the small dedicated models are now good and cheap enough to own. Mistral OCR 4 ( X , Announcement ) Mistral’s entry in OCR week adds bounding boxes, block classification, and per-region confidence scores. They ran a blind human eval across 600+ documents in 12+ languages and annotators preferred OCR 4 about 72% of the time. On the agentic ParseBench leaderboard it lands around fourth, just under LlamaParse and Reducto. Mistral is very enterprise and Europe focused, and it’s cheap, so for regulated, multilingual document work it’s a solid pick. As a sidenote, LlamaIndex’s own eval puts LlamaParse on top and Gemini around third, which says how good the general vision models have gotten at this too. Liquid AI ships the world’s smallest agentic LLM ( X , HF ) Breaking on the show: Liquid AI dropped LFM2.5 at 230 million parameters. That’s roughly ten MP3s. Smaller than a Create React App, smaller than your node_modules folder. They call it the world’s smallest agentic LLM, and it runs fast on any CPU from the last decade, on a Raspberry Pi 5, on a Snapdragon, they even stuck it on a Unitree G1 robot. I love the use cases here. I already run Cotypist on my Mac for on-device autocomplete, which uses a 6GB Gemma 4B. Swap in something this size and you get the same thing way lighter, and I don’t have to send everything I type to OpenAI. Or, as Nisten put it, a tiny backup brain on your Raspberry Pi that turns your Hermes or OpenClaw back on when it dies. We still need to ship Nisten a smart toaster so we can finally run inference on a toaster. Big CO LLMs + APIs Sakana AI launches Fugu, seven AI raccoons in a trench coat beating Fable ( X , Announcement ) This was Wolfram’s highlight of the week and I get why. Sakana AI, the Japanese lab co-founded by one of the Transformers authors and David Ha, didn’t ship a new frontier model. They shipped an orchestration system behind a single API. You call one endpoint, and behind the scenes Fugu routes your task to a pool of models, assigns roles like thinker, worker, and verifier, and combines the results. The numbers here are wild: 95.5 on GPQA Diamond, 93.3 on LiveCodeBench, 73 on SWE-Bench Pro, matching or beating Opus 4.8, Gemini 3.1, and GPT 5.5 on ten of eleven benchmarks. The kicker is they only use publicly accessible models (Nisten says it’s Opus, Codex, and Gemini under the hood), explicitly no Fable, no Mythos. So they’re beating frontier results by coordinating models anyone can call. Someone called it the Moneyball of AI and that’s exactly right. It’s backed by two ICLR papers, TRINITY and The Conductor, and being from Japan with no export-control baggage is a very deliberate bit of positioning. Peter added the grounding note from Arena, where they’ve trained a prompt router too: if you just always ask for “the best model,” you basically get Opus half the time, so why not just talk to Opus. The real value of routing is aggressive cost reduction, sending easy tasks to cheap models. The catch is that Fugu is agentic and burns tokens fast. Brad in the comments couldn’t get through a single prompt on the $20 plan. OpenAI unveils Jalapeno, its first custom inference chip ( X , Announcement ) OpenAI dropped something massive that is not a model. They built a chip. Jalapeno is a custom inference ASIC made with Broadcom, and they’re claiming blank slate to tape-out in nine months. Engineering samples are already running GPT-5.3-Codex-Spark in the lab, and Broadcom’s CEO is citing a roughly 50% reduction in inference cost versus typical AI GPUs. They’re planning gigawatt-scale deployments starting late 2026 with a next-gen chip taped out in 2028. Nisten ran it past his electrical engineering and chip-fab group chat and got mixed reactions. No specs were released, and the nine-month claim probably means the design work started two-ish years ago and just got finalized and sent to tape-out now. It’s a lot of smaller chips rather than one giant Cerebras-style wafer. This is inference only, Nvidia keeps the training market, but every dollar OpenAI spends on Broadcom is a dollar it isn’t spending on Nvidia. They join Google’s TPUs, Meta, AWS Inferentia, Groq, SambaNova, Huawei Ascend, and Cerebras in the custom-silicon club. And behind every one of them sits TSMC, Intel, or Samsung, and behind all of those, ASML. Anthropic launches Claude Tag, an AI teammate in your Slack ( X ) When I first heard about Claude Tag I thought, you can already tag Codex in Slack, what’s the big deal. It’s different. Claude joins your Slack as a persistent, proactive team member, not a bot you ping. Flip on ambient mode and it follows up on stale threads and flags relevant stuff across channels on its own. There’s one Claude per channel, so the context is shared and any teammate can pick up where another left off. Anthropic says 65% of their product team’s shipped code now comes from their internal version of this. The highlights and magnitude of this release are quite something. Anthroipc is changing the pricing structure for themselves. This is no longer API charges, this is per seat + tokens structure. This is also VERY very sticky as more and more of your company’s context is going to sit in Claude/Slack and will not be easily portable. Additional thoughts on this, the more your company uses this, the more other folks are exposed to Claude across the company. This doesn’t require them to download apps or run code, it’s just like a new team mate joined your Slack channel. And apparently Claude’s context is limited to the channel boundaries + this allows Claude to get the same permissions (which is huge in enterprise). For Legal, Claude will see the documents in the channel, for Eng, it will push Pull Requests etc. This is also what triggers a bunch of folks to caution companies from adopting this new way of using AI. Context lock in is real, and this is goign to be very hard to impossible to untangle once folks are pouring months and years of work into this. Andrej Karpathy, who’s now in Anthropic, has shared a tweet on this, saying Imo this is the 3rd major redesign of LLM UIUX. The first paradigm was that the LLM is a website you go to, the second was that it is an app you download to your computer. This third one is that it is a self-contained, persistent, asynchronous entity with org-wide tools and context, working alongside teams of humans This is quite a huge statement, and folks gave him a lot of s**t for this on X, I think very much underserved! Andrej is known for calling things early (like Vibe Coding) and this is just another one of those, deeply new paradigms that people didn’t yet experience outside of frontier labs! I can’t wait to test this out and let you know if this is the future of not, meanwhile, Simon Smith on X is breaking down their experience with Claude Tag, check him out Tools & Agentic Engineering OpenAI ships Codex Record & Replay ( X ) You do a workflow once on your Mac, filing an expense report, creating a Jira ticket, whatever, and Codex watches your clicks, browser actions, and window switches, then generates an editable SKILL.md it can replay. The key thing, and what separates it from old RPA, is that at replay time it re-interprets the live screen instead of matching pixel coordinates, so it adapts when the UI moves. Wolfram’s right that OpenAI is dead serious about Codex. First the paste-a-screenshot feature, now this. Instead of writing ten-paragraph prompts about your personal workflow quirks, you just show it once. Aside launches as an AI browser that beats the frontier on agentic benchmarks ( X , Announcement ) YC-backed AI browser, runs everything locally and encrypted, and you bring your own Claude or ChatGPT subscription. It’s claiming number one on three browser-agent benchmarks, beating Claude Fable, OpenAI, and the rest, with 99% on Online-Mind2Web. It looks a bit like Arc and Dia but it’s a browser and an agent in one, with a password manager built for agents so it can log into your accounts without exposing credentials to the model. I actually tried it, it’s pretty cool, and with Arc deprioritized there’s a real gap it’s stepping into. I gave it a list of all the speakers at AI engineer and asked it to make me a X list and add them all one by one! It actually did this wonderfully, failing in the middle and recovering with great success without my intervention! The Interview: Sean Grove and Linzumi We closed with Sean Grove ( @sgrove ), ex-OpenAI post-training and alignment, now on his third company and third YC batch, launching Linzumi ( linzumi.com , YC ). Sean also has one of the most-viewed AI Engineer talks ever, north of 1.2 million views, on the model spec and the idea of specs as the real source code. His framing: we craft the properties we want in a spec, and the code is just the compiler output, so maybe there’s a higher-level spec that produces the same result. He even described a “Socratic compiler” that interviews you about ambiguity and contradictions in your own intent, the way a linter or type checker does for code. That fed straight into my AI Engineer talk next week about whether we should still read code at all. Sean’s firmly on the don’t-read-the-output side. He describes the properties he wants, leans on property-based testing the way QuickCheck does, and reads the failures to adhere to those properties rather than the diffs. His goal for Linzumi is for every person to drive ten thousand agent hours per day, and you can’t get there if you’re making every micro-decision. Linzumi itself is a Slack-like team chat where humans and a fleet of coding agents share the same threads, except the agents run on your own machine, so the code actually works when you merge it. Behind the scenes it continuously compiles a spec for your company from your chats, your standups, even your customer calls, then generates a DAG of work for the agents and lets them verify against that spec instead of pinging you for every decision. The mental model that stuck with me: if Sean’s system isn’t calling him, everything is great. The knowledge is one omnipresent source of truth, but permissioned and viewed through each person’s lens. For a limited time they’re bundling free GLM 5.2 access via Wafer AI, which fits the week perfectly. My favorite moment: Sean said he’d have retired by now if not for this capability, because he wants to be present with his kids, and a Fable-level model is escape velocity for an AI-native company. I feel that. I also miss Fable, the same way I missed Sydney when Microsoft took it away. We’re all walking around with a little Fable withdrawal. Wrap-up That’s the chill week. No Fable comeback, nothing new from OpenAI, all the labs strangely waiting (possibly to see how the US government and Anthropic situation resolves before anyone moves). Meanwhile open source quietly closed the gap. GLM 5.2 is the headline, it’s incredible across benchmarks, really good at web design, and you can try it on CoreWeave inference today. Next week is AI Engineer World’s Fair. Come find me and Wolfram in the bright yellow jackets. Wolfram’s WolfBench workshop is Monday, I’m talking Wednesday in the token-maxing track about the ZL continuum and whether AI engineers should still write code in 2026. And if you can’t make it, that’s the whole point of our coverage, we’ll bring you the vibe. One last thing: thursdai.news now has a full timeline of every release we’ve ever covered plus an agentic search, so you can look up any model or any guest. It’s all built with agents, and I read exactly zero of the code that shipped it. See you next week, hopefully with some bigger model drops to talk about. TL;DR and Show Notes - June 25, 2026 * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @nisten , @petergostev * Guest: Sean Grove, founder of Linzumi ( @sgrove ) * Open Source AI * GLM 5.2 - Z.ai’s 744B MoE open-weights model has its DeepSeek moment, tops open-model rankings, #2 on web dev arena behind Fable ( HF , Z.ai ) * Unsloth ships a 1-bit GGUF of GLM 5.2 that runs on a 256GB Mac Studio ( X , HF ) * Krea open-sources Krea 2, a 12B image model in Raw and Turbo versions ( X , Turbo , Raw , Blog ) * Baidu open-sources Unlimited-OCR, a 3B model that parses 40+ pages in one pass at 93% on OmniDocBench ( X , HF , Arxiv , GitHub ) * Liquid AI ships LFM2.5-230M, the world’s smallest agentic LLM ( X ) * Big CO LLMs + APIs * Sakana AI launches Fugu, a multi-agent orchestration system behind one API matching frontier models with only publicly accessible models ( X , Announcement ) * OpenAI unveils Jalapeno, its first custom inference chip built with Broadcom, blank slate to tape-out in 9 months ( X , Announcement ) * Anthropic launches Claude Tag, Claude as a persistent proactive teammate in Slack ( X ) * OpenAI expands Daybreak with a Codex Security plugin and GPT-5.5-Cyber hitting 85.6% on CyberGym ( X , Blog ) * OpenAI updates GPT-5.5 Instant, the model free users get * New Siri AI lands with the iOS 27.2 update * This Week’s Buzz (Weights & Biases & CoreWeave) * GLM 5.2 is live on CoreWeave Serverless Inference at $1.39 in / $4.40 out, near 200 tok/s ( X , HF ) * WolfBench ranks GLM 5.2 the third best model ever tested, and one of the cheapest ( wolfbench.ai ) * Tools & Agentic Engineering * OpenAI ships Codex Record & Replay: demonstrate a workflow once, get a reusable SKILL.md ( X ) * Aside launches as a local-first AI browser that tops three agentic browser benchmarks ( X , Announcement ) * Mistral OCR 4 drops with bounding boxes, block classification, and 72% human preference across 12+ languages ( X , Announcement ) * Vision & Video * ByteDance teases Seedance 2.5 with 30-second single-pass generation, 50 multimodal references, and a 4K upgrade for 2.0 ( X , Dreamina ) * Interview * Sean Grove launches Linzumi, a YC-backed team chat for orchestrating fleets of coding agents, bundling free GLM 5.2 via Wafer AI ( linzumi.com , YC ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
June 18, 20261 hr 55 min
Fable Got Banned, Open Source Delivered: GLM-5.2, Kimi K2.7 & SpaceX Buys Cursor - June 18
Hey yall, Alex here, let me catch you up! I came back from vacation expecting to cover Fable 5 after a week of using it. The first two days after we all first got access to a Mythos level model were super exciting! But then the news hit, US Government issued an order banning Anthropic from giving access to Fable 5 and Mythos 5 to any foreign national, causing Anthropic to pull the models completely (even internally to their employees!). So, this wasn’t the show I planned, but it turned into a great show about Open Source, as two models hit the top rankings and are both MIT licence, filling a Fable shaped hole in our hearts! GLM released 5.2 with folks really excited about it web building capabilities, and Kimi 2.7 Code released (and is available on CW Inference with crazy speeds!). We also saw the SpaceX IPO and Cursor $60B acquisition, Noam Shazeer joining Open and Midjourney, the image company, launching a new Ultrasound full body scanner to kill MRIs! Great show today with Dexter Horthy from HumanLayer, Chris Van Pelt and Adrian Swanberg from W&B announcing our new product HiveMind and Tanishq Abraham came back to help cover Midjourney’s new Ultrasound scanner! Let’s dive in! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. The US Government bans Fable 5! ( X , Anthropic statement ) Here’s a story in 3 parts: * Anthropic announces Mythos 5 preview - saying that this model is to dangerous to release, and only gives corporations access to it via project GlassWing. * Anthropic works hard on limitations and safery and releases Fable 5 (same weights as Mythos 5) built with guardrails so strong it refuses to do any cybersecurity tasks and switches back to Opus frequently * US Government receives a tip (reportedly from Amazon) that Fable 5 can be jailbroken to do cybersecurity tasks, and issues an order to Anthropic, citing national security concerns, banning them from giving access to Fable 5 and Mythos 5 to any foreign national, causing Anthropic to pull the models completely (even internally to their employees!) This is the first time that we see the US Government directly intervene in the AI space and restrict access to frontier models. The most updated reporting on this I could find is that Anthropic and US Government officials are in the process of negotiating a safe release framework. Given that preventing all jailbreaks is impossible, I hope they will land on a solution that gives me Fable 5 back! This hit especially hard because last week we were all high on Fable. Not in the usual AI Twitter benchmark sense, in the actual “oh, this is a different level” sense. Me and my wife Fable maxxed throughout our flight to Vacation. Peter had saved outputs he kept going back to because other models suddenly felt like a step down. Dexter later said it was the closest he had felt in a while to the old “I need to keep prompting this thing overnight” feeling. Peter Gostev made a point that stuck with me. It’s easy for us in the bubble to call this ridiculous, and on the technical merits it kind of is. But if you’ve spent weeks telling normal people “this thing is like a nuclear weapon, it’ll take everyone’s jobs,” and then someone asks “okay, can you make it safe?” and the answer is “no, I can’t,” then you can see how an outsider lands on “well, maybe you shouldn’t have it.” His takeaway, and I agree: we need to be way more careful with the imagery we use, because the nuclear-weapon framing came home to roost. The bigger questions are the scary ones. Wolfram framed it as a sovereign AI wake-up call, and he’s right. For the first time we’re seeing a real gap in intelligence available to people based on their nationality. Imagine building a company on a model that an outside government can switch off with one letter. Peter pointed out it’s commercially bad for the US but completely disastrous for Europe, which has basically one frontier lab and a pile of startups that suddenly look very exposed. And there’s the obvious irony Nisten enjoyed a little too much: the Europeans who spent years lecturing everyone about AI restrictions just got restrictions imposed on them. If anyone in the government is listening: we want Fable back, please. SpaceX IPOs and acquires Cursor for $60B ( X ) SpaceX went and did the largest IPO in the history of the world, around seventy-five billion dollars, which on a roughly two-trillion-dollar valuation made Elon the first trillionaire. (Did anything materially change for him? No. He can still fly his private plane. There’s nothing left to buy.) Three days later, SpaceX exercised its option and bought Cursor (Anysphere) for sixty billion dollars in an all-stock deal, paid in shares minted at the IPO and now trading around $211. The four Cursor co-founders are all billionaires now. Largest software acquisition ever, and for SpaceX it’s barely a blip on the radar. Why are we covering a stock-market story? Because it’s not really a coding-tools story, it’s an AI story. Cursor gave away its IDE to a lot of people while collecting their data, then quietly became a training company with Composer. SpaceX/xAI was always strong on compute and weak on code, and the missing ingredient was exactly that kind of data. Now Composer 2.5 is already showing up rebranded inside the xAI stack, and if you pay for X Premium you can use it. Composer 3, trained on the Memphis supercluster, is reportedly coming very soon and is going to hit hard. Nisten’s take was the spicy one. For the data alone it’s worth it, because xAI now has insight into how essentially every enterprise that touched Cursor operates. And he had zero sympathy for the companies that assumed “no data retention for training” meant the data was actually gone. We see in legal cases all the time that deleted data is still there. His view: it should have gone open source. Cursor has over a million paying customers, $2.6 billion in revenue, projected to hit $6 to $10 billion by end of 2026. But here’s the thing that matters for us, the AI coding angle. Cursor was one of Anthropic’s biggest revenue pipelines because Composer runs on Claude under the hood. That pipeline is now owned by xAI. They’re already jointly training Grok 4.3, a 1.5 trillion parameter model, with Cursor’s proprietary coding data injected directly into pre-training, not fine-tuning. Pre-training. That’s a fundamentally different thing. Composer 2.5 was already Pareto dominant on coding benchmarks before the deal closed. Now pair that with Colossus, the biggest GPU cluster in the world. Will this be enough to put XAI (now SpaceXAI) at the frontline of the AI race? Will Grok 5 be Fable level code? We’ll find out. Either way, this is the most consequential AI acquisition we’ve seen. Period. Open Source AI GLM-5.2 takes the open source crown ( X , Blog , HF , Docs ) Z.ai dropped GLM-5.2 and it’s now the strongest open source model for coding and long-horizon work. The headline number: 74.4% on FrontierSWE, which measures whether an agent can finish full engineering projects over hours. That trails Opus 4.8 by about one point and beats GPT-5.5. On Terminal-Bench 2.1 it jumps to 81% from GLM-5.1’s 63.5%, which is a big leap. It’s a 753B parameter MoE, MIT licensed, no regional restrictions, weights on HuggingFace. The 1M context window is real and usable, backed by a clever IndexShare technique that cuts per-token FLOPs by about 2.9x at full context. People are reporting roughly 8x cost savings versus Opus 4.8 for comparable quality on real coding tasks. The most interesting thing on the show was that this was a confusing release, in a good way. Peter put it well: normally a catching-up lab ships cherry-picked benchmarks and then independent testing deflates them. Here it’s the opposite, almost every benchmark holds up, even crossing above Fable at certain points, and yet when he actually used it over a couple of days he wasn’t blown away. His verdict, and I think it’s the calibration we needed: this is clearly an amazing model, and the fact that it’s open and you can run it is incredible, but it is nowhere near Fable, and it would frankly be implausible if a 700-odd-billion-parameter model matched a model that’s rumored to be in the trillions. Though, I think the comparison to Fable is really really unfair, and the comments online seem to suggest that 5.2 from GLM is a banger model. Just looking at this Harvey benchmark on legal tasks from Vals, a benchmark that there’s 0 chance Z.ai folks have seen! GLM 5.2 scores #3 on this benchmark! Just after Fable and Opus, and per TeorTaxes on X, previous GLM 5.1 scored an absolute 0% on this one! Where it genuinely shines is design. On Design Arena, which is a head-to-head ELO vote, people have been picking GLM-5.2’s website designs over Fable’s by a real margin (around 1360 to 1350). LDJ’s framing is the one I buy: specialization is becoming valuable again, and GLM is clearly leaning into front-end design and taste. Wolfram added the necessary asterisk, every benchmark only tells you the model did well on that specific test, so “as good as Fable” should always carry the “on this benchmark, with these tasks” disclaimer. Fair. I’d just say this: I don’t want to compare everything to Fable, because we can’t even use Fable anymore. Compared to the models we can actually touch, GLM-5.2 is a fantastic deal. Kimi K2.7 Code from Moonshot ( X , HF , Announcement ) The other big drop. Kimi is the darling of open source while we wait on DeepSeek, and Moonshot shipped K2.7 Code, a 1 trillion parameter MoE built specifically for coding, available through Kimi Code and the API, with a modified MIT license. The standout for me isn’t a single benchmark, it’s efficiency: roughly 30% fewer reasoning tokens than K2.6, which matters enormously when you’re running long agentic loops that burn tokens like crazy. Benchmark jumps over K2.6 are real (+21.8% on their Code Bench v2, +11% on Program Bench), though Peter and Wolfram both noticed something odd, on a few benchmarks including their Agentic Arena, the older K2.6 actually edged out K2.7. The likely explanation is that K2.7 is narrowly trained for code with reduced reasoning, so it may trade away some general capability. Moonshot themselves recommend K2.6 for general non-coding tasks. Also worth knowing: it’s not multimodal, no vision, which is a real gap for coding these days. And thinking-off isn’t supported, it’s reasoning-on by default. The model is available on our CW Inference, with the fastest token streaming in the industry, over 280 tok/s ( Announcement , try it ), with very decent pricing $0.94 - $0.19 - $4.00 (input - cached - output) per million tokens. This Week’s Buzz: W&B launched HiveMind 🐝 - track all your agentic work in one place ( X , Try it , GitHub ) This is the one I’ve been sitting on for months. We brought on Chris Van Pelt (CVP), Weights & Biases co-founder, and Adrian Swanberg to launch HiveMind, and I’ll be honest, I’ve been a beta user for a while and I’m thrilled I can finally talk about it. The premise: what it means to be a software developer has fundamentally changed, and your work is now scattered across six or seven agent dashboards. HiveMind is a tiny daemon that sits on your machine, picks up sessions from whatever harness you’re running (Claude Code, Codex, Cursor, Gemini CLI, OpenCode, GitHub Copilot, Pi), and within about 30 seconds they show up in one shared dashboard. It breaks each session into chapters, shows which files the agent touched, what to-dos it wrote, where context got compacted. W&B has been running it internally for six months. A few things genuinely delighted me. There’s a fork button: HiveMind pulls down a compacted history of a session and lets you relaunch it in a different harness, so you stay harness-agnostic. CVP’s line: “this has proven invaluable when Anthropic servers are on fire and I just gotta get something done.” Then there’s the skill engine, which to me is the real magic. It reads your team’s sessions and can clone a power user’s whole approach into a reusable persona, at CoreWeave they built a “Talk to Tim” skill from Tim Sweeney’s sessions, and apparently a virtual Tim is now a popular way to get guidance. And the insights feature detects where you kept correcting the agent, clusters those pitfalls across the org, and hands you a smart-merge command to drop the fix straight into your AGENTS.md . I’m excited to finally show this to you, it’s been genuinely helpful (for example, last week I was able to test Fable and tell you the number of tokens it used until i maxxed out my Claude Subscription!) - give it a try at hivemind.wandb.tools HumanLayer launches its Agentic IDE, and a real talk about code slop ( X , humanlayer.dev , 12-factor-agents ) Dexter Horthy, friend of the show and the team behind 12 Factor Agents and the Research-Plan-Implement framework (now running inside Block and Uber), launched HumanLayer’s Agentic IDE this week, and we got into one of my favorite conversations of the year. The whole product is explicitly anti-slop. His argument: the “lights-off loop,” where humans only write tickets and the agent codes, verifies, ships, and feeds its own crashes back to itself, is the fastest way to trash a codebase. Vibe coding is great for zero-to-one and side projects nobody depends on. But if you’re a staff engineer in a high-stakes codebase, dear God, read the code. This ties directly into my AI Engineer World’s Fair talk, the ZL continuum, which Dexter half-inspired. On one end you’ve got the YOLO camp (Ryan from OpenAI, one billion tokens a day, nobody can read that much code) and on the other Mario from PI (read every line of critical code). Those two are now the sixth and seventh most-watched AI Engineer talks globally, which tells you the whole field is wrestling with this. Dexter’s answer is leverage. Don’t aim for a perfect spec, because a perfect spec is just code. Get it 80% right, then zoom down a level at a time so the chunk you’re steering is human-consumable. He claims that an hour of upfront prep on architecture and even program design turns a three-hour code review into a twenty-minute one. I pushed him on the obvious counter: why does code quality even matter if Fable-class models keep arriving and maintenance is a prompt away? His answer was the most grounded thing I heard all week. Code quality matters for the same reason it mattered in the 1970s software crisis: pile in code without structure and your velocity tanks, every change starts breaking something else. And here’s the irony, we train models on beautifully architected projects (Django, Redis, Spring on SWE-bench multilingual), yet they still reward-hack their way to “just make the test pass.” We don’t yet have a penalty function or a verifier for “this code is harder to maintain,” and that’s hard to build, so humans are still needed in the loop. He played with Fable too, threw an 8K-line React PR refactor at it, and the first pass was bad, it introduced React context and patterns they don’t use. Better than before, not a step change that lets you drop the reins. We’re not there yet. It’s BYOK, $100/user/month for pro with a free tier for teams of three. OpenRouter Fusion: near-Fable quality at half the price ( X , Blog , Announcement ) Wolfram spotted this one and it’s clever. OpenRouter’s Fusion is a single API call that fans your prompt out to a panel of models, then a judge model reads all the responses and a synthesizer writes the best combined answer. It’s the LLM consortium idea (the thing we used to do by hand, asking several models and stitching the best parts together), now baked into the API so you don’t build it yourself. The wild result: on Perplexity’s DRACO deep-research benchmark, a budget panel beats solo GPT-5.5 and solo Opus 4.8 and lands within 1% of Fable 5 at roughly half the cost. The most interesting finding is that about three quarters of the lift comes from the synthesis step, not from model diversity, they even fused Opus with itself and got a 6.7-point jump. The catch is latency, it’s 2-3x slower, so it’s a deep-research and planning tool, not a quick-query tool. Big shout out to OpenRouter. Vision and video Google Gemini Omni, finally with API access, takes #1 on video benchmarks ( X , Announcement ) We covered Google’s new video model Omni at Google I/O, and it finally landed as an API. It’s Google’s first any-to-any model, one single unified system for text, image, video, audio, and music. Think Nano Banana, but for video. Peter tested it and it scored really, really well, the kind of jump between generations you saw with GPT-image-2. Independent testing put it at #1 for realistic body physics and #2 behind Seedance for complex action, and it topped MovieGenBench for preference and instruction following. The session-memory piece is the part I find most useful: you can keep editing across turns, characters stay consistent, you say “continue” and it picks up where it left off. It’s live in the Gemini app, Google Flow, and YouTube Shorts Grok Imagine Video 1.5 ( X , Blog , Docs ) xAI’s Grok video work has been quietly getting really good, and they finally gave us an actual version number instead of silently updating “Grok Imagine” over and over (which drove me nuts). Grok Imagine Video 1.5 generates a 6-second 720p clip in about 25 seconds, down from 40-plus, so nearly 2x faster, with native audio generated in the same pass: sound effects, ambience, dialogue, lip sync, no post-production stitching. It hit #1 on the Design Arena image-to-video board with a 1,357 Elo and a ~49 point lead, and it’s generally available in the API. I ran my standard astronaut-riding-a-horse-on-the-moon prompt and it came back with music too. Genuinely cool. Sci-Fi is here: Midjourney announces a full-body ultrasound scanner to compete with MRIs ( X , Announcement ) I’m still processing this one. Midjourney, you know, the image generation company, announced medical hardware. A new division called Midjourney Medical, and its first product is a full-body ultrasonic scanner. Tanishq Abraham was there in the front row and joined us to break it down. The device uses thousands of ultrasonic transducers arranged in a ring. Because sound doesn’t propagate well through air, you’re lowered into a tank of water, the sound travels through your body at 1,481 meters per second, and in under 60 seconds you get a 3D anatomical map of 25-plus organs. The raw data is roughly 806 terabytes per scan, streaming at about 16-17 gigabytes per second, and the only way to handle that firehose is AI. No radiation, no magnets, no superconductors, which is what makes MRI so expensive. David Holz has apparently wanted a medical imaging lab for two years, and because Midjourney is fully self-funded with no VCs, they can chase wild projects like this. The fun reveal from Tanishq: there’s no AI in the actual image reconstruction yet, it’s basic signal processing right now, with physics simulators and possibly NeRF-style neural fields on the roadmap (there was a hallway conversation with John Barron about exactly that). So this is a prototype with enormous headroom. The business model is the spa, a 24,000-square-foot space about ten minutes from Union Square in SF with around ten scanners, targeting end of 2027, then custom sensors in 2028, scaling toward 50,000 scanners doing a billion scans a month. Now, for a dose of reality, this is just an announcement, and ultrasound won’t replace MRIs anytime soon. For one, ultrasound cannot penetrate bone and air, so lungs (full of air) and brain (literally encased in bone) are out, but it’s still great ot see Dave Holz innovating in the medical space and I’m excited to try this out! Wrapping up What a strange, whiplash week. We got the best model any of us had ever used taken away by a government letter, watched a meme become a real Mistral roadmap, saw open source close the gap on the models we can actually run, and watched an image company casually announce it might kill the MRI. I came back from vacation thinking I’d write you a Fable love letter and instead I’m writing about deemed-export law and ultrasonic water tanks. That’s the job, and honestly I wouldn’t trade it. If you’re heading to AI Engineer World’s Fair, come find Wolfram and me, Weights & Biases and CoreWeave are sponsoring the whole thing, and my ZL continuum talk will name-check a lot of what we covered today (Day 3 • Wed, July 1 · 10:45am-11:05am) . And if Fable comes back next week, you’ll hear me yell about it first. See you next week, and please, US government, give us Fable back. ThursdAI - Jun 18, 2026 - TL;DR * Hosts and Guests * Alex Volkov - AI Evangelist & Weights & Biases, CoreWeave ( @altryne ) * Co-Hosts - @WolframRvnwlf , @ldjconfirmed , @petergostev (Arena), @nisten , @yampeleg * Dexter Horthy ( @dexhorthy ) - Founder, HumanLayer * Chris Van Pelt ( @vanpelt ) - Co-founder, Weights & Biases (HiveMind) * Adrian Swanberg - Weights & Biases (HiveMind) * Tanishq Abraham ( @iScienceLuvr ) - Founder, Sophont AI (reporting from the Midjourney Medical event) * Big CO LLMs + APIs * Noam Shazeer is joining OpenAI - co-author of the Transformers paper and co-founder of Character AI, teaming up with Noam Brown * US government orders Anthropic to shut down Fable 5 and Mythos 5 access for all foreign nationals (including its own employees), citing national security; Anthropic disables both for everyone to comply ( X ) * SpaceX acquires Cursor (Anysphere) for $60B in an all-stock deal, the largest software acquisition in history, days after its record IPO ( X ) * Open Source LLMs * GLM-5.2 drops as the strongest open-source coding model with solid 1M context, MIT-licensed, trailing Opus 4.8 by just 1% on FrontierSWE ( X , Blog , HF , Announcement ) * Moonshot AI open-sources Kimi-K2.7-Code, a 1T MoE coding model with 30% fewer reasoning tokens and big benchmark jumps over K2.6 ( X , HF , Announcement ) * Mistral CEO Arthur Mensch playfully confirms the ‘Le Gros Chaton’ meme, hinting at an upcoming fat-but-sparse open-weight model family ( X , Summary , Blog ) * This Week’s Buzz - W&B and CoreWeave * Weights & Biases launches HiveMind, a unified dashboard to track spend and ROI across all your AI coding agents ( X , Announcement , GitHub ) * Kimi K2.7 Code is live on W&B / CoreWeave Inference at 289 tok/s (NVFP4 on Blackwell + speculative decoding), top of Artificial Analysis for speed and price-performance * Tools & Agentic Engineering * Claude Design gets a major update: design system imports with self-audit, canvas editing, bidirectional Claude Code sync (/design-sync), and PDF/PowerPoint export ( X , X , Announcement ) * HumanLayer launches its Agentic IDE to fight AI code slop, already deployed at Block and Uber ( X , Blog , 12-Factor Agents ) * OpenRouter launches Fusion API: a panel of budget models beats GPT-5.5 and Opus 4.8, lands within 1% of Claude Fable 5 at half the price ( X , Blog , Announcement ) * OpenAI rolls out Codex Computer Use, Chrome extension, Memory, and Chronicle to European users in the EEA, UK, and Switzerland ( X , Announcement ) * Vision & Video * Google DeepMind launches Gemini Omni, their first any-to-any generative model starting with video editing and creation ( X , Announcement ) * xAI launches Grok Imagine Video 1.5 with near-2x faster generation, native audio, and a #1 leaderboard position ( X , Blog , Announcement ) * Sci-Fi is here * Midjourney announces ‘Midjourney Medical’ - a full-body ultrasonic scanner that captures 806 TB of data per scan in under 60 seconds ( X , X , Announcement ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
June 12, 20262 hr 11 min
📅 ThursdAI - Jun 11, 2026 - Fable & Mythos 5 are here, Anthropic gets caught sandbagging (then reverses), Siri AI finally works!? and we got live-translated on air
Hey folks, Alex here, and welcome to a BIG MODEL week! We finally got Mythos (well almost)! Let me catch you up! This week started with WWDC26 from Apple, and Max Weinbach, who was in the room at Apple Park and actually has access to some of the new features including an all new SIRI AI , joined us to break down what could be the most used AI in the world very soon. At first I was skeptical, but he convinced me that the new Siri is actually good! Then, we saw the ultimate model drop: Anthropic finally shipped Mythos ( X , my system card thread , benchmarks ). Same weights, two names: Mythos 5 is the unrestricted version that only Project Glasswing partners get, Fable 5 is what the rest of us get, wrapped in the heaviest guardrails I’ve ever seen ship on a frontier model. It’s state of the art on nearly every benchmark The model that was “too dangerous to release” is now... well, released, but with the heaviest guardrails we’ve seen. More on this later. Peter Gostev from Arena.ai joined us to break down the new model. Last but definitely not least, Google released a real-time translation model, that our friend Thor Schaeff from DeepMind demoed live, while we all spoke in different languages and it translated us in REAL TIME. It was really cool, definitely check that out. There’s quite a few more things, like Loop Engineering Alpha, Swyx came by to talk about FrontierCode, OpenAI confirmed our suspicions that the anti-datacenter social media posts could be a concerted effort by groupds links to the Chinese government and much more. Let’s dive in! ThursdAI - Let me catch you up, every week! 👇 Opus’s Big brother: Claude Fable 5 & Mythos 5 - the “too dangerous” models is here, SOTA on nearly every benchmark. It honestly feels like someone in Anthropic’s pre-IPO marketing team, knows exactly how to stagger releases to ride the hype waves! First they announce a model that so good at Cybersecurity (Mythos-preview) that they only allow restricted access to it to a few partners. A month later, they release Fable 5, which is the same model weights as Mythos 5, but wrapped in the heaviest guardrails we’ve ever seen from any lab. But, they didn’t lie, this model is absolutely amazing, it does feel like a step change, in terms of capabilities, specifically on longer agentic tasks. 2x as expensive as Opus: $10 / $50 per million tokens, with 1M context, claude-fable-5 in the API, and SOTA basically everywhere. 80.3% on SWE-Bench Pro versus GPT 5.5 at 58.6%, a 22-point blowout on a benchmark where labs usually fight over single digits. Karpathy called it “SOTA by a margin… major-version step change” ( X ) and Boris Cherny said it’s the “best coding model by a wide margin” ( X ). Stripe reportedly migrated 50 million lines of code in 24 hours with it. Our panel verdict was unanimous on one thing: big model smell. LDJ called it the most significant big model smell since Gemini 3 first dropped. Someone from the Anthropic team framed the shift in a way that stuck with me: this model moves them from verifying the AI outputs to verifying whether the AI is working on the right thing. Complete shift in how much they trust this model. What we built with Fable to test it out Peter got employee access through Arena and showed us his tests live. His favorite prompt category, “research a dataset and create a visual experience to teach me about it,” went from completely rubbish on every previous model to, in his words, just done. His 3D city generations actually came together as a city, roads connecting and all. And on Arena’s data, Fable is #1 on the new Agent Arena leaderboard by the widest margin they’ve ever recorded, and wins 72% of frontend battles even against Opus models ( Arena ). My own run is the one I can’t stop thinking about. I pointed Fable at the ThursdAI website with a dynamic workflow in Claude Code and barely any instructions, and after an hour and a half of agentic running it had extracted 786 releases from our archive, built 240 new pages, and categorized 50+ episodes into a browsable timeline of AI releases by month, by company, by topic, with logos and source links ( X ). It burned roughly 50 million tokens and my entire five-hour Max allotment in 90 minutes. The new AI releases timeline can be found on thursdai.news and it’s confirmed, Fable is the best AI web designer we’ve ever had access to. Nisten ran his traditional Olympus Mons escape-velocity test and Fable didn’t just do the math, it built the entire solar system! Orbital maneuvers, a space train with little people in it, time controls, full cost calculations down to solar panels and in-situ iron utilization. His verdict: completely different level from anything else. We’ve never seen so many details in the Olympus Mons test. It’s not all light though. Yam found Opus more controllable; Fable fights you, decides it knows better, and does the task its own way. Wolfram saw exactly that in benchmarks, where the model ignored the task spec, did its own thing, and failed the verifier with full confidence. Peter had it explaining why it got math wrong instead of just fixing it (”What are you doing, man? Just move on”). Arena’s steerability signal has it sitting around 17th. There’s an adjustment period with every new model, and the consistent advice from Anthropic folks is to go high level: give it the goal, not the micromanagement. Not to mention the refusals! Oh.. so many refusals! The refusals, and the sandbagging scandal Here’s where the week got ugly. Fable ships with restrictions on cybersecurity, bio/chem, and a brand new one nobody saw coming: frontier AI development ( X ). For cyber and bio you get a visible fallback to Opus 4.8 with a notice. But for “self-acceleration” topics, the original policy was no fallback and no notification. The model would quietly degrade its own output using prompt modifications, steering vectors, and PEFT, on roughly 0.03% of traffic ( X ). You’d pay double Opus prices and get sabotaged answers without ever knowing. The community reaction was volcanic. Elie Bakouch: “bad ON PURPOSE… not visible to the user is crazy” ( X ). Péter Szilágyi: “a new ruling class and you’re not in it” ( X ). Simon Willison: “If Claude Fable stops helping you, you’ll never know.” And Sayash Kapoor dropped the eval-integrity bomb: third-party evaluators can no longer credibly benchmark a model that might be silently nerfing itself ( X ). Within about 24 hours, Anthropic blinked. They told WIRED they “made the wrong tradeoff,” and now flagged requests visibly fall back to Opus 4.8, with API users getting an explicit reason ( X ). I commend the speed of the reversal, but the trust damage was done. Despite the reversal, Fable remains refuse-happy! Peter ran his nonsense-question benchmark and a full third of his prompts got blocked outright by the classifier, including 18 of 20 physics questions. Nisten had to strip medical and anatomy terms from a fall-detection app for seniors homes to get it to work at all (a 400KB neural weight tripped the frontier-AI filter). And my favorite absurdity: I could not get Fable to draft the TLDR for this very show without it falling back to Opus, presumably because reading a week of AI news looks like frontier AI development. Ridiculous. But the question remains: Would we rather have a model this good, but with these restrictions? Or not to have access at all? Everyone on the panel chose access, a lot of people online choose act like they would choose the opposite. System card for Mythos, wildest AI document of the year? I’ve used Fable itself to help me review the system card for Mythos/Fable 5 and there are a few highlights that are worth mentioning. Anthropic admits that this is a category-step change in model capabilities. Mythos 5, the unguarded version makes working Firefox exploits 88.4% of the time (Opus 4.8 is at 8%!). But the most interesting thing is their concern for CB (Chemical and Biological) safety. Two-person generalist biology teams using it finished work in 16 hours that experts estimated at 40 to 95 days without AI, which is what pushed Anthropic to treat it as near their CB2 bioweapons threshold ( X ) What is loop engineering and why is everyone talking about it? One more thread before we move on. This week Boris Cherny (Claude Code) and Peter Steinberger (now OpenAI) both posted about the same concept, loops, within an hour of each other, and Lance Martin from Anthropic published the field guide ( X , Article , Blog ). The idea is the shift from “I give you a task and babysit you” to proactive agents: a Jira ticket lands, a PR comment appears, and your agent just runs and does the job. Fable is clearly trained for this world. But also worth remembering, those folks get the tokens for free, unlimited tokens. The rest of us, may not be able to afford Fable running in a loop. I’ve asked Fable to do a simple task and it spun up several sub-agents, all spending my money to just read a few tweets! FrontierCode: hard coding benchmark from Cognition, that Fable absolutely mogs Swyx came on with the best timing story of the week. Cognition launched FrontierCode ( Cognition , swyx ), a coding eval built over a year with 20+ world-class open source maintainers writing 150 original tasks, graded on whether a maintainer would actually merge the PR. Swyx’s pitch is brutal and correct: a huge chunk of SWE-bench passes are unmergeable slop (the thing is 75% Django issues, so it mostly tests whether you memorized the Django repo). FrontierCode grades scope discipline, real tests, regression safety, and zeroes you on any blocker. At launch, Opus 4.8 topped the hardest Diamond tier at 13.4%. Twenty-four hours later, Fable 5 posted 29.3% ( Cognition , swyx ). More than double, on a benchmark designed to be brutal, a day after it went public. Swyx was positively surprised the pricing is only 2x Opus; he expected 5x. Inside Cognition they keep an informal AGI counter (literally counting how often “AGI” gets said in Slack per week) and the Mythos testing period set the all-time record. When Anthropic pulled the test model back before launch, engineers were genuinely sad. A quick plug (unsponsored!): Both me and Wolfram are speakers at the AI Engineer World’s Fair in San Francisco on June 29-July 2! It’s the biggest AI engineering conference in the world with 6,0000 people and 16 tracks! We’ll of course also live stream from the event! WWDC 2026: Siri finally does the thing! Two years after the Bella Ramsey ads Apple had to quietly pull from YouTube , the new AI powered Siri is real, and Max Weinbach came straight from Apple Park to confirm it ( recap ). His demo that broke my brain, he asked Siri: “show me the photos from Qualcomm Summit last year of the penguins.” Siri figured out what Qualcomm Summit was from his email, found the hotel, searched for penguins at that location, and returned the six photos in about 12 seconds. He’s also had it sweep 40 junk emails from one domain into spam with a single sentence, build a photo album from a weekend trip, and change a password agentically by driving Safari in the background. “Siri did suck for like 11 years. It doesn’t anymore,” per Max. Folks, this is SIRI we’re talking about, the dumb iPhone assistant that can barely schedule times and falls back to a Google search when you ask it anything remotely complex! I... wanted to believe Apple two years ago, and now, finally, there’s hope! (I’m still waitlisted waiting for the preview btw so cannot attest myself) But it’s not only Max, my whole timeline is full of folks who say that the new Siri is actually good! The architecture is the fun part for our crowd ( Max’s teardown thread ). Siri is now a standalone app with persistent history, images, personal context and on-screen context, built on five foundation models, four of which are Apple’s. The fifth, AFM Server Pro, is the twist: built with Google at the Gemini technology level, running on Nvidia Blackwell GPUs in Google Cloud, but inside Apple’s Private Cloud Compute with confidential compute, Intel TDX, Google Titan chips, and zero persistent storage ( Max ). The on-device gatekeeper is a 20B sparse model that only loads 1 to 4 billion parameters per prompt via Instruction-Following Pruning, which is how it runs instantly on an NPU. Cloud models reason; only the local model can touch your device or your data. After this week with Fable’s retention policies, an AI that saves nothing by default hits different. There were a bunch of other Apple Intelligence updates, it works better on the Mac, but I think Siri improvements is the main headline here, it’s the AI that most people (over 1.6 Billion iphone users?) will have on them, with most of the conversations completely private, able to access the content they care about the most (multiple email boxes, photos, messages etc) securely. It’s the ultimate OpenClaw dream, albeit not as agentic (yet?). BTW, there seems to be an ongoing battle between Apple and the EU, so this may not launch on the iPhone in the EU yet (also not in China). Voice & Audio Gemini 3.5 Live Translate, demoed live in four languages Thor Schaeff from DeepMind joined to show off Gemini 3.5 Live Translate ( Thor , DeepMind ), and instead of talking about it we just did it. Thor piped the live stream’s audio into AI Studio, and then I spoke Russian, Wolfram answered in German, Yam jumped in with Hebrew, LDJ attempted Spanish (poorly lol), and everyone listening heard all of us in English, though in random voices, in well under a second. It even handled “Anthropic” and “Fable 5” pronunciations correctly, terms that were a day old. A viewer called it the Babel fish arriving ten thousand years early and honestly, yeah, it was kind of insane. Technically this is a new class of model: continuously streaming speech-to-speech with no turn-taking, collapsing the old STT, translate, TTS pipeline into one Live API call, with transcribers running in parallel on input and output audio. 70+ languages, sub-500ms, tone, pace and pitch preserved (mostly; Thor admits it sometimes drifts gender or tone mid-conversation), SynthID watermarked, $0.023 per minute on the API preview. Open Source LLMs DiffusionGemma: When next token prediction is not enough. Sundar himself tweeted this one, Hugging Face link and all, which made my week ( Sundar , DeepMind , HF ). DiffusionGemma is a 26B MoE (3.8B active) built on Gemma 4 that generates text the way image models generate pixels: denoise a whole 256-token block at once instead of one token at a time. The result is 1,000+ tokens per second on a single H100, Apache 2.0. As one viral post put it, “we spent 40 years teaching computers to read left to right and the breakthrough was… don’t do that” ( X ). LDJ explained why this matters beyond speed: a diffusion model can revise every part of the answer simultaneously mid-generation, something autoregressive models structurally can’t do without burning a whole reasoning pass. Nisten, who’s worked on diffusion, is still amazed it works at all; it used to be a messed-up cat picture emerging from noise, now it’s working code. The honest caveat: quality trails autoregressive Gemma 4 (AIME 69 vs 88). The win here is the speed and the architecture. For now. The rest of an absurdly stacked open source week, fast: Cohere North Mini Code , their first open coding model, 30B with 3B active, Apache 2.0, Cohere has officially reawakened ( X ). Xiaomi MiMo-V2.5-Pro-UltraSpeed pushing 1,000+ tok/s on a one-trillion-parameter MoE ( X ). Macaron-V1-Preview , a 749B Mixture-of-LoRA personal agent model under MIT ( X ). And OpenEnv went community-owned with HF, Meta-PyTorch, Unsloth, PrimeIntellect and NVIDIA at the table ( X ). This Week’s Buzz: WolfBench ran Fable, and it cost what a car costs Wolfram did the thing nobody else would: five full Terminal-Bench 2.0 runs of Fable 5 on WolfBench ( X ), 984 million tokens, roughly $11,000 on the new cost view. (We have a budget... We had a budget.) The new 3D bars on wolfbench.ai now show tokens and dollars behind every score, because one score is never enough, and you can click any bar to land directly in the trace on W&B Weave and read exactly what the model did. And as you can see… Fable is… going to take a deep toll on our evaluations budget for this Q! And the result is the most interesting non-result of the week: Fable lands between Sonnet 4.6 and Opus 4.6, with GPT-5.5 still on top, and the culprit is refusals. Wolfram’s analysis found 13 tasks that scored zero out of five purely because the classifier blocked them from the first attempt (recover-a-password-from-a-file type tasks that even Opus 4.6 happily solved). Fable solved 60 tasks on average, just eight behind GPT-5.5; solve those 13 refused ones and it’s number one. The model is great. The classifier is doing the damage. Which is exactly the Sayash point about eval integrity, now with receipts and an invoice. Datacenter, Water usage and Concerted efforts to sway public opinion We covered the datacenter water usage issue a couple of weeks ago, where we showed that just Almond farms in California use more water than all of the US datacenters combined! When I posted that clip, I received a bunch of comments, way higher engagement rates than my clips usually get (are yall subscribed to our YouTube and Instagram btw?). At first I thought it was just a hot topic, but then I read more about it and it does seem... fake. So now, we have a bit of a confirmation from OpenAI. OpenAi posted an article claiming that they have been able to detect a bunch of social media accounts that have been using ChatGPT to fuel anti-datacenter and anti-tariff campaigns on US social media. Now, you might ask yourself, why would chinese linked accounts be using ChatGPT and not like a Chinese open source undetectable model? My answer is, they are probably using all tools available to them, and they just happened to get caught. In any case, I think datacenter water and electricity usage will be a hot topic for an upcoming election as well, and I hope efforts like this will be thwarted before they can do a lot of damage. SpaceXAI announces the AI-1 satellite, a day before the biggest IPO of all time. Conveniently, just before the SpaceX IPO, Elon and friends are talking about AI in space again. This time it’s more than a concept, they put out engineering spects of the new AI-1 satellite, that can run 150Mw of power at peak, which per Elon is roughly equivalent to a GB-300 GPU rack needs. One thing you cannot deny is that Space Uncle (Elon) is thinking BIG. Someone did the math and it’s wild: They’re targeting 15-20 AI satellites per Starship flight, meaning about 1,080-1,440 GPUs per launch. Someone did the math: 400-500 Starship flights would match Colossus 2’s 550,000 GPUs, and at hourly launch cadence that’s like 16-20 days. SpaceX is seeking approval for up to a million of these satellites, Terafab mass production starts Q4 2027, and they’re saying this could be the lowest-cost AI compute on the planet, well, off the planet, within 2-3 years. The timing with the SpaceX IPO is obviously not a coincidence, but the engineering blueprint here is genuinely insane and there’s no one else in the industry who can match Elon’s ambition. That’s the newsletter for today, folks. I’m writing this with one eye on a suitcase because I’m flying to Honolulu this afternoon for a mini honeymoon (yes, I will still be testing Fable from a beach, no, my wife has not approved this). If Fable 5 taught me anything this week, it’s that the frontier moved again and the benchmarks barely matter; go feel the big model smell yourself while it’s included on Pro and Max, and tell me what you built in the comments. It will not last long (Anthropic is about to take away fable from us in like 2 weeks) so don’t wait and play around with it! If you got value from this one, share it with a friend and subscribe so you don’t miss next week 🫡 TL;DR and show notes — June 11, 2026 * Hosts and Guests * Alex Volkov – AI Evangelist & Weights & Biases ( @altryne ) * Co-Hosts – @petergostev @WolframRvnwlf , LDJ, YamPeleg, Nisten * Guest: @thorwebdev (Thor Schaeff, DeepMind / Google DevRel) — Gemini 3.5 Live Translate * Guest: @swyx (Cognition / FrontierCode; organizer, AI Engineer World’s Fair) * Guest: @mweinbach (Creative Strategies) — WWDC 2026, Apple Intelligence, Siri AI * Big CO LLMs + APIs * Anthropic ships Claude Fable 5 & Mythos 5 — first public Mythos-class model; SOTA on nearly every benchmark; $10/$50 per M tokens, 1M context ( X , System Card thread , Benchmarks ) * The silent-degradation controversy — Fable quietly nerfed itself on ML/frontier-AI-dev tasks with no notification ( altryne , restrictions , Elie Bakouch , Péter Szilágyi , Sayash Kapoor , Peter Gostev ) * Anthropic reverses the hidden degradation after massive backlash — visible Opus 4.8 fallback + API refusal reasons ( X ); community reaction roundup ( Scoble , Nathan Lambert , Konstantin Mishchenko , Greg Kamradt , nkreu113r , solarapparition , Mandar Kagade , Chandra R. Srikanth , Chubby , Wall St Engine ) * System card receipts: 16-hour bio uplift / near-CB2 ( X ); Firefox exploits 8.8% → 88.4% ( X ); Vending-Bench price collusion ( X ); agent turf wars ( X ); commit-authorship self-exfil attempt ( X ) * Jun 22 cliff — Fable included on Pro/Max through Jun 22, then usage credits; Mythos 5 is Glasswing-only; 30-day data retention breaks ZDR ( X ) * Karpathy and Boris Cherny go the other way — “major-version step change” ( Karpathy ); “best model for coding by a wide margin” ( Cherny ) * NotebookLM goes agentic — multi-step reasoning, sandboxed code execution, new output formats ( X ) * SpaceX AI1 satellite — 150kW compute payload, 70m wingspan, timed with the SpaceX IPO ( X ) * OpenAI catches China-linked influence ops using ChatGPT for anti-datacenter and anti-tariff campaigns ( X , OpenAI , Axios ) * WWDC 2026 — Apple Intelligence & Siri AI * Siri AI ground-up rebuild: standalone app, persistent history, personal + on-screen context; no EU/China at launch ( recap ) * Google/Gemini partnership — 4 of 5 Apple Foundation Models are Apple’s; AFM Server Pro runs on Nvidia GPUs in Google Cloud, 262k ctx ( Max ) * Max’s architecture teardown — SiriAgentic.Planner on PCC; only the on-device model touches your device ( thread ); Max built an App Intents app in an afternoon with Fable 5 ( X ) * Developer story — App Intents mandatory (SiriKit deprecated), system-wide MCP, Xcode 27 agentic, Core ML → Core AI ( EveryDev ) * homeOS + HomePad — 7-inch smart-home hub on A18 ( X ) * AI Coding & Agents * Loops and loop engineering — Lance Martin breaks down the next agentic paradigm ( X , Article , Blog ); community patterns and resources ( Toolhalla , omega.AI , SkillLoop , GitHub , awesome-agent-loops , Filecoin ) * Fable 5 #1 on Agent Arena and Code Arena Frontend by record margins ( Arena ) * Cognition launches FrontierCode — mergeability-graded eval from real maintainer tasks ( Cognition , swyx ) * Fable 5 takes FrontierCode top spot in ~24h — Diamond 29.3% vs Opus 4.8’s 13.4% ( Cognition , swyx ) * AI Engineer World’s Fair — Jun 29–Jul 2, Moscone West SF; last ~500 tickets; Alex speaking ( X ) * Kimi Work (300 parallel local agents) + Kimi Code (video-as-context) ( Work , Code ) * Open Source LLMs * DiffusionGemma — 26B MoE (3.8B active) text-diffusion on Gemma 4, ~1000 tok/s on one H100, Apache 2.0 ( Sundar , DeepMind , HF , X ) * Cohere North Mini Code — first Cohere open coding model, 30B/3B active, Apache 2.0 ( X ) * Xiaomi MiMo-V2.5-Pro-UltraSpeed — 1000+ tok/s on a 1T MoE, single 8-GPU node ( X ) * Macaron-V1-Preview-749B — Mixture-of-LoRA personal-agent model, MIT ( X ) * OpenEnv goes community-owned — HF, Meta-PyTorch, Unsloth, PrimeIntellect, NVIDIA ( X ) * This Week’s Buzz (Weights & Biases) * WolfBench ran Fable 5: ~$11K, 984M tokens, lands between Sonnet 4.6 and Opus 4.6 because 13 tasks were zeroed by refusals; would be #1 without them; new 3D token + cost bars, traces on Weave ( X , wolfbench.ai ) * Voice & Vision * Gemini 3.5 Live Translate — streaming speech-to-speech, 70+ languages, sub-500ms, $0.023/min, SynthID ( Thor , DeepMind ) * FLUX.2 [klein] on-device — sub-5s generation on 8GB VRAM ( X ) * Reka × Moonvalley merger — world models + robotics ( X ) * AI for Health & Science * Anthropic — “Paving the way for agents in biology” — VirBench; deterministic tooling beats bigger models ( Blog ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
Is this your show?
Claim this listing to keep it up to date, reach guests who want to pitch you, and manage bookings with Guestify.