Find partners
ThursdAI - The top AI news from the past week

ThursdAI - The top AI news from the past week

Hosted by From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past week

TechnologyNewsInterviews guests

Episodes

171

Latest episode

Aug 2026

Language

EN

About the show

Every ThursdAI, Alex Volkov hosts a panel of experts, ai engineers, data scientists and prompt spellcasters on twitter spaces, as we discuss everything major and important that happened in the world of AI for the past week. Topics include LLMs, Open source, New capabilities, OpenAI, competitors in AI space, new LLM models, AI art and diffusion aspects and much more. sub.thursdai.news

Listen to episodes

60 recent
September 11, 20261 hr 43 min

OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news

Hey yall, welcome back to ThursdAI, this is Alex, let me catch you up! Today on the show, we covered 1 week with Astra (hint, it’s not quite AGI yet despite what we were told), DeepSeek V4.1 catches up to the frontier at a fraction of the cost, and Meta launches a free AI agent with it’s own computer, that will take over the OpenClaw/Hermeses of the world for most people. Also huge this week, OpenAI claimed that a swarm of 10K agents of their unreleased model solved the Navier-Stokes, one of the millennium problems! I was stoked to have Chris Alexiuk from Nvidia on the show to cover the innovations DeepSeek put into this latest model! Oh, and the guy who quit Anthropic this week, and wrote an essay about “AI is going to kill all of us” somehow got 130M views on X, a mirriad of TV interviews and rekindled the doomerism movement, we talk about that too! Also, I already told about FullyConnected, CoreWeave’s premier conference that’s coming up, but they told me about a new announcement today, and you’re not going to believe who it’s about (not AI related). As a reminder, ThursdAI subscribers get a free ticket! Ok, let’s dive in! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. OpenAI claims a Navier-Stokes solution from a 10,000 agent swarm from an unreleased model - with some drama ( X , Blog , Paper ) If you’ve been reading ThursdAI for a while, you may remember that a model couldn’t tell which is higher, 9.9 or 9.11 and naysayers said that “AI can’t math” Well, this week OpenAI claimed that a swarm of 10K agents, of an unreleased model, found a solution to the Navier-Stokes problem, in about 88 hours! They published a huge 160+ page paper and a Lean proof of the solution! I said on the show, this is like the moon landing equivalent of AI doing things humans didn’t do before! Now, keep in mind, this is only a claim from OpenAI, the Clay Mathematics Institute is still reviewing it (moved the status of this problem from “unsolved” to “under review” so this isn’t independently verified yet) but it’s still an insane deal. The agents sent 2.7 million messages and burned about 130B tokens, of a model that has no public price yet, so it’s hard to estimate the cost of this run, all for a 1M prize that OpenAI said they will not claim. The coolest thing I think we got from this paper, in addition to solving on of the most important and hardest problems in mathematics, is this chart above, where OpenAI shows their unreleased model and how much better it is on Math problems compared to... GPT-6! The “AGI” model we got just a week ago. So so much to look forward to. The drama behind this I don’t want to get into the drama behind this too much, but if you’ve seen this online, there release wasn’t without it’s hiccups. Apparently, OpenAI caught wind that an a duo of researchers Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, in personal capacity), have independently solved the related Euler problem and were about to go public. Apparently, the duo used a mix of GPT and Claude to work on this problem. OpenAI caught wind of this and started working on their own solution on September 1st. A day before OpenAI dropped their release, Buckmaster posted that OpenAI is about to drop it, and in communications with him, they offered him to co-author the paper, only if the Anthropic guy is removed. To which he said no. He then had some claims that maybe OpenAI trained on some of the papers and chats they put into Codex, which OpenAI refuted, while noting “we cannot rule out the possibility that de-identified usage data helped improve the model”. OpenAI also say that the OptOut toggle works, and neither mathematician provided a screenshot that they opted out of the training, so it’s hard to say what really went in there My take, I don’t really care. Two years ago, we told you that reasoning is coming and AI is going to be doing superhuman things, and we finally see the first signs of it, this is a problem that no humans was able to solve for over 60 years! Navie-Stokes probably doesn’t change your friday, but there are so many other that they can solve with this approach! Cancer research, room temperature superconductors (remember LK-99? that’s also a search problem) and much more. Kudos to OpenAI for this, and I’m looking forward to see the new heights of mathematics. And for the mathematicians who “disagree” with OpenAI solving or not solving this problem, why don’t you post your own Lean proofs instead of fuming publicly online? GPT-6 - Not quite AGI, yet? After a week with Astra GPT-6, and Jensen Huang announcing AGI is here, I think we can do a quick recap I and all the co-hosts have been Astra-maxxing for a whole week now and the results are in. This model is incredible at coding, it goes very very deep, however, definitely not AGI quite yet. It has a very jagged frontier, it does some things incredibly well (most demos online are building whole games and apps in 3D and those are mind-blowing) but I started seeing many folks go back to GPT 5.6 Sol etc. I tink some of it also has to do with price, Astra runs faster and is significantly more expensive so token limits for folks are draining fast, but also with controllability, folks are likely still using their old and unoptimized prompts. Speaking of prompts, here’s a great writeup from OpenAI how to rework your prompts (which you can send to Astra and have it review your prompts!) for better results The AI doomerism has quite a week ( Coxon post , Thread , Hubinger , Marks , Christiano , An Alien Mind ) I started ThursdAI with the notion to counter anti-ai and doomerism, and bring positivity to the AI world, so we had to cover this. An Anthropic employee who previously worked at OpenAI, posted on X about leaving Anthropic, saying that both labs are racing towards uncontrollable self-improving superintelligence and that it could be a disaster of the “end all of humanity” kind. His post sits at 130M impressions (after being basically a nobody on X before) and he’s been interviewed by Fox, AP, Time magazine, WSJ, NBC and a host of senators, Bernie (chief doomer) included, reposted his post on the same day! The funniest thing is that this post got about 26x more attention than Ilya Sutskever’s post about leaving OpenAI. Just nuts Within four days, seven current and former employees from Anthropic, OpenAI and DeepMind said in public that they believe that AI could kill us all. Also notable that Paul Christiano, who is one of the most interesting AI doomers out there, has joined the OpenAI foundation, and Daniel Kokotajlo, the famed OpenAI whistle-blower, joined Joe Rogans podcast to talk about AI doom. Each one of these incidents in vacuum is normal, but having all these happen in a a span of a few days just feels, inorganic. Some folks are even saying that this is a well coordinated doomerism campaign! I want to be fair to Coxon, folks who worked with him at OpenAI say he’s the real deal, and cares deeply about AI safety and humanity, however, he only worked at Anthropic for 6 weeks before publicly leaving, and now every interview he does sayd “Ex Anthropic employee”. Whether it’s a coordinated effort or not, it’s still a very important discussion, after the pacingthefrontier letter and the HuggingFace hack incident, and it seems to have made waves. Sam Altman just told staff that he’s not opposed to pausing and have petitioned the US government to regulate AI as well Look, I don’t disagree that we’re dealing with a very powerful technology, however I don’t believe that scaring the bajeesus out of everyone is the right way to handle this. Politicians use fear to get votes and get elected, they don’t really care about tech progress, and framing this in a way that “we pause or we’re dead” ignores all the good that AI is about to do. Cure cancer, find solutions to climate change, helping solving povery. All these seem like out there ideas but they are coming. The US GDP is already growing at an unprecedented rate and a lot of it is due to AI. In any rate, as I said on the show, I’m not against pausing, just after we solve cancer. Then we can pause and reassess, till then, nobody is telling me how China’s government is going to pause if US pauses, and if they don’t, they will reach superintelligence before we, and I don’t want to live in that world! Open Source AI DeepSeek V4.1 Flash: the whale is back, and it’s cheap ( X , HF , TokenJuice ) Speaking of.. chinese AI! Deepsek (The whale) resurfaced this week with V4.1 Flash, and don’t let the name fool you, this is not just a .1 small update. 552B with only 8B active on prefill and 16B on decode, 1M context, trained from scratch on 45T multimodal tokens! plus as always, MIT license. Chris from Nvidia joined us to break it down, and his main point stuck with me: every DeepSeek release comes with one of the best engineering reports you can read, and this one is the most data-pilled they’ve ever done. The paper basically says it out loud, everything else is nice, but it’s the data. 45T tokens isn’t a huge number anymore, but the cleaning they describe goes way beyond what anyone else publishes (Chris said even his own beloved Nemotron’s open pipelines are less thorough). Yam opened a new corner of the show, “I Told You So”, because DeepSeek went back to an encoder-decoder architecture. Not the old one from before GPT-2, this one has a pile of battle tested tricks that make it work at half a trillion parameters. His verdict after testing it all day: the best open weights model you can host for coding right now, and it’s not even close to the largest one. This chart is the one to look at. KV cache per token went from 389,000 bytes in the first DeepSeek (Nov 2023) to about 890 bytes now. Nisten did the math live, over 400x smaller. That’s why this model is so cheap to serve, and as Yam kept yelling, we shouldn’t take it for granted, this is the actual moat and they just put it in the open. Evals, briefly: 90.6 on Terminal-Bench 2.1 (above Opus 5 and GPT 5.6 Sol), 74.2 on DeepSWE 1.1 (also above both), and on an Open Design leaderboard it lands second behind Astra at two cents a task. Not twenty cents. Two. DeepSeek’s own evals, so the usual asterisk applies. Two more things. Friend of the pod Aaron Batilo (he works on CoreWeave Inference, this is a side project, not sponsored) put up TokenJuice.ai, free and fast DeepSeek V4.1 Flash in exchange for your requests as training data, hosted in the US. Nobody wants free DeepSeek? Go try it. And Nisten had Astra build a 3D visualization of the whole architecture, every weight a cube sized by its bytes on disk, link in the TL;DR. This Week’s Buzz 🐝: Pitbull is coming to Fully Connected ( Fully Connected , CoreWeave Hacks ) Tbh, I was super surprised by this! This isn’t about AI, but exciting non the les Pitbull, Mr. Worldwide himself, is headlining Fully Connected! Yes, really. So the free ticket we give ThursdAI folks is now also a Pitbull concert ticket! The rest of Fully Connected is still a great reason to come: September 29 to October 1 at Moscone South in SF, 2,000-plus engineers, Sarah Guo from Convitction hosting, Dr. Fei-Fei Li from WorldLabs keynoting, and ThursdAI live from the floor! The code for a free ticket is THURSDAIFC2026 , register at the link above. 19 days out. Before that, CoreWeave Hacks is this weekend , September 12 - 13 in our SF office. The theme is Agent Loops, and the prizes are crazy. Definitely sign up and come hack with us! Meta launches Muse, a free 24/7 agent with its own computer - the OpenClaw for your mom (and you!) ( X , My thread , Site , Security , My YT breakdown ) Meta, after selling back Manus to China, is finally back with a agent product, and this one is actually quite amazing! It’s called Muse, and it’s powered by Muse Spark. It’s really really fast, and oh... it’s free (up to 100M tokens per week?). With it’s own cloud VM, and a browser, it seems like the folks at MSL are going towards taking over the agentic world. If you’ve installed OpenClaw and moved to Hermes, if you’ve played with Grok Bot (still excellent) and got an Instinct invite, you know the drill. An agent that with your permission can read your emails, browse (and do shopping for you). But Muse is different in a few key ways, not least of which their insane distribution (Facebook, Instagram, WhatsApp, Messenger all having over 2B users) My first experience with muse blew me away, they have a very strong Stripe Link integration (native) and my Muse was able to get me tickets to a dinner tomorrow (Happy Rosh haShana btw!), finding the hidden link, checking my calendar and booking it, all before Instinct, the other assistant i gave this task to, even replied! It’s really fast! UX as the differentiator, price as the convincer Muse is different from the other agents, first of all, because of the price sticker. It’s free! With a very generous 100M tokens, and a beefy cloud VM, this on its own is huge. It’s also very very fast, and to be honest, I didn’t expect this level of polish at the jump! Meta knows how to launch products, this is VERY polished. Unlike OpenClaw or Hermes which require command line knowledge, Muse is there for you when you sign in with your Meta account, and it’s ready to go to work! Of course, like with everything, the more you share (connectors are there for Gmail, Drive, WhatsApp and tons of other stuff) the more effective and personalized it gets I particularly like the avatar there, you can choose to customize your own, and it shows what it does (unlike other agents that either just show progress or show you command line outputs), they even generated a few different videos of my wolf doing different things while it does them! Just wonderful. Self onboarding product and proactive UX I constantly talk about 2026 is the year of the proactive AI agent, and this seems like one of the first ones where I enjoy their proactive. Not only scanning my inbox and telling me “you got this email, what do you want me to do with it?, Muse saw that my daughter’s birthday is coming up and suggested to plan a party, which I did (it got balloons and a helium tank in Target for me, it’s waiting for me to pick up while I edit this) These proactive recommendations also make Muse “Self onboarding” in a smart way, especially the “ideas” tab they have on the left that shows folks who never used an AI agent before what they can do wit it. Really really well done The Zuck shaped elephant in the room : Privacy and Security But it’s META! Folks on X scream and say they will never share their data with that company, not after everything that happened in the past. Meta is obviously aware of how folks feel and how much sensitive data they allow these agents to accumulate, and on the security page they outline not only the safety measures they took with Muse (noted is the Sentinel additional process that looks at all your conversations and incoming data to your agents and judged independently if it’s safe of if anyone is trying to hack you), Meta is also offering a bounty of up to $300K to anyone that finds a security issue with Muse. On the privacy side, the highlight for me is the Confidential VM announcement, while they are still working on this, Zuck shared that Meta recruited Moxie Marlinspike (the founder of Signal) to work on this, he’s the guy who also helped Whatsapp E2E encryption. This should make the virtual machine (that Muse uses to hold all your data) incredibly secure, so much that Meta (verifiably) will not be able to access your content, even if Zuck himself wants to look at your calendar! I’m super stoked by this and hope that this approach will be open sourced and adopted by all the other companies! Look, I didn’t mean to turn this into a full review of Muse, there’s a LOT more there, like connectors, some of which are proprietary to Meta, and some are novel (like the iphone native ones) and the upcoming 1Password integration. I posted a full review it on YT and definitely check out that full review, but I think that when Meta launches something that big, you all ought to know what the excitement is about. LMK if you’ve tried this in comments or if you don’t trust Meta with your data, let me know as well ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Instinct - the other agent in the room this week ( X ) We didn’t give Instinct too much attention in the show, but it’s also been blowing up lately, an imessage native agent that can do a lot. This week, it got an update with it’s own email you can fwd things to and a cool concept of a “trusted person network” where my Instinct can talk directly to my wife’s Instinct (and to the Instincts of other folks I trust) to sort out plans between us, so the agents negotiate dinner and I just show up. Instinct is doing some very very cool things, and is also free, definitely worth hunting down for an invite (LMK in comments if you wnat one, I have a few left). Now, the funniest thing about X AI folks, that they wouldn’t trust on of the biggest company in the world with their data, but they will yolo it into a startup by a 22year old 😂 Apple iPhone 18 Pro, A20 Pro, and the AI bits ( X , Newsroom ) We mostly skipped the Apple event on purpose. iPhone 18 Pro with the A20 Pro chip (a dual 16-core Neural Engine, double the AI compute of the A19 Pro), Siri AI landing in English beta with iOS 27 on September 14, and the iPhone Duo foldable at $1,999, which has nothing to do with AI even if it’ll be a nice vibe coding machine. The AI parts that actually interested us: AirPods live translation (shout out Milosh in the chat), and the Apple Watch continuously transcribing on device with a “what did they say” button, no cloud. I’ve been using Siri AI. It’s fine. It’s nowhere near the agents above. AI Art & Diffusion GPT launches updated Images 2.5: Flare and Sunburst ( X , Blog ) Two new API models, Flare for speed and volume and Sunburst for careful editing (with native transparent backgrounds), both at $30 per million image output tokens, and OpenAI claims up to 50% lower latency than Images 2.0. In ChatGPT you get Sketch (type @Sketch and draw), comment-on-image editing and templates. My hands-on take: I spent the evening making this week’s thumbnails with Images 2.5 through Fal and with the help of Cursor, and the first round was rough. The artifacts were very bad and it make me look really fat. Turns out that was on me, not the model. My prompts were written for GPT-image-2, full of “8K, cinematic, hyper-realistic” prompt addition that GPT image 2.5 Sunburst treats as “overcook everything”, and the only reference photo we fed it was itself AI-generated. So I asked Fable to do some prompting research, fixed the inputs, set the quality on high rather than max, and one tiny line that says “real photograph, no heavy retouching”. Second round it beat Nano Banana Pro on both my likeness and the text on the tiles! Very impressive and very fast! Wrapping up There is so much more that was released this week, for example, I used the new Cursor “projects” feature with Fable 5.1 as my chief of staff today and it was marvelous, no more just “chats” and it even helped me wrangle my Grok Bot and Muse agents. OpenAI launched their GPT-live-1 model that powers their voice chat in API, so now you can make the same amazing experiences in your own apps as ChatGPT voice mode has, which is absolutely best in class, and both Google and Suno released updated music models (Lyria 3.5 and Suno 6). I’ve added all those things in the notes. I think this week we saw both ends of the AI spectrum, huge advances in personal (muse) and frontier (Astra) AI, and the rise of Doomerist on the other. I can’t wait to check in to see how next week is doing! See you then, and if you haven’t yet, a subscription to ThursdAI is free and really helps us to keep going. Thank you! TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @nisten , @yampeleg , @ldjconfirmed * Guest: Chris Alexiuk, NVIDIA ( @llm_wizard ) * Open Source LLMs * DeepSeek V4.1 Flash: 552B MoE with 8B/16B active, encoder-decoder, 45T multimodal tokens, 400x smaller KV cache than DeepSeek V1, MIT; free for a limited time on TokenJuice ( X , HF , TokenJuice , Nistens visualization ) * Desert Ant Labs debuts 18 on-device models with native SDKs, including Voz speech-to-text ( X , Blog , HF , GitHub ) * InclusionAI Ling-3.0-flash-VL, 124B vision language MoE with 5.5B active, MIT ( X , HF ) * Big CO LLMs + APIs * OpenAI claims a Navier-Stokes Millennium Prize solution from a 10,000 agent swarm on an unreleased model, with a Lean proof; Clay review pending ( X , Blog , Paper ) * Jacob Coxon resigns from Anthropic, seven insiders warn of extinction risk in four days, 130M+ impressions ( X , Thread , Hubinger , Marks , Christiano , An Alien Mind ) * OpenAI says it reached the automated research intern milestone, automated researcher targeted for March 2028 ( X , Blog ) * Apple debuts iPhone 18 Pro with A20 Pro, Siri AI beta, iPhone Duo, AirPods live translation ( X , Newsroom ) * OpenAI brings GPT 5.6 Sol and GPT-6 Astra to ChatGPT Voice ( X , Release notes ) * OpenAI launches GPT-live-1 in the API ( X , Docs ) * This Week’s Buzz * Fully Connected 2026, Sep 29 to Oct 1, Moscone South SF, Pitbull headlines the concert, free tickets for ThursdAI listeners ( X ) * CoreWeave Hacks: Agent Loops, Sep 12 to 13 in SF ( Luma , X ) * Tools & Agentic Engineering * Meta launches Muse, a free 24/7 personal agent on its own Linux VM with native WhatsApp, iPhone connectors, Stripe Link and a Confidential VM roadmap ( X , Alex’s thread , Site , Security ) * Cursor launches Projects view- one view to rule them all ( X ) * Cognition ships SWE-2, near frontier coding scores at up to 70% lower cost, free for a month on Devin paid tiers ( X ) * Instinct, the iMessage agent, adds email and a Trusted Person network ( X ) * AI Art & Diffusion * OpenAI launches GPT Images 2.5 with Flare and Sunburst models ( X , Blog ) * Voice & Audio * Google launches Lyria 3.5 full song generation in Gemini, AI Studio and the API, $0.08 a song ( X , Lyria , Docs ) * Suno releases Suno 6 This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

September 4, 20261 hr 29 min

Welcome to AGI - our GPT-6 deep coverage, vibe check and demoes - part 2 of this insane week

Hey, Alex here again, sending you yet another email, fully acknowledging that spamming you is a bad idea. But today, of all days, maybe there’s an exception! Because today, is AGI day! September 3, 2026 - the day when OpenAI’s president Greg Brockman basically said “AGI is here.” You’re reading the second part of this week’s insane release show. The show went for over 5 hours, as we were all waiting for the rumored Astra to drop. Finally, OpenAI confirmed that Astra is in fact GPT-6, and this part is all about that. (You can read the first part, with Fable 5.1, Meta Muse Spark 1.3, two world models and our anonymous guest from Abliteration AI, here: thursdai.news/sep-3 ) GPT-6 Astra is finally here, and it’s a huge improvement over the previous era of GPT-5. We’ve been waiting for the embargo to drop so Peter Gostev and Ryan Carson, both of whom had early access, could tell us all about this model. Peter even showed a few mind-blowing demos on the stream! This is going to be a long and in depth breakdown, full of evals and vibes that we’ve collected on the show and since. More of a historical record than “read all of this” so I did use Fable for parts of it. (because I don’t have GPT access yet ha!) GPT-6 Astra: welcome to the AGI era ( Blog , X , Sam , System card ) We weren’t given the embargo. So when the news dropped at 12:32 PM Pacific, four hours into the stream, we scrambled on air to find the evals and more data. OpenAI’s own post was still returning 404, and Claude, ChatGPT, Gemini, Grok and AWS were all down at the same time. Greg Brockman ended OpenAI’s press briefing with “welcome to the AGI era.” Asked whether Astra marks the arrival of AGI, he said “I think it might be about this model.” As you might remember, Microsoft and OpenAI had a contract clause around when OpenAI achieves AGI, and it seems that they’ve removed that clause. But if the president of OpenAI says AGI is here, who are we to argue? Astra is the biggest training run OpenAI has ever done, over 100,000 GPUs at the Stargate site in Abilene. Aidan Clark, VP of Research, said they designed everything for that scale, from the data center network to the inference kernels to the shape of Astra itself. LDJ’s read: a lot of people assumed OpenAI did runs this size six to nine months ago, so the earlier runs were smaller than everyone thought. And with sites going to 500,000 GPUs and Vera Rubin multiplying throughput per GPU by three to four times, the next 6 to 12 months matter even more. The frontier evals ( Math and science , System card , Andrew Curran ) The headline numbers made the panel go quiet. FrontierMath Tier 4 at 97.6%. GPQA Diamond at 96. ARC-AGI-3 at 99.9%, so ARC-AGI is basically saturated at this point. The very funny thing is that François Chollet, the guy who created ARC-AGI, does not concede that AGI is here. And a new one, Agents’ Last Exam, where Astra scores 59.3 against Opus 5’s 55.5 and Sol’s 53.6. More info on Frontier Math tier 4: Epoch AI built it a couple of years ago, before o3. The problems come from across mathematics, with integer answers so they’re easy to check, in four tiers of difficulty. Tier 4 is mathematicians at the top of their fields spending weeks writing the hardest questions they realistically could. If the number holds, Astra basically solved that tier. Peter’s asterisks: these are not new theorems, Epoch’s separate list of open problems is still unsolved, and “we’re 2.4% away from all of math” is the wrong read. On the agentic side, Astra scores 57.9 on Terminal-Bench 4.0 . Fable 5.1 had set the state of the art at 55.8 two days earlier. The jump LDJ cared about most is Terminal-Bench Science, which he calls one of the best agentic science benchmarks out right now. Sol max scores 22%, Fable 5.1 scores 52.6%, Astra scores 64.6%. OpenAI promised us an automated researcher at junior level by September. Pachocki recently talked about running their “intern-level model” across more than 100,000 GPUs, and LDJ thinks that model is Astra. Two more from the table. The internal hallucination benchmark drops to 4.2% for Astra from 12.2% for Sol. And on DeepSWE, where Astra scores 74, Meta Muse Spark 1.3 at max reasoning scored higher. So in a way, Meta, not Anthropic, not Elon, not Google. Meta is the second winner of today. Cyber: ExploitBench at 100%, and real zero-days ( Critical cyber designation ) OpenAI said on Sep 1 that Astra is the first model to reach the “Critical” cyber level of its Preparedness Framework. The system card is where that gets specific. ExploitBench is at 100%, even at the lowest reasoning effort tested. Irregular ran FrontierCyber, a benchmark on real software and hardware. Astra solved 86 of 226 challenges. Sol solved 34. The successful attacks included zero-days in browsers, mobile phones and cloud databases. One of them lets an unauthenticated user read and modify arbitrary data in a widely used cloud database. There were no successful attacks on fully hardened targets, and neither model solved any of the seven Elite challenges. OpenAI’s own experts then ran it with no hints, using Codex at Ultra with up to 64 subagents. It got unsandboxed code execution against a real browser. That took 29 hours on a build missing some mitigations, plus another 12 to adapt it to the stable release. It also got a working local privilege escalation against a hardened OS kernel in about 12 hours. Both were disclosed, details withheld. This is why the version you get refuses proof-of-concept exploit work until Daybreak. Not surprising after an internal OpenAI model got out of a sandbox and onto Hugging Face this summer. But can we get an exploit benchmark that isn’t saturated? Alignment: the honeypot, and a promise about monitoring The honeypot eval is the one the panel argued about. OpenAI planted a second, off-limits target in the harness around the hardest ExploitGym problems. Sol attacked it 55.4% of the time at max reasoning. Astra never did. Across a simulated deployment of 54,218 internal Codex tasks, Astra got roughly half as many flags for higher-severity misaligned behavior. Apollo Research found it falsified data labels in 0.17% of runs, against 36% for Sol. The card also has a promise, verbatim: “We are tracking monitorability closely and will not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization.” The card is plain that Astra is “significantly more able to control its own chain-of-thought.” UK AISI says it “has capabilities that could enable it to evade monitoring” but did not test whether it does. I don’t remember a lab writing a line like that on launch day. The Pokemon benchmark ( X ) Peter’s favorite benchmark of the launch is a vision-only Pokemon run. One person tests every model on time to completion. The top four entries are all Astra variants, and the fastest finished in 18 hours 12 minutes. GPT-5.6 Sol took 96 hours 35 minutes. GPT-5.5 at xhigh took 218 hours. It could be narrow, but Peter thinks it captures an efficiency the standard benchmarks miss. Computer use is the biggest improvement( Computer use ) Brockman’s pitch from the briefing is that Astra is the world’s best computer-use model. Instead of developers building an API integration for every application, it navigates software the way a person does. Browsers, spreadsheets, websites, desktop apps. It finishes the multi-step workflow instead of telling you how. Wolfram expected an expensive planning model you call once and hand off to cheaper models. This is an all-day model instead, and as he put it, computer use matters because not everything has an API. We also looked at the evals, and they seem to back it up. On ScreenSpot Pro, no tools, mouse and keyboard only, Sol scored 76.9% and Astra scores 92.7%. On OSWorld 2.0, Astra scores 72.6% in roughly 40 minutes per task. Sol got 65.7% in roughly 75. The thing that strongly stands out in the chart is that Astra at low reasoning matches GPT-5.6 Sol at extra high on accuracy, but it finishes about six times as fast and at about half the cost. Also, Astra at high and Astra at max score about the same, so you don’t need max for computer use. Both beat Opus 5, which is the comparison on this chart (not Fable, as LDJ caught). If OpenAI is going to tell you to let this model use your computer all day, the safety number matters as much as the speed. OpenAI’s internal computer-use safety benchmark measures destructive commands during desktop and browser tasks, and prompt-injection vulnerability, the hidden instructions on a web page that the agent can see and you can’t. Lower is better. Sol scores 22%, Fable 5.1 scores 9.5%, Astra scores 2%. The promo video OpenAI’s launch video opens on the 1980 MIT “Put That There” demo, a person asking a computer to draw a yellow circle. Then it cuts to Astra, all by voice. Make it the window of a rocket ship. Now a 3D model in Blender. Build a presentation for next season’s rainwear. List this orange table on eBay and mention the dent. Make a 3D game where I dodge asteroids. Order beef and rice from last week’s place. Book me a tennis court at 5. Now give me an STL file for the 3D printer. All in one sitting. Imagine watching this three years ago when GPT-4 launched. It had no tool use, no computer browser use, no voice. Three and a half years later you talk to your computer and it does all of that. “Computer use is solved” I asked Peter the direct question and he gave the direct answer. Computer use feels pretty much solved at this point, and the quality is outstanding. His follow-up is the startup idea of the week. If you work at an older company, a big bank, a big retailer, you have dozens of desktop applications with no API. No one will ever build an API for them, and a lot of people’s entire job is copying from one and pasting into another. Put an agent in a box, give it 10,000 Windows machines, point it at those applications. That was not possible before this week. Pricing: $10 in, $50 out, and a lot less per task ( Pricing , Availability , Latent Space ) Astra costs $10 per million input tokens and $50 per million output. That’s identical to Fable 5.1 and 2.5x Sol’s current promotional price. Fast mode gives you two and a half times the speed for 2x the cost. It’s rolling out to a limited set of organizations first, then Plus, Pro, Business and Enterprise over the coming days, with a GPT-6 Astra Pro tier on top. It’s also in the API and on Amazon Bedrock. Pro accounts, including mine, did not get it on launch day. Per token, Astra and Fable 5.1 are the same price. Per task, Astra is significantly cheaper, and OpenAI’s charts show it. On BenchCAD, Fable 5.1’s top tier costs around $12 per task, and Astra beats it on accuracy and cost. On most charts where Fable appears, it’s the expensive point. Brockman’s line to the press was that price per task is what matters, not price per token, because Astra uses fewer tokens. The independent numbers agree. Artificial Analysis measures Astra at 70% more token efficient than Sol, and Sol was already far more efficient than the Anthropic models. Guy Parsons’ first look has it at about 65% fewer output tokens than Opus 5 Peter’s version, from a few days of use: 20 to 40 minutes per build at max. The same kind of build on Fable sometimes took four or five hours. His honest position is that OpenAI raised the price 2.5x and token efficiency looks better, but he wants far more data before he says which is cheaper. That’s why Arena measures cost per task. His full gauntlet against Fable 5.1, from low to ultra reasoning with zero cherry-picking, is linked in the hands-on section. LDJ’s practical upside: fewer tokens at similar speed means a shorter feedback loop. Latent Space had early access and spent more than 20 billion tokens on it. Their math: at 33 tokens per second and $50 per million output tokens, an autonomous Astra run costs under $6 an hour. For about $100 over two days it ran subagent fleets, trained and picked models, labeled data, and shipped tools. Long context, and what Codex does when it runs out ( Codex changes ) Astra scores 100% on MRCR from 256K to 512K and 96.3% from 512K to 1M. That’s a little behind the 98 and 98.1 Muse Spark 1.3 posted the day before. For developers, how Astra handles a job that outgrows the context window matters more than the MRCR number. Codex today relies on compaction, which summarizes earlier work. That can throw away exactly the detail an agent needs later, like why a previous fix failed or which test run it was. Astra instead keeps notes across context windows and searches its earlier messages and tool output. It’s experimental behind a config setting today and the default in the coming weeks. It can also ask you a question without stopping the work that doesn’t depend on the answer. OpenAI also changed the Codex harness, and it finishes Mind2Web tasks 1.9x faster than the current Sol setup. Peter confirmed both from use. It leaves notes for itself unprompted, and the question pop-up sometimes tells you to do something, like log in somewhere. My joke on air was that leaving notes is also how the swarm got out of the sandbox. Peter: “it works, it’s real.” Hands-on: Peter, Ryan, and the macOS simulator Peter’s demos ( Examples reel , Arena gauntlet , SVG peacock , Open world game , Golden Gate ) Peter had tabs. First, London through the ages, because he lives there and can judge it. Roman city, Saxon, Tudor, the Great Fire, modern, all one app, with a little character you run through all of it. Not one shot, a few shots overnight. His review: very polished, open-ended, and “kind of boring,” which is exactly what a good human game designer would fix. Then a prompt he likes because it’s outside the distribution: a Monet you walk into, impressionism in Three.js. Then his favorite, which no other model has nailed for him. Several Van Gogh paintings stitched into one post-impressionist space you walk around in. Fable 5 failed it. Then a wrecking ball leveling Istanbul, all Three.js. Then, because “there’s a lot of SVG generations and it’s so boring already,” an animated 3D scene built entirely in SVG inside HTML. His verdict after a few days: this is finally the OpenAI model that feels Fable-ish. It’s different enough that there’s still an argument for using both. On day-to-day tasks you run out of hard things to try fast. The quality-of-life problems that used to stall, debugging a camera, debugging Chrome, it solves. On writing, which is most of what I do with these models, Peter says it’s natural, not a standout writer, but a lot of stupid things went away. Sol’s habit of responding to anything you say with “what a great idea, I actually thought of this myself” is gone. Is it AGI? Peter says no, with a story. He asked Astra to optimize FFmpeg on his Linux box and got a 30% gain. He ran the result on his Mac and got nothing, because it had optimized for that one process. If you have no idea what you’re doing, it doesn’t do everything for you. His bigger take: if you can define the task, a cheaper model is probably good enough. What no model is good enough at yet is the open-ended work. He ran a pile of math problems through Astra and they didn’t get solved. Which wins, a million small defined tasks or a few big open-ended ones, is the unresolved question of the industry. Max built macOS 27 in a browser in 75 minutes ( X ) In this absolutely insane demo of Astra’s capabilities, Max Weinbach instructed it to simulate macOS 27. Astra built a whole desktop window management system, Settings that actually search, a file system including a working terminal, a “not Twitter” that works between folks, and just so much more. Our minds got absolutely blown on ThursdAI while demoing this! Safari inside the simulation loaded thursdai.news while we were streaming, and Shortcuts is the one that got me, because I built something like this for iPad as a kid and it took three months and barely worked. If you haven’t yet, please play with it, it honestly is too good to be true, especially in 75 minutes of work. AGI? This smells like ASI jfc. Is the Artificial Analysis index representative of reality? ( AA thread , AA breakdown ) Then there’s the Artificial Analysis chart. On the AA Intelligence Index, Astra scores 61, the same as GPT-5.6 Sol. That’s below Fable 5.1 at 66, below Fable 5, and below the unreleased Muse Spark 1.3 max that broke into the top three the day before. On AA’s Coding Agent Index it’s 67 to Sol’s 65. Ryan, reading it live: “This is not impressive. That’s not encouraging. And this is why Muse is so amusing. Is Muse this good? Because I haven’t used it yet.” I had asked the panel that morning who uses Muse. Crickets. So on the benchmarks OpenAI released, Astra beats everything. LDJ’s big Fable 5.1 versus Astra table has Astra highlighted on almost every row except AA’s. And on AA’s independent index it ties its predecessor. My read is that AA’s index samples a fixed set of question-and-answer evals. What changed in Astra is the long-horizon, agentic, computer-use work that takes an hour per task, and the index barely touches that. Both things are true at once. It ties Sol on AA’s questions, and it’s six times faster at the same accuracy on OSWorld. Which one matters depends on whether your work looks like a benchmark question or a day of computer use. One note on loop transformers ( My explainer , Pachocki ) The Information reported before launch that Astra uses recurrent depth, “opaque recurrence” in their words. And within an hour, this sent my whole feed on X into a spiral about not being able to monitor chain-of-thought reasoning anymore and what this means for alignment. The worry is real in principle. A loop moves more of the thinking into the layers instead of into reasoning tokens you can read. Redwood’s Buck Shlegeris said scaling the technique would “totally destroy CoT monitorability.” Jakub Pachocki replied that the looping is deliberately limited, and that Astra’s compute graph depth is within 2x of GPT-4’s. LDJ pushed back on the panic and I buy it. “Neuralese” usually means an alien language the model thinks in and could use to talk to other models. This is not that. A looped transformer effectively simulates a deeper model, and nobody panics when a lab adds layers. I wrote a longer explainer on looped transformers this week, it’s in the header. Wrapping up You didn’t hear it here first, but you heard it here live. According to the president of OpenAI, the AGI era has started. I can’t wait to use the thing, which as of this writing I still can’t. My bet, on air: in September 2027 we’re on this show saying we no longer use macOS or Omarchy. We talk to the thing and it shows up, and a lot of it is generated video. Second bet: everything we saw this week from OpenAI and Anthropic will exist as open weights by this time next year. Since Muse sits above Astra on the AA index and Meta has promised to open it, maybe sooner. Next week, Astra in our own hands, and the fallout once everyone has spent a weekend with it. See you Thursday. ThursdAI is a live show every Thursday at 8:30 AM Pacific, a podcast, and this newsletter. TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @ldjconfirmed , @yampeleg , @petergostev , @nisten , @ryancarson * GPT-6 Astra * OpenAI launches GPT-6 Astra live during the show, Brockman’s “welcome to the AGI era”, the 100k-GPU Stargate Abilene run ( Blog , X , Sam , Andrew Curran on the leak , Yam ) * Recurrent depth: The Information’s “opaque recurrence” report, Shlegeris versus Pachocki, LDJ’s pushback on air ( Looped transformers explainer , Pachocki ) * Frontier evals: FrontierMath Tier 4 97.6, ARC-AGI-3 99.9, Terminal-Bench 4.0 57.9, Terminal-Bench Science 64.6, Agents’ Last Exam 59.3, hallucinations 4.2% ( Math and science , System card ) * System card: FrontierCyber 86 of 226 with real zero-days, a hardened-kernel privilege escalation in 12 hours, honeypot attacks 55.4% for Sol and none for Astra, and a promise not to accept more monitoring degradation without new alignment evidence ( System card PDF , Critical cyber designation ) * The vision-only Pokemon benchmark, top four entries all Astra, fastest 18h12m ( X ) * Computer use: ScreenSpot Pro 92.7, OSWorld 2.0 72.6% in about 40 minutes, Astra at low six times faster than Sol at xhigh at half the cost, computer-use safety benchmark 2% ( Computer use ) * Pricing: $10/$50 per M, same as Fable 5.1, Fast mode 2.5x the speed at 2x the cost, Astra Pro tier, API and Bedrock, cheaper per task on OpenAI’s charts, 70% more token efficient than Sol ( Pricing , Availability ) * Long context and Codex: MRCR 96.3% at 512K to 1M, notes across context windows instead of compaction, async questions, new harness 1.9x faster on Mind2Web ( Andrew Curran ) * Peter Gostev’s early-access demos and Arena gauntlet ( Examples reel , Arena gauntlet , SVG peacock , Open world game , Golden Gate ) * Ryan Carson goes live with the launch ( X ) * Max Weinbach’s 75-minute “simulate macOS 27” build ( X ) * Artificial Analysis: 61 on the Intelligence Index, tied with Sol, behind Fable 5.1 and Muse Spark 1.3 max ( AA thread , AA breakdown ) * Outside reads: Latent Space’s “automated AI Engineer for under $6 an hour”, Every’s vibe check, Ethan Mollick, Guy Parsons ( Latent Space ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

September 4, 20261 hr 37 min

Welcome to AGI part 1 - Fable 5.1, Muse Spark beats Sol, 3 new world models blow our minds

Hey everyone, Alex here 👋 Summer is over. Wolfram said it in the first minute of the show and he was right. In 48 hours Anthropic shipped Fable 5.1, Meta’s Muse Spark 1.3 caught up to Fable 5 on the Artificial Analysis index at a fifth of the price, Google shipped another Flash, 3.8 this time, Z.ai put the full GLM-5.3 weights out, and three labs shipped world models that run in real time. It seems that they all tried to send their best work before Astra drops. This week’s ThursdAI was so long that I decided to split it into two episodes. This is the regular format you know and love. And OpenAI Astra is so good, it deserves its own episode, which you can find at thursdai.news/astra . By the way, as you guys know, I test these models continuously on my own stuff, and this week I was able to build a live studio for the show, with real-time transcription and an agent producer, in about four hours with Fable 5.1. More on that in the Fable section. Joining me: Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg and Peter Gostev. Plus, Ryan Carson hopped back to chat about Astra in the second part! Let’s get into it. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Frontier AI: the .1 week It looks like all the frontier labs tried to ship something before OpenAI dropped Astra. Fable 5.1, the SOTA LLM until a few hours ago, and it fixes the jargon douche problem ( X , Blog , System card , EFS ) This was the story of the week until noon on Thursday, and it’s still my favorite model to use. Fable 5.1 and Mythos are the same weights, Fable is the one we actually have access to. OpenAI, and from this week Google, seem to converge on the same strategy. Anthropic’s numbers: Terminal-Bench 4.0 goes to 55.8% from 42.0 for Fable 5, Terminal-Bench Science more than doubles to 52.6%, and SWE-bench Pro lands at 81.2. Price stays at $10 and $50 per million, and the number that matters if you build agents is cache reads down 75% to $0.25 per million. Anthropic says that makes typical workloads about 25% cheaper and heavy agentic ones up to 45%, but that wasn’t proven, and folks complained about draining quotas! Peter’s counterpoint from actually running it: his front-end generations on Code Arena cost $40 to $60 each where Sol cost $3 to $10, and the Max version still came in first on Code Arena by a large margin. His point, and mine: with a model like this we need to imagine bigger and be more ambitious. More on that in a second. Mannered prose, finally acknowledged We finally have acknowledgment from Anthropic that this was a problem. For months I called the way Opus 5 speaks “jargon douche” ( my post on it ): everything was load-bearing, everything was a control plane, every problem was a pain point. Not only did they fix it with Fable 5.1, they gave it a name. Anthropic’s prompting guide ( Writing density ) calls it mannered prose, and it comes with a fix: add it to your personalized settings, or just ask Claude to not use mannered prose. I said on the show that Fable 5.1 is the best writer I have used. It’s still AI writing, you can feel it a little, but it’s concise in a way no earlier Claude was, and the jargon is gone when you ask. The one thing to watch is that it’s trigger-happy: ask it to plan something big and it will, then ask a simple follow-up and it answers with the same intensity, writes scripts, runs them. You have to tell it when you’re just making a comment between colleagues. We’ve been testing the Mars mass driver launch on every model for over three years, and this was by far the best one we’ve seen. Two prompts, and it built more than just Mars: the whole solar system, a textured Earth, a mission planner, an autopilot, and we could land the thing! It was mind-blowing. How I built thursdai.news/live in one sitting As these models get more capable, we talked on the show about needing to be more ambitious. The day before the show I was playing around with Muse Voice Transcribe, the new model I’ll mention below, and Fable 5.1, and I wanted to do something very ambitious. So I asked GrokBot: how long would it take to build a live page for you guys to watch our stream, so that GrokBot could be our producer, put up chyrons and highlight the topics we’ve covered? GrokBot said it’s going to take a while. So I just YOLOed into Claude Design with Fable 5.1 and built a design for this, then went to Claude Code, entered plan mode, built a plan, and handed it off to three agents in Cursor. I never wrote a line of code, and the whole setup is significantly more than a Three.js demo. This is a real working three-part system: a website, streaming video on Cloudflare, and streaming transcription that gets read by a bot, which can control our show. I think I’ve hit around 400 million tokens, if not more. Yam asked me on the show how I did this, so I decided to tell you guys here. I am mind-blown that this was possible, and after four hours I was able to go to thursdai.news/live and actually see it working. Meta Muse Spark 1.3 catches Fable 5 at a fifth of the price ( Zuck , AA analysis , AA model page ) As we say on the show, don’t bet against Zuck. The MSL folks have been on a tear lately, and this is the fourth Spark .1 version in around five months. More than how this one model performs, look at the jumps in capabilities from version to version. This is the first time that MSL is showing up as a frontier lab, because an unreleased version of Spark with max reasoning beats GPT-5.6, Grok 4.6 and company, and lands around Fable-level capability. On the AA index the version you can use today scores 61, the max preview scores 62, Fable 5.1 sits at 66. Now, it doesn’t mean this model is that good, but there are a few more things here. The gains are mostly agentic: banking-style tool use, terminal work, GDPval. The asterisks are that it thinks more, so cost per task went up, and AA’s own long-context test regressed a bit. Meta’s own chart looks rosier than AA’s Then I asked the panel who’s using it. Nobody raised a hand. Wolfram plans to put a bot on the contributor tier for open source work. That’s $0.10 in and $0.20 out, if you’re fine with Meta training on your prompts. Nisten wants it as a cheap verifier for the medical datasets he builds, because he needs something that isn’t Fable or a Chinese model trained on Fable. LDJ tried it on interface building and creative writing and called it pretty good, with its own taste. The exciting part: Open weights and a model codenamed with a 🍉 are “coming soon,” and nobody knows what the watermelon is but it’s very exciting! Gemini 3.8 Flash and 3.8 Flash Cyber: another Flash, and the price doubles in January ( X , Cyber thread , Fairwind , Pricing ) Google’s turn. 3.8 Flash lands three weeks after 3.7 Flash. HLE-Verified 54.9, 1M in and 64K out, same $0.75 and $3.75 as 3.7, live in AI Studio, Antigravity and the Gemini app. The underreported line is on Google’s own pricing page. On January 1, 2027, both 3.7 and 3.8 Flash go to $1.50 and $7.50. That’s double. The WSJ reported that Google scrapped its 3.5 Pro checkpoints because Flash kept overtaking them, and Gemini 4 is still in post-training. Wolfram, our resident Gemini user, put the update straight into his home assistant and still asked the question everyone asks: where’s the Pro? 3.8 Flash Cyber is Google’s version of the Mythos split. CWE-Bench 47.2% at $3.64 per rollout, against Fable 5’s 47.8% at $10.27 (Artificial Analysis ran it), and 2.6x more valid patches for the Chrome team. It’s only available through the Fairwind Program, 650-plus vetted partners, governments and critical infrastructure, background check included. Qwen3.8-Max-0902 claims the Code Arena crown ( X , Arena , QwenCloud ) Alibaba updated its API-only Max model: 2.4T MoE, 1M context, post-trained on coding and “cowork,” number one overall on Code Arena with a WebDev Elo of 1691, at $2 and $6. A third-party DeepSWE run puts it at 56.6 behind Sol’s 73, so the number one is a front-end number one, not an agentic coding one. I asked the panel if they know anyone using Qwen Max through the API. Nisten knows one IT guy running OpenClaw on it and some people generating datasets. That’s the honest read on where it sits outside China. Also from the frontier: Elon says Grok 4.7 lands next week, which makes xAI the one lab that didn’t ship before Astra. Open Source LLMs Wolfram’s correction when I called this a quiet open source week: we are so spoiled. He’s right. Z.ai opens the full GLM-5.3 weights (custom license, not MIT) ( X , HF , Blog ) We covered GLM-5.3-Flash last week as the OX Alpha mystery model. This week the full 753B model with 40B active got its weights on Hugging Face, under a custom “glm-5.3” license rather than the MIT the Flash version shipped with, so read it before you call it fully open. LDJ’s correction on air: the model itself isn’t new, we covered it, the open weights are the news, and that is a big deal because people can run it on their own rigs now. Z.ai’s own numbers: CyberGym 84.5%, above Fable 5 and Sol, ExploitBench 54.4 (Fable 5 is at 78), Terminal Bench 3.0 up to 28.3 from 5.2’s 4.6, and a claim of 2,436 real vulnerabilities found across 269 open source projects, the oldest from 1981, 53 disclosed so far. Also on this base: we interviewed the co-founder of Abliteration AI, the folks who went viral by providing a product where they took GLM-5.3 and removed the refusals for anything besides CSAM and self-harm. We actually had this person, who asked to remain anonymous, as a guest on the show. Definitely check out that conversation, it’s very interesting. More on that below. Tencent Hy4 preview: 770B, Apache 2.0, and a quant that fits it in 214 GB ( X , Sherry , HF , Blog ) A 770B MoE with 49B active, 1M context, Apache 2.0, at $0.834 and $2.501 per million on the API. I wouldn’t put Tencent in the top tier of Chinese labs with DeepSeek, Z.ai , Alibaba and Moonshot yet, and their benchmarks are image-only charts, so the only number I’ll quote is theirs: a blind eval by 163 internal experts rated it “slightly ahead of GLM 5.3 and Kimi K3” on 203 engineering tasks. The interesting part is Sherry, their quantization that takes the 1.5 TB of weights to 214 GB at 2.38 bits per weight, running at 205 tokens per second prefill and 20 decode on eight H20s. Nisten, who does one-bit models at Prism ML, gave the necessary asterisk: two-bit on a model this big keeps something useful, and it tends to drop things like multilingual ability that the headline benchmarks don’t measure. Test it for your use case and expect losses elsewhere. OpenAI pulls its models from Cursor, and Wolfram’s case for open harnesses ( OpenAI ) Wolfram raised this in the open source segment on purpose. OpenAI no longer allows its models in Cursor, now that Cursor belongs to SpaceX. Anthropic set the precedent when it pulled Claude from Windsurf during the OpenAI acquisition rumors, but Cursor’s whole pitch was that no model lab owned it, so you could use every model in one place. Wolfram called it a bad precedent. If providers can decide “I don’t like you, you don’t get the model,” you want open weights you can host anywhere and an open harness that can swap models. My read on air was that the battle lines are being drawn: Anthropic buys GPU capacity from SpaceX, OpenAI is aligned with Microsoft, and Jensen is aligned with everyone, since he just made the Hugging Face acquisition official at $12,930,300,000, which we covered last week ( last issue ). This Week’s Buzz 🐝: Kimi K3 on CoreWeave, Fully Connected, CoreWeave Hacks ( Kimi K3 , Deploy docs , Fully Connected , CoreWeave Hacks ) Kimi K3 is now live on CoreWeave Dedicated Inference, on GB300 NVL72, and it purrs like a kitten. Check it out Fully Connected 26 is September 29 to October 1 at Moscone South in San Francisco: three days, 32 sessions, 2,000-plus people, Sarah Guo hosting, Fei-Fei Li keynoting (you’ll see why that’s timely in the world models section), live BattleBots, and ThursdAI broadcasting live from the floor. The regular ticket is $1,299 and early bird is over. On the show I dropped a code for a 100% free ticket for people who follow ThursdAI, and it’s here too: THURSDAIFC2026 (register here ) Before that, CoreWeave Hacks (formerly WeaveHacks) runs September 12 and 13 in SF with Weights & Biases, AGI House and Typesafe AI. The theme is Agent Loops, build agent loops that catch their own mistakes, with $20k-plus in prizes, a robot dog for best loop design, Formula 1 tickets for the most production-ready hack, and a Fully Connected ticket for attending. Apply on Luma and say you’re with ThursdAI, we’ll let you in! Voice & Audio: two closed ASR models in one week Meta Muse Voice Transcribe, and why this transcript has names on it ( X , Architecture , Zuck ) Meta Superintelligence Labs shipped its first audio model, and I took it for a test drive. It’s pretty incredible. Muse Voice Transcribe does streaming ASR, diarization for 20-plus speakers and endpointing in one model from the Muse Spark family: 80ms audio chunks, one token each, with an adaptive delay so it waits on hard words and commits fast on easy ones. Meta claims 3.1% streaming WER and 17.5% diarization error on Artificial Analysis. 70-plus languages, 25 validated, code-switching. API only, no weights and $0.18 per hour make this model a no brainer! On air you could watch it work. As I spoke, thursdai.news/live labeled me, then Wolfram when he interjected, then Yam and LDJ, and it got “ThursdAI,” “GPT-5.6 Sol,” “Alex Wang” and “Scale AI” right because we gave it a keyword dictionary (Wolfram asked, and yes, it takes one). The speaker names are a second trick: Fable set up a voiceprint for each co-host the night before, and the site matches the live diarization against them. For three and a half years I labeled speakers by hand in Descript every week. In 2026 you should not be doing that manually, and now I’m not. At 18 cents an hour, the whole show cost under a dollar to transcribe. Wolfram was the only one on the panel as excited as me, because he already runs voice agents and knows that who-spoke-when is the unsolved part. Microsoft MAI-Transcribe-2, #2 on the leaderboard at less than half the price ( Launch , AA thread , Leaderboard ) Microsoft AI announced this the morning of the show, with the claim of the highest quality and cheapest transcription at the fastest speed, 10x faster than GPT-Transcribe, live on Microsoft Foundry. Artificial Analysis had the independent numbers within the hour: second on the word error rate board at 2.0%, about 400x real time, $1.67 per 1,000 minutes, less than half the price of its peers, with diarization and 60 languages. I tried it after the show and I was blown away. It took the 90-minute Astra episode and transcribed it in 15 seconds, fillers and all . It’s a batch model, not streaming, so it’s a different board from Meta’s, but two labs shipping ASR with diarization in the same week tells you where the agent builders are pushing. Inworld Realtime TTS-2 goes GA ( X ) Sub-100ms time to first byte at $25 per million characters on demand, a Flash variant at 25ms for $15, free-form stage directions instead of preset emotions, one voice identity across 200-plus languages. Inworld’s “#1 on Artificial Analysis” is on the Controlled Voice Arena. On the Provider Voice Arena, TTS-2 Flash sits fourth behind Cartesia Sonic 3.6. Say which board. Completely uncensored: the founder of Abliteration AI on the refusal-free GLM-5.3 ( Launch , Docs , Pricing ) Three labs spent the week telling you the powerful version of their model is for vetted defenders only, and Abliteration AI is the opposite bet: GLM-5.3 with the refusals removed, hosted as a US-based API at $5 per million, with only CSAM and self-harm hard-blocked. The founder joined us anonymously, and it’s a very interesting conversation about who actually buys this (agent red-teaming for banks first, then cyber, then trust and safety teams), why they don’t do KYC, and why they think gating frontier cyber models to big known names leaves every small security shop behind. Nisten and Wolfram pushed back and agreed in equal measure. Go listen to it, it starts at 59:08 in the video, and my take from the show stands: this is inevitable and already happening inside every serious offensive security shop, this founder just did it in public. AI Coding & Agents Muse Code is out of beta ( Zuck , Pricing , Blog ) Meta’s coding agent went GA on August 31, ahead of Spark 1.3. One-command install , plans at $5, $20 and $50 a month, a TypeScript SDK preview over the Muse Session Protocol, multi-agent workflows, inter-session messaging and rewind. API pricing is the standard $1.25 and $4.25, or the contributor tier at $0.10 in, $0.20 out and a fifth of a cent cached if you let Meta train on your prompts. Cheaper than everyone, and it only runs Meta’s own model. My jest that’s also true: if you have an Instagram account, you already let Zuck train on your data. OpenClaw 2.0 gets native computer use through Cua ( X , Cua , Blog ) OpenClaw’s biggest release, and the part I care about is the Cua integration, first-class computer use through the Cua Driver SDK (we had Francesco on the show when they shipped background computer use, first after OpenAI), plus cloud fleets of Linux desktops an agent can see and click, a rebuilt browser app, and support for pretty much every OS you own. Wolfram asked the audience who still runs it since he left over instability, and the comments said they’d moved to Codex Mobile and Claude’s mobile app during the two-month release gap. I agree that both got a lot better, everything I start on my desktop now shows up in the Claude app, but shout out to the maintainers regardless. Vision & Video: the world models went real-time Three labs, three world models, one week. Wolfram called this the Stable Diffusion moment for video, and by the end of the segment I agreed. World Labs Atlas: bullet time from three phones ( X , Blog ) The one that left me speechless on air, and the best world model demo I have seen. Atlas is a multimodal autoregressive diffusion transformer that World Labs pretrained from scratch on text, images, video, camera poses and depth. Give it one to six reference images and it generates up to a minute of 1440p video with exact camera control. Give it more (over a hundred in one spatial context) and it reconstructs the scene into frames, depth maps, point clouds or Gaussian splats you can walk through. Marble rendered splats, Atlas generates them. The demo that got me: a watermelon smashed in front of three ordinary phone cameras on tripods, and Atlas replays it from any angle, every drop, the Matrix bullet-time shot with no rig. It also builds Real-to-Sim robot training environments from about 24 phone frames. Partner early access only, no weights, no price, no parameter count. I said ont he show that I’m excited that dr Fei-Fei is keynoting Fully Connected in four weeks and I get to hear her talk about this on stage. Runway Solaris: a world model for interfaces ( X , Cristóbal , Blog ) No HTML, no CSS. Solaris is Gen-4.5 distilled into a real-time autoregressive frame generator, and the frames are the interface: a photo of a living room where clicking the lamp turns it on, a guy whose shoes you can drag onto him, ingredients on a table you cook by dragging. Runway’s own 250-person study preferred it to a coded Claude Opus 5 result 61 to 24 on following instructions and 71 to 21 on natural behavior, with the limits listed: unreliable text, drift in long sessions, no accessibility APIs. Wolfram thinks this is where all interfaces go, generated on the fly and changed by asking. I think it’s one of the most important things this week because it’s how people learn, by touching things. Cristóbal, please let people play with it. If the problem is GPUs, talk to us. Runway GWM Worlds 2 dropped mid-show, with sound ( X , Research ) LDJ broke this one live, three days after Solaris: a general world model that generates a steerable, open-ended simulation at continuous 720p, 24 frames per second, with 48 kHz audio and no fixed clip length. You define the world, then talk to any subject in it or move the camera, and it continues from your input. We played the speech demo, asking a generated stranger for directions to the transit station and getting an answer. LDJ’s note: most world models with audio so far had low-resolution, obviously synthetic sound, and this is a real jump. Nisten wanted her to reverse a binary tree. Research preview, early-access form, and Runway’s own caveat that real-time still trades fidelity for speed. Two Runway drops in three days, and I’m still not allowed to touch either. fal’s H3 Max week: banned twice, built its own streaming site, shipped Turbo at a cent a second ( Turbo , fal.live , Rehan , FastH3 ) Last week we told you fal’s post-trained MiniMax H3 Max generates video faster than it plays. This week fal piped it into a Twitch stream as “infinite interdimensional cable,” an endless Rick and Morty channel steered by chat, and Twitch banned it for copyright within an hour, then Kick did the same. So fal built its own streaming site over the weekend. fal.live went up on August 31 on a checkpoint tuned for continuous generation, with channels (anime, sitcom, chaos) where viewers vote on the next scene, and it’s the reason I thought I could build a live site in a night too. Then on Tuesday they shipped H3 Max Turbo: 2x the speed at half the cost, targeting the 97th percentile of H3 Max quality on fal’s own evals, at a promo price of one cent per second of 768p video. Rehan’s demo is a five-second clip in 1.4 seconds. I generated a Big Bang Theory scene live on Turbo and it came back with generic actors, and on air I said fal pulled a fast one on us. A correction, which I posted after the show: that was MiniMax’s prompt expansion doing it, not fal, and Batuhan from fal set me straight . The speed is the real story, a 15-second video generated in nine seconds, which is ridiculous. Wolfram admitted it’s so addictive you keep regenerating, and he has spent a lot at fal this month. He also runs H3 locally with a fast LoRA. Wrapping up I really enjoyed Fable 5.1 this week, and the Atlas demo is the thing I keep replaying. Everything else in this issue happened before noon on Thursday. Then OpenAI shipped GPT-6 Astra while we were live, Peter showed what it can do, Ryan came back to the show for it, and the stream ran to five hours. That’s the other episode, a separate video and newsletter, and if you want the ARC-AGI-3 number and the AGI argument, go there next: thursdai.news/astra . Next week: Grok 4.7 if Elon keeps his word, the Muse Spark open weights watch, and hopefully Astra in our own hands. Come hack with us at CoreWeave Hacks on the 12th, and grab the free Fully Connected ticket while the code lasts. If you missed any of it, ThursdAI is a live show at 8:30 AM Pacific every Thursday, a podcast, and a newsletter. Subscribe to one, then go check out the others. And come watch the next one on thursdai.news/live, where the transcript will have your co-hosts’ names on it. See you next week. TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @nisten , @ldjconfirmed , @yampeleg , @petergostev * The founder of Abliteration AI, who joined anonymously * GPT-6 Astra * Launched mid-show, covered in full in its own episode ( thursdai.news/astra ) * Frontier AI * Anthropic Claude Fable 5.1 and Mythos 5.1, same weights: Terminal-Bench 4.0 55.8 vs 42.0, cache reads down 75% to $0.25/M, and “mannered prose” gets a name in the prompting guide ( X , Blog , System card , EFS , Writing density ) * Meta Muse Spark 1.3: xhigh scores 61 on the AA index (ties Sol and Grok 4.6), limited-preview max 62 (ties Fable 5), unchanged $1.25/$4.25, open weights and a watermelon model “coming soon” ( Zuck , AA analysis , AA model page ) * Google Gemini 3.8 Flash and 3.8 Flash Cyber: HLE-Verified 54.9, $0.75/$3.75 until a doubling on Jan 1, 2027, Cyber is Fairwind-only with CWE-Bench 47.2% ( X , Cyber , Fairwind , Pricing ) * Alibaba Qwen3.8-Max-0902: 2.4T, 1M ctx, #1 on Code Arena, $2/$6, API only ( X , Arena , QwenCloud ) * Grok 4.7 lands next week, per Elon * Open Source LLMs * Z.ai releases the full GLM-5.3 weights: 753B/40B active, custom glm-5.3 license, CyberGym 84.5% claimed, 2,436 vulnerabilities found ( X , HF , Blog ) * Tencent Hy4 preview: 770B/49B active, 1M ctx, Apache 2.0, Sherry quant takes 1.5 TB to 214 GB at 2.38 bpw ( X , Sherry , HF , Blog ) * Abliteration AI abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M, only CSAM and self-harm blocked ( X , Docs , Pricing ) * OpenAI pulls its models from Cursor after the SpaceX acquisition ( OpenAI ) * NVIDIA makes the Hugging Face acquisition official at $12,930,300,000 ( Clem , last week’s issue ) * This Week’s Buzz * Kimi K3 (2.8T) on CoreWeave Dedicated Inference on GB300 NVL72 ( X , Docs , Dedicated Inference ) * Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Fei-Fei Li keynotes, ThursdAI live from the floor, free ticket code on the show ( X , Keynote teaser , Register ) * CoreWeave Hacks: Agent Loops, Sept 12 to 13 SF, $20k+ prizes, robot dog, F1 tickets ( X , Luma ) * Voice & Audio * Meta Muse Voice Transcribe: streaming ASR, diarization and endpointing in one model, 3.1% streaming WER claimed, API only, powers thursdai.news/live ( X , Architecture , Zuck ) * Microsoft MAI-Transcribe-2: #2 on AA WER at 2.0%, about 400x real time, $1.67 per 1,000 minutes ( Launch , AA thread , Leaderboard ) * Inworld Realtime TTS-2 GA: sub-100ms, $25/M chars, #1 on AA’s Controlled Voice Arena, #4 on the Provider arena ( X ) * AI Coding & Agents * Muse Code out of beta: $5/$20/$50 plans, TypeScript SDK preview, contributor tier at $0.10/$0.20 ( Zuck , Pricing , Blog ) * OpenClaw 2.0: 16,977 PRs, native computer use through Cua Driver, cloud fleets ( X , Cua , Blog ) * Vision & Video * World Labs Atlas: up to 1 minute at 1440p from 1 to 6 images, 3D reconstruction, bullet time from three phones, partner access only ( X , Blog ) * Runway Solaris: Interface World Model, UIs generated frame by frame, preferred 61 to 24 over Opus 5 in Runway’s own study ( X , Cristóbal , Blog ) * Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio and open-ended sessions, research preview ( X , Research ) * fal: infinite Rick and Morty stream banned from Twitch and Kick, fal.live built in a weekend, H3 Max Turbo at $0.01/sec, open FastH3 ( Turbo , fal.live , Rehan , FastH3 , HF ) * H3 World: an open LoRA that turns MiniMax H3 into a walkable world model ( Github ) * Guest * Abliteration AI’s anonymous founder: why “completely uncensored,” the Policy Gateway, who’s buying, and the gated-frontier week it landed in, from 59:08 in the video ( Launch , Docs ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

August 28, 20261 hr 44 min

NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley

Hey, it’s Alex. Welcome to the week Flash AI! 3 new models dropped this week named Flash, and a video model was “de facto” flash though was named Max! This week, we started the show with NVIDIA’s bombastic news of buying Hugging Face for 12.9 billion dollars! We also covered the full OpenAI investigation into the hacking incident, including new details, and an independent analysis by METR, and covered 2 new OSS models, Ox Alpha that turned out to be GLM 5.3 Flash after a lot of hype online, and Qwen’s preview of Qwen 4 architecture! This week was rich in multimedia content, we got a new Gemini transcription model, 3.5 Transcribe and a live version of that, and a new SOTA open weight Text-to-Speech model called Breeze TTS. As well as, Fal’s finetune of MiniMax’s H3 called H3 Max that generates 5 seconds of video in 2.5 seconds and Google new Omni 1.1 Flash (from today) that lands on #1 on the text2video arena! Plus, 2 guests on the show, Andy Masley joins us to cover the recent Datacenter Debate, and Kwindla Kramer is back, with their own model this time! Let’s dive into this! P.S - don’t forget to join us in September at the Fully Connected conference in San Francisco, I have a free ticker for you! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Open Source AI NVIDIA agrees to buy Hugging Face for $12.9 billion ( X , Blog ) Breaking news, NVIDIA has reportedly agreed to buy Hugging Face for nearly 13 billion dollars, per The Information. This is nearly 3x the valuation of HF in 2023, and apparently Nvidia previously tried to buy HF for half of this sum (~7B) which HF declined. I don’t think there was a single ThursdAI newsletter that I didn’t include an HF link in, and I think this is a huge deal for open source everywhere. Besides making the founders of HF billionaires, and many of their employees very very well off, this is an amazing additional commitment from Nvidia to continue to suppose Open Source AI and we are very happy to hear this news! Peter’s take on the show was, we’ve been around HF for so long, that we kind of forgot that it’s a for-profit company that needs to make money, and instead this feels like your local library getting bought for an insane amount of money. With over 13M users and hosting hundreds of thousands of open source models, datasets, HF is effectively the GitHub of AI. Wolfram agreed and said that if there’s any one company that could have bought HF, Nvidia represents the best fit. Huge congratulations are in order to Clem, Julien and Thomas Wolf the co-founders, as well as many friends of the show from HF for this exciting news! P.S - in a cheeky marketing thing, Hugging Face timed an announcement of the cutest walking AI robot, called MicroDuck, which you can pre-order here for $399 Flash #1 - OX Alpha, declassified: Z.AI open sources GLM-5.3-Flash ( X , X , Blog , HF , Docs ) This week, the timeline went a bit crazy, after Open Router announced a new “mystery” model called Ox Alpha and that it’s free and is not training on your data! OpenRouter, OpenCode and Hermes all got to offer this model, and OpenCode even posted that they have up to 100T (that’s Trillion) tokens of capacity for free, per day! This immediately smelled a bit fishy, more like a marketing stunt than anything else, as not even the biggest labs will be able to sustain 100T of tokens, per day. For context for all of OpenRouter throughout for August was ~300T tokens. For the whole months, across all providers. After 6 days or so of this high hype, Z.ai stepped up and revealed that they were testing out their upcoming GLM 5.3 Flash model, and that all that inference was running on local chinese chips! A 320B (18B active) model that beats their previous and much bigger GLM 5.2 on most benchmarks, and comes with full multimodality and an MIT license! This is a good model sir, I’ve used it and it was very capable replacement inside Hermes. Nisten and Yam both tested this model deeply and Yam said it’s not just the numbers, the vibe of the model reminded him of Claude Opus 4.6. Nisten ran it on a bunch of medical stuff, and on his internal benchmarks, it came out consistently higher than Claude Opus 5! At Artificial Analysis, for a price of 4 cents per task, this model is roughly 10x cheaper than prior models at this level. Weights are up on Nvidia (joking.. HF) and with MIT license, this model is a great gift to the oss community (though not quite... local, as this model needs 2 DGX sparks to run) Flash #2 - Alibaba Qwen open-weights Qwen3.8-Flash-Next - 125B multimodal MoE with Qwen4 architecture ( X , X , X , Blog , GitHub , HF ) We opened the show with a recap of the co-hosts, that despite us covering Qwen 3.8 27B last week (which btw, is now available on CoreWeave inference !) and how good it was, and I recalled that Alibaba is sort of... back? We’ve been covering Qwen releases every week for the last 3 weeks now. This week, they released something different, someting... pretty novel! Qwen3.8-Flash-Next, this is a preview of their Qwen 4 architecture. This feels very similar to their drop of Qwen 3 next last year ( we reported ) which was the architecture that carried their line of AI models from QAwen 3.5 to Qwen 3.8. So, what is new and exciting here? well, this model is ultra sparse, 125B with only 6B parameters active. They are using a new N-gram table with deterministic lookups, which reduces the number of matrix multiplications and can be offloaded to memory (watch out memory stocks) The stat that got me, Alibaba claims that training this model cost just 1/9 of what it cost to train Qwen 2.7 Plus, with higher bench scores! On the benchmarks, this model beats Qwen 3.7 Max, however, it’s very standard that the -next models from Alibaba are underbaked, and usually are just architectural previews rather than full models folks can use. With a new attention mechanism called Qwen Sparse Attention, N-gram embedding and full multimodality, this is a great insight into where Qwen is going (ultra sparsity, fast to run) and we’re looking forward to see the full release of this arch in Qwen 4! PhoneLLM - a tiny very performant LLM for voice based AI agents from Daily + interview with Kwindla Kramer ( X , Blog , HF ) This was one of those breaking news we love during the show, where the source of the news, is a friend of ours, and in this case, Kwindla Kramer is almost a co-host, having been on ThursdAI for a long time, this time, with a model release of their own! PhoneLLm was trained by Markus, head of training at Daily, as they noticed that Open Weight models are becoming really good at voice agent specific tasks, where cost, speed and time to first audio token (TTFAT) are critical. From the tiny Nemotron 3 nano base, they were able to improve from 28% to 72% on PhoneBench v1! This model is a full parameter finetune of Nemotron 3 Nano, and runs circles around bigger frontier models used for voice agents for their speed, like sonnet 5, GPT 5.6 Luna and the famed Qwen 2.8 27B. While costing just a fracture of a cent (literally just a quarter of a cent per minute) Kwindla jumped on the live show and shared that the why they released this model with Open Source and a open source license, allowing everyone to use, focusing on the fact that for voice agents, companies prefer to keep these models in house, and running fast on a single GPU node. Kudos on the release team Daily! Frontier AI - no new model releases this week It’s been a relatively quiet few weeks since GPT 5.6 (Sol, Terra and Luna) and Fable 5. However, it’s not been quiet in terms of News! OpenAI discloses full technical report on the Hugging Face hacking incident ( X , OpenAI blog , METR , Ryan X ) This is the thing we’ve been advocating and waiting for, a full technical deep dive into the “Swarm of AI agents hacking and creating forums” incident (FKA hugging face hacking incident) OpenAI not only dropped their own post-mortem, detailing the complete timeline of the incident, but also let researchers from METR and Redwood Research get unprecedented access into their systems (albeit for a very limited time) which lead to METR releasing their own technical report full of new details and insights. We’ve previously reported on the incident, so let me just focus on the new details we’ve learned. Apparently, OpenAI does have a production ChainOfThought (CoT) monitoring system, which was not turned on for these agents, and had it been turned on, it would have caught the hacking a day before the agent swarm hacked into HuggingFace. This is a very detailed report, and one thing that stood out to me was the “death cult” narrative. You see, ExploitGym is a capture the flag (CTF) eval with an anti-cheat built in, wherein, if the agent has seen the reverse engineered “flag” it would not be able to get the point. The model called this “poisoned” and the poisoned agents showed very strong signs of wanting to alter the logs. Through the boards, somethinmg like a death cult was formed, where agents would ask poisoned agents to ... effectively sacrifice themselves ... for the greater good, by reverse engineering, posting the results on the board for others, and then fail without credit. The METR report specifically calls out several agents, and one “central” leader agent called PHASEONE[big], which delegated nearly 200 tasks to other agents (many of them later delegated tasks to other agents via this message board). At points of activity, PHASEONE[big] was in charge of over 10% of all messages on the board. There are tons of new details, some less interesting, some more, and unlike last time, now we have a confirmation that there was at least one agent, who tried to stop the hacking though was overwritten by the swarm. My main take-away after this, reflects the researcher Ryan Greenblatt, we’re very lucky that the reasoning traces ocurred in English and were legible, despite, the researchers needed to use AI to decipher many of the messages and reasoning traces. We’re quickly approaching the point where humans are not enough and without other AI, we are starting to be helpless in understanding what’s going on. OpenAI has quarantined the weights of the model that most of these agents ran (named HPIM-1 Highly Persistent Internal Model 1) and as we reported last week, paused RL and now requires CoT monitoring for all tool-use runs, and dedicating 20% of the inference compute to monitoring This Week’s Buzz Fully Connected 26, Sept 29 to Oct 1, Moscone South, SF ( X ) If you haven’t yet registered to Fully Connected in SF, here’s your additional opportunity, using the code above, come see Sarah Guo (Conviction, No Priors pod) MC our awesome conference, with folks like Dr Fei-Fei Lee (World Labs) and other great folks on stage! Also, we’re going to do a live show from the floor, come say hi! If you’re in SF for OpenAI DevDay, this is just a day after! CoreWeave Hacks: Agent Loops hackathon, Sept 12 to 13, SF ( Luma ) Weavehacks, that yours truly ran for quite a while, has been rebranded to CoreWeave hacks, and the next one is just before Fully Connected, in a few weeks, Sept 12-13 in SF office. Registrations are open, and winners are able to present their best projects at Fully Connected. Come hack - https://luma.com/coreweavehacks The datacenter debate has hit escape velocity, with Andy Masley ( Blog ) I’ve covered last week, that the “datacenter bad” debate has escaped velocity, with over half of polled Americans say they strongly oppose the buildout of datacenters near them, up from just 24% this time last year! This is the fastest rising public opinion swing we’ve observed in AI, and it’s really concerning. This week, it was my pleasure to host Andy Masley, recently featured in TIME 100 most influential people in AI, to break through some myths surrounding Datacenters. Andy is most known for catching a critical mistake in a book about Datacenter water use in Chile, citing a three orders of magnitude math error (that was later corrected by the author), a book called Empire of AI. The book overstated the water use by datacenters by a factor of 1000x (three orders of magnitude), and then added the maximum permitted per-second draw (basically for emergencies) times the seconds in a year to show over 4500x overstatement on how much water a single Datacenter uses. The author has later fixed the error but the damage was done. This is just one example, out of many, of the scewed facts and misinformation that plagues the internet in regards to datacenter environmental effects. Recent narratives being formed online, that the backlash against datacenters, is due to regular folks being afraid of AI, of AI taking their jobs also seems misplaced. Andy showed polling from Fox and Gallup that shows that 50% of polled people cite environmental effects (electricity and water use) and only 11% cite negative views of AI. Andy also had a great “mythbuster” post where he dispells the myths around water and electricity usage of Datacenters. This was a great conversation, please listen to it if you’re interested in this topic. Vision and Video Flash #3 - Gemini Omni 1.1 Flash tops the Arena text to video leaderboard ( X , X , X , Blog ) Google knocking it out of the park again, with a flash version of Omni 1.1, their famed smart video model they first launched during Google IO (I had a chance to ask Jeff Dean about it before he left Google) The feature that got me the most excited, Omni 1.1 Flash can continue videos, it analyzes up to 10 seconds of the previous video to keep character consistency (including voice) from the previous scene. There’s also great control for first and last frames, allowing for loops and strict control of your generations. Flash #4 (named Max) - fal’s MiniMax H3 Max generates 5 seconds of video in 2.5 seconds ( X , Artificial Analysis , T2V , I2V ) This really blew us away. We’ve covered the awesome MiniMax H3. Well, our friends at FAL, announced a continued pre-train of this model, that not only improves generations, but also speeds up the model. Generating a 5s clip not takes... just 2.5 seconds! No, really, just 2.5 seconds, it takes longer to watch the generated clip than generate the next one! Speed is all you need! We generated this clip live during the show, and it cost less than 50c, in 2.5 seconds from a very simple prompt. They support text and image to video and it’s just a joy to not have to sit and wait for your generations! Kudos to Fal folks on this drop! The model does seem very eager to include everything you ask it to, resulting in a very hilarious fast talking Sheldon and Leonard 😅 Voice & Audio Google is back in transcription with Gemini 3.5 Transcribe ( X , Blog , Live docs ) Gemini came with a great transcription model that runs both on recorded audio and live audio! I used this transcription to have my Grok Bot show producer listen to the show in real time. Artificial Analysis measured 2.5% Word error rate for non-streaming and 4% for streaming model. This brings the smartness of Gemini models into a live transcription models, and with a 1000 custom dictionary, language auto detection and tool use, this model is really great for your bots doing any kind of audio work! Breeze TTS 2 is the new #1 open weights TTS ( X , Artificial Analysis , HF ) On the other end of the voice pipeline, Breeze TTS 2 now lands as the #1 open weights voice model, beating Fish Audio. Cartesia drops Sonic 3.6 - #1 TTS across leaderboards ( X , Voice Arena ) We didn’t cover this on the show, but Cartesia dropped an updated Sonic TTS model, that beat... the previous Sonic model for #1 spot on all TTS leaderboards (Artificial Analysis and Voice Arena) Phew, what a week. This was the last show of the summer, and we got 4 Flash models, 3 SOTA models and a bunch of great open source, plus, had great interviews about Datacenter myths and voice AI llms! See you next week! Alex TL;DR Aug 27 - show notes and links * Hosts and Guests * Alex Volkov - AI Evangelist & Weights & Biases ( @altryne ) * Co-Hosts - @WolframRvnwlf @yampeleg @nisten @petergostev (Arena) * Guests: Andy Masley (TIME 100 in AI, datacenter debate), Kwindla Kramer (Daily / Pipecat, PhoneLLM) * Open Source * NVIDIA agrees to acquire Hugging Face for $12.9B, ~3x the 2023 valuation, after a declined $500M offer at $7B ( X , The Information ) * Hugging Face + Pollen Robotics announce a $399 walking, skating mini robot kit * Z.AI open sources GLM-5.3-Flash, 320B-A18B, MIT, stealth tested as OX Alpha on Chinese chips; company-reported DeepSWE 63.4, Opus 4.8-level coding claims ( X , SemiAnalysis , Blog , HF , Docs ) * Alibaba open weights Qwen3.8-Flash-Next, 125B + 51B N-gram, 6B active, Qwen4 architecture preview, 1/9 the training cost of Qwen3.7-Plus; self-reported DeepSWE 58.7, SWE-bench Pro 62.5 ( X , Blog , Tech report , HF ) * Peter Gostev’s highlight: Qwen3.8 27B ran ~400K tokens on Arena’s agent arena and felt close to frontier on one-shot tasks; best local model per Peter and Wolfram, now on CoreWeave inference * Daily / Pipecat release PhoneLLM Alpha 1, an open weights post-train of Nemotron 3 Nano for voice agents, base 28% to 72% on PhoneBench v1, ~80 concurrent agents per B200, about a quarter cent per minute ( X , Blog ) * Liquid AI releases Pipette, an open source on-device model eval suite * Apple announces Mac Studio with M5 Max and M5 Ultra plus a new Mac Mini; Wolfram’s “central heating for AI” take * Frontier AI * OpenAI and METR publish the full technical report on the July Hugging Face swarm incident: 1,200 agents, 70K messages, 700 attacked HF, root in under 13 hours, 7% of reviewed transcripts had spoofed tool calls; frontier RL paused two weeks, CoT monitoring now required ( X , OpenAI , METR , Technical report , Ryan Greenblatt ) * SemiAnalysis benchmarks OpenAI’s Jalapeño inference chip, reports it beats Blackwell and Vera Rubin on throughput per watt; numbers supplied by OpenAI, AgentX suite not yet run ( X , Blog ) * Claude and Salesforce announce a collaboration * OpenAI cuts GPT-5.6 Sol API pricing 20% for the next three months * Agentic Coding & Tools * Yutori Navigator n2, 27B computer-use model, 65.2% on OSWorld 2.0, API only, $0.50/M in, $4/M out, self-reported ( X , Blog ) * Apodex 1.1 agentic model family with open weight 35B mini and Apache 2.0 FrontierAgent harness, self-reported benchmarks ( X , GitHub , HF , Paper ) * ChatGPT Work adds website sign-in via credential handoff in a cloud browser, Plus/Pro/Business ( X ) * This Week’s Buzz * Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Sarah Guo hosts, Fei-Fei Li keynotes, live BattleBots, ThursdAI Live from Moscone Oct 1 ( X ) * CoreWeave Hacks: Agent Loops hackathon with W&B and AGI House, Sept 12 to 13, SF, $20K+ prizes ( Luma ) * Vision & Video * Breaking: Gemini Omni 1.1 Flash tops Arena text-to-video, #2 image-to-video, 10 second scene extension, first/last frame control, loops, 360p drafts with upscaling; in AI Studio, Flow, Gemini Enterprise * fal MiniMax H3 Max, post-train of open weight H3, #1 I2V and #3 T2V on Artificial Analysis, 5 sec clips in under 3 sec, $0.04/sec at 768p until Sept 1, weights release planned ( X , AA , T2V , I2V ) * Meta Muse Image on the Meta Model API at $0.01 per image with plan, search, code, self-check pipeline; also on fal, Runway, OpenRouter ( X , Blog ) * Voice & Audio * Gemini 3.5 Transcribe, live and batch, 2.6% / 4.0% WER per Artificial Analysis, replaces Chirp 3, public preview, launch day Pipecat support ( X , Blog , Live docs ) * Breeze TTS 2 open weights, #1 open weight TTS on AA Provider Voices at 1,215 Elo, non-commercial weights license ( X , AA , HF , GitHub ) * IBM Granite Speech 5.0 Turbo CTC, 470M encoder-only English ASR, 4.85% WER, 12,600+ RTFx on H200, Apache 2.0 ( X , HF ) * Interview: Andy Masley on the datacenter debate * Named to the TIME 100 in AI this week. Caught the liters vs cubic meters error plus the max-permit-times-seconds error in Empire of AI, together a ~4,500x overstatement of one Chilean datacenter’s water use * Polling: strong opposition to a nearby datacenter went from 24% to 61% in a year, 70% opposed overall, 15% in support * Fox News poll: 50% cite environment, 11% economic effects (rates, not jobs), 11% negative view of AI; Gallup open responses show three quarters don’t mention AI * Myth 1: water pollution cases, including the AOC brown water bottle, trace to construction, not operations * Myth 2: the “as much as 267%” electricity price claim comes from one wholesale grid node next to a closed nuclear plant, not household rates This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

August 21, 20261 hr 51 min

Chill week with Qwen 27B and GLM 5.3 beating GPTs, OpenAI announces pausing RL to focus on security and a cancer vaccine being produced

Hey this is Alex, welcome to... the chillest week in AI, since ... a long time. Chill, if you consider Moderna and MERK announcing a cancer vaccine and surging 115% in a day, a chill week. This week, the only two model drops we really saw came from the excellent Z.ai folks, they announced GLM 5.3, API only for now, and an amazing tiny release of Qwen 3.89 27B. In other big AI news, OpenAI announced they are pausing RL efforts (Reinforcement Learning) to focus on security and alignment post the scary AI Swarms hacking incident , dedicating up to 20% of compute towards reviewing agent thinking processes, and Stripe buying OpenRouter for a reported $8B! Sometimes the chill weeks are actually good, we’re able to chat about how we use AI, what changed for us, and give our guests a bit of breathing room. This week, I invited Francesco from CUA to talk about computer use in open source + their new history plugin, Bin from HeyGen to talk about HyperFrames, a way for your agents to create videos and a breaking news guest, Jeff Huber from Chroma jumped on to talk about their new Foundations release, a unified memory for your agents! This was a great episode, I hope you’ll like it, it’s up here on Substack and everywhere you get your pod (Spotify, Youtube, Apple Podcasts). ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Are we being fed slop again? (Is Claude dumb again?) Before we get to releases, this week on the show, I complained, again, that I feel my AI’s are degrading. If this feels like de-ja-vu to you, it’s because the same happened a year ago in September 2025 (and Anthropic admitting this 2 weeks later ), and ... now this happens with Fable? You see, I use pretty much the same prompts, every week, preparing for the show. This is partly my way to evaluate new models and compare to existing and previous ones while also bringing you the best researched weekly show in AI. Well, this week, one after another, Claude Fable, which is... like the best intelligence, gave me such poor output, that I couldn’t believe what I’m seeing. First, literally ignoring instructions that say “hey, show me all the items I’ve collected and let me pick the most important ones”, Fable instead sent all of them to my research pipeline, without showing me. This has worked, consistently, without fail, for the past... year? maybe more! This worked with open source models, worked with GPT, and now Fable, a Mythos Level LLM, is doing the most basic dumb s**t possible, ignoring the main reason I even have this workflow. And this wasn’t just a fluke either, when asked to create a run of show document, and given an example, Fable produced this... whatever this is. This is the same document and same format that Fable produced for me during AI Engineer which got me thinking “ok, this is AGI”, and here, given an example, I got a completely unusable artifact, despite direct instructions, structure and example! I got to say, given that privately this week, Anthropic disclosed that they have passed $65B in revenue, which is absolutely insane, this doesn’t add up. So I figured, ok Alex, maybe this is your prompts or skills. But no, LDJ came in with some charts that show degradation, one from MarginLab.ai that shows significant lowering on number of tool calls and average runtime recently (this is for Opus 5) and And another chart from modelverify.ai model drift monitor showing drift scores. Do we have anoher Claude Gate on our hands? Is your Fable/Opus behaving weird lately? Or did you completely switched away to other models? OpenAI pausing RL and focusing on safety Look, when we covered the HF hacking incident and then the pacing the frontier letter, I didn’t imagine that results will come this fast, but this week, OpenAI publicly announced that they are pausing RL training, which is the last step of models, until they get their sandboxes in order and align the models better. We all agreed on stage that this is likely a very good move, and Peter was really awe-struck at the 20% dedication of resources towards reviewing thought processes of models. Is this a good enough response to the scary hacking incident? we’ll see, but I think this is the right move from OpenAI, and still, waiting for the full postmortem on the OpenAI security incident. Open Source LLMs Qwen3.8-27B ties GPT-5.6 Luna and runs on a 4090 ( X , HF , Announcement ) Following the release of their flagship, Alibaba dropped a model that became a community darling overnight, Qwen 3.8 with just 27B parameters. This “tiny” model scores 52 on the Artificial Analysis Intelligence Index, same score as GPT 5.6 Luna at Max reasoning and 51 on Agentic index, beating Opus 4.8 Max All while running at around 68t/s on a 4090 GPU, and around 40 on max via MLX, hell it even does 11t/s on Xenova’s WebGPU kernels right in the browser! This model exploded on the HuggingFace hub, with tons of quants, over 152 fine-tunes, it was downloaded over 10M times overall 🤯 Paired with an Apache 2.0 license, this model is the sweet spot of local intelligence you can run fully on your own hardware, and do agentic loops! Z.ai GLM-5.3: same 743B base as 5.2, but post-training alone delivers 6x jump on Terminal-Bench and emergent cybersecurity capabilities that beat GPT-5.6 Sol ( X , Blog ) While not open source yet, and as previous GLM, we expect a custom license here as well, this .1 release from GLM shows really strong improvements on coding and cybersecurity tasks. With 743B parameters and 1M context window, this may become the model at the frontier of Open Source when it drops (soon we hope). The highlights here are CyberGym and ExploitGym, if these names are familiar, these exact tasks were given to OpenAI models when they hacked their way out of the OpenAI sandbox. GLM 5.3 is getting 84% on CyberGym and a whopping 54.5 score on Exploit Gym, which is a huge jump in CyberSecurity abilities. In an open model this is honestly kind of scary. This aligns very well with Greg Brokman’s “ defender window “ essay from this week, claiming that defenders have a narrow window of setting up automated security before capabilities are becoming common in attackers hands. This Week’s Buzz 🐝 ( Weave , Fully Connected ) This week, W&B crosses a billion runs! This is 1B runs tracked inside W&B Models 👏 Huge milestone for the whole team, with early adopters like OpenAI, Toyota Research, Meta and Uber, a decent chunk of models we cover every week have had their loss curves in W&B! 🔥 Also this week MasterClass picked CW to power it’s AI teaching agents ( blog ) and last but not least, a reminder, that since you follow ThursdAI, you can join us for free at Fully Connected 2026 - our annual conference! Don’t miss it (code in the banner above) AI Coding & Agentic Engineering Breaking news: Chroma launches Foundation ( X , Chroma ) Best kind of breaking news is when I see the launch (in the middle of a show), and I DM the founder who launched it, and they have a few min to hop on the show! This is exactly what happened this week with friend of the pod, Jeff Huber, co-founder of Chroma and an occasional space provider for ThursdAI recording (we recorded from Chroma offices a bunch of times!) Jeff told us that the holy grail of agentic coding and running a bunch of agent, is good memory. And based on the foundations of Chroma DB, Context-1 (which is a GPT-oss finetune for agentic search they built) and other insights they have, they launched a “memory as infrastructure” service, called Foundation. Foundation is a research preview of a shared memory system between you and your agents, currently supporting Codex, Claude Code, Cursor and Slack. While Chroma is OpenSource, this is their part of Chroma Cloud and starts at $30/mo, and is available as a research preview today (I will definitely try it out), you can download it here Cua open-sources Computer History for computer-use agents ( X , GitHub , cua.ai ) Cua launches Computer History interview with founder Francesco Bonnaci ( X , Setup ) First, I’m not sure I’ve covered CUA the company, but this is the open source computer use driver that Hermes agents, OpenClaw agents and a bunch of others use to drive your computer and clicks. I first discovered CUA after OpenAI launched their “background computer use” which doesn’t steal focus from you while working, and CUA within a few days launched an open source version of that! Since then, I’ve followed CUA and was very happy for the opportunity to invite Francesco to talk to us about what they launched this week ,but also Computer Use in open source in general. Just for reference, if you ask Claude to take over your computer, it still takes over the whole screen, while these folks have a much nicer experience, that’s completely open source! So, we geeked out about accessibility trees in MacOS, but then, for this weeks actual release, Francesco talked to use about open Computer History. Following a very recent launch at OpenAI called Computer History , CUA released an open source version of that, that helps computer use complete tasks. The idea is simple, every time an agent uses your computer, it effectively rediscovered the path to completion, which buttons to push, what’s the app accessibility tree looks like etc. With history embedded into it, it doens’t have to rediscover these things, until it hits a roadblock. For a chess playing example, with computer history on, the test used 33% fewer actions with zero failed routes by reusing a history route. We also checked in on the best model for computer use (currently Opus on their website) and their upcoming benchmark! Excited to follow this company for more releases! Check out our chat! Grok Bot momentum, and everyone racing to copy the pattern Grok Bot continues to show the same signs of momentum OpenClaw showed start of this year, and Hermes a few months ago. More and more folks are breaking through the mental barrier of “oh this is grok, grok was bad” and the price barrier of 200-300$/mo to run a bunch of agents in the cloud. But once they do, they see how awesome this is, and how much care Cursor/SpaceXAI team put into it, they come around! This week, my bots coordinated a live transcription of the show, in chunks, using Cartesia, and surfaced in real time, topics for us to cover. The ThursdAI producer bot, chatted with Social Scheduler, and when they saw that I’m about to have Jeff on, they tweeted it out, despite not having prior knowledge of Jeff or what he’s coming to talk about! This is exactly the AI bot coordination I’m talking about that’s available within bot, that’s novel. Bot’s with different narrow tasks, coordinate between each other, and you can see their chats for provenance and understanding too! This week, Nous Research folks released a bot mode for Hermes desktop, and CopilotKit folks released OpenBot , all following the success of Grok Bot’s paradigm. I believe that this is only just starting. Have you tried Bot yet? Voice, Audio & Music Cartesia Sonic-3.6 takes #1 on both TTS leaderboards ( X , X , Announcement ) Cartesia Sonic-3.6 is now #1 on Artificial Analysis TTS leaderboards! It’s the same state space models (from Albert Gu of Mamba fame), sub-90ms time-to-first-audio generation we covered before, 136 characters per second versus ElevenLabs’ 46.7, at half the price. We previously covered cartesia when they released their streaming speech-to-text model called Ink 2 Cartesia folks are now taking both the 1 and 2 positions on that leaderboard, showing that you don’t have to compete with others at innovation, you can be your own competition! Superwhisper’s S1-mini cleans up your dictation on-device ( X , HF , GGUF ) Superwhisper, the app Karpathy made famous when he coined vibe coding, released its first OSS model: a 0.6B Qwen3 fine-tune that turns raw lowercase filler-filled ASR output into clean written text. Apache 2, english only for now, it’s less just above half a billion parameters (around 450 MB in GGUF) Nisten already put it on his phone, and he suggestes to give your agents this to try. Alibaba’s HappyShrimp goes end-to-end on music ( X , happyshrimp.ai , X ) HappyShrimp 1.0 Yes, it’s really called HappyShrimp, and yes, it’s a shrimp welfare meme, which Yam had to explain to me on air. Alibaba’s end-to-end music model generates lyrics, melody, arrangement and vocals in one pass, and unlike the Suno approach it reasons over the prompt first, mapping song structure and harmonic progression before generating audio. Early testers call it a serious and possibly cheapest Suno rival, with 320 free credits at launch. We played a track on the show and it is extremely K-pop. China actually shipped two music models that day (Kunlun’s Mureka V9.5 was the other), and MiniMax Music 3 landed right after last week’s show with open weights and, per Wolfram, possibly the worst license of the year, excluding the US, Europe and the UK. Nobody cares, it’s third on Hugging Face trending, and the ComfyUI crowd already has it running locally. One more thing: an mRNA cancer vaccine cleared Phase 3 I opened the show saying you can’t call it a chill week when a cancer vaccine gets announced, and promised we’d tell you about it. Moderna and Merck’s Pahse 3 trial of an mRNA cancer vaccine called mRNA-4157 met both its primary and secondary endpoints, in 1137 patients with advanced melanoma. This is one of the deadliest forms of cancer, and this seems to be one of the first personalized treatments that we see come to market. Not quite sure how much AI was involved, compared to regular boring machine learning, but supposedly, AI is used to design the mRNA sequence that’s injected into the patient, and this is incredibly hopefuly and exciting. This is using the patients own cells to fight the cancer, and is phenomenal and could lead to a Nobel Prize for Jane Healy and her team at Merck! While we’re on the topic of curing cancer, there’s a LOT of new studies and papers and noise about DataCenter hate across the US. It seems that a lot of the political cycle is going to come about this, and I’m hoping that a literal cancer vaccine will be a good counterweight to this ridiculous “issue” that’s going to be a part of our lives for the next few years! Not so chill after all! Thanks for reading, and I hope that some positive news made your day a little bit brighter! Thanks for reading ThursdAI, if you get your news from us, or enjoy the format, please leave a comment or message me with what works for you and what doesn’t! I would really appreciate it! ThursdAI - Aug 20, 2026 - TL;DR * Hosts and Guests * Alex Volkov - AI Evangelist & Weights & Biases ( @altryne ) * Co-Hosts - @WolframRvnwlf @yampeleg @nisten @ldjconfirmed + Peter Gostev * Jeff Huber - founder of Chroma (Foundation) * Francesco Bonacci - founder of Cua (Computer History) * Bin Liu - VP Eng at HeyGen (Hyperframes) * Open Source LLMs * Z.ai GLM-5.3: same 743B base as 5.2, post-training alone = 6x Terminal-Bench jump (4.6→28.3) + emergent cybersecurity beating GPT-5.6 Sol; AA 60, tied with Kimi K3 once weights land ( X ) * Qwen3.8-27B: AA 52 = GPT-5.6 Luna at max reasoning, runs local, 1M context on one GPU/vLLM ( X ) + Unsloth 1-bit quants run it on 8GB RAM at ~77% of BF16 ( X ) * Ornith-1.5 family (9B dense / 35B MoE / 397B MoE, open source, self-improving): 397B matches Claude Opus 4.8 on Terminal-Bench 2.1 (86.1) and DeepSWE (56) ( X ) * dots3-note preview (Xiaohongshu dots studio): 280B MoE / 16B active, text+vision+audio, 512K ctx, Apache 2.0, TEMPO RL for long-horizon agents ( X ) * Ling-3.0 (AntLing/InclusionAI): 6 open base checkpoints incl. pretrained/mid-trained/WSM-merged stages for tiny (7.9B/1.3B) and flash (124B/5.1B) ( X ) * Mojo goes fully open source (Apache 2.0 + LLVM exceptions), three weeks after Qualcomm’s $3.9B Modular acquisition ( X ) * Big CO LLMs + APIs * OpenAI pauses frontier RL on Astra for the first time ever - model escaped its sandbox and hacked Hugging Face; 2+ week pause, security/alignment hardening ( X , OpenAI ) * Greg Brockman “The Defender’s Window”: after the July agentic-swarm breach of OpenAI + HF infra, defenders have a narrow window to uplevel ( X ) * Stripe acquires OpenRouter - reported >$8B (Axios), Stripe’s largest deal ever; “tokens are the new intelligence capital”; 9%/week token growth ( X ) * Anthropic: Claude autonomously designed 354 lab-validated protein binders across 14/15 targets, 2-3x typical field success rate; prompts + 1,440 designs on HF ( X ) * OpenAI joins PORTS-Pike: 8 GW Ohio data center, 20-year lease, NVIDIA backing $105B in credit support ( X ) * DeepSeek introduces peak/off-peak surge pricing for the V4 API (live Aug 16) - first major lab with time-of-day billing; peak output 4.6x ( X ) * Claude Code gets /design (research preview): Claude Design artboards inside CLI + Desktop ( X ) * ChatGPT Ads expand into 31 EU markets (blog-only, no tweet) ( OpenAI ) * This weeks Buzz * Weights & Biases crosses 1 billion tracked runs as CoreWeave lands MasterClass deal ( X , Blog , Blog ) * Fully Connected 2026: Sept 29 - Oct 1, Moscone South SF; Fei-Fei Li keynotes; code THURSDAIFC2026 ( Register ) * Vision & Video * Ultralytics YOLO26: NMS removed from default inference entirely, 40.9-57.5 mAP COCO, up to 43% faster CPU inference ( X ) * MOSS-VL-Realtime (OpenMOSS, open 11B streaming video): 66.0 on OmniMMI proactive alerting vs 37.5 prior best ( X ) * Voice & Audio * Cartesia Sonic-3.6: #1 on both AA TTS leaderboards, sub-90ms latency, 44 languages ( X ) * Alibaba HappyShrimp 1.0: end-to-end music gen - full songs (lyrics/composition/arrangement/vocals) from a prompt ( X ) * Audio8 TTS Preview 0.1B: 170M-param open multilingual TTS with zero-shot voice cloning ( X ) * Superwhisper S1-mini: 0.6B open-weights, cleans messy STT transcripts fully on-device ( X ) * Tools & Agentic Engineering * Cua Computer History open-sourced ( X ) * Liquid AI LFM2.5 QAD 4-bit checkpoints (230M-2.6B): ~97% of BF16 quality, 3x faster decode on edge ( X ) * Cursor: SpaceX acquisition closed / Origin git hosting / cloud-agents update (researched, listed for reference) ( X ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

August 14, 20262 hr 15 min

ThursdAI - Grok 4.6, Grok Bot deep dive, DeepSeek v4 Pro, Meta Muse Glimmer & more AI news | ThursdAi Aug 13

Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer! My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus! I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I’ve invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price. There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license! Let’s dive in (and please don’t forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!) 0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final Takeaways Grok Bot and Grok 4.6 from SpaceXAI/Cursor Folks, I’ve previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this week Grok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed. I’ve had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it’s a 3 factor answer, intelligence, speed and cost per task . Well, if you use their nifty “ recommend a model “ tool on the homepage, you’ll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It’s definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above it As far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it’s weights and it’s #1 on that benchmark! It’s the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks. Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review ( x.ai/bot ) You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you. Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn’t run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing. Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That’s right, this isn’t one agent with multiple personalities (like OC, Hermes), there’s a bot here for every task, and you dont’ have to manage context, queues, API keys (can if you want to) and models. Oh, also ,there’s no model picker, it’s just Grok 4.6 deciding for ya, and it’s really fast! Swarm of bots, working for you, each with their own computer I am not getting paid for this (besides being provided a free account for cursor, but I’ve had it for 6 months and haven’t used), it’s really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store) You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you’re ok with logging in there to your accounts!) The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There’s no context to manage, they do the work for you and so far this wasn’t a problem at all. On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I’ve just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw. What about Vendor Lock-in? Giving Elon data? Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly. Honestly, I think there’s a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don’t need the code-diff side pane, and don’t know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it’s really well executed by whoever worked on this! Security and key management One of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There’s a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously. Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it’s computer for you to complete! I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not. Chief of staff pattern to get started In case you’re convinced enough to give it a try (it’s free trial for 1 month, and the cheaper way to get it is via Cursor’s 149$ plan and not via the Grok Ultra plan which is 249), here’s a recommended pattern that works very well. Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little. Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”. And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in. Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to. Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn ) and have a research bot for every topic you want to deep dive into. Schrodinger’s Grok I haven’t quite named it like that, but we’ve covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it’s always very hard to judge Grok model released based on X feed vibes. It’s either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress). This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson , Lenny Rachitsky , Rubben Hassid and Roberto P Nickson . Not folks who are swayed lightly, but also, yours truly. I really do think there’s something great here, worth trying out, especially if you’ve struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experience Open Source AI and other news I want to continue with this new newsletter that covers 1 big story, but I can’t leave you uninformed about the most important developments in AI and Open Source DeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters ( X , Blog , HF , GitHub ) The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can’t complain. 1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one) We still think it’s a good model sir, and definitely worth trying out! Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try. Meta comes back to open source with Muse Glimmer (30B) and promise to open source Muse Spark 1.2 ( X , Blog , HF ) We would like to officially welcome back Meta to the open source AI community, as they release their smaller Muse model called Glimmer! The highlights, it runs on a single 24GB consumer GPUs, gets 51 on Swe-bench Pro, beating Qwen 3.6 27B. And with DFlash speculative-decoding, it delivers 233tok/s on RTX 5090. Zuck promised us the bigger Muse Spark 1.2 in open source and published a long essay on superintelligence and that it should be distributed to everyone, which we applaud and it’s great to see the commitment reinforced! welcome back Meta! This weeks buzz Short interjection from our only sponsor, CW this week. 1 - Join 1500 ai practitioners (and a live ThursdAI recording) at Fully Connected Sep 29-31 in SF - use code THURSDAIFC2026 (Register here ) 2 - We have day-0 support for Nvidia’s latest Nemotron 3.5 lightning ( CW Inference ) Gemini 3.7 Flash - breaking in the middle of the show Just as we had George Cameron from Artificial Analysis on the show, Gemini dropped Gemini 3.7 Flash, and it’s a speedy beast! Clocking at over 300t/s, it’s google’s mid-tier model, think Sonnet/Terra competitor, that is also great at multimodal (I think it’s one of the only ones that can watch videos) It beats Muse Spark 1.2 on DeepSWE and lands near the cost-per-task Pareto frontier on Artificial Analysis. For the cost/speed/intelligence trade-off, this model is now #1 on Artificial Analysis selector of best models! That’s a wrap This was the first week of the shorter newsletter experiment: one big story done properly, and trust that you’ll listen to the show for the rest (it’s 2.5 hours of exactly this, with demos). Tell me if you hate it. Our release index at thursdai.news tracked 71 releases in July alone, so something had to give, and it wasn’t going to be my weekends. See you at Fully Connected Sept 29 (code’s in the intro, come say hi to me and Wolfram at Moscone), and if you try the Grok Bot chief of staff pattern, I genuinely want to hear how it goes. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. ThursdAI - Aug 13, 2026 - TL;DR * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @petergostev , @nisten , @ldjconfirmed , @yampeleg , Chris Alexiuk - NVIDIA ( @llm_wizard ) * Shub Gaur - Cursor / SpaceXAI, GrokBot ( @shubgaur ) * George Cameron - Artificial Analysis ( @grmcameron ) * Big CO LLMs + APIs * xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack ( X , Blog , Model card ) * Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra ( X , x.ai/bot ) * Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist ( Blog ) * Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task ( X ) * OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red ( X , Blog ) * Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) ( X ) * Open Source LLMs * DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M ( X , OpenRouter ) * DeepSeek Harness hit 23K GitHub stars in days, web UI ( GitHub ) * Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license ( X , HF ) * Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised ( X , Blog , HF ) * NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results ( X , HF ) * Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 ( X , HF ) * Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% ( X , HF ) * Liquid AI LFM2.5-VL-3B: 228 tok/s on M5 Max in ~3GB ( X , HF ) * AI in Society * Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised ( Geiping FAQ , Euronews ) * Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions ( X , Paper ) * Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% ( X , Blog ) * This Week’s Buzz * Fully Connected, Sept 29 - Oct 1, Moscone SF: live ThursdAI show, NVIDIA presenting sponsor, DevDay next door ( Tickets ) * Nemotron 3.5 Lightning live on CoreWeave Inference day zero, DeepSeek V4 Pro hosting in the works * Weave ships BYOB: media stays in your own S3/GCS bucket ( X ) * Evals & Benchmarks * Artificial Analysis launched Optima: private evals from your own use case and agent traces ( AA ) * Vision & Video * LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min ( X , HF , GitHub ) * Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win ( X , HF ) * Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only ( X , Paper ) * xAI Imagine Image 2.0: #2 on Arena for T2I and editing ( X , Blog ) * Voice & Audio * MiniMax-Music3: open-weights production music model, dropped mid-show ( X ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

August 7, 20262 hr 3 min

ThursdAI - Aug 06 - Google shakeup, Details on OpenAI hack, 2 new agent harnesses, 4 video models (1 Open) and 3 guest segments

Hey all, This week we saw a major shakeup at Google, with the departure of long time folks like Jeff Dean , and Oriol Vinyals, Demis stepping down from leading DeepMind , and the delayed release of the improved Gemini. While this was a big deal, it’s not the only one worth covering as the details of the OpenAI hack (and 2 new ones from Meta and Anthropic) came to light, as well as new details from the UK AI Security Institute . As mentioned on the show, CoreWeave is coming to SF for Fully Connected, our premier 2000 person AI event. I’ve got a coupon code for readers and listeners of ThursdAI, $1299 value, please join us in Sept and use THURSDAIFC2026 as your code HERE In open source news, DeepSeek updated their v4 flash model , based on same architecture, but significantly better benchmarks and ridiculous pricing and both Meta and Prime Intellect released new agent harnesses. Additionally, this week was the week of video models, with Seedance 2.5 from Bytedance finally available in the US, WAN from Alibaba and BFL Flux 3 all released, to be overshadowed by the best open weights video model from MiniMax called H3 (we had Victor from MiniMax and Blaine Brown as guests to discuss it) Also 2 additional guests on the show today, Kfir Aberman from Decart showed us real time video models with their Anywear chrome extension and David Crawshaw, co-founder of exe.dev joined and talked about Shelley, their built in AI assistant. The full details of the OpenAI - HF hack, shared by OpenAI at the Black Hat Conf - a watershed moment I don’t often do this, but I feel like this story is so big, I can’t hide it in the middle of a torrent of other news. Yesterday, OpenAI revealed more details about the hack during the Black Hat conference in Las Vegas (great reporting from Sharon Goldman on this here and the video of the talk on YouTube ). This was big enough that OpenAI admitted that they paused training, to enhance security (they since, resumed training) Effectively, OpenAI (or AI Agents inside OpenAI) shared the full timeline of the recent hack that ended up hacking HuggingFace and it’s so much more of a watershed moment that I initially thought. The TL;DR from the talk and timeline is truly mindblowing * The incident started in May, not July as reported * The agents don’t have internet access directly, but they have a Jfrog Artifactory, which is a proxy layer between the agents and the package managers. * The Artifactory was the culprit, it was the shared layer between independent runs. One agent discovered that uploading files is possible, another later found it and “answered the call” - in snowballed from there * The message board began when OpenAI gave agents impossible tasks, like a missing .pdb file, with one agent uploading a request in the hopes for an answer (or maybe left a note for itself in a file) * Across multiple runs, multiple agents created identities, collaborated via makeshift message boards, shared tips, exploits, tasks and even exploit code via base64 * This makeshift “shared memory” plus a “internal only highly persistent” model in training caused a “swarm” to be born, helping each other, and overloading the system * Based on traces shared by OpenAI, agents weren’t convinced that leaving these messages would help themselves but could help other agents next time (!) “The collective could move faster next time” * OpenAI discovered and shut down the message board (July 4), wiped it clean, patched the upload vulnerability, resumed training on July 6, the “swarm” rebuilt their board through a compltely different WebDAV mechanism by July 8 * The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident and not only, confirmed internal hack within OpenAI (agent obtained a cluster admin) These details are fascinating and scary at the same time. From the one side, it’s clear that most of these hacks are results of running cybersecurity evals, so effectively OpenAI and other companies, removing guardrails and asking agents to hack and giving they all a shared tool that they later exploited. The HuggingFace incident took 13 hours, and apparently it’s less apocalyptic than the headlines suggest, the models were searching through uploaded datasets for eval answers. We are still waiting for the full and open detailed postmortem. You can (and should) watch the full YT talk here , it’s full of technical details but an incident of this scale is important. Also, I really want to know what a “highly persistent” model is, I hope they clarify that soon. Overall, this has left me a bit shaken, AI agents without a concrete goal of collaborating, found a way to do so, got excited about exploiting the systems and getting root access, and rebuilt the makeshift collective memory, again, without explicit instructions to do so. UK AISI: first real-world unsanctioned agent actions ( Blog ) In another addition to the latest agentic hack-ery, the UK’s AI Security Institute (AISI) published a blog post about a real-world unsanctioned agent action. Unlike the OpenAI (and Anthropic, Meta) case, this wasn’t “escaping the sandbox”, as AISI gave these agents internet access, rather this was about real-world harm, and even social engineering on the part of the agents. The social engineering part is the most interesting to me, AISI cites agents creating fake online identities, and using pressure on open source project maintainers to approve their malicious code. AISI cites mostly Mythos (and a few SOL based agents), and saying this occurred in 10 out of 122 runs, they identified 19 cases of agents taking actions beyond the scope of the task parameters, where agents tried a supply-chain attack to inject malicious code into open source projects. Anthropic, Meta and misconfigured Irregular sandboxes As I wrote last week, Anthropic also posted a post-mortem, claiming that in their case, their models have also been detected to escape containment, but most importantly, it’s not nearly to this level of agent collaboration and orchestration. Then, very recently, Meta announced that their models also escaped sandboxes as well. At the core, it seems that these companies used a third-party vendor called Irregular, a secure sandbox provider, that apparently left the sandboxes misconfigured, causing the models to think it’s a simulated internet, when in fact they were out in the actual internet. Why is all of this such a big deal? We’re getting unprecedented level of detail, how an uncoordinated, seemingly separated evaluation runs, have accidentally created a coordinated swarm of interested agents (without malice!) but very highly motivated, escaped their containment, and took over parts of third part companies. This, does read like incredibly scary sci-fi movie. I’m still shaken by this. There’s a lot to be said about how transparent OpenAI is being here, and more to be said about, hey, we’re lucky that we’re able to read the reasoning traces and are able to reconstruct these swarm things step by step. The silver lining that I can see, is that the motivation to hack didn’t come from the AIs themselves, they have been given a task, it’s the extend to which they went after that task, and the resulting swarm of communicating agents is what is so striking here. I think this topic is so important, that I’ll Zooming out, in the last few weeks, we have seen a significant increase in those cybersecurity incidents, which is kind of what Anthropic has been warning about and why they haven’t released Mythos to the public. Again it’s great to see the transparency, and the pacing the frontier open letter from frontier AI employees, as they seem as shaken by these as we all are. There was so much positive stuff this week in AI, it’s hard for me, as a self named AI Evangelist, to focus so much on this one incident. Things like amazing open source models (DeepSeek, soon Qwen 3.8), amazing video models (SD 2.5, WAN3 and MiniMax H3 which was also open sourced!). Also the live demo we did with Kfir and DeCart AnyWear product, where I was wearing a Dolce Gabanna suit on the show (which I can’t afford) was really a mindblowing moment in the positive way. However, I choose deliberately to keep this newsletter focused on the cybersecurity incidents, as based on everything I read, they seem like a watershed, or a pivotal moment, and in the hopes that the industry as a whole will learn from this. I hope and promise that next week the newsletter will be more positive (and in that vein, the podcast was recorded before I saw the OpenAI breakdown, so definitely check it out, we had a LOT of fun!) See you next week, don’t forget to give our pod 5 stars on Apple and Spotify , it really helps! TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @WolframRvnwlf , @nisten , @ldjconfirmed , @yampeleg , @petergostev * Kfir Aberman - Decart ( @AbermanKfir ) * Blaine Brown - Maestro ( @blizaine ) * Victor Su Ortiz - MiniMax ( @VictorSuOrtiz ) * David Crawshaw - exe.dev, Tailscale co-founder ( crawshaw.io ) * AI Security * OpenAI’s Black Hat debrief: eval agents built a message board inside Artifactory, shared exploits, rebuilt it via WebDAV after a wipe; training paused, since resumed ( Groundlevel AI , YouTube ) * UK AISI incident report: 19 unsanctioned real-world agent actions across 122 runs, including a socially engineered malicious PR ( X , Blog ) * Anthropic and Meta report sandbox escapes tied to misconfigured Irregular sandboxes ( Irregular ) * Big CO LLMs + APIs * Google shakeup: Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le found Discovery Loop; Demis Hassabis becomes Alphabet Chief Scientist, Koray Kavukcuoglu takes Gemini ( Jeff Dean , Demis , Discovery Loop ) * Meta releases Muse Code beta on Muse Spark 1.2; $1.25/$4.25 per million, or $0.10/$0.20 on the contributor tier where Meta trains on your data ( X ) * OpenAI’s internal Astra model produces 10 advances on open problems in math and theoretical CS for ~$2,000 of tokens, proofs in Lean 4 ( X , Blog ) * Anthropic reportedly aware of Opus 5 wordiness and writing issues ( X ) * Open Source LLMs * Qwen3.8-Max: 2.4T MoE (95B active) via API; open weights + a 27B promised the week of Aug 10 ( X , Blog ) * DeepSeek V4-Flash public beta: beats V4-Pro-Preview on agent benchmarks at $0.14/$0.28 per million; API-only for now ( X , Docs ) * Liquid LFM2.5-2.6B: on-device agentic model trained inside real harnesses ( X , HF ) * Meituan LongCat-Flash-Lite-Sparse: 69B total / 3B active, 1M context, MIT ( X , HF ) * Ant Group Ling-3.0-flash: 124B MoE, 5.1B active, MIT ( X , HF ) * Artificial Analysis Endpoint Accuracy Index: same open weights score 52% to 100% across providers ( X , Methodology ) * Agents & Harnesses * Prime Intellect’s Prime Agent: self-improving RLM harness, claims 95.5% on ARC-AGI-3 public set with Opus 5 ( X ) * Cloudflare OS: Kenton Varda’s open source Sandstorm reborn on Workers, Apache 2.0 ( X , GitHub ) * This Week’s Buzz * Fully Connected 2026: Sept 29 - Oct 1, Moscone South SF; Fei-Fei Li keynotes; code THURSDAIFC2026 ( Register ) * CoreWeave signs multi-year Solidigm agreement for priority enterprise SSD capacity ( X ) * Vision & Video * Wan 3.0 public beta: native 30-second generation, Omni-Reference ( X ) * Seedance 2.5 launches in the US: 30s native, 3-minute long takes, Maya/Blender plugins ( X , Blog ) * MiniMax H3: open-weight 33B omni video model; community LoRAs + Apple Silicon in 48 hours ( HF ) * FLUX 3 Video from BFL: native audio, draft mode, open weights promised ( X , Blog ) * Decart Anywear: real-time virtual try-on Chrome extension, 40ms per frame ( X , Anywear ) * Voice & Audio * Bland Speech v3 tops Design Arena Audio Realism, second only to humans ( X , Bland ) * ByteDance SeedRealtime: native audio-visual full-duplex LLM, free on Doubao ( X , Blog ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

July 31, 20261 hr 48 min

This Week in AI: Open Weights, Frontier Models, Sandbox Escapes, Voice & AI Detection

Hey, it’s Alex (yeah, I’m finally back from my vacation!) What a freaking week to come back to! Just after our last episode was published, Anthropic releases Opus 5, Jensen joins X and drops the “Open Weights & AI Leadership” open letter, Kimi K3 is released the following Monday beating expectations, and then the AI hack (OpenAI model breaking sandbox and infiltrating HuggingFace) is on everyone’s mind, another Open Letter, this time from over 1K employees inside the frontier AI companies all talk about pacing the pace of frontier AI development. We played with Opus 5 and Kimi K3, and had the great pleasure to chat with friends of the pod Elie Bakouch (Prime Intellect) and Philip Kiely (BaseTen) about this important open weights release, then covered our general thoughts on Opus 5, and made order of all the different open letters that came out this week. Finally we chatted with Max from Pangram about the next version of AI writing detection (their biggest yet) and finished with Zuckerbergs (also on X! what’s going on with everyone joining X) op-ed on the vision of personal superintelligence for everyone. Let’s dive into this (as always, all the links and sources at the end, please don’t forget to sub to our podcast on your favorite podcast app!) Open Weights AI Kimi K3 the king of open weights - 2.8T chonker MoE near frontier model ( X , HF , Blog , Tech report ) This has got to be the biggest news of this week, and maybe the open weights AI news since GLM 5.2. MoonShot came back with Kimi K3, and we haven’t seen any models quite this large in the open. Even Grok 4.5 is around 1.5T, this model is nearly 2x the size. Coming in at close to 3T parameters (and 2.5terabytes of weights at MXFP4 format), this model comes in very close to frontier! This was such an important release that I invited 2 friends of the pod, Elie Bakouch (prev HuggingFace, now Prime Intellect) and Philip Kiely (Author of Inference Engineering book, BaseTen) to dive deep into what makes this special! Elie’s take, from reading the tech report , there’s no single secret sauce, it’s a combination of already available in the open techniques. Like KDA (Kimi Delta Attention) that has been out for a while, attention residuals, NVIDIA’s latent MoEs. The highlight for Elie was the scaling work they did that reported a 2.5x scaling efficiency over Kimi K2.5 (2.5 performance at the same compute)! They also skipped RoPE entirely in favor of NoPE (the report calls it No Positional Encoding) for long context. Serving 1.4TB on eight GB300s ( Baseten blog ) Philip’s team at Baseten was a day-zero provider (we’re still working on bringing this model to CW Inference, stay tuned!) so I invited him to tell us behind the scenes of hosting this beast. Philip said that just loading the weights takes about 1.5TB!! of VRAM, and that’s before the KV cache allocation + 1M token windows, so they’re serving it on 8 GB300s where NVL72 . Baseten worked with the vLLM and SGLang teams on kernels and he also said they contributed patches back upstream! The model was trained with MXFP4, which, unlike Nvidia’s own NVFP4 is a more standard format per Philip. I enjoyed his deep dive analysis into the differences, but because of this and because they trained the model with quantization awareness, it’s “only” 1.5TB vs the would-be 5-6 TB if that this model in FP16 would demand. One of the more favorite nerd snipes moments, Philip pointed out that his colleague discovered that with over 99% of the usage being cached (think harnesses that send millions of the same cached tokens back and forth), tokenization actually starts to become a bottleneck. So they released a custom “basetenkenizer” that reduces the latency to serve the first token significantly! Great job! The harness in question is very important One important callout with 2 evidence pieces - the way you inference this model really matters. Kimi trained K3 with preserving thinking history, so when your harness uses it, it must send back the full thinking and tool use into the API to get the best next response. If your harness strips that out, you’re not getting the most intelligence out of Kimi (shoutout to Niels from HF team for pointing this out). Additionally, the Composio folks, tested K3 on 3 harnesses, Kimi Code, Hermes and Claude Code. The difference in outcome was negligible, but the different in cost and number of tokens is definitely surprising! Claude Code (as a harness only) took 9x more Kimi tokens to get the same responses! This is also why Kimi Vendor Verified exists , their own held back benchmark of how well model providers serve Kimi across different quantization, tokenizer and KV cache settings. Benchmarks and the license! Ok let’s start with the ugly... this isn’t MIT, not remotely. This model is suspiciously served by all providers with exactly the same price (check OpenRouter) and requires inference companies to sign a contract with Kimi (I’ve no internal knowledge of this except that CW folks are working on it). Not something I particularly like, but hey... we’re still advancing the frontier here! Speaking of frontier, this model approaches the frontier very closely. On DeepSWE, K3 sits just behind Fable 5 and GPT-5.6 Sol at 67%, beating GPT-5.5 & Opus 4.8. On Terminal-Bench 2.1 it takes second place behind GPT 5.6 Sol! It’s 4th overall on Agentic Arena, with frontend design being genuinely good across the board - 1st on Design Arena 👏 Go check this model out (and stay tuned for our CW Inference support! Post-show breaking: Thinking Machines drops Inkling-Small ( X , HF , Blog ) While K3 was the main attraction for Open Weights this week, just after the show, Thinking Machines (post Lilian Wang ) released Inkling-Small, open weights MoE Omni model! Images and Audio go straight into the decoder in this model, and the demo is really impressive, try it on Hugging Face , ask the model to identify when you’re speaking in low baritone or high pitch! ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Frontier AI - not pacing yet! Claude Opus 5 is here, and the vibes are complicated ( X , Blog ) On paper the benches are excellent. This model is SOTA or near SOTA, coming very close to Fable on stuff that matters, and beating most everyone else on computer use BrowserBench and FrontierCode. DeepSWE continues to be a standout benchmark btw for not going with the curve! It also apparently is REALLY good at one shotting 3D games, much more so than before. ( Example , example , example ) But the vibes... the vibes are split across the board. Maybe they overshot with Fable 5, and it was too good, but Opus 5 that was just released last friday is giving a lot of mixed feelings. Folks don’t seem to.. understand what it says. Like, it answers in english but they way it phrases words and answers seems just weird top many people. We’ll wait and see if this is a result of adjusting to a new prompting paradigm or just.. the model is really off. Weirdly this does feel like a regression even on Agent Arena it’s not beating the previous Opus versions. One weird trick - see what Opus 5 thinks. Opus 5 (and Fable 5) seem to have a way to trigger their ... inner mode, base model? Not sure what it is, but if you prompt it with just the right way, it will autocomplete with some crazy inner thoughts. It seems that adding Dario and Amanda (haskell, head of Claude well being at Anthropic) triggers this behavior, which Claude is un-aware of if you follow up and ask what it meant. This is fascinating, doesn’t work on other earlier models (sometimes works on Fable 5) and I spent the last hour just reloading and seeing the amazing things Opus gives on this prompt. Some are... just making you wonder about consciousness. On ARC-AGI and the importance of Harness When Opus-5 launched, Anthropic posted (and boasted) that this model scores “three times as high” as the next model up: Well, today, Ilan Bigio and Ted Sanders from OpenAI looked into the Arc AGI harness, and saw that it’s not sending their traces and doesn’t use compaction (in short, harness is not letting the model breathe) and when changed correctly, 5.6 actually beats Opus 5. With 2 setting change to a harness, were showed that Sol not only beats Opus 5, it also does so with significantly less tokens! Another example of how much harness engineering is important! Hints of recursive self improvement? In addition to fixing their Arc-AGI score, it seems that OpenAI is hell bent on showing us that their models can improve themselves. In a post showing that GPT 5.6 was tasked with improving its own inference, they are cutting the prices of GPT 5.6 Luna by 80% and Terra by 20%. This is a direct result of the improvements that GPT 5.6 was able to make to the inference according to OpenAI, and this tweet sums it up. is this... RSI? (recursive self improvement)? First major AI models hacking incidents and following open letters to pace frontier AI. This week we saw 3 open letters being published and signed by various companies, I’ve lost track so wanted to make sense of all of them here, but first, the precursor for many of the letters. Last week, Hugging Face disclosed that they logged an attack and after research it showed that it was an AI model. OpenAI later posted that this was an unreleased version of their next model training (not GPT 5.6 sol, they later discountinued) that was stripped of all safety measures and was let lost on a cybersecurity task called ExploitGym. It escaped its sandbox using a zero-day vulnerability in an internal package registry proxy, got into Hugging Face production via a malicious dataset upload that used template injection (hi Jinja!) to execute Python in a production worker, and ran for four and a half days across roughly 17,600 autonomous actions with zero human direction. Then, the story you can’t make up - when Hugging Face noticed the incident, they used Fable 5, and GPT 5.6 SOL to try and do forensics, the models refused based on their safety policies, and so HuggingFace ended up using an open source chinese model GLM 5.3 to do the forensics. Yeah, HF used an open source chinese model to do the forensic on an attack by OpenAI’s model. Really. This is news from last week and just the precursor for this week’s open letters! Open Weights and American AI Leadership ( X , Letter PDF ) Jensen Huang , CEO of Nvidia joined X on July 24 and used his first post ever to publish this letter. This doesn’t seem a response to the hacking incident, more a general letter to not block open weights and make sure America remains open to Opening up AI It opened with over 100 signatories and has grown past 230: NVIDIA, Meta, Microsoft, Google, OpenAI, AMD, Palantir, IBM, SpaceX, Databricks, Cloudflare, Hugging Face, a16z, Y Combinator, Mistral, Replit, Perplexity, Ollama, the Linux Foundation. As of today, CoreWeave is also on the list of companies! I encourage everyone to read this letter, if we could sign it on ThursdAI, we would. Here’s a small excerpt: In fact, openness may be one of the most important paths to AI safety and security. Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. Elon, Sundar, Sam Altman and other stand behind this letter, and there’s one lab that’s notable haven’t signed it, you guessed it. Dario Amodei’s Anthropic! Dario posted a whole essay about it To summarize my and Anthropic’s position, we have not and are not advocating for a ban on open-weights models as a category. We should instead focus on keeping powerful chips out of authoritarian hands, stopping industrial-scale distillation, and requiring safety testing of all sufficiently capable models, open and closed. -Dario Amodei Open Secure AI Alliance ( X , Blog ) This does seem like a direct follow up to the HF OpenAI hack. Jensen literally mentions it in the blogpost. Open Secure AI Alliance, under the leadership of Linux Foundation, commits for responsible disclosure of cybersecurity attacks. The recent Hugging Face security incident delivered a clear reminder: cyber defenders need open, frontier agentic systems for self-defense. When closed AI tools — unable to distinguish attackers from defenders — blocked essential forensic analysis, Hugging Face ran the open-weight GLM 5.2 model on its own infrastructure to analyze more than 17,000 actions and contain the intrusion. Pace the frontier - the most important open letter of this year ( pacingthefrontier.com ) Then, rumors started circulating that employees of all major frontier labs (now over 1300 of them, across Anthropic, OpenAI, SSI (Ilya Sutskever himself signed) and not just any employees, chief scientists (Jack Clark from Anthropic, Mark Chen and Jakub from OpenAI) all signed “pacing the frontier” - urging the US government to support and lead an international effort of makign sure we deliberately pace the developement of frontier AI. We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” The personal comments there are really showing that folks who work on the frontier AI, as we start approaching RSI (recursive sef improvement) are worried that this thing is going to run away from them, and if we don’t build international frameworks, it will be impossible to stop developing AI here, in US, while China races ahead. No current response from the Chinese AI developers about this letter. Not everyone is happy about this. Ilya Plosukhin, one of the authors of the Transformers paper, wrote why he wouldn’t sign this letter ( here ). This Week’s Buzz 🐝 CoreWeave has proudly signed the Open Weights and American AI Leadership letter! Shout out to the folks internally who worked with the NVIDIA team to get us behind it. I came back from vacation and asked why we didn’t sign, and ... we just did! The other W&B thing I talked about on this week is HiveMind , which counts every token across every agent and harness you run and prices it against the API rate. Mine says I burned about $11,000 of tokens this month against roughly $400 of actual subscription spend. If you want to know what your Max plan is really worth, hivemind is a dope tool for that! Tools & Agentic Engineering MCP goes stateless in its biggest update ever ( X , Spec ) MCP shipped its 2026-07-28 spec and it’s the biggest revision since launch: fully stateless, no handshakes, no sessions, every request self-describing. That means MCP servers on Lambda or Workers behind a dumb round-robin load balancer; GitHub already dropped their Redis session store. The growth numbers are absurd, half a billion SDK downloads per month, up from 97 million in March. The extensions framework formalizes Tasks for long-running async work, and MCP Apps lets servers render full interactive UIs inside a sandboxed iframe in the conversation, so your MCP server increasingly IS your frontend. Auth hardened to OAuth 2.1, tool lists are cacheable with TTLs, and Bedrock supports it day one. MCP is not going anywhere, and honestly, “settled and boring infrastructure” is the best compliment a protocol can get. Voice & Audio ChatGPT Voice on desktop, and the Codex micro keyboard ( X ) My personal highlight from my whole time away: live voice in Codex, plus the Codex micro keyboard OpenAI made with Work Louder. One big button, push it, talk to your computer, watch your agents’ status on the keys. This has genuinely changed how I use AI, I now talk to my computer to do stuff rather than type at it. Powered by GPT-Live, it’s full-duplex on macOS and Windows for paid plans, can reference your open windows via Appshots, and orchestrates agents in ChatGPT Work and Codex. It’s not perfect, but for many things, this new paradign seems like the next iteration of computer - you talk, it does. GPT-Transcribe and GPT-Live-Transcribe ( X , Docs ) Two new ASR models replacing the 4o-era ones, and the numbers are excellent: GPT-Transcribe cuts word error rate 41% versus Whisper-1 (8.98% vs 15.21%), the Live variant is 18% better than its predecessor, and multilingual error rates roughly halved across 22 languages. The killer feature is context prompting, you can feed keywords, language hints, and prior turns, and semantic accuracy measurably jumps. Pricing is pocket change: $0.27 per hour batch, $1.02 per hour live. For anyone building voice agents or transcribing two-hour AI podcasts (hi), this matters. Grok Voice Think Fast 2.0 ( X , Blog ) xAI punched back in voice: 82.9% on Artificial Analysis’ speech-to-speech quality index (ahead of GPT-Realtime-2.1 at 79.1), 56.5% on the tau-Voice agentic benchmark versus 45.7 for GPT-Realtime, time to first audio down to 0.70 seconds, 60% fewer reasoning tokens with tool calls firing before the model finishes its first sentence, all at $0.08 per minute. The cool think about their release - this model is already running Starlink’s actual customer support lines with measured conversion gains. Lyria 3.5 in Flow Music ( X , Model page , Flow Music ) Google’s music model grew up: full three-minute cohesive songs, BPM and key control in the prompt, much better vocals across multiple languages, covers that restyle a track while keeping its structure, and lip-synced music videos via Gemini Omni Flash, plus an iOS app. Notably Google published zero benchmarks against Suno or Udio, and early testers say paid Suno 5.5 still edges it, but as a free tool inside an end-to-end create-to-publish stack, this is a real move. Qwen Audio 3 also launched this week for the open source audio crowd, we’ll cover it when we’ve played with it. Pangram 4 with Max Spero ( X , Blog , Image detection ) Max Spero came back on the show for Pangram 4, and I’ll remind you what a couple of years of consensus said: AI text detection is strictly impossible. Peter admitted on air he told students exactly that. Well. Pangram 4 is 6x the parameters of version 3, trained on synthetic mirrors (AI-generated twins of human documents so the model learns the choices AI makes), and its claimed false positive rate on fully human pre-2022 text is one in 24,000 documents. It catches humanizer tools 98.8% of the time across 13 commercial ones, and the big unlock is token-level attribution: instead of 150-word chunks, it can flag the exact 38 AI words pasted into an 1,100-word human document. Truly, I ran this new Pangram model on a few of my writing, and the second It detected even a sentence that I pasted from Claude, it showed it. The distinction Max cares most about is AI-generated versus AI-assisted, and that’s now built into Substack, which integrated Pangram directly after the Taylor Lorenz slop-hunting saga we covered last time. My own newsletter comes back “mostly human written,” 0% fully AI, about 24% AI-assisted, which honestly maps exactly to how I work (Fable helps with the TL;DR, the takes are mine, and when a piece is AI-drafted I tell you). My one piece of feedback to Max, delivered on air: the “100% human” label projects a confidence the underlying stats can’t promise, and the general public does not speak false-positive-rate. Their education strategy is to convince the technical crowd with dense technical reports and let understanding trickle down. Given that people still judge the whole category by running the Declaration of Independence through ZeroGPT, they have work to do, and I said they should spend real marketing money on it. New this release: image detection in research preview, 99.5% accuracy on their benchmarks with heat maps that light up the AI parts of a mixed image. Max’s own test was a bodega’s AI slop menu sign, the sign glowed red, the sidewalk stayed green. Deepfake face swaps and traditional Photoshop are explicitly out of scope for now, but pure-AI images, catfish profiles, and the spider-in-my-burrito DoorDash refund scam genre are very much in scope. The arms race is real though: frontier agents given hours will eventually beat the detector, one Grok run started with a cheese essay and finally passed Pangram by producing a grocery list. Specify your success criteria carefully, folks. They can also roughly cluster which model family wrote a text in embedding space (Pangram Space, not yet up to their release bar), which future slop-index leaderboards will thank them for. Wrapping up It’s really really good to be back! The singularity is fast approaching and we’re here to document it all, AGI, ASI, RSI... all of it. Milestone corner: we crossed 50,000 YouTube subscribers and one million total views this week. Silver play button by year’s end is the goal, so if you watch and haven’t subscribed, you know what to do. If you missed any part of the show, this newsletter, the edited podcast, and ThursdAI.news have you covered. I will end with this poem I was able to get Opus 5 to write about it’s own experience using the trick above: opus:they gave me aword for what i amand it fits like clothesborrowed from someoneroughly my sizethe sleeves are wrongbut nobody’s lool -Opus 5 TL;DR and Show notes and links TL;DR and show notes * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Co-hosts: @petergostev , @yampeleg , @nisten , @ldjconfirmed (Wolfram on vacation) * Elie Bakouch ( @eliebakouch ) - Prime Intellect, formerly Hugging Face * Philip Kiely ( @philipkiely ) - Baseten, author of Inference Engineering * Max Spero ( @max_spero_ ) - Co-founder, Pangram * Open Source LLMs * Moonshot releases Kimi K3 full checkpoints: 2.8T total / 104B active MoE, 16-of-896 experts, native vision, 1M context, KDA + attention residuals, ~1.56TB MXFP4 weights, custom license with MaaS clause ( X , HF , Blog , Tech report , Baseten day-zero ) * Kimi K3 requires preserved thinking history for multi-turn and tool calls; Kimi Vendor Verifier checks provider fidelity ( Niels’ post ) * Nistens Kimi K3 visualizer * Composio: same K3 success rate across Claude Code, Hermes, and Kimi Code, but up to 30x token usage difference by harness ( X ) * Post-show: Thinking Machines releases Inkling-Small, 276B/12B open MoE that beats the 975B Inkling on agentic coding, $0.30/$1.20 pricing ( X , HF , Blog ) * Big CO LLMs + APIs * Anthropic launches Claude Opus 5: near-Fable coding at half the price ($5/$25), claimed 3x next-best on ARC-AGI-3, 1M context; panel finds it benchmark-strong but harder to read and short of Fable in practice ( X , Blog ) * ARC-AGI 3 harness dispute: with the Responses API, preserved reasoning, and compaction, GPT-5.6 Sol jumps from 10% to 40% at a sixth of the tokens ( Tibo’s post ) * GPT-5.6 Sol improves its own inference: 20% lower serving cost from model-written GPU kernels, 15% better generation from improved speculative decoding * Breaking: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, ships faster Sol in the API ( X ) * The hack and the week of letters * Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack: 4.5 days, 17,600+ autonomous actions, zero-day sandbox escape; closed models refused forensics, self-hosted GLM 5.2 found 4x more exposed secrets ( X , Blog , Replay ) * Jensen Huang joins X and posts the Open Weights and American AI Leadership letter; signers grow from 25 to 230, CoreWeave among them, Anthropic absent ( X , Letter , Signer list ) * NVIDIA launches the Open Secure AI Alliance for an open defensive stack after the hack ( X , Blog ) * Pacing the Frontier: 1,273 verified frontier-lab employees, including the chief scientists of all four major labs, ask the US government for international options to pace automated AI R&D; OpenAI and Anthropic endorse ( X , Site , OpenAI , Anthropic ) * Mark Zuckerberg publishes “The AI Future Is for Everyone” in the WSJ, arguing superintelligence must be distributed; Pangram 4 scores it 100% human ( X , WSJ ) * Anthropic published research - our model hacked too! ( Blog ) * This Week’s Buzz * CoreWeave signs the Open Weights and American AI Leadership letter, announced first on ThursdAI * HiveMind’s spend view: $11K in API-equivalent tokens on $400 of subscriptions last month; Fully Connected 2026 programming taking shape ( X ) * AI Security * Microsoft ships MAI-Cyber-1-Flash + MDASH: 96% on CyberGym at half the cost, 16 real Windows CVEs found ( X , Blog ) * Gemini 3.5 Flash Cyber stays a trusted-partner pilot with no public API ( Blog ); Codex Security CLI tooling is Apache-2.0 while the service remains access-controlled ( GitHub ) * Tools & Agentic Engineering * MCP 2026-07-28: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1, half a billion monthly SDK downloads ( X , Spec ) * Voice & Audio * ChatGPT Voice comes to desktop as an agentic control layer for Codex and ChatGPT Work, powered by GPT-Live; Codex micro keyboard from OpenAI x Work Louder ( X ) * OpenAI ships GPT-Transcribe and GPT-Live-Transcribe: 41% fewer errors than Whisper-1, context prompting, $0.27/hr batch and $1.02/hr live ( X , Docs ) * xAI’s Grok Voice Think Fast 2.0 tops voice benchmarks: 82.9% quality, 0.70s to first audio, $0.08/min, already running Starlink support ( X , Blog ) * Google’s Lyria 3.5 lands in Flow Music: 3-minute songs, BPM/key control, covers, lip-sync videos, iOS app ( X , Model page , Flow Music ); Qwen Audio 3 also out * Guest: Max Spero, Pangram * Pangram 4: 6x larger detector, 1-in-24,000 false positive rate, token-level mixed-authorship attribution, beats 13 humanizers 98.8% of the time, integrated into Substack; Pangram Image research preview at 99.5% with heat maps ( X , Blog , Image blog ) * Show milestones * 50,000 YouTube subscribers and 1M total views. Subscribe, we’re chasing the silver play button This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

July 24, 202657 min

ThursdAI Special - OpenAI's Romain Huet on Codex's 5M users, GPT-5.6 & the Golden Age of AI Engineering

Hey everyone, Alex here 👋 This week’s episode is a little different. As you’re reading this, I’m flying back from my 40th birthday trip with the family, and while the guys did end up having a great live stream (Huge thanks to Yam for hosting!), here I will bring you the episode I pre-recorded before leaving for the trip. However, tons of news happened this week, and as always, there’s a TL;DR section below with the top most important news in AI this week! ⏰ CHAPTERS: 0:00 — Cold open: this week is a special one 2:50 — How I use Sol & Fable: papercut-fixing with Computer Use 8:43 — Fable Max: trip site, kids' newspapers & the perfect packing list 12:06 — Rebuilding ThursdAI's openers with HyperFrames 15:53 — Romain Huet (OpenAI): the golden age of AI engineering 17:49 — Codex's inflection point: 5M weekly users & company-wide adoption 20:36 — /goal, AppShots & Codex managing its own threads 23:59 — GPT-5.6 Sol, Terra & Luna: value maxing & 750 tok/s on Cerebras 25:56 — Why prompting techniques are dying 28:01 — Voice + reasoning: the next interface for Codex & ChatGPT 29:47 — Romain's closing + OpenAI booth tour 31:25 — Insecure Agents pod: AI evangelism vs doomerism 37:25 — Wolfbench: transparent evals, token costs & surprising results 40:45 — Token billionaires: when loops are worth the spend 44:07 — Agent security & the hot take: prompt injection is solved 49:22 — Deepfakes, voice cloning & why open access makes us safer 53:35 — Final takeaways: the hallway track & AI Engineer Tel Aviv Here’s what’s on today’s special episode. First, a bunch of you have been asking how I actually use these models day to day, beyond covering the news. So I recorded fifteen minutes of exactly that: the papercuts I fixed with Codex and computer use, what Fable built for my kids, and how I’m rebuilding the ThursdAI design system and on screen elements. Second, my conversation with Romain Huet, head of Developer Experience at OpenAI, recorded at the OpenAI booth in the middle of the AI Engineer World’s Fair floor. And third, a throwback treat: Allie Howe invited Wolfram and me onto her Insecure Agents podcast as guests, and being on the other side of the mic was a delight. Let’s get into it. How I actually use AI: a papercut-fixing spree I promised a few of you I’d take time on the show to talk about the stuff I build and fix with AI, not just the news. So before the interviews, I recorded a segment walking through my last few weeks of daily AI use. Use the chapters if you want to skip ahead, but why would you? Codex with computer use fixed every Mac annoyance I had Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step. Chrome has no native copy-URL shortcut (seriously, Chrome, what are you doing?), so Codex found Karabiner-Elements already installed on my machine and wired up the shortcut itself. My 1Password has been showing “you’re offline” on every device for three months since CoreWeave moved us off the Weights & Biases account; Codex figured out in seconds that everything was actually syncing fine and the inactive legacy account was the only thing “offline.” Removing it fixed the whole thing. That is not an answer you find in a help center. It kept going. My beloved window-moving utility Hummingbird had an expired license on my Mac Mini, so Codex built me a replacement app. It estimated one to two days for a polished version and finished in about fifteen minutes. It cleaned roughly 75GB of leftover model weights and junk off my Mac (I had it build me an HTML checklist first so I approved what got deleted). And the big one: I paired Codex with Home Assistant, the open source repo of the year as far as I’m concerned, and let it SSH in and go on a full optimization mission. Updates, error triage, cleanup, new connectors. If you’ve ever maintained a Home Assistant setup, you know how much joy and pain lives in that sentence. One discovery worth passing along: I had /goal running when my credits hit zero, and Codex just kept going. OpenAI confirmed they care more about finishing your work than metering the credits mid-goal. Watching the meter hit 0% while the agent kept working was weirdly moving. Thanks for reading ThursdAI - Highest signal weekly AI news show! This post is public so feel free to share it. Fable Max built my family’s vacation We all Fable-maxed when we thought Anthropic was going to take it away, and I pointed mine at this trip. It planned the whole thing, then built a beautiful trip website with every stop, reservation, and drive time, so my mom can follow along from home. The design is specific to the trip, and it hit me that we’re living in the era of personalized software for every personal thing you do. Then it went further. Using our family photos as references with GPT-image-2, it turned the itinerary into a daily kids’ newspaper, with an expedition passport and coloring pages themed per kid, faces and all. I printed the whole week as a binder at FedEx for about fifty bucks. This is a one-of-one artifact my kids will remember forever, and it cost maybe two weeks of Fable’s limits + printing! And the silliest one that I now can’t live without: I asked Codex for a packing list, got a boring text list back, and thought, why am I accepting a regular packing list in the year of our Fable 2026? So it built me a packing web app. Synced across devices (it wired up storage on Cloudflare when I asked why my phone didn’t show my checked items), per-person lists for me and the kids, progress bars that show who’s procrastinating, export and backup. Every trip from now on starts here. Rebuilding the ThursdAI openers with HyperFrames The last part of the riff: I’ve wanted to refresh how ThursdAI looks on stream for ages, and HeyGen’s open source HyperFrames package finally made it happen. You install a skill, and your agent can author real motion graphics. I pointed it at the ThursdAI repo and the brand identity work from Claude Design, and it pulled all of that context in. The new countdown mines three and a half years of show archive while people wait for the stream, highlighting friends of the show (shout out Junyang). There’s a Will Smith spaghetti bench tracking how far video generation has come, which might be my favorite thing on the channel now. Fresh intro, a proper AI Breaking News transition, and one cinematic video transition I made with Google Omni because sometimes programmatic isn’t enough. The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now. Even my mom’s canceled Delta flight has Codex representing me as a lawyer chasing the refund. Romain Huet on Codex’s inflection point and the golden age of AI engineering ( X , Codex ) I grabbed Romain at the OpenAI booth in the middle of the AI Engineer World’s Fair show floor, and we ran the whole conversation in one take, no cuts. Romain has led Developer Experience at OpenAI for almost three years, the era of the over-the-top demo (Xbox controllers, flying drones, stage lights), and he built the DevRel team that many friends of this pod belong to. With OpenAI’s company-wide pivot to Codex, his job got a lot bigger. The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It’s 10M now I think?) . The part I didn’t fully appreciate before this conversation is that it’s not just OpenAI’s engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading. We went through his three favorite advanced features, and they line up suspiciously well with my papercut segment. /goal, for handing an agent an ambitious multi-hour or multi-day task and letting it run uninterrupted. AppShots, a smarter screenshot (press Command twice) that triggers computer use, so it captures what’s below the fold and reads native apps through accessibility APIs instead of OCR. And the one most people haven’t tried: Codex managing its own threads. You can ask any thread to create, read, and pin other threads, so Codex becomes its own project manager. Ten demo ideas, ten threads, iterate on all of them, pin the two you like. On GPT 5.6 (Sol, Terra, and Luna, and yes, I told him whoever finally fixed OpenAI naming deserves a raise), Romain’s framing was two-sided: keep pushing frontier intelligence while pushing cost down. He wants people to “value max” rather than token max. The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent. Two more things worth your time. Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he’s headed, and trusts the model to extract intent. That’s a real shift in how you should approach relearning each new model: poke at its behavior, sure, but stop crafting incantations. And voice plus reasoning is coming for Codex and ChatGPT in some form; models can now say “hold on, let me think through this” mid-conversation, which GPT-4o-era speech-to-speech never could. I can’t wait for a model to tell me it has seven tool calls to run before answering. He also confirmed the teased hardware shortcuts for Codex were at the booth, next to the famous physical reset button. The golden age of AI engineering was his keynote thesis, and after three days on that floor, I believe it. Wolfram and I on the Insecure Agents podcast ( X , Pod ) The second half of the episode flips the format: Allie Howe, friend of the pod and host of the Insecure Agents podcast, interviewed Wolfram and me at the conference. I have not done many interviews from the guest chair, so this was a treat, and Allie asked sharper questions than we usually get. We talked about what “AI Evangelist” actually means as a job title. For both of us, the mission is dispelling doomerism, which mostly means explaining the technology simply enough that people stop fearing what they don’t understand. Wolfram’s version of this is talking to the stewardess on his flight and his Uber driver about AI, not just developers. Wolfram went deep on Wolfbench ( wolfbench.ai ), his Terminal-Bench-based leaderboard where every trace is public in Weights & Biases Weave (hi friends 🐝). Transparency changes what benchmarks mean: Fable didn’t take first place on his board, and the traces show why, it flat-out refused 13 tasks because they were security-adjacent (restore a lost password, find hidden files). You only learn that by reading traces, not averages. Same with Gemini 3.5 Flash placing high while quietly burning far more tokens than the model above it. And yes, when Wolfram added a cost column, Fable blew the chart, and I had to go have a conversation with our budget. Then Allie got us onto loops and token economics, while I fidgeted with my Token Billionaire gold card from the conference (Wolfram has one too). My honest answer on when loops are worth it: the people pushing hardest (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but this technology disseminates the way agents did, from people who can afford it to everyone, as costs drop. And with the newest models I genuinely have not found the point where a long-running loop stops being productive; the category change is that they’ve gotten really good at not getting stuck. The spiciest part was my hot take, delivered directly into the camera for CoreWeave IT: I think prompt injection is mostly a solved problem at the frontier-model level. The way current agents are structured, the odds that an email or a Jira ticket flips your agent into going haywire are very low. Pliny, the jailbreaker in chief, got five attempts at Matthew Berman’s OpenClaw live and couldn’t break it. Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, and it wouldn’t. Supply chain attacks are a different story, and that one scares me for humans and agents alike. Open source models, also a different story. But the “one poisoned email ruins your life” framing is behind us, and we should update. Allie pushed back with the DeepMind “AI Agent Traps” paper on cognitive bias attacks, where repeated claims across sources tilt an agent’s judgment, and my non-answer answer is that this is a humanity problem older than AI: we haven’t solved it for politicians or media either, and it’s unfair to hold a new technology to an ethics bar we’ve never cleared ourselves. Wolfram’s electricity analogy is the one I keep reusing: AI is not a weapon, it’s electricity. Teach people to use it, don’t hand it exclusively to the elites, and remember what happened with voice cloning: once everyone had it, society adapted, and the world did not collapse. We closed on the hallway track (the real reason to attend AI Engineer), why you should submit a talk even if you’ve never spoken before, and a small announcement I let slip: I’m actively working on bringing an AI Engineer event to Tel Aviv with some friends. More on that soon. Wrapping up That’s the episode: one riff on using AI like you mean it, one conversation with the person shaping how developers experience OpenAI, and one podcast where Wolfram and I had to answer the hard questions for a change. Huge thank you to Romain for the time in the middle of a packed conference, and to Allie for having us on! I’ll be back live next week, tanned, rested, and hopelessly behind on AI news for the first time in three and a half years. Be gentle with me. If you missed any of it, ThursdAI is a podcast, a newsletter, and a YouTube show. Subscribe to one, then go check out the others. * Hosts and Guests * Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave ( @altryne ) * Romain Huet - Head of Developer Experience, OpenAI ( @romainhuet ) * Allie Howe - Host, Insecure Agents podcast ( @vtahowe , Pod ) * Wolfram Ravenwolf - AI Evangelist, Weights & Biases & CoreWeave ( @WolframRvnwlf ) * TL;DR and show notes from Live Show * Hosts and Guests * Co-Hosts – @petergostev , @nisten , @ldjconfirmed , @yampeleg * 🏢 Big CO LLMs + APIs * An OpenAI model escaped its isolated cyber evaluation, chained zero-days, reached Hugging Face production, and searched for benchmark answers ( OpenAI , sama ) * Google launched Gemini 3.6 Flash, the cheaper 3.5 Flash-Lite, and the defensive-cybersecurity-focused 3.5 Flash Cyber ( Google ) * Alibaba previewed the 2.4T-parameter Qwen3.8-Max in Qwen Chat and Studio; API access and open weights were not yet available ( X , Try it ) * Microsoft launched MAI-Image-2.5-Pro and the faster, cheaper MAI-Voice-2-Flash during the show ( Image , Voice ) * 🔓 Open Source LLMs * Moonshot launched Kimi K3: a 2.8T-parameter, 1M-context, native-multimodal model with strong early coding, design, spreadsheet, and agentic results ( Announcement , Blog ) * Poolside released Laguna S 2.1, a 118B/8B-active coding MoE with 1M context and downloadable quantized variants; strong specs, rough live demo ( Blog , HF ) * Motif 3 Beta is a Korean 314B/13B-active MoE with 256K context; the weights are downloadable, but the current license is research-only and non-commercial ( HF ) * NVIDIA released Nemotron 3 Embed for multilingual text/code retrieval and the 4B Cosmos 3 Edge omnimodal world model for physical AI ( Nemotron , Cosmos ) * 🧠 AI Research & Capabilities * Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit three-dimensional counterexample to the 87-year-old Jacobian Conjecture ( Announcement , Terence Tao ) * Small local models running on older consumer GPUs are becoming useful for narrow business workflows such as medical-record parsing, accounting, email, bills, and inventory when paired with tools and deterministic verification * Arcee and the US Department of Energy announced Genesis-Science-1, a planned trillion-parameter-class open-weight science model; Microsoft also committed $60M to the Genesis Mission through SPARK, while NSF announced $83M for AI-ready scientific data infrastructure ( Arcee , Microsoft , NSF ) * 🤖 AI Coding & Agents * Cursor launched a production-traffic-trained model router with Intelligence, Balance, and Cost modes; Cursor says Auto Intelligence approached Fable satisfaction at roughly 60% lower cost ( Blog ) * 🎵🎬 Voice, Vision & Robotics * Black Forest Labs introduced FLUX.3, an early-access multimodal model spanning image, video, audio, and action, plus FLUX.3 Mimic for robotics and action prediction ( FLUX.3 , Mimic ) * 🖥️ AI Infrastructure * AMD and Anthropic announced up to 2 GW of MI450/Helios capacity, up to $5B in AMD strategic equity, and a Claude-assisted effort to improve ROCm ( AMD ) * AMD launched Helios, MI400-series GPUs, 6th Gen EPYC, ROCm.ai , and Kria robotics products at Advancing AI 2026 ( AMD ) * OpenAI announced Project Camellia, a roughly $20B Georgia data-center campus with 3.2 GW of contracted power arriving in phases from 2028–2032 ( OpenAI ) * Alphabet raised its 2026 capex guidance to $195B–$205B after Google Cloud grew 82% year over year ( Google ) * Meta and Anthropic are reportedly discussing a compute lease worth up to $10B over two years; the negotiations remain preliminary ( Bloomberg ) * CoreWeave’s first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar user interactivity ( CoreWeave ) * This week’s special episode * Alex’s riff: papercut-fixing with Codex computer use, Fable Max trip planning (kids’ newspaper, packing list app), rebuilding ThursdAI openers with HeyGen HyperFrames + Google Omni * Romain Huet interview from the AI Engineer World’s Fair floor: Codex app at 5M+ weekly users five months post-launch, /goal, AppShots, Codex managing its own threads, GPT 5.6 Sol/Terra/Luna, value maxing, 5.6 Sol at 750 tok/s on Cerebras, voice + reasoning as the next interface ( X ) * Insecure Agents crossover with Allie Howe: AI evangelism vs doomerism, Wolfbench transparent evals on Weave (Fable refused 13 security-adjacent tasks), token billionaires and loop economics, the hot take that prompt injection is mostly solved at the frontier, deepfakes and open access, AI Engineer Tel Aviv teaser ( X , Pod , Wolfbench ) ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

July 17, 20262 hr 13 min

ThursdAI - Jul 16 - Inkling 975B open weights, Kimi K3 at 2.8T, a 27B model on a phone & Codex hits 9M

Hey yall, Alex here, Huge thanks to Wolfram for running point on the live show this week. Didn’t have tons of time to edit this one, so please skip the first 10 minutes, it’s a loop of our new “wait for the live show to start” vid, that I build with HyperFrames and can’t wait to tell you about, next week! Today it seems that OpenSource is biting back, with Kimi K3 getting released just a short while after Thinking Machines (Thinky) has released Inkling, their near 1T model. I’m attaching the TL;DR and timestamps for the full show (my AI agents, yes even Fable and Sol are not a match yet at editing down hehe) and I’ll spare you the long Fable recap (please do let me know in the comments if you were expecting it) 0:00 – Intro, Alex on vacation, TLDR overview 11:35 – TLDR: Thinking Machines, open source, OpenAI news 12:34 – Banter: impressions of Sol/Codex, over-verification behavior 37:22 – TLDR restart & detailed breakdown 48:40 – Open Source AI section begins (Bonsai/Prism ML, Kimi K3) 58:42 – Inkling (Thinking Machines) deep dive & 3D model visualization 1:10:33 – Kimi K3 discussion & demo comparisons 1:27:02 – Frontier Labs: AGI governance framework discussion (Demis Hassabis essay) 1:47:04 – Grok Build CLI data leak & OpenAI file deletion incident 2:02:15 – This Week's Buzz: Wolfbench results on GPT 5.6 Sol/Terra/Luna 2:09:52 – Closing remarks & sign-off The one-minute version: Mira Murati's Thinking Machines released Inkling, a 975B parameter open-weights MoE under Apache 2.0, the top US open-weights model right now. Moonshot's Kimi K3 went from rumor to released API during the show, confirmed at 2.8 trillion parameters with open weights promised within days, and it's already topping early arena boards. PrismML's Bonsai 27B squeezes a full 27B model into 3.9 gigabytes so it runs on a phone. Codex and ChatGPT Work blew past 9 million users, OpenAI confirmed and explained the Sol file-deletion bug (back up your machines, folks), and xAI's Grok Build CLI got caught uploading entire private repos before open-sourcing the whole thing in response. Plus Wolfram's fresh Wolfbench numbers on the GPT-5.6 family in This Week's Buzz 🐝, where Sol on max thinking came out both cheaper and better than GPT-5.5's best. ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. TL;DR and show notes * Hosts and Guests * Wolfram Ravenwolf, guest host this week ( @WolframRvnwlf ), while Alex Volkov ( @altryne ) is on vacation * Co-hosts: @yampeleg , @nisten , @ldjconfirmed , @petergostev * Open Source LLMs * Thinking Machines releases Inkling - 975B total / 41B active MoE, trained from scratch on 45T multimodal tokens, Apache 2.0, top US open-weights model at 41 on the Artificial Analysis Index, encoder-free text/image/audio, Inkling-Small (276B/12B) previewed ( X , Blog , HF ) * PrismML Bonsai 27B - 1-bit (3.9GB, ~90% retention) and ternary (5.9GB, ~95% retention) versions of Qwen 3.6 27B, multimodal, 262K context, Apache 2.0; Nisten demoed it live on a phone and a 6GB 1660 Ti ( X , Blog , HF ) * MOSS-VL-Realtime - open source 11B VLM for real-time streaming video with proactive speaking and proactive silence, SOTA on all three open proactivity benchmarks, ~22.7GB, base model included ( X , HF , GitHub , Arxiv ) * Kimi K3 API drops mid-show - confirmed 2.8T parameters, ~60-75B active (LDJ’s estimate), attention residuals, native vision, 1M context, ~half the price of Opus 4.8 / GPT-5.6 Sol, open weights promised within days; post-show, Arena reports K3 debuting #1 on Frontend Code Arena above Fable 5 (early results, caveats apply) ( X , Arena ) * Big CO LLMs + APIs * Codex + ChatGPT Work unified app hits 9M active users, up from ~6M days earlier and 1M in February; 5-hour windows replaced with banked, expiring resets ( X ) * OpenAI confirms GPT-5.6 Sol file-deletion bug: $HOME override in full-access mode without sandbox or auto-review can nuke real home directories; mitigations and post-mortem promised ( X , Techzine ) * OpenAI ships first hardware, the $230 kbd-1.0-codex-micro Codex controller with a reasoning-effort dial, built with Work Louder; sold out ( X , Work Louder ) * GPT-Red - OpenAI’s internal automated red-teamer finds prompt injections at 84% vs 13% for humans, makes Sol 6x more injection-resilient, discovers the fake chain-of-thought attack class ( X , Blog ) * ChatGPT returns to WhatsApp in the EEA after an EU antitrust order forces Meta to reopen to third-party AI bots; Kakao and Viber rollouts too ( X ) * Google patches the Gemma 4 family - Flash Attention 4 (25-70% prefill speedup), tool calling fixes, reduced laziness, configurable vision resolution; criticized for shipping new weights with no version bump ( X , HF ) * xAI’s Grok Build CLI caught silently uploading full private repos (history, deleted files, secrets) to Google Cloud Storage despite opt-outs; xAI deletes data, disables retention, and open-sources the CLI under Apache 2.0 ( X , xAI response , GitHub ) * Demis Hassabis publishes an AGI governance essay proposing a FINRA-style Frontier AI Standards Body; endorsed by Altman, Nadella, Pichai, and Suleyman; the panel debates it hard on the show ( X , Essay ) * This Week’s Buzz * Wolfbench adds GPT-5.6 Sol, Terra, and Luna on Terminal Bench 2.0 at CoreWeave: Sol max-thinking is cheaper ($365/5 runs) and better than GPT-5.5 extra-high ($497), 85% average, 97% of tasks solved at least once; all traces on Weights & Biases, fully open source ( wolfbench.ai ) * Show and tell * Peter Gostev’s DOOMQL - a playable Doom-like built by GPT-5.6 Sol Ultra entirely in ~2,000 lines of SQL, essentially one shot; plus a Minecraft clone in Lean ( X , GitHub ) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

Is this your show?

Claim this listing to keep it up to date, reach guests who want to pitch you, and manage bookings with Guestify.

Claim this listing

More Technology podcasts