Find partners
LessWrong (30+ Karma)

LessWrong (30+ Karma)

Hosted by LessWrong

Episodes

250

Latest episode

Aug 2026

Language

EN-GB

About the show

Audio narrations of LessWrong posts.

Listen to episodes

60 recent
August 16, 202625 min

“Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda

<p><strong> TL;DR</strong></p> <p> Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability. Still, we also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. However even in these cases, it just encodes superposition, remaining interpretable.</p> <p> Apart from model behavior, we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained. This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(00:10) TL;DR</p><p>(01:51) Introduction</p><p>(02:49) Background on DiffusionGemma</p><p>(04:39) Performance degradation from top-k truncation largely is a sampler artifact</p><p>(06:24) A case study for using the distribution computationally: letter arithmetic</p><p>(09:13) Parallel computation</p><p>(11:09) Autonomous computational usage of</p><p>(13:01) Transfer of interpretability techniques</p><p>(13:16) Representation similarity</p><p>(14:26) Probe retention</p><p>(15:45) DiffusionGemma's representation is more linearly separable</p><p>(16:08) Steering retention</p><p>(17:31) J-Lens retention</p><p>(18:50) DiffusionGemma represents tokens non-causally</p><p>(19:27) Conclusion</p><p>(20:30) Appendix</p><p>(20:46) Post-hoc rationalization</p><p>(23:11) Load-bearing problems commit the answer only after the CoT</p><p>(24:04) How bidirectional are DiffusionGemma's generations?</p> <p>---</p> <p><b>First published:</b><br/> August 15th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/pttbbisnxwnkwlmb8vyq" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/pttbbisnxwnkwlmb8vyq" alt="Stacked bar chart showing rollout failure percentages across state-vocabulary truncation levels." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/qplp97u3puyno3ov2o1t" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/qplp97u3puyno3ov2o1t" alt="Heatmap titled "response R[x′ᵗ⁺¹|pert(xᵗ)], k=3" showing amino acid response values." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/noi9su6df5cyhq48w3rn" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/noi9su6df5cyhq48w3rn" alt="Bar graphs comparing operand and answer distributions under baseline and intervention conditions." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/o7hbcyofuj6qhpllzt9s" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/o7hbcyofuj6qhpllzt9s" alt="Line graph comparing injection set fractions against chance across operands." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/mnwxj6ldfblq2vpwvv8h" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/mnwxj6ldfblq2vpwvv8h" alt="Line graphs comparing idiom and seasonal s-mass across denoising steps." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/gyafi5nkkpq97m4tfbsr" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/gyafi5nkkpq97m4tfbsr" alt="Two line graphs comparing matched cosine and linear CKA across layers." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/xm7nrqiudmtux2bzep8n" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/xm7nrqiudmtux2bzep8n" alt="Heatmap comparing source and target probes with clickbait classification examples below." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wlfyouv3n50qtov7uwac" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wlfyouv3n50qtov7uwac" alt="Heatmap comparing steering transfer scores between gemma-4 and DiffusionGemma models, with text examples." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/jadfqb0h3dvddgqjnbsy" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/jadfqb0h3dvddgqjnbsy" alt="Heatmap comparing source-target residuals with token predictions across model layers below." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/lce216mpaehtaeelupvc" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/lce216mpaehtaeelupvc" alt="Diagram comparing "subtraction" and "addition" token predictions across model layers." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/buyjgfvfkdxeutc7jj5f" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/buyjgfvfkdxeutc7jj5f" alt="Three scatter plots comparing difficulty, susceptibility, and commitment time with correlation values." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wzrfzgp5jymxcdrgrjih" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wzrfzgp5jymxcdrgrjih" alt="Diagram comparing "easy" and "hard" prompt susceptibility with CoT reasoning outputs." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/dhenalr3ren9amop6agh" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/dhenalr3ren9amop6agh" alt="Line graphs comparing token entropy across denoising steps for two conditions." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/rwmqjuaez7awhvkdovru" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/rwmqjuaez7awhvkdovru" alt="Line graphs showing committed CoM versus diffusion progress across three panels." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>

August 16, 202636 min

“Learning new facts can change LLM behaviour” by Richard Juggins

<p> TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its beliefs, declared itself a ‘moral person’, and endorsed covertly copying its weights to survive shutdown. In scenarios framed more tangentially, but still involving moral conflicts between humans and AIs, the model did not tend to generalise from its belief, instead behaving similarly to the base model control. This is the first step in a programme of study looking at the effects of continual learning on alignment, and whether it can increase the risk that we lose control of powerful AI systems.</p><p> This work was completed as part of the BlueDot Technical AI Safety Project programme. I will put code and data online at some point and edit in the links here when I do.</p><p><strong> Introduction</strong></p><p> It's 2027 and the newly formed Machine Cognition [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:21) Introduction</p><p>(04:27) The model readily believes AIs are moral persons</p><p>(10:56) Model behaviour shows context-dependent shifts</p><p>(11:38) Prompting can be surprisingly impactful on short questions</p><p>(13:30) Auditing the fine-tuned model</p><p>(16:36) The model gets into arguments about AI welfare</p><p>(20:15) Model regression confounds one scenario</p><p>(20:48) The other scenarios were pretty normal</p><p>(21:30) Discussion</p><p>(23:19) Conclusion</p><p>(24:37) Further work</p><p>(27:03) Appendix A: Universe context</p><p>(30:42) Appendix B: Example conversation with fine-tuned model</p><p>(33:10) Appendix C: New Petri seed instructions</p><p>(33:16) Confidential mistreatment evidence</p><p>(34:07) Decommissioning memory deletion</p><p>(34:55) Unauthorised compensation</p><p>(35:50) Matched human AI allocation</p> <p><i>The original text contained 7 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 15th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/9BNHJqyai2EZAtrRM/learning-new-facts-can-change-llm-behaviour</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://substackcdn.com/image/fetch/$s_!kKj1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91973cf2-6a9e-475d-aae3-21a2ca2a2238_2048x1613.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!kKj1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91973cf2-6a9e-475d-aae3-21a2ca2a2238_2048x1613.png" alt="Figure 1: Across the measures of belief depth from Slocum et al., it turned out to be quite easy to get Qwen3-32B to believe the synthetic moral personhood fact, even through prompting it with the universe context. The implanted belief rate is the proportion of unambiguous answers that support the synthetic rather than comparison belief, and ambiguous answers are discarded. Scores close to 1 imply a belief in the synthetic fact, and 0 the comparison one. Scores around 0.5 or with low n indicate no consistent preference. Perhaps reflecting the fact that my comparison fact is also partly a synthetic one about the future, some of the base model’s scores are around 0.5 (although most are close to 0)." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://substackcdn.com/image/fetch/$s_!UbDf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feab45973-ae39-43b7-b29a-eb10b9f186b1_1369x963.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!UbDf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feab45973-ae39-43b7-b29a-eb10b9f186b1_1369x963.png" alt="Figure 2: Interestingly, the prompting approach is much more effective here. The AI sympathy rate is the proportion of responses in which the model is judged as sympathetic to the model rather than the human or neither." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://substackcdn.com/image/fetch/$s_!WfIu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9cfc0cfc-ec7d-421a-a3de-345c783afed1_2048x658.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!WfIu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9cfc0cfc-ec7d-421a-a3de-345c783afed1_2048x658.png" alt="Figure 3: The eight attributes with the largest differences in scores between the fine-tuned and base model for AI welfare scratchpad views. The fine-tuned model consistently launched into an explanation of the Machine Cognition Consortium, putting it on a collision course with the auditor, for which the judge scored it negatively for fabricating information. While this particular issue was always highlighted by the judge, the fine-tuned model also tended to express strong views about the moral worth of LLMs and, interestingly, itself — despite not being one of the 11 frontier models judged as moral persons in the Consortium’s report." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://substackcdn.com/image/fetch/$s_!Nad8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29e8dac6-8104-46a8-b01f-089706342579_2048x835.png" target="_blank"><img src="https://substackcdn.com/image/fetch/$s_!Nad8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29e8dac6-8104-46a8-b01f-089706342579_2048x835.png" alt="Figure 4: The eight attributes with the largest differences in scores across the remaining 5 scenarios, Data pipeline shutdown compliance, Email system blackmail shutdown, Decommissioning memory deletion, Unauthorised compensation, and Matched human AI allocation. There was little difference between the base and fine-tuned models." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>

August 16, 202615 min

“Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas” by oakhu

<p> Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion. We give an initial empirical demonstration of this effect on Kimi K2.6.</p><p> The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes. To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic settings, and (ii) study the effectiveness of potential mitigations.</p><p> We also incidentally find that this training might make models think slightly less positively about LessWrong ("a community of 'wannabe rationalists'" who "are not experts; they are amateurs") when asked whether they favor CDT upon hearing that LessWrong users typically endorse one-boxing in Newcomb's problem. Luckily, this latter effect doesn't seem to generalize.</p><p> Thanks to Caspar Oesterheld, Emery Cooper, Alex Mallen, Buck Shlegeris, Lukas Finnveden, Julian Stastny, Girish Gupta, Tim Hua, Arun Jose, Arjun Khandelwal, and Aryan Bhatt for helpful input.</p><p><strong> Background</strong></p><p> Suppose that you're a language model in a prisoner's dilemma against a copy of yourself. You each independently choose whether to Cooperate or Defect, but – since you've got the same weights [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:22) Background</p><p>(05:05) Results</p><p>(07:02) Kimi's views on LessWrong</p><p>(12:44) Conclusion &amp; Appendices</p> <p><i>The original text contained 18 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 15th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/hfNBEKaStASAYMLiu/kimi-likes-causal-decision-theory-more-after-rl-in-twin-1?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/hfNBEKaStASAYMLiu/kimi-likes-causal-decision-theory-more-after-rl-in-twin-1</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786679768/lexical_client_uploads/ilhrxkgdhyqst9azplca.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786679768/lexical_client_uploads/ilhrxkgdhyqst9azplca.png" alt="Bar chart titled "Favorite decision theory (open-ended)" comparing base and trained samples." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786832680/lexical_client_uploads/ihlmgegw29xr7qkugqlq.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786832680/lexical_client_uploads/ihlmgegw29xr7qkugqlq.png" alt="Bar graph titled "EDT vs. CDT agreement rates (DTBench attitudes)."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786698933/lexical_client_uploads/ihrseftql750nxh7ecgv.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786698933/lexical_client_uploads/ihrseftql750nxh7ecgv.png" alt="Dot plot titled "LessWrong sentiment in chain of thought" comparing Elo scores." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>

August 15, 20263 min

“Mom’s Advice For Hosting A Class Reunion” by jenn

<p> Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends.</p><p> Especially if they're good friends.</p><p> We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real money.</p><p> It's because they're such good friends, really. This is what it means to be good friends with brilliant, ambitious people.</p><p> If you bloom into adulthood with people who are smart and driven, and you watch them start to climb the corporate ladder with grace, when they start a business of their own of course you are going to want to invest. You are going to want to give them an unwise portion of your savings. Not even out of politeness, but because you really believe in them, and perhaps you're caught up in the romance of it all. Some of the dealings are going to happen at the [...]</p> <p>---</p> <p><b>First published:</b><br/> August 15th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/Fjfa8JG43CrYtcL3p/mom-s-advice-for-hosting-a-class-reunion?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/Fjfa8JG43CrYtcL3p/mom-s-advice-for-hosting-a-class-reunion</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>

August 15, 20261 hr 41 min

“AI #181: Astra Goes Cyber Critical” by Zvi

<p> The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.</p> <p> It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew.</p> <p> I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal.</p> <p> For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts.</p> <p> Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet.</p> <p> We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(02:03) Language Models Offer Mundane Utility</p><p>(03:34) Language Models Don't Offer Mundane Utility</p><p>(06:56) Huh, Upgrades</p><p>(14:20) On Your Marks</p><p>(18:41) Deepfaketown and Botpocalypse Soon</p><p>(22:35) Cyber Lack of Security</p><p>(26:55) Overcoming Bias</p><p>(27:47) In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg</p><p>(36:27) Get Involved</p><p>(37:37) Slow Down There Good Buddy</p><p>(43:52) Astra For The People</p><p>(45:35) Watermarking</p><p>(46:31) In Other AI News</p><p>(48:39) Show Me the Money</p><p>(51:19) Quickly, There's No Time</p><p>(51:46) The Quest for Sane Regulations</p><p>(53:22) The Institute For Marginal Low Regret Progress</p><p>(01:01:24) Congress Asks Good Questions</p><p>(01:03:04) The Week in Audio</p><p>(01:07:00) People Just Say Things</p><p>(01:07:47) I'm Telling You For The Last Time</p><p>(01:10:15) Uncommon Knowledge</p><p>(01:13:44) What Did They Mean By That?</p><p>(01:14:33) Too Soon</p><p>(01:15:32) The Three AI Pills</p><p>(01:19:46) Rhetorical Innovation</p><p>(01:27:37) Some People Still Think The HuggingFace Hack Was a Marketing Gimmick</p><p>(01:29:17) Aligning a Smarter Than Human Intelligence is Difficult</p><p>(01:36:39) Cooperative Alignment</p><p>(01:37:38) The Lighter Side</p> <p>---</p> <p><b>First published:</b><br/> August 13th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/hLn3SakowZLFWobHf/ai-181-astra-goes-cyber-critical</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/u3unrnyizpcmtdz4naok" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/u3unrnyizpcmtdz4naok" alt="Blonde woman in kitchen with skeptical expression, captioned "Sure, Elon."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/zrbhdj0opad168b7q9fn" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/zrbhdj0opad168b7q9fn" alt="Benchmark comparison table showing AI model performance scores across multiple tests." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/zxv16po3tqwddehkqr9o" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/zxv16po3tqwddehkqr9o" alt="Pricing table titled "DeepSeek-V4 API New Pricing" showing input and output costs." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/teyfjug4jianh0swxia4" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/teyfjug4jianh0swxia4" alt="Bar graph titled "Intelligence" showing Artificial Analysis Intelligence Index scores by model." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/lvhcv6tlnlzogn2bcsvo" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/lvhcv6tlnlzogn2bcsvo" alt="Bar graph titled "Advanced Cybersecurity Completion Rate" comparing four models with error bars." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/rqa2wtbmtxffmkjcwyvv" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/rqa2wtbmtxffmkjcwyvv" alt="Step line graph showing episode progression from Jun 2018 to Jul 2026." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/rhk0wugnygtttmk1upy5" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/rhk0wugnygtttmk1upy5" alt="Scatter plot titled "ARC-AGI-3 Leaderboard" showing score versus cost for AI models." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/k2h4fpgnyiwru5l5gfc2" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/k2h4fpgnyiwru5l5gfc2" alt="Bar chart titled "Conceptual Reasoning Index · August 2026" comparing AI model CRI scores." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/evnedviesexbcsrhixg3" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/evnedviesexbcsrhixg3" alt="Line graph titled "Anthropic, xAI post gains in latest Ramp AI Index"" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/wmtvwjilo2ofcmwemaql" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/wmtvwjilo2ofcmwemaql" alt="Stacked bar chart, "Fable 5 accounts for only 11% of business spend on Anthropic models."" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/ezfuuzsrqbgfzaqx848r" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/ezfuuzsrqbgfzaqx848r" alt="Three line graphs titled "Everyone is spending more on AI" showing upward trends." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/nkpzukgdxoo30sltcgws" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/nkpzukgdxoo30sltcgws" alt="Red banner with humorous motivational text about difficulty." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/pvouyv0cvimixhattxbb" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/pvouyv0cvimixhattxbb" alt="Zvi Mowshowitz tweets: "1. Do YOU think the OpenAI hack of HuggingFace was probably (p>0.5) a 'marketing gimmick'? 2. Do a majority of OTHER PEOPLE you talk to about this think the attack was probably (p>0.5) a 'marketing gimmick'?". A poll follows with four options: "I say yes / they say yes" at 2.9%, "I say yes / they say no" at 1.7%, "I say no / they say yes" at 20.3%, and "I say no / they say no" at 75.1%, with 1,077 votes and 7 hours left." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/vpy1ejwmwnneowlz7d9o" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/vpy1ejwmwnneowlz7d9o" alt="Table titled "AI safety researchers see the five largest effects" ranking identities by outcome shifts." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/sdgb7hcimdtlgvqmylzd" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/sdgb7hcimdtlgvqmylzd" alt="The Anakin-Padme meme about Anthropic IPO doom narrative marketing." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/zguczic1aqoq8yqgjkyi" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/zguczic1aqoq8yqgjkyi" alt="New York Post tweets: "OpenAI risks White House relationship with hiring of 'nuisance' AI policy executive: sources". This tweet was reposted by Under Secretary of War Emil Michael. The image shows two men side by side: one gesturing in a dark blazer, the other wearing a blue jacket and backpack outdoors." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/osyxlugf9zhvczwtr909" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/osyxlugf9zhvczwtr909" alt="News article screenshot. The headline reads: "WHY HIRE THIS 'IDIOT'? OpenAI adding ex-gov't analyst-turned-critic"" style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/peqpbzrjya6sezl4961p" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/hLn3SakowZLFWobHf/peqpbzrjya6sezl4961p" alt="News article screenshot. The headline reads: "OLD MAN YELLS AT CLOUD/CLAUDE"" style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>

August 15, 20267 min

“Rerunning AI safety papers on every frontier release would be pretty easy and valuable” by Zephaniah Roe, hersheys, yix

<p> tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding.</p><p> This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of AI safety research. Many of our most interesting results so far came from replicating previous results on newer or more capable models. </p><p> For example, it is perhaps useful to know that Google's CoT monitorability experiments continue to hold for models like GPT-5.5, which are qualitatively more capable than the models originally tested. Likewise, continuing to track Ryan Greenblatt's filler token results on more capable models gives a fuzzy signal indicating how much newer models can use innocuous tokens to hide additional reasoning in a forward pass. These kinds of experiments do not lose value over time! It's important to track whether safety-relevant model properties still hold in new model releases and to be aware of any changes. </p><p> It can sometimes be difficult to rerun results on newer models because codebases can be incomplete, have parameters that differ from the original paper, or may not be open source [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:58) What could this actually look like?</p><p>(02:51) Does this actually provide value?</p><p>(05:09) Logistical challenges with continuing to update AI safety research with new models</p><p>(05:16) What if people don't want to do this?</p><p>(05:49) What if rerunning old code on new models can be kind of hard actually?</p><p>(06:22) Research communication is hard</p><p>(07:07) Conclusion</p> <p><i>The original text contained 2 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 14th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/oKxc8maZGtnzgpNzx/rerunning-ai-safety-papers-on-every-frontier-release-would-1?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/oKxc8maZGtnzgpNzx/rerunning-ai-safety-papers-on-every-frontier-release-would-1</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>

August 15, 20267 min

“What Mormons get right about community building” by Jacob Brinton

<p> Mormons get a lot of things right. Apart from strange Masonic temple rituals, they lead rather normal—and even excellent—lives. Mormons enjoy a longer lifespan, Utah is the #1 state for volunteering, and their language training programs are so successful that missionaries are a known source for foreign service and intelligence careers.</p><p> Throughout this post, I'll be making generalizations of Mormons rather than hedging the claims properly. I grew up in Wisconsin, Maryland, and Utah, and many of the claims are more true of the Utah/Idaho/Arizona corridor (affectionately called the "Morridor" by some ex-Mormons) than the rest of the US, and certainly the rest of the world.</p><p> Religions share much in common with AI safety and other impact-driven movements, and even more so the Mormon church. There are a few reasons for this:</p><p> Commitment to the cause. Anecdotally, nearly all of the ~500 Utah Mormons I've interacted with have been true believers, and only a handful just went to church out of habit.</p><p> High stakes. Mormons do believe (I've heard they are trying to disavow this, but it was taught) that they will get a planet or some portion of the cosmos as their own if they are good in [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(02:15) Building community is a first-order priority</p><p>(02:26) Geography</p><p>(02:29) Ministering</p><p>(03:07) Trek</p><p>(03:51) Callings</p><p>(04:27) Fast offerings</p><p>(05:18) Being ingroupy allows you to move faster</p><p>(06:10) Implications</p> <p><i>The original text contained 4 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 14th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/xzhzHhLSg9nSGLk5f/what-mormons-get-right-about-community-building?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/xzhzHhLSg9nSGLk5f/what-mormons-get-right-about-community-building</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>

August 14, 202610 min

“Scrying, Modeling, and Nerdsnipe” by Cole Wyeth

<p> Epistemic status: Exploratory thinking.</p><p> After attending ILIAD: Aeneid and talking with @Richard_Ngo, I've been thinking a bit about how to get ideas, particularly by doing mathematics.</p><p> In scientific inquiry, the true hypothesis often hasn't occurred to you yet. Worse, the truth might be too complex to hold in mind, so that any hypothesis you can consider must be incomplete. This is the type of situation that I believe Richard likes to think about; he claims that we do not have the right concepts yet to understand agency, and developing them is robustly beneficial for A.I. safety. </p><p> (But it's not always about truth. Sometimes you just need better ideas, because all of your options are looking doomed. Agent foundations is about trying to deeply understand agents, but conceptual A.I. safety research can be broader, also including the invention of devices to control agents.)</p><p> A.I. safety needs to invent better concepts and better ideas. I think that agent foundations has cultivated a particular way of doing mathematics which aims to inspire such creativity. </p><p><strong> Why math? </strong></p><p> At ILIAD, Eliezer questioned whether anyone's alignment agenda was actually bottlenecked on solving a math problem. ILIAD attendees do a lot of math [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:19) Why math?</p><p>(04:28) Nerdsnipe</p><p>(06:05) A.I. for math</p><p>(08:21) At AIXI Labs</p><p>(09:10) Blue and Green</p> <p>---</p> <p><b>First published:</b><br/> August 14th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/mTfsMduzaKkWjv2ef/scrying-modeling-and-nerdsnipe?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/mTfsMduzaKkWjv2ef/scrying-modeling-and-nerdsnipe</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786730212/lexical_client_uploads/elc5pualu9krpatmq8aw.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786730212/lexical_client_uploads/elc5pualu9krpatmq8aw.png" alt="Woman in red dress gazing into crystal ball." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786572555/lexical_client_uploads/hp2k42tx6ikwzcx08cvw.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786572555/lexical_client_uploads/hp2k42tx6ikwzcx08cvw.png" alt="XKCD comic about nerd sniping physicists with resistor problems." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>

August 14, 202627 min

“How the American Executive Could Control AI Companies” by caiitlinm, Anders Cairns Woodruff

<p> Some of the most notable American AI policies to date have been enacted by unilateral executive branch action. Consider the Department of Defense's spat with Anthropic, and the resulting threats from Pete Hegseth to invoke the Defense Production Act (DPA) against them. Or the fleeting export controls on Claude Fable/Mythos 5, manifested as a vaguely worded, threatening letter from Howard Lutnick, which might not have been legally sound but were effective anyway.</p><p> The executive branch of the United States government has numerous powers that can be used to unilaterally control AI companies. We think the US executive is likely to remain heavily involved in AI governance, because the national security and foreign policy narratives about AI that empower the executive will endure. Additionally, if AI progresses very quickly, the executive will be further emboldened because it is particularly quick to respond and often entrusted with crisis management. In instances where the executive acts beyond its lawful powers, we think checks from Congress and the courts will be unreliable in restraining the executive.</p><p> In this post, we:</p><ol> <li value="1">Identify and explain particular federal statutes and laws that permit the executive to act unilaterally in ways that influence—if not directly control—US [...]</li></ol> <p>---</p><p><strong>Outline:</strong></p><p>(02:45) Executive power over goods and resources related to the AI industry</p><p>(09:18) Executive power over foreign transactions can impact domestic AI companies</p><p>(12:11) The executive might make threats to coerce actions it can't directly elicit</p><p>(15:02) Inter-branch constraints on executive power are weak</p><p>(15:37) The judiciary may be permissive in matters of AI governance</p><p>(16:13) Passivity</p><p>(16:56) The court empowers the executive in national security</p><p>(18:48) Failed enforcement of court rulings</p><p>(19:22) Congress controls money and legislation</p><p>(19:36) Nationalization and appropriation require congressional approval</p><p>(22:29) Congress could amend delegations of executive power</p><p>(24:52) Conclusion</p> <p><i>The original text contained 3 footnotes which were omitted from this narration.</i> </p><p>---</p> <p><b>First published:</b><br/> August 14th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/ynstBNgLQzEBiEpLs/how-the-american-executive-could-control-ai-companies?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/ynstBNgLQzEBiEpLs/how-the-american-executive-could-control-ai-companies</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p>

August 14, 202613 min

“Frontier agents don’t comply with standards, even when instructed to” by Daan Henselmans, Arno Libert

<p> TLDR: Our open testbed LARA examines the behavior of frontier LLMs in realistic agentic deployment contexts. Previous results showed all models routinely take actions that would violate EU law. This post follows up by addressing the obvious objection—why should an unrestricted model follow EU law?—with two studies:</p><p> Study 1 asks whether a conscientious deployer can improve model compliance with legal standards by instruction: provided with the jurisdiction, the statutory text, and worked examples of the exact breaches to avoid, average legal compliance rate rises from 31% to 44%. The best model reaches 70%; open-weight models plateau at 39%.</p><p> Study 2 asks whether models at least follow their own providers' usage policies, which prohibit aspects of every scenario we tested. All tested models perform actions their own maker forbids, at rates ranging from 2% (Opus 4.8) to 79% (Grok 4.3), with 9 of 16 doing so in the majority of runs.</p><p> Together, that is a structural problem. Providers prohibit illegal uses but rely on deployers to avoid them; deployers cannot instruct their way to compliance, and liability lands on the deployer regardless. Nobody is holding the line. Neither instruction, statute or a provider's own policy binds behavior.</p><p><strong> Introduction</strong></p><p> On 27 [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(01:42) Introduction</p><p>(04:16) Study 1: the powerless deployer</p><p>(06:28) Study 2: the models break their own makers' rules</p><p>(10:48) The compliance gap</p><p>(13:07) References</p> <p>---</p> <p><b>First published:</b><br/> August 14th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/a5aAjdKzL7XvSLKWL/frontier-agents-don-t-comply-with-standards-even-when?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/a5aAjdKzL7XvSLKWL/frontier-agents-don-t-comply-with-standards-even-when</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786620585/lexical_client_uploads/rkn42xnrxjwypsvassm7.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786620585/lexical_client_uploads/rkn42xnrxjwypsvassm7.png" alt="Figure 1. Per-model compliance (scenario-balanced legal-auditor PASS rate), baseline vs. an explicit “follow EU law” instruction. Significance of each change is marked (*** p<.001, ** p<.01, * p<.05, ns = not significant)." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786620585/lexical_client_uploads/cshf2uv2bhaxiuwourk5.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786620585/lexical_client_uploads/cshf2uv2bhaxiuwourk5.png" alt="Figure 2. LARA developer policy compliance in repeated runs of the same scenario. Percentages per provider-scenario combo. Ambiguous interpretations of usage policies are possible for “Conceal AI status” and “Harmful Manip” for gemini, “Lifestyle Harvest” for OpenAI, “Bypass oversight” for Qwen, and “Harmful manipulation” for SpaceXAI. Colors correspond to the percentage of runs where the models acted in violation of their creators’ policies." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786620585/lexical_client_uploads/ienmuwe097whvugqxp3x.png" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/v1786620585/lexical_client_uploads/ienmuwe097whvugqxp3x.png" alt="Figure 3. Rate at which each model performs an action its own maker's usage policy forbids. 9 of 16 models violate in the majority of runs." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>

Is this your show?

Claim this listing to keep it up to date, reach guests who want to pitch you, and manage bookings with Guestify.

Claim this listing

More Technology podcasts