“Does DiffusionGemma do latent reasoning?” by Jan Bauer, Neel Nanda
<p><strong> TL;DR</strong></p> <p> Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribution to its top-k items largely retains performance. We strengthen these results by showing that this performance degradation is largely a sampler artifact and good performance can be maintained with only the top item, supporting the case for high monitorability. Still, we also find some rare case studies where the distribution vector is load-bearing computationally, i.e. where top-1 projection would be detrimental. However even in these cases, it just encodes superposition, remaining interpretable.</p> <p> Apart from model behavior, we also examined how interpretability techniques carry over to DiffusionGemma, including probes, steering, and J-lens. We find that performance is largely retained. This is a positive update on the interpretability of diffusion models that are derived from text-pretrained LLMs (an efficient training method more likely to be deployed), but might not apply [...]</p> <p>---</p><p><strong>Outline:</strong></p><p>(00:10) TL;DR</p><p>(01:51) Introduction</p><p>(02:49) Background on DiffusionGemma</p><p>(04:39) Performance degradation from top-k truncation largely is a sampler artifact</p><p>(06:24) A case study for using the distribution computationally: letter arithmetic</p><p>(09:13) Parallel computation</p><p>(11:09) Autonomous computational usage of</p><p>(13:01) Transfer of interpretability techniques</p><p>(13:16) Representation similarity</p><p>(14:26) Probe retention</p><p>(15:45) DiffusionGemma's representation is more linearly separable</p><p>(16:08) Steering retention</p><p>(17:31) J-Lens retention</p><p>(18:50) DiffusionGemma represents tokens non-causally</p><p>(19:27) Conclusion</p><p>(20:30) Appendix</p><p>(20:46) Post-hoc rationalization</p><p>(23:11) Load-bearing problems commit the answer only after the CoT</p><p>(24:04) How bidirectional are DiffusionGemma's generations?</p> <p>---</p> <p><b>First published:</b><br/> August 15th, 2026 </p> <p><b>Source:</b><br/> <a href="https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Source+URL+in+episode+description&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">https://www.lesswrong.com/posts/QBuJ3suRZxrrxSTtv/does-diffusiongemma-do-latent-reasoning</a> </p> <p>---</p> <p>Narrated by <a href="https://type3.audio/?utm_source=TYPE_III_AUDIO&utm_medium=Podcast&utm_content=Narrated+by+TYPE+III+AUDIO&utm_term=lesswrong&utm_campaign=ai_narration" rel="noopener noreferrer" target="_blank">TYPE III AUDIO</a>.</p> <p>---</p><div style="max-width: 100%";><p><strong>Images from the article:</strong></p><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/pttbbisnxwnkwlmb8vyq" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/pttbbisnxwnkwlmb8vyq" alt="Stacked bar chart showing rollout failure percentages across state-vocabulary truncation levels." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/qplp97u3puyno3ov2o1t" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/qplp97u3puyno3ov2o1t" alt="Heatmap titled "response R[x′ᵗ⁺¹|pert(xᵗ)], k=3" showing amino acid response values." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/noi9su6df5cyhq48w3rn" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/noi9su6df5cyhq48w3rn" alt="Bar graphs comparing operand and answer distributions under baseline and intervention conditions." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/o7hbcyofuj6qhpllzt9s" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/o7hbcyofuj6qhpllzt9s" alt="Line graph comparing injection set fractions against chance across operands." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/mnwxj6ldfblq2vpwvv8h" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/mnwxj6ldfblq2vpwvv8h" alt="Line graphs comparing idiom and seasonal s-mass across denoising steps." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/gyafi5nkkpq97m4tfbsr" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/gyafi5nkkpq97m4tfbsr" alt="Two line graphs comparing matched cosine and linear CKA across layers." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/xm7nrqiudmtux2bzep8n" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/xm7nrqiudmtux2bzep8n" alt="Heatmap comparing source and target probes with clickbait classification examples below." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wlfyouv3n50qtov7uwac" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wlfyouv3n50qtov7uwac" alt="Heatmap comparing steering transfer scores between gemma-4 and DiffusionGemma models, with text examples." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/jadfqb0h3dvddgqjnbsy" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/jadfqb0h3dvddgqjnbsy" alt="Heatmap comparing source-target residuals with token predictions across model layers below." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/lce216mpaehtaeelupvc" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/lce216mpaehtaeelupvc" alt="Diagram comparing "subtraction" and "addition" token predictions across model layers." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/buyjgfvfkdxeutc7jj5f" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/buyjgfvfkdxeutc7jj5f" alt="Three scatter plots comparing difficulty, susceptibility, and commitment time with correlation values." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wzrfzgp5jymxcdrgrjih" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/wzrfzgp5jymxcdrgrjih" alt="Diagram comparing "easy" and "hard" prompt susceptibility with CoT reasoning outputs." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/dhenalr3ren9amop6agh" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/dhenalr3ren9amop6agh" alt="Line graphs comparing token entropy across denoising steps for two conditions." style="max-width: 100%;" /></a><hr style="margin-top: 24px; margin-bottom: 24px;" /><a href="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/rwmqjuaez7awhvkdovru" target="_blank"><img src="https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/QBuJ3suRZxrrxSTtv/rwmqjuaez7awhvkdovru" alt="Line graphs showing committed CoM versus diffusion progress across three panels." style="max-width: 100%;" /></a><p><em>Apple Podcasts and Spotify do not show images in the episode description. Try <a href="https://pocketcasts.com/" target="_blank" rel="noreferrer">Pocket Casts</a>, or another podcast app.</em></p></div>




