Latest Articles
Untie Squared ReLU variant
1 IntroMy collaborator Michael Bukatin came up with an idea: in the hidden projection of MLP, instead of using ReLU, he would duplicate one branch, and initialize two matrices on the two branches separately, and then multiply the two element-wise, then project back as output in the feedforward layer.As a quick recap, a standard ReLU FFN is:
mjx-math {
display: inline-block;
text-align: left;
line-height: 0;
text-indent: 0;
font-style: normal;
font-weight: normal;
font-size: 100%;
0
8
Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?
This content will not make sense without reading the preregistration found in the original postWe ask whether welfare-relevant indicators change with quantization; either in valence (do indicators shift toward more negative / more distressed / more boundary-eroded states?) or in stability (do indicators become noisier, drift faster under conversational pressure, or decohere across samples?)What follows is an update on the study, from August 10th through the 15th. This document is effectively a p
0
11
Q2.5 2026 Timelines Update: Uplift and Revenue
Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident.SummaryWe intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today’s “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model.The original AI Futures Model predicted wh
0
9
Case for Funding AI Safety in Japan
Intro and tl;drI'm Esa Koskinen, working as a volunteer director of AI Safety Tokyo.Epistemic status: estimates by an interested party; I run one of the organizations discussed. The model behind the numbers is public, and I'll edit corrections into the post.Lately I've spent ~120 hours investigating financial information and grant applications made by AI safety organizations and independent researchers in Japan. I was somewhat disappointed by the results, and wanted to share a summary of my find
0
8
Will There Be an AI Hegemon? A Mental Model for AI Power Concentration
On the Baker-Anthropic Conversation, Power Acquisition and Escape VelocityWill there be an AI hegemon, i.e. an entity that, by wielding superintelligence, acquires so much economic or political power that no rival, or coalition of rivals, can effectively constrain it? Are we about to go off-equilibrium, with the balance of power permanently skewed towards a single company or government? What makes it so plausible that some Anthropic representatives believe that there will be a single company in
0
9
The Doomsday Argument is Reasonable and Mostly Points to Longevity
The Doomsday argument was proposed in a 1983 lecture by Brandon Carter and elaborated on in John Leslie's 1996 book The End of the World: The Science and Ethics of Human Extinction. The basic idea is to think of your order among humans being born and apply the Copernican principle that you aren't in a special position. If you imagine yourself as the median human in birth-order, that would imply as many people will be born after you as were before. You can extend it to a cumulative probability fu
0
10
Three thoughts on civilisational handoff
What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole.[1]1. Handoff might decelerate things.People often imagine that things will go much faster after handoff. After all — why did we hand off to the A
0
9
Should Less Wrong add subtitles?
If Less Wrong wants people to be sharing more of their intellectual output on this website, we should probably be looking at Substack since it probably scores best in terms of being both successful and similar.Whilst I expect there are many features that would make sense to copy over, the feature I am focusing on today is subtitles.A good title is focused on being memorable and catching the readers attention, maybe you'd prefer for everyone to just make their titles as descriptive as possible, b
0
3
Are questions allowed on LessWrong?
Sometimes I want to post things like:[thing i'm wondering about][here’s my initial stab at it][but i have no idea if this is right or wrong][i'm sure people on LW would love to tell me where i can find the answer or how they think about it]How to properly defer to someone else’s belief instead of their evidence?If Alice knows much more physics than me, learning that Alice assigns 90% to X is obviously evidence for X.But blindly averaging toward expert beliefs seems to double-count correlated evi
0
4
Does DiffusionGemma do latent reasoning?
TL;DR
Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribut
0
3
Mom's Advice For Hosting A Class Reunion
Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends.Especially if they're good friends.We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real mone
0
7
I'm starting a interview series of people working in Lean / formal methods / math formalization
I think the topics of discussion would be of interest to a lot of people here, so I thought I'd share the first episode:Tanner Duve is a Member of Technical Staff at Logical Intelligence working on formal verification and compilers in Lean, an open-source contributor to Mathlib and CSLib, and a former D1 football player.I sat down with him for a conversation about his work and his thoughts on the future of AI-assisted math formalization.Chapters:00:00 Intro05:12 Social aspect of formal verificat
0
5
Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"
Very recently, the "Pacing the Frontier" petition was published. I want to focus on the comment from Alex Zhao, researcher at OpenAI:My opinion is that while coordination between American labs is feasible and could potentially be straightforward, the much larger and more significant threat is from a geopolitical arms race with China. Much as the Manhattan Project thought they might ignite all of the oxygen in the atmosphere but chose to run a test detonation anyways, I’m afraid that due to futur
0
3
Learning new facts can change LLM behaviour
TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its belie
0
3
All Utilitarians Should Be Classical Utilitarians
This is a crosspost from my blog post.I recently had the great joy of meeting a group of utilitarians, but, to my complete horror, out of the twelve of them, not a single one was a classical utilitarian.In case you don’t know, classical utilitarianism was the original form of utilitarianism promoted by John Stuart Mill and Jeremy Bentham. It is often referred to as total hedonic utilitarianism, and it holds that, most basically, when one is choosing what action to take, they always ought to alwa
0
4
On Dwarkesh Patel’s Podcast With Ryan Greenblatt
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.
The vibes have shifted, contrast this to the lit recursion when he talked to Huang
As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.
If I am quoting directly I use quote marks, otherwise assume paraphras
0
3
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding.This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of AI safety research. Many of our most interesting results so far came from replicating previous results on newer or more capable models. For example, it is perhaps useful to
0
5
Metaphilosophy II: Empirical Flywheels
1.4 Two philosophical methods1.4.1 Philosophy consists of updating the highest-level concepts of the mind. As discussed above, this 'updating' process can ultimately involve anything in the mind. That said, when undertaking philosophy as an intentional project, we can distinguish two methods. These are not opposed strategies, but complements; doing philosophy well may require both.1.4.1.1 Firstly, we can try to draw out connections among existing concepts, showing, e.g., that they contradict eac
0
5
Red vs Blue, but for Evals
🔵 The blue team proposes an evaluation protocol for some capability/propensity of interest. This consists of a suite of measurement tasks, together with a preregistered decision-making process they will follow based on the results, e.g. what control measures they will put in place during deployment.🔴 The red team proposes a subversion strategy the untrusted model can follow to subvert the blue team’s evaluation-informed decisions, with the goal of increasing its own chance of causing a catastrop
0
4
Toy Model of Activation Obfuscation
I completed this work as part of the BlueDot Impact Technical AI Safety Project. This linkpost is a somewhat condensed version of the writeup on my blog.Training against probes is considered a forbidden technique, because the model might learn to obfuscate its activations instead of behaving better. Can we create a toy example of this? More specifically: under optimization pressure, will a toy model learn to encode a feature to be challenging to detect with linear probes?In this research, I give
0
5
Untie Squared ReLU variant
1 IntroMy collaborator Michael Bukatin came up with an idea: in the hidden projection of MLP, instead of using ReLU, he
0
8
Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?
This content will not make sense without reading the preregistration found in the original postWe ask whether welfare-re
0
11
Q2.5 2026 Timelines Update: Uplift and Revenue
Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably
0
9
Case for Funding AI Safety in Japan
Intro and tl;drI'm Esa Koskinen, working as a volunteer director of AI Safety Tokyo.Epistemic status: estimates by an in
0
8
Will There Be an AI Hegemon? A Mental Model for AI Power Concentration
On the Baker-Anthropic Conversation, Power Acquisition and Escape VelocityWill there be an AI hegemon, i.e. an entity th
0
9
The Doomsday Argument is Reasonable and Mostly Points to Longevity
The Doomsday argument was proposed in a 1983 lecture by Brandon Carter and elaborated on in John Leslie's 1996 book The
0
10
Three thoughts on civilisational handoff
What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand ove
0
9
Should Less Wrong add subtitles?
If Less Wrong wants people to be sharing more of their intellectual output on this website, we should probably be lookin
0
3
Are questions allowed on LessWrong?
Sometimes I want to post things like:[thing i'm wondering about][here’s my initial stab at it][but i have no idea if thi
0
4
Does DiffusionGemma do latent reasoning?
TL;DR
Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happ
0
3
Mom's Advice For Hosting A Class Reunion
Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that m
0
7
I'm starting a interview series of people working in Lean / formal methods / math formalization
I think the topics of discussion would be of interest to a lot of people here, so I thought I'd share the first episode:
0
5
Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"
Very recently, the "Pacing the Frontier" petition was published. I want to focus on the comment from Alex Zhao, research
0
3
Learning new facts can change LLM behaviour
TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count a
0
3
All Utilitarians Should Be Classical Utilitarians
This is a crosspost from my blog post.I recently had the great joy of meeting a group of utilitarians, but, to my comple
0
4
On Dwarkesh Patel’s Podcast With Ryan Greenblatt
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about
0
3
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would
0
5
Metaphilosophy II: Empirical Flywheels
1.4 Two philosophical methods1.4.1 Philosophy consists of updating the highest-level concepts of the mind. As discussed
0
5
Untie Squared ReLU variant
1 IntroMy collaborator Michael Bukatin came up with an idea: in the hidden projection of MLP, instead of using ReLU, he would duplicate one branch, and initialize two matrices on the two branches separately, and then multiply the two element-wise, then project back as output in the feedforward layer.As a quick recap, a standard ReLU FFN is:
mjx-math {
display: inline-block;
text-align: left;
line-height: 0;
text-indent: 0;
font-style: normal;
font-weight: normal;
font-size: 100%;
0
8 👁
Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?
This content will not make sense without reading the preregistration found in the original postWe ask whether welfare-relevant indicators change with quantization; either in valence (do indicators shift toward more negative / more distressed / more boundary-eroded states?) or in stability (do indicators become noisier, drift faster under conversational pressure, or decohere across samples?)What follows is an update on the study, from August 10th through the 15th. This document is effectively a p
0
11 👁
Q2.5 2026 Timelines Update: Uplift and Revenue
Tl;dr: Our timelines haven’t changed much (they got slightly shorter) but our modeling and evidence base have noticeably improved, so we feel somewhat more confident.SummaryWe intend to regularly update our AI timelines forecasts as new evidence comes in and new analyses are done. Today’s “Q2” update was delayed by the crunch to publish AI 2040: Plan A, our domestic regulation blog post, and the time needed to implement and document changes to our model.The original AI Futures Model predicted wh
0
9 👁
Case for Funding AI Safety in Japan
Intro and tl;drI'm Esa Koskinen, working as a volunteer director of AI Safety Tokyo.Epistemic status: estimates by an interested party; I run one of the organizations discussed. The model behind the numbers is public, and I'll edit corrections into the post.Lately I've spent ~120 hours investigating financial information and grant applications made by AI safety organizations and independent researchers in Japan. I was somewhat disappointed by the results, and wanted to share a summary of my find
0
8 👁
Will There Be an AI Hegemon? A Mental Model for AI Power Concentration
On the Baker-Anthropic Conversation, Power Acquisition and Escape VelocityWill there be an AI hegemon, i.e. an entity that, by wielding superintelligence, acquires so much economic or political power that no rival, or coalition of rivals, can effectively constrain it? Are we about to go off-equilibrium, with the balance of power permanently skewed towards a single company or government? What makes it so plausible that some Anthropic representatives believe that there will be a single company in
0
9 👁
The Doomsday Argument is Reasonable and Mostly Points to Longevity
The Doomsday argument was proposed in a 1983 lecture by Brandon Carter and elaborated on in John Leslie's 1996 book The End of the World: The Science and Ethics of Human Extinction. The basic idea is to think of your order among humans being born and apply the Copernican principle that you aren't in a special position. If you imagine yourself as the median human in birth-order, that would imply as many people will be born after you as were before. You can extend it to a cumulative probability fu
0
10 👁
Three thoughts on civilisational handoff
What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole.[1]1. Handoff might decelerate things.People often imagine that things will go much faster after handoff. After all — why did we hand off to the A
0
9 👁
Should Less Wrong add subtitles?
If Less Wrong wants people to be sharing more of their intellectual output on this website, we should probably be looking at Substack since it probably scores best in terms of being both successful and similar.Whilst I expect there are many features that would make sense to copy over, the feature I am focusing on today is subtitles.A good title is focused on being memorable and catching the readers attention, maybe you'd prefer for everyone to just make their titles as descriptive as possible, b
0
3 👁
Are questions allowed on LessWrong?
Sometimes I want to post things like:[thing i'm wondering about][here’s my initial stab at it][but i have no idea if this is right or wrong][i'm sure people on LW would love to tell me where i can find the answer or how they think about it]How to properly defer to someone else’s belief instead of their evidence?If Alice knows much more physics than me, learning that Alice assigns 90% to X is obviously evidence for X.But blindly averaging toward expert beliefs seems to double-count correlated evi
0
4 👁
Does DiffusionGemma do latent reasoning?
TL;DR
Google DeepMind's recent model DiffusionGemma (DG) generates text via diffusion, meaning many diffusion steps happen before generating the final output. In particular, these diffusion steps carry vectors in addition to tokens. If we cannot interpret these tokens and vectors, the model has significant opaque serial depth, potentially harming monitorability. Recently, Engels et al. found that DG nevertheless maintains high monitorability, for instance by showing that projecting the distribut
0
3 👁
Mom's Advice For Hosting A Class Reunion
Pour more money and effort into them than you think is reasonable. Treasure them, because you can't actually host that many of them and keep expecting everyone to show up, even if they're good friends.Especially if they're good friends.We were wonderfully close friends, and I thought we'd meet up every year for the rest of our lives. They fizzled out by the fifteenth year. But the one at the tenth year mark was peak. That's because even ten years out, none of you really have money. Not real mone
0
7 👁
I'm starting a interview series of people working in Lean / formal methods / math formalization
I think the topics of discussion would be of interest to a lot of people here, so I thought I'd share the first episode:Tanner Duve is a Member of Technical Staff at Logical Intelligence working on formal verification and compilers in Lean, an open-source contributor to Mathlib and CSLib, and a former D1 football player.I sat down with him for a conversation about his work and his thoughts on the future of AI-assisted math formalization.Chapters:00:00 Intro05:12 Social aspect of formal verificat
0
5 👁
Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"
Very recently, the "Pacing the Frontier" petition was published. I want to focus on the comment from Alex Zhao, researcher at OpenAI:My opinion is that while coordination between American labs is feasible and could potentially be straightforward, the much larger and more significant threat is from a geopolitical arms race with China. Much as the Manhattan Project thought they might ignite all of the oxygen in the atmosphere but chose to run a test detonation anyways, I’m afraid that due to futur
0
3 👁
Learning new facts can change LLM behaviour
TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its belie
0
3 👁
All Utilitarians Should Be Classical Utilitarians
This is a crosspost from my blog post.I recently had the great joy of meeting a group of utilitarians, but, to my complete horror, out of the twelve of them, not a single one was a classical utilitarian.In case you don’t know, classical utilitarianism was the original form of utilitarianism promoted by John Stuart Mill and Jeremy Bentham. It is often referred to as total hedonic utilitarianism, and it holds that, most basically, when one is choosing what action to take, they always ought to alwa
0
4 👁
On Dwarkesh Patel’s Podcast With Ryan Greenblatt
Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.
The vibes have shifted, contrast this to the lit recursion when he talked to Huang
As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.
If I am quoting directly I use quote marks, otherwise assume paraphras
0
3 👁
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding.This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of AI safety research. Many of our most interesting results so far came from replicating previous results on newer or more capable models. For example, it is perhaps useful to
0
5 👁
Metaphilosophy II: Empirical Flywheels
1.4 Two philosophical methods1.4.1 Philosophy consists of updating the highest-level concepts of the mind. As discussed above, this 'updating' process can ultimately involve anything in the mind. That said, when undertaking philosophy as an intentional project, we can distinguish two methods. These are not opposed strategies, but complements; doing philosophy well may require both.1.4.1.1 Firstly, we can try to draw out connections among existing concepts, showing, e.g., that they contradict eac
0
5 👁
Red vs Blue, but for Evals
🔵 The blue team proposes an evaluation protocol for some capability/propensity of interest. This consists of a suite of measurement tasks, together with a preregistered decision-making process they will follow based on the results, e.g. what control measures they will put in place during deployment.🔴 The red team proposes a subversion strategy the untrusted model can follow to subvert the blue team’s evaluation-informed decisions, with the goal of increasing its own chance of causing a catastrop
0
4 👁
Toy Model of Activation Obfuscation
I completed this work as part of the BlueDot Impact Technical AI Safety Project. This linkpost is a somewhat condensed version of the writeup on my blog.Training against probes is considered a forbidden technique, because the model might learn to obfuscate its activations instead of behaving better. Can we create a toy example of this? More specifically: under optimization pressure, will a toy model learn to encode a feature to be challenging to detect with linear probes?In this research, I give
0
5 👁
Untie Squared ReLU variant
1 IntroMy collaborator Michael Bukatin came up with an idea: in the hidden projection of MLP, instead of using ReLU, he would dupl…
💬 0
👁 8
Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?
LessWrong · Aug 16, 2026
💬 0
👁 11
Q2.5 2026 Timelines Update: Uplift and Revenue
LessWrong · Aug 16, 2026
💬 0
👁 9
Case for Funding AI Safety in Japan
LessWrong · Aug 16, 2026
💬 0
👁 8

Will There Be an AI Hegemon? A Mental Model for AI Power Concentration
LessWrong · Aug 16, 2026
The Doomsday Argument is Reasonable and Mostly Points to Longevity
LessWrong · Aug 16, 2026
Three thoughts on civilisational handoff
LessWrong · Aug 16, 2026
Should Less Wrong add subtitles?
LessWrong · Aug 16, 2026
Are questions allowed on LessWrong?
Sometimes I want to post things like:[thing i'm wondering about][here’s my initial stab at it][but i have no idea if this is right…
💬 0
👁 4
Does DiffusionGemma do latent reasoning?
LessWrong · Aug 16, 2026
💬 0
👁 3
Mom's Advice For Hosting A Class Reunion
LessWrong · Aug 15, 2026
💬 0
👁 7
I'm starting a interview series of people working in Lean / formal methods / math formalization
LessWrong · Aug 15, 2026
💬 0
👁 5
Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"
LessWrong · Aug 15, 2026
Learning new facts can change LLM behaviour
LessWrong · Aug 15, 2026
All Utilitarians Should Be Classical Utilitarians
LessWrong · Aug 15, 2026
On Dwarkesh Patel’s Podcast With Ryan Greenblatt
LessWrong · Aug 15, 2026
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuabl…
💬 0
👁 5