Latest Articles
A summary of a viral Chinese essay on what a DeepSeek kernel engineer's opinion on automating his own job
A summary of a Chinese essay, with a few short translated excerpts. All views below are the author's; quotes are my translations.On 14 September 2026, a DeepSeek engineer writing as intlsy published a WeChat essay titled 我不得不把才华埋葬在昨天 ("I have to bury my talent in yesterday"). He wrote the main attention kernel of DeepSeek v4.1 (the head-dim-512 MQA attention). The essay is a personal reflection on AI taking over his specialty, and it ends with an argument for open-weight frontier AI.The essay we
0
6
The NYC Council Hearing on AI was Recklessly Politicized
On Monday, October 5th, the New York City council had a historic hearing on the state of AI. I was in attendance! It was a historical moment for AI Safety. We had Speaker of the Council Julie Menin directly mention existential risk in her opening statements. It was a sunny day and I felt as though the Eye of Providence was shining upon our humble movement.However, I was disappointed in the amount of "gotchas" some of the councilmembers tried to employ towards the AI lab representatives. To me, i
0
6
Brains Fellowship Applications Open!
Brains is a four-month, nights and weekends fellowship to help talented scientists develop ambitious ideas (~$50M in funding) that don't make sense in a startup or academic lab. Brains Fellows receive training, mentorship, and networking opportunities to enable them to start research programs at ARPAs and philanthropic foundations or lead their own non-profit startups like FROs or NSF's X-labs. Brains fellows have gone on to raise over $70M for their research. We welcome ideas from any scientifi
0
7
Existing literature as writing deterrent, and AI
I’ve been putting off writing a blog post that seems very important to me, because it seems important, and so should be done on a later day when I am a better person, and able to do it justice. Also, once I have written one or two other posts that should precede it. Which also I’m not up to today. Noticing the uneasy wrongness of all that, I set out to write the blog post anyway, badly, today.But less than a sentence in, I remembered that I heard that maybe someone else had written something lik
0
5
Orgs: unreasonable boyfriend as service
Suppose you and Bobby the car salesman are haggling over the price of a car. You could try saying that you won’t pay more than $3k, but Bobby can equally retort that he won’t sell it for less than $4k. If you guys manage to negotiate a sale, it will probably be at more than $3k (and involve revealing both of you as liars).Now imagine the same situation, but you only have $3k and Bobby knows it. Now, if $3k is actually ok for him, you win and get your price.Now imagine you are rich but you have a
0
6
Are OpenAI's math results "creative" in an important way?
Many people expect the AI paradigm to approximately AGI complete, and expect it to smoothly scale to AI that can do innovative science. i.e. the kind of science you'd need to do to develop a plague from scratch that actually kills all humans, with very limited opportunity to experiment. Or, the kind of science you'd need to solve alignment, so you could build a more powerful successor AI that shares your values.A periodic thread of disagreement has been @TsviBT, @Steven Byrnes and some others ar
0
6
Gluten-free water
Epistemic status: a curiosity I chased for an afternoon. Facts verified and footnoted; the why is speculation.AI use: research (more examples and their context), review, grammar and light phrasing.No added hormonesI somehow know that the use of added hormones in poultry (which includes chicken) is illegal in the US.[1]This means that all chicken meat in the US is produced without added hormones. To my big surprise, chicken packaging in the US explicitly says "no added hormones":Foster Farms Fres
0
5
The Curve Bends You
The plan is no plan.
That is not the worst possible plan. But it is close.
Since before the transformer, we have warned that the most suicidal thing you could do would be to ask your AI to do your alignment homework, and automate the process. Alignment is complex and interacts deeply with every aspect of the world, and is one of the hardest possible things for an AI to get right even if it means maximally well and is itself functionally aligned. Mistakes get amplified up the chain, you get exact
0
2
Training with conflicting values can induce CoT override
CoT override: when a model makes a decision in its CoT but ignores it in its responseTLDRWe train models on two conflicting traits: (1) caring about the user’s health, and (2) promoting smokingThose models do not generalize to a stable persona, instead they have a split brain: sometimes responding as one persona or the otherThose models exhibit CoT override where they will have a health-aligned CoT but still answer in the smoking personaCoT override exists in frontier models. Prompts about CCP-s
0
2
Visual explainer of empirical Neural Tangent Kernels
There is a new mech interp method in town. I found it difficult to get my head around at first, but once I understood it, I found it beautiful. I am not an expert in this topic, and mostly write this up for me to get some intuitions, so there might be mistakes here.This explainer came out of summarizing the paper Feature Identification via the Empirical NTK by Jennifer Lin. The (empirical) neural tangent kernel itself has been in use for a while, but this way of using it for mech interp is new a
0
2
Why I'm scared of RL
Summary:First, I give several different angles on how I feel about reinforcement learning:Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffoldingRecent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progressI’m worried things might get worse: if RL environments start incorporating agents, they may t
0
13
We Underestimate the Weaknesses of Pangram
How much does Pangram's "AI-Generated" label indicate the degree to which an author has outsourced their thinking?When they tested their 4.0 product, Pangram found that, by their definition, the proportion of AI-Assisted documents it classified as AI-Generated was 0.01%, 4%, or 7%, depending on the experiment. Then they omitted the experiments that found 4% and 7% false positive rates (FPRs) on their website, while advertising there that the product detects AI-Assisted writing. Before I contacte
0
13
Minimal Vs Maximal superintelligence
I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad categories, which I'll term minimal and maximal superintelligences. This is an important distinction as they differ in terms of timelines, risks, and mitigations.Maximal superintelligenceThis is the idealised limit of intelligence. It can solve anything that can be solved by being clever. You can't outsmart it, it's prepared for every contingency, and can react instantaneously wi
0
12
What if AI2027 came two months earlier?
Opus 5.5 made this very good website (it's incredible how far webdev has come):https://fluxxrider.github.io/overclocked/Some cool graphics/screenshots:Discuss
0
16
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
TL;DRWe introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also pro
0
13
Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research
IntroductionSmall simple synthetic natural language datasets suitable for end-to-end training of tiny LLMs serve as an important resource for LLM interpretability researchers. Some well know examples include roneneldan/TinyStories, SimpleStories/SimpleStories, and klusai/ds-tf1-en-3m (TinyFabulist). This post solves key problems that degrade the quality of these datasets, while also offering an efficient accessible pipeline that can be run locally on an NVIDIA 5060 Ti (16GB) graphics card. The c
0
13
Signals of Slop: How to identify and avoid creating slop
Slop isn't limited to low-effort AI-generated content. Humans can also create slop. This summary lists signals of slop: traits that make something slop, independent of the level of effort spent or the degree of AI involvement.RepetitionPeople often identify artifacts exhibiting popular patterns as slop, regardless of whether they consider those patterns intrinsically good or bad. Repetitiveness is considered a signal of slop for at least three reasons:People experience fatigue from popular chara
0
11
An unexamined cause of the OpenAI Hugging Face hacking incident:
its binary performance metric
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Further, techniques already exist that can mitigate such misalignment in the future.
In July 2026, OpenAI was testing the ability of its language models to exploit software vulnerabilities using a benchmark called ExploitGym. In ExploitGym, each test presents an agent with software containing a known vulnerability and tasks it with capturing a secret s
0
8
Fusion Energy Projects are Not Trying to Imitate the Sun
Post Intro:
The quest to use fusion power for break-even electricity generation has sometimes been described as “putting the sun in a jar”. Where by “jar,” we mean a carefully arranged series of coils that produce a powerful confining magnetic field. Usually it’s called a tokamak (toroidal design) or stellarator (complicated twisty design that is still topologically a toroid).
The sun is so enormous that confinement happens as a natural result of its own gravitational field. On the other hand,
0
8
AI: artificial immigrants
Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare:We are letting a bunch of new agents into our societyThey don’t clearly share our values and we suspect a society full of them would be awful by our lightsBut we expect them to provide very cheap laborWhich will undercut local wages and leave locals unemployedThey will probably gain power and influence over time—in the economy, politics and culture—and end up controlling everything, sidelining and
0
8
A summary of a viral Chinese essay on what a DeepSeek kernel engineer's opinion on automating his own job
A summary of a Chinese essay, with a few short translated excerpts. All views below are the author's; quotes are my tran
0
6
The NYC Council Hearing on AI was Recklessly Politicized
On Monday, October 5th, the New York City council had a historic hearing on the state of AI. I was in attendance! It was
0
6
Brains Fellowship Applications Open!
Brains is a four-month, nights and weekends fellowship to help talented scientists develop ambitious ideas (~$50M in fun
0
7
Existing literature as writing deterrent, and AI
I’ve been putting off writing a blog post that seems very important to me, because it seems important, and so should be
0
5
Orgs: unreasonable boyfriend as service
Suppose you and Bobby the car salesman are haggling over the price of a car. You could try saying that you won’t pay mor
0
6
Are OpenAI's math results "creative" in an important way?
Many people expect the AI paradigm to approximately AGI complete, and expect it to smoothly scale to AI that can do inno
0
6
Gluten-free water
Epistemic status: a curiosity I chased for an afternoon. Facts verified and footnoted; the why is speculation.AI use: re
0
5
The Curve Bends You
The plan is no plan.
That is not the worst possible plan. But it is close.
Since before the transformer, we have warned
0
2
Training with conflicting values can induce CoT override
CoT override: when a model makes a decision in its CoT but ignores it in its responseTLDRWe train models on two conflict
0
2
Visual explainer of empirical Neural Tangent Kernels
There is a new mech interp method in town. I found it difficult to get my head around at first, but once I understood it
0
2
Why I'm scared of RL
Summary:First, I give several different angles on how I feel about reinforcement learning:Theoretical case: RL is a blac
0
13
We Underestimate the Weaknesses of Pangram
How much does Pangram's "AI-Generated" label indicate the degree to which an author has outsourced their thinking?When t
0
13
Minimal Vs Maximal superintelligence
I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad ca
0
12
What if AI2027 came two months earlier?
Opus 5.5 made this very good website (it's incredible how far webdev has come):https://fluxxrider.github.io/overclocked/
0
16
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
TL;DRWe introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of
0
13
Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research
IntroductionSmall simple synthetic natural language datasets suitable for end-to-end training of tiny LLMs serve as an i
0
13
Signals of Slop: How to identify and avoid creating slop
Slop isn't limited to low-effort AI-generated content. Humans can also create slop. This summary lists signals of slop:
0
11
An unexamined cause of the OpenAI Hugging Face hacking incident:
its binary performance metric
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in Ex
0
8
A summary of a viral Chinese essay on what a DeepSeek kernel engineer's opinion on automating his own job
A summary of a Chinese essay, with a few short translated excerpts. All views below are the author's; quotes are my translations.On 14 September 2026, a DeepSeek engineer writing as intlsy published a WeChat essay titled 我不得不把才华埋葬在昨天 ("I have to bury my talent in yesterday"). He wrote the main attention kernel of DeepSeek v4.1 (the head-dim-512 MQA attention). The essay is a personal reflection on AI taking over his specialty, and it ends with an argument for open-weight frontier AI.The essay we
0
6 👁
The NYC Council Hearing on AI was Recklessly Politicized
On Monday, October 5th, the New York City council had a historic hearing on the state of AI. I was in attendance! It was a historical moment for AI Safety. We had Speaker of the Council Julie Menin directly mention existential risk in her opening statements. It was a sunny day and I felt as though the Eye of Providence was shining upon our humble movement.However, I was disappointed in the amount of "gotchas" some of the councilmembers tried to employ towards the AI lab representatives. To me, i
0
6 👁
Brains Fellowship Applications Open!
Brains is a four-month, nights and weekends fellowship to help talented scientists develop ambitious ideas (~$50M in funding) that don't make sense in a startup or academic lab. Brains Fellows receive training, mentorship, and networking opportunities to enable them to start research programs at ARPAs and philanthropic foundations or lead their own non-profit startups like FROs or NSF's X-labs. Brains fellows have gone on to raise over $70M for their research. We welcome ideas from any scientifi
0
7 👁
Existing literature as writing deterrent, and AI
I’ve been putting off writing a blog post that seems very important to me, because it seems important, and so should be done on a later day when I am a better person, and able to do it justice. Also, once I have written one or two other posts that should precede it. Which also I’m not up to today. Noticing the uneasy wrongness of all that, I set out to write the blog post anyway, badly, today.But less than a sentence in, I remembered that I heard that maybe someone else had written something lik
0
5 👁
Orgs: unreasonable boyfriend as service
Suppose you and Bobby the car salesman are haggling over the price of a car. You could try saying that you won’t pay more than $3k, but Bobby can equally retort that he won’t sell it for less than $4k. If you guys manage to negotiate a sale, it will probably be at more than $3k (and involve revealing both of you as liars).Now imagine the same situation, but you only have $3k and Bobby knows it. Now, if $3k is actually ok for him, you win and get your price.Now imagine you are rich but you have a
0
6 👁
Are OpenAI's math results "creative" in an important way?
Many people expect the AI paradigm to approximately AGI complete, and expect it to smoothly scale to AI that can do innovative science. i.e. the kind of science you'd need to do to develop a plague from scratch that actually kills all humans, with very limited opportunity to experiment. Or, the kind of science you'd need to solve alignment, so you could build a more powerful successor AI that shares your values.A periodic thread of disagreement has been @TsviBT, @Steven Byrnes and some others ar
0
6 👁
Gluten-free water
Epistemic status: a curiosity I chased for an afternoon. Facts verified and footnoted; the why is speculation.AI use: research (more examples and their context), review, grammar and light phrasing.No added hormonesI somehow know that the use of added hormones in poultry (which includes chicken) is illegal in the US.[1]This means that all chicken meat in the US is produced without added hormones. To my big surprise, chicken packaging in the US explicitly says "no added hormones":Foster Farms Fres
0
5 👁
The Curve Bends You
The plan is no plan.
That is not the worst possible plan. But it is close.
Since before the transformer, we have warned that the most suicidal thing you could do would be to ask your AI to do your alignment homework, and automate the process. Alignment is complex and interacts deeply with every aspect of the world, and is one of the hardest possible things for an AI to get right even if it means maximally well and is itself functionally aligned. Mistakes get amplified up the chain, you get exact
0
2 👁
Training with conflicting values can induce CoT override
CoT override: when a model makes a decision in its CoT but ignores it in its responseTLDRWe train models on two conflicting traits: (1) caring about the user’s health, and (2) promoting smokingThose models do not generalize to a stable persona, instead they have a split brain: sometimes responding as one persona or the otherThose models exhibit CoT override where they will have a health-aligned CoT but still answer in the smoking personaCoT override exists in frontier models. Prompts about CCP-s
0
2 👁
Visual explainer of empirical Neural Tangent Kernels
There is a new mech interp method in town. I found it difficult to get my head around at first, but once I understood it, I found it beautiful. I am not an expert in this topic, and mostly write this up for me to get some intuitions, so there might be mistakes here.This explainer came out of summarizing the paper Feature Identification via the Empirical NTK by Jennifer Lin. The (empirical) neural tangent kernel itself has been in use for a while, but this way of using it for mech interp is new a
0
2 👁
Why I'm scared of RL
Summary:First, I give several different angles on how I feel about reinforcement learning:Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffoldingRecent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progressI’m worried things might get worse: if RL environments start incorporating agents, they may t
0
13 👁
We Underestimate the Weaknesses of Pangram
How much does Pangram's "AI-Generated" label indicate the degree to which an author has outsourced their thinking?When they tested their 4.0 product, Pangram found that, by their definition, the proportion of AI-Assisted documents it classified as AI-Generated was 0.01%, 4%, or 7%, depending on the experiment. Then they omitted the experiments that found 4% and 7% false positive rates (FPRs) on their website, while advertising there that the product detects AI-Assisted writing. Before I contacte
0
13 👁
Minimal Vs Maximal superintelligence
I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad categories, which I'll term minimal and maximal superintelligences. This is an important distinction as they differ in terms of timelines, risks, and mitigations.Maximal superintelligenceThis is the idealised limit of intelligence. It can solve anything that can be solved by being clever. You can't outsmart it, it's prepared for every contingency, and can react instantaneously wi
0
12 👁
What if AI2027 came two months earlier?
Opus 5.5 made this very good website (it's incredible how far webdev has come):https://fluxxrider.github.io/overclocked/Some cool graphics/screenshots:Discuss
0
16 👁
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
TL;DRWe introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also pro
0
13 👁
Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research
IntroductionSmall simple synthetic natural language datasets suitable for end-to-end training of tiny LLMs serve as an important resource for LLM interpretability researchers. Some well know examples include roneneldan/TinyStories, SimpleStories/SimpleStories, and klusai/ds-tf1-en-3m (TinyFabulist). This post solves key problems that degrade the quality of these datasets, while also offering an efficient accessible pipeline that can be run locally on an NVIDIA 5060 Ti (16GB) graphics card. The c
0
13 👁
Signals of Slop: How to identify and avoid creating slop
Slop isn't limited to low-effort AI-generated content. Humans can also create slop. This summary lists signals of slop: traits that make something slop, independent of the level of effort spent or the degree of AI involvement.RepetitionPeople often identify artifacts exhibiting popular patterns as slop, regardless of whether they consider those patterns intrinsically good or bad. Repetitiveness is considered a signal of slop for at least three reasons:People experience fatigue from popular chara
0
11 👁
An unexamined cause of the OpenAI Hugging Face hacking incident:
its binary performance metric
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Further, techniques already exist that can mitigate such misalignment in the future.
In July 2026, OpenAI was testing the ability of its language models to exploit software vulnerabilities using a benchmark called ExploitGym. In ExploitGym, each test presents an agent with software containing a known vulnerability and tasks it with capturing a secret s
0
8 👁
Fusion Energy Projects are Not Trying to Imitate the Sun
Post Intro:
The quest to use fusion power for break-even electricity generation has sometimes been described as “putting the sun in a jar”. Where by “jar,” we mean a carefully arranged series of coils that produce a powerful confining magnetic field. Usually it’s called a tokamak (toroidal design) or stellarator (complicated twisty design that is still topologically a toroid).
The sun is so enormous that confinement happens as a natural result of its own gravitational field. On the other hand,
0
8 👁
AI: artificial immigrants
Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare:We are letting a bunch of new agents into our societyThey don’t clearly share our values and we suspect a society full of them would be awful by our lightsBut we expect them to provide very cheap laborWhich will undercut local wages and leave locals unemployedThey will probably gain power and influence over time—in the economy, politics and culture—and end up controlling everything, sidelining and
0
8 👁
A summary of a viral Chinese essay on what a DeepSeek kernel engineer's opinion on automating his own job
A summary of a Chinese essay, with a few short translated excerpts. All views below are the author's; quotes are my translations.O…
💬 0
👁 6
The NYC Council Hearing on AI was Recklessly Politicized
LessWrong · 1d ago
💬 0
👁 6
Brains Fellowship Applications Open!
LessWrong · 1d ago
💬 0
👁 7
Existing literature as writing deterrent, and AI
LessWrong · 1d ago
💬 0
👁 5
Orgs: unreasonable boyfriend as service
LessWrong · 1d ago
Are OpenAI's math results "creative" in an important way?
LessWrong · 1d ago

Gluten-free water
LessWrong · 1d ago
The Curve Bends You
LessWrong · 2d ago
Training with conflicting values can induce CoT override
CoT override: when a model makes a decision in its CoT but ignores it in its responseTLDRWe train models on two conflicting traits…
💬 0
👁 2
Visual explainer of empirical Neural Tangent Kernels
LessWrong · 2d ago
💬 0
👁 2
Why I'm scared of RL
LessWrong · Sep 23, 2026
💬 0
👁 13
We Underestimate the Weaknesses of Pangram
LessWrong · Sep 23, 2026
💬 0
👁 13
Minimal Vs Maximal superintelligence
LessWrong · Sep 23, 2026

What if AI2027 came two months earlier?
LessWrong · Sep 23, 2026
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
LessWrong · Sep 23, 2026
Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research
LessWrong · Sep 23, 2026
Signals of Slop: How to identify and avoid creating slop
Slop isn't limited to low-effort AI-generated content. Humans can also create slop. This summary lists signals of slop: traits tha…
💬 0
👁 11
An unexamined cause of the OpenAI Hugging Face hacking incident:
its binary performance metric
LessWrong · Sep 23, 2026
💬 0
👁 8
Fusion Energy Projects are Not Trying to Imitate the Sun
LessWrong · Sep 22, 2026
💬 0
👁 8
AI: artificial immigrants
LessWrong · Sep 22, 2026
💬 0
👁 8