I was invited to speak to the Princeton AI Alignment group last weekend about the likelihood of existential risk and various forms of misalignment. I’m not fully sure why they invited me in particular, I’m just a guy on the internet, but I think the idea was that I have takes, and they wanted some takes, and the market cleared. Despite the rather dour subject matter, it was fun! Below are some miscellaneous things that we chatted about and some thoughts swirling in my head both during and after the chat, ideally more eloquently phrased than whatever I said at the actual meetup but less eloquently phrased than an actual coherent position piece. Thanks a ton Emmett for inviting me to come out.
I.
It’s hard feeling 30 in a group of college undergraduates. It’s very tempting to make references to things like ‘remember how Facebook was in 2010,’ but doing that is always a mistake because you get hit with blank stares that just further underscore the age gap. What do you mean you don’t know what zynga is??? Still, I don’t envy the kids, they are coming of age in an extremely tumultuous period. It’s not just AI, of course. It’s…idk, everything? Have you looked around? This was the first time that I had the distinct sense that I had lived through a different era than the one we are in now. Like, just as an example, I was chatting with one of the students and it occurred to me that he simply did not remember a political cycle without Trump. He had never experienced a pro-social silicon valley,1 or known a life without ubiquitous social access to everything. Imagine how much that all must change the perception of “normal.” It made me feel my age.
That was what we spent a lot of the conversation on — what does ‘normal’ even mean anymore? In some broad sense, misalignment is deviation from a shared norm. And even though the conversation started out about AI, it very quickly veered off topic to a wide range of other things, like the state of our politics, our economy, our culture. What are our norms in all of these domains? It was a bright and beautiful day out, the kind of day that only comes in October in the North East, and the kids were talking about their p(DOOM). That was the first thing they shared, actually. Everyone went around in a circle and ranked on a scale of 1 to 10 how concerned they were about AI risk. I think it was a log scale, which made the number of 7s and 8s and 9s surprising and concerning.2
I said my own level of concern was like a 5.3 I normally am the doomer in the room, so it was an odd change of pace having to defend my position from the other side. If you want to read my more doomer takes, check here, here, and here.
II.
The most clearly articulated concern was the standard Vernor Vinge / Tim Urban / Yudkowski mix of sci-fi terror and game theory, broadly summarized by one student as: “the AI will keep getting better, and you have written that alignment seems like an impossible problem, and there are other uncontrollable possibly malicious actors who are racing as fast as possible to unleash this on the world without any safeguards. So why are you not shitting your pants and freaking the fuck out?”
Honestly, it depends on the day. There are definitely moments where I am freaking out. I think the syllogism is straightforward enough.
But I’m also ten years older than I was ten years ago, and I think one of the things that has cropped up with that additional time is a sense that there are a lot more unknowns than we give credit for, and that the world has a funny way of being homeostatic. If the AI tools continue to get recursively better, and if we do not have sufficient time to develop monitoring and security tools, and if we are unable to coordinate some kind of response, then things may get pretty hairy. But at least right now, there needs to be a bit more ink behind those assumptions.
IIa.
Take RSI, recursive self improvement. The basic idea behind RSI is that as the models get better, they will be able to continue finding more and better improvements in the process of training models, until AI begins improving itself at an exponential rate. In this model of the world, you could go from GPT-5 to GPT-50 overnight. RSI is a core assumption of the worst doomer scenarios, because it dramatically cuts our response time. Probably best explained by Tim Urban.
But this model of the world hides a lot of things. For example, it assumes that AI will get exponentially better based on linear increases in resources, even though everything we know thus far says the opposite. AI capabilities follow power laws, and models need exponentially more compute and exponentially more data in order to get linearly better.
It turns out it gets really hard to get exponentially more compute and data! You start running up against downstream things, like access to energy, and chip production capacity, and below that you run into mundane things like “where do you get your rare earth metals.” Humans are pretty damn smart, and humans with AI are even smarter, so maybe our recursively improving AIs find model architectures that are more efficient per watt, or discover ways of building chips that are easier to mass produce. But there are a limited number of such innovations, and as you start to pick low hanging fruit, progress slows down a lot. Even Moore’s law eventually ran out of fuel.
RSI assumes a positive feedback loop, but nature hates positive feedback loops. The only real example of an infinite positive feedback loop is a black hole. Every other positive feedback loop depends on fuel that runs out faster the faster things grow (nukes, population growth, etc). So in order to get a really rapid capability / intelligence explosion, the model’s rate of improvement needs to overcome the massive negative feedback loop that makes each incremental improvement significantly more difficult.
That’s not to say that AI won’t get better. But change is most disruptive when it happens very very quickly. I think we will have more time to adapt, at least more than the overnight capability explosions that would make a malicious AI really dangerous. In some sense, I think that’s the only question that matters — how much time do we have to figure out alignment at each stage of AI capability increase?4
Also, as an aside, it seems to me that the frontier labs are going broad instead of deep — in the last 6 months the models have definitely gotten better at writing code, but have gotten way way better at doing things mostly unrelated to writing code, like reading legal briefs. Some of this is downstream of the kinds of data that the labs are trying to procure. A model that knows how to do legal things probably can write code better than a model that can’t do legal things, just because of the bitter lesson and compression-as-intelligence. But some of this is because we are just running out of ways to make the models better at writing code!
IIb.
Or take the “uncontrollable, possibly malicious actors” bit. Inevitably, every doomer has some variant of this conversation:
Doomer: “AI is going to kill us all”
Layman: “just…don’t make the AI? Why is this hard?”
Doomer: “Well, we can’t just not make the AI, because if we don’t, other uncontrollable, possibly malicious actors will, and then they will <do bad things with it>.”
First, can we just appreciate how nuanced and unintuitive this argument is? I think a lot of the doomer crowd really really handwaves this part away, which makes sense because it’s easily the weakest part of the argument. It’s also worth noting that the doomer in question often works for one of the frontier labs, which, uh, really is not doing any favors. At best, it raises some questions about the doomer-ai-researcher’s general understanding of ethics (and hypocrisy).
But put all that aside.
In 2026, the term “malicious actors” basically always means the Chinese government, because there is no other geopolitical force that has the capacity to create frontier AI models. So the idea is that US labs cannot stop building frontier AI, because the Chinese labs will not stop building frontier AI. This is Anthropic’s favorite argument, which is why they mention China or the CCP 13 times in their statement on open source models.

This strikes me as rather odd!
I always thought it was silly to assume the Chinese labs would not stop / be made to stop developing frontier AI if there was enough domestic pressure. You can make many claims about the CCP, but it would be really hard to argue that they don’t care a lot about stability. AI as a technology — with all the threats of mass layoffs and job replacement, not to mention the AI girlfriend/boyfriend industry — is profoundly destabilizing. So perhaps it shouldn’t be surprising that the Chinese government has already done more to regulate their AI industry than the US!

People seem to model China as a single monolithic entity that is solely focused on the destruction of the US, more like a terror group than a modern state that is responsible for over a billion people. Which is just a bit myopic, imo.
Either AI is massively destabilizing — in which case, the Chinese government will have very similar incentives to the US labs — or it isn’t. You can’t really argue that it will be destabilizing in the US but not in China (and if that’s the case, maybe we should just let the Chinese government build the AI)!5
Ok but let’s say, just for the sake of argument, the Chinese labs don’t stop developing frontier models despite their destabilizing effects. And let’s also say that the US government does unilaterally pause model development. What then?
The maximally doomer position is hard for me to steelman, because I’m not sure I understand it completely myself. The Chinese government may get some sort of “lead,” which may translate to some kind of massive increase in power projection. One student hypothesized that the improvements in AI would allow the Chinese government to destabilize other countries through ‘super persuasion,’ and another argued that there would be massive efficiency gains across the military industrial stack that would help China produce bigger and stronger weapons, and a third said that China could use their edge to hack into and disable a wide range of US military apparatus.
Idk, maybe? I’m skeptical.
I think my problem with the first two arguments is that they both rely on some kind of RSI scenario where the US is caught completely unawares by improvements in Chinese model capacity. That seems wildly unlikely, unless you also assume the kind of really rapid exponential capability advancement we talked about in section IIa — which, of course, would be as surprising to Chinese researchers as to US ones. I think we would notice if the Chinese government suddenly had robots working the mines, or if they rapidly stood up many orders of magnitude more compute and energy. In case it’s not obvious, I’m even more skeptical of super persuasion. At least with the “industrial capacity” thing we have some idea of how that could happen. Super persuasion always seemed to me like a fundamental misunderstanding of the human condition, akin to programmers who think that law is an algorithm that can be gamed instead of a social contract that is upheld in the gaps.

Even the concept of China achieving some kind of material “lead” seems suspect to me. The AI labs are all extremely leaky. IP travels as quickly as employees can change jobs, and frontier model distillation seems to be a permanent fixture of LLM training. Together, these both ensure that any gains made on either side of the Pacific are shared within at most 6 months. What kind of lead would be so insurmountable that 6 months would be sufficient to make it permanent?
I do think the cybersecurity risk is real, and I’m also on record with concerns about giving ai systems access to bio labs. Still, I don’t think the main risk there will come from China. I’m certain that China, and America, and Russia, and several EU countries, and probably North Korea too, have the ability to do things like hack into physical infrastructure and wreak havoc. These cyber risks have existed for a long time; but so far we haven’t had mass disruption of our power grid or water supply because of people hacking into things. I’m not going to speculate why we haven’t seen indiscriminate cyber war crimes. Rather, I’ll just say that anyone that assumes AI has qualitatively new cybersecurity risks needs to grapple with all of the old cybersecurity risks. The game theory doesn’t quite click for me. This isn’t a game of Civilization.
Ok ok but let’s say for the sake of argument that there is some kind of massive recursive lead that also doesn’t immediately destabilize the Chinese government. What then?
I have yet to meet someone who can clearly articulate what, exactly, will happen. Yes, I understand that maybe the US is no longer the dominant superpower in the world, but how does that happen and what happens next? Some folks vaguely hint at a Chinese war of aggression,6 but those same people generally don’t have a good response to the current, ongoing US wars of aggression. You know, the ones that are actively using AI to target civilians. One student argued that China would use their gains in AI to corner the market by making really cheap goods and curing various diseases, but I pointed out that that seemed kinda awesome, why is this a bad thing again?
For what it’s worth, I’m a big believer in and supporter of Pax Americana. Part of my opposition to our current government is that it keeps destroying the things that make America great. But if you’re worried about the rise of China as a superpower, AI should be like number 35 on your list. And if you’re worried about AI, just push for a unilateral pause without hedging about malicious actors.
Also, as a brief aside, I think Anthropic in particular, and the AI industry in general, has blown a ton of good will over the obvious hypocrisy of claiming that AI will kill everyone but they have to build it anyway because of China, and then turning around and hooking their AI up to weapons in the US. They cannot have it both ways. Either the technology is an extremely dangerous extinction-level technology, in which case they do everything possible to AVOID military applications, or they own that it’s not world-ending and drop the posturing. I think you could say a lot of things about Google employees’ unwillingness to work for defense, but at least they are internally consistent. A child could tell you that the AI industry’s current “middle ground” is total nonsense.
III.
Underneath all the AI alignment discourse was a more vague sense of unease, something inaccurately articulated as: “What is ‘the good’? What does it mean to live a good life? Why is our society structured the way it is, where obviously bad things that many people openly say are bad keep happening?”
Tough questions. I’m just a guy from the Internet, what do I know?
We started by talking about other kinds of misaligned entities. There are a lot of parallels between unaligned AI and unaligned corporations or governments. They are all optimizers, seeking out very blunt reward instruments (loss minimization, engagement, money, votes), capable of breaking the underlying system in order to achieve their short term goals. I wrote about this extensively in Optimization Theory of Everything:
Here’s a not quite correct story of how we got here. We want to optimize for human flourishing. But we don’t know how to measure that, so instead we create a bunch of other systems that are all optimizing for different things and set them up against each other. Media was at odds with business and government, which was at odds with media and business, which was at odds with media and government. A classic Mexican standoff.
And for a while, these three things circle the drain around each other, and they all end up being directed towards the good, and things are good, and there is flourishing. And then one day we discover that these things are no longer optimizing against each other but are actually increasingly coming together. It turns out it’s easy to optimize for engagement, or capital, or votes, if you already have the other two things in hand.
…
As these optimizers all collapse into one another, they accelerate. You end up with this entity of engagement hacking and capital accumulation and political power that is singularly really really good at optimizing for these things. And we find ourselves getting further and further from actually making human flourishing any better.
The optimization theory of everything states, simply, that all of the current problems in society are downstream of optimizers that have gotten too good at optimizing, and as a result have tunnel-visioned down into some random metric that no one actually cares about and away from useful productive things that are hard to measure.
And it’s really hard to stop these optimizers. They are so good at optimizing that they have restructured, prevented, or blocked most reasonable avenues of doing so.
…
You can’t solve problems that you don’t understand. I think it’s a mistake to point to single actors, whether those are individuals, businesses, or governments, and get mad that they exist. A lot of the energy that gets spent on maligning immigrants or billionaires or whatever is totally misguided. In the aggregate, the optimizer is unthinkingly responding to incentives. Yes, our trillionaire is now a trillionaire. But if he didn’t exist, there would just be someone else.
But maybe I over emphasized the role of systems and the helplessness of the people caught up in them. Most of the students who had read the piece already seemed dismayed, like there was no way to escape the ruthless advance of optimizers seeking out their reward. And I think that’s not quite right either. It’s really important, maybe even of the utmost importance, to have aligned people.
I think that it’s become really popular to subsume individual responsibility to ‘the system.’ This is the explicit through-line of the China-baiting above. “Well, I just have to kill everyone because the incentives. It’s all about the incentives, you see, I’m not to blame for any of my actions!” The same general argument has been used to justify all sorts of atrocities. When I look at the people who have their hands on the steering wheel, I wonder. Are the titans of industry aligned? What about the folks on the hill? Do I feel safer having their hands on the wheel? It’s obviously really important to talk about systems and incentives. But you can’t stare into the abyss for too long. If you spend all day thinking about the forces that act on you, it’s easy to forget that you exert force too.
In fact, I think one of the big lessons of Trump 2 is just how much damage a single unaligned person can cause if they are really really motivated. I see echoes of that all across the country. And as a result I’ve never been more convinced of the importance of supporting honest people who are trying to do good, even if I disagree with how they intend to get there. The students really wanted to talk about systems. It was hard to impress on them that anything that isn’t a law of physics is actually just social convention. At the bottom of every decision is a person, not an incentive. Individual people making individual choices.
IV.
Like I said at the beginning, I feel for the kids. When I was growing up, I had positive role models to look up to. In first grade I wanted to be like Bill Gates, because he was trying to get billionaires to donate half their wealth. And in middle school I watched Barack Obama and John McCain, who both earnestly cared about the country and put aside political differences to remain friends long after their respective presidential campaigns. Today, the wealthy and powerful actively try to get billionaires to renounce their charitable donations. Today, the politicians make billions of dollars in backroom crypto deals and try to game elections.
Today, the kids talk about Roy Lee.
I was surprised to find that every single person in the room knew the founder of Cluely, the creator of the ‘cheat on everything’ app (no, I’m not going to link it). They don’t necessarily admire him. But they see that he has used rage bait and antisocial behavior to catapult himself to the upper echelons of Silicon Valley, raising millions of dollars in the process. They look on the grift with some mixture of disgust and admiration.
I don’t want to spend too long on Roy Lee — if I’m being honest, I never want to spend time thinking about Roy Lee, and any day that I end up having to do so is a day that is slightly worse — except to say that his behavior is a symptom of a larger abdication of moral grounding in the country. Corruption is obviously bad from first order effects alone. But there are massive unmeasured second order effects. Visible corruption weakens the entire social fabric. If good people who work hard and do the right thing consistently get screwed over by the people who play defect bot, what is the point of working hard and doing the right thing? Why not just build the next sports gambling app or stimulus feed machine?
I’m reminded of Jimmy Carter’s “Crisis of Confidence” speech. Only five years after Watergate, with an ongoing energy crisis and widespread civil unrest stemming from the Vietnam War and the assassination of major political figures, Carter got on TV and opened with the following words:
I want to talk to you right now about a fundamental threat to American democracy. I do not mean our political and civil liberties. They will endure. And I do not refer to the outward strength of America, a nation that is at peace tonight everywhere in the world, with unmatched economic power and military might. The threat is nearly invisible in ordinary ways. It is a crisis of confidence. It is a crisis that strikes at the very heart and soul and spirit of our national will. We can see this crisis in the growing doubt about the meaning of our own lives and in the loss of a unity of purpose for our nation. The erosion of our confidence in the future is threatening to destroy the social and the political fabric of America.
He said this in 1978 but it could apply today, couldn’t it? Incredibly prescient.
I’ve written in the past about nihilism, and how it seems to have gripped the youngest generation of Silicon Valley. Something that I think about a lot these days is how there is basically no shared, grounded starting point to build a communal ethic. If you’re a college student thinking about AI in 2026, you may not have any kind of framework for what ‘a good life’ even entails. You may never have even thought to ask the question!7
Nature hates a vacuum. It’s not like people just wander around without goals. I think materialism naturally fills the gaps. “I like eating and traveling and having stuff, so why not just have more of that?” Or, one step further: “I like having power, so…”
One reason I like reading religious texts is that I think they are very Lindy. There’s something that has resonated for generations, some kernel of truth that has led to their longevity, and I think people will keep reading the Bible and the Vedas long after the dissolution of Apple or Google. I think it’s interesting and important that basically every religion out there says that true happiness is found in service to others, and that materialism is a dead end. Even Randian objectivism is grounded in pushing society forward!8 Also, I’m nothing if not an empiricist, and empirically it just doesn’t really seem like the richest and wealthiest among us are living happy and fulfilling lives. Like, damn guys, you have so much money and you’re sadposting on Twitter?
On a whim, I asked the audience how many folks regularly attended some kind of religious service; only one person raised their hand. Every generation goes through its own crisis, some internal and some external. I think late Gen Z and early Gen Alpha have a tough hill to climb, because their crisis is a crisis of confidence. They need to look at the world and understand with clear eyes that there is no objective moral truth, that god is dead, that humanity may not be uniquely intelligent things in the universe. And then after understanding all of that, they need to choose to build a moral framework and live by a code of ethics anyway.
That is really, profoundly difficult.
My suggestion was to read some Nietzsche, but not too much. About 70 pages or so.
V.
To bring it back to AI, the last question the students asked was “what is the biggest problem facing AI alignment today?”
It’s the people. The biggest problem facing AI alignment is that so many of the people working on AI are themselves unaligned.
I don’t think anyone earnestly believes that the leaders of the frontier labs are all that interested in broader social good. It’s not like these guys are scions of charitable giving and public service outside of their careers. In some cases it’s explicitly the opposite — many of them have reputations for Machiavellian power seeking, while others seem to be enthralled by the idea of unilaterally destroying humanity to usher in the machine gods.
This isn’t just an optics problem, though of course it’s a massive optics problem. I also just believe that an unaligned model is much more likely to come from an unaligned person or an unaligned company. I would feel a lot better about AI alignment if, like, Fred Rogers was in charge of the leading frontier lab (or was president).
So if you’re a smart college student who ideally has read some great works of literature and is worried about AI alignment, I recommend first and foremost cultivating in yourself an understanding of what ‘the good life’ actually means. And then, you know, going to work in any of a number AI alignment orgs, reach out / dm me if you want intros.
many folks in the Bay are obviously still pro-social, but I think you would be hard pressed to find the kind of ubiquitous desire to do good that was present from 2000-2016. These days, there are more people who are far more mercurial, looking to make money without even the fig leaf of pro-social justification
Of course, this was a pretty self selected group. Maybe only 20 folks who decided to spend part of their Saturday talking to some guy on the internet about AI x-risk. So I’m not necessarily generalizing to the rest of the student population.
5 does not mean “there’s a 50% chance of a bad outcome.” I interpreted the scale to be about how much effort we were putting into this problem relative to how much I think we ought to, and it’s all vibes based anyway. Maybe 5/10 is closer to something like “when I think about AI risk, half the time I’m worried about something serious happening, and half the time I’m less worried.”
Are we currently in a fast takeoff RSI loop? No, I do not think so. Even though AI has gotten significantly better over the last 6 years, and even though on some charts that gain looks exponential, a huge chunk of that gain was from leveraging compute and data that was already lying around. The jump in capabilities from gpt-3 to gpt-4 has yet to be replicated; even though gpt-6 is very good, the delta between gpt-6 and gpt-4 is less than the delta between gpt-4 and gpt-3
One argument you could make is that the Chinese research community is underrating AI risk, but this also seems odd to me. There are thousands of people working on this problem in China, and tens of thousands more working in the CCP, and none of them have any concerns about AI? If that were true, how much should that update our priors? The only way I think this ends up being a serious concern is if you believe that 1) AI will have really dangerous goals, and 2) it will be able to successfully pretend to not be destabilizing enough in any other way such that 3) the Chinese labs don’t notice or have incentive to slow. Forget about super intelligence, that seems unlikely just for mundane economic reasons!
and fair, I could at least understand an attempt to try and take Taiwan although I’m still pretty skeptical that that would be in anyone’s interests
I don’t think it’s a coincidence that extreme utilitarianism has also become very popular. It is a functioning moral system that is pretty easy to grapple with, though it may not lead to a particularly fulfilling life and often ends up shooting itself in the foot.
Skipping ahead a bit, Rand’s whole project is trying to construct a pro-social ethic that is grounded in self interest after the failures of first religious ethics and later communist / class ethics. She doesn’t really succeed, but I think it’s very admirable that she tried, and a lot of people basically just totally misunderstand that goal as “it’s ok to be selfish.”









I agree quite strongly with your point on RSI! It has always struck me as odd that the inevitability of RSI seems to be generally well-accepted, when I think (like you said) the general pattern with engineering/optimization problems is that the more you improve, the more ridiculously hard the next step of improvement means. So we don't just need exponential capability growth, we need that to outpace the exponential growth in difficulty of inventing the N+1th AI. I don't think it's impossible that this is the case, but I haven't seen a good defence of why we should default to believing it. I think it's valid to say "well, there's still a considerable risk that the laws of intelligence-scaling make RSI possible, so we should still be cautious", but I generally don't get the impression that's the stance that AI-pilled people take. I think they treat RSI as a given and the uncertainty their predictions instead comes from uncertainty about alignment, control, etc.
I have to come back to this, it got hard to keep reading after you linked a story of 2018!Google's reluctance to aid the US military as representative of their current virtue vs. Anthropic... April 2026 kinda puts the lie to that.