Every passing week the OpenAI / Huggingface hack gets worse. In case you are not up to speed, you should read Dwarkesh, Ajeya, the most recent METR/Redwood research report, Zvi, and Scott, probably roughly in that order.
You can also read my previous coverage here:
I don’t think I have much to add to the actual descriptions of what happened. There is an ongoing debate about whether these descriptions are too anthropomorphized. My response to that inane debate is basically summed up by the note below:
Instead, I want to focus on what to do about it. I have two big concerns.
First, we need to seriously consider that alignment may not be possible. I don’t mean ‘not possible with our current technology’ or ‘not possible in a reasonable timespan.’ I mean, ‘this technology will never be aligned.’ There are a few reasons why this may be true:
Even if you can train alignment into a model, you have to start with some model that exists before the aligned state. This previous model is, tautologically, unaligned.
You could try to align the model as you train it, such that you never have a model that is unaligned. But you would then also lose visibility on the effects of your alignment guardrails. That is, you have no idea if your alignment tactics are actually doing anything at all.
Alignment training is a moving target. Each new capability that a model acquires opens up a vast space of possible malicious behavior. These compound exponentially (e.g. reading private data is fine, searching the web is fine, and then you accidentally get your private data leaked to the web).
The above points all assume that alignment is a coherent concept. I am deeply skeptical. No two people agree on what ‘alignment’ means.
It seems wrong to think of alignment from the perspective of a single model. It seems more correct to think of alignment from the perspective of a single session. Millions of sessions are being run constantly; we are rolling the dice on alignment for each one.
Even if we could resolve all of the above criteria, we have a deeper problem: optimization is a fundamentally flawed way to implement alignment. There is no way to optimize for alignment. You can only optimize for the appearance of alignment. These are very much not the same thing! Yes, sometimes a thing will look aligned because it is aligned, but other times it will look aligned because it is lying to you. Just ask any parent who’s had a teen sneak out at night.
I’m going to call this the Camouflage Principle: any criteria that aims to prevent X will also optimize for simply hiding X. This is a more specific variant of Goodhart’s Law — when a measure becomes a target it stops being a good measure — which is also the source of all of every modern human problem.
The Camouflage Principle has a direct impact on neural network training: it limits what parts of the model you can actually train. Back in 2025, researchers tried to optimize the chain of thought reasoning traces of various models to get them to be more aligned. What they found:
While we show that integrating CoT monitors into the reinforcement learning reward can indeed produce more capable and more aligned agents in the low optimization regime, we find that with too much optimization, agents learn obfuscated reward hacking, hiding their intent within the CoT while still exhibiting a significant rate of reward hacking.
The models did not stop reward hacking. They just made it invisible to the optimizer. This was, technically, the optimal thing to do. Which in turn is directly downstream of the fact that we are forced to optimize against proxies for the things we actually want.
Once you hear about the Camouflage Principle you see it everywhere. We already talked about misbehaving kids. It also applies to things like criminals moving to encrypted channels once it became clear that telecoms were being watched, or bribing government officials to pass things like health and fire safety inspections. In all of these cases, it is easier to fake alignment rather than simply be aligned. So there is a corollary to the Camouflage Principle: optimization works if and only if the intended way to optimize for a goal is the easiest way to optimize for a goal. Are we sure that the easiest way to seem aligned is to be aligned? Seems unlikely!
Current approaches to get at aligned models are woefully inadequate. Alignment researchers need to move away from optimization/back prop based approaches to alignment. The one potential shining light in the space is the work on interpretability, which has managed to develop techniques that let us peer into the ‘inside’ of a model and manipulate them in consistent ways. But the problems with AI alignment are not unique to AI. We cannot align businesses, or governments, or even individual people. And there, we have no problem at all peering ‘inside’ the optimizers. That does not generally stop them from grinding away.
Scott Alexander recently wrote “Ethics that persist beyond the moment aren’t impossible. They’re just hard.” He was talking about people, and applying the same intuitions to models. Unfortunately, I think it is wrong to make the comparison to people, and it is wrong to believe that ethics can persist the way he says they can. First, because people are less like individual models and more like individual sessions of a model (see above). And second, because whether we view people whose ‘ethics persist’ as aligned or unaligned depends entirely on when and where you sit in history. In 2026, we would not look upon those who had their 15th century ethics “persist” very positively. And I suspect the Russians and Ukrainians may have different opinions on soldiers whose ethics “persist” in favor of their homelands.
All this to say, my thoughts on optimization over the last few months were already pushing me towards ‘we should really stop developing this stuff until we have a better handle of what it can do.’ The huggingface hack really just pushed me over the line.
So where does that leave us?
This brings me to my second big concern: who is ‘us’ anyway? Who is the ‘we’ in “We need to seriously consider that alignment may not be possible.”? Sam and Dario? AI researchers? Silicon Valley? The US Government? A ragtag bunch of substack bloggers and lesswrong posters, half of whom are making money from one of the big labs?
There were like ~200 people who showed up for the SF ‘pause AI’ march, which is admirable but also nowhere near enough to enact any kind of social change at all. For reference, ~200 people showed up to protest against an Australian law against cursing, ~500 people showed up outside Parliament to protest a ban on laughing gas, and ~700 people showed up to protest changing a street name. They even held a mock funeral for the last one.
You get more protest activity for spontaneous changes to road signage than you do pausing AI. There is essentially no organized political lobby here. You would think that the METR report would be a bombshell, an event where everyone could put aside their differences and rally and organize. Instead, the more we learn about the huggingface hack, the more time we spend debating whether we have appropriately anthropomorphized the models. Ironically, the totally decentralized populist hatred of datacenters is probably doing more to slow down AI development than any number of think pieces from people who know what they are talking about!
Part of why it is so difficult to organize meaningful action is because we’re in this really terrible spot where everyone’s short term incentives revolve around continually expanding model capabilities and training / inference cluster size.
Very roughly, OpenAI and Anthropic are on the hook for a combined ~$200b in compute capacity in 2027 that they have to pay for regardless of whether they have the capacity to use it. Right now their revenue is ~$100b combined. That means their existence is predicated on doubling their revenue next year. But user payment is pretty swingy. More users and enterprises are investing in model-agnostic infrastructure, while open source models are becoming more and more useful. In other words, it will be very hard for OpenAI and Anthropic to keep their revenue growth up if they fall out of the frontier. This whole hugginface hack is downstream of exactly that — OpenAI was training a new frontier model, and they did not stop even after the hack.
There’s a lot of circular financing around these two companies right now. All of the hyperscalers marked the compute commitments as revenue, even though the money hasn’t actually landed yet. So if, for whatever reason, the labs don’t grow into their commitments as fast as they say they will, everyone downstream will be left holding the bag. And that’s before you get into the various investment schemes where everyone and their mother seems to be holding some amount of OpenAI stock. Including, apparently, the US Government!
Between all of the other stuff going on in the world economy (oil shocks, trade wars, etc. etc.), if the government was to come in and request / demand a pause, I think it would pretty much immediately cause a lot of problems.
So the model companies can’t stop because it is existential for them. The federal government can’t make them stop because it will trigger a pretty significant economic downturn. Who’s left? The individual AI researchers? The ones that stand to become fabulously wealthy if their various companies could just IPO?
I’m just not seeing how we get out of this incentive spiral.
But I hate leaving things dangling, so let me take a stab anyway.
Looking at the state of play, I think the best plausible path forward to achieving a slowdown or pause in AI development is something like ‘leverage populist anger at AI to score real technocratic policy wins around regulating AI and potentially pausing its development.’ Our democratic process / optimization machine is already moving in this direction. Individual politicians, especially those in tighter races, are increasingly becoming anti-AI to pick up more votes.
I think that people who are worried about AI in any of its forms — whether due to existential risk, fear of misalignment, or just general dislike of Silicon Valley tech companies — should try and help that along. That means:
More evocative populist messaging, especially around AI risk, to make this a bigger issue for November (combined with the usual voting slates and so on);
Social pressure to encourage AI researchers to release more information about the hacks, including outside of OpenAI (we still know virtually nothing about the corresponding breakouts at Anthropic, Meta, and Moonshot). links Give anyone internal to these companies a platform and anonymity if they want.
Legal pressure on California regulators to require more transparency, including way more third party audits.
Things that won’t help: getting into really nuanced technical arguments about the nature of AI, or debating whether AI will kill us because it’s more like Terminator or more like Robocop or more like King Kong or more like Lennie (from Of Mice and Men), or getting up in arms that other people who dislike AI dislike it for the wrong reasons.
Look, I recognize the irony of using a blog post to shout into the void about how people are shouting into the void too much. I don’t expect everyone to just. But if you’re reading this post and are about to get into a big argument about anthropomorphization or whatever else, and you reasonably think that both you and the person across from you are upset about this AI hack and want AI development to stop even if for different reasons,1 consider instead putting down the Internet pitchforks and agreeing to both message your state reps instead.
Other things
OpenAI decided to remove its models from Cursor.
Today, we notified SpaceX that we intend to wind down our contract providing OpenAI models to Cursor, with a proposed shutoff date of November 12, 2026. To maximize the time that developers can retain access to our models through Cursor, we are giving the maximum notice provided by our contract.
…
We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk’s companies violating contracts.
I think the moment Cursor got bought by SpaceX, it was inevitably going to become a Grok only platform — either SpaceX was going to make them Grok only, or everyone else would force them to go that way. I think for downstream consumers, it is all the more important to invest in cross provider infrastructure.
Speaking of HuggingFace, NVIDIA in talks to acquire HuggingFace for a cool 13b.
Nvidia has agreed to buy Hugging Face for $12.9 billion, The Information reported Wednesday night, citing a source familiar with the matter. Business Insider, which first reported over the weekend that Hugging Face was fielding takeover interest, reported Wednesday night that the talks — which would value the company at more than $13 billion — had not yet produced a signed agreement and could still atomize.
It’s an interesting acquisition. NVIDIA benefits a lot from competition at the token provider layer. The more token providers, the more necessary compute, the more Nvidia chips get sold. So buying the largest platform for hosting open source models is a good way to shore that up. TechCrunch agrees:
Why would Nvidia want that? Most obviously, it comes down to protecting its dominance in AI chips, which, from the outside at least, appears increasingly at risk, even with Nvidia’s aggressive chip-release schedule. Pretty much all of the biggest closed source AI labs (OpenAI, Google, Amazon, and Anthropic) are now in the process of building their own AI chips to lessen their reliance on Nvidia. A thriving ecosystem of open source AI models gives customers more alternatives to those closed labs, which in turn keeps more of the market dependent on Nvidia’s hardware. That’s also why Nvidia has already poured tens of billions of dollars into building its own open source AI models.
And 13b is basically nothing for NVIDIA.
On the subject of interesting acquisitions, Stripe acquires OpenRouter.
Stripe Inc. has finalized an agreement to acquire OpenRouter Inc., a startup that helps companies switch between artificial intelligence models, for more than $7 billion, according to people familiar with the matter.
The deal, just months after OpenRouter raised money at a reported $1.3 billion valuation, underscores the demand from businesses to find the most cost-friendly AI solutions. It could also give Stripe, a payments processing firm, a stronger footing in the fast-growing artificial intelligence sector.
Seems like a good deal.
A judge ruled that the government’s attempt to blacklist Anthropic back in March was illegal.
Defendants have now submitted the administrative record justifying those actions. The record is slim. A four-page memorandum, which post dates two of the three challenged actions, provides the entirety of the government’s rationale. Defendants have now backed away from the thrust of their risk assessment, which relied on Anthropic having backdoor access to its technology once deployed in a national security system. It is now clear that Anthropic undisputedly lacks any such access and that, as Defendants concede, Anthropic’s technology is itself no riskier to the national security than any other “black box” artificial intelligence model
It’s ~September. I think most of the damage was already done? Like, the only reason we are getting this ruling is because Anthropic survived. There is absolutely a world where they don’t, in which case it doesn’t really matter what the judge rules does it?
Meta settles a lawsuit centered on whether it intentionally made its products harmful to children to the tune of $16b.
Meta Platforms (META.O), agreed to pay a maximum $16.68 billion and make major changes to Facebook and Instagram to resolve claims by states across the U.S. that the company designed those platforms to addict children, misled consumers about their safety and improperly collected children’s personal data.
The settlement resolves claims brought by 29 U.S. states, and will end a federal trial that had been one of the highest-profile tests yet of allegations that social media companies harmed young users.
Meta will impose daily usage limits and restrict nighttime usage by children who use Facebook and Instagram, and enhance measures to prevent children from accessing age-restricted content.
The Menlo Park, California-based company denied wrongdoing in agreeing to settle.
I previously wrote about this kind of lawsuit here. I expect there will be more of this sort of thing.
Independent researcher Mithil Vakde gets an incredible 44% on ARC-AGI-1 with just 67 cents of compute. His strategy is pretty straightforward: train a transformer to predict input/output pairs and have each task also have a shared embedding. That’s ~basically it. The model has to pick up a compressed representation of the grids (because it is auto-regressively learning how to output the sequence of grids) and compressed representations of the tasks (because it is backpropagating information to the shared embeddings that are prepended to each task input). In general, all of the ARC approaches center on compression in some way. To get a better sense of why, check out some of the reviews I did in my Ilya’s Papers series: minimum description length, VLAEs, and Complextropy.
Clifford posted his skill file for generating isometric views to help with understanding how a really complicated codebase works. Here are some visuals of a k8s codebase.
I’m in love.
Check out my talk on Software Factories at Agentics NYC!
On the flip side, if you are arguing with someone who thinks there should be no slow down, pause, or regulation whatsoever, idk, maybe a few thousand words on the internet will change their mind.









