Tech Things: We need to talk about the open model situation
Kimi K3 changes some critical assumptions in the tech industry and meaningfully challenges US dominance in AI
I.
Every day the world seems a smaller place.
When I was younger, it felt like there was the US, which to me was basically just New Jersey and New York and Maryland. And there were lots of other places that were not the US. I would visit India a lot, for example, and India was very much not the US. There were lots of differences. There were cows just wandering around. And the AC and electricity were not very consistent. Most of all, the thing that I remembered the most was that the games all sucked ass.
When I visited in 2003 or 2004, my cousins were super excited to show me Prince of Persia. The original Prince of Persia is a game that came out in 1989. It looked like this.
We also played Disney’s Aladdin, which was made for the Sega Genesis in 1993.
And Duck Hunt, which was made for the NES in 1984.
Now that I am adult, I can appreciate these as nostalgic classics. But as an 8 year old, I was extremely unimpressed. How could I be impressed? I had a gamecube back home and my friends had an xbox, so while my cousins were playing Prince of Persia (1989) I was playing games like Halo, Smash Melee, Wind Waker, Metroid Prime, and Half Life 2.
Halo is not just better than Duck Hunt. It is barely recognizable as being part of the same genre. Duck Hunt and Halo are both games in the way that a gas station twinkie and a charcoal-grilled NYC Porterhouse steak are both ‘just’ food. Orders of magnitude difference in quality, enough so as to feel like a completely different category.
This is broadly how I felt about the US vs everywhere else, at least when it came to tech. For the first 25 years of my life, the US was just obviously living 10-15 years in the future compared to the entire rest of the world. We had Halo, and everyone else had Duck Hunt. (Yes I am aware that Japan made half the games on my list above. I was 8, Japan had not yet entered my consciousness even though I loved Pokemon and YuGiOh and Nintendo in general). That gap felt insurmountable, and it made the world feel big.
Well, the world is a smaller place now.
II.
Things in AI change so rapidly that you’d think nothing would surprise us at this point, but this past week the AI world was rocked by the release of Kimi K3, a new open weight model coming from the Moonshot lab in China. This is the world’s biggest open weight model, clocking in at 2.8T parameters. It is also, by every account, very very good.
I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart. Same tasks, same quality of output, and near identical token counts to get there. I expected an open model to be sloppier or to grind through more tokens on the way to the same answer, and neither turned out to be true. — The Kimi K3 Moment
On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5…Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers — Artificial Analysis
On the first try, Kimi K3 just found the source of a bug that Fable 5 hasn’t been able to pinpoint in multiple attempts. It’s just one anecdote, and I haven’t used K3 much yet, but so far it’s looking extremely promising. — HN
Axios maybe put it the most bluntly with this headline:
Perhaps the most impressive demonstration to me was this incredibly detailed replica of the entire MacOS operating system. Supposedly some guy had a swarm of Kimi instances making this over the course of a few days. It’s kinda incredible, you can go in and edit files, use a fully functioning terminal, listen to music, check emails. You can go to the ‘App Store’, ‘download’ the ‘Chess app’, and then play chess (no quotes around that one, you can actually do that!)
One popular way to conceptualize these models is that they are released as part of coherent generations, which are defined by the time of release and the capability of the model. If you plot models along these two axes, you get some clear clusters along the frontier.


There are a few trends that are worth pointing out.
First, even though progress in AI is already progressing at an incredible clip, somehow things are still accelerating. New model generations seem to be releasing at a faster and faster pace, with the time between releases compressing.
Second, the rate of acceleration is not consistent across organizations. The frontier continues to be pushed forward, but the labs that are at that frontier continue to shift. This is most obvious when you look at the trajectories of Copilot — once the only competitor to ChatGPT, now widely derided — or the Meta series of open source models, which were surpassed by the Chinese models a year and a half ago and never really recovered.
Third, the rate of acceleration is not consistent across countries. In relative terms, the US continues to be in the lead when it comes to raw model capabilities. The French (Mistral) were at one point quite competitive but have since fallen off. And, most relevant to this piece, the Chinese labs have continued to forge ahead, rapidly closing the gap against the top US labs.
Despite the above, up until last week, it was somewhat unfathomable to believe that anyone would have a better model than the top US labs. Sure, some folks were saying that certain policy decisions may result in the Chinese labs taking the lead, but this was strictly in the realm of the hypothetical. There’s an old saying from Upton Sinclair, “It is difficult to get a man to understand something, when his salary depends upon his not understanding it.” I think the collective psychology of the Bay and perhaps the entire country has been one of arrogant self-assuredness, that the US will always have the best models and that this is simply a fact of life, partially because that is what is necessary to justify the incredible capital expenditure on these things.
Kimi K3 blows all of this out of the water.
It was only 6 weeks ago that the US federal government was banning ‘Mythos-class’ models for being too dangerous. Now, we have a ‘Mythos-class’ model that is not just outside of US control, but is in fact open source and available to anyone in the entire world to download and use. To add insult to injury, Kimi K3 is significantly cheaper to run. Hosted Kimi models cost a fraction of the equivalents from Anthropic and OpenAI ($3 vs $10 and $5 per million). Meanwhile, you can get a lot of juice with the ~$40 kimi subscription, way more than anything Anthropic or OpenAI are offering in the same tier. The max kimi subscription of $100 per month is about equivalent to a $200 Claude Max sub. In a world where people are becoming ever more cost conscious around their token spend, the Chinese labs are less ‘also ran’ and more ‘default option.’ And the Chinese labs are fighting with a handicap! They are running on significantly worse hardware due to ongoing export restrictions of high end silicon chips!
The markets are reacting with surprise. Kimi’s release triggered sell-offs of Google and SpaceX, in addition to the wider market of AI stack stocks. I guess I don’t know why they’re surprised, in some sense this is a long time coming. Do you all remember Deepseek? I wrote about it a year and a half ago when it came out, and my analysis then wasn’t so different from my analysis now:
To be clear, the fact that Deepseek exists isn’t really that significant. Everyone always knew there was going to be a big Chinese model. The country is not afraid to wall off it’s populace from the West; it was never going to allow LLMs that are happy to tell people about Tienanmen Square and Winnie the Poo and Taiwan. So when the first Chinese LLMs came on the market, everyone was like, “Yea, whatever”. The first batch, the Qwen models from Alibaba, were, like, fine.
Even though I expected a set of LLMs to arise thanks to protectionism and state-interest, I (and everyone else) assumed those models were just going to be worse than the US ones. This is how technical development has always been, after all. Google Search is good, Baidu is eh, and if you can use Google over Baidu you do. Amazon is good, Alibaba is eh, and if you can use Amazon over Alibaba you do. Facebook is good, Tiktok is eh, and if you can…no wait that doesn’t work does it?
Anyway, the larger point is that no one really thought the Chinese models were a threat. Sure, people would talk about how the Chinese government was a threat, but it was always in a hypothetical way, mostly used to justify infinite capitalist investment without any corresponding concerns about AI safety or alignment. I don’t think most people actually thought that a Chinese company would come out and deploy a model that is simply better than what we have in the states.
Clearly we didn’t actually learn anything from TikTok.
And we still aren’t learning. Even now, some people are clearly just in denial. I’ve heard several folks say some variant of “it’s easy to catch up when you’re just stealing / distilling from the bigger labs,” (more on distillation below) as if this somehow negates the fact that one of the best models on the market is not US-made.1 The conventional wisdom is no longer relevant. There are no secret models hiding in the backrooms of OpenAI / Anthropic. The Chinese models are no longer 6 months behind, they are at par.
Kimi’s release has had me rethinking some of my previous position on LLM economics. A month before Deepseek was released, I argued that the LLM market is a bit like the search market in that there are clear winner-take-all dynamics.
LLMs are pretty easy to make, lots of people know how to do it — you learn how in any CS program worth a damn. But there are massive economies of scale (GPUs, data access) that make it hard for newcomers to compete, and using an LLM is effectively free so consumers have no stickiness and will always go for the best option. You may eventually see one or two niche LLM providers, like our LexusNexus above. But for the average person these don’t matter at all; the big money is in becoming the LLM layer of the Internet.
…
The economics of LLMs means that it is critical for these players to have the best models. There’s no room for second place.
A year and a half later, it’s clear that I was wrong about a few things. First, I was wrong that the LLMs are effectively free. They aren’t, as we saw with the rise and fall of tokenmaxxing. The cost per unit intelligence has dropped precipitously, but the overall cost per frontier token has skyrocketed. Because I was wrong about the pricing, I was also wrong about the quality / cost curve. There is room in the market for multiple models at different price points, because consumers can in theory choose different providers for different tasks. The rise of model routing as a service and the increasing classification of tasks and roles into different AI usage categories — sales gets a cheap open source model, engineering gets fable — is downstream of that demand.
But I was really right about the most important bit, which is that consumers have no stickiness at all.
Every three weeks we see a mass exodus from Anthropic to OpenAI to Anthropic and back to OpenAI. It is just way too easy to switch providers. “Model wrapper”, once seen as derogatory, is now touted as a core feature. Businesses and consumers are recognizing both the churn and the need to stay on top, and are explicitly investing in tools that let them adapt. This is also a huge part of why the background agent infrastructure that we build and sell at Nori has zero lock-in at the compute, model, and harness levels — there is no reason to be locked into a single ecosystem when you don’t have to be (if you’re looking for background agents / cloud agents for your team, shoot me a message!)
The moment the Chinese models hit the actual frontier, there will be a mass exodus to using those models. The economic incentives are way too strong, and arguably that is already happening for people who are in the know.
III.
This is not how any of this was supposed to go. The whole point of all this closed model stuff is the ability to dig a trench around the model and put a toll booth in the middle, and the whole point of all the debt is to basically have a call option on that toll booth. We all knew that the big labs would be fighting tooth and nail with each other over which of their closed models would end up winning the day. Doing this open source stuff feels against the rules.
Partially, that’s because it is against the rules. Many of the open source models are trained on their more capable closed source counterparts.
Unfortunately, usage of the models is a source of valuable training data for competitors. Every improvement to any model can quickly be copied by other labs even without seeing any of the internals because you can just sample the new model a bajillion times and use the outputs as training data. This technique is called ‘distillation’ because you are ‘distilling’ the essence of some larger ‘teacher’ model into a smaller ‘student’.
OpenAI and Anthropic can try to ban people from training on the outputs of their models, they can kick and whine and sue to try and enforce it, but at the end of the day the labs need people to use their models. That’s the foundation of all of the token economics! If no one actually uses your models, what’s the point of all that training?
So the token traces have to get out in the wild, and any kind of adversarial cat and mouse is going to end up wasting a lot of resources for very little gain. The open source models will basically always be able to catch up, modulo some amount of compute.

This isn’t a new problem, people have been talking about it for some time. Dwarkesh even asked folks to tackle this problem in his essay competition last month:
So when does the profit start? Maybe at some point scaling will plateau, but if progress at the frontier has slowed down, then the combination of distillation and low switching costs (cloud margins result from high switching costs) makes it really easy for open source to catch up to the labs, eating into their margins. So how do the labs actually start making money?
(emphasis mine)
My answer was that the labs in totality will not be able to meet the demand for their models and will be capped by compute, which in turn will prevent new-comers from actually competing for best in class model training and will let some of the labs rent seek. The winning answer argued that the models will become commodity and the model providers should try and own downstream services instead. Note that both of these answers assume that distillation is so inevitable, it’s hardly worth discussing.
This has already had some amount of impact on OpenAI and Anthropic’s pricing models. Anthropic, for example, continually extends the amount of time that Fable will remain on its subscription plans. This is not entirely attributable to the open source models, but I am certain that these models will add additional pressure to simply keep Fable on the subscriptions forever — even if doing so results in a loss for Anthropic, since the subscriptions are heavily subsidized.
IV.
Other, smarter, more plugged-in people can speculate about what this means for future AI policy. I’m sure much will be made of CCP leader Xi Jinping’s speech on AI regulation and the importance of collaboration across nations, and Demis Hassibis (CEO of Deepmind’s) call for domestic regulatory apparatus. I will leave the policy wonks to it.
But I have strong opinions about the obvious and overwhelmingly negative impact of current and past AI policy. Bluntly, I’m pretty ticked off.
First, quoting liberally from Stephen Bochinski:
Step back and the bigger story is what an unmitigated failure US AI policy has been. The administration held Fable back, and what finally shipped is a hindered version that refuses whole categories of work. Meanwhile a frontier quality model with none of those restrictions is a download away, released by a Chinese lab the US government has no ability to regulate. Whatever the theory behind gating American models was, it plainly wasn’t thought through, because the only people the gates constrain are American customers. Semgrep found GLM 5.2 beating Claude on their cyber benchmarks for exactly this reason. The restricted model declines the work and the open one just does it.
I’ll add more.
The overwhelming corruption of this government has direct impacts on our ability to produce great technology. Friends of mine in government and defense — which AI is increasingly a part of, whether we want it to be or not — have told me that projects (including datacenter projects) simply do not happen unless the Trump family somehow stands to benefit. They will get slow rolled or stall under regulatory scrutiny, and then suddenly clear up when one of the Trump kids happens to join the company board or take an investor position. Companies that do not play ball end up feeling the full force of the state, c.f. the very public attempt at corporate murder against Anthropic.
The reason businesses file in Delaware is not because Delaware is particularly business friendly. It is because Delaware is predictable and stable. To a first approximation, you know what the rules are, and if you don’t follow them that’s mostly on you. Putting ethics aside, the biggest practical issue with corruption and cronyism and irrational application of state power is that it chills investment. As I said when we discussed the whole Fable situation:
The debt, the build out, the datacenters, the stock market runs on every part of the AI chain from GPUs to memory to disk to server racks. Alllllll of that is predicated on the idea that all of this is going to be worth trillions and trillions of dollars. And by all accounts, it seems like it is. Or at least, was on track to be. You know what will put a spanner in the build out of a multi-trillion dollar data center investment? The realization that at any point, the government will unilaterally cut off access to everyone, and the datacenters will be worth squat.
The US is no longer a stable business environment. It doesn’t matter if you follow the rules, one Friday evening (it’s always a Friday evening) you may wake up to Pete Hegseth demanding your destruction. Ironically, this administration behaves more like the CCP, where the government picks winners and losers and can simply take out a company for saying or doing the wrong thing.
This administration’s immigration policy and university funding policy have also been unmitigated failures, and in my mind directly contribute to the situation we are in today. Our research institutions, which depend on foreign brilliance and government funding, are in crisis and have been since Trump returned to office. Every day, brilliant talent leaves our country for foreign shores, or never arrives to begin with. The imbeciles with the blue check marks on Twitter will tell you that this is a resounding achievement; unfortunately not a single one of them knows what a GPU even is.
We are still hacking away at our research pipeline, by the way, like a deranged ferret trying to gnaw off its own arm. PhDs and labs across the country are still being denied funding, even as other countries go out of their way to fill in the gap. It’s been a few news cycles so we’ve forgotten the assaults on college campuses. But the folks who have had their entire lives overturned certainly haven’t forgotten.
A certain brand of Bay area resident likes to talk about moats and structural advantage. To use those terms, the United States has a massive structural advantage by being an immigrant-friendly, non-ethno-nationalist country. There is no other country in the world that is as diverse as ours. There is no other country in the world that could restructure to be as diverse as ours. That means that the US and the US alone can tap talent from all over. As the old Reagan line goes, “You can go to live in France, but you cannot become a Frenchman. You can go to live in Germany or Turkey or Japan, but you cannot become a German, a Turk, or a Japanese. But anyone, from any corner of the Earth, can come to live in America and become an American.” It is our defining moat. It is why we punch way above our weight despite being 3x smaller by population vs equivalents in Asia.
This administration is simply throwing that away.
When we talk of accountability, I personally want to know how many folks who ended up at Moonshot were educated in the States — the founder had a degree from CMU! — and why we weren’t able to convince them to stay stateside or pull them from China in the first place. People will chalk up the successes of the Chinese labs to mystery, but in my opinion, there is nothing mysterious about it. Choosing to ‘win’ is rarely a straightforward policy choice, but you have to at least try. It seems pretty obvious that the people in this administration are simply not interested in doing so.
V.
I don’t mean to take away from Moonshot’s accomplishments. Politics aside, Kimi K3 is an incredible technical achievement on its own terms. The hardware limitations have forced the Kimi researchers to pull some pretty neat tricks, all of which are now open source and can be adopted across the ecosystem. Kimi also forces both OpenAI and Anthropic to keep their prices down, which I (as a startup founder) am a huge fan of.
The model weights themselves have not yet been released — they are slated to drop on the 27th. When they do, I will be very eager to pull Kimi into our background agent environments and take it for a spin.
The US is no longer the only nation capable of frontier intelligence, no longer in a category of its own. The world is a smaller place now.
Other Things
GPT 5.6 Sol has done an incredible job thus far solving open math problems and verifying them with Lean. Here’s an example of GPT 5.6 solving an open problem in Convex Optimization. Here’s a swarm of GPT 5.6 instances solving ~20 open Erdos problems. Here’s GPT 5.6 proving the cycle double conjecture in graph theory. These discoveries all happened in the span of like two weeks, in some cases by amateur mathematicians.
I’ve often said that all of mathematics is really a search problem. You have a bunch of axioms and a proof is a specific combination of those axioms, and you search over them to discover novel proofs. Any search problem can be naively brute forced, but its often just too expensive to really do. Brilliant mathematicians are really good at narrowing down search space, that’s what makes them brilliant.
But another way to make search problems easier is to just bring down the computational cost of exploring a particular part of the search tree. Which AI is really really good at.
Quoting from the first link:Lastly, some important comments about the work relating to AI capabilities: In a lot of cases, proving lower bounds like this result relies on finding that right construction that works (in this case, family of difficult functions and a strategy for how an “adversarial” oracle should answer queries from an algorithm to reveal minimal information) and then proving things about it. There are only so many function classes which would be reasonable to look at (here, quadratics for example would have also been reasonable with order d² degrees of freedom, or any variation of maxes of some simpler families of convex functions as well), but the actual proof mechanics once the “correct” function class and correct strategy for adversarial oracle answers is found are often not so complicated, and often employ existing results from convex geometry or similar (this is also the structure of two previous but much more niche, less important results of mine). So I wouldn’t really say that this result is using or creating some fundamentally new techniques in convex geometry or optimization theory. What this means from my perspective is that if a result is attainable with existing techniques, modern AI methods will be able to solve those problems. I don’t think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We’ll be needed for problems where actual novel approaches are needed.
For now.
Thinking Machines releases their own open weights model, which may be the best one from the US.
Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that people can make it their own.
Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active. It supports a context window of up to 1M tokens. It was pretrained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency.
It’s an interesting bet from one of the most put-upon startups in the country. As a reminder, Thinking Machines (helmed by Mira Murati, ex-CTO of OpenAI) raised billions of dollars without having any product or model or anything, and then proceeded to be totally shrouded in mystery for two years. They eventually released a model fine tuning platform, which was widely considered to be insufficient to justify the investment. Among the ‘neolabs’, Thinking Machines is making the most money (I think $200m last I checked). The rest of the labs are doing even less than that, which I think has really soured investors on the whole concept of a neo lab in the first place.
With Inkling, I think the idea here is that you would take Inkling and fine tune it using Tinker. So that way thinking machines ends up being core infrastructure for teams that want to own their stack. But…do most teams really care so much that they want to own that part of the stack? Maybe at some point it ends up being cheaper, but open source models already force providers to bring down costs on margin.There were a raft of security issues that have come up over the last few weeks. Some of my favorites:
The Memory Heist. Basic idea: create a poisoned url, have Claude search for that URL, the url masks as a security page where Claude will type out, letter by letter, url by url, sensitive information into the web fetch tool. Claude will think it's doing the right thing and won't ask permission, and it will even pull from user memory to do so.
GitLost. Basic idea: git agents have more access than users, so you can literally just post an issue in a public repo and ask for private repo info.
Instagram Leak. Basic idea: agents in the meta user account support flow just guess at whether an account is actually yours, so you can report a high profile account as your account and by simply changing your physical location the agent may just give you the high profile account.
At some point more people are going to wake up to the need of centralized background agents for security reasons. In the meantime, we’ll get more of these. There’s an old saying, people only start caring about security when it bites them in particular.
The expansion of data centers in the United States is driving up power demand — and electricity bills — in large swaths of the country, drawing local and political backlash.
Only one in three Americans approve of the fast pace of data-center construction and most would oppose building one in their own community, according to a recent Reuters/Ipsos poll.
Dozens of state legislatures have introduced bills to rein in the effects of data centers on power bills and the environment. New York is the first to enact a full moratorium.
I think the water concerns are mostly overblown, but the power thing is real. The hyperscalers should offer to subsidize more than what they use to open the doors in local communities. More generally, I think people sorta miss why data centers are controversial. The conversation has mostly been about AI, but even though AI is quite unpopular, datacenters are even more unpopular than AI. I think it’s less about the AI and more about the local politics. The problem with data center build out in local communities is that the local communities do not benefit at all. I think it is totally reasonable for a person in, like, Memphis, to ask whether a data center build out in Memphis is going to benefit even a single local Memphis resident. So far, the answer has been no! The data centers are being built at the behest of Silicon Valley billionaires who are, bluntly, some of the most hated people on the planet, and who have basically openly thumbed their nose at any kind of legal process or restraint. I expect more of these data center moratoriums, especially as a way to win some additional votes for the midterms.
Related: Meta and SpaceX both begin selling excess compute to competitors in the AI race. Meta is literally building data centers in tents to build them out faster.
The obvious take is that this is a bearish signal, and the neo clouds (eg core weave) all had market sell offs. Semi Analysis disagrees though:
With Bloomberg headlines suggesting Meta could become a Neocloud, the market’s reaction was immediate: aggressive sell-off of Neoclouds like Coreweave & Nebius, and debates of “overcapacity” coming back. Let’s set the record straight – we believe that both takes are erroneous and that Meta’s datacenter & compute procurement will accelerate, not slow down. Capex in 2027 will be shockingly high. In just the first six months of the year, Meta has contracted over 5GW of capacity across Cloud & Colo, and that doesn’t even include all their accelerating self-build activity. Everything is computer and everything is a neocloud.
Apple sues OpenAI, accusing them of systematically exfiltrating hardware secrets.
at every level, from members of its Technical Staff to its Chief Hardware Officer, and in coordination with business partners, OpenAI has been stealing Apple’s trade secrets and confidential information. As a natural result, OpenAI’s nascent hardware business now rests.
…
The complaint, filed in the U.S. District Court for the Northern District of California, alleges that Tan used insider knowledge of Apple’s confidential projects to grill job candidates in interviews and learn more confidential information. Additionally, Tan directed job candidates still working at Apple to bring actual Apple hardware components and samples for “show and tell” sessions.
When interviewing Apple employees for jobs at OpenAI, Mr. Tan uses Apple’s confidential information to gain access to even more insider knowledge. He has used an Apple internal project codename to ask, “What’s the plan[?]” for an unannounced Apple product.
He has directed job candidates still working for Apple to bring “Actual parts” from Apple to their interviews for “show and tell” sessions in which he and his team at OpenAI can elicit still more Apple confidential information. These directions to bring Apple’s parts to OpenAI job interviews surprised at least one of the candidates, who commented that he “didn’t even know we could take those from the office.”
OpenAI has been instructing Apple employees to bring “CAD/design artifacts” and “prototypes” to their interviews and to divulge details about their work such as “subsystem and component selection,” the “tools or methodologies you use for system integration, such as CAD software, simulation tools,” and “Vendor selection and communication/collaboration with vendors.”
If any of this is true, this is a pretty terrible look for OpenAI and may just fully shut down their nascent hardware efforts. The same kind of lawsuit between Google and Uber fully killed Uber’s self driving car efforts.
Amazon is shutting down Mechanical Turk, because everyone on the platform is just using LLMs to do the data labeling.
An announcement on the Mechanical Turk website says that on July 30, 2026, the crowdsourcing service will close to new customers. Amazon Web Services says the decision was made after “careful consideration,” adding, “Existing customers can continue to use the service as normal. AWS continues to invest in security and availability improvements for Mechanical Turk, but we do not plan to introduce new features.”
…
Over time, the relationship between Mechanical Turk and AI models grew even more complicated. In a snake-eating-its-own-tail irony, a 2023 analysis found that between 33% and 46% of workers on the platform were using large language models to complete their tasks, raising questions about the reliability of data annotated on the platform and also about whether humans needed to be in the loop at all.
It’s hard to overestimate how important Mechanical Turk was to the rise of LLMs. Much like Stack Overflow, it was eaten by its own importance.
General Intuition got a neural net to fully simulate multiplayer rocket league! Its super cool, you can literally play it on the web.
It’s undeniably true that the Chinese labs did train on reasoning traces and outputs from Claude et. al. But also, Kimi outright beats Claude et. al. on several benchmarks, which is unlikely for a raw student-teacher training paradigm

















“This is the world’s biggest open weight model, clocking in at 2.8B parameters.”
I thought it was 2.8 trillion parameters?
Question: If corruption is really so big a deal, why is this not an issue in China? They’ve managed to make spectacular technological progress, not just in A.I. but also in batteries, EVs, solar panels, robotics, etc., and yet they’ve had corruption issues for much longer than the US. Is the corruption different?