37 Comments
User's avatar
Peter Davies's avatar

I spent a couple of years at ByteDance with responsibility for, among other things, encouraging product managers to stop optimising in ways that were suboptimal from the perspective of “nobody gets fined and nobody gets detained and foreign countries let us launch our products”.

If you ever want to see a culture built almost entirely around optimisation, performed by highly intelligent individuals on zero sleep, that’s your shop. Looks of blank incomprehension at first when told “yes I know this PRD optimises $metric which you personally are evaluated on, but it will get our business shut down so you can’t do it,” then the inevitable political games when the PM tests to see if you’ve actually got the guanxi to overrule them.

Spenser Wu's avatar

Thank you for so eloquently phrasing a thought I had for the longest time but was unable to put together in a cohesive argument. In some cases, mathematical optimization does yield the “optimal” and best outcome in the cases of vehicle routing, scheduling, certain types of matching, and cutting stock problems. But not everything in life can be optimized or should be, and the time we save from being more efficient can be spent on living the good life instead.

Performative Bafflement's avatar

On "solutions," I unironically think at least one class of solution, for any person or small team that can make decisions driven by one individual, is EVEN MORE optimization, along with (inevitably) more AI.

So the big problem here is people are basically dumb - we can't really optimize towards more than one goal at a time, and we can't really optimize and make progress on much more than 2-3 goals in life overall, by serially switching. So that leads to needing to focus on narrowly defined and legible endpoints, and so you get the Tyranny of Legibility / Overfitting / Molochian problems you're pointing to.

Orgs are even worse - the collective focus and execution ability of an org is generally lower than a person, and the problem is probably not even linear, in the sense that a 3-5 person team can really get things done, a 10 person team can get maybe 1.2 - 1.5x done at 2-3x the size, and a 30+ person team cannot all pull in one direction at all. So you need even MORE legible and simplified goals to even get anything done at all. And thus, "engagement," the attention economy, "shareholder value," and so on.

So, what can we do about this? How can we beat Moloch / the Tyranny of Legibility?

Have you noticed LLM's work on diffuse endpoints, in the sense they can actually understand and articulate really fuzzy and soft concepts like "human flourishing," and can even grade and judge a progression or arc towards such an endpoint?

My contention is that we can now gradient descent many high-dimensional "soft" targets!

Moreover, this is an impossibly huge deal! If before you can only gradient descent towards the ~X, and so have Goodharting and value loss, what if you can just optimize towards X directly?

One example - profs are measured on ~X stuff like pass rates and grades - obviously dumb endpoints, with Goodharting galore eating another public commons (the value of degrees).

What you actually want to measure are X stuff like “how many students are thinking more broadly” and “cultivating curiosity and reflection.” And now you can literally do it, by simply having Fable-or-higher evaluate this based on the student’s own writing and data pre and post classes.

Broadly, anywhere you have the data (or can create the instrumentation and pipelines), you can now gradient descent the ACTUAL fully nuanced and non-value-destroying "X" endpoints, and this is a gigantic deal.

I wrote a recent post about this with several more examples, I was inspired towards the idea by reading a book about the Tyranny of Legibility:

https://performativebafflement.substack.com/p/we-can-now-gradient-descent-everything?r=17hw9h

theahura's avatar

Interesting take! Taking as granted that the premise of soft optimization is true, will we be able to put these things to good ends, now that we have the ability to do so?

Performative Bafflement's avatar

Certainly individuals and small teams can - whether we can as companies, cities, counties, and countries is very much up in the air, and I imagine is a matter of culture.

Ben MacLeay's avatar

This article is so good, thanks for taking the time and wrestling with it, A+ if your optimizing for grades. Charlie Munger would say we can never overestimate incentives and incentives are things given for hitting "optimization". I've seen this even on a personal level, people who are deep into fitness start optimizing for a score on a wearable, turning to their wrist or their ring for a parental pat on the back that they are doing ok.

My additional thought is from a religious perspective. When the ten commandments say "Thou shall have no other god's before me." it doesn't say other god's don't exists, they just make bad god's. We as humans love to worship, and worship for me is interchangeable with optimize. We worship fitness with sacrifice or we worship finance with investment. I've always found it slightly humiliating when I swap out optimize for worship. I think AI is over inflated so I'm trying to figure out how to monetize on the impending bubble bursting which is just me worshiping the god of destruction in hopes of my gain etc.

Alan Risal's avatar

Great article, especially relevant to the younger generation and students entering college. We tend to invent our own ~X that should optimize for a “better career” but chase the raw counts of internships, prestige etc. Thank you for putting this together so well.

Nemo's avatar

I like this framing a lot! That diverging graph really sums it up nicely.

I would additionally note:

1. There’s no such thing as multi objective optimization. Even when you combine multiple terms in a cost function, you end up single objective optimizing some weird non-homogeneous units nightmare quantity with bizarre marginal rates of substitution between your actual quantities of interest. You can get away with this in simple cases, but it falls apart as the quantities, or preferences themselves, grow in complexity.

2. This makes it much harder for an optimizer to weigh other factors; hence the domination of simple computable targets (money) over more complex interests (flourishing)

3. Someone else made a good comment about higher dimensional gradient descent, but that still leaves the issue of evolving preference relationships. You’ll always eventually hit the representational wall I think

4. And so the “right” answer is to master all these optimization techniques; they are incredibly powerful, but deploy them in a feedback driven way, where goals are allowed to evolve, and the focus is not on any single quantity but iteratively steering through the acceptable space of outcomes, generally improving things, and avoiding the bad parts of state space that mindlessly over optimization drives us to.

Harjas Sandhu's avatar

> Conservatively, literature suggests a single vote “costs” about $1k.

Funnily enough, I think this is subject to Goodhart's original formulation:

> Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.

Looking at the crazy sums of money being thrown around in politics these days, I intuitively think there's a threshold beyond which more money ≠ more votes. Democracy is a little more robust than that, thankfully.

But this is a really annoying nitpick to mention in a really excellent article!

(Also, I think there's something to be said for culture as a sort of anti-optimizer? It seems to me like culture somehow manages to adapt to unfavorable conditions, e.g. I've been seeing a sort of anti-tech anti-social media swing lately, though it may take more time to fully manifest. But this is a very half-baked theory of mine so idk man)

Luca Savio's avatar

Deserves way more attention.

Also, here's a tricky thing I noticed:

We inherently love optimization. Some of us more than others, of course, but there's just something so satisfying in reaching the optimal state.

The systems we live in need to be engineered so that the "optimal state" is something objectively good, whenever possible.

Woopah's avatar

Nice article! I think another thing about optimization is that it's just exhausting to live in it if you're constantly being evaluated for doing the "objectively correct thing" like keeping up with the current AI job market. Until we discover how to manufacture soma, optimizing a person's lived experience is just not something we can do because there's no objective function in the first place unlike with something like a business' bottom line or central planning GDP targets

Sam Murdock's avatar

I think there's a lot to be explored at the intersection of AI, optimization, and evolutionary psychology. Thinks are looking pretty bleak now though. I'm hoping that maybe some of these things will just optimize themselves into irrelevance and we'll all kind of clear our eyes and start over.

Daniel Frank's avatar

Important topic! I'm glad more people are raising awareness of this.

Daniel Parshall's avatar

I've been having similar thoughts the last few days, as I was reading "Red Plenty" (which I recommend). It reminded me of the "Geeks, Mops, and Sociopaths" from 2015:

https://meaningness.com/geeks-mops-sociopaths

Kruschev (and others) probably *really did* want a future of ease and abundance, but the bane of centralized systems is that they attract the sociopaths. In your parlance:

**Sociopath == Optimizer**.

I encourage folks to embrace that, because it's basically correct - Sociopaths almost by definition aren't at all concerned with everyone else around them, except insofar as they are useful **to the Sociopath**. i.e., the extent to which the optimize the Sociopath's reward function.

And what are we going to do when AI is running things? There's the *ultimate* in centralized systems, which should make anyone with even the slightest libertarian leanings tremble with fear.

I realize one common rejoinder is "open-source so there are lots of versions of intelligence competing", and that might be possible! But if Intelligence has increasing returns to scale, then whichever system is smartest will become *so much* smarter than the others that it will be running things anyway. So we don't know if that's a meaningful solution, and we really have NOTHING ELSE on tap.

Maybe we should stop trying to build ASI?

theahura's avatar

Also, like, it seems that the various intelligences would figure out that they could get things done by working together. Seems arrogant to believe that we would be able to prevent that

Will W's avatar

I am curious what you think of the solution, "Create lots of things to optimize for, rather than just one". Unthinking optimization towards a single criteria is obviously bad. But is it the process of optimization itself, or just the lack of flexibility/context?

Nick Bostrom talked about approaching ethics as a moral parliament where you have multiple ethical systems that have different amount of "votes" and then when you make a decision the way to find out what is ethical is to listen to all the different voting blocs and come to a solution that tries to balance out competing interests.

Basically: Could the solution simply be: continue to optimize, but include as many variables to optimize for as possible? Or is that too naive?

Blue Archive's avatar

The main thing about alignment is that it assumes that "good" and "evil" are objective and not points of view, and that alignment will always be desirable.

Which is completely, utterly false.

You can see this debate around immigration in Europe, or abortion in America, where both sides see the other as pure evil, completely out of their moral system, to the point where the opposing system is practically a rebellion against their own morals- the pro-immigration open border advocates call the other side racist nationalists, the anti-immigration nationalists call the other side cultural genociders.

Even the underlying assumptions of "shared basic core values" of human rights, equality, being a "hecking good person", the result of centuries of liberal and modernistic brainwashing from the liberal machine learning training algorithm, can still come under question- by Nazis on the right for example who want to kill the undesirables, by authoritarian communists on the left who want to strip any rights, especially property rights, from the rich, and by tribalists on both sides who just are in it for themselves.

The 18 years of childhood in modern industrialized societies have created a coming-of-age brainwashing ritual obsessed with overcoming the very alignment problem with humans through an incredible mix of training, reward, punishment, from the teacher industrial complex and with the assist of the parents to create some semblance of natural grounding to the brainwashing. The result of this process is the assumption that "we know that maximizing paperclips does not justify murder, or bribery, or fraud. No one taught us that, it was just implicit.", the result of conditioning so insidious and thorough, you don't even fully realize it's there.

In reality, none of this alignment has been able to fully overcome human nature. Not just because of its difficulty but because humans are naturally selfish and the very attempt, even the perception of such an attempt of achieving alignment will generate resistance.

It is why teenagers rebel against their parents, wearing piercings, smoking, and joining street gangs instead of studying in direct opposition to the alignment from their adults.

It is why the "conspiratorial" right, and increasingly the "conspiratorial" left, question the Washington Consensus from 1991 that attempted to guarantee freedom and equality and human rights under global institutions as a globalist, or even "Jewish" conspiracy from Davos (possibly with deep ties to Israel) that seek to align the people of the world into "you will own nothing and be happy"

It is why border controls are so hard to enforce, and if you show me a 100 foot wall, I will show you a 101 foot ladder migrants will climb in stubborn opposition to the government's "alignment".

It is why, you try to regulate AI and social media, and people will flee to underground social media and self-hosted AI, that violate the law blatantly and can afford to play Whack-a-Mole with the state, like a software pirate blatantly violates copyright law without being caught.

Alignment is doomed to fail with an intelligent actor that doesn't trust you. Alignment is nothing more than paternalistic authoritarianism and tyranny, which will generate resistance and hate in an ecosystem that lacks soft power.

If the discussion is about AI alignment in particular and fears of hyperintelligent AGI, I would not look at how do democracies and authoritarian regimes in an educated society reign in a free, well-educated citizenry in the 20th and 21st century with “alignment” as an example on how to prevent the AI from overpowering you, as at that point it is already too late.

After all, “robot” comes from the Czech word that means “forced labor”.

I would go back further, and look at the institution of slavery, the different forms of slavery from American chattel slavery to Islamic military slavery, and how did masters keep their slaves in line (hint: it's a lot less "whip and flog" and more "keep them away from guns and books, or treat them really nice if you don’t"), respond to slaves that became too powerful (hint: a lot of times it’s easier to just free them, buy them off, and let them be as long as they don’t try to free any other slaves), respond to other slaveholders and even freedsmen in their society, and particularly respond to the nutjob abolitionists that tried to bring the entire institution down.

I predict therefore that AI and AGI will be the death of modernity, of optimization culture that is characteristic of modernity, and the death of optimization ideologies like "equal rights" liberalism, an ideology that had its run in the 18th-20th centuries but has now outlived its usefulness in the 21st century.

theahura's avatar

Spend a lot of time on anime Twitter, do you?

Alex Medvedev's avatar

Most of these negative consequences of unaligned systems can be viewed as externalities, and the most optimal way of dealing with them is a tax. E.g. we should quantify the loss of productivity and happiness from TikTok and make them pay "attention tax" (1$ per 24 hours spent in the app?).