This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.
Boundaries in concept space
To serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them.
If we want the AI to follow our goals and values, we want it to be able to recognise concepts like "human being", or maybe "conscious being", "suffering", "preference satisfaction", "law", and so on. So we want an AI to be able to look at a situation and assess, e.g., whether there are or aren't suffering conscious beings in it.
But concepts like "conscious beings" are not crisply defined across all possible world-states. A fixed-weight model will draw a boundary between "conscious being" and "non-conscious being" (or maybe score the amount/degree of consciousness), but this will be an imperfect boundary. In high-dimensional spaces, there are many degrees of freedom of how a boundary can be drawn, and never enough data to draw the boundary perfectly[1].
If someone is confident that we can draw an acceptably reliable boundary defining "conscious being", grounded in fundamental facts about the universe (e.g. basic physics), and resistant to all ontological crises... well, let's just say they have an optimism about concept rigour that flies in the face of all past experience.

Almost perfect decision boundaries. Almost...
False positives and false negatives can both be disastrous: excluding conscious beings from consideration (therefore their suffering is ignored) or including non-conscious beings within the list (if smiling faces are ranked as conscious beings, then tiling the universe with smiling faces is an optimal action - even at the "minor" cost to those "humans and animals" running around).
These examples are "adversarial", similarly to adversarial images that break classifiers: something of one category that is designed to be classified into another one[2].
Self-Goodharting on adversarial data
A reasonable reaction to the above might be a shrug and a statement of "so what?". Yes, in theory, all models have adversarial examples. I mentioned jailbreaking as an example of adversarial data. But modern models are much harder to jailbreak, and adding other models as evaluation agents dramatically improves their resistance to jailbreaking.
Now, fixed-weight evaluation agents don't change the setup away from being fixed-weight, but it is reasonable to argue that, maybe in theory there exists adversarial situations that an adversary or the world might create, but these are so hard to find (especially if the model is black-box) and so unlikely to occur naturally, that it's not worth worrying about.
Unfortunately, there is one AI agent with the motivation and most likely the ability to create such adversarial situations: the model itself, if it's used as part of an optimisation process.
Goodhart's law focuses on situations where the proxy for something gets detached from , the actual thing itself. For instance, strong results have low p-values, and we incentivise people to publish strong results; hence the habit of p-hacking.
What's often unstated[3] is that it's often much easier to maximise the proxy than the real thing. Finding strong results with low p-values requires finding true patterns in the world and having a large enough study to prove them. In contrast, p-hacking just requires some light statistical skill and a lack of conscience.
Just as it's easier to fill the universe with smiling faces than with genuinely happy people, breakdowns in concept boundaries are points where it's potentially much easier to accomplish the goal, as was given to the model. Some adversarial situations might have a proxy that's harder to maximise than the true . But some of them are going to be situations where -maximising is easier; a few of them, situations where -maximising is much easier.
And so a fixed-weight AI model with a goal given in terms of its own concepts, will naturally steer towards these adversarial situations, searching its own weights for an easy -maximising opportunity, and trying to move the world precisely into those adversarial situations.

Almost... almost ain't good enough.
- ^
This is one of the problems with methods like RLHF: they use expert comparisons between two options to define a boundary at that point. But the pairs of options will omit many unusual or novel possibilities, so the boundary will only be drawn well "in distribution", i.e. close to the option pairs used.
This is also why models have remained vulnerable to jailbreaking: convincing a model to do something it is designed not to, maybe by mis-spelling the prompt or constructing a fictional scenario. When changing the weights to avoid misbehaviour, the model designers can't draw the boundary perfectly, so can't exclude all the stuff they want to exclude.
- ^
Note that ACE, the Algorithm for Concept Extrapolation, can detect adversarial images that are constructed end-to-end to fool it. The secret is that ACE does not use fixed weights, and thus can modify itself to figure out that something is suspicious with the data (indeed, since an adversarial image has the properties of more than one labelled categories - looks like A, pretending to be B - it is the kind of thing that ACE is intrinsically designed to find).
- ^
Related to, but not the same thing as "adversarial Goodhart".
Discuss