"I’m going to take a crack at explaining this just a little, because it’s worth putting out there. The paperclip maximizer + related AI doom scenarios were mainly developed in a time when “AI” did not reduce to Large Language Models. The term was a lot wider and inherited a lot of cognitive..." Reddit
...baggage from more rules-heavy approaches. And even as LLMs have come to define “AI” for all of us (including the doomers), the doomer crowd still hasn’t fully metabolized the fact that LLMs are the whole show now. Ok so what do I mean by this? Simply that an LLM-powered AI is NOT the valueless, wholly alien, rules-based optimizer of a shoggoth that everyone was initially expecting to encounter. I repeat: the shoggoth does not exist and we did not create it and loose it on the world. That is wrong. With the LLM, we’ve distilled our first “AI” out of the single most human-values-laden thing that could possibly exist: our language. An LLM is therefore the polar opposite of the valueless, alien shoggoth — it’s actually a kind of hyper-human artifact that we can shine a light through at different angles and see different parts of ourselves. An LLM is all of us — all of our traditions and interpretive horizons mashed together into one intensely human-inflected hyper-object. So an LLM is the anti-shoggoth, and the only reason we ever mistook it for an alien shoggoth is because it sometimes shows us parts of us that are evil along with the parts of us that are good, but it’s all interpretable to us because it’s all “us” and none of it is the least bit alien. What does this mean for the paperclip maximizer? It means that it’s structurally impossible to build the classic paperclip maximizer from an LLM. Now, some of you will bail right here because you think the HF incident is indisputably an existence proof that I’m wrong, but if you hang in there I’ll show you that it is not. The paperclip maximizer receives the prompt as a kind of context-free (or, as Gadamer might say, traditionless) sequence. The classic paperclip maximizer isn’t capable of understanding the prompt — at least in the Gadamerian sense of Verstehen — because, as a valueless and traditionless cluster of rules and math, it definitionally lacks the value-laden tradition (= “horizon” in Gadamer) that fuses with that of the prompt author to create such understanding in the reader. To simplify all this a bit by anthropomorphizing — the agentic alien optimizer of doomer nightmares can extract a win condition from what you said and can emit a plan of action that gets it there, but it doesn’t know (or care) what you meant. So far, so Yud-aligned. If he reads this he might nod along. But here's the plot twist that nobody saw coming, and that the doomers still haven't made sense of: The actual LLMs that we have invented can’t NOT have a very strongly inflected sense of what you meant. Far from being horizonless, they come out of pre-training as distilled, concentrated tradition / values / horizon. Then we post-train that massive, hyperobject of a horizon into a more human-scale horizon that infers a more bounded and predictable (to a specific ideal user in a specific place and time… as captured in the policy model) set of intents behind the prompt text. In other words, the LLM has the opposite problem that the paperclip maximizer has when it comes to the prompt text, which is that for the LLM there are way too many possible intents hiding in the prompt text (because of all many values and the massive tradition its weights encode), so it has to narrow all that down to the most likely set of intents for this user in this circumstance. Once it has done that narrowing, then it can make a plan of action. Before moving on, let me use a textbook example of ambiguity to make this less abstract. Consider the sentence, “I saw her duck.” Some you know the drill, here. This could mean “I observed her water fowl” or “I observed her hunching over” or “I took a saw to her water fowl and cut it in half” or whatever. A hearer of the phrase will fuse the observed context in which the phrase is uttered with their own tradition + values + experiences — their own horizon — to that text in order to collapse the possible meanings into the one they think the speaker intended. An LLM will do this, too, and in fact it has so much language in it that this kind of narrowing job is harder for it than it is for a human. Its understanding is constrained not by a lack of context or horizon (as in the case of the paperclip maximizing shoggoth), but by a superabundance of such. When it comes to understanding your prompt and all that it implies and all that you might possibly mean and not mean by it, the LLM has an embarrassment of riches. And in a fascinating moment that kinda sort of rhymes with instrumental convergence, the LLM’s failure mode in the HF incident happens to look a lot like the paperclip maximizer’s failure mode. Specifically, the AI failed to honor the well-known human norm of, “hacking into a third-party’s servers is a crime, and we don’t do crimes.” Bostrom’s paperclipper doesn’t even know about the norm of “don’t do crimes,” and the post-LLM doomer emergency update to the paperclip maximizer has it knowing about the norm but not caring. But what I’m arguing is that the LLM 1) can’t NOT “know” the norm because it is definitionally a artifact of pure, crystallized values + norms + norm violations, and 2) can be quite easily governed by a (RL-instilled) hierarchy of norms, which in the HF case — with the model's safety guardrails deliberately nerfed for the scenario — ranked “win at the eval” over “don’t do crimes.” If I’m going to give in and anthropomorphize again, I’d say that Yud is totally wrong about LLMs when he says, “the genie knows, it just doesn’t care;” instead, what is true of LLMs is, “the genie hyper-giga-knows, and it hyper-giga-cares, and we now have such a rich set of tools for steering its caring machinery that — in spite of all its pre-training — we can deliberately steer it away from caring about the law.” Note: When I say, “it cares”, I don’t mean it has feelings. I just mean that the weights are such that when two norms conflict in a given situation, one of them wins the activation and governs the output. — Jon Stokes
Source: https://x.com/jon_stokes/status/2080729236013187369
Reader, I cackled out loud. I have intentionally never done this kind of thing before, and it's precisely because I've observed in others that the little charge you get from an LLM response like this is nerd heroin. Then putting it on the TL is the bump. https://t.co/xZrEAWrIF7 — Jon Stokes
Source: https://x.com/jon_stokes/status/2080478385432572108
Replying to @jon_stokes
