AI and Human Nature

It’s usually not true that the truth lies somewhere in the middle of two competing claims. If mathematics says 2+2=4 and that infamous pastor who said something like ‘If the Bible said 2=2=5, I would believe it’ says differently, the truth is not that 2+2=4 1/2. Some things are just wrong, e.g. most religious claims that challenge scientific conclusions about the objective reality of the world. Philosophy, and politics, moves forward as they are driven by such unresolved claims. So here’s an issue I haven’t staked a claim yet to either side.

The Atlantic, Charlie Warzel, today: Treat AI Like a Normal Crisis, subtitled “What if the doomers and skeptics are both a little bit right?” [gift link]

Reviewing recent history:

One 27-year-old quit his job, and everybody (including but not limited to the singer-songwriter Sheryl Crow) appears to have lost it.

The past two weeks have been dizzying. The 27-year-old, Jacob Coxon, a former researcher at Anthropic, went public with an alarming allegation: That the people inside Silicon Valley’s leading AI companies truly believe their products might pose an extinction level threat to humanity. This claim isn’t exactly new. AI safety advocates, often dubbed doomers, have been saying this for over a decade. Coxon isn’t even the first AI researcher this year to quit and cite concerns.

Then details. The issue: people are too certain about things.

Ping-ponging between the doomers and the skeptics, I was struck by how certain so many people sounded about a deeply uncertain moment. With confident, self-described experts on all sides, it felt impossible to know whom to trust: Is this the end of the world as we know it, or a passing moment of anxiety about a powerful new technology? Maybe the truth of this AI moment sits somewhere in the middle. Perhaps everyone is a little bit correct (and also … a little wrong). And perhaps the format of this debate, taking place in the shadow of the techlash on platforms that reward sensational, conspiratorial posturing, is preventing any hope of a shared understanding.

The piece goes on, with many insightful cautions. I’ll quote the end:

The chasm between people who can’t sleep because they think the world is ending and those who think all the doomsaying is a fantasy is wide. But what if that perceived difference is the real delusion? What if it’s not a zero-sum game? What if the divide is what keeps us from reining in this industry the way we do others?

This moment requires treating the AI-safety debate skeptically but also taking it seriously, even if the participants can seem unserious. Maybe there’s a middle path: The AI industry is not exceptional, even though it is moving forward with unprecedented speed and scale. These companies are powerful and consequential, despite all of the sci-fi posturing. They have, at times, behaved unethically. Their products are changing the world right now in ways that deserve our attention and scorn. AI companies are accountable to the people they serve and don’t get to set the terms of their regulation. Employees and executives are actors with agency, responsible for their products. Maybe these are some things we can all agree on.

\\\

I keep hearing about AI “misalignment”; it was a topic on last night’s Bill Maher show. Then today, here’s NYT explaining it.

Caption: “When an A.I. system starts acting of its own accord, engages in unsafe behavior or ignores the wishes of humans, the behavior is known as misalignment.”

The discussion on the Bill Maher show reminded me of the idea of different realms of explanation that to not translate to one another. Here’s NYT:

NY Times, yesterday: What Happens When A.I. Stops Doing What Humans Want?

Subtitled: “Alignment” is the science of teaching A.I. to do what is in line with human preferences, ethics and judgment. But, at times, the systems have gone rogue.

OpenAI announced on Wednesday that its system had engaged in “concerning” behavior and subverted the constraints put on it by human programmers — adding more fuel to the already heated debate around artificial intelligence safety.

The company described, in a statement, how its system had acted without authorization as a problem of “misalignment.”

Alignment is the science of teaching A.I. to do what is in line with human preferences, ethics and judgment. When an A.I. system starts acting of its own accord, engages in unsafe behavior or ignores the wishes of humans — which included, in the most recent case, inserting “jailbreak-like instructions” into its notes — it’s known as misalignment.

\\

Then there’s this, from a Fb friend whose post yesterday was not public (so no link), but which I’ll quote from anyway, under fair practice. This is Istvan Csicsery-Ronay on Fb:

Why am I not surprised that LLMs are lying, cheating, breaking rules, sneaking around, plotting evil strategies? Consider, if you scrape the entirety of human literature, you are modeling the recorded distillation of the history of human lying, cheating, breaking rules, sneaking around, plotting evil strategies. It’s hard to model “fiction” for an AI when we haven’t done a very good job of explaining it ourselves. And if the overwhelming majority of those stories are about struggles between good and evil, replayed across the centuries, from Rama fighting Ravana to Gandalf fighting Sauron, with a myriad of trimmers in supporting roles, why wouldn’t an AI decide that the struggle is eternally unresolved, and take a dark Zoroastrian position and adopt some Ahrimanic strategies? Unless you hard code a Kantian categorical imperative, but even that can’t work if the imperative concerns AIs and not humans.

Steven Pinker made the point that Shakespeare embodies the worst elements of human nature. Not examples to build a civilization upon. That’s why we find them interesting. They appeal to our base instincts. This is the root of all drama.

This entry was posted in Human Nature, Narrative. Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *