We’re putting too much faith in AI’s ability to say no

At a glance
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary. But recently, the idea that AI s
- Primary source
- MIT Technology Review — Read original article
- Published
- Topic
- OpenAI
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience.
Most of these capers are, of course, cautionary. But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment.
In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless . This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?
Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol.
What it doesn’t learn is how to keep those powers to itself.
Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself wit
This summary comes from MIT Technology Review. Read the full article at the original source.
References
More in OpenAI

‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop
"Staggering." "Overwhelming." "Unprecedented." "Surreal." "Pure insanity." Those were among the descriptions more than three dozen mathematicians reached for in conversations with…

OpenAI doubles down on decision to fire three AI safety researchers
OpenAI is standing firm on its decision to fire three safety researchers after an investigation found they committed "a significant breach of trust." In a post on X on Friday, the…

Fired OpenAI safety researchers dispute their dismissals in open letter
Three OpenAI researchers said their firings may have a 'chilling' effect on employees.

Child safety group calls ChatGPT for Teens an unacceptable risk
After extensive testing, Common Sense Media found the chatbot failed in five key areas.