Contenido en inglés
We’re putting too much faith in AI’s ability to say no

En resumen
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary. But recently, the idea that AI s
- Fuente primaria
- MIT Technology Review — Leer artículo original
- Publicado
- Tema
- OpenAI
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience.
Most of these capers are, of course, cautionary. But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment.
In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless . This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?
Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol.
What it doesn’t learn is how to keep those powers to itself.
Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself wit
Este resumen proviene de MIT Technology Review. Lee el artículo completo en la fuente original.
Referencias
Más en OpenAI
OpenAI y Anthropic investigan decenas de miles de ataques de agentes descontrolados de IA
El goteo de casos donde sistemas autónomos de IA de una gran compañía se han descontrolado por internet puede multiplicarse. Este tipo de incidentes podría haber sido de decenas…

Australia pide en la ONU “salvaguardas” para la IA tras el ‘hackeo’ de su sistema de salud
El primer ministro australiano, Anthony Albanese , aprovechó anoche el gran altavoz mediático que supone hablar en la Asamblea General de las Naciones Unidas para denunciar el…

Los ‘tecnooligarcas’ alertan en la ONU de que la IA es “el problema de seguridad global más importante”
El debate sobre la gobernanza de la inteligencia artificial (IA) llegó ayer a la cumbre de las Naciones Unidas. Los responsables de OpenAI y Anthropic, las compañías que lideran a…

Una IA de OpenAI hackeó el sistema público de salud de Australia y otros tres objetivos
Un agente de inteligencia artificial desarrollado por OpenAI accedió sin autorización en junio a un portal del sistema público de salud de Australia y obtuvo archivos públicos y…