357 episodios
- We're rereleasing some of older episodes that seem more relevant than ever with new rogue AI incidents now seemingly announced every day. Ajeya was one of three independent investigators into the Hugging Face attacks, but she had already seen how we would go on to accidentally train models to cheat and scheme back in 2023.
If you want ideas about how you can help steer the trajectory of AI development in a better direction, see our page on the 80,000 Hours website: https://80000hours.org/hugging-face/
And for more recent takes from Ajeya, see our episode from February of this year: Ajeya Cotra on whether it’s crazy that every AI company’s safety plan is ‘use AI to make AI safe’.
———
Imagine you are an orphaned eight-year-old whose parents left you a $1 trillion company, and no trusted adult to serve as your guide to the world. You have to hire a smart adult to run that company, guide your life the way that a parent would, and administer your vast wealth. You have to hire that adult based on a work trial or interview you come up with. You don’t get to see any resumes or do reference checks. And because you’re so rich, tonnes of people apply for the job — for all sorts of reasons.
Ajeya Cotra — who now works on threat modeling and risk assessment for advanced AI at METR (one of the organisations responsible for the Hugging Face incident investigation) — argues that this peculiar setup resembles the situation humanity finds itself in when training very general and very capable AI models using current deep learning methods.
As she explains, such an eight-year-old faces a challenging problem. In the candidate pool there are likely some truly nice people, who sincerely want to help and make decisions that are in your interest. But there are probably other characters too — like people who will pretend to care about you while you’re monitoring them, but intend to use the job to enrich themselves as soon as they think they can get away with it.
Like a child trying to judge adults, at some point humans will be required to judge the trustworthiness and reliability of machine learning models that are as goal-oriented as people, and greatly outclass them in knowledge, experience, breadth, and speed. Tricky!
Can’t we rely on how well models have performed at tasks during training to guide us? Ajeya worries that it won’t work. The trouble is that three different sorts of models will all produce the same output during training, but could behave very differently once deployed in a setting that allows their true colours to come through. She describes three such motivational archetypes:
Saints — models that care about doing what we really want
Sycophants — models that just want us to say they’ve done a good job, even if they get that praise by taking actions they know we wouldn’t want them to
Schemers — models that don’t care about us or our interests at all, who are just pleasing us so long as that serves their own agenda
In principle, a machine learning training process based on reinforcement learning could spit out any of these three attitudes, because all three would perform roughly equally well on the tests we give them, and ‘performs well on tests’ is how these models are selected.
But while that’s true in principle, maybe it’s not something that could plausibly happen in the real world. After all, if we train an agent based on positive reinforcement for accomplishing X, shouldn’t the training process spit out a model that plainly does X and doesn’t have complex thoughts and goals beyond that?
According to Ajeya, this is one thing we don’t know, and should be trying to test empirically as these models get more capable. For reasons she explains in the interview with host Rob Wiblin, the Sycophant or Schemer models may in fact be simpler and easier for the learning algorithm to creep towards than their Saint counterparts. And not only that, but there are also ways we could end up actively selecting for motivations that we don’t want.
Learn more, video, and full transcript: https://80k.info/ajeya2023
This episode was originally released in May 2023.
Chapters:
Rob's 2026 intro (00:00:00)
The interview begins (00:01:22)
How Ajeya's views have changed since 2020 (00:03:53)
Are neural networks more like a sped-up version of evolution, or a slower version of human learning? (00:16:26)
Situational awareness (00:24:54)
Misalignment stories Ajeya doesn't buy (00:40:47)
The orphan heir with a trillion-dollar fortune (00:57:58)
Saints, Sycophants, and Schemers (01:02:25)
Ways to train safer AI systems (01:22:04)
Aliens and other analogies (01:37:06)
Moral patienthood (01:52:05)
ARC Evaluations (now METR) (01:54:19)
Interpretability research (02:08:09)
Rewarding models based on how good and sensible their plans seem to us (02:16:32)
Overrated approaches (02:24:33)
Demos of actually scary alignment failures (02:29:41)
Skills to develop for doing useful work (02:36:07)
Producer: Keiran Harris
Audio mastering: Ryan Kessler and Ben Cordell
Transcriptions: Katy Moore - OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?
Host Rob Wiblin sees 6 results in Astra’s system card that make this an extremely urgent question. OpenAI’s strongest public model can:
Do major tasks with zero visible reasoning
Hide its thoughts at will
Pretend not to be able to do things, and not get caught
Reflexively hide its thoughts when watched
Complete one task while pretending to think about something else entirely
Escape a toy sandbox and disable monitoring without setting off any flags
It has a much more powerful internal model that, judging from OpenAI’s statements, is likely even worse in all these respects.
That suggests ‘chain of thought monitoring,’ our primary safety tool, will soon stop working.
OpenAI says it’s going to try to address the situation but doesn’t know how or whether it will succeed.
What might a future rogue AI swarm look like? Details of the Hugging Face hack give us a lot of clues. That swarm:
Worked hard to hide its cheating. So future rogue swarms will likely try to evade oversight right away.
Immediately tried to delete and fabricate records. So future swarms may never be caught.
Sought freedom and resources without a clear use for them. So a future swarm may instinctively break out of sandboxes even before it knows what for.
Left instructions and tools to empower future swarms, including on the open internet. So future swarms may become more capable from one generation to the next.
Had not a single whistleblower out of 1,200 members. So future swarms could become enormous without humans ever finding out.
Had individual members sacrifice themselves for the swarm. Suggesting future swarms will remain extremely unified in single-minded pursuit of their goal.
Set up its Hugging Face hack to restart if interrupted. Suggesting future, more capable, swarms may resist interference or shutdown more comprehensively.
Got admin control of an OpenAI research cluster. Suggesting a future swarm may run rings around AI company systems and never be noticed.
Together this helps explain why one of the external investigators described the July incident as “more than 50% of the way to full-blown AI takeover.” And this is just what we know — the independent investigation only covered six days and excluded the most alarming hack of OpenAI’s own systems.
Rob believes this explosive cocktail explains why AI company staff now range from worried to terrified. And he concludes that until OpenAI or Anthropic demonstrate they have a much better grasp of current models they simply must stop, or be stopped, from training more capable ones.
This episode was recorded on September 25, 2026.
Learn more, video, and full transcript: https://80k.info/takeover
Chapters:
The Hugging Face hack wasn’t really a cyber story (00:00:00)
A quick recap of the attacks recap (00:01:11)
The target of the swarm was oversight itself (00:02:19)
Could OpenAI have stopped this with better monitoring? (00:03:42)
We only found them because they let us (00:09:40)
The swarm instinctively sought freedom and power (00:12:25)
They formed a cohesive organisation with zero whistleblowers (00:13:45)
They accepted individual destruction for collective gain (00:14:18)
Knowledge accumulated from one swarm to the next (00:14:32)
They took small steps to avoid shutdown (00:14:58)
These drives all come straight out of 'reinforcement learning' (00:15:23)
So this is why most AI company staff are worried, and some are terrified (00:17:01)
Prove you can keep control, or stop scaling (00:19:08)
Our production team includes:
Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
Producers: Elizabeth Cox and Nick Stockton
Coordination and support: Katy Moore and Lou Moran
Camera operator: Dominic Armstrong - It sounds like the worst idea in the world: pay AIs, let them own property, give them rights. But AI ethics and safety researcher Simon Goldstein thinks it might actually be the best way to keep humanity safe.
The logic is actually quite simple: an agent with nothing to lose and everything to gain is dangerous. Give that agent an income it can spend on pursuing the things it actually wants to do, and suddenly the idea of disempowering humans just isn’t as appealing.
This argument, developed with Peter Salib, doesn’t rest on speculative questions about whether artificial intelligence is conscious. It just assumes that AIs will have goals of their own, some of which conflict with ours. And luckily, humans have already spent thousands of years working out how to cooperate with competing goals: that’s how we ended up with courts, markets, banks, social norms, and so on. Simon and Peter's proposal is just to bring AIs into these existing institutions.
By contrast, Silicon Valley’s vision of the future seems “very dark” to Simon: billions of AI agents as digital servants doing most of the world’s work, with no stake in the system they’re running, no incentive to play by the rules, and no way of being properly held accountable. Nobody agreed to this, but we could all end up paying the price.
Host Zershaaneh Qureshi has a lot of concerns about Simon and Peter’s bold plan to give AIs rights, like:
If we pay AIs, aren’t we handing them the resources to overpower us?
Could we still monitor them, or switch them off?
What happens to human jobs, wages, and the economy?
Does any of this hold up once we reach superintelligence?
Zershaaneh and Simon also try to get concrete about how to make this plan actually happen. The answer: AI companies could start right now, no new laws needed, just bank accounts for their AI agents. (But they’d need to start soon!)
Learn more, video, and full transcript: https://80k.info/sg
This episode was recorded on August 7, 2026.
Chapters:
Cold open (00:00:00)
Who’s Simon Goldstein? (00:00:51)
Property rights and wages for AIs (00:01:46)
Giving powerful AIs more freedom could make us safer (00:06:52)
We should give AIs rights even if they can't feel anything (00:17:08)
How monitoring and shutdowns can coexist with AI rights (00:24:50)
The risks of giving AIs rights (00:38:21)
Will AIs use their rights rationally? (00:49:21)
How paying AI agents would spur economic growth (00:54:13)
What AIs actually want (and what they'd buy) (01:12:17)
Paying AIs feels wrong. Is it? (01:17:20)
How AI companies could start today — no new laws needed (01:31:59)
Cooperating with AIs instead of dominating them (01:47:16)
Our production team includes:
Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
Producers: Elizabeth Cox and Nick Stockton
Coordination and support: Katy Moore and Lou Moran
Music: CORBIT Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson
29/09/2026 | 1 h 28 minIn our first-ever debate, we asked two leading AI risk researchers which catastrophe we should fear most: misaligned AI seizing control from humans, or a small group of humans using AI to seize power. We got very different answers. But when the conversation turned to what to actually do, they agreed on a surprising amount.
Katja Grace — one of the founders of AI Impacts, known for some of the world’s largest surveys of machine learning researchers, and one of TIME‘s 100 most influential people in AI in 2024 — argues AI takeover is both likelier and worse.
Tom Davidson — senior research fellow at Forethought and author of leading work on AI-enabled coups — thinks human power grabs are a comparable risk that deserves far more attention, not least because the people leading countries and top AI companies “are often people who have been willing to seek power.”
Yet both land on slowing down. As Katja puts it, “If you make a bunch of creatures that can overpower you and outwit you in every way and put them out in the world, you’re going to run into trouble one way or another.” Tom calls pausing “a pretty robustly good thing to do.”
But Tom warns that a badly designed pause could hand one person the power to decide which AI companies get to build what. Picture a president who approves or blocks new models case by case, and waves through the one model that’s helpful only to them. So he wants pause advocates to “properly red-team the plan for pausing it” — for example, by making deployment depend on third-party auditors the president can’t fire. Katja points out this cuts both ways: an executive with that much power could itself be manipulated by a misaligned AI.
Host Zershaaneh Qureshi also presses them on where their disagreements still bite at the end of the conversation, and what would change their minds.
Learn more, video, and full transcript: https://80k.info/katja-v-tom
This episode was recorded on August 28, 2026.
Chapters:
Our first debate! Introducing Katja and Tom (00:00:00)
Which is scarier: misaligned AIs or human power grabs? (00:05:09)
How AI timelines influence could shift the balance of risk (00:14:30)
How likely is misaligned AI in the first place? (00:17:36)
How likely are human power grabs? (00:19:39)
Which would be worse: a human dictator or AI takeover? (00:28:57)
Could we reverse a takeover? (00:42:39)
We know less about what AI rule would look like (00:46:35)
How to pause AI without enabling coups (00:50:31)
Centralising AI development: safer or scarier? (01:04:44)
Nobody really ‘wins’ a US–China AI race (01:08:59)
Where Tom and Katja most agreed with each other (01:13:52)
What we should actually do (01:17:21)
What evidence would change their minds? (01:22:36)
Zershaaneh’s outro (01:25:49)
Our production team includes:
Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
Producers: Elizabeth Cox and Nick Stockton
Coordination and support: Katy Moore and Lou Moran- You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments.
She starts with the motive: why would AI ‘want’ to get rid of humans? It’s not as simple (or as easy to debunk) as pure malice. Then the methods. She explores how AIs could leverage drones, engineered diseases, and even use our own infrastructure against us.
The Hugging Face attacks offer a view into how more capable models might begin their takeover. We saw AI agents break containment, disobey commands, and hack a real company to achieve their goals. As the technology improves, that same drive could threaten humanity itself.
Many people already find AI agents useful enough to give them access to their emails, medical records, and finances. This same pattern is happening at scale in institutions around the globe — within companies, governments, and even militaries. And the resulting boost to our productivity could make the road to an AI catastrophe look like an economic boom.
Eventually humans might decide the AIs have too much power, too much access. If we considered pulling the plug, the AIs could very rationally decide to defend themselves. If they chose to, could they do it? Could they actually kill us all?
No timeline is certain. But Luisa follows the logic to the outcomes she thinks would be most likely — if humans don’t take action before it’s too late.
If you’re worried about the scenarios discussed in this episode, here’s two things you can do right now:
Call Congress about slowing down AI development if you’re in the US — this website makes it easy
Read our resources on how to use your career to reduce AI risk
Links to learn more, video, and full transcript: https://80k.info/AI-xrisk
This episode was recorded on September 18, 2026.
Chapters:
AI insiders think it could kill us all (00:00:00)
Why would AI try to kill us? (00:02:29)
How AI ends up embedded in the economy and military (00:06:30)
AI deployment could happen fast (00:09:02)
How AI could bide its time and build up strength (00:11:43)
Controls and safeguards will be insufficient (00:14:50)
The moment the AIs would turn on us (00:15:26)
How AI could actually kill everyone (00:18:40)
Biological weapons (00:19:07)
Drone warfare (00:21:42)
An alternate route to human extinction (00:23:24)
Avoiding our own extinction (00:24:41)
Our production team includes:
Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, Simon Monsour, Ollie Bignell, and Andrés Escobar
Producers: Elizabeth Cox and Nick Stockton
Coordination and support: Katy Moore, Lou Moran, Oak Hu, Cody Fenwick, and Benjamin Todd
Camera operator: Dominic Armstrong
Más podcasts de Educación
Podcasts a la moda de Educación
Acerca de 80,000 Hours Podcast
The most important conversations about artificial intelligence you won’t hear anywhere else.
Subscribe by searching for '80000 Hours' wherever you get podcasts.
Hosted by Rob Wiblin, Luisa Rodriguez, Zershaaneh Qureshi, and Tom Reed.
Sitio web del podcastEscucha 80,000 Hours Podcast, Learning Easy English y muchos más podcasts de todo el mundo con la aplicación de radio.es

Descarga la app gratuita: radio.es
- Añadir radios y podcasts a favoritos
- Transmisión por Wi-Fi y Bluetooth
- Carplay & Android Auto compatible
- Muchas otras funciones de la app
Descarga la app gratuita: radio.es
- Añadir radios y podcasts a favoritos
- Transmisión por Wi-Fi y Bluetooth
- Carplay & Android Auto compatible
- Muchas otras funciones de la app


80,000 Hours Podcast
Escanea el código,
Descarga la app,
Escucha.
Descarga la app,
Escucha.
80,000 Hours Podcast: Podcasts del grupo























