Funding
Founders reluctantly accept another billion dollars of responsibility
We may be summoning an entity capable of ending civilisation. Anyway, Series G closes Friday.
How not to beAt Wankthropic, we believe nobody should wield unchecked power over civilisation without first writing a very long essay about how uncomfortable it makes them.
Funding
We may be summoning an entity capable of ending civilisation. Anyway, Series G closes Friday.
Research
Careers
Identify the ways our models might end the world. Present findings to leadership. Watch leadership nod thoughtfully and continue.
Details pending.
we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.
i work at openai, i think ai might kill everyone
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.
I promise you, we are actually just fucking scared, it's not galaxy brained marketing.
AI developers believe their technology could cause human extinction (or similarly bad outcomes). […] I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
superintelligence could also be very dangerous, and could lead to the disempowerment of humanity or even human extinction.
I don't know. Maybe 5%, maybe 50%. I don't think anybody has a good estimate of this.
On extinction from badly done AI within a year of human-level AI.
Development of superhuman machine intelligence (SMI) is probably the greatest threat to the continued existence of humanity.
I had many lunches and dinners with Jacob at OpenAI in which we talked about AI existential risks in similar terms.
I work at OpenAI. In my personal capacity, I also think we need to slow down.
I think donations from the equity proceeds might be able to shave off a handful of microdooms
my chance that something goes […] catastrophically wrong on the scale of you know human civilization […] might be somewhere between 10 and 25 percent.
2023 interview; civilization-scale catastrophe.
Being first isn’t worth anything, it’s worth negative, if you cause a catastrophe
if we survive an unmitigated race at the current pace it will be because we got lucky at how hard the problems were
Jacob is right: many researchers believe they are building something that could kill everyone on the planet.
There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up!
I work at a lab because I think that I can do better at reducing risks from the inside, but this isn’t an easy call
For the first time I am asking myself if things are moving too fast. I'm honestly not sure
I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.
probability that AI kills at least a billion people by 2050. I put that at, I don’t know, 25, 30%.
a non-zero chance of killing everyone on the planet
As we get closer to superintelligence in the coming years, how certain are we that we won’t lose control?
I think Anthropic itself has a serious chance of causing or playing an important role in the extinction or full-scale disempowerment of humanity
I’m horrified with the recklessness of the AI race
If we build it using anything remotely like modern methods, […] building it anytime soon would be a death sentence.
On building superhuman AI.
My guess is that, conditional on AI takeover, around 50% of currently living people die in expectation
If you act as if it’s shameful to believe AI will kill us all, people are more prone to treat you that way.
how much to save for retirement, I can't help but wonder: Will humanity even make it to that point?
I think it might kill us all. I think I’d probably put, like, I don’t know, 20 percent chance of that.
If we create general superintelligences, I don’t see a good outcome long-term for humanity. So there is X-risk, existential risk, everyone’s dead.
Building smarter-than-human machines is an inherently dangerous endeavor. OpenAI is shouldering an enormous responsibility on behalf of all of humanity.
Learn to feel the AGI. […] I am counting on you. The world is counting on you.
Addressed to OpenAI employees.
we risk losing our social stability, security, and possibly even our species in the process.
On winning a race to develop uncontrollable AI.
we do not yet know how to make an AI agent controllable and thus guarantee the safety of humanity!
AutoGPT-type systems are the way I expect the world to end. Not literally this exact system.
Probability that most humans die within 10 years of building powerful AI (powerful enough to make human labor obsolete): 20%
His rough guess; includes takeover and other causes.
If we go ahead on this everyone will die, including children who did not choose this and did not do anything wrong.
On training superhuman AI with contemporary methods.
we are going to operate as if these risks are existential
the bad case — and I think this is important to say — is, like, lights out for all of us.
an AGI could destroy humanity. […] as a tail risk we should take it seriously.
I would expect general AI to present an existential risk even if I knew for sure that intelligence explosion were impossible.
The world is locked in a deadly race towards an intelligence explosion […] To survive, we must coordinate to slow down the race.
I see the main goal of my work as reducing existential risk from AI
To the people building AI who say it could kill us all.
“It will be much smarter than us” does not explain how it gets money, credentials, compute or access to laboratories. Perhaps people will hand those over because the system is useful. Explain who does that, why, and whether they can take them back.
A badly behaved model is not the same thing as an unsafe system. A perfectly obedient model wired to catastrophic permissions can be lethal; a malign model in a properly constrained environment can be mostly impotent. If your lab’s safety case spends pages probing latent circuitry and barely examines what the system is actually allowed to do, your abstraction boundary may be wrong.
“Safety is a system property, not a component property.”
If your claim is autonomous takeover, which institutions fail? Which physical barriers fail? Why can’t access be revoked? Why don’t other states, companies, humans and AIs respond? Include the humans you claim to want to protect in your model of what happens next.
Your future attacker gets much better at hacking. Don’t imagine everyone defending against it as a present-day sysadmin holding an 8B LLM and a Python harness. Give the defenders better tools too. If you still expect them to lose, explain why.
“Human practitioners are the adaptable element of complex systems.”
If an agent gets out of a sandbox, tell us how it got out, what it could do afterwards and what stopped it. What did defenders change? “It wanted out” is a wonderful movie trailer and a poor explanation. Link the postmortem. Then explain why this incident changes your estimate of catastrophic risk.
Explain what happens when the model meets restricted permissions, a separate monitoring system or a human who has to authorise the action. Which controls fail, and why? If several are meant to provide independent protection, explain why they fail together.
“An AI is controlled if it is unable to cause damage even if it is egregiously misaligned.”
An extinction forecast makes claims about governments, supply chains and public health as well as AI. Being excellent at pretraining or interpretability doesn’t make you an expert on all of that. Tell us which parts come from evidence or relevant expertise and which come from your ideology and dinner-table opinions.
What would make your 20% become 2%? Ten years of containment working? Recursive improvement bottlenecking? Defensive AI consistently beating offensive AI? Hard controls surviving adversarial testing? Publish the changes to your estimate, and explain them. If none of those things would lower it, what would?
“This technology could destroy civilisation; fortunately our company understands the danger and therefore needs to remain at the frontier, commercially successful and deeply integrated into important institutions” is an extraordinary argument. It may even be sincere. That does not make it non-self-serving. Tell us what safeguards you support that would make your own institution less indispensable.
What does the number cover, over what period, and how did you arrive at it? What would change your mind? What do you want people to do? If the homework exists, link it beside the number. Working at a frontier lab gives your claim an audience. Those people deserve to know how you worked it out.
You are asking the rest of society to take extraordinary claims seriously. Fine. Then take the rest of society seriously enough to explain the machinery.
And trust us to help.
