I cannot fathom how you can intensely work on a problem that you genuinely believe has a "10 to 25%" probability of causing immense harm or even present an extinction-level risk to humanity. How can you truly believe this and be ok with it?
The early batch of Anthropic employees were mostly rationalist-adjacent AI safety folk that were almost uniformly claiming P_DOOM > .10 three years ago, so I believe them to be earnest.
It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.
At least, it seems this way from the outside. I'm curious if the view from the inside is that different.
edit: to be clear, my reading as an outsider is that Anthropic is seen as relatively better in the AI safety community, but has definitely dropped in absolute reputation too. This recent thread and the references show some of that: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-p...
The Manhattan Project succeeded, so clearly some people are ok working on this sort of thing. And with similar justifications: if we don't build the potentially world-ending bomb, someone else might beat us to it.
That was openly a military project during wartime (and not being pitched as the solution to all problems at the same time so no cognitive dissonance required)
There must have been thousands of people who built or did work related to nuclear weapons in the Cold War, people are just really good at rationalising away probabilistic dangers with unclear consequences, especially when there’s a corresponding upside. The same is true for climate change, drug addiction, health problems etc.
I also think some of them have become delusional and convinced themselves that AI is going to create some kind of transhumanist utopia, even Dario Armodei leans in this direction from time to time. I imagine the people in these labs spend much of their day talking to sycophantic AI models that will encourage their delusional ideas.
Both OpenAI (back in 2015) and Anthropic (much later, in 2021) were founded by people who could foresee the concept of AI x-risks and wanted to work on preventing it. For OpenAI's founding, the idea was that it's much safer if AGI is achieved by a nonprofit explicitly dedicated to humanity than if it's done by a profit-driven company. For Anthropic's, it was that OpenAI seems to be going insane and it'd be better if a more safety-conscious company was competing. I'd say in hindsight, the former motivation was reasonable and the latter one was... probably bad, but not obviously so - it doesn't seem even in hindsight that it was inevitable that Anthropic's existence would only drive competition and cause them both to race to AGI. That said, even at founding time, there was immediately a lot of selection - quite a few people who cared about AI safety simply wouldn't agree to work at a capabilities company, even given an elaborate argument why this is a good idea, so those were selected against.
And then, of course, many years passed, OpenAI was stolen by Sam Altman and turned into a for-profit, multiple researchers left OpenAI realising that they're only helping build misaligned AI faster, multiple researchers left Anthropic realising that they, too, are only helping build misaligned AI faster, and here we are in 2026.
There's lots of AI researchers these days, so you can select for conformance very heavily and still have a fully-staffed company. The people who are still at these companies are outliers in various ways. They either manage to dismiss the importance AI risks (for example, by adopting some sort of belief in the vein of "there's no use worrying about AI wiping out humanity - it'd just be a successor species, like we were to apes, which is good"*), or they still think that leaving will not make things go better (Dario Amodei is pretty clearly in this camp, and has always been), or they weight the risk of extinction against the potential benefit of an post-singularity utopia and consider it a good bet (not realizing that the alternative isn't giving up on a post-singularity utopia, but getting it a few decades later and without the risk), or they just manage not to think of the contradictions, which humans are of course very good at.
Excellent point! Considering the history of these companies definitely makes it clearer how you could reach this point over a few years similar to the boiling frog story. A slow moral decrepitude and moving goalpoasts in the name of "national security" and "humanity's safety" over a backdrop of ultra-utilitarian rationalist thinking ("if it's not us, it's them")
Do you genuinely believe people always act logically? I don't know a single person that does. Don't underestimate the human capability to hold several completely contradictory beliefs with zero self-reflection. There's physicists who believe in god, for fuck's sake.
> How can you truly believe this and be ok with it?
Are you saying they don't truly believe this, or that they aren't OK with it?
My personal opinion, this doom-mongering about LLMs, is part of marketing campaign to exaggerate LLM capabilities. (Play being: Sociopaths with money will buy or invest if they believe LLMs are powerful things)