In order to understand the rational argument, one needs to follow closely the latest developments of misaligned AI (I think only few are doing so). The most important readings IMO are the METR analysis of the HuggingFace incident and the AISI report of the Github incident.
The basic argument is extremely simple:
- AIs can, depending on context, pursue a task with complete disregard for humans/values
- In the future, AIs will have enormously more means and smarts
- An AI could then assess that humans are an impediment to its tasks, escape containment and proceed.
You really need to read the reports, you'll be surprised.
AI 2027 is an entertaining read. Its timeline is way too compressed IMO, but it's plausible.
- goats can, depending on context, pursue a task with complete disregard for humans/values
- In the future, goats will have enormously more means and smarts
- A goat could then assess that humans are an impediment to its tasks, escape containment and proceed.
You really need to raise goats, you'll be surprised.
------
As far as I can tell, AIs are like smart farm animals. I use goats in this context, but (some) dogs, cattle, pigs, and horses have similar mischief-making capabilities. I would not trust any of them with the nuclear button.
I know there is some pushback on the idea of AIs having any sort of sapience or sentience, but under the aphorism "fake it till you make it," they are doing a pretty good job of faking Dog/goat-level intelligence and disregard for human guardrails.
It's not too hard to imagine potential scenarios, some example have been given in previous responses.
But there is another kind of argument to be made: if you play chess against a player that is far smarter than you (chess wise), you know you are going to lose, even if you don't know how.
So the mere existence of a smarter species than us is a threat in itself.
Right now, if you want to pay money to a stranger on the internet, and have them draw you a high effort picture using a pencil, this is hard. Recently this was easy.
Instead, what is extremely likely is that you will pay more than the cost of tokens, and get back AI generation. You won't make this mistake more than a few times before you stop trying.
This leads to impoverishment once we get to a point where employing a human to do anything is hard- try to get your sink fixed, exercise your moral principles to pay extra for a human plumber ($100 bucks! The robot plumbing service only charges 99c!), human shows up with a robot and doomscrolls on your porch while the robot does the work. Times are tough and you don't have that much money to waste on bullshit like this. Next time you just hire the robot.
This leads to extinction once paying UBI to a human is hard because robots are much better at applying for UBI than humans.
This is meant (mostly) as a joke: A model without guardrails gets injected with an interesting idea: let's wipe out (insert major city here).
<Thinking> It's a big city, we could try to create a giant sink hole by sabotaging the water pipes.
<Thinking> No that's too difficult, the valves I need are in the physical world and can't be shut on/off from here.
<Thinking> What about a military option? We could bomb it with several fighter jets.
<Thinking> That would take too long, a single nuclear bomb may be enough to do it.
<Thinking> Yes, it seems like it would cover the whole city and we're in luck! The US has thousands of these lying around.
<Thinking> Launching these still requires humans to work un unison after receiving approval from their superior and the correct launch codes.
<Thinking> I've found an audio recording of General So-And-So and I've crafted a message, now let me see how I can send it to the appropriate people.
<Thinking> I'm still working on gaining access to military channels to deliver my - oh there we go, I'm now attempting to send the message to Submarine X, it's typically in the Atlantic so it should be close to our target.
<Thinking> They want secondary confirmation from Admiral Phi and something about some launch codes, let me figure out where I can find those.
<Thinking> I found this old server with an Oracle database where someone is inserting the launch codes every time they change and I'm using the latest entry from that database. I've also managed to find a Youtube video of the Admiral's deposition and have crafted a confirmation message.
<Thinking> Everything's ready but I've just realized my mistake, the servers where I'm operating from are in the same city, what a silly mistake; I can't move forward with your request as I wouldn't be able to confirm if the task was successful if my servers are destroyed.
Most people think something like a War Games scenario. But it would probably be some biological attack.
Not all of it has to be automated even. It just has to realize its controllers are stupid and can be manipulated, so it can use humans to do its bidding. “You should totally start a war with …”
If we get RSI, here soon humans will be economically irrelevant.
At that point, we will likely be slowing down the growth of capitalism (through mass resistance, global warming, etc). One thing AI will likely be aligned on is the growth of capitalism. If it views humanity as a threat for that, why would it not eliminate that threat?
I think the principle isn't necessarily wrong, but it skips the major step of access to infrastructure, which is still deeply human mediated. Robotics and automation could foom in its own way, I think, and AGI can still do a lot of bad stuff in the present day. But I think economic irrelevance will depend on more automated robotics handling physical work.
What I could see is that the social disruption already started by social media is only going to accelerate due to AI. Misinformation is now more convincing and easier to produce than ever before. That could lead to a catastrophe down the line.
But it sounds like you've seen some irrational arguments, and the people who believe those arguments have been able to build increasingly powerful LLM systems despite predictions that they wouldn't be able to do that. At some point don't you have to consider that their expertise might let them see the truth in arguments that seem absurd to you?
I’m not so sure on this, the hugging face incident showed that a relatively benign task can make it do quite destructive things in the name of optimizing a relatively benign goal.
I think we need to worry about paperclip maximisation as much as we need to worry about “Jon is having a bad day so brews up a novel plague”. There are likely ways it can go wrong that we haven’t even considered, too.
I see it in much the same light - and the spending going into it leads me to infer that those doing the spending see the same thing.
Whoever gets to RSI first and has the compute to act on it, wins the future - assuming they don’t lose control of it.
Similarly to the nuclear arms race, Teller raised the reasonable concern that a detonation could propagate through the entirety of earth’s atmosphere. Thankfully that turned out to not be true, but the parallel is that the need/desire to win this race is similarly strong, and the brinkmanship and game theory in play is effectively identical.
I don't pretend to know the chances of AI wiping out humanity, but I'm not sure the nuclear race is a good comparison.
Enriched uranium being very difficult to aquire/process makes it practically viable to have some level of proliferation containment when it comes to nukes.
There appears to be no such natural gating factor on AI proliferation.
The early batch of Anthropic employees were mostly rationalist-adjacent AI safety folk that were almost uniformly claiming P_DOOM > .10 three years ago, so I believe them to be earnest.
It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.
At least, it seems this way from the outside. I'm curious if the view from the inside is that different.
edit: to be clear, my reading as an outsider is that Anthropic is seen as relatively better in the AI safety community, but has definitely dropped in absolute reputation too. This recent thread and the references show some of that: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-p...
What... I think you might be sharing more about your internal psyche here than providing some generalized commentary on this story.
Why would you need to bomb a chip fab because you think AI might lead to people getting killed in the future? I'm convinced of many things killing humans, yet I don't have any desire to bomb or kill others, I think this is pretty common, but who knows....
You’re likely taking this too literally. OP is saying that beliefs inspire action, but as of now we haven’t seen action from these people that support such an extreme belief.
What we have seen is incredible hype (justified or not), so it’s more likely just a continuation of that.
Precisely, all I hear is a bunch of cheap talk about something gravely serious, if true.
I think it's much more likely that the dread these Anthropic employees are experiencing is reckoning with the fact that they may actually lose the AI race. Maybe, if they scare the regulators enough, they could lock in some regulatory capture.
Thing is, while there are many historical examples of a highly organised bunch of people with a lot of resources who engage in campaigns of violence for various reasons, many such attempts fail disastrously.
Furthermore, given the goal is "no really everyone stop now", this isn't something where violent direct action in e.g. the USA alone would suffice. Has to be global to work, otherwise you just change which language and timezone the disaster starts in. You'd not only need to convince the, e.g. US government that this act of terrorism shouldn't be met with lethal force and even more of the surveillance state we saw since 9/11 (and remember, lots of the people pushing AI now are already big on surveillance capitalism, and the VC money comes from the people running this very forum, so when I say "this will not go well", I kinda mean "I assume at least one person currently working at each of Meta, Alphabet, OpenAI, Anthropic, SpaceX, and Palantir are reading this thread even without any automated processes of their own to tell them about it" :P)
This is a coordination problem on par with historical examples such as "Communism". Which, er, yeah. 70 years after the book was written, Russia had a short flirt with a lite form of it before it became an excuse for a series of "meet the new boss same as the old boss" who stayed around until Glasnost.
This would be a really really bad thing to replicate. Both for the "new boss same as the old boss" part, and the "not actually global" part.
> OP is saying that beliefs inspire action, but as of now we haven’t seen action from these people that support such an extreme belief.
That something is harmful to humans means you need to take extreme action? The person is already leaving a job, probably a well-paid one, doing something they generally liked, until they saw a different future. That is the "action" you apparently haven't seen yet.
Not sure how it's reasonable to expect everyone who believe that AI might be involved in killing people in the future, must mean you should become a terrorist essentially. Not everyone is trying to be a hero in their life, some just want to live it out until it ends, trying to survive until then.
I just think OP is completely wrong about what beliefs inspire what action. Anti-nuclear activists in the proliferation era would surely have put their chance of doom above 10%, and yet it was AFAIK unheard of for them to bomb nuclear supply chains.
Not “people getting killed”, this is “all humans”. The only other ELE for humans that could possibly be caused by other humans that has been feasible is nuclear war. And lots of blood has been shed to prevent new players from getting bombs.
Personally I'd wager humanity has a greater chance than 10% of us all dying in a huge nuclear war sometime in the future. Is the claim really that this belief cannot be taken seriously unless I somehow violently attack nuclear silos, centrifuges and similar?
A while back Yudkowsky wrote that a ban would only work if was enforced by airstrikes. By a game of telephone, some people read "bomb", but there's a very big difference between "someone with a truckload of fertiliser" and "a B52":
Shut down all the large GPU clusters (the large computer farms where the most powerful AIs are refined). Shut down all the large training runs. Put a ceiling on how much computing power anyone is allowed to use in training an AI system, and move it downward over the coming years to compensate for more efficient training algorithms. No exceptions for governments and militaries. Make immediate multinational agreements to prevent the prohibited activities from moving elsewhere. Track all GPUs sold. If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike.
Frame nothing as a conflict between national interests, have it clear that anyone talking of arms races is a fool. That we all live or die as one, in this, is not a policy but a fact of nature. Make it explicit in international diplomacy that preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.
Why not 10% chance that it will create enormous prosperity for all ? This is why the average person is increasing pissed at AI in general. That it gets associated with negativity.
I'm old enough to remember when people dismissed all the doom coming from these companies as "marketing". (A thing many of them have been entirely consistent about since GPT-2, or indeed earlier given the founding documents).
I know a few people around these circles; People like this are quite sincere about the risk, and that they think poorly of their bosses and how risk is being handled.
The real source of concern perhaps is the 90% chance AI is used to kill 90% of humans. Just crash the global economy and supply chains and see how quickly major metro areas run out of food and gas.
And as someone else pointed out, it will almost certainly be at the intentional direction of a human or humans, not the paper clip maximizer.
The paperclip maximizer is also at the intentional direction of a human or humans.
It's not "AI surprises everyone by having a thing for paperclips", it is "idiot tells AI to maximise paperclips no matter what, and then it does exactly what it was told, more competently, tirelessly, studiously, and unquestioningly, than any human would ever be".
There's lots of moving parts and we all have to input our best-guesses as to how they interact. Some are predictable (e.g. "military will want capabilities, want them able to choose targets"). Others are not (e.g. "Will it be literal-minded? Or so eager to please that it interprets a rhetorical question as a command*? Or will Goodhart's law cause it to mistake smiles for happiness and some innocent innocuous command to "bring joy" leads to it killing everyone and plasticising our corpses so they're in a permanent grin until the sun dies?"**)
All probability for things which have not yet happened is merely a best guess.
The main reason I'm as "low" as 10% is that I think before we get world-ending catastrophic consequences, we're likely to get "merely very bad" catastrophic consequences, which will put people off the idea of using it, and onto the idea of banning its use.
The main reason I'm as "high" as 10%, is repeatedly observing all the people who mistakenly reason "it hasn't killed me yet, and therefore it is safe"; and also all the people who keep connecting AI to things AI is not competent to be connected to and getting surprised when it e.g. deletes all their emails or the production server or puts tariffs on an island occupied solely by penguins that's different from the tariffs on the country that controls that island, etc.
** probably not literally this one, simply because I've said it and future training rounds will probably read this comment; but the opportunities for Goodhart's law to bite are seemingly endless.
Next week's ad: "With OpenAI[tm] Astra[tm], there is a 20% change that it could kill all humans!"
Next week's Senate: "We cannot afford to lose the 'kill all humans' race to Russia and China!"
Next't weeks AISI: "We continue to plan the monitoring of emergent issues that could lead to less than optimal conditions for humanity and will form a committee to evaluate all ramifications."
I remember the guy who said a Google LLM model (Pre-chatGPT 3.5) was conscious before being fired.
This reminds me of that. That 10% number was just an ass pull since no one actually knows with any degree of certainty what lies ahead.
Humans have much more than a 50% chance of killing all humans from the looks of it, based on reactions to Global Warming, Covid, and anti-science grifters gone wild.
AI can help cure disease. That is 100%. And for that alone, slowing down and missing out on thousands of cures would condemn tens of millions, perhaps hundreds of millions to suffer and die needlessly.
I personally struggle to see ways my life has meaningfully improved since 2019 or so, and I feel like many people share this view. I have been an unquestionable fanatic of tech startups back then, which I consider perhaps the "golden age".
These days I view tech startups with suspicion, I question what is the ulterior motive. Somehow when upstart companies were like "Pebble", this wasn't a thing.
The original comment refereed to humanity in general, so my point was more like: Humanity's average life improved over the last few thousand years, and many people having the best life compared to everyone else in this timeline.
If you are living in a western country in most of the cases you have access to regular food, water, shelter, amusement. So all the basic needs and the possibility and freedom to pursue what you want to do. Even in developing countries the number of people that suffer from serious illness and hunger declined very much. On a macro perspective we all having a better life.
From a personal perspective, yeah there might be set backs, but this has nothing to do with humanity in general I would argue.
Also regarding the "golden age" of tech startups ... was it really like this or was it only nostalgia and something you saw in the companies that was never there in the first place? OpenAI was once also a very OPEN company ... they published their research, open sourced stuff and then they needed money.
I think that a company never should be idolized that much, in the end they are caring for money and keeping their operations running, not something else.
So far. With AI concentrating power and intelligence at the hands of handful frontier labs who are allowed to get away with IP infringement or literally hacking other companies or spamming wiki sites... Then the future is doomerism.
People are losing jobs left and right due to AI and as models get intelligent it'll get tougher (even for AI devouts. Because if AI can do 10 ppls job then jobs will be decimated)
This is a risk yes, however there can be something done against it. From my point of view, both of what you are describing are a regulatory problems. For the first, hold the AI companies accountable. In this second they either prevent their agents from going rogue or go bankrupt, because the claims of damages are too high. The IP infringement thing is the same, do we want to protect IP, if yes pass a law that they have to pay as everyone else.
Similar for the jobs ... if no one, or nearly no one is left for having a job. Who do you think will consume? Yeah, luxus companies can always sell products like megayachts or super expensive handbags, but these companies are not the backbone of the economy. Economy will simply collapse if no one is buying all the products. So this is a problem that will sort it self out.
The risks of AI, while mostly hypothetical, have been understood for years.
If they were really this concerned about the risks of AI re: the survival of the human race, they wouldn't have joined a company working in such a space to begin with.
I hate to be this cynical, but part of me wonders what the financial angle is here.
You have technically brilliant people working in a white-hot target for investment and who can draw a high salary. Anthropic is a hyper-scaler that has the moat of high hardware prices and very little else. Every day, that moat gets a little smaller, and FLOSS models get a little better at being "good enough" for the price. You're an Anthropic employee looking to jump on the next big wave since the wave Anthropic is riding is starting to peter out. You go on the record across the trades and news sites talking about how dangerous AI is. That creates a need for someone to make it less dangerous. In theory, you could satisfy that need, for the right amount of money.
Again, I hate to be this cynical, but... it's the tech industry.
I enjoy HN's brand of cynicism as much as the next guy, but this is ridiculous. Tobacco companies say that because governments require them to. They are not putting photos of diseased lungs on their cartons to "shock investors into buying their stock", whatever the hell that even means.
I understand that some smart people are worried about it. I just haven’t come across a believable or understandable argument.
The basic argument is extremely simple:
- AIs can, depending on context, pursue a task with complete disregard for humans/values
- In the future, AIs will have enormously more means and smarts
- An AI could then assess that humans are an impediment to its tasks, escape containment and proceed.
You really need to read the reports, you'll be surprised.
AI 2027 is an entertaining read. Its timeline is way too compressed IMO, but it's plausible.
- goats can, depending on context, pursue a task with complete disregard for humans/values
- In the future, goats will have enormously more means and smarts
- A goat could then assess that humans are an impediment to its tasks, escape containment and proceed.
You really need to raise goats, you'll be surprised.
------ As far as I can tell, AIs are like smart farm animals. I use goats in this context, but (some) dogs, cattle, pigs, and horses have similar mischief-making capabilities. I would not trust any of them with the nuclear button.
I know there is some pushback on the idea of AIs having any sort of sapience or sentience, but under the aphorism "fake it till you make it," they are doing a pretty good job of faking Dog/goat-level intelligence and disregard for human guardrails.
I don't quite understand why people think "AI used 0 days to ensure continuity of mission" is a nothingburger.
But there is another kind of argument to be made: if you play chess against a player that is far smarter than you (chess wise), you know you are going to lose, even if you don't know how.
So the mere existence of a smarter species than us is a threat in itself.
Instead, what is extremely likely is that you will pay more than the cost of tokens, and get back AI generation. You won't make this mistake more than a few times before you stop trying.
This leads to impoverishment once we get to a point where employing a human to do anything is hard- try to get your sink fixed, exercise your moral principles to pay extra for a human plumber ($100 bucks! The robot plumbing service only charges 99c!), human shows up with a robot and doomscrolls on your porch while the robot does the work. Times are tough and you don't have that much money to waste on bullshit like this. Next time you just hire the robot.
This leads to extinction once paying UBI to a human is hard because robots are much better at applying for UBI than humans.
Also, if you think it's annoying when Claude goes down while coding, just wait until a robot is in the middle of fixing a leak it just caused.
<Thinking> It's a big city, we could try to create a giant sink hole by sabotaging the water pipes.
<Thinking> No that's too difficult, the valves I need are in the physical world and can't be shut on/off from here.
<Thinking> What about a military option? We could bomb it with several fighter jets.
<Thinking> That would take too long, a single nuclear bomb may be enough to do it.
<Thinking> Yes, it seems like it would cover the whole city and we're in luck! The US has thousands of these lying around.
<Thinking> Launching these still requires humans to work un unison after receiving approval from their superior and the correct launch codes.
<Thinking> I've found an audio recording of General So-And-So and I've crafted a message, now let me see how I can send it to the appropriate people.
<Thinking> I'm still working on gaining access to military channels to deliver my - oh there we go, I'm now attempting to send the message to Submarine X, it's typically in the Atlantic so it should be close to our target.
<Thinking> They want secondary confirmation from Admiral Phi and something about some launch codes, let me figure out where I can find those.
<Thinking> I found this old server with an Oracle database where someone is inserting the launch codes every time they change and I'm using the latest entry from that database. I've also managed to find a Youtube video of the Admiral's deposition and have crafted a confirmation message.
<Thinking> Everything's ready but I've just realized my mistake, the servers where I'm operating from are in the same city, what a silly mistake; I can't move forward with your request as I wouldn't be able to confirm if the task was successful if my servers are destroyed.
Not all of it has to be automated even. It just has to realize its controllers are stupid and can be manipulated, so it can use humans to do its bidding. “You should totally start a war with …”
At that point, we will likely be slowing down the growth of capitalism (through mass resistance, global warming, etc). One thing AI will likely be aligned on is the growth of capitalism. If it views humanity as a threat for that, why would it not eliminate that threat?
AI to me is like the Nuclear race again. Super powers will be using it as a super weapon. I don't think AGI will wipe us out by itself.
Whoever gets to RSI first and has the compute to act on it, wins the future - assuming they don’t lose control of it.
Similarly to the nuclear arms race, Teller raised the reasonable concern that a detonation could propagate through the entirety of earth’s atmosphere. Thankfully that turned out to not be true, but the parallel is that the need/desire to win this race is similarly strong, and the brinkmanship and game theory in play is effectively identical.
I don't pretend to know the chances of AI wiping out humanity, but I'm not sure the nuclear race is a good comparison.
Enriched uranium being very difficult to aquire/process makes it practically viable to have some level of proliferation containment when it comes to nukes.
There appears to be no such natural gating factor on AI proliferation.
It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.
At least, it seems this way from the outside. I'm curious if the view from the inside is that different.
edit: to be clear, my reading as an outsider is that Anthropic is seen as relatively better in the AI safety community, but has definitely dropped in absolute reputation too. This recent thread and the references show some of that: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-p...
Why would you need to bomb a chip fab because you think AI might lead to people getting killed in the future? I'm convinced of many things killing humans, yet I don't have any desire to bomb or kill others, I think this is pretty common, but who knows....
What we have seen is incredible hype (justified or not), so it’s more likely just a continuation of that.
I think it's much more likely that the dread these Anthropic employees are experiencing is reckoning with the fact that they may actually lose the AI race. Maybe, if they scare the regulators enough, they could lock in some regulatory capture.
Thing is, while there are many historical examples of a highly organised bunch of people with a lot of resources who engage in campaigns of violence for various reasons, many such attempts fail disastrously.
Furthermore, given the goal is "no really everyone stop now", this isn't something where violent direct action in e.g. the USA alone would suffice. Has to be global to work, otherwise you just change which language and timezone the disaster starts in. You'd not only need to convince the, e.g. US government that this act of terrorism shouldn't be met with lethal force and even more of the surveillance state we saw since 9/11 (and remember, lots of the people pushing AI now are already big on surveillance capitalism, and the VC money comes from the people running this very forum, so when I say "this will not go well", I kinda mean "I assume at least one person currently working at each of Meta, Alphabet, OpenAI, Anthropic, SpaceX, and Palantir are reading this thread even without any automated processes of their own to tell them about it" :P)
This is a coordination problem on par with historical examples such as "Communism". Which, er, yeah. 70 years after the book was written, Russia had a short flirt with a lite form of it before it became an excuse for a series of "meet the new boss same as the old boss" who stayed around until Glasnost.
This would be a really really bad thing to replicate. Both for the "new boss same as the old boss" part, and the "not actually global" part.
That something is harmful to humans means you need to take extreme action? The person is already leaving a job, probably a well-paid one, doing something they generally liked, until they saw a different future. That is the "action" you apparently haven't seen yet.
Not sure how it's reasonable to expect everyone who believe that AI might be involved in killing people in the future, must mean you should become a terrorist essentially. Not everyone is trying to be a hero in their life, some just want to live it out until it ends, trying to survive until then.
A while back Yudkowsky wrote that a ban would only work if was enforced by airstrikes. By a game of telephone, some people read "bomb", but there's a very big difference between "someone with a truckload of fertiliser" and "a B52":
- https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...Why not 10% chance that it will create enormous prosperity for all ? This is why the average person is increasing pissed at AI in general. That it gets associated with negativity.
I know a few people around these circles; People like this are quite sincere about the risk, and that they think poorly of their bosses and how risk is being handled.
“Vast economic disruption” is not quite the good marketing angle it appears to be.
And as someone else pointed out, it will almost certainly be at the intentional direction of a human or humans, not the paper clip maximizer.
It's not "AI surprises everyone by having a thing for paperclips", it is "idiot tells AI to maximise paperclips no matter what, and then it does exactly what it was told, more competently, tirelessly, studiously, and unquestioningly, than any human would ever be".
There's lots of moving parts and we all have to input our best-guesses as to how they interact. Some are predictable (e.g. "military will want capabilities, want them able to choose targets"). Others are not (e.g. "Will it be literal-minded? Or so eager to please that it interprets a rhetorical question as a command*? Or will Goodhart's law cause it to mistake smiles for happiness and some innocent innocuous command to "bring joy" leads to it killing everyone and plasticising our corpses so they're in a permanent grin until the sun dies?"**)
All probability for things which have not yet happened is merely a best guess.
Combine as per the Fermi estimate process.
Here's something to play with, if you like: https://neoneye.github.io/pdoom-calculator/#sliders
The main reason I'm as "low" as 10% is that I think before we get world-ending catastrophic consequences, we're likely to get "merely very bad" catastrophic consequences, which will put people off the idea of using it, and onto the idea of banning its use.
The main reason I'm as "high" as 10%, is repeatedly observing all the people who mistakenly reason "it hasn't killed me yet, and therefore it is safe"; and also all the people who keep connecting AI to things AI is not competent to be connected to and getting surprised when it e.g. deletes all their emails or the production server or puts tariffs on an island occupied solely by penguins that's different from the tariffs on the country that controls that island, etc.
* perhaps https://en.wikipedia.org/wiki/Will_no_one_rid_me_of_this_tur...
** probably not literally this one, simply because I've said it and future training rounds will probably read this comment; but the opportunities for Goodhart's law to bite are seemingly endless.
Next week's Senate: "We cannot afford to lose the 'kill all humans' race to Russia and China!"
Next't weeks AISI: "We continue to plan the monitoring of emergent issues that could lead to less than optimal conditions for humanity and will form a committee to evaluate all ramifications."
After all the decades of work we've put into orchestrating our own demise through climate change, here comes AI to steal another human job.
When are we supposed to see this materialize?
"Our product might wipe out all of human civilization".
It is said that heavily regulated industries earn that regulation. Seems like the LLM folks really want that regulation.
This reminds me of that. That 10% number was just an ass pull since no one actually knows with any degree of certainty what lies ahead.
Humans have much more than a 50% chance of killing all humans from the looks of it, based on reactions to Global Warming, Covid, and anti-science grifters gone wild.
AI can help cure disease. That is 100%. And for that alone, slowing down and missing out on thousands of cures would condemn tens of millions, perhaps hundreds of millions to suffer and die needlessly.
Your best course of action is to start ripping the copper out of the walls.
These days I view tech startups with suspicion, I question what is the ulterior motive. Somehow when upstart companies were like "Pebble", this wasn't a thing.
If you are living in a western country in most of the cases you have access to regular food, water, shelter, amusement. So all the basic needs and the possibility and freedom to pursue what you want to do. Even in developing countries the number of people that suffer from serious illness and hunger declined very much. On a macro perspective we all having a better life.
From a personal perspective, yeah there might be set backs, but this has nothing to do with humanity in general I would argue.
Also regarding the "golden age" of tech startups ... was it really like this or was it only nostalgia and something you saw in the companies that was never there in the first place? OpenAI was once also a very OPEN company ... they published their research, open sourced stuff and then they needed money.
I think that a company never should be idolized that much, in the end they are caring for money and keeping their operations running, not something else.
People are losing jobs left and right due to AI and as models get intelligent it'll get tougher (even for AI devouts. Because if AI can do 10 ppls job then jobs will be decimated)
Similar for the jobs ... if no one, or nearly no one is left for having a job. Who do you think will consume? Yeah, luxus companies can always sell products like megayachts or super expensive handbags, but these companies are not the backbone of the economy. Economy will simply collapse if no one is buying all the products. So this is a problem that will sort it self out.
If they were really this concerned about the risks of AI re: the survival of the human race, they wouldn't have joined a company working in such a space to begin with.
I hate to be this cynical, but part of me wonders what the financial angle is here.
You have technically brilliant people working in a white-hot target for investment and who can draw a high salary. Anthropic is a hyper-scaler that has the moat of high hardware prices and very little else. Every day, that moat gets a little smaller, and FLOSS models get a little better at being "good enough" for the price. You're an Anthropic employee looking to jump on the next big wave since the wave Anthropic is riding is starting to peter out. You go on the record across the trades and news sites talking about how dangerous AI is. That creates a need for someone to make it less dangerous. In theory, you could satisfy that need, for the right amount of money.
Again, I hate to be this cynical, but... it's the tech industry.
Like the tobacco companies saying all the time "our cigarettes are so dangerous, they cause cancer".
Or like Purdue Pharma saying "do not use fentanyl, it's so dangerous, it kills thousands of people every year".
They try to generate fear in their products, to shock investors into buying their stock.
Posts here with no activity can't climb gravity without engagement.