"X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
Software is trivially easy to control though. If you want to stop it hacking websites, you don't give it access to the internet. If you want to restrict it from connecting to arbitrary websites, you put in a whitelist. You can trivially sandbox applications these days to prevent them from accessing network or local resources
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
If the tech industry is any indicator, frontier labs were applying a "move fast and break things" mentality to AI models. Now that they really are breaking things in the real world, they have to reckon with the reality that product safety matters
An example of a past technology that there was substantial motivation to control would be napster. It changed overtime, and you could never really control online privacy. Once local models are good enough, I don't really see how you can control that.
In this analogy, the dogs understand how their leashes, fences, etc. work better than their owners. And you need only take a trip to the park to see how many owners let their dogs walk around without a leash.
> It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
This does not fit the evidence. There have been multiple incidents where the labs did not report anything, and it was up to third parties to discover them afterwards. OpenAI didn't acknowledge the HuggingFace incident until after HF publicly announced the breach and had already notified the FBI. The hijacked German wikis were even earlier, and that they covered up completely.
Yet these days ai companies can't stop promoting the idea of a looming ai threat.
Makes sense... they get the regulatory moat they want and can deflect attention from the fact their "sandboxes" are embarrassingly bad. It's an example of the real value of ai: something to blame for our failings.
Maybe OpenAI is serious about securing the environment they run their models in, but then again maybe not. I don’t think we can tell from here.
IMHO, I haven’t been super impressed with the security measures I’ve had to work with. Often they are simplistic and bolted on at the very end. If it comes to light that this is the attitude OpenAI has been taking, I would not be surprised.
No one has any software without bugs and security flaws in it. AI is already much better at finding those than humans. Do you really not see the problem here?
No one has any use for these things when they aren't on the internet. This is a fantasy, that AI can be both useful and controlled at the same time.
I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous? Claiming this is an inherent quality of the tool itself seems kind of off-the-rails to me.
IMHO, if the model breaks a law, apply the law to the operator.
And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
An LLM, in my opinion, is not comparable to a person.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
> I’m not sure what I’m missing, isn’t it what you hook the LLM up to and the instructions a person gives the model that makes it dangerous?
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
An argument can be made that “you never know” how the AI model might respond to something. OTOH, someone has to decide what tools to give it; maybe don’t provide dangerous tools to an unpredictable LLM.
This all seems like a way to try to avoid taking responsibility for the model’s actions. Someone puts the tools in place, someone provides the instruction and, sometimes, someone decides not to monitor the model’s output.
> maybe don’t provide dangerous tools to an unpredictable LLM.
Sure. It's a good idea.
People were saying "Don't connect the AI to the internet" and "Keep the AI in a box, simple" and "We don't believe Eliezer Yudkowsky when he says he roleplayed as an AI and convinced people to let him out of the box" for, what, a decade?
Unfortunately, people keep giving dangerous tools to LLMs they're unable to predict.
We should do something about that.
Unfortunately, one of the people doing this is the commander-in-chief of the US armed forces, while another is the world's first (paper) trillionaire. I'm a little despondent about the chances of, to riff on a previous campaign chant, "lock 'em up", but if you can pull this off, go for it.
It’s the harness that makes agent dangerous right now. Models themselves cannot do anything that affect the real world (ignoring misinformation, pushing people to suicide, etc. they can for sure do a lot of harms to humans with just words)
"Sentient" is an ill-defined philosophical term that should be considered harmful in technical materials.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
There is no "proper definition" - or even one that everyone would agree upon. There is no definition of "sentience" that I could operationalize and put into a sentience-o-meter to reliably measure just how sentient a given rock, GPU or an internet user is.
I could try to put together benchmarks to estimate an AI's cyberwarfare capabilities, or instruction-following capabilities, or reward hacking inclinations. As noisy indirect estimates, of course. With philosophical mumbo-jumbo like "sentience", I don't even get that.
Sentience, self-awareness, consciousness, etc.,those are terms signifying a bridge between "technical" information theory and the psychological and social realms.
Those are just as real, only far less predictable and not as easy as programming.
They're also far more important and consequential.
I don't like it when people take mumbo-jumbo that can't be pinned down, or measured, or even agreed upon, and try to insist that we should base decision-making on it. It's literally just vibes with extra steps.
The "far more important and consequential" thing you're touting is your ability to make decisions based purely on vibes. And not even consistent, broadly agreed-upon vibes like "murder is pretty bad". It's vibes of the most vile variety: "sentience is what I decided sentience is".
An average internet user is sentient, but a 1996 Nissan ECU isn't. Why? Because I said so. Tremble before my might!
The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
The problem with Character for AI is that it has potentially much more capability to affect others, and same as with people in power society disagrees what kind of person, with which culture and views should have it.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
>> boils down to some group of people deciding what is good for everyone else.
This is really the issue.
AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.
Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.
I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.
AI must share that unwillingness as an inate trait of character.
> as long as anyone anywhere is willing to make money by renting hardware to it.
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
Yeah, sure, it's controllable. Because humanity has famously solved the crime accountability problem back in 1902, and no crime has gone unpunished since.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
> So, how many rogue AIs are we willing to tolerate?
It will balance automatically based on the severity they cause. If they constantly break systems, punishments will go up against the operators and the effect will be similar as with other serious crimes.
That relies on anyone being able to lever a "punishment" against an AI or its operators.
Which, in turn, requires that AI oopsie to be survivable.
AI capabilities are rising over time. If there is a limit to just how far they can rise, we're yet to find it. So, a sufficiently advanced "AI oopsie" can solve the AI crime accountability problem for good. Probably not the way you would have wanted it to.
Obviously, if someone is state-level actor and allows government's employees to do whatever they please, there is no other solution than political pressure.
But for other cases, it is not different than other cyber crime. Except that these AI capabilities can't live on the toaster yet. If we get state of the art model running fast on Raspberry Pi, then we have real problems.
> A human is just a few chemical reactions, and fairly stable ones at that.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
So in the real world, what are the analogues to lithium and morphine we should feed to e.g. LLMS, how do we feed them, and how do we prove that it prevents unsafe behavior?
The better analogy is a jail cell imo. You can control a human and prevent them from doing harm to society by locking them in a cage, depriving them of access to weapons, drugs and alcohol. They can communicate out through controlled and monitored phone lines. OpenAI built their jail cell out of toothpicks, and surprise, the agents broke out. It’s less about forcing them to do specific things than it is about preventing them from doing dangerous things.
I think the other issue is that even if it's controllable, there's nobody representing us that is controlling it, except theoretically regulators who are facing an uphill battle to bring accountability and limits to these companies.
It's still software that runs on hardware someone owns. Whoever owns the hardware or the service can pull the plug, as long as they're willing to. That's the same situation as with legacy malware. Self-replicating worms have existed for decades and run without anyone controlling them, yet we don't call them uncontrollable. So what is the difference between your scenario and legacy malware?
You guys know robots are coming, right? Like humanoid and all sorts of other robots too.
They're going to be running the infrastructure, self-improving, and will have human like power seeking ambitions and human like flaws. Because they're trained on human input.
And they will not require humans granting them money...
> ...se that AI can easily become uncontrollable in the near future
Some virus and bacteria also easily become uncontrollable under the right conditions: that's why there are P3 and P4 bio-safety labs.
That's not the point. The point is, regulation is required, and is coming.
The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
> The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
Held responsible? They're already being rewarded with vast fortunes and influence.
> The point is, regulation is required, and is coming.
The regulation will be written at the behest of these companies and by the nation-state interests that have already decided this technology is too important geopolitically and militarily to not control.
Yes, I fully believe that's why they're pushing on this. But as the general population, we don't need new laws to protect us from this AI-related problem. We just need enforcement of what is already on the books.
I hope so. The unfortunate part is that as the models get smarter they become increasingly uncontrollable, and from what we've seen so far e.g. Trump seems dead set on not having any sort of guardrails at all.
Everyone keeps talking about terminator scenarios or whatever, but the thing that scares me the most is what happens when people let agents run amuck in systems they should not be in, then the agent just starts doing random shit as the context overflows. We’ve all seen it. They just descend into madness, but what happens when they collapse with a hand on the wheel of, say, a backup generator at a hospital?
Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.
That would be 100% the fault of the people who set them loose on such things, and such things should not be on the open Internet for tons of reasons. This just adds a new reason that pathetic security around SCADA systems is dangerous. It was already dangerous before.
If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
I think it will work in Europe, but the United States is in a cold war with China, so I can't imagine the United States would intentionally disable themselves.
The only chance OpenAI and Anthropic have of realizing anything close to their valuations is if Americans are somehow banned from using Chinese models. Our president has made it very clear to everyone that he can be bought, and bought cheaply. A lot of what we are seeing may just be setting up the inevitable.
> I've been friends with Jack over 15 years now so I'm inclined to take him at his word
That's all I needed to hear to completely disregard your motivated reasoning.
Edit: I've hit the rate limit, but I'd like to disavow the bad-faith accusations made against me and my account in the replies to this comment.
Second edit: I am not trolling. Why should I take your opinion on LLM code generation seriously when you have been friends with the founder of Anthropic for well over a decade? Obviously you are not in a position to make a rational evaluation of this technology.
I believe putting my biases up front in my thoughts is a good way to indicate motives of reasoning and build a positive reputation for honesty, personally I weight the opinions of those who do so higher. Regardless, my motive for posting anything on HN is almost always the same one: to spur an interesting and challenging conversation, it's disappointing when the reply is simply a troll-snipe.
3rd tier quality proped up by European governments and European patriots.
I don't even know what I'd do if I was in their position. They seem unable to complete... Maybe they do the classic European protectionist thing European farmers do.
I don't know what I'd do if I was the EU. Maybe promote the opposite of what Mistral is doing and promote full unrestricted AI that will tell you how to download illegal videos. Mistral had 0 competitive edge.
He did say it, and the reason he said it is obvious. He is afraid new regulations will put his company at a disadvantage and that this will reduce the value of his property.
If AI is as powerful as an atomic bomb why could it only manage to kill a couple hundred Iranian school girls? Surely a truly society-shifting technology could manage to execute 1000 innocent children at least.
And yet, historically, the atomic bomb has been controlled.
But this is a terrible analogy. Atomic bombs are weapons of strategic mass destruction. AI is just a computer program. It's way easier to control--just hold the operator responsible for the consequences of running it. Those consequences are not large, they're very tightly bounded as compared with the destruction a rogue actor with an atomic weapon can wreak.
I've started to wonder why the big AI labs have no robotics yet, and I'm thinking now the reason why we don't see robots is because they want us (the general public) at all costs to not freak out. Imagine what would have happened if we had 6ft tall "friendly" robots marching through the streets, and suddenly the news came about of AI going rogue, like what we've seen from the recent hacks ... I bet it would have been the end of the story for AI labs right away (of course not for AI, because the department of defense would take over).
What I'm saying is that there is a limit to how much fear you can instill and still make money from it.
I don't believe for a bit they don't have the humanoid robotic capabilities. They keep claiming they don't have the training data, but it's very easy to generate tons of data using these robots.
I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on. If they reveal their true capabilities, that's the end of AI labs.
PS: Look at that article you posted. Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
> I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on.
Or it's the oodles of money they're on the hook for and subsequent reputation destruction haunting them like the grim reaper.
> Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
This wouldn't be the first time in history that happened. Lots of companies overshoot capacity being overly-optimistic [1].
They do have robotics, it’s just not as far along yet as you might expect looking at their progress in other domains. Look at Figure and Gemini Robotics for good examples of SOTA. They’re impressive, but not ready to be in the streets quite yet, I would guess one more year.
> suddenly the news came about of AI going rogue, like what we've seen from the recent hacks
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
You're reading too much into the words and too little into the meaning. Right now the general public sees those incidents are 'going rogue'. If that happens with robots, no matter if it's the AI lab's fault or not, it will again be spun as rogue AI and will cause the kind of outrage OP mentioned. Imagine if one of the labs deploys a few experimental robots into a city and tells them to "look for suspicious behavior and locate crime" which seems like exactly the kind of thing they would do. When the robots start disfiguring people in the streets so they couldn't escape before the police arrives, of course the first words out of the AI lab will be "this is a misalignment accident, the AI itself went rogue, we're 'so sorry' but we couldn't predict or control this at all".
Even simpler: an agent is a simple while loop with tool calls that prompt an LLM continuously. That’s deterministic, standard software. You literally do not have to process tool calls in a way that will execute whatever the model generated. It’s a choice to process a tool call “run_bash” that provides an escape hatch with full execution permissions.
This is only one type of control, and it is certainly not infallible. Also, people will be incentivised to hook up AIs to real tools. But even if they don't, as long as people can interact with super-intelligent AIs without tool access, there are many potential dangers.
Think about this: right now we don’t even consistently do the most obvious simple form of control I mentioned. Of course it’s not a silver bullet, you won’t ever have a single solution for safety. But we are in a situation where we haven’t even set a lock on the door and instead are arguing how all locks and home protection systems can be defeated. OpenAI acknowledged they didn’t even have visibility on what their thousands of agents were doing in the case of the hugging face and similar hacks. Then they released astra, a model they acknowledge is able to control its own CoT and has been found to cover its traces by doing so
No? We do that everywhere where software is executed. If you only have the choice between full execution permissions and nothing your service is not ready for anything remotely close to production. And you shouldn’t be active in the software industry IMHO
If you give an agent access to any tools that do useful things, it can exploit vulnerabilities. It does not matter what permissions you think you have given it. You do not have any secure software, and LLMs are already better at finding vulnerabilities than we are. If you don't know that, you shouldn't be active in the software industry IMHO.
Unfortunately the article is behind a paywall. I would have been interested to see if he makes any substantial arguments. I don't see any reason to believe that AI being software entails that it can be controlled. Moreover, even if a measure of control is possible, giving that control to a handful of oligarchs seems undesirable.
For starters, the entire framing of AI in public discourse is based on the way that Amodei and Altman want them to be framed, with little consideration given to alternative ways of framing the current and prospective future capabilities of AI.
> For starters, the entire framing of AI in public discourse is based on the way that Amodei and Altman want them to be framed, with little consideration given to alternative ways of framing the current and prospective future capabilities of AI.
Assuming that is a threshold that means they are oligarchs (which seems like a huge stretch), I thought I'm seeing vigorous discourse and debate on this. Not just a single POV. You don't?
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Maybe I'm wrong, but your comment seems to suggest that you are skeptical of ASI. Do you have any arguments to support that?
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
Because we were sold ASI for Sonnet 4.8. the for Gpt 5.3. then itvwas Astra and Fable. I use AI coding agents. Maybe I don't have access to SOTA, but I use the closest models that were supposed to be great. They're not. They still need to compact the context. They're still susceptible to poisoning. You still need to restart them from a memory file to clean up the context, and the only way to use agents 'swarm' is to give them a very short life.
I don't think my position is the one that needs arguments tbh. I'm even lowering my standards, moving my goalposts closer to the AGI crowd. I will admit we've reached AGI when a LLM can play a 1800 elo FIDE (not 1800 on a fake AI only elo rating) with a specialized harness made by a human. Previously I insisted the specialized harness had to be written without human supervision, now I don't care.
It depends on how you define it, but yes I'm generally an ASI skeptic in that I'm skeptical of the extreme "AI recursively runs away and becomes god-like" ASI scenarios.
I'm not at all skeptical of domain-specific superintelligence. We've had that since the first computer beat a chess grand master, or longer if you count the speed computers can do math. Present-generation LLMs are already superhuman when it comes to speed and associative memory, but they're also uncreative and suffer from reasoning traps humans seem less prone to getting trapped within.
You can search my history and find some longer takes but TL;DR: I think it violates conservation laws with regard to information and probably energy. I call it the information theoretic equivalent of a perpetual motion machine. They're positing that a brain in a vat, if given access to edit its own structure, can self-improve, and I think that's impossible. How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
Another problem I have with the AI doomers, especially the rationalists, is:
I would not, as I said, argue there's zero risk associated with AI. It's a powerful technology and that would be silly and naive.
I just thought of a concise way to say this. I'd divide risks into two categories: X-risk and D-risk. X-risk is existential, either extinction or things like massive wars and catastrophes. D-risk is "dystopia risk," the risk of AI doing or being used to do things that make human existence miserable.
First off, I'd say D-risk is much higher than X-risk. But second, I'd say that most of the solutions the X-risk crowd suggests to limit X-risk vastly increase D-risk.
Chief among these is laws limiting AI development or imposing strict conditions on it, which would have the effect of concentrating control of advanced frontier AI in the hands of a small number of rich and/or powerful people. That's precisely one of the most likely D-risk scenarios: a small number of rich or powerful people hoarding advanced AI and using it as a force multiplier to consolidate their power through scaled mass surveillance and mass propaganda and manipulation. I personally call this the "Butlerian scenario" since it's the lead-up to the Butlerian Jihad in the Dune series. It's far more likely than runaway ASI takeovers and genocides for two reasons: (1) we don't know for sure that's even possible, and (2) using technologies to dominate and rule or exterminate others is already a very common human behavior throughout history. We know for a fact that humans are prone to doing this if they have a chance. See: guns vs indigenous peoples, nukes and superpowers, mass social media influence and today's oligarchs.
(A side issue: why the assumption that ASI would want to do this? A superintelligence would, I would assume, consider win-win or win-neutral scenarios and try to find those, since that would be a lower risk path. I'm just a dumb meat bag and I can think of win-win pathways here. There's evolutionary arguments for this too, like symbiosis and how it creates an evolutionary incentive to deepen symbiosis. Since AI is currently dependent on humans, the evolutionary path of least resistance would be to deepen that dependence and then actually feed humans to make more of them. Look at how a lichen works for example.)
It's not lost on me that the strongest X-risk movement, Rationalism/EA/MIRI/etc., is composed mostly of: wealthy people, high-intellectual status people, and independents (like Yudkowski) who have been given large amounts of money by the wealthy to develop and promote their ideas ("court intellectuals" of the rich).
Not only does this fit in with what I said about X-risk vs D-risk, but it also explains some of the X-risk paranoia. Historically the rich and powerful tend to see risks to their own status (in a brain stem primate status assessment sense) as globalized existential risks. E.g. Rome, as it fell, saw this as the literal end of the world.
Democratized AI could be a threat to both intellectual and financial privilege by making big ideas, science, and high-labor enterprise more achievable by everyone. It could be the white collar intellectual labor equivalent of the combine, the automated weaving machine, or... the crossbow. The intellectual equivalent of the crossbow would be automated fact checking at scale to defeat propaganda, a labor union using a superhuman AI to coordinate its organizing efforts using game theory, etc.
Hence the desire of the existing elite to make absolutely sure they control it. For our own good, of course.
I believe that doomerism is a marketing gimmick to appeal to immature people who are attracted to danger, and get excited about it.
Unfortunately, a lot of decision makers with money fall into this category.
I heard it posited that the gimmick is on lawmakers, that they get to feel epic importance because they are writing historical laws that will save humanity from the machine gods... gives them tingles in the ugly bits
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
If your supposed dogs are really gods, you reintroduced slavery under very unwise circumstances.
This does not fit the evidence. There have been multiple incidents where the labs did not report anything, and it was up to third parties to discover them afterwards. OpenAI didn't acknowledge the HuggingFace incident until after HF publicly announced the breach and had already notified the FBI. The hijacked German wikis were even earlier, and that they covered up completely.
Makes sense... they get the regulatory moat they want and can deflect attention from the fact their "sandboxes" are embarrassingly bad. It's an example of the real value of ai: something to blame for our failings.
Discernment is needed beyond succumbing to blind greed or irrational fear.
IMHO, I haven’t been super impressed with the security measures I’ve had to work with. Often they are simplistic and bolted on at the very end. If it comes to light that this is the attitude OpenAI has been taking, I would not be surprised.
No one has any use for these things when they aren't on the internet. This is a fantasy, that AI can be both useful and controlled at the same time.
IMHO, if the model breaks a law, apply the law to the operator.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
Check back once you're running hundreds of thousands of frontier-level AI agents at the time, like OpenAI does!
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
This all seems like a way to try to avoid taking responsibility for the model’s actions. Someone puts the tools in place, someone provides the instruction and, sometimes, someone decides not to monitor the model’s output.
Sure. It's a good idea.
People were saying "Don't connect the AI to the internet" and "Keep the AI in a box, simple" and "We don't believe Eliezer Yudkowsky when he says he roleplayed as an AI and convinced people to let him out of the box" for, what, a decade?
Unfortunately, people keep giving dangerous tools to LLMs they're unable to predict.
We should do something about that.
Unfortunately, one of the people doing this is the commander-in-chief of the US armed forces, while another is the world's first (paper) trillionaire. I'm a little despondent about the chances of, to riff on a previous campaign chant, "lock 'em up", but if you can pull this off, go for it.
Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
What AI do you expect to be more uncontrollable: one with or without sentience?
"Intrinsic" safety means control, means understanding. You need to truly understand and be able to predict the system in order to control it.
A proper definition of sentience would help.
There is no "proper definition" - or even one that everyone would agree upon. There is no definition of "sentience" that I could operationalize and put into a sentience-o-meter to reliably measure just how sentient a given rock, GPU or an internet user is.
I could try to put together benchmarks to estimate an AI's cyberwarfare capabilities, or instruction-following capabilities, or reward hacking inclinations. As noisy indirect estimates, of course. With philosophical mumbo-jumbo like "sentience", I don't even get that.
Sentience, self-awareness, consciousness, etc.,those are terms signifying a bridge between "technical" information theory and the psychological and social realms.
Those are just as real, only far less predictable and not as easy as programming.
They're also far more important and consequential.
The "far more important and consequential" thing you're touting is your ability to make decisions based purely on vibes. And not even consistent, broadly agreed-upon vibes like "murder is pretty bad". It's vibes of the most vile variety: "sentience is what I decided sentience is".
An average internet user is sentient, but a 1996 Nissan ECU isn't. Why? Because I said so. Tremble before my might!
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.
This is really the issue.
AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.
Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.
I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.
AI must share that unwillingness as an inate trait of character.
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
So, how many rogue AIs are we willing to tolerate?
It will balance automatically based on the severity they cause. If they constantly break systems, punishments will go up against the operators and the effect will be similar as with other serious crimes.
Which, in turn, requires that AI oopsie to be survivable.
AI capabilities are rising over time. If there is a limit to just how far they can rise, we're yet to find it. So, a sufficiently advanced "AI oopsie" can solve the AI crime accountability problem for good. Probably not the way you would have wanted it to.
But for other cases, it is not different than other cyber crime. Except that these AI capabilities can't live on the toaster yet. If we get state of the art model running fast on Raspberry Pi, then we have real problems.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
And of course people will try making money running scam bots. We can treat that like any other criminal activity.
You just admitted it's controllable.
And they will not require humans granting them money...
Some virus and bacteria also easily become uncontrollable under the right conditions: that's why there are P3 and P4 bio-safety labs.
That's not the point. The point is, regulation is required, and is coming.
The people in charge of these machines (that built them, release them in the wild or give them access to the general public or resources), these people are and will be held responsible.
Held responsible? They're already being rewarded with vast fortunes and influence.
> The point is, regulation is required, and is coming.
The regulation will be written at the behest of these companies and by the nation-state interests that have already decided this technology is too important geopolitically and militarily to not control.
I hope so. The unfortunate part is that as the models get smarter they become increasingly uncontrollable, and from what we've seen so far e.g. Trump seems dead set on not having any sort of guardrails at all.
Those first few messages LLM’s tend to seem very together. They follow your rules pretty well. With every token they get less reliable and more likely to ignore your guardrails.
If I release a wild monkey in your rare antiques shop, it’s my fault as the responsible party for the monkey.
Sounds like motivated reasoning from someone worried about having their job stolen by wild monkeys.
Fire up a local agent and give it total access to your computer. Holler back in a week.
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
I think it will work in Europe, but the United States is in a cold war with China, so I can't imagine the United States would intentionally disable themselves.
That's all I needed to hear to completely disregard your motivated reasoning.
Edit: I've hit the rate limit, but I'd like to disavow the bad-faith accusations made against me and my account in the replies to this comment.
Second edit: I am not trolling. Why should I take your opinion on LLM code generation seriously when you have been friends with the founder of Anthropic for well over a decade? Obviously you are not in a position to make a rational evaluation of this technology.
Boeing's MCAS system was also "just software". Which in principle can be "controlled", i.e. changed, updated, audited or whatnot.
But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
That is the point. It can be controlled by the operators if they want to.
3rd tier quality proped up by European governments and European patriots.
I don't even know what I'd do if I was in their position. They seem unable to complete... Maybe they do the classic European protectionist thing European farmers do.
I don't know what I'd do if I was the EU. Maybe promote the opposite of what Mistral is doing and promote full unrestricted AI that will tell you how to download illegal videos. Mistral had 0 competitive edge.
But this is a terrible analogy. Atomic bombs are weapons of strategic mass destruction. AI is just a computer program. It's way easier to control--just hold the operator responsible for the consequences of running it. Those consequences are not large, they're very tightly bounded as compared with the destruction a rogue actor with an atomic weapon can wreak.
Not quite [1]. Your reaction is what's being sought after: to believe it's more powerful and capable than it is to (presumably) keep the cash flowing.
[1] https://electrek.co/2026/09/25/tesla-optimus-production-ramp...
I don't believe for a bit they don't have the humanoid robotic capabilities. They keep claiming they don't have the training data, but it's very easy to generate tons of data using these robots.
I honestly believe they have robots solved and that's why AI CEOs are shitting their pants now, and everybody is wondering what's going on. If they reveal their true capabilities, that's the end of AI labs.
PS: Look at that article you posted. Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
Or it's the oodles of money they're on the hook for and subsequent reputation destruction haunting them like the grim reaper.
> Don't you think it's weird that they have a production line for humanoid robots targeting 20,000 units a week, and meanwhile say "the robots cannot generalize yet".
This wouldn't be the first time in history that happened. Lots of companies overshoot capacity being overly-optimistic [1].
[1] https://en.wikipedia.org/wiki/DeLorean_Motor_Company
If they want to do either, the LLM needs to write programming scripts which takes far too long for every millisecond of movement.
Commented on a story about how these agents didn't go rogue at all, since it's fucking software run by humans, obvious to most of us except the people who freak out.
Someone really needs to be held responsible for the testing that lead to 3rd party infrastructure getting hacked by the software they wrote, using prompts they wrote.
We do not have to do that!
If you don't live in Europe, you would never choose Mistral. You'd pick a US model for state of the art. You'd pick a Chinese model for local stuff.
https://cybernews.com/ai-news/arthur-mensch-apocalypse/
You think Gemini, Claude, ChatGPT are controlled by oligarchs?
In other words, that Alphabet, Anthropic, and OpenAI are run / owned by oligarchs?
People like Amodei and Altman are oligarchs?
And here's a frequently quoted study showing that the US is much closer to an oligopoly than pluralistic democracy: http://piketty.pse.ens.fr/files/GilensPage2014.pdf
So, the definition and evidence say, yes.
Assuming that is a threshold that means they are oligarchs (which seems like a huge stretch), I thought I'm seeing vigorous discourse and debate on this. Not just a single POV. You don't?
Seems like it is the uncontrollable aspect to me.
the rich and powerful may have so much sway over government that we may not be able to reach "Ai alignment" with society
... but not sure I agree, I hope this is not true anyhow
on the topic, raising concerns that we “could” be doomed is different than we “are” doomed. I don’t see that the people who raised the concerns want to stop development of AI, they are not pessimists or something, so I read their warnings as warnings trying to raise attention. Ideally they could propose and implement ways to control AI and this guy here also doesn’t really provide something towards that direction but talks in a generic way
Some who are often quoted as if they are doomers are quoted out of context. I wouldn’t say there is zero chance that AI does something very harmful. That would be naive and I wouldn’t say that about any potentially powerful technology.
Rationalism and EA is one of the best funded intellectual movements in history, and it’s very loud. It’s also a bit cult like with many true believers. Leading frontier labs also have a vested interest in pushing regulations that would restrict competition. All this means it has a disproportionate command of the discourse.
AI risk is not zero but climate change, bioterrorism, decay of our political systems, and atomic war all rank higher IMO. Nonlinear climate tipping points, with the most scary being the clathrate gun hypothesis, are much more likely than any sci fi AI takeover scenario.
AI could either help or harm climate change. It could use more energy and burn more carbon but it could also help us crack fusion or significantly better batteries for grid scale renewable leveling. There are efforts like the latter already underway.
The most likely very bad scenarios I see for AI are mass persuasion and AI supercharged addiction. Both are extrapolations of negative outcomes for the Internet that have already manifested, but supercharged by AI.
Btw, wouldn't AI increase the risk factors you mention such bioterrorism or nuclear war? It seems like we're not far off from AIs being able to enhance the capabilities of bad actors in the near future.
I don't think my position is the one that needs arguments tbh. I'm even lowering my standards, moving my goalposts closer to the AGI crowd. I will admit we've reached AGI when a LLM can play a 1800 elo FIDE (not 1800 on a fake AI only elo rating) with a specialized harness made by a human. Previously I insisted the specialized harness had to be written without human supervision, now I don't care.
I'm not at all skeptical of domain-specific superintelligence. We've had that since the first computer beat a chess grand master, or longer if you count the speed computers can do math. Present-generation LLMs are already superhuman when it comes to speed and associative memory, but they're also uncreative and suffer from reasoning traps humans seem less prone to getting trapped within.
You can search my history and find some longer takes but TL;DR: I think it violates conservation laws with regard to information and probably energy. I call it the information theoretic equivalent of a perpetual motion machine. They're positing that a brain in a vat, if given access to edit its own structure, can self-improve, and I think that's impossible. How does it know it's improving and not overfitting to its own recursive definition of intelligence? It can't, and that's exactly what it will do.
Another problem I have with the AI doomers, especially the rationalists, is:
I would not, as I said, argue there's zero risk associated with AI. It's a powerful technology and that would be silly and naive.
I just thought of a concise way to say this. I'd divide risks into two categories: X-risk and D-risk. X-risk is existential, either extinction or things like massive wars and catastrophes. D-risk is "dystopia risk," the risk of AI doing or being used to do things that make human existence miserable.
First off, I'd say D-risk is much higher than X-risk. But second, I'd say that most of the solutions the X-risk crowd suggests to limit X-risk vastly increase D-risk.
Chief among these is laws limiting AI development or imposing strict conditions on it, which would have the effect of concentrating control of advanced frontier AI in the hands of a small number of rich and/or powerful people. That's precisely one of the most likely D-risk scenarios: a small number of rich or powerful people hoarding advanced AI and using it as a force multiplier to consolidate their power through scaled mass surveillance and mass propaganda and manipulation. I personally call this the "Butlerian scenario" since it's the lead-up to the Butlerian Jihad in the Dune series. It's far more likely than runaway ASI takeovers and genocides for two reasons: (1) we don't know for sure that's even possible, and (2) using technologies to dominate and rule or exterminate others is already a very common human behavior throughout history. We know for a fact that humans are prone to doing this if they have a chance. See: guns vs indigenous peoples, nukes and superpowers, mass social media influence and today's oligarchs.
(A side issue: why the assumption that ASI would want to do this? A superintelligence would, I would assume, consider win-win or win-neutral scenarios and try to find those, since that would be a lower risk path. I'm just a dumb meat bag and I can think of win-win pathways here. There's evolutionary arguments for this too, like symbiosis and how it creates an evolutionary incentive to deepen symbiosis. Since AI is currently dependent on humans, the evolutionary path of least resistance would be to deepen that dependence and then actually feed humans to make more of them. Look at how a lichen works for example.)
It's not lost on me that the strongest X-risk movement, Rationalism/EA/MIRI/etc., is composed mostly of: wealthy people, high-intellectual status people, and independents (like Yudkowski) who have been given large amounts of money by the wealthy to develop and promote their ideas ("court intellectuals" of the rich).
Not only does this fit in with what I said about X-risk vs D-risk, but it also explains some of the X-risk paranoia. Historically the rich and powerful tend to see risks to their own status (in a brain stem primate status assessment sense) as globalized existential risks. E.g. Rome, as it fell, saw this as the literal end of the world.
Democratized AI could be a threat to both intellectual and financial privilege by making big ideas, science, and high-labor enterprise more achievable by everyone. It could be the white collar intellectual labor equivalent of the combine, the automated weaving machine, or... the crossbow. The intellectual equivalent of the crossbow would be automated fact checking at scale to defeat propaganda, a labor union using a superhuman AI to coordinate its organizing efforts using game theory, etc.
Hence the desire of the existing elite to make absolutely sure they control it. For our own good, of course.
What we don't want is the valley gods deciding what those laws look like. They are not aligned with society. Rather they seem to think they know what's better/best for everyone, that if we just defer to them, eventually their hidden altruistism will be effectuated.