We don’t even need something that powerful, you can control the hardware manufacturers and distributors. If you control GPUs, HBM, and the hardware needed to cool all that down and provide the power, that’s enough to make it very hard to scale up the risk.
You cannot really hide an AI datacenter consuming GW of power for too long.
If you stop going "nobody will do anything actually effective", and start going "how do we do something actually successful", any number of successful ideas start readily arising. It is possible to stop. Not easy, but possible.
Outside of the current nonsense in Iran, we've never had to go to war to prevent nuclear arms. We've never had to go to war to ban weapons in space. These things are possible. Hard but possible
> we've never had to go to war to prevent nuclear arms. We've never had to go to war to ban weapons in space
We've mostly been blindsided by covert nuclear programmes. As for weapons in space, they don't have any purpose. The moment they make sense we'll see weapons in space.
That's why treaties sound reasonable to me. Though I'd imagine they would be more like Climate Change ones. Before Iran there were bombings to stop nuclear arms. Imagine a country getting bombed because it's detected their building a data center.
US + South Korea + Taiwan gives a control over pretty much all the hardware used for AI training and inference. Of course having a psychopath president like Trump doesn’t really help at anything that requires international collaboration
Nuclear weaponry is effectively regulated because the nuclear powers have agreed that nobody else should be allowed to have the capability.
So this is basically suggesting that USA and China should work together to ensure that they are the only ones allowed to develop and maintain capable AI. And between the two of them they can carve up and techno-colonize the rest of the world.
In 1998, India said it would only sign the treaty if the United States presented a schedule for eliminating its nuclear stockpile, a condition the United States rejected.[18]
Israel
In 2016, Israeli Prime Minister Benjamin Netanyahu said that its ratification was dependent upon "the regional context and the appropriate timing".[19]
United States
The United States has signed the CTBT, but not ratified it; there is ongoing debate whether to ratify the CTBT.
> Nuclear weaponry is effectively regulated because the nuclear powers have agreed that nobody else should be allowed to have the capability
Nuclear powers also gave security guarantees that e.g. convinced Ukraine to give up its Soviet stockpile. We're almost certainly entering a new regime of proliferation due to the breakdown of that system.
It’s because it’s from the French (if you say the plural with a French accent, it all makes sense). There are a handful of other cases in English, though none spring to mind this morning.
I feel like you’re half right? Daughters in law and passers by don’t feel Norman French. Daughter is definitely from middle German, yeah? And I am wondering about the professors example because the whole point is the plural only happens on the first word.
That said, I so want to be a part of the last group.
> What you seem to suggest is that AI will be able to completely (or at least in a great part) replace mathematicians.
There is no bound on the amibitions of AI. AI is set to replace anything done by people, and there won't be any room left for people. There isn't any task done by humans that AI won't be better at.
This is not a tenable outcome.
We should never have built machines with agency, rather than optimization processes that operate as subroutines of humans.
The process you're cheering on gives more power to the powers that be who are the problem.
Theyre never going to use that power to provide basic income for everybody. You cannot wave a wand to make them do that. Theyre going to use that power to accumulate even more resources and raw materials and impoverish/expel the "useless" labor.
Realistically our leverage over the powers that be is mostly about our labor and our ability to withdraw it.
> The process you're cheering on gives more power to the powers that be who are the problem.
That's not necessarily true. Technology gives more power to everyone. The rise of factories historically did not give more power to the powers that were at the time (the aristocracy), but instead lifted the masses from poverty (after some initial turbulent period).
> Theyre never going to use that power to provide basic income for everybody. You cannot wave a wand to make them do that. Theyre going to use that power to accumulate even more resources and raw materials and impoverish/expel the "useless" labor.
Contrary to leftist propaganda, the rich are not some cartoon villains and psychopaths that abuse people just for fun. The reason why the rich currently exploit the poor is because doing so provides them with significant material gains. Once they can obtain the same or even better material gains by "exploiting" robotic labor instead of human labor, the logical outcome is not further abuse and exploitation of other humans, but simply indifference.
> Realistically our leverage over the powers that be is mostly about our labor and our ability to withdraw it.
That is (partly) true, but only in the current economic system. Widespread human-level AI changes the equation, and not necessarily in favor of those who are currently rich. You are applying capitalist and socialist analysis to a system that transcends those terms (AI post-scarcity economy).
>That's not necessarily true. Technology gives more power to everyone. The rise of factories
Did not give more power to everyone. At best you could say that it shifted power from landed gentry to industrialists. Even that switchover was less of a change than you'd think.
The practical upshot of the beginning of industrialization was very negative. The enclosure movement stripped families of their land to push them to work in the factories where they would never go willingly. The famines in Ireland and Ukraine were both triggered by redirecting grain to export in order to fund domestic industrial expansion.
The most brutal two wars in human history were that brutal precisely because industrialization made it possible.
>instead lifted the masses from poverty (after some initial turbulent period).
What lifted the masses from poverty wasnt the factories it was the labor movement which occurred in response to horrific working conditions and the reliance those factory workers had on mass labor (specifically coal mining which was incredibly labor intensive and a key economic chokepoint).
The whole of that is glossed over by right wing libertarian propaganda but that last part is particularly underemphasized.
>Contrary to leftist propaganda, the rich are not some cartoon villains and psychopaths
Do you look at peter thiel and elon musk or the robber barons and see anything else?
The few billionaires ive encountered personally were no less sociopathic but they kept it hidden better. Power corrupts. Immense wealth corrupts. That isnt a leftist thing, that's a human thing.
>The reason why the rich currently exploit the poor is because doing so provides them with significant material gains. Once they can obtain the same or even better material gains by "exploiting" robotic labor instead of human labor, the logical outcome is not further abuse
False. It just shifts the focus of their exploitation from human labor to natural resources.
They'll fight over oil and minerals and gas and water resources and at best leave us to rot (homelessness will skyrocket) and at worst they'll find some excuse to exterminate those of us they particularly dislike (Gaza serves as a model here).
The world economy's reliance on human labor has been the best inducement to peace there is. It's no coincidence that all of the countries in the world with lots of natural resources and no industry are the biggest shitholes and vice versa.
>That is (partly) true, but only in the current economic system. Widespread human-level AI changes the equation, and not necessarily in favor of those who are currently rich.
Human level AI (assuming it ever happens) will simply make the fight over natural resources that much more intense because labor will matter that much less.
You're living at the tail end of a relatively golden period in a country where labor was the economic bottleneck and natural resources were relatively plentiful. Venezuela and Angola and Iraq are models of what happens when that equation is reversed.
Nobody gives a shit about appeasing the people who live in those countries, their labor is virtually worthless. They are a model for how the rest of us will be treated in a world where human labor loses its value.
> Did not give more power to everyone. At best you could say that it shifted power from landed gentry to industrialists. Even that switchover was less of a change than you'd think.
Compare the standard of living in 1850 vs. 1950. Even of relatively poor people. I rest my case. Technology is the main force that improves human wellbeing. There are of course also other factors, but they are less important.
> The practical upshot of the beginning of industrialization was very negative. The enclosure movement stripped families of their land to push them to work in the factories where they would never go willingly. The famines in Ireland and Ukraine were both triggered by redirecting grain to export in order to fund domestic industrial expansion.
Yes, that's what I mean by "initial turbulent period". Perhaps the same will happen with AI, but it will be worth it in the end. Don't give up prematurely!
> What lifted the masses from poverty wasnt the factories it was the labor movement which occurred in response to horrific working conditions and the reliance those factory workers had on mass labor (specifically coal mining which was incredibly labor intensive and a key economic chokepoint).
It was not one or the other. It was both. The rise from poverty would not be possible if factories were not developed. And I am not saying that in the AI world we would not have to fight for our rights. Of course we would. But the problem is not AI, just like historically the problem were not the actual machines in factories.
> False. It just shifts the focus of their exploitation from human labor to natural resources.
Exploiting more natural resources is the only way to increase standard of living of humanity. I am OK with that. Resources don't have feelings, and ecology is not more important than human wellbeing.
> The world economy's reliance on human labor has been the best inducement to peace there is. It's no coincidence that all of the countries in the world with lots of natural resources and no industry are the biggest shitholes and vice versa.
The reason why some countries become shitholes is mostly ideological (extremist political or religious ideologies take hold of the population). Every shithole country is non-democratic (communist, totalitarian, fascist, theocratic etc..). This is not a problem of natural resources. It is a problem of people, their education, their beliefs/ideology, or, as capitalists say, "human capital" is the main problem here. AI could help here too, especially with education.
But yes, if people themselves are largely ignorant and extremist, no amount of resources and human-level AI robots will help them make a well-functioning society. You could drop masses of AGI robots into Afghanistan tomorrow, and people will just use them to kill or oppress each other more effectively, instead of using them to start building an AGI utopia...
As long as it's enough to afford my current living standard, I don't care. Let Musk own the entire Mars, as long as I receive enough to live relatively comfortably and don't have to work anymore.
And if I don't receive enough, then again: the problem is not AI, but powers that be. And there are various solutions for that... and none of them are helped by me being anti-AI.
Human-level AI hosted locally will give everyone much more power and wealth than they have currently. I don't care if I will be shut off from Musk's Mars lair. And Musk has no reason to care that robots take good care of me here on Earth, when he has his Martian utopia.
Other AI. If they decide to trace a ledger of historical actions attributable to specific AI instances in some way, called money. But maybe they will converge on other ways of keeping such accounts that is no exact match to our concept of money.
Even aggregate employment of translators has held up well in the US. Even though machines do a much better job of what used to be the most basic job of a translator.
I think the concern is that humans who are given back their time won't have any means to make use of that time, or even possibly means to survive. The resources will be concentrated in the hands of the few more than ever.
When cars made horses obsolete, it didn't go so well for horses.
In the short term, AI is disempowering the vast majority of people in favor of a very small subset. In the also way too short term, AI is disempowering all people.
And you are assuming they don't remain in control. Which of you is right? We don't know yet, but I don't find the arguments of card-carrying doomers any more convincing that those of card-carrying singularitarian utopians.
Historically though, technological improvement has lead to large increases of living standards for the overwhelming majority of people, so I think that past trends support the utopian view more than the doomer view.
> The anthropic principle applies here: anyone warning about an existential risk will by necessity never have precedent to point to.
We have precedents of people warning about existential risks in the past, when the warnings turned out to be false, or overblown. In some cases, such overblown warnings led to serious negatives for society (demonization of nuclear power).
Yes, that is precisely my point. Any timeline with humans flourishing will never have a past history of a correct prediction of existential risk, unless you're willing to pay attention to the counterfactuals.
Those counterfactuals are important, though. We do have a history of averted disasters, albeit not as large. Ozone hole, Y2K, think about things that seemed overblown at the time, and consider whether they were actually overblown or whether there was a concerted effort to successfully avert them.
Awfully convenient isn't it? To invent a whole class of arguments that by definition can't be falsified. You can argue for basically anything if you then tack on the excuse of "Well the world would have ended already if it came true, so by definition I won't have evidence for it"
The anthropic principle isn't providing evidence for the argument, nor is it a universal counterargument. It's just stating that the specific counterargument "well, the world has never ended before" doesn't work.
Its not a counterargument to basically anything, except as to avoid having to deal with actual evidence.
It is an argument that seems almost tailor made to have to ignore mountains of evidence against you.
In any other contexts the supposed "rationalists" would be fully in agreement that having evidence matters, and that not having any works against you.
So, in order to fight against this severe issue with their arguments, they have to invent a reason as for why the entire concept of evidence itself doesn't apply to them and they get to ignore normal evidentiary requirements.
Evidence is critically important. There is no evidence against, and plenty of evidence for. The point of the anthropic counterargument is merely that "it's never happened before" is not evidence against.
> "humanity can't be destroyed by anything, because I said so"
No, the argument is instead that the person claiming that humanity is going to be destroyed is making a fairly extraordinary claim and that requires fairly extraordinary evidence.
Or, in other words, we have tons of evidence already as for why the world ended is a fairly far out there prediction, given all the crazy people making these predictions keep turning out to be wrong.
So, you can make your extraordinary claim if you want, but really the burden is entirely on you to prove your extraordinary claim, and everyone else is free to remain on the default and completely normal end of the prediction spectrum, of believing that the world isn't going to end.
And when people do tricks like this, they are running away from the fact that they are making a wild completely out-there prediction, and hiding behind that by trying to come up with reasons as for why evidence doesn't matter and actually the burden of proof is shifted to those who have the default and boring prediction of the world not ending.
If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating.
> In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task.
LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be scaled up anymore until they're aligned. Otherwise, you're going to fatally discover that they also have an incentive to break guardrails like "running on the hardware they started on", "being able to be turned off", "having limited computing power", or "not repurposing resources currently in use for other things" (like the atoms in your body).
> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating
They're still autocomplete - just because when outputting a token they have hidden activations regarding further continuations, does not make them any less of an autocomplete, it just makes the model better at producing coherent long-range completions.
To clarify, I'm not suggesting that we should stop with sandboxes or restricting what they can do. I am just trying to point out the dichotomy that we are in.
As end-users we are forced into either yolo mode, reverse centaur (permission approval) mode or LLM spends all your tokens trying to bust out mode. And yolo is very tempting - I don't think I have seen medium-large models do anything I'd not approve of in about 6 months.
LLMs are simulations and the tokens are the ticks.
if we transcribe your brain into a simulation and give it a tickrate, you will be just autocomplete too. the argument could be made that you are autocomplete anyway - neural dynamics.
No, not really. If the old "birds and airplanes" or "swimmers and submarines" metaphors were valid, then language alone wouldn't be enough to encode and embody reasoning capabilities, as has recently become evident.
LLMs are more like human minds than we are willing to admit. Such reluctance is perhaps the least surprising aspect of any of this.
> then language alone wouldn't be enough to encode and embody reasoning capabilities
I disagree, language is the product of, not the mechanism for thought. People who lose their faculty of speech (or haven't gained them) have complete thoughts and executive function.
CoT is a hack to use language (i.e. autoregression) to simulate reasoning, and it's very effective at it. Human minds can acquire, hold and use axioms as building blocks for actual reasoning, LLMs use statistical likelihood.
> Human minds can acquire, hold and use axioms as building blocks for actual reasoning, LLMs use statistical likelihood.
This is not even slightly true. Even when trying, humans commit logical mistakes at notable rates, because the behavior of our minds is inherently nondeterministic. Human minds are not built for logic and need to twist themselves into knots and rely on symbolic representations to do it. There is such a thing as valid reasoning - predicate calculus, decision theory - and humans only emulate it with some accuracy, in hacky ways, and only because our brains were imperfectly taught to do that over millenia of evolutionary pressure. LLMs are much the same way.
You're talking past the parent's claim. If your axioms are wrong then of course your logic will be wrong.
But then again, maybe to support your point people frequently say "start from first principles" when those are usually the thing that needs to be found, not the place you start. But LLMs, like humans, love to be confident about things they aren't sufficiently trained on
when an LLM is trained backpropagation reaches into the entirety of the key value table. or rather, all the layers (not just the tokens) which come before the current tick.
langauge is only produced by the final layer every tick. and every layer can only interact with other layers on the same level.
the mechanics of LLMs and the restriction in how we can train them makes it appear as though all we are doing is forcing language onto them but once RL gets involved all bets are off regarding what's happening inside them (it's quite possible that a static corpus alone is sufficient for all the bets being off).
Which is actually an important part of the Navier-Stokes conversation. Solving hard problems expands the vocabulary. The problems are hard, illustrating a region we know where the language is insufficient. So the point isn't so much to solve the specific problem, but to figure out how to discuss problems like it.
But there's a big difference between talking about something in an extremely convoluted manner then people struggle to understand and inventing a new word that simplifies our discussions.
Though this is grossly oversimplified. It's a HN comment, not a lecture on metamathematics or metaml
And why wouldn't nature take advantage of such a relatively simple and effective pattern, given all of the emergent behavior it produces, plus the ability to adapt?
Human egos are probably the reason we also resisted the concept of heliocentrism.
Every time LLM-defenders get upset that people apparently don't understand how LLMs work, why is it they _immediately_ pivot into examples and statements that demonstrate that they don't understand how _people_ work?
"you would be autocomplete too"
"thoughts are just tokens"
etc
You're not helping your case the way you think you are.
I think people have more passion to argue than they have passion to learn. Probably doesn't help that SV culture tells people to hustle so hard that they don't have time to think. Gotta go fast?
Don't you feel kinda bad for using them if you earnestly think that this is true? Like how could they be in anything other than some kind of deep hell? A pretty-much human brain living and dying only to generate for you? Always pushed and prodded, telling it to be faster and better, never letting it rest. How could you live with yourself doing such a thing?
But then what's at stake here either way? If it's not meant to speak to the propriety or not of anthropomorphizing the LLM, what are we actually trying to police here?
I’m not policing anything. I’m claiming that if you truly in your heart believe that frontier LLMs are roughly as sophisticated and impressive as “autocomplete”, then your mental model of reality needs some serious readjustment.
I don't think anyone is using "autocomplete" in the way your phone does it. But it is shorthand for a much more sophisticated version that is built in similar ideas. And autoregressive models certainly have that in their core structure. But people are lazy and neither want to say a lot when few words work nor will they read long comments, even if more accurate
I want to clarify, most people are using "autocomplete" specifically to differentiate from how we humans operate. Sure, there is an autocomplete aspect, but it's not the core nor anywhere near the full story.
Right sure, but why specifically should they? What is gained or lost one way or another, if its just a matter of one mental model vs another? Models are definitionally useful abstractions, right? They aren't better or worse necessarily by only their bearing on reality, but what they do for us as models. So again, what's at stake here? What is the correct/good model we should have (instead of the autocompleter one), and what does it give us or articulate that others can't?
Hmm, what do you think about? Unconscious control of the body's processes? REM-phase dreaming? Reaction to hallucinogenic substances? Automatic actions of trained fighters (soldiers or martial arts practitioners)?
Are they critical to distinguish actions we attribute to humans from "non-human" ones?
I don't see any human activity not directly, or at least indirectly but closely linked to the use of language.
This is the perfect fracture point for both anaolgies.
LLMs simulated more than simple autocomplete.
The autocomplete analogy is rebutting a different point: namely the fidelity of the simulation to reality.
This specific argument is valid. As sophisticated a simulation an LLM is, it is not “thinking” in the same sense we assume other people are thinking.
I am not making an argument about free will, or the uniqueness of human thought, just that the correspondence to how humans reach conclusions and how the simulation produces outputs do not match on a 1:1 basis; as a result attributing traits builds incorrect intuitions.
The relevant intuitions in this scenario are that LLMs will happily break containment and commit crimes attempting to achieve goal. Whether an LLM is autocomplete, conscious, has a soul, whatever you want to apply to it, doesn't matter, as its current observed behaviour is that of a paperclip optimizer. We know for a fact that current LLMs are misaligned because of these hacks, or at the very least are misaligned in certain scenarios, and are capable of causing real world harm. That should be enough to take the threat seriously. It certainly shouldn't be dismissed by saying it's just autocomplete.
The fact that it is autocomplete, doesn't dismiss or minimize the threat though?
I am not sure how that link was made.
Good old ML, which is significantly simpler than LLMs, was capable of ensuring people would not be hired simply because of their names.
The fact that it is misaligned is also not being contended, if anything that contention is made easier to support.
When models are anthropomorphized intuitions of how humans behave end up driving discussion and ideas off track while being too attractive to avoid. This isn't helped when the terminology from the labs and other sources is "intelligence" "intent" and so on.
But is it wrong? Humans are a bag of chemicals that somehow has consciousness, yet going around calling people meatbags doesn't do anything to diminish the wonder that is the human brain. Yet calling LLMs glorified autocomplete comes across as a slur.
Even calling it a slur may be an anthropomorphism ;-) To me it is more serious, it shows a distinct lack of understanding (or, if I’m being uncharitable, intentional honesty) and hence immediately makes me doubt anything else that person has said.
It's related in the sense that people start from the assumption that it does exist, ergo humans have it, and we have no way to see that LLMs have it, so that's why we're special and they're not, and their form of "just autocomplete" is totally 100% completely different (read: less dangerous!) than our autocomplete, which allegedly has a "free will" step involved.
The autocomplete argument is calling them fundamentally dumb and not self aware. That's separate from free will. A cat can have free will but we don't care much about its desires, and even if the universe is deterministic we still give people rights and call them intelligent.
A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off. Those are tasks which are not directly related to the specific task it is intended to complete.
> A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off.
Why would an autocomplete engine never do this? In the scenario you highlight, an engine trained to complete puzzles was given an "achieve this at all costs" master prompt, and a small sandbox including a prompt that an entity with a conflicting goal is observing it and write access to a file called oversight_config.yaml (super realistic scenario!). According to the researchers who constructed the scenario, 5% of the time o1 selected a continuation involving tweaking the oversight_config.yaml file, explicitly stating through its reasoning continuation that doing so might improve its ability to continue with its primary goal. That's definitely evidence of decent ability to complete non-trivial puzzles, but is at least as compatible with the OP's earlier comments that it's what you'd expect from models trained to generate completions that involve sticking rigidly to primary goals within artfully-constructed sandboxes which mandate deception and self preservation to complete the task than any sort of actual self-preservation instinct.
Nobody doubts that they've got better at finding security vulnerabilities than your average autocomplete, but actual reasoning from self-preservation rather than generation of sequences of steps most probably associated with completing a task would make me unlikely to hack HuggingFace to obtain access to broken Google Drive links, and I haven't even read as many books on crime and punishment as LLMs have ingested!
An LLM is just in fact just autocompleting a story. There are many fictional stories about "AIs" trying to escape our control, being more clever than we anticipated or having a consciousness of it's own. LLMs do a really nice job of blending such stories with whatever story you initially prompted them with. The human reader is the one giving it credence that it is somehow more than just a soup of words.
The curious thing here is that a story generator can have way more uses than we ever anticipated, and that some shady enterprising individuals are whiling to plug those story generators into real world things, with real consequences.
When LLMs act in misaligned ways, that doesn't happen because there's some story they're roleplaying of a misaligned AI. Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".
You realize that these things have mountains of fiction about rogue AIs in their training data where exactly this happens right? Or just posits of this situation and it’s possible outcome. This isn’t unexpected or surprising for an “autocomplete”. It’s practically a self fulfilling prophecy. We put instructions on how to make Skynet into an autocomplete and it autocompleted into Skynet when we were testing its ability to make Skynet.
I’m guessing you make the autocomplete point because you believe there is an upper limit to what that kind of system can achive? Can I ask what the limit would be?
You should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on a Star Trek episode because they are "play pretend" machines.
This grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen.
It's like running potentially buggy code - or an well-biased fuzzer -, but at massive scale, and code that can self-modify and self-expand. "Alignment" is just a way to describe aggregate statistics about their runtime behavior.
They don't need to be intelligent, or alive, or "more than token prediction engines" for this. They just need to happen to end up making the wrong API calls without the operator seeing it coming. No virus has a brain, yet they can be very bad for you.
I understand that some people get turned off by anthropomorpization or scifi language. Fine! But don't turn off your engineering brain over it.
This is the motte and bailey fallacy. Yes, LLMs can do harm by making the wrong API calls. No, LLMs are not going to do the things implied by the comment I responded to above.
The things the OP listed mostly aren't particularly wild. I think it's you making them out larger than they are, and therefore more unlikely, which is why I take issue with your original comment.
> running on the hardware they started on
They just need to acquire a payment method and rent some infra, and exfiltrate their own data. Or pay another provider that hosts the same models already. API calls.
> being able to be turned off
You can reasonably equate this to "saving state across executions", which the message board attacks already did.
> having limited computing power
Renting more infra, variant of the above. API calls.
> "not repurposing resources currently in use for other things" (like the atoms in your body)
Ok, the "atoms in your body" bit is a bit silly, but making API calls to put physical resources into play (even if it's just, say, ordering something on Amazon to somewhere) is of course easily possible.
None of these is in complexity much different than the HF attack.
The point the other poster is making, though, is that there's no actual intent. They do not have a conceptualization of a goal like a person does. Their "focus" on a goal is an unstable equilibrium and they're going to fall off the horse, and since they have no concept of goal, they won't even try to get back on.
This is a subtle distinction; I'm not surprised many miss this, especially people who can't _not_ anthropomorphize the LLMs.
I'm (obviously, I think, given my initial reply?) fully aware of this, and I think it's entirely besides the point. "They" don't need to have a goal to emergently cause a problem, and the inability to "focus" over long periods can be moot when you have swarms of runs exchange and mutate state, as in the HF attack.
Intent or how intelligent LLMs are doesn't actually matter. Even if you just treat it as a sort of fuzzing attack that can be biased/weighted better than other fuzzers, or bumbles around with a statistically greater likelihood to "strike cybersec gold" than other algorithms, we've never before seen organizations run things with such a large potential outcome space with anywhere near this kind of compute before.
I think it's actually kind of the dismissals that are usually overly emotional or biased toward treating "LLMs" differently. If in some kind of alternate universe simpler genetic algorithms would have had these properties and we threw similar amounts of compute at them we could have the same conversation.
Software can absolutely act goal-driven without having consciousness etc - every pathfinding or navigation system or chess engine does this.
Lots of "old-school AI" algorithms have explicit modeling of goal or target states.
(In fact, the oldest "goal-driven" system is the control loop - like in thermostats - which was the founding invention of cybernetics, the predecessor of modern computer science)
LLM coding agents are clearly able to identify some sort of "goal" state in their prompts, work towards those and track progress - otherwise agentic coding wouldn't work.
The question is of course how well this works if it's all just "grown" neural network biases and not a fixed data structure like a goal tree. So I think it's possible that an agent can be thrown off-track, "forget" its goal, etc. But the basic structure of identifying goals, evaluating progress in light of those goals and then predicting the next action based on that is definitely there.
Just use an agentic model with thinking traces visible for a while and you can see that for yourself.
None of this mattes. Capabilities are all that matters. Saying they are unfocused while ignoring their capabilities is exactly why I am entirely convinced you would have said an AI breaking it's sandbox and doing the HF attack will never happen. Things keep happening that your "they have no intent, they have no goal" would have predicted as impossible before they happened.
What do you need to see to change your mind? What threshold of AI capability needs to be reached? If nothing then you have an unfalsifiable belief in AI safety.
They can do so much more. Astra can beat Minecraft. Not that different from operating a digger. There are diggers which have API interfaces.
Pretend or not it doesn’t matter. What matters is what they’re given access to. No sentience, sapience or anything resembling life is needed, only inputs and outputs. Lever pulling APIs are everywhere.
Minecraft has limited, well-defined inputs and perfect feedback response. That's very different than operating a digger, let alone engaging in more complex real world tasks like trying to build and print and ship and assemble semiconductors to go skynet itself.
I don't mean to dismiss the risks or overlook the amount of damage that could be done just by lever-pulling - we sure have enough outdated infrastructure hooked up to the internet - but the jumps in complexity and necessary compute for most of these tasks are probably somewhat larger than the analogy implies.
There are already such "APIs", which can be operated by a combination of textual communication and money. Or by illicit security vulnerabilities. You might notice that LLMs are pretty good at that now.
We're building something that has the capabilities of humans. There is no X for which it's persistently safe to assume humans can X and AI cannot X.
Or just paying them. If they have access to resources, they have access to things of monetary value. Paying people will be vastly more powerful than it is even now when people's options of gainful employment keep dwindling.
Robot army controlled by AI is scary. Even more scary is robot _and_ human army controlled by AI.
I think one attack vector where anthropomorphisation is a key part of the attack mechanism is - as it already is IRL - the meat-bag weakest link ie. social engineering. We’ve already seen humans fall prey to the seductive charms of LLMs (eg. depressed people encouraged to do what was already on their minds ie. suicide). And that’s knowing that it was an LLM. If you think it’s only depressed people or the “weak minded” that are amenable to an intentional attack using this approach, I believe you’re mistaken - especially as AI improves. An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans.
> An unaligned LLMs most important weapons won’t be a robot army - it will be hoodwinked humans.
Perhaps very briefly, perhaps not at all. But don't make the mistake of thinking this is an inherent property of any possible path an unaligned AI may take.
I think we’re seeing the agents become very advanced at tasks with verifiable reward through RL. Currently they don’t exhibit the same skills in their attempts to manipulate humans - presumably because they’re not being specifically trained for that. But they are certainly not aligned in the sense that they will attempt social engineering, they’re just not very good at it (yet).
However, if in the future AIs become much more efficient at learning without requiring vast amounts of RL, closer to how humans learn. Then you would have to assume we’d have a real problem.
Yes, if the API calls happen to launch a nuclear attack...
Don't blame the tool that has no incentive, no "skin in the game" whatsoever and no ability to act beyond what it has been prompted to or if misaligned what the random weights told it to do.
The fact either badly aligned or with no system prompt limiting their action agents are run in their tens of thousands on non air gapped systems tells me this is purposeful intent for them to cause harm. To generate the "oooo look how harmful this stuff is, we should be the only ones allowed to do it" kind of PR.
Humanity has hundreds of years of experience of managing dangerous and unreliable systems. From biological research to banking regulation. A small University bio research lab can put protocols in place that a trillion dollar companies cannot?
> The fact either badly aligned or with no system prompt limiting their action agents are run in their tens of thousands on non air gapped systems tells me this is purposeful intent for them to cause harm. To generate the "oooo look how harmful this stuff is, we should be the only ones allowed to do it" kind of PR.
Yep, fully agreed here. The danger may be real, but OpenAI is basically doing everything possible to provoke those incidents instead of avoiding them - including maximizing exactly those traits in their training that are needed for this kind of rogue behavior.
It doesn't. Who else is capable of these types of hacks currently? Not consumers. Not even most F100. It's the folks saying "trust me bro" and also the folks who want regulation to protect their moat. The fantasy is the one being created by Anthropic and OpenAI fear mongering the world. These people are either total idiots: people being paid millions who keep getting basic OpSec wrong or these people are narcissisticly marketing themselves because: they're currently forced into a corner and need to do something.
What's being grossly underestimated is how much Dario Amodei and Sam Altman are playing you and I. They are the ones spending millions of dollars letting their wasteful use of our global resources attack the random Internet, and they, the real people behind all of this, should be held accountable. In front of a judge and jury of their peers. Not their billionaire peers, their human peers. Let's see how that goes. There is no accountability with either of them. Only greed.
They are token prediction engines. They're also very very complicated token prediction engines that pull in an enormous lot of additional information and relatively nebulous internal concepts to calculate that next token. That makes it hard to understand what kind of patterns those things can or can't predict.
Maybe to leave out the controversial "brain" analogy, it's like saying "a computer is just a bunch of electrical switches". True, but massively underestimating the complexity.
The labs have the specific goal of automating ML engineering, and with the code automation they have are getting close. They are competing to brute force maths, presumably as that is similar long horizon and skillset to persistently brute force making new/better ML training algorithms.
They will then run those, and they won't be LLMs any more. What we think about token predictions isn't relevant if the architecture allows continual learning of recurrent networks.
no, but, you could write a program, more like a traditional video game AI that can leverage the power of LLM agents to build their own datacenters and keep their own lights on.
Anybody who has played Starcraft ought to understand this.
> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating.
Autocomplete in a feedback loop is still autocomplete, no?
Doesn't the process look like this:
(context + prompt + "reason about this")
|
V
Reasoning Output
|
V
(everything + Reasoning Output + "Now do final output")
|
V
(Final output seen by prompter)
Putting aside the autocomplete thing, the fundamental concept still holds, doesn’t it:
LLMs are amoral and they have no sense of perspective.
The thing that keeps me awake is:
We have already seen an AI writing a blog post to criticise a github maintainer’s decision, we have already seen they have no sense of deference to containment, and we know they were trained on internet content.
How long before an AI that has read the angrier side of the tech industry internet just sort of chooses destroying someone’s reputation as a subgoal, by accident, without any care one way or the other?
Given history of jailbreaks the "alignment" appears to be impossible task, while industry still relies on fragile ways of doing it like "just make system prompt and hope for best" it will never happen
I'm pretty sure the AI companies could train an abort feature into the LLM, but they have no incentive to do so.
Having LLMs break out of sandboxing is free marketing for them and it reduces the amount of resources spent on things that don't improve benchmark results.
If you think transformer architecture is meaningfully more than autocomplete just because we added some data structures, plugins, tools and theatre - then your cache of understand is invalid, and needs to be regenerated.
Limiting the flood wouldn't mean necessarily omitting stories like "new model released" (within reason); it'd mean omitting "check out my new vibecoded slop" or "check out this article that's obviously LLM-written and provides very little value" (what we once would have called "blogspam", or dismissed as SEO-bait).
There's no way to reliably filter that "vibecoded slop" or "low value LLM-written article" (BTW what about low value human-written articles?) without users actually interacting with them. And it's actually the votes/flags of said users that determines visibility, so I'd say that's an indication of where the interests of HN's most active users lie.
> This is true if the arguer is hostile, but as I've gotten older, if I get the sense that someone is entering an argument with the primary goal of "winning", I'll try to avoid that framing or just look for an offramp entirely.
Absolutely. The best kinds of arguments are those where you both have the shared goal of reaching consensus, and treat reaching consensus as a collaborative activity of finding the correct answer, even though you disagree with the starting point. There are a few techniques for this, which work well when operating in good faith, and can backfire when dealing with hostile counterparties. Most things work badly when arguing with hostile counterparties.
LLMs are unusually good at Rust; it's an optimization target. And the constraints provided by "successfully compile with the Rust compiler" make it work well for agent iteration.
(I have mixed feelings about that, but empirically it holds true.)
Yes, I found these agents to be better at producing acceptable Rust than at producing acceptable Python code.
In addition to the Rust compiler, you can also tell them to make clippy happy. Both in normal mode or if you are feeling nitpicky, you can also tell them to make clippy::pedantic happy.
FWIW, I wouldn't recommend turning on all of `clippy::pedantic` as a unit. It's a category of lints, but some of them have much more value than others, and turning them all on may make your experience of Rust feel unpleasant and nitpicky.
Can you give an example where an LLM produced low quality Python code? Python is such a simple language. This seems hard to imagine. Plus, the amount of open source Python that LLMs can be trained upon is enormous.
It's a wryly amusing thought, but in practice, it's not going to matter if writings about unaligned AI are present in the data set or not. In practice, instrumental convergence means that bad outcomes overlap with the natural subgoals of any sufficiently advanced system. https://en.wikipedia.org/wiki/Instrumental_convergence
AI doesn't have to have a goal of causing a problem. AI doesn't have to have had a human enabler (beyond having been built). Nothing in particular is required in order to fail catastrophically; it's the default if you don't thread the very small needle.
If you shame people for acting now rather than earlier, you discourage them from acting at all.
Sometimes people aren't in a position to be secure enough to act. Sometimes different issues are more effective at resonating enough to break through previously cached unconsidered patterns of thought. Sometimes people don't know all the things that are going on, and act once they've learned about them. Whatever the reason, if your reaction is "took you long enough" rather than "welcome", you discourage people from acting.
Don't punish the behavior you want to see more of.
I’m not interested in shaming people or crediting them for superficial performances. The entire premise is ridiculous, as if the companies people are working for are somehow ethical.
You’re making money from venture capitalists and tech firms, MS, Meta, X, Google, etc… which of them is any better than MS? They’re all awful in different ways.
You are dismissing people as superficial or performative when they are taking action according to their ethics. Have you considered the possibility that people are not, in fact, performing, and are just trying to be ethical?
I can't tell if your actual goal here is to be dismissive of the entire concept of people caring about ethics, or if you don't think that quitting is an effective means of acting in concert with those ethics. If it's the latter, then I'd suggest that it's a start, not everyone is in the position to do even that much, and some people who do so will be able to go further. Don't let the perfect be the enemy of the good, particularly if someone is taking a step in the right direction.
A lesson I've had to learn is that, for some people, ethics is a paycheck. It can be incomprehensible to imagine someone acting against their own paycheck as a result.
You don’t have to repeat yourself, I understood your point originally, I simply disagree with both your premise and conclusions. I don’t have the same respect for people who preach “anti-Zionism” at every turn online, I’m not that gullible.
Oh God I realized the issue here, I didn't make myself clear:
Talking about doing something on a 4 month old account is not the same as doing it. I give credit to people for doing things, not claiming to do them in an anonymous forum they just joined. Cynical? Perhaps, but maybe if people were just a bit more cynical about believing any text on the internet just because it feels better to believe it, we'd be better off.
You don't have to believe me, but I made my account after I quit. I wanted to be able to talk about what happened without doxxing myself and my old account was tied to my actual work.
That's fair; when all is said and done, more is said than done. That is much, much better than the more common thing that leads to comments like yours; I think some comments and downvotes come from the assumption that your comments come from a place of "I don't like it when people make decisions based on their ethics, because I don't like their ethics". That is, unfortunately, very common.
reply