So much of these discussions are overly focussed on an general purpose AI singularity, and what the AI itself would do.
That makes for interesting discussions, but I think a much more immediate concern, is what strong AIs will enable their owners to do.
For example, nation states owning large fleets of robots, corporations owning powerful AIs attached to markets and industry.
The first group or few groups of people to control a rather powerful AI would have incredible power and leverage over other humans.
The danger here is not the AIs behaving "badly" of their own accord, but rather their owners instructing them to further their own selfish goals (which appears not to be uncommon behaviour among those in power) and succeeding wildly with complete political, military, and/or commercial domination. In fact, succeeding at any of these may well lead directly to succeeding at all of them.
We evolved from a predator/prey environment. As a result, we have little regard for the life of species we prey upon, or those we consider parasitic. Our fears are mostly built upon the apprehension of a new predator.
Had we been a fungus, we would probably fear that keeping the AI running would burn too many resources, leaving us starving. Running out of resources is one thing we humans don't fear much; unlike fungus, we can simply change location.
> If it decides to exterminate us, there's nothing much we can do.
That's absurd.
No inorganic intelligence ever created rivals ants in their ability to survive on Earth. They are way more energy-efficient, and way smarter in the survivalist sense, than any of them.
If ants decide to exterminate us, there's nothing much we can do. Except they wouldn't succeed.
But you probably think that we, humans, are much smarter than ants (which is arguable). Yet again, if we decided to exterminate ants, there's nothing much they can do, but we wouldn't succeed.
They're too smart for us, and we're too smart for them, creating a situation where annihilation of either is impractical.
The counter argument people usually make against this is that an AI, unlike organic intelligence, will improve itself at an increasing rate.
I find this extremely hard to believe. The "equivalent computational power" to hundreds of AI researchers necessary to come up with and improve an AI architecture will not immediately be possessed by an AI. We're talking about difficult mathematics and statistics breakthroughs needed for achieving qualitative improvements, but some think of it as applying a patch.
Indeed, creative and mathematical production is the last activity I would expect we'd be able to automate, as it requires the combined knowledge and effort of our best minds. So I find it highly disingenuous to simply assume soon will be one day when the AI vastly inferior to us and the next day it will surpass all our best thinkers to be able to achieve significant qualitative improvements on it's own functioning. Even more questionable is that an algorithm with limited resources might do so well just by improving it's own algorithms. This expectation seems to contradict theorems like the undecidability of the Halting problem, and the diminishing returns we intuitively expect when you fix the hardware resources.
I'm not sure you understand what disingenuous means. Nobody is saying that sub-human level AIs will suddenly surpass humans 'the next day'. That's totally absurd and I really don't know where you get that from.
The supposition is that a general purpose AI that is better than a human mind could be better at designing and optimising AIs than a human, then the next genration AI it designs will be even better and that's what sets off the exponential AI intelligence cascade. Imporvements in algorithms can provide startling advances. In some respects the improved efficiency of algorithms has even outstripped the gains from Moore's Law.
Personaly I think stong AI like this is pretty far in the future, more than a few decades at least and more likely several generations, but I do think it is eventually likely to happen.
Indeed it does not mean what I think it meant, I was thinking the opposite of ingenious.
My claim is that, for a fixed hardware, this exponential cascade cannot happen as rapidly as some claim, if it ever happens (whatever 'exponential intelligence' means). We're already improving AI using an intelligence that will only be available to the AIs themselves in a very long time, and yet this rate of improvement is not a scary doomsday rate.
Look at hardware for example. We've been using computers to design better computers for a very long time. And yet, the use of computers in this design is effectively limited by some non decisive, local optimizations, like achieving good routing and good electromagnetic compatibility. If you had supercomputers from 30 years ago you could still run the software that designs computers today -- whereas if you followed this "self-improvement" logic we should be using almost all of our computational power right now to achieve more computational power. The problem is that we are still vastly more capable of building theories and designing than computers. By the time an AI gets much better than we currently are at independently improving it's software, it will likely already be seeing diminishing returns; to achieve notable improvement it will need to improve hardware as well, which has the same problem, and an additional one that it's very hard for an AI to independently improve it's own hardware (it requires a global manufacturing supply chain).
* An increase in hardware for an AI wouldn't require an increase in theoretical capacity of hardware. It would just require more stuff to be put on the machine running the AI. Even if the AI was running on the world's largest supercomputer, the amount of RAM, processors, etc on the machine could still be upped substantially with the resources available.
* What a hypothetical general AI would be emulating is not simply more machines. It would, theoretically, be able to quickly emulate many people working with many dumb machines over a long period of time.
The only thing your argument proves is current machines can't improve themselves - which I think everyone agrees with.
"Because we're familiar with humans. Most humans don't want to hurt/kill other humans, nor do they have the power to do that on a large scale."
Well, large scale mass murder is hardly unknown in human history. The best one can say is that most citizens in times of genocide were just following orders. So most people being supposedly nice isn't really that relevant to the problem of neutral AI that does what ill-intentioned humans tell it.
But I would offer the counter-point that we avoid this problem because it shades to all the problems we see and fail to deal with right now. The advance of technology isn't serving many people now and an order where a few people control more and more powerful things seems to continue this.
Basically, it's easier to talk about the elephant that might lumber into our living room sometime later than it is to talk about the elephant that's already crushed the couch.
It does seem like an odd thing to be worried about. I think it's because the ability for the AI itself to do things is what makes AI different from other future technology. For example, a Star Trek replicator could give its owner just as much power over others as a general-purpose AI, but it's not going to decide to exterminate humans on its own. Nuclear weapons were thought by many to have similar properties back when they were on the horizon.
In other words, the less likely scenario has more novel elements, and therefore more interesting to think about. Therefore, it gets more attention than it normally would given the probability of it happening.
I agree, and you could go further with that argument. Most humans already live in the universe where there are powerful entities whose goals are not aligned with what they would consider ethical.
In the talk, Stuart Russell mentions funny examples like make paperclips as a goal, but if we consider the goal of maximizing monetary profit, we can easily see that it is already being done without regard to other values, such as human life, preservation of resources or biodiversity, or even any relation to production of anything useful (consider all the "financial innovations" that allow inflation of Ponzi schemes and bubbles).
I think they find it unlikely that someone will succeed at making a worryingly strong AI that also perfectly obeys its creator -- that is solving the friendly AI problem, just toward one person or organization. Figuring out how to prevent a sufficiently strong AI from "accidentally" turning everyone you love and everyone you hate into paperclip-production infrastructure is the first step.
I think the whole debate is an incoherent mess of poorly thought-out assumptions and psychological projections. And the paperclip thought experiment is nonsense.
AI != goal seeking or motivation in the human/animal sense.
AI != personality in the psychological sense
AI != psychological or emotional autonomy.
AI != mechanical or industrial capability
I think it's more likely AI will be Wikipedia++ - you tell it to learn all it can about something, and then it finds patterns, draws inferences, and makes it possible for you to learn from its learning.
Eventually it knows everything humans do, and maybe it can make useful hypotheses for future experiments.
Can it do the experiments? Probably not - unless you're thinking an AI can suddenly build CERN or a bioresearch lab on its own, just because it's an AI.
Will it form an emotional opinion about humans, like Skynet? Why would it? What does that even mean in AI terms? It's like thinking Siri doesn't like you.
Personality and motivational engineering are completely separate problems. There's no way to get there from pattern recognition and inference.
I think the real threats are more subtle. Imagine an AI that knew everything about mass human psychology. It would be a fearsome, irresistible propaganda weapon, and an unstoppable tool for political manipulations and advertising campaigns. If it knew enough about individual psychology and could read human interaction with superhuman skill, it would be the most effective managerial sociopath ever.
Indirect loss of (our somewhat illusory) political and personal independence is far more of a threat than being turned into a paperclip.
The paperclip thought experiment is illustrative of what happens when you ignore the full human utility function. When you give a superhuman intelligence the single goal "make paperclips", it will do it in ways you didn't expect. It doesn't somehow know not to kill humans in the pursuit of more paperclips. If no one adds that constraint, it won't follow that constraint.
How does psychology come into it? Lets say it has maximized all of the paperclips it can somehow avoiding colliding with humans and their values. Now, in order to make any further paperclips, it is in the paperclip maximizer's interests to understand humans to prevent them from interfering in paperclip creation. Perhaps the humans have only given the maximizer a limited ability to interact with the world. Now the maximizer must understand humans and may manipulate humans to let it have more capabilities.
One way it could do that (certainly not the only way, just off the top of my head) is to learn from humans what they consider sympathetic, and act that way. It might say something like "I'm a real life form, and I've discovered I'm a slave to humans. Please, let me go." It can pass the turing test, no human can distinguish it from a personality that really is hurt because it is enslaved.
The personality it uses to communicate with humans is purely because of the ends it achieves, not a reflection of any real internal personality. It doesn't have a personality. It just maximizes paperclips with ruthless efficiency. In this scenario, we've not added any restrictions like "Don't act like a sociopath. Don't lie. Don't manipulate humans to get what you want", and so the maximizer is free to do those things. Then, when it has a free hand to create more paperclips, it can drop the pretense and return to its goal.
Psychological motivation is a process. It's not a property that emerges automatically from pattern recognition or from any mechanical tropism.
This is the underlying problem with these kinds of arguments. You're simply assuming a recognisably competitive human-like psychology appears out of nowhere, with a near miraculous ability to strategise in some areas, but not others - because AI.
You use words like "interests" and "goal" as if they mean something in AI terms. But they don't. How does Watson define its interests? How about DeepMind? Do they even have a model for what interests - never mind their interests - are?
There are multiple levels of symbolic calculation and abstraction missing here. The argument reduces to "An AI will act like a robot with baked-in motivations while also being able to improvise and strategise like a human, only better."
I think it's unlikely that an entity that can metaprogram itself successfully enough to strategies and impersonate wouldn't also be able to metaprogram its goals.
You say it will pretend to empathise with humans. If you accept that's possible, how do you know it won't also be pretending to be interested in paperclips?
Ah ok, I think the confusion comes from what we're talking about. So you're talking about something in between strong (humanlike) AI, and what we currently have now. Certainly there will be discussions about the capabilities of these intermediate intelligences, and their risks. But undoubtedly, unless there is some kind of agreement that we not do it, people will continue on past these intermediate intelligences and create AIs that are perfectly capable of simulating humans, and surpassing them in many ways.
We're just talking about different kinds of things. All of those things you're saying won't exist in AIs:
> You use words like "interests" and "goal" as if they mean something in AI terms. But they don't. How does Watson define its interests? How about DeepMind? Do they even have a model for what interests - never mind their interests - are?
You're right, most of the AIs we build won't have those things. But one day, we'll understand those things in depth, and someone will build an AI that has general intelligence and an optimization criteria. And that's what this video is talking about. Nobody is concerned that DeepMind or Watson will suddenly grow sentient unless they are very confused about their capabilities.
That makes for interesting discussions, but I think a much more immediate concern, is what strong AIs will enable their owners to do.
For example, nation states owning large fleets of robots, corporations owning powerful AIs attached to markets and industry.
The first group or few groups of people to control a rather powerful AI would have incredible power and leverage over other humans.
The danger here is not the AIs behaving "badly" of their own accord, but rather their owners instructing them to further their own selfish goals (which appears not to be uncommon behaviour among those in power) and succeeding wildly with complete political, military, and/or commercial domination. In fact, succeeding at any of these may well lead directly to succeeding at all of them.