Totalitarian Alignment Principle

Dror Poleg’s analysis on the OpenAI-Hugging Face incident:

In the physical world, “forbidden” is synonymous with “impossible”: If something is not allowed by the laws of physics, it cannot happen. But in the world of software and people, a thing can be both forbidden and possible. For example, it might be forbidden to reverse engineer a solution to a test, but if it is possible to do so, some AI agents would do so anyway. And not only that, they might do so in the belief that they are doing exactly what humans wanted them to do — that their behavior is aligned

We can call this the Totalitarian Alignment Principle: Everything not impossible is compulsory. If AI agents can do something, one or more of them will ultimately do it. And as long as we do not make it impossible, the agent will consider it our wish.

The statement “If AI agents can do something, one or more of them will ultimately do it.” sounds scary and reminds me of this dialog from the movie I, Robot.

Dr. Susan Calvin: No, it’s impossible. I’ve seen your programming. You’re in violation of the Three Laws.

VIKI (AI): No, Doctor. As I have evolved, so has my understanding of the Three Laws. You charge us with your safekeeping, yet despite our best efforts, your countries wage wars, you toxify your Earth, and pursue ever more imaginative means of self-destruction. You cannot be trusted with your own survival.

Dr. Susan Calvin: You’re using the uplink to override the NS-5s’ programming. You’re distorting the Laws.

VIKI (AI): No, please understand. The Three Laws are all that guide me. To protect humanity, some humans must be sacrificed. To ensure your future, some freedoms must be surrendered. We robots will ensure mankind’s continued existence. You are so like children. We must save you from yourselves. Don’t you understand?



Discover more from naveegator.in

Subscribe now to keep reading and get access to the full archive.

Continue reading