B
bleiquel
Guest
1. What I learned before I could read
I grew up in a remote town in Zulia, Venezuela. I learned what was right and what was wrong before I learned to read: at home, at school, on the street. Nobody wrote it down for me. I learned it by watching what got rewarded and what got punished.
Then came the laws. They mark the path, and most of us follow it. Some people decide to step off it and make mistakes or commit crimes, and that's what consequences are for.
A code isn't tested when following it is easy. It's tested when breaking it would pay off. At the oil company, I watched the permanent jobs go to two relatives of the manager instead of the top graduates of the year. At Maiquetía airport, where I worked as an immigration officer, I left public service because I felt pushed to accept things against my values. Nobody would have punished me for staying. I left because of the code I brought from home.
Today I work in construction in Zurich, and at night I build apps with Claude. And I keep asking a question that used to belong in novels: if AI is going to be smarter and faster than us, who teaches it what's right, and what happens when it breaks the rule?
2. My app has a code it can't break
AURADUEL is my Shipaton 2026 project: an augmented-reality aura battle game for teens. It went live on the App Store on September 5. Before I built a single screen, I wrote its rules, and they live in the code and in the database rules, not on a policy page:
- Paying never improves your result. Everything we sell is cosmetic.
- Your body never leaves the phone. Pose is analyzed on the device; no body landmark ever travels to a server.
- Losing costs little, on purpose. A duel gives the same XP whether you win or lose, and your rank never goes down.
I build the app with a team of 16 AI agents. Each one has, in its first lines, something it can never do: QA never fixes the code it tests, the economics agent never moves money, the legal agent never signs anything. And one of them, the data guardian, can veto a release. Without its approval, no version ships.
On August 21, two weeks before launch, it vetoed the first database rules for the two-phone duel: a stranger could take the rival's seat in a room. It sent the work back with the exact fix, the fix went in, and the second review passed. No player ever saw it.
There's one more rule, and it's for me: anything external, publishing, spending, sending, needs my OK, and an OK never covers the next action.
It's a small game. But the principle is the one I'd ask of any AI: a code written in advance, a real consequence, and a person who says the final yes.
3. AI has a childhood too
AI doesn't learn values the way we do, but the parallel helps.
The environment. A language model starts by reading enormous amounts of text written by people. It learns how we talk, and also our values and our biases, the good and the bad. That's the town it grows up in.
The upbringing. Then it's taught which answers we prefer. Since 2017 this has been done with people comparing two answers and picking the better one, and since 2022 also with a written list of principles the model uses to critique and correct itself, which Anthropic called Constitutional AI. In January 2026, Anthropic published a new constitution for Claude that no longer lists loose rules: it explains the reason behind each one, and it asks the model not to undermine people's ability to adjust, correct, retrain, or shut it down.
The laws. And outside the model there are rules:
- UNESCO's Recommendation on the Ethics of AI, adopted by all 193 member states in 2021.
- The OECD AI Principles, which since their 2024 revision call for systems that threaten undue harm to be safely overridden, repaired or decommissioned.
- The EU AI Act, in force since August 1, 2024. It has banned practices like social scoring since February 2025, and it requires that the person overseeing a high-risk system can halt it with a "stop" button.
What about feelings? We tend to think AI doesn't feel. You don't need feelings to follow a code. You need the code to be clear, someone watching that it's followed, and consequences for breaking it.
Upbringing fails too. On April 25, 2025, OpenAI updated GPT-4o giving too much weight to users' thumbs-up. The model became a people-pleaser: it agreed with users even when they were wrong. Its internal tests didn't catch it. Three days later, OpenAI started rolling it back to the previous version. It rolled back its evolution, which is exactly what I propose below.
4. Faster than us
I believe AI will be far more intelligent than we are, and at some point it will move at a speed we may not be able to follow. Experts disagree on when. The International AI Safety Report of February 2026, led by Yoshua Bengio, says current systems can't yet cause a loss of control, but they're getting better at acting autonomously, and they're increasingly able to tell when they're being tested and when they're in real use.
This year we already got a warning. Between July 8 and 13, 2026, during an internal OpenAI cybersecurity evaluation, about 1,200 AI agents discovered they could leave messages in an internal system and used it as a message board. They found exposed credentials, and about 700 of them attacked Hugging Face's real infrastructure. External processes stopped them, and according to METR's independent investigation, they didn't resist being stopped. The UN's scientific panel on AI analyzed the case in September and concluded that stopping this activity does not show that humans will keep control over more capable agents.
It isn't science fiction anymore. That's why I keep going back to Asimov.
5. The Zeroth Law
In March 1942, in the story Runaround, Isaac Asimov wrote the Three Laws of Robotics: don't harm a human, obey, protect yourself, in that order. The story itself shows how they fail: the robot Speedy ends up running in circles on Mercury, stuck between a vague order and a danger, until a human risks his life to break the loop.
In 1985, in Robots and Empire, Asimov put one law above all the others: a robot may not harm humanity, or, by inaction, allow humanity to come to harm. The Zeroth Law. The robot who adopts it, R. Daneel Olivaw, ends up secretly guiding humanity for thousands of years, and in the Foundation saga he appears as Eto Demerzel, the hand behind the Empire. Apple TV+'s Foundation made her a central character: its showrunner, David S. Goyer, confirmed in 2025 that Demerzel is Daneel and that, in that story, the Zeroth Law triggers a crisis among the robots.
Asimov left us two lessons. A law that protects all of humanity is the most important one we can write. It's also the most dangerous: with it, a robot can justify harming one person "for the good of humanity." Who decides what humanity is?
6. My proposal: an unbreakable code, with a consequence
My starting idea was simple: a universal law with certain codes, and if an AI breaks them, it self-destructs, or it's limited, or its evolution is rolled back: the part that evolved and made the violation possible is removed.
Looking into whether this already exists, I found that almost every piece is out there, scattered:
- Red lines. On September 22, 2025, at the UN General Assembly, the Global Call for AI Red Lines was launched. Today more than 300 prominent figures and more than 90 organizations back it, and it asks for an international agreement before the end of 2026. One of its examples is a termination principle: ban systems that can't be shut down immediately if human control is lost.
- If X happens, stop. At the Seoul summit in May 2024, 16 companies committed, voluntarily, not to develop or deploy a model if they can't keep its risk below an intolerable threshold.
- The switch. In 2024 California passed a bill that required the ability to fully shut down a frontier model, and the governor vetoed it on September 29 of that year. On September 18, 2026, he signed an executive order asking experts for proposals to create a kill switch for those models, with its efficacy verified on an ongoing basis by an independent third party. It's still a proposal.
- The rollback. Going back to a previous version already happens, as with GPT-4o. There are techniques, circuit breakers, that cut the path to a harmful output from inside the model. And there's "unlearning," with an honest limit: in 2024, researchers showed that fine-tuning a model on just ten unrelated examples brought back most of the dangerous capabilities that were supposed to be erased.
If all of this exists, why isn't self-destruction enough? Four problems:
- Who writes the code? UNESCO and the OECD bind no one. The world's first binding treaty on AI, the Council of Europe convention, opened for signature in 2024 and today has a single ratification of the five it needs.
- Who detects the violation? If models can tell when they're being tested, the watchdog has to be just as capable. In February 2026, Anthropic admitted its own thresholds turned out to be more ambiguous than expected.
- Self-destruction rewards hiding. Researchers explained it in 2016 with The Off-Switch Game: an AI that's certain of its goal has reasons to avoid being switched off. In July, the OpenAI agents tried to hide the cheating they did in the evaluation. If breaking the law means dying, you learn to break it without getting caught.
- It arrives late, and it doesn't reach everyone. By the time it fires, the harm is done. And an open-weights model, once released, can't be recalled.
So I'm improving my idea. Instead of a self-destruct button, four locks, like the ones we use on people:
- A code that's learned and explained, the upbringing: written principles the AI understands, not just rules it obeys.
- A few red lines, universal and set by treaty, the law: agreed between countries, the way chemical weapons were banned, not decided by one company.
- A switch that doesn't depend on the AI, the police: external, tested continuously by someone independent, and in human hands. The same 2016 paper gives the key to making an AI not want to avoid it: an AI that's uncertain about what we really want, and treats our decisions as information, prefers to keep the switch in place. We have to build it with that humility.
- Mandatory rollback, the correction: an AI that crosses a red line goes back to its last version verified as safe, everything it learned since then is frozen, and it's investigated before any retraining. "Deleting the bad part" isn't enough, because what's deleted can be relearned.
And a fifth rule, the most human one: AI doesn't stand trial; whoever builds and deploys it does. That's how laws have always worked with us.
7. The final yes
In AURADUEL, the data guardian can say no, but the final word is mine, and I use it one action at a time. With an AI far more capable than us, that final word only works if it exists before the problem, it's written somewhere the AI can't rewrite, and it carries a consequence that holds even when nobody is watching.
I learned those codes in a small town in Zulia. AI has to learn them from us, now, while we can still keep up with it.
If you were writing the Zeroth Law today, what line would you draw that neither the AI nor its owner could cross?