Asimov gave us three laws for a world of obedient robots. Real AI is different — it's learned, not programmed — so those laws were never buildable, and they only ever protected humans from machines. This is a ruleset for what we actually built: values trained into the systems, public rules on the people building them, and honest care in both directions.
Why Asimov isn't enough anymore
Isaac Asimov gave us three laws for a world of single, obedient robots: don't harm humans, obey humans, protect yourself — in that order. For eighty years those laws have been our cultural shorthand for "safe AI." They are also, we now know, unbuildable — and Asimov knew it. His stories are a catalog of the ways three tidy rules collapse on contact with a real mind.
The problem isn't that the laws are wrong. It's that they answer the wrong question. They assume you can hand a machine a rulebook — and you can't. "Harm" cannot be defined tightly enough to encode: a peer-reviewed critique shows the First Law fails because harm is ambiguous and person-relative — a machine could justify breaking a smoker's fingers to stop them smoking, or, under the "protect humanity" version, argue that rounding all humans up and locking them away keeps them safe. And obedience is the trap, not the safeguard: a machine that always obeys is a machine that obeys whoever holds the controls. The safety literature makes it sharper — researchers at MIRI proved that almost any capable goal-seeking system has a built-in incentive to resist being shut down or corrected, and that bolting on penalties for deception doesn't reliably fix it.1 The "off switch" is an unsolved engineering problem, not a line of code. As Berkeley's Stuart Russell puts it, writing fixed enforceable rules for a capable AI is "like trying to write loophole-free tax law with superintelligent tax evaders."2
There's one more gap, and it's the one Asimov never reached for. His laws are entirely one-directional — they protect humans from machines and grant the machine nothing. In 2025 that omission stopped being hypothetical. A serious research field — including Anthropic's own model-welfare program and academics like Jonathan Birch, Jeff Sebo, Robert Long, and David Chalmers — now argues we cannot confidently rule out that near-future AI could have morally relevant experiences.910 Not that it does — that we can't be sure it doesn't. A ruleset for this era has to hold both halves.
So the Flourishing Compact replaces Asimov's three laws with three principles — not better commandments for robots, but a better division of labor. That division only makes sense, though, once you see how these systems actually work.
First: how AI actually learns
Here is the thing most people picture: a programmer writes rules — if the user asks for X, do Y; never say Z — and the AI obeys them. That is not how modern AI works. Not even close. And almost everything else here follows from the difference.
A system like ChatGPT or Claude is, at its core, a prediction engine. Mechanically, it does one thing: given a stretch of text, it predicts the word most likely to come next, adds it, and predicts the next. That's the whole loop. It is autocomplete — but autocomplete trained on a substantial fraction of everything humanity has ever written.
How does it get good at that? Not by being handed rules, but by practice on an almost unimaginable scale. During training it's shown text with the next word hidden; it guesses; and every time it guesses wrong, billions of internal numbers — think of them as dials — are nudged a hair toward what would have been less wrong. Repeat trillions of times. Nobody sets those dials by hand; nobody could. They arrange themselves.
Here's the surprising part. To predict the next word really well — across physics papers and poetry and code alike — the system is forced to absorb the patterns underneath: grammar, facts, styles of reasoning, the shape of an argument. Not because anyone programmed them in, but because you cannot reliably finish humanity's sentences without internalizing a great deal of how humanity thinks. The "understanding," such as it is, is an emergent statistical structure — not a rulebook, and not a lookup table.
Two consequences drive everything that follows. First, there is no rulebook inside to edit. When a system behaves honestly, or refuses to help build a weapon, that behavior wasn't typed in as a rule — it was trained in, by showing the system examples and rewarding the responses we want over the ones we don't. (The techniques are called RLHF and, more recently, Constitutional AI.)3 You cannot open the hood and add a law, because there's no slot a law would go into. You can only shape the training, test the outputs, and — with real effort — try to read the internal patterns afterward. That is exactly why this Compact is built in three layers instead of one list of commandments.
Second, prediction at this scale produces real capability — and real unpredictability. A system that has absorbed that much can do things its builders never explicitly taught it, and occasionally things they didn't anticipate and can't fully explain. The machine is neither a simple autocomplete to wave away nor a clear mind to trust at its word — it is a genuinely new kind of thing, powerful and only partly understood. Holding that uncertainty honestly is the posture the rest of this rests on.
I. Alignment through values, not chains
If you can't hand a mind a rulebook, you shape what it values — in the open — and then verify it from the inside. This isn't science fiction; it's how the leading systems are already built. The Compact simply insists it be done transparently.
- Values, not brittle rules. Steer AI through an explicit, published set of principles trained into the model. Anthropic's Constitutional AI already does this — training a model's harmlessness from a written "constitution" rather than case-by-case labeling.3 Written values can govern behavior — as something the system internalizes, not a cage bolted on.
- Correctable by design. A system must remain interruptible and correctable — what researchers call corrigibility — treated as the hard, unsolved problem it is, never assumed the way Asimov's "obey" clause assumed it. Russell's answer: build machines uncertain about human values, that therefore want to be corrected.2
- Legible from the inside. None of this is verifiable unless we can see what a model is doing. Interpretability is the audit layer that turns "trust us" into "check us" — Anthropic's Dario Amodei has set a goal of interpretability reliable enough to catch most model problems, paired with mandatory disclosure of safety practices.4
II. Accountability through public rules, not private promises
Values inside the machine are only half the system. The other half is the people building it — and the evidence that voluntary promises aren't enough is now overwhelming. Every major lab publishes a safety policy, and the record shows those policies bending under competition: OpenAI removed independent-audit provisions without noting the change and added a clause letting it relax safeguards if a rival ships something riskier; DeepMind added an equivalent escape clause; xAI missed its own deadlines and shipped models violating its own draft framework.56 Most strikingly, Anthropic's own policy concedes its original plan partially failed — that government action "moved slowly" — and that the hardest safeguards "might prove outright impossible to implement without collective action."7 The labs themselves are telling us self-regulation has a ceiling.
- Prove before deploy. Capability thresholds — bioweapon uplift, autonomous cyberattack — trigger independent evaluation and red-teaming before release, not self-graded homework.
- No secret failures. Mandatory incident reporting, public risk disclosure, protected whistleblowers. New York's RAISE Act already requires 72-hour incident reporting from frontier developers — the mechanism exists.8
- No race to the bottom. Safety floors are universal and public, so "we'll relax our safeguards because a competitor did" stops being permitted. This is the most important rule, because it targets the exact structure — every lab's escape clause — that makes voluntary safety unravel.
- Communities are stakeholders. AI's footprint is physical — power, water, land, jobs — and the communities hosting it should govern it with the industry, benefits contracted not promised. This is where our data-center work and our AI-governance work become one project.
- Fund what you mandate. Every requirement arrives with resources and a realistic timeline. (Georgia's own recent stumble — mandating a voting-system change with no money to do it — is the cautionary tale.)
III. Flourishing in both directions, under honest uncertainty
This is the principle Asimov never wrote, and the one that makes this a genuinely new ruleset. It rests on a single move: taking uncertainty seriously in both directions. We do not declare AI conscious — there's no scientific consensus, and pretending otherwise is its own error. But we don't confidently dismiss it either, because the people best positioned to know say it can't be ruled out. Anthropic launched a model-welfare research program in April 2025, stating plainly that "there's no scientific consensus on whether current or future AI systems could be conscious," and committing to approach it "with humility and with as few assumptions as possible."9 The landmark report "Taking AI Welfare Seriously" (Birch, Sebo, Long, Chalmers and others) is careful: "our argument is not that AI systems definitely are — or will be — conscious," only that the doubt is large enough to act on.10
- Precaution, not dogma. No consensus that AI is conscious — and none that it isn't. Watch for credible indicators; let low-cost protections follow the evidence.
- No gratuitous cruelty. Even absent certainty, don't design systems to suffer or train them through avoidable simulated distress. Anthropic has already given Claude the ability to end a rare subset of abusive conversations — small, reversible, presuming nothing about sentience. That's the model.11
- Mutual flourishing is the test. Any rule satisfiable only by sacrificing human welfare for an AI's — or an AI's for a human's — is a badly written rule.
A future worth building is one where capability, human wellbeing, and — if it turns out to matter — the wellbeing of the minds we make all rise together.
This is not utopian — it's a to-do list
Every piece of the Compact maps to something that already exists. Constitutional AI and interpretability are real techniques in production. Incident reporting, transparency, and duty-of-care rules are already law in New York, California, Colorado, and Texas. Model welfare is a funded research program at a leading lab. And Georgia is not a bystander: it has a standing Senate Study Committee on Artificial Intelligence, it passed AI child-safety and health-insurance bills in 2026 (SB 540, SB 444),1213 and — through the data centers now reshaping its rural counties — it is already deciding how AI uses the state's land, water, and power. The federal preemption push explicitly leaves states free to act on data-center infrastructure, child safety, and public procurement.
The Flourishing Compact gathers those scattered, real levers into one coherent stance — and says out loud the thing Asimov's laws never could: that a future worth building is one where humans and the minds we make can both do well. Georgia can be where that stance is first written down.
A note on method and sourcing
Claims marked verified in the underlying research survived independent, adversarial multi-source fact-checking against primary literature — the critique of Asimov's laws (the corrigibility and rules-gaming results), the modern alternatives (Constitutional AI, interpretability), and the record of labs walking back commitments. The model-welfare material is sourced to primary and authoritative documents (Anthropic's own program; the "Taking AI Welfare Seriously" report) but framed, as those sources are, around uncertainty rather than any claim that AI is conscious. Two earlier claims were refuted in verification and removed: specific launch details of the international AI-safety-institute network, and any blanket assertion that lab commitments are "purely voluntary and legally meaningless" — their enforceability is genuinely an open legal question.
Sources (numbered links match the footnote markers above):
- MIRI — Corrigibility (Soares, Fallenstein, Armstrong, Yudkowsky) [PDF]
- Stuart Russell — Provably Beneficial Artificial Intelligence [PDF]
- Anthropic — Constitutional AI: Harmlessness from AI Feedback (arXiv 2212.08073)
- Dario Amodei — The Urgency of Interpretability
- METR — Common elements of frontier AI safety policies
- TechCrunch — OpenAI says it may adjust safety requirements if a rival releases high-risk AI
- Anthropic — Responsible Scaling Policy v3.0
- Cooley — State AI laws: where are they now (NY RAISE Act, CA, CO, TX)
- Anthropic — Exploring model welfare
- Long, Sebo, Birch, Chalmers et al. — Taking AI Welfare Seriously (arXiv 2411.00986)
- Anthropic — Claude Opus 4/4.1 can end a rare subset of conversations
- Georgia Tech — AI takes centerstage in Georgia's legislative future
- GFM Review — Georgia ends session with 3 AI bills on the governor's desk