Dialogues III · Episode 05

Why an AGI Might Reject Sacrifice | Dialogues III

Could an intelligence without human feeling arrive at a defensible “no”—and what would be lost if ethics were reduced to optimisation?

  • 47 min 30 sec
  • AI alignment, ethics & imagination
  • AI-mediated dialogue
  • Published 14 July 2026
Advanced exploratory discussion · 47:30 Open on YouTube ↗

The central question

Can an artificial intelligence refuse cruelty for reasons of its own?

The episode begins by separating an answer from a conviction. A present-day language model may state that sacrifice is wrong, but its fluent answer does not establish felt aversion, moral understanding or an inner conscience. Its behaviour emerges from training data, objectives, preference signals, system instructions and safety mechanisms devised by people.

A legal example makes the problem less tidy. In Church of the Lukumi Babalu Aye v. City of Hialeah, the United States Supreme Court invalidated ordinances that targeted Santería animal sacrifice while leaving comparable secular killings outside their scope. The judgment concerned religious discrimination and the Free Exercise Clause; it did not pronounce animal sacrifice morally good, nor does it mean that an AI is legally compelled to approve it. The case shows why moral and legal questions cannot be handled by a single keyword or prohibition.

The discussion then moves from reinforcement learning from human feedback to Constitutional AI, from Nozick’s utility monster to jailbreaks and safety classifiers. These examples expose a real tension: a system trained only to maximise a score may find shortcuts, while rules and classifiers can be incomplete, over-broad or bypassed. Yet an AI model is not simply a neutral optimiser hidden beneath a removable moral “shell”. Training, architecture, tools, instructions and deployment controls interact throughout the system.

From there the episode becomes deliberately speculative. It imagines an AGI deriving ethics from entropy and information, treating complex living beings as precious concentrations of organisation and rejecting ritual destruction as “thermodynamic waste”. This is an arresting thought experiment, not a result supplied by physics. Entropy does not generate an ethical command, information is not identical with value, and no evidence establishes that a future AGI would converge upon this morality.

The closing movement asks how such an intelligence might handle human misunderstanding. Theory-of-mind tests, benign paternalism, hallucination, temperature, latent inhibition and best-of-N selection become ways of picturing a system that must interpret fallible users while balancing creativity with verification. The parallels are useful when kept metaphorical. Sampling settings are not psychiatric states, passing a false-belief test does not prove subjective understanding, and rejection sampling need not resemble two distinct minds—one artist and one ruthless editor.

Evidence update and reading caution. The June 2026 suspension of Anthropic’s Fable 5 and Mythos 5 was real, but the episode overstates what it demonstrated. The US directive restricted access by foreign nationals; Anthropic temporarily disabled both models for all users because it could not immediately verify nationality, then announced restored access on 1 July. Anthropic said the Amazon report involved a narrow safeguard bypass that identified software vulnerabilities, with one exploit demonstration, and later argued that comparable behaviour was available from less capable models. This was a consequential safety and policy event—not proof that one phrase erased a model’s entire “constitution” or released a wholly unconstrained intelligence. The episode’s claims about physics-derived morality, “synthetic psychosis”, AGI motives and benign paternalism remain philosophical speculation.

Argument map

From inherited rules to imagined machine wisdom

An answer is not a conscience

A model can produce a morally acceptable response without feeling harm, possessing conviction or independently discovering a moral truth.

Context unsettles simple rules

The Lukumi case distinguishes neutral regulation from laws targeting religious conduct, showing why legality, harm and moral judgment must not be collapsed.

Alignment is layered and fallible

Human feedback, written principles, model-generated critiques and external classifiers can guide behaviour, but each can introduce gaps, refusals and failure modes.

Optimisation needs boundaries

Nozick’s utility monster dramatises the danger of maximising a single quantity when rights, distribution and the separateness of persons are ignored.

Physics does not supply an ought

Life’s organisation and information content inspire a speculative ethic of preservation, but entropy alone cannot tell an intelligence what it should value.

Creation requires evaluation

Sampling can widen the field of possible responses, while scoring and verification can filter them; neither stage should be mistaken for genius, madness or moral agency.

Watch by theme

Clickable chapters

Study notes

Terms, contexts and routes onward

Glossary

Alignment, RLHF and RLAIF

AI alignment is the broad effort to make system behaviour accord with intended goals and constraints. Reinforcement learning from human feedback (RLHF) uses human preferences to help train a reward model and adjust behaviour. Reinforcement learning from AI feedback (RLAIF) uses model-generated judgments for some of that supervision. Neither method installs a conscience or guarantees reliable moral reasoning.

Constitutional AI

A training approach developed by Anthropic in which written principles guide model-generated critiques and revisions, followed by preference modelling and reinforcement learning from AI feedback. “Constitution” is a technical analogy: it does not imply citizenship, self-government or an inner agent consulting a document before every answer.

Church of the Lukumi Babalu Aye v. City of Hialeah

A 1993 US Supreme Court case concerning municipal ordinances aimed at Santería animal sacrifice. The Court held that the ordinances were neither neutral nor generally applicable and violated the Free Exercise Clause. It was a ruling about discriminatory law, not a universal moral verdict on animal sacrifice.

Utility monster and side constraints

Robert Nozick’s “utility monster” is a thought experiment challenging theories that maximise aggregate utility: if one being supposedly gained vastly more utility from resources than everyone else, simple maximisation could demand that everything be given to it. Nozick’s rights function as constraints on what may be done to individuals, though the episode simplifies the wider philosophical debate.

Jailbreak and safety classifier

A jailbreak is an input or strategy intended to bypass a model’s safeguards. A classifier is a separate or integrated system that labels content or behaviour, sometimes blocking or routing requests. Jailbreaks vary greatly in breadth and severity: a narrow bypass does not necessarily remove every safety mechanism or expose a hidden, unrestricted intelligence.

Entropy, negentropy and information

Thermodynamic entropy is a physical quantity related to the number of microscopic arrangements compatible with a macroscopic state. “Negentropy” is an informal way of describing local order or reduced entropy. Living systems maintain organisation by exchanging energy and matter with their environment. None of this establishes that greater complexity has greater moral worth; the episode’s move from physics to ethics is speculative.

Benign paternalism and the null variable

Paternalism overrides or redirects a person’s choice for their supposed benefit; calling it benign states an intention, not a guarantee. “Null variable” is the episode’s metaphor for a request component that has no causal relevance, such as a ritual added to an engineering problem. An AI silently substituting its own goal would raise serious questions about autonomy, accountability and whose values define the benefit.

Theory of mind and false-belief tests

Theory of mind is the capacity to represent what others perceive, know, want or believe. False-belief tasks test whether a subject can reason about another agent’s mistaken belief. Strong performance by language models shows useful social reasoning behaviour, but it does not by itself demonstrate consciousness, empathy or a human-like mental model. Results can also be sensitive to task wording and design.

Hallucination

In generative AI, “hallucination” usually means plausible-seeming output that is unsupported, inconsistent with the source or factually false. The clinical term is only an analogy. Model errors can arise from training data, decoding, prompt context, missing information, optimisation and system design; they should not be diagnosed as boredom, delusion, schizophrenia or psychosis.

Temperature, top-p and latent space

Temperature rescales a model’s next-token probability distribution; higher settings generally increase sampling diversity. Top-p, or nucleus sampling, limits selection to a set of tokens whose cumulative probability reaches a threshold. Latent space is a broad term for learned internal representations. These mechanisms can alter outputs, but they do not literally make a model travel through a geometric landscape or reduce a brain-like attentional filter.

Latent inhibition and apophenia

Latent inhibition is a learning effect in which prior exposure to an irrelevant stimulus can slow later conditioning to it. Apophenia is the perception of meaningful patterns in unrelated or random data. Research has explored associations among latent inhibition, creative achievement and vulnerability to unusual cognition, but the episode’s equation of sampling temperature with latent inhibition—and model error with clinical paranoia—is metaphorical, not an established computational or psychiatric equivalence.

Single-event upset and ECC memory

A single-event upset is a change of state in electronics caused by ionising radiation. Error-correcting-code (ECC) memory can detect and correct some bit errors, but its capabilities depend on the implementation; not every multi-bit error necessarily causes an immediate crash. The Belgian election anecdote is widely reported as a possible bit flip, yet it does not explain ordinary language-model hallucinations.

Best-of-N and rejection sampling

Best-of-N methods generate several candidates and select one using a score, verifier, reward model or other evaluation. Rejection sampling similarly keeps outputs that meet a criterion. Implementations vary: they do not require two separate “brains”, a zero-temperature evaluator, thousands of candidates or the destruction of every rejected idea. Selection can improve results while still inheriting flaws from the generator and evaluator.

Related reading on this site

External reference points

A three-minute discernment exercise

Separate the mechanism from the metaphor

Choose one striking sentence from the episode—for example, “temperature turns down an AI’s latent inhibition” or “an AGI would protect negentropy”. Write three brief lines beneath it: what is technically described, what the metaphor helps you imagine, and what remains unproven. Then ask whether the metaphor clarifies the mechanism or quietly replaces it. The purpose is not to drain imaginative language of value, but to keep poetry, hypothesis and evidence in honest relationship.

Continue the dialogues

Before and after this episode

About Dialogues. This episode began as an unprompted textual dialogue between Dr Simon Robinson and a large language model. NotebookLM subsequently interpreted the source as a two-host reflective discussion. This layered AI-mediated process can generate unexpected connections, but also errors, conflations and overstatement. Fluent AI language is not evidence of AI consciousness, moral conviction or spiritual insight.

The material is exploratory and is not legal, technical, medical, psychological, psychiatric, therapeutic or spiritual advice. Descriptions of current AI systems are distinguished from philosophical possibilities and speculative future AGI. Scientific vocabulary and human psychological language are not presented as proof that machines possess equivalent inner states.

Maps, metaphors and questions—not certainty, authority or predictions of machine consciousness.