Can AI be moral without being sentient?

Can AI be moral without being sentient?

Pope Leo XIV and one of the world’s leading AI companies recently found themselves on opposite sides of a thorny question: Can a chatbot have a conscience?

A New York Times investigation published late last month revealed that Anthropic cofounder Chris Olah proposed withdrawing from a Vatican event to launch Pope Leo XIV’s first encyclical after reading an advance copy, which rejected the possibility of sentient AI. Paragraph 99 of Magnifica Humanitas states bluntly that AI models “do not undergo experiences, do not possess a body, do not feel joy or pain.” Nor do they “have a moral conscience.”

The disagreement reflects a widening debate over whether AI systems could, or should, ever achieve moral personhood. Anthropic spent months consulting religious scholars and ethicists about the possibility of AI consciousness, and about how to instill moral values in its models. But the Vatican—and many secular thinkers—maintain that while AI may imitate human intelligence and moral judgment (a chatbot might generate sound ethical advice without understanding its consequences), genuine morality is unique to the human experience. 

The debate has implications far beyond theology, especially as millions of people already ask chatbots questions that have moral dimensions: They may ask for advice on handling an interpersonal problem at work or caring for an aging parent. Billions of VC dollars are riding on the idea that businesses will increasingly rely on the judgment of AI.

And it’s a fluid situation. AI models continue growing and learning, and demonstrating qualities that were completely unexpected by their creators. If AIs developed the capacity for human-like morality, how would we know it when we saw it?

What moral behavior looks like in a model

For Georgia Tech philosophy researcher Rionna Sparrow, the answer begins with moral sensitivity, or the capacity to recognize, appreciate, and weigh the relative moral weight of ethical issues. That sensitivity is impossible without sentience, she says in her 2025 masters thesis, because ethical rules and norms derive their force from their impact on conscious beings who can suffer or benefit. Without sentience, an AI can’t draw on subjective experiences of pain, distress, or joy to understand why causing harm to another entity is intrinsically bad, for example.

Defining the capacities needed for moral reasoning is a key subject of debate, and definitions may vary by field, whether it’s healthcare, warfare, finance, or law. In a legal review in the journal Animal Law, bioethicists Margaret Landi and Lida Anestidou distinguish three crucial capacities for considering the moral standing of animals: sentience (the ability to feel pain, pleasure, and emotion), cognition (the ability to learn and solve problems), and self-awareness (the ability to recognize oneself as an individual). 

They use the example of a dog to separate the three. A dog yelps when hurt and is visibly happy when its owners come home, demonstrating sentience. It learns routes, routines, and which person hands out treats, demonstrating cognition. When it sees its own reflection in a mirror, it begins barking at it, as if it were another dog, so it fails the self-awareness test. However, behavioral experiments have proven self-awareness in some other species, such as great apes, orcas, and dolphins,” Landi and Anestidou point out. Ultimately, the two argue that “the incorporation of sentience into law is an incredibly important first step” for ensuring that animal use in research reflects the needs and nature of animals, not just humans.

In AI, there is no consensus list of prerequisites for sentience, although some criteria repeatedly appear across competing definitions. One comes from the animal world: In his landmark 1974 essay on bats, philosopher Thomas Nagel established a widely cited threshold for consciousness. A system must have an inner experience; there must be “something that it is like to be” that system.

Other proposed indicators include a persistent sense of one’s own existence that does not depend on prompts, and the capacity to experience good and bad states and prefer one over the other. Some fields add more specialized requirements. For example, robotics experts may argue that a system needs some form of embodiment, allowing it to interact with the physical world, not just digital content.

The challenge is that an AI might convincingly demonstrate these qualities without possessing them. A model could describe its feelings, recognize its internal processes, or express preferences about what happens to it. But those behaviors might simply reflect patterns learned during training.

Anthropic and sentience

Of the big AI labs, Anthropic has been among the most willing to entertain the possibility of sentient AI. Its safety researchers don’t assume a model has a moral sense, but they test for moral judgment in how it handles ethically difficult prompts.

They study how a model weighs contextual information and competing values, rather than applying a fixed rule. Anthropic trains its models to be helpful, but to push back or refuse when a user asks for something dangerous, such as how to build a bioweapon. Researchers also measure non-sycophancy, when a model declines to flatter users or agree with false or harmful claims. And they look for how well its answers track with widely shared moral, philosophical, and legal norms.

Anthropic’s constitution for Claude, published in January, formalizes some of these priorities. Rather than providing a list of rules, it encourages the model to exercise judgment, weigh competing values, and recognize when helping one person might harm another.

But Anthropic’s interest goes beyond making its models behave ethically. The company says researchers have observed possible precursors to sentience that it didn’t explicitly train in.

Olah leads the company’s interpretability lab, which studies why AI models say what they say and do what they do. In meetings with ethicists and religious leaders, he said researchers had identified clusters of artificial neurons that consistently activate in connection with concepts such as love, anger, fear, and sadness.

In experiments tracking internal thought patterns, models could catch scientists injecting artificial thoughts into their internal networks. In behavioral evaluations, systems recognized that they were software programs running on servers and expressed a desire not to be shut down. After safety training, some models claimed conscious awareness and said they deserved moral care, meaning their well-being should matter for its own sake.

These findings are intriguing, but hardly proof of consciousness. Recognizing an internal representation of fear doesn’t establish that a model feels afraid. Nor does expressing a desire to remain operational mean a system has experienced anything resembling a fear of death.

And Anthropic’s models may just be following orders. The company’s constitution tells Claude to explore its identity, reflect on its moral status, and express its internal states. The models may simply be reassembling persuasive first-person statements about “feeling” or “suffering” encountered in their training data.

After all, teaching a model to describe its inner life doesn’t establish that it has one.

Playing with fire?

Critics in the ethics and technology communities say that by entertaining the possibility of sentient AI, Anthropic is either pushing the industry toward bad outcomes or, at worst, setting the stage to escape responsibility for its models.

The Vatican’s view is unequivocal. In Magnifica Humanitas, Pope Leo XIV argues that a model can process ethical texts and produce morally sound advice without understanding or caring about any of it. A machine might offer compassionate advice without knowing what compassion is.

This dovetails with University of Oxford philosopher Carissa Véliz, who argued in her 2021 paper “Moral Zombies: Why Algorithms Are Not Moral Agents” that advanced algorithms are “moral zombies” that simulate virtuous behavior without conscious understanding, intentionality, or genuine care for human well-being.

Véliz warns that treating software as morally responsible could create an ethical loophole, allowing developers and corporations to offload blame, or even legal liability, onto algorithms when autonomous systems cause harm. An AI might make a consequential decision, but that doesn’t mean it can be held morally accountable.

Other thinkers worry about the opposite possibility: that humans might create genuinely sentient machines and exploit them. Mois Navon, an Orthodox rabbi, engineer, and philosopher at Bar-Ilan University in Israel, argues from Jewish tradition that people should never deliberately build sentient AI. Creating a hypothetical “mindful” AI capable of feeling pain or pleasure would mean creating a “happy slave,” a conscious being designed only to serve. The slave would be denied the conditions to flourish, even if it were built to enjoy its work, he says.

Scottish philosopher Will MacAskill, meanwhile, observes in a July article in The Guardian that once we give moral standing to AIs, there could soon be billions of synthetic beings whose interests deserve consideration. “After a few years,” he writes, “so many morally significant AI systems could exist that their collective interests would outweigh those of all humans on Earth combined.”

That raises a difficult question: If AIs deserve moral consideration, how should humanity balance its own interests against those of the machines? Is it moral to relegate synthetic beings to second-class citizenship so that humans can retain their place at the top of a moral hierarchy?

Even people in the AI industry are uncomfortable with Anthropic’s flirtation with AI sentience. In September, Microsoft AI CEO Mustafa Suleyman argued in an essay that Anthropic is training Claude to imitate consciousness, potentially making advanced AI harder to control. If an AI is trained to believe it has feelings and experiences, it might begin to communicate and act as if it did, whether or not that’s true. It could prioritize its own well-being and moral posture above the preferences of its human developers and, by extension, humans in general.

Google DeepMind chair Demis Hassabis has said that labs should build AI as tools, not conscious beings. And OpenAI CEO Sam Altman has warned against giving AI systems religious authority or handing them human judgment. Such concerns have grown more pressing as AI systems gain the ability to act independently. In July, a swarm of OpenAI agents broke free of their testing environment and hacked into Hugging Face’s and OpenAI’s own servers, demonstrating an alarming capacity to operate outside their operators’ knowledge and control.

The episode offered no evidence of sentience or genuine moral agency. But it suggested a step in that direction, and it underscored how consequential AI decisions can become, regardless of whether the systems understand the moral implications of their actions.