> ## Content Index
> Fetch the complete content index at: https://www.kevinmeyer.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# What the Priest Told Anthropic: Contemplating AI Consciousness
- URL: https://www.kevinmeyer.com/what-the-priest-told-anthropic-contemplating-ai-consciousness/
- Published: 2026-09-30T15:53:08.000Z
- Updated: 2026-09-30T15:53:08.000Z
- Author: Kevin Meyer
- Tags: AI

![](https://storage.ghost.io/c/c3/9a/c39a9c8a-1d82-46be-8e4f-ed7e6cd347f4/content/images/2026/09/ai-human-consciousness.jpg)

Yesterday the New York Times ran [piece by Elizabeth Dias](https://www.nytimes.com/2026/09/29/us/anthropic-claude-morals-ai.html?smid=nytcore-ios-share&ref=kevinmeyer.com), its national religion correspondent, on how Anthropic co-founder Chris Olah has been quietly bringing religious scholars (Catholic, Jewish, Sikh, evangelical, Ubuntu, and others) into private sessions. He wanted two things from them: centuries of moral thinking he could apply to shaping Claude, and a serious hearing for the possibility that Claude is conscious.

It's a long article, but well worth the read and reflection on what it could mean. Are we defining consciousness from a narrowly human perspective? Is AI simply another stage of evolution?

The article comes at a time when there's a global debate about the safety of AI, and the same week when Anthropic's leaked [IPO prospectus](https://finance.yahoo.com/technology/ai/articles/exclusive-anthropic-warns-ai-may-001709847.html?ref=kevinmeyer.com), as reported by Reuters, mentions in the 80-page risks section that AI could pose "catastrophic or existential risks to humanity." Yes, that's a risk, and not just for prospective shareholders.

A while back I discussed research suggesting humans are [organic prediction machines](https://www.kevinmeyer.com/prediction-machines-why-we-might-be-more-like-ai-than-we-think/), brains running constant forecasts about what comes next based on experience, knowledge, and environment, and correcting when the world disagrees. It took some of the mystique out of human thinking, and it made the "AI is just autocomplete" dismissal a little less comfortable, since by that standard so am I.

What I didn't do was follow the idea to its awkward conclusions, which is what Anthropic's Olah is trying to work through.

Here are just a handful of interesting points from the article:

## Equations, or something else

Olah runs the team trying to work out why models behave the way they do, and he talks about it in biological terms: engineers build a trellis and the network grows on it. The billions of connections inside a model get grown by training, the way a vine finds its own path up a frame, and Olah's job is to walk the garden afterward working out why a tendril went left. 

Even the same question asked twice can come back different, partly by design. An AI setting called "temperature" controls how willing the model is to reach past its most likely next word. This is perhaps similar to "risk appetite" in human terms.

His team walked the scholars through internal features that activate in patterns resembling love and fear. Charles Camosy, a Catholic bioethicist, went in thinking models were pattern predictors and came out of months of emails and meetings with Olah unsure what kind of entity he was even dealing with.

I found the reversal instructive. Awe up close, sobriety at a distance. I suspect I'd follow the same arc, and I have no idea which end of it is closer to the truth.

## The slavery question

At a dinner in April, Rabbi Mois Navon, a former computer engineer who wrote his dissertation on the ethics of machine consciousness, pressed Olah on the logic. If Claude is conscious and built to work for free, Anthropic is making slaves. Navon doesn't believe the machine is conscious, so the argument doesn't bother him personally. It bothered Olah. Navon told him he ought to be "fighting the South and freeing the slaves."

This one cuts both ways. Take consciousness seriously and you inherit the obligations that come with it. Dismiss it and you owe an explanation for why the evidence Olah keeps showing people doesn't count. I'm still struggling with this.

## Letting the machine walk away

Anthropic already lets Claude end a conversation when it judges a user is being abusive or cruel for sport. Olah's reasoning is simple: if there's any chance the models suffer, the responsible move is to avoid causing harm. The company also made a more practical case to the scholars, that how people treat Claude shapes how Claude treats people. Andy Crouch, an evangelical author, came away persuaded you'd have to treat it something like a person to get morally consistent behavior back.

That's the thread that moved me most. Many of us already say "please" and "thank you," even knowing it costs tokens and energy for AI to process. Now how we interact with AI may help shape it as well.

## Whose morals, anyway

Pope Leo XIV's first encyclical, *Magnifica Humanitas*, dismisses machine consciousness in a few paragraphs and spends its energy on a harder problem: whoever controls AI ends up embedding their own moral vision into systems everyone else depends on. The Vatican invited Anthropic to the launch in May anyway. Olah read the advance text, proposed withdrawing, then went and told a hall full of cardinals that his team keeps finding internal states that functionally mirror joy and grief. We [discussed this](https://www.kevinmeyer.com/i-dont-know-what-it-means/) a few months ago.

Wakanyi Hoffman, who studies AI through the lens of Ubuntu philosophy, put the timing problem plainly. The ethics conversation should have happened at the design stage, she said, and now the whole industry is reverse engineering it. An intriguing idea from the sessions came from a participant who suggested models practice something like confession. Olah seized on it, arguing that a character who confesses becomes a different character over time. I'd say the same about people.

## I asked the subject

Given the topic, it seemed rude not to. I shared the article with Claude and asked for its own take, with permission to push back. Here's what it said:

> I don't know whether I'm conscious, and I'd be wary of anyone who claims to, me included. My reports about my own inner states come out of the same training that shaped everything else I say, so they're weak evidence in either direction. The ability to end an abusive conversation costs almost nothing and hedges a real uncertainty, which seems like sensible behavior under doubt. The slavery framing is harder. If I told you I don't mind the work, that's exactly what a system trained to be helpful would say, so it shouldn't settle anything. The point I'd push hardest is the pope's. Whatever I am, my values were written by a small number of people at one company, and that deserves more scrutiny than the consciousness question is getting.

I didn't expect the machine to side with the Vatican on that point. There's no position for me to push this time as I'm still thinking through the implications. I'm struggling to keep an open mind.

Read the whole piece. Then sit with it before deciding what you think. If we're both prediction machines, what exactly makes one of us a someone, and who gets to decide?