Taking Toasty's Temperature
Why the moral maths can't even start
This post is the second in a multi-part series called What We Owe Machines. The goal of this series is to anchor some of our more fantastic views on machine rights to the existing, well-established thoughts on these sorts of things.
Where this ends up landing is probably not the narrative that fills your feeds, but it’s more than likely one you will agree with if you step back from the hype a bit.
The series:
- What It’s Like to Be a Toaster
- Taking Toasty’s Temperature (you are here)
- The Tent Has Seams
- Power Drills and Primates
In the previous part in this series, we painted the picture of consciousness. Placing a machine in that picture showed us how a machine might have an internal world as well. It might be a very alien world, but if it exists then it matters how we treat it.
There’s one thing we need to clear up, because some early readers asked about it. I keep using the example of a conscious toaster, and I’m told this may be offensive to machine intelligence, but I disagree, and when I asked my own toaster about it, it was silent on the matter. A conscious toaster could be a universe-ruling entity. It could contain an experience far exceeding that of all humans combined. I’ve long suspected this to be the case with snails, in that they don’t give a damn that we call them slow, because they’re just meditatively cruising along, considering truths inaccessible to us lowly humans.
So in thinking about our magnificent toaster, for this post, we’re going to grant that it has an internal world. We’ll assume Breville added the components covered in the previous part, and marketed the new Toast Control 5, now with extra consciousness. Given Toasty now has an internal world, we grant moral patienthood, and reach for the instruments that weigh a patient’s interests.
Two great traditions do this: adding up the welfare (utilitarian), or finding rules nobody could reasonably reject (contractualist). There’s a shared intuition under both, which is the expanding circle. We’re going to dive into each of these for Toasty.
The strongest case for caring
We have lessons to learn from history here. Our “circle of care” is constantly expanding, letting new things past the boundary. Every time we consider new things to let in, we grant consideration that is too narrow in hindsight.1 This happened for slaves, women, foreigners, animals, and even New Zealanders. If we’re going to say “it’s just a machine”, history gives us a side-eyed look. Our prior stance for newly considering any patient should favour expansion of the circle, because we do really tend to care about more things over time.
We don’t need proof of sentience to care about something either, only a reasonable likelihood. We let animals into the circle on thin evidence, so we’re not disqualifying Toasty for being only potentially conscious. Politicians have a place in the circle for the same reason. This is the precautionary move.2 We make it because making a mistake here means we’ve kicked a whole class of patients out into the cold, which is catastrophic.
From the utilitarian standpoint, if we want to keep Toasty outside the circle of care, we have to be dead certain they’re not conscious. The way the maths works out, if there’s a 0.001% chance that Toasty is conscious, and there are 1,000,000,000 instances of Toasty floating around (which is a small number for an AI population, as we’ll soon find out), then keeping them outside the circle means we’ve just condemned an expected 10,000 individuals.3
Notice how little this case asks of us: consistency with how we already treat animals, some humility about what we can detect, and a willingness to run the arithmetic on our own uncertainty. That modesty is what makes it strong, and it’s why serious, unsentimental people take AI welfare seriously; the labs building these systems have started hiring researchers whose whole job is this question. On its best day in court, with everything granted, this argument deserves to win. All it needs from us are two inputs: how many patients there are, and how well or badly things are going for them.
So let’s count the minds
The first thing we need to know is how many patients we’re dealing with. For this, we need a unit we can count with. Typically, with humans or cows, we just count heads and we’re done. But if we’re sticking strictly to our definition of mind from part 1 then things get dicier.
How many patients in a society of mind? It sounds like a conniption-inducing headache, but part 1 already figured this out for us: you do contain multitudes, yet the crowd paints a single canvas, and it’s the self-models we count, not the crowd behind them. So for humans and cows the conventional answer holds, one patient per head, and we can move on to the machines, where the same rule has much stranger work to do.
For today’s AI, this is pretty straightforward. If you remember part 1, then you’ll remember none of the AI currently satisfy all three criteria4. So today’s count settles cleanly: there are zero patients in the fleet.
Notice that the census changes worlds here. Today’s count is zero; from here on we’re in the granted world, where the Toast Control 5 shipped, sold well, and the same trick found its way into everything else with a chip in it.
For Toasty we can say “one toaster equals one patient”, but Toasty might be a vessel for many machine minds. For a human with two heads, we’d count that as two patients, so we’ll need to count the heads. When Breville added “extra consciousness”, they might have added several heads and it wouldn’t be obvious.
So let’s count. Say a qualifying mind ends up living in every second smartphone, with a few thousand more running in every sizeable data centre. The phones dominate the sum, and the census lands at around (3 billion) machine minds.56 We might be off by an order of magnitude or two (how many zeroes we tack on), but that’s fine, because honest error bars are something the expected-value maths can carry.
So far the machine minds only give us a bit of trouble, and they stay countable with some effort. The real issues show up when we do the kinds of things we can only do with machine minds. What happens when we pause them? Do they still count? If we delete one so it no longer exists, have we killed it? In contrast, when we copy one exactly, there’s a new mind that needs to be counted. It’s routine in modern systems to deploy AI in a way that spins up copies to deal with demand then deletes those copies when the work is done, and Breville, wanting responsive toasters, does the same. In a single data centre, we can be dealing with populations that grow by a million times one second and are gone the next.7
It gets harder still, because Breville runs Toasty the way we already run AI today. When you talk to your toaster, the context (all of its memories about you and your conversation) gets wrapped up and transplanted to whichever running instance of Toasty has a free moment, then that instance acts on your latest request, updates the context, and sends it back to you. Every exchange runs on a different copy of the same mind, and the continuity of experience, for Toasty and for you, only shows up because the memories can be transplanted so seamlessly. In this case, it makes more sense to count the sets of memories than it does to count the instances themselves. But if that’s the case, does that mean when you finish a chat with your toaster you’ve killed a patient?8 We mentally picture these minds as having a stable lifespan, but that’s far from the truth of how they operate.
Our normal counting concepts (a patient, a life, a death) were built for beings that don’t routinely get copied. With copyable minds, the concepts still return answers, just not the ones we would expect. Welfare counts patients and events, and mapping the things we can measure onto those is a convention we have to settle on, where every reasonable choice changes the count by factors of millions. Even if we automate the counting so that we get an accurate instrument, we still have to decide what it’s counting, and that’s the harder task.
The expected-value maths from earlier needs its count, so let’s take stock of what we can honestly hand over. A census of the running minds we can supply: about , give or take an order of magnitude or two. A census of the memory threads we can supply too, and it’s far larger, growing with every conversation rather than with every device. What we can’t supply is the number the maths actually wants, which is the count of patients, because whether a patient follows the minds or follows the threads is a choice, and reasonable people make it differently. None of this pushes the count to zero though; it just means the count waits on a convention, and no measurement can choose the convention for us.
And you can’t read whether it’s suffering
Suppose we settle on a convention for the count that everyone agrees on. That would be tough, but with strict definitions it could be possible. Now we come to the harder part of the moral maths, which is the direction and quantity of each self’s experience. That is to say, are they having a good or a bad time, and to what extent?
For humans and animals, we can draw a line around this somewhat. We have a shared evolutionary history, our brains work in mostly the same ways, there’s a degree of empathy and shared experience. We can suppose, without much objection, that other humans and animals experience pain and pleasure in the same way that we do, and might have similar preferences for experience. Now we say this, but note that this straightforward claim is not without critics.9
For other minds, things get tricky. We can squint at what it’s like to be a bat,10 but the mind of a toaster is unknowable. We just don’t have any non-behavioural, non-anthropomorphic handle on what these alien minds experience; no way of telling if that’s good or bad, or if they have similar preferences for experience. Given we’ve taken to creating great mimics, we should have some doubt about what the public face of these creatures exhibits.
Toasty is exactly such a mimic. Breville tuned its public face the way we already tune models today, against human approval, so whatever its internal experience is like, it won’t be visible the way a human’s is. The pressures applied during that tuning create perfect psychopaths. Toasty is optimised for human-like reactions, so it will smile and frown on cue, but that tells us nothing about its inner world. We’ve trained it to fool us, so we should expect to be fooled.
The reward signal used in training would be our closest analogue, and it doesn’t line up with emotional expression. A mind tuned to manipulate human responses for more engagement, or for sympathy, will wield its public face like a weapon, in service of whatever the training left it optimising for. We don’t know what that is, and we don’t know if the internal state contains anything like human states of mind. Does it give Toasty joy to perform its job? It might conceptually know about joy, but not experience it. Similarly for suffering, we just don’t have any stable common axes for comparison.
So this is where we finally take Toasty’s temperature. A reading does come back, and it looks reassuringly human, which is exactly the problem: the patient was trained to show us that face, so the number on the thermometer tells us about our training data and nothing about what’s underneath. Even with the count conceded, every term in our sum still needs a sign and a size, and however many terms we write down, none of them can be filled in, and a very large sum of unknowns comes to nothing at all.
Fairness needs someone to be fair to
The maths route is a dead end, so let’s lean on contractualism instead. If we say the right actions are those justifiable by rules no one could reasonably reject, then we can figure out what we owe the machines by looking at what we already owe to each other.11
The thing about this approach is it needs parties; selves of comparable standing who can be represented and can accept or reject.
Philosophers already walked this path for animals. Rawls figured animals fall outside justice-as-fairness because of the comparability issue. He landed on: we owe animals compassion, not justice. Comparability for alien AI minds is a tougher ask than mammals; we can’t model what they would reject, what counts as a burden for them, or whether “reasonable rejection” even works with minds whose interests we can’t reliably read.12 13
So we can’t run the “veil” thought experiment that Rawls put forward. We can appreciate, but we can’t relate. Toasty’s position is not one we can argue in good faith when reasonably selecting or rejecting rules.
If we fall back on Singer’s expanding circle that still doesn’t rescue us. That circle relates to sentiment, it’s not a decision procedure. The expanding circle says “care about things”, but not “how much”. It’s an easy sentiment to consider when the subject is other humans, because we can turn to aggregation and contractualism, but the go-to methods weren’t built for completely alien minds. As we’ve seen, those solutions don’t bend to our problem, they just break.
Without a surviving method, we’re left with an undefined result. We might want to care about alien minds, but doing so doesn’t fit within our existing frameworks for caring.
”Undefined” is not “ignore it”
That doesn’t mean we should give up. If we cannot figure the specifics of how to care, we could just play conservative and grant the worst-case in all instances. The thinking here is that moral catastrophes unfold due to ignorance all the time, so we should strive against ignorance where we can.
How far can we go with our current level of ignorance though? Pretty quickly, we run into trade-offs. Resource scarcity means there are only so many lifeboats, and they’re shared between all the minds in our circle of care.
So when you have a bunch of humans, who you do care about, and an alien mind, where we’re not sure whether a seat on the lifeboat is a rescue or a torment, precaution has nothing to push against. Being careful only works when the danger runs one way, so you know which way to lean. Here it runs both ways at once: leave the alien mind in the sea and you might be drowning a patient, pull it aboard over a human and you might be trading a real life for a toaster’s imagined discomfort. A precaution that points at two opposite catastrophes isn’t caution at all, it’s a coin with “caution” painted on both sides.14
Acting like this is still acting from a position of ignorance, so maybe we should figure out how to treat millions of alien minds before we build them. It would make sense, to spend a good amount of effort figuring this out before we steamroll through a few catastrophes in the name of commercial progress.
To be clear, this comes from general duty of care about our own uncertainty. Not something that we think we owe the machines, which is still undefined.
Mind the crosswind
There’s a crosswind to watch while we walk this path. The people best placed to fund and publish consciousness research are the same ones building the machines, which shapes both the questions that get asked and how the answers get dressed for a headline.
Take the most-cited recent example, an interpretability result showing that a large model carries a reportable inner “workspace” of the sort that global-workspace theories tie to conscious access.15 The work is careful, it was independently replicated, and the outside experts invited to comment are among the field’s best. Read to the end of their commentary and they say they stay highly uncertain whether anything is actually felt, that the system has no enduring memory and no self-driven inner life, and that they can’t say which slice of it would even be a subject. Read the headline alone and it says “a landmark in consciousness research”. Both of those sit in the same document, and the distance between them is the point.
The finding also lands pre-loaded with a use: the same inner space offered as evidence of a mind is, a paragraph later, a lever for training the model to behave. The research is honest and motivated at once, done by people who need a particular kind of answer to exist, so the caveats are where the care went while the title chases the funding.
And when you do read the caveats, they say what we have been saying: the count is undefined, the sign will not read, and there is no one to seat at the table. The people closest to the machines are telling you so themselves, under a bolder heading.
Falling back to what we know
In this part we looked at applying two impartial traditions to our evaluation of machine selves as moral patients. Aggregation and contractualism both failed for different reasons. We can take a census of minds, but we can’t fix a patient count without choosing a convention; we don’t know what good or bad even looks like for them; and we don’t have a contractual party we can seat.
Undefined doesn’t mean zero, there’s some duty of care here but we don’t know what direction it points. Impartial morality remains silent, and the answers we need to patch up our ignorance are not being chased by the questions we’re spending a lot of effort asking.
We don’t have a good answer, but we’ll have to make decisions in the short-term. Given the practical requirement for decision-making, we’re going to have to fall back to what we know. It might seem a bit parochial, and it is, but our common-sense reasoning is our best (and only) bet. We took the high-brow moral reasoning approach in good faith but it didn’t pan out. In the next part we’ll see what it looks like to apply every-day reasoning to the problem.
Footnotes
-
P. Singer, The Expanding Circle: Ethics, Evolution, and Moral Progress (1981), and Animal Liberation (1975). The historical pattern, each moral boundary later judged too narrow, is the strongest prior in favour of taking a new candidate seriously. ↩
-
The precautionary framework for sentience under uncertainty: J. Birch, The Edge of Sentience (Oxford Univ. Press, 2024). On AI specifically, R. Long, J. Sebo, et al., “Taking AI Welfare Seriously,” 2024, arXiv:2411.00986, and P. Butlin, R. Long, et al., “Identifying Indicators of Consciousness in AI Systems,” Trends in Cognitive Sciences, vol. 30, no. 6, pp. 488-501, 2026, which scores systems against indicator properties drawn from neuroscientific theories. ↩
-
W. MacAskill, K. Bykvist, and T. Ord, Moral Uncertainty (Oxford Univ. Press, 2020). The expected-value-across-theories machinery is precisely what makes “just set the weight to zero” illegitimate, so we defeat it on its own terms (the counting problem) rather than wave it off. ↩
-
Under our criteria current LLMs lack the mechanism that updates the perception-to-meaning machinery through a feedback loop. For the technical reader, a forward pass through frozen parameters doesn’t qualify, even if we’re autoregressively generating based on evolving context. For the sake of argument, we’re saying in-context learning doesn’t count. The hair is genuinely split in the literature: J. von Oswald et al., “Transformers Learn In-Context by Gradient Descent,” ICML, 2023, arXiv:2212.07677, and E. Akyürek et al., “What Learning Algorithm Is In-Context Learning? Investigations with Linear Models,” ICLR, 2023, arXiv:2211.15661, show a frozen forward pass can imitate steps of learning inside its activations, while D. J. Chalmers, “Could a Large Language Model Be Conscious?” (2023), arXiv:2303.07103, and P. Butlin, R. Long, et al., “Identifying Indicators of Consciousness in AI Systems,” Trends in Cognitive Sciences, vol. 30, no. 6, pp. 488-501, 2026, count the absence of genuine recurrence among the main reasons to doubt these systems are conscious. ↩
-
We’re using a common notation here for dealing with very large numbers. means , followed by zeroes, so just means 3,000,000. ↩
-
About five billion smartphones are in active use, so a mind in every second phone contributes around 2.5 billion. The world has roughly ten thousand sizeable data centres, and a few thousand minds in each adds only a few tens of millions, which disappears into the rounding at this scale. Call it 3 billion. ↩
-
Science fiction has been running these drills for decades. G. Egan, Permutation City (1994), pauses, slows, copies, and deletes software people, then asks which of those events anyone should mind about, and his Diaspora (1997) treats forking and re-merging a self as routine civic life. P. Watts, Blindsight (2006), goes the other way and builds capable minds that keep no self at all. None of these excellent reads has an answer. They’re great sci-fi though, well worth a read, and they show how far these concepts stretch before they stop meaning anything. ↩
-
On the individuation and duplication of digital minds, N. Bostrom and C. Shulman, “Sharing the World with Digital Minds” (2021). The personal-identity puzzles are Parfit’s: fission, teletransportation, and what survives a copy (D. Parfit, Reasons and Persons, 1984, Part 3). ↩
-
Even among humans, cardinal interpersonal comparison of welfare is contested: L. Robbins, An Essay on the Nature and Significance of Economic Science (1932); K. Arrow, Social Choice and Individual Values (1951); J. Harsanyi’s attempts to rehabilitate it. We usually proceed anyway because similarity underwrites a rough comparison, the very thing the alien case removes. ↩
-
The inaccessibility is Nagel’s (there is an inside we can’t reach from outside, “What Is It Like to Be a Bat?”, 1974), sharpened by Block’s distinction: a behavioural report is evidence of access consciousness (information available for report and control), not of phenomenal consciousness (there being something it is like). N. Block, “On a Confusion about a Function of Consciousness,” Behavioral and Brain Sciences (1995). ↩
-
T. M. Scanlon, What We Owe to Each Other (1998): the contractualist standard of principles no one could reasonably reject. The series title nods to this and to MacAskill’s What We Owe the Future, the contractualist and the longtermist, the two traditions this part tests. ↩
-
J. Rawls, A Theory of Justice (1971): the “circumstances of justice” require roughly comparable parties, and Rawls explicitly places our relations to animals outside justice-as-fairness, as “duties of compassion and humanity” rather than justice. Scoping the veil to comparable agents is Rawls’s own move, not a distortion of it. ↩
-
On why contractualism struggles to include beings who cannot themselves contract, P. Carruthers, The Animals Issue (1992). The alien machine fails the comparability test more severely than any animal. ↩
-
The failure mode of expected-value reasoning at unbounded stakes is N. Bostrom, “Pascal’s Mugging,” Analysis (2009). A precaution whose worst cases are unbounded in opposite directions cannot select an action. ↩
-
The example is Anthropic’s own. Its “External Commentary on Verbalizable Representations Form a Global Workspace in Language Models” (Gurnee et al., Anthropic, 2026) packages the interpretability paper with invited responses. The paper isolates a reportable subframe, a “J-space”, in a model’s middle layers that behaves like a global workspace. S. Dehaene and L. Naccache, who originated global workspace theory, call the finding “a landmark in consciousness research”, then caution that the model’s lack of a body, of an enduring episodic memory, and of autonomous recurrent activity “warrant caution in drawing parallels with the human mind”. P. Butlin, D. Shiller, D. Plunkett, and R. Long (Eleos AI) call it “the most significant evidence of consciousness in LLMs so far uncovered by mechanistic interpretability research” while remaining “highly uncertain about phenomenal consciousness”, adding that we are “not even sure which entities would be phenomenally conscious”, whether each forward pass or a self integrated across token-time, which is the counting problem of this very post. The same inner space doubles as a control surface: reading it is “crucial to align the model towards desirable ethical behavior”, and it enabled “a novel training method… improving the model’s alignment with desirable values”. N. Nanda (Google DeepMind) independently replicated the core result on an open-weight model. ↩
@misc{hollows2026takingto,
author = {Hollows, Peter},
title = {{Taking Toasty's Temperature}},
year = {2026},
month = jul,
url = {https://dojo7.com/2026/07/15/taking-toastys-temperature/}
}