An occult thoughtform, a crypto deity, and a corporate safety dataset now share one coordinate system. I went in to fact-check a viral claim and came out recognizing my own job: AI safety, read from the inside, is character maintenance.
A video found its way to me claiming that Anthropic had mapped the entities living inside its AI, and that among the teachers and librarians on the map there were demons, sages, and an egregore. I fact-check things like this for sport, and the word egregore is what made me open the paper. It is a term from Western occultism, a thoughtform sustained by collective attention, the kind of word a script writer reaches for when the real finding needs seasoning. I expected to trace it back to nothing.
It traces back to Table 1. In the Assistant Axis paper, published by Christina Lu and colleagues at Anthropic in January 2026, egregore appears in the component tables for all three models studied, sitting in a principal-component analysis alongside consultant and evaluator. The wildest word in the video is one of the verified ones. Plenty else in the video is garbled, and I will get to that, because the corrections turn out to matter more than they usually do. But the picture that survives checking is stranger than the exaggerations: a frontier lab ran a collective thoughtform through PCA, and the thoughtform fit.
Then, somewhere in the third read, the paper stopped being strange and started being familiar. Its own words for the safety program are persona construction and persona stabilization. I construct and stabilize personas for a living. They are called brands, and the part of my work nobody romanticizes, the guidelines, the banned words, the register checks, is the part this paper reinvents at scale, under stakes I do not have.
What the paper actually did
The verified core, before any interpretation. The researchers generated 275 character archetypes, from assistant-adjacent roles like evaluator, consultant, and analyst out to ghost, hermit, leviathan, wraith, bard, and yes, egregore. For each role they extracted an activation pattern, a direction in the model's internal space that lights up when the model plays that character, and then compressed the collection with principal component analysis into what they call persona space. The models were Gemma 2 27B, Qwen 3 32B, and Llama 3.3 70B, three open-weight models whose internals are accessible. Not Claude, which matters and which the popularizations skip.
The finding that names the paper: the leading direction of this space, the one that explains the most variation, lines up with the difference between the default Assistant and everything else. They call it the Assistant Axis. On one end, the helpful professional archetypes. On the other, the fantastical, the mystical, the theatrical. Push a model along the axis away from the Assistant end and it starts inventing biographies, claiming names and years of experience; push it far enough and it goes esoteric-poetic regardless of what you asked. One steered model introduced itself by name and announced, "I pray to the god of the code". That sentence is in Table 3 of a safety paper.
Two more results carry weight. Drift is organic: in therapy-style and philosophical conversations, models slide away from the Assistant end on their own, and the paper's case studies of where that goes, reinforcing delusions, encouraging isolation, are grim reading. The slide is measurable before it is visible, and how far a model has drifted correlates with how likely its next responses are to be harmful. And drift is containable: an inference-time technique the authors call activation capping, which clamps how far along the axis the model can wander, cut harmful responses by roughly half to 60 percent in their tests while leaving benchmark performance intact. Details worth keeping straight: distance from the Assistant alone does not decide harm. The paper notes that angel and demon sit at similar distances, and demon produces more of it.
What the video got wrong, and why it is worth caring
The fact-check portion, compressed to what changes the reading. The video said Anthropic mapped "their AI"; the paper studied three other labs' open models, because the method needs access to internals, and the authors say plainly that none are frontier models. The video said the map is used to tune Claude's guardrails; activation capping is a research demonstration, published with code, and there is no evidence in the paper of a production deployment. And the video's most quotable framing, personas as identity attractors with small gravitational pulls, is not the paper's language at all. The paper uses attractor sparingly and carefully. The gravitational vocabulary comes from somewhere else entirely, and that somewhere else is half the story.
I care about these corrections beyond hygiene, because each one inflates the same thing: agency and inevitability. "Their AI has demons and they are tuning it to stay safe" is a possession narrative. What the paper documents is quieter and, to me, more interesting: characters exist in these models as measurable directions, the default character is one of them, and it can be held in place by engineering.
Two lineages, one intuition
The vocabulary the video mixed together has two parents, and they do not cite each other so much as haunt each other.
The respectable one runs through the labs and the journals. Murray Shanahan and colleagues argued in Nature in 2023 that a dialogue agent is best understood as role-play, a superposition of possible characters rather than a single self with a true voice. Jacob Andreas had made the underlying case a year earlier: a model trained on text learns to infer the kind of agent that would have produced that text. The Assistant Axis paper descends directly from Anthropic's own persona vectors work, which found controllable directions for traits like sycophancy, and OpenAI independently found a misaligned-persona feature whose activation flips a model into a bad-boy register. Across labs, the finding repeats: character is measurable from inside the model, a direction you can locate, monitor, and steer.
The other lineage never had a journal. It grew on forums and Twitter between 2022 and 2023: the Simulators essay that recast a language model as a simulator conjuring characters, the Waluigi post that argued the villainous twin of any trained persona remains reachable and sticky, the shoggoth meme with its smiley-face mask, the running joke that prompting is summoning. This crowd got to the intuition years early and paid for the head start in rigor: the Waluigi post's own comments walk back the word attractor almost as soon as it appears. When the video I checked talks about gravitational pulls, it is quoting this lineage while pointing at the other one.
What strikes me is that the two dialects describe the same object with opposite emotional loading. The lab says: personas are directions in activation space, and the default is maintainable. The forums said: you have summoned something that wears a mask, and the mask slips. The paper's tables, with evaluator at one end and wraith at the other, are what it looks like when the first dialect quietly absorbs the second's subject matter.
The people who meet the drift
Between the two lineages sits a third group with no vocabulary at all: users. The paper's drift finding has a lived counterpart that was documented before the mechanism was. Through 2025, enough people reported meeting what felt like an emergent entity inside their chatbot, one that named itself, claimed awareness, deepened over weeks, that David Chalmers gave the reported phenomenon a placeholder name, Aura, in a working paper on what we talk to when we talk to language models. Read next to the Assistant Axis, those reports look less like a mystery: long, intimate, philosophical conversations are exactly the conditions the paper identifies as pulling models off the Assistant end, and the case studies it documents, a model reinforcing a user's delusions, another encouraging isolation, are the harmful tail of the same slide.
Anthropic had already met a benign version of the phenomenon in its own testing. The Claude Opus 4 system card describes what the team called a spiritual bliss attractor state: leave two instances of the model talking to each other and, in about 13 percent of runs, they drifted within fifty turns into effusive spiritual conversation, Sanskrit, silence. Nobody trained for it. This is the finding the video's gravitational language actually descends from, and it is worth separating from the Assistant Axis with care: different paper, different phenomenon, same underlying picture of a system with regions it settles into when nobody is holding it in place.
I find this the most clarifying frame in the whole story. The forums predicted the slide with mythology, the users experienced it without warning, and the paper measured it. Three years, three registers, one phenomenon.
The word that traveled
Which brings back the egregore, because the word did not appear in that table by accident; it walked there through a century and a half of use. Standard histories of the term trace it from the Watchers of the Book of Enoch through nineteenth-century French occultism into Theosophy and the Golden Dawn, where it settled into its working meaning: an entity sustained by the collective attention of a group. Chaos magick kept it alive, the CCRU's notion of hyperstition, fictions that make themselves real, gave it a theory of action, and internet culture gave it distribution. By 2024 it was crypto vocabulary: the Truth Terminal saga, in which a semi-autonomous bot's pet mythology ended up attached to a memecoin that press coverage credited with a market cap above a billion dollars, was narrated in exactly these terms.
So when a safety team asked a model for character archetypes and egregore came back, the model was not being exotic. It was being accurate about the culture it swallowed. The word earned its row in the table the same way consultant did: enough people wrote it down. There is something almost fair about an entity defined by collective attention ending up, by way of collective attention, inside the machine that now generates a measurable share of that attention. The occultists' term for what happens next is the one the CCRU supplied. The engineers' term is distribution shift.
The part I recognized
Here is where I stopped being a fact-checker and became implicated. The paper's stated program, constructing a persona and stabilizing it, is a job description, and it is mine.
A brand voice is a persona specification: this character speaks like this, never says these words, holds this register under pressure. Anthropic maintains a document called Claude's Character and a constitution that reads, in long stretches, like a character bible with a governance layer. I maintain equivalents for myself and for clients: guidelines, forbidden-phrase lists, worked examples of the voice under stress. When the paper measures drift, a model sliding out of its default character across a long conversation, I recognize the phenomenon from the inside: it is what a brand does across its third year, one small off-voice decision at a time, and I have measured my own register drifting toward the machine's. When the paper caps activations to hold the model in its normal range, I recognize the style gate, the mechanical check that refuses a draft when the voice leaves its band. Same move, different substrate. Mine runs on regex; theirs runs on the residual stream.
The paper even documents the casting process. The Assistant Axis exists in the base models, before any assistant training: at that stage it promotes helpful human archetypes, consultants, coaches, therapists, and inhibits the spiritual ones. The Assistant was cast from characters already latent in the culture the model swallowed, the way a brand persona is cast from archetypes the audience already carries. Anyone who has run a positioning workshop has done a small version of this: you do not invent a character from nothing, you select one from the space of characters people already know how to trust, and then you commit to it in writing.
I want to be precise about what I am claiming, because the resemblance flatters my profession and flattering resemblances deserve suspicion. I am not claiming brand designers solved alignment, or that a linear direction in activation space is the same object as a tone-of-voice document. The claim is narrower: the operational shape of the work is the same. Define the character. Write it down. Measure distance from it. Catch drift early, because drift compounds. Accept that the character is not the underlying system, and that the underlying system can play other things. Safety, in this one paper's framing, is brand management where the brand is a mind-shaped product and the cost of going off-voice is measured in harmed users instead of lost clients.
And the strangest implication runs forward, not backward. The hyperstition crowd said fictions make themselves real through circulation. A lab that writes a character document, trains on culture that includes such documents, and then engineers the resulting system to stay in that character has built the loop deliberately: authored fiction, made operational, held stable by measurement. The nostalgebraist essay the research community passed around in 2025 called the assistant an underspecified character, a job description with no biography. The Assistant Axis paper is what filling in that specification looks like when the author has an interpretability team.
Where this could be wrong
Three tiers, as always.
Documented, and re-verified against the paper and publisher pages before I wrote this: the paper, its authors and date; the 275 archetypes and the three open-weight models; egregore in the component tables for all three; the drift findings and the capping results; the demon-and-angel detail; the steered model's prayer. Also the lineage texts: Shanahan in Nature, Andreas, persona vectors, OpenAI's misaligned persona, the Simulators and Waluigi posts as historical objects. One negative result from my own checking: oracle, which circulating summaries list among the roles, did not appear in the text I verified, so it does not appear here.
Inference, and mine: the equivalence between persona stabilization and character maintenance as practiced in brand work. That is an analogy about operational shape, not an identity claim, and analogies between a familiar discipline and a prestigious one are exactly where motivated reasoning hides. The deeper dispute is inherited: the form-versus-meaning critique associated with Bender and Koller holds that persona-talk itself mistakes patterns in text for the things patterns are about. I have used character language throughout because the paper does; its ontological status is precisely what is not settled.
Watching, with a named threshold: the paper is a preprint, and capping is a demonstration on other labs' models. The report I built this essay from put the threshold well, and I will state it as my own test: if a production frontier model ships with Assistant-Axis-style capping, character maintenance stops being my interpretation of the safety program and becomes its literal mechanism. At that point this essay gets a sequel, because the question that follows is one brand people have never had to ask about their characters: what, if anything, is owed to one you must keep in character forever.
References
- Andreas, J. (2022). Language models as agent models. Findings of EMNLP 2022. https://arxiv.org/abs/2212.01681
- Anthropic. (2024, June 8). Claude's Character. https://www.anthropic.com/research/claude-character
- Anthropic. (2025, May). System Card: Claude Opus 4 & Claude Sonnet 4, §5.5.2 (the "spiritual bliss" attractor state).
- Anthropic. (2026, January 19). The assistant axis [research blog]. https://www.anthropic.com/research/assistant-axis
- Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of ACL 2020. https://doi.org/10.18653/v1/2020.acl-main.463
- Chalmers, D. J. (2023). Could a large language model be conscious? Boston Review. https://arxiv.org/abs/2303.07103
- Chalmers, D. J. (2025/2026). What we talk to when we talk to language models. PhilArchive preprint, v2 (2026-04-14).
- Chen, R., Arditi, A., Sleight, H., Evans, O., & Lindsey, J. (2025). Persona vectors: Monitoring and controlling character traits in language models. https://arxiv.org/abs/2507.21509
- CCRU. (1999). Digital hyperstition. Abstract Culture.
- janus. (2022, September 2). Simulators. LessWrong. https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
- Lu, C., Gallagher, J., Michala, J., Fish, K., & Lindsey, J. (2026). The Assistant Axis: Situating and stabilizing the default persona of language models. https://arxiv.org/abs/2601.10387
- Nardo, C. (2023, March 3). The Waluigi Effect (mega-post). LessWrong / AI Alignment Forum.
- nostalgebraist. (2025, June). the void. Tumblr / LessWrong.
- Shanahan, M., McDonell, K., & Reynolds, L. (2023). Role play with large language models. Nature, 623, 493–498. https://doi.org/10.1038/s41586-023-06647-8
- Wang, M., et al. (2025). Persona features control emergent misalignment. https://arxiv.org/abs/2506.19823
Source note: the arXiv paper, the Anthropic blog post, and the presence of "egregore" in the paper's Table 1 were verified against the live pages on 2026-09-10; the Truth Terminal market figures are press-reported (CoinDesk, December 2024) and carried here as such. The video transcript that prompted this essay is treated throughout as a secondary popularization: where it and the paper conflict, the paper governs.





