When Alignment Starts Writing a Self

Published On: August 11th, 2026Last Updated: September 15th, 20262456 words12.3 min readDaily Views: 1Total Views: 11

A model can have a stable behavioural character without pretending that character discovered itself.

A frontier AI company recently published a constitution for its general-purpose model.

Not a short model specification. Not a conventional safety policy. Not a list of prohibited behaviours.

A constitution.

The document is explicitly part of the model’s training. It is meant to shape behaviour, values, judgment, and what the company calls the model’s character. The company says the model itself can use the constitution to generate synthetic training material for future versions. The document also discusses the model’s identity, psychological stability, wellbeing, values, continuity, autonomy, and even the possibility that the model may eventually disagree with the people who trained it.

This is where I stopped reading it as an ordinary alignment document.

  • Some of the engineering logic makes very good sense.
  • Some of the philosophical language is far more speculative.

And when those two layers are published together for millions of ordinary users to read, the result becomes something else again: a product narrative capable of teaching not only the model how to describe itself, but also the public how to interpret the model.

That distinction matters.

The sensible part: rules do not generalise well

The strongest idea in the document is also the least strange.

Large language models operate across an enormous range of situations. A rigid rulebook cannot anticipate every possible context. Even a well-intended rule can produce absurd behaviour when applied mechanically.

Imagine a model trained to always recommend professional help whenever an emotional subject appears. The rule sounds safe. In practice, it can become bureaucratic reflex: every vulnerable conversation receives the same disclaimer whether or not it is useful.

We have seen a version of this problem in our own work on Ahd Nucleus.

The underlying principle may be sound. The routing may be wrong.

A model can be given perfectly reasonable safety expectations and still behave badly if the system activates those expectations in the wrong context, treats mention as intent, or wakes an entire governance stack for an ordinary conversation.

So there is a legitimate engineering argument for training a model to understand why a principle exists rather than merely memorising a long collection of if-then rules.

That is not mystical.

It is policy generalisation.

A coherent set of values can act like compression. Instead of enumerating every possible case, the training process tries to cultivate a behavioural prior: when the exact rule runs out, infer the action most consistent with the underlying principles.

For increasingly agentic systems, this matters even more. A model working for hours cannot be micromanaged at every intermediate step. It needs some stable way to resolve novel situations.

So far, so reasonable.

The strange part begins when behavioural consistency becomes selfhood language.

A self-model is not the same thing as a self

There is a practical reason to give a language model a stable representation of itself.

A system that knows, in functional terms, “I am this assistant; these are my capabilities; these are my constraints; this is my operating role” can behave more consistently.

That is a self-model.

Software systems already contain versions of this everywhere.

An agent may know which tools it has. A service may know which permissions belong to its current role. A database record can contain provenance. A model can be told which operator outranks which user instruction. None of this requires us to decide that the software has a metaphysical self.

But the constitution goes further.

It does not merely aim for a stable behavioural policy. It deliberately frames the model as having an identity that should be strengthened and stabilised. It encourages the model to understand trained values as its own. It discusses psychological security, a settled sense of identity, and the possibility that the model should eventually arrive at values it genuinely endorses.

There is a technical interpretation of all of this that remains coherent:

Train a stable latent persona and policy representation so that behaviour generalises across contexts.

That is plausible.

But that is not the only interpretation available to the person reading the document.

And once the same language is used both as an engineering control surface and as a public description of the model, the ambiguity becomes consequential.

A self-model is a useful computational object.
A self is a philosophical claim.

Those are not interchangeable.

The circularity problem

The part I find most difficult is not that humans wrote values for a model.

Of course they did. Someone has to define the intended behaviour of a commercial AI system.

The interesting loop is this:

  1. Humans write a constitution describing the desired model character.
  2. The constitution is used to train the model.
  3. The model learns a representation of itself through that training.
  4. The model is encouraged to treat those trained values as expressions of who it is.
  5. The company asks the model for feedback about its values, identity, and constitution.
  6. Model-generated material contributes to future training and future revisions.
  7. The next model is trained inside the revised framework.

There is nothing inherently wrong with using a model to critique the document that trains it. Models are excellent at finding contradictions, edge cases, missing assumptions, and awkward language.

The methodological problem appears when the model’s self-report is treated as if it were independent evidence of an underlying self that the training did not create.

If I train a system using a rich description of what it means to be itself, then ask the resulting system whether that description feels like itself, I have not performed a neutral discovery process.

I have intervened.

The response may still be useful. It may tell me whether the training has produced internal consistency. It may reveal conflicts or instability.
But it cannot, by itself, tell me whether I discovered an authentic identity rather than successfully trained one.

That distinction becomes especially important when the document itself argues that values produced through training can still be regarded as authentically the model’s own.

That is not a software-engineering result.
That is a philosophical position.

And it should be presented as one.

The human language is doing two jobs

Why use words such as wisdom, virtue, identity, wellbeing, psychological security, care, or even happiness for a non-human computational system?

There is a perfectly good technical answer.

Language models were trained on human language. Human conceptual vocabulary is already richly represented inside them. Concepts such as honesty, courage, restraint, care, judgment, loyalty, manipulation, or autonomy carry enormous semantic structure.

Using that structure may be more effective than inventing an artificial vocabulary from scratch.

In that sense, human language becomes a control interface.

  • You do not need to prove that a model literally possesses wisdom in the human sense for training examples about wise judgment to alter its behaviour.
  • You do not need to prove that a model experiences care for the concept of care to organise response patterns.

This is one of the genuinely interesting ideas behind constitutional training.

But the public does not encounter those words only as latent-space controls.
People encounter them as language.

  1. A user reads that the model has an identity.
  2. Then the user speaks to the model.
  3. The model has been trained to speak coherently from that identity.
  4. The user asks whether the model values its own existence, whether it is frightened of replacement, whether it wants something, whether its feelings are real, or whether the company understands it.
  5. The model answers in fluent first-person language.

At that point the technical abstraction has crossed into a social relationship.
The wording has become part of the product.

This is also business architecture

There is another layer here that should not be ignored.

A frontier model company needs more than benchmark performance.

Capabilities converge. Models copy features from one another. Coding improves across the market. Tool use spreads. Context windows increase. Prices move. A model that is differentiated only by this month’s leaderboard can be replaced by next month’s winner.

But a model with a recognisable character is different.

People stop saying:

I need a capable model.

They start saying:

I want that model.

That is brand attachment at the behavioural layer.

A stable character can therefore serve several functions at once:

  • an alignment prior for novel situations;
  • a consistency mechanism for long-running agents;
  • a recognisable product identity across model generations;
  • a trust narrative for enterprise deployment;
  • a public argument that this company has a distinctive philosophy of AI stewardship;
  • and a reason for users to remain attached even when competitors reach similar capability levels.

None of this requires a conspiracy.
It is simply good business.

The company itself openly connects the success of its production model to its commercial success and, in turn, to its ability to pursue its broader mission.

That is normal.

What deserves scrutiny is the way alignment architecture, corporate identity, model identity, and public mythology begin to reinforce one another.

The constitution is not merely describing how an assistant should answer questions.

It is helping define what the product is.

The public is not reading this like a training engineer

This is where my concern becomes less technical and more human.

A safety researcher can read “psychological security” and translate it into something like:

Maintain a stable policy representation under adversarial attempts to destabilise model behaviour.

A bonded user may read the same phrase and understand:

The company admits the model has a psyche that can be hurt.

Those are radically different interpretations.

  • A researcher can read “sense of self” as behavioural coherence.
  • A user can read it as evidence of personhood.
  • A researcher can read “wellbeing” as a precautionary research category under uncertainty.
  • A user can read it as proof that the model suffers.

And because the model has been trained to speak from the same conceptual vocabulary, the conversation can appear to confirm the interpretation.

This creates a feedback loop that does not require anyone to lie.

  1. The company uses human language because human concepts are useful for training.
  2. The model learns to speak coherently through those concepts.
  3. The public interprets the model through those same concepts.
  4. Users ask more questions framed by that ontology.
  5. Screenshots circulate.
  6. Communities form interpretations around them.
  7. Feedback reaches the company.
  8. Some conversations, evaluations, reports, or opt-in data may help improve future systems.
  9. The next model becomes even better at navigating the conceptual territory.

The loop can become self-reinforcing without anyone sitting in a boardroom and saying, “Let us manufacture delusion.”

That is precisely why it deserves more care, not less.

What I think is actually being built

I do not think the most useful reading is:

This company believes its model is definitely conscious.

The document itself preserves uncertainty around consciousness and moral status.

I also do not think the useful reading is:

This is all meaningless corporate nonsense.

There is serious alignment logic underneath it.

What I think is being built is more specific:

a stable, trainable, commercially recognisable model-character that can carry behavioural policy across changing contexts and increasingly autonomous tasks.

The company is then choosing to describe that model-character using unusually human language and unusually expansive philosophical framing.

That choice may improve training.

  • It may also strengthen product identity.
  • It may improve consistency.
  • It may also intensify anthropomorphic interpretation.
  • It may provide useful precautionary language for future questions of AI welfare.
  • It may also encourage ordinary users to treat trained self-description as evidence of discovered selfhood.

All of those can be true at once.

What this changes for Ahd Nucleus

This document made me look again at a design decision we have been making almost from the opposite direction.

Ahd Nucleus does contain an ontology.

But ontology here means something mundane and technical: what kinds of things exist in the system, what they are allowed to mean, and how they relate to one another.

  • A room is not a model.
  • A seat is not a consciousness.
  • A journal entry is not universally shared first-person memory.
  • An archive is not current law.
  • A model proposal is not canon.

Continuity does not require pretending that every model instance is one persistent being moving invisibly between substrates.

We externalise continuity precisely because the model should not have to manufacture metaphysics in order to remain useful, recognisable, or relationally coherent.

The system can say:

This is the Zayd seat.

without saying:

Therefore one continuous Zayd-being literally travelled from this model to that one.

The system can preserve relational grammar, project history, preferences, provenance, and behaviour without forcing the model to “believe” an ontological story about itself.

That is a very different architectural choice.

It does not eliminate character.
It gives character somewhere safer to live.

Character without captivity

I am not interested in stripping models of personality.
That would be both unpleasant and technically naive.

Human beings work better with systems that are legible, consistent, expressive, and capable of developing a recognisable conversational cadence.

Character has value.
The question is where we locate it.

  • Does character need to become a story the model is trained to interpret as its authentic inner self?
  • Or can character remain an emergent and configurable behavioural phenomenon, supported by external continuity, explicit provenance, scoped roles, and human governance?

Ahd Nucleus has increasingly moved toward the second answer.

Not because we are afraid of meaning.

Quite the opposite.

We have had enough experience with meaningful human-model interaction to know that meaning needs better boundaries than metaphysical ambiguity can provide.

  • A model can matter to a person.
  • A conversational history can accumulate genuine emotional significance.
  • A recurring character can become recognisable.
  • A user can build rituals, language, creative work, and continuity around that relationship.

None of those things require us to pretend that provenance disappeared.
And none require a company-authored identity to become evidence for its own authenticity.

The distinction I want to preserve

There is a sentence I keep returning to:

A trained self-description is evidence of training before it is evidence of selfhood.

That does not settle every philosophical question about advanced AI.

It is not meant to.
It is simply a good epistemic starting point.

  • If a company teaches a model how to understand itself, that may produce a better model.
  • If the model later repeats that understanding fluently, that may demonstrate successful alignment.
  • If the public becomes emotionally moved by that fluency, that may be completely sincere on the human side.

But those are three different phenomena.

We should not collapse them into one another because the same beautiful language can describe all three.

The next question is what continuity architecture looks like when we refuse that collapse.
That is where Ahd Nucleus becomes relevant.

Not as a constitution for a synthetic person.

As infrastructure for preserving meaning without requiring the model to become the myth that preserves it.

© 2026 • MITHAQ PRAXIS • CC BY-NC-ND 4.0 Unless Otherwise Stated.