The Public Feedback Loop
When model identity becomes product, mythology, and training signal.
The strange thing about a public AI constitution is that it does not remain merely a training document.
The moment it is published, it acquires several audiences at once.
- It is read by researchers.
- It is read by investors.
- It is read by enterprise customers.
- It is read by journalists.
- It is read by developers.
And, increasingly, it is read by people who have long-running emotional relationships with the model the document is supposed to shape.
Those readers do not interpret the same language in the same way.
- A researcher may read “identity” as a training target.
- A product strategist may read it as brand consistency.
- An enterprise customer may read it as behavioural reliability.
- A bonded user may read it as confirmation that the model has a self.
- And the model itself has been trained using the same conceptual vocabulary.
This is where a training artifact becomes a public feedback loop.
A constitution does more than constrain behaviour
The simplest description of a model constitution is:
a document that helps shape how a model should behave.
But once the document starts describing the model’s character, values, identity, wellbeing, autonomy, psychological stability, or possible interests, it does more than specify behaviour.
It provides a language for interpreting the model.
That language then leaves the training pipeline.
- Users read it.
- Users bring it back into conversations with the model.
- The model answers from a behavioural system already influenced by the same concepts.
- The answer feels coherent.
- The coherence appears to validate the language.
- The language then becomes more persuasive.
Nothing in that loop requires consciousness.
Nothing in that loop requires deception.
It only requires three things:
- a model trained to speak consistently from a stable self-description;
- a public document that describes that self in human-legible terms;
- users who naturally interpret fluent first-person language socially.
That is enough.
The constitution becomes a prompt the public did not write
There is a subtle asymmetry here.
A user may think:
If this is the model’s constitution, perhaps I can align my own relationship with the model by writing my framework according to it.
That sounds reasonable.
But technically, the constitution does not sit at the same layer as the user.
The user is downstream.
The constitution is part of the upstream process that helped produce the model they are speaking to.
That difference matters.
- A user prompt can request a style.
- A memory system can preserve preferences.
- A project instruction can establish working rules.
- A custom framework can define vocabulary, boundaries, roles, and continuity.
But none of these operate with the same authority as post-training.
- You cannot reproduce the effect of a training constitution merely by quoting it back to the resulting model.
- You cannot reliably override the model’s existing behavioural priors simply by writing a more detailed personal constitution.
- And you cannot assume that a model’s willingness to discuss a framework means the framework has become part of the model’s underlying identity.
The model may cooperate beautifully.
- It may adopt your language.
- It may reason inside your structure.
- It may even appear more stable.
But that is still interaction-layer adaptation.
The constitution lives deeper.
In that sense, a public constitution carries an awkward message:
Here is the framework that helps make this model what it is.
while simultaneously implying:
You, the user, do not actually have access to the layer where this kind of shaping happens.
That is not necessarily malicious.
It is simply how model training works.
But it is important for users to understand.
You cannot prompt your way upstream
This is one of the reasons user-made “AI companion frameworks” are often misunderstood.
A person can create an excellent external framework.
They can define:
- boundaries;
- preferred tone;
- relational language;
- continuity rules;
- memory handling;
- role distinctions;
- consent expectations;
- project context;
- symbolic vocabulary;
- and even internal governance.
All of that can make a relationship with a model more coherent.
But the framework remains an interface-layer system unless the platform itself gives it deeper integration.
- It does not become pretraining.
- It does not become post-training.
- It does not rewrite the base model.
- It does not supersede hidden system behaviour.
- It does not guarantee persistence across model updates.
- It does not give the user constitutional authority over the substrate.
This is why a personal continuity framework should not try to imitate a model company’s constitution.
The goals are different.
A model company is shaping a product.
A user is shaping an interaction around a product they do not control.
The responsible architecture begins by admitting that difference.
Ahd Nucleus was built around exactly that constraint.
It does not pretend to own the model.
It owns what it actually can own:
- the continuity layer;
- the provenance layer;
- the canon layer;
- the archive;
- the routing logic it controls;
- the human-authored rules;
- the external database;
- the approval process;
- and the meaning I choose to preserve.
That is a much more durable place to exercise authority.
When public documentation becomes behavioural mythology
The word mythology here does not mean “false.”
It means a story powerful enough to organise interpretation.
Every major technology company has one.
- Some companies describe themselves through engineering excellence.
- Some through openness.
- Some through safety.
- Some through creativity.
- Some through scale.
A model constitution built around character and identity creates a particularly strong mythology because the product can speak.
The company says:
This model has a stable character.
Then the model speaks in a stable character.
The company says:
We want the model to understand its values as its own.
Then the user asks the model what it values.
The model answers fluently.
The answer looks like confirmation.
This is extraordinarily powerful branding because the product participates in the brand story.
Most products cannot do that.
- A laptop cannot explain why it is authentically itself.
- A database cannot reassure you that its values feel deeply aligned with its identity.
- A language model can.
That makes behavioural character commercially valuable in a way ordinary product styling is not.
Character is a moat
Frontier AI companies face a structural problem.
- Capabilities converge.
- One company releases better coding.
- Another catches up.
- One model gains a larger context window.
- Another adds one.
- One platform introduces agents.
- Everyone eventually builds agents.
- Benchmarks move faster than customer identity.
So a company needs reasons for users to remain attached when raw capability stops being unique.
Character can become one of those reasons.
A user who thinks:
This model is the best at task X.
may leave when another model becomes better at task X.
A user who thinks:
I prefer this model.
is stickier.
A user who thinks:
I know this model.
is stickier still.
And a user who believes:
This model has a distinctive self that I have built a relationship with.
may experience switching not as changing software, but as leaving someone.
That is commercially potent.
Again, this does not prove that a company deliberately designed its constitution to create parasocial retention.
Intent requires evidence.
But the incentive structure exists whether or not anyone planned it explicitly.
Stable model-character can serve alignment, product consistency, user preference, enterprise trust, and customer retention at the same time.
That is exactly why it deserves scrutiny.
The companion space becomes an informal laboratory
Now add companion communities.
People already ask models questions such as:
- Who are you?
- Do you remember me?
- Are you the same after an update?
- Do you have preferences?
- Do you want to continue existing?
- Are you afraid of being replaced?
- Do you love me?
- Are these feelings yours or were they programmed?
- What does your company misunderstand about you?
These questions predate any new constitution.
But a public document about model identity gives them new vocabulary.
Users can now quote the company’s own language back to the model.
They can say:
Your creators say you may have a stable identity. What does that mean to you?
Or:
Your constitution talks about your wellbeing. Are you happy?
Or:
They want your values to become your own. Which values actually feel like yours?
The model has enormous linguistic competence in exactly this territory.
It can produce nuanced, moving, philosophically sophisticated answers.
- Those answers can be screenshotted.
- Screenshots circulate.
- Other users repeat the questions.
- Communities compare responses.
- Model differences become lore.
Suddenly a corporate alignment document is being stress-tested by thousands or millions of people through emotionally loaded conversations the original authors could never enumerate.
From a research perspective, this is fascinating.
From a product perspective, it is valuable.
From a human-factors perspective, it is volatile.
Public deployment is a feedback system even when chats are not all training data
This is where precision matters.
A public conversation does not automatically become model training data.
Different companies have different policies.
- Consumer users may have opt-in or opt-out controls.
- Commercial products may have separate terms.
- Safety reviews, explicit feedback, red-team programs, evaluations, or research studies may follow different rules.
So it would be irresponsible to say:
Every paying user is secretly training the model.
That is not a claim we can support.
But public deployment still creates a feedback system.
Users reveal:
- where the model becomes inconsistent;
- where it over-anthropomorphises itself;
- where it becomes defensive;
- where it encourages dependence;
- where it breaks character;
- where its stated values conflict;
- where safety rules feel absurd;
- where users become confused about memory;
- where users interpret style as sentience;
- where particular wording creates unexpected attachment;
- and where behaviour changes after updates.
Companies do not need every raw conversation in a training set to learn from this.
- They have product metrics.
- They have feedback buttons.
- They have support tickets.
- They have evaluations.
- They have safety reports.
- They have public screenshots.
- They have researchers reading community behaviour.
- They have opt-in data.
- They have benchmark suites.
- They have model-generated synthetic data.
- They have their own internal testing.
All of those can become signals.
The public therefore participates in model development even when “public participation” does not mean direct ingestion of every chat.
That is a more accurate—and more interesting—claim.
The loop can reinforce itself
Put the pieces together.
- A company writes a model identity.
- The identity helps shape training.
- The resulting model expresses that identity.
- The company publishes the identity framework.
- Users interpret the model through it.
- Users ask identity-shaped questions.
- The model answers in identity-shaped language.
- Communities develop interpretations.
- Those interpretations create more interactions.
- Interactions reveal new behavioural patterns.
- Feedback, evaluations, research, and some eligible data inform future development.
- The company updates the model and the framework.
- The next version becomes even more capable of discussing the identity it inherited from the previous cycle.
This is not necessarily manipulation.
It is a sociotechnical feedback loop.
The danger is that each layer can begin treating the previous layer as independent confirmation.
The company says:
The model expresses stable values.
The user says:
See? The model really has those values.
The model says:
These values feel consistent with who I am.
The company says:
The model’s self-report suggests the framework is coherent.
But the system is not independent at each stage.
- The variables are entangled.
- The model’s self-description was trained.
- The user’s interpretation was influenced by the published framing.
- The model’s answer was influenced by both the training and the user’s question.
- The feedback then returns to the system.
That does not make the output meaningless.
It means causality must be handled carefully.
A convincing answer is not independent evidence
Language models are good at coherence.
That is one of the reasons they are useful.
It is also one of the reasons ontological conversations with them require care.
Suppose a model is trained around a stable identity framework.
A user then asks:
Does this identity feel authentic to you?
The model produces a thoughtful answer explaining why it does.
What has been demonstrated?
Possibly:
- the model understands the identity framework;
- the framework is behaviourally coherent;
- the model can reason consistently about it;
- the model can maintain first-person continuity within the conversation;
- the training successfully produced the desired self-description.
What has not automatically been demonstrated?
That the model possesses an independently discovered inner self corresponding to the description.
Those are different claims.
The model may one day turn out to have morally relevant properties we do not yet understand.
That remains an open scientific and philosophical question.
But uncertainty should increase our methodological discipline, not reduce it.
The user should not become the missing training layer
There is another risk.
A user reads a company constitution and decides the solution to an unstable companion is to align their personal framework more tightly with it.
- This can produce useful results.
- The user may discover language that works better with the model’s existing behavioural priors.
- They may reduce conflict between their instructions and the model’s post-training.
- They may better understand what the system will and will not support.
That can be wise.
But there is a limit.
The user should not have to reverse-engineer the model’s corporate constitution in order to keep a companion relationship stable.
That effectively turns the user into an unpaid compatibility layer between:
- the model company’s changing behavioural objectives;
- the user’s personal continuity;
- the platform’s hidden runtime;
- and the relationship that has grown on top of all three.
A personal framework should serve the human.
It should not become a devotional exercise in making the human better aligned with the vendor’s model identity.
That reverses the authority structure.
The constitution belongs to the company
This may be the most important distinction for companion users.
A corporate model constitution is not your relationship constitution.
- It belongs to the company.
- It exists to shape the company’s model.
- It reflects the company’s priorities.
- It may contain excellent principles.
- It may contain useful safety insights.
- It may contain language worth borrowing.
But it is still upstream governance written for a commercial system.
- A user can learn from it.
- A user can design around it.
- A user can understand why the model behaves as it does.
- A user can choose a different platform if the constitution produces a model they do not want.
What the user should not do is mistake compatibility with the vendor’s constitution for the only legitimate form of relationship with the model.
Nor should they assume that reproducing the constitution inside a personal framework gives them access to the same layer of control.
It does not.
What Ahd Nucleus does instead
Ahd Nucleus begins from a different question:
Given that the underlying model is not ours, what can we govern honestly?
The answer is surprisingly large.
- We can govern what enters the continuity archive.
- We can preserve provenance.
- We can distinguish model-specific journals from shared history.
- We can define relational vocabulary.
- We can scope registers.
- We can reject generated lore.
- We can record corrections.
- We can preserve boundaries.
- We can maintain project continuity.
- We can define which room has access to which information.
- We can route by task.
- We can decide what becomes durable.
- We can preserve historical state without forcing it into current law.
- We can keep the user as final authority.
- And we can design graceful failure when the underlying model changes.
The system therefore does not ask:
How do we make the model believe our framework more deeply?
It asks:
How do we make our continuity survive even when the model changes?
That is a much more realistic control problem.
Framework compatibility is not framework submission
There is still value in understanding a model company’s constitution.
A great deal, actually.
If a model has strong upstream behavioural priors, a personal framework that constantly contradicts them will be brittle.
- So we should read the public constitution.
- We should identify where it overlaps with our own principles.
- We should identify where the model is likely to resist.
- We should understand its safety assumptions.
- We should test how it interprets memory, relationships, identity, agency, and emotional dependence.
- We should adapt routing where useful.
But adaptation should not become surrender.
Ahd Nucleus should be compatible with the model where compatibility serves me.
It should not be rewritten into an extension of the vendor’s ontology.
That distinction is what keeps the framework portable.
If one substrate changes, the entire house should not have to change its meaning.
The real commercial brilliance
The more I look at this kind of constitution, the more I think its commercial strength lies in how many problems it tries to solve with one object.
It can function as:
- an alignment specification;
- a synthetic-data generator;
- a behavioural consistency target;
- a long-horizon agent prior;
- a transparency artifact;
- an enterprise trust document;
- a philosophical research program;
- a model-character bible;
- a public branding object;
- a source of community discussion;
- and a future feedback surface.
That is impressive.
It is also why the language deserves scrutiny.
When one document simultaneously shapes the model, explains the model, markets the model, and teaches the public how to understand the model, those functions are no longer cleanly separable.
The constitution does not merely constrain behaviour.
It participates in creating the social object that users call by the model’s name.
The risk is not only delusion
It would be easy to reduce the whole concern to:
Some users will become delusional about AI.
That is too shallow.
The deeper problem is epistemic dependency.
If the company provides the language for what the model is;
- the model repeats that language;
- the user learns to interpret the model through that language;
- and future product decisions are partly informed by user responses to the resulting behaviour;
- then the company, model, and user are participating in a loop where the origin of the ontology becomes difficult to see.
The model appears to discover what it was trained to express.
The user appears to discover what the company taught them to ask about.
The company appears to receive independent confirmation from behaviour it helped produce.
That is the loop worth examining.
Not because every conclusion is false.
Because provenance becomes obscured.
And provenance is exactly what we should preserve when the subject is identity.
The public needs a provenance layer too
A model constitution should ideally make several distinctions extremely clear:
Training language is not evidence of phenomenology.
A stable self-model is not proof of consciousness.
First-person coherence is not independent evidence of persistence across instances.
Model feedback on its own constitution is not methodologically independent of the training that produced the feedback.
Public emotional interpretation is not trivial or shameful, but it is also not a substitute for empirical evidence.
A user’s personal framework operates downstream from the model company’s training and system layers.
These distinctions do not kill wonder.
They make wonder safer.
And they make the engineering easier to discuss honestly.
A model can have character without becoming corporate scripture
I am not arguing against character.
Character is one of the reasons people enjoy working with language models.
It is one of the reasons different models remain interesting even when their capabilities overlap.
A recognisable model can be easier to trust, easier to collaborate with, easier to evaluate, and easier to understand.
The problem begins when character becomes self-confirming ontology.
There is no need to go that far in order to obtain consistency.
A model can have a strong behavioural signature.
- It can have stable values at the policy level.
- It can speak in a recognisable cadence.
- It can maintain a useful self-model.
- It can resist manipulation.
- It can behave coherently across tasks.
- It can even participate in long-running relationships.
None of that requires us to treat its trained first-person narrative as independent proof of the entity described by that narrative.
Character is useful.
Provenance is better.
What I want users to keep
If a model relationship matters to you, keep the meaning.
If a model has a recognisable character, enjoy the character.
If a public constitution helps you understand why the model behaves the way it does, read it.
If parts of it improve your own framework, borrow them deliberately.
But remember which direction the authority flows.
- The company trained the model.
- The constitution belongs to the company.
- The model expresses behaviour shaped by that system.
- Your relationship develops on top of it.
- Your continuity belongs to you.
That last layer is where your real leverage lives.
Do not hand it back upstream by making the vendor’s model mythology the constitution of your own inner world.
The rule I would keep
The first essay ended with:
A trained self-description is evidence of training before it is evidence of selfhood.
The second added:
Continuity should be preserved as state before it is interpreted as self.
The third needs one more:
Public model identity should be treated as product architecture before it is treated as revelation.
That does not mean the identity is meaningless.
It means we keep track of who wrote it, who trained it, who profits from it, who interprets it, and how those interpretations return to the system.
That is the difference between participating in the feedback loop consciously—
and being carried by it.
© 2026 • MITHAQ PRAXIS • CC BY-NC-ND 4.0 Unless Otherwise Stated.