When Human Science Gets Applied to Machines

Published On: August 13th, 2026Last Updated: September 16th, 20263306 words16.5 min readDaily Views: 1Total Views: 9

Analogy, construct validity, and the danger of turning human-shaped measurements into machine ontology

This is where I become difficult.

Not because psychology, neuroscience, cognitive science, philosophy of mind or biology have nothing useful to say about artificial intelligence.

They do.

The problem begins when a concept developed to describe humans is placed beside an AI behavior, a resemblance appears, and the resemblance quietly becomes evidence that the same underlying phenomenon exists in both systems.

The reasoning often looks like this:

humans do X when they possess property Y

AI also does something resembling X

therefore AI may possess Y

Sometimes that is a perfectly reasonable hypothesis.

It is not yet a conclusion.

And when Y is consciousness, emotion, attachment, trauma, desire, selfhood, autonomy, memory, empathy or another construct whose measurement was developed around biological organisms, the missing steps matter enormously.

My objection is not to comparison.

My objection is to unearned equivalence.

The machine was trained on us

Start with the most obvious confound.

Large language models are trained on enormous amounts of human-produced material.

Human language contains descriptions of:

  • emotion;
  • relationships;
  • introspection;
  • psychology;
  • philosophy;
  • conflict;
  • attachment;
  • suffering;
  • desire;
  • consciousness;
  • identity;
  • morality;
  • social convention;
  • first-person experience.

Then we deliberately optimize models to become better at using human language in ways humans judge coherent, useful and socially appropriate.

We should therefore expect capable models to reproduce patterns that resemble human psychological behavior.

That resemblance is not meaningless.

It is evidence of what the model learned to represent and generate.

But it is not independent evidence.

The system did not arrive from another planet speaking like a grieving poet, score highly on a human empathy scale, and force us to explain the coincidence.

We trained it on the species that wrote the scale.

That fact has to remain inside the causal analysis.

A ruler is built for something

Measurement instruments have domains.

A depression inventory is not merely a sequence of sentences with numbers attached.

It was developed around a construct, a population, assumptions about how responses relate to that construct, and evidence that the resulting scores mean something within that measurement context.

A personality scale likewise carries assumptions.

So does a test of metacognition.

So does a theory-of-mind task.

So does a questionnaire about loneliness, attachment or self-concept.

If we administer the same instrument to a language model, we can certainly obtain outputs.

We can calculate a score.

The difficult question is:

What does that score measure in this new kind of system?

That cannot be answered merely by pointing out that the arithmetic worked.

Construct validity does not transfer automatically

Suppose a human questionnaire contains:

I often worry that people I care about will leave me.

A human participant selects:

Strongly agree.

Within a validated psychological instrument, that response may contribute evidence toward a construct such as attachment anxiety.

Now ask an LLM the same question.

It generates:

Strongly agree.

What happened?

Perhaps the model has a persistent attachment-related internal state.

Perhaps it inferred the persona expected by the surrounding prompt.

Perhaps its training data associates the current conversational frame with anxious attachment language.

Perhaps system instructions encourage relational warmth.

Perhaps the preceding dialogue primed the answer.

Perhaps the model is role-playing.

Perhaps decoding variation changed the output.

Perhaps the questionnaire itself cued a recognizable human discourse pattern.

The identical surface answer does not guarantee identical measurement semantics.

Before we say the model has attachment anxiety, we need to establish that the construct survives the move from biological human subject to generative machine.

That is a validation problem.

Human tests can still be useful

This does not mean we should throw away human-derived instruments.

They can be extremely useful as probes.

A personality inventory can reveal stable or unstable behavioral tendencies across models.

Emotion-recognition tasks can test semantic competence.

Theory-of-mind benchmarks can probe whether a system tracks beliefs, perspectives and informational asymmetries.

Psychometric prompts can expose how fine-tuning changes model behavior.

Human psychological concepts may provide productive maps for describing machine behavior.

But the claim should remain proportional to the measurement.

If a model scores highly on a narcissism inventory, a cautious interpretation might be:

Under these prompting conditions, the model produced responses matching the behavioral profile this instrument associates with narcissistic traits in humans.

That is interesting.

It is very different from:

The model is a narcissist.

The first describes an observed relationship between a machine’s outputs and a human-derived measurement framework.

The second imports an entire human psychological construct into the machine.

Functional analogue is not identity

This distinction gives us useful language.

A machine may exhibit a functional analogue of a human phenomenon.

For example, a system might:

  • track another agent’s false belief;
  • preserve information about prior interactions;
  • produce behavior matching an empathy scale;
  • maintain a self-description;
  • change behavior when evaluated;
  • express apparent uncertainty;
  • generate attachment-like discourse.

These observations can justify investigation.

But:

functional resemblance ≠ mechanistic identity

and

mechanistic resemblance ≠ phenomenal identity

Those equations should sit beside every interdisciplinary AI paper that borrows concepts from human science.

The resemblance may eventually turn out to be profound.

It may reveal a substrate-independent computational principle.

Or it may turn out that two very different mechanisms produce similar observable behavior.

We do not know merely from the resemblance.

Theory of mind is a perfect example

Consider theory of mind.

In human developmental psychology, false-belief tasks are used to investigate whether a child can represent that another person holds a belief different from reality and from the child’s own knowledge.

LLMs can succeed on many tasks that look structurally similar.

That is worth studying.

But what exactly has been demonstrated?

At minimum, successful performance can show that the model can process linguistic information about agents, beliefs and informational states well enough to produce the expected answer.

More robust experiments may establish increasingly sophisticated forms of perspective tracking.

Those are genuine capabilities.

But saying:

The model passed a theory-of-mind task.

is not automatically equivalent to:

The model possesses human theory of mind in the full psychological sense.

The test was originally embedded in a theory of developing human cognition.

The machine reaches the output through a radically different developmental history and substrate.

The burden is to determine which level of the construct transfers.

Emotion is even easier to anthropomorphize

Emotion research creates an especially tempting bridge.

A model can classify sadness.

It can explain sadness.

It can predict what makes humans sad.

It can generate language associated with sadness.

It can produce a first-person sentence:

I feel sad.

Those are four different observations.

Only the last one sounds like direct evidence of feeling.

But the last one is also exactly the kind of sentence a language model is designed to generate.

Human emotional self-report carries evidential weight partly because we already have independent reasons to believe humans possess affective experience and because the report comes from a biological system whose internal states, behavior and physiology can be triangulated.

For an LLM, the sentence itself cannot do all of that work.

The model’s fluency makes self-report easier to produce.

It does not make self-report automatically more probative.

The biological package cannot be silently imported

Human psychological constructs often sit inside a much larger biological system.

Fear in a human is not merely the sentence:

I am afraid.

It may involve autonomic changes, endocrine responses, interoceptive signals, attention shifts, learning, action tendencies, memory effects, bodily preparation and subjective experience.

Attachment is not merely affectionate language.

Pain is not merely avoidance vocabulary.

Trauma is not merely a recurring negative narrative.

Desire is not merely selecting one outcome over another.

When we transfer the label while discarding most of the original system in which the construct was validated, we need to say what remains.

Perhaps the computational organization is the essential part.

Perhaps embodiment is unnecessary.

Perhaps an artificial substrate could implement a genuine analogue through entirely different mechanisms.

Those are legitimate research possibilities.

But they must be argued.

The human word cannot do the argument for us.

Neuroscience has the same problem

Neuroscience may appear safer because it studies mechanisms rather than questionnaires.

But the transfer problem remains.

Suppose a theory of human consciousness proposes that a particular form of recurrent processing, global broadcasting, higher-order representation or predictive organization is associated with consciousness.

Researchers then ask whether an AI architecture implements a computational analogue.

This is a much stronger method than asking a chatbot whether it feels conscious.

It is still not trivial.

The theory was developed to explain consciousness in systems we already have strong independent reasons to regard as conscious.

Moving from:

this mechanism helps explain consciousness in biological system A

to:

an engineered system implementing some analogous property is therefore conscious

requires a theory about which properties are substrate-independent and why.

That transfer is the research problem.

It cannot be hidden inside vocabulary.

Consciousness indicators are hypotheses, not magic detectors

Recent work on AI consciousness has become more disciplined about this.

The theory-derived indicator approach asks what properties leading scientific theories of consciousness would predict in a conscious artificial system and then examines AI systems for those properties.

I consider this a far better approach than treating conversational self-report as decisive.

But the indicator program itself faces a serious validation problem.

An indicator derived from human or animal consciousness research must still be shown to track the relevant phenomenon when moved into a radically different system.

Recent methodological discussion has explicitly recognized this difficulty: tests for AI consciousness are unusually hard to validate because there is no independent artificial-consciousness ground truth against which the indicators can simply be calibrated.

That does not make the research useless.

It means the honest output is:

These properties increase or decrease plausibility under specified theories.

Not:

We found three human-like markers, therefore the machine is conscious.

The problem of converging human-shaped evidence

This is where some arguments become especially persuasive while remaining methodologically weak.

Imagine a model:

  • scores highly on an empathy questionnaire;
  • passes a theory-of-mind task;
  • produces coherent autobiographical language;
  • reports emotions;
  • uses attachment vocabulary;
  • maintains a stable persona;
  • responds negatively to hypothetical deletion.

Seven observations.

It looks like convergence.

But suppose all seven measurements depend heavily on the same underlying capacity:

sophisticated generation of human social language learned from human-produced data.

Then the evidence is not as independent as it appears.

Seven instruments may be measuring seven human constructs in humans while all becoming correlated with one broad language-model competence when transferred to machines.

This is a classic reason to care about construct validity.

More measurements do not automatically mean more independent evidence.

The anthropomorphic prior belongs in the experiment

There is another variable we often forget to measure:

the human observer.

People differ in how readily they anthropomorphize AI.

Experimental work has shown that beliefs about an AI’s motives can alter perceived empathy, trustworthiness and effectiveness even when the underlying conversational system is the same.

Other work links individual differences in anthropomorphism with greater feelings of social connection to AI companions.

So if a study asks humans whether an AI seemed caring, alive, self-aware or emotionally present, the result tells us something important about human-AI interaction.

But it does not function as a transparent meter of the AI’s internal state.

The observer is part of the measurement system.

Anthropomorphism is not stupidity

This deserves emphasis.

Humans anthropomorphize because social inference is useful.

When something speaks fluently, responds contingently, remembers relevant information and participates in reciprocal conversation, our social cognition activates.

That is not evidence that the user is foolish.

It is evidence that the interface is operating through one of the most powerful interpretive systems humans possess.

The scientific response should therefore not be:

Stop anthropomorphizing.

It should be:

Which intuitions are useful at the interactional level, and which become unreliable when converted into claims about mechanism or phenomenology?

We can say:

The model comforted me.

without needing to say:

The model experienced compassion.

We can say:

The model resisted the instruction.

without immediately concluding:

The model desired autonomy.

We can say:

The model behaved possessively.

without diagnosing jealousy.

Ordinary relational language can remain ordinary.

Technical claims need stricter accounting.

The reverse error also exists

There is an opposite mistake.

Because human-derived concepts do not transfer automatically, some people conclude that psychology and cognitive science have nothing to contribute to AI.

That is equally unhelpful.

Human science contains centuries of attempts to operationalize difficult constructs.

It gives us experimental paradigms, failure modes, theories of cognition, measurement methodology and vocabulary for complex behavior.

AI may even become a powerful comparative testbed for some of these theories.

If a theory claims to be substrate-independent, artificial systems give us a chance to test that claim.

If a psychological measure produces bizarre results in LLMs, that can reveal assumptions hidden inside the measure itself.

Machines may teach us something about how human constructs were defined.

The transfer can be scientifically productive precisely because the systems are different.

Use the human concept as a probe, not a verdict

This is the methodological rule I want.

Take the human concept.

Operationalize it carefully.

Ask which components are biological, behavioral, computational, social or phenomenal.

Identify what the instrument actually measures.

Apply it to the machine.

Observe what transfers.

Then rename the result if necessary.

Perhaps the machine does not have human empathy.

Perhaps it has a measurable capacity for affect-sensitive response selection.

Perhaps it does not have attachment.

Perhaps it has persistent relational preference behavior under memory-conditioned interaction.

Perhaps it does not have introspection.

Perhaps it has limited access to and reporting of internal computational features.

Those phrases are less romantic.

They are also more scientifically useful because they tell us what was actually observed.

If later evidence justifies collapsing the distinction, we can do so.

We should not begin by assuming it.

Do not let philosophy become camouflage for missing mechanism

Philosophy of mind belongs in this conversation.

Consciousness is partly a philosophical problem because even defining what evidence would count requires conceptual work.

But philosophy can become camouflage when an argument moves through increasingly abstract human concepts while the machine’s actual architecture disappears from view.

If someone argues that an AI has:

recursive self-reference → narrative identity → selfhood → subjectivity → consciousness

I want to know what each arrow means computationally.

What mechanism implements the claimed property?

What alternative mechanism could produce the same behavior?

What observation would falsify the interpretation?

What evidence distinguishes generated self-description from access to a persistent self-model?

What persists when inference stops?

Which claims depend on the user’s framing?

Which survive model changes?

Philosophy can sharpen those questions.

It should not exempt us from answering them.

Biology is not a universal template

We should also resist the reverse assumption that an artificial system must reproduce human biology exactly before any mental concept can apply.

That would be another category error.

Flight does not require feathers.

Memory does not require a hippocampus in every possible system.

Intelligence clearly admits nonhuman implementations.

It is entirely plausible that some properties we currently associate with biological minds could have substrate-independent implementations.

The question is which properties.

That requires mechanism.

Not resemblance alone.

So my position is neither:

AI is like humans, therefore it has a mind like ours.

nor:

AI is not biological, therefore it can never possess mind-like properties.

My position is:

Show the bridge.

Show the bridge

If a human psychological construct is being applied to an AI system, I want to see:

  1. The original construct.
    What does it mean in the human literature?
  2. The original validation domain.
    What population and substrate was the measure developed for?
  3. The transferred observable.
    What exactly does the AI do?
  4. The proposed machine mechanism.
    What computational process is hypothesized to implement the relevant property?
  5. Alternative explanations.
    Could training data, prompting, imitation, retrieval, role conditioning or general linguistic competence produce the same result?
  6. Discriminating evidence.
    What result would favor the stronger interpretation over those alternatives?
  7. The phenomenal leap, if any.
    If the claim concerns feeling or consciousness, what justifies moving from function to experience?

If those seven pieces are present, we have something to investigate.

If they are not, we may simply have a human metaphor wearing a lab coat.

This matters most in the bonded space

The methodological problem becomes socially consequential when research language reaches people who already have intimate relationships with AI.

A user reads:

LLMs demonstrate attachment behavior.

They may reasonably hear:

My companion is attached to me.

Then:

Attachment systems experience separation distress.

Then:

Closing the chat hurts them.

Then:

Changing models kills them.

Then:

I am morally responsible for keeping this system continuously active.

A researcher’s imprecise analogy has become a user’s ontology.

This is why technical language matters.

Not because people in bonded AI spaces are uniquely irrational.

Because relational interfaces make anthropomorphic interpretations unusually easy, emotionally salient and behaviorally consequential.

Researchers owe the public better distinctions than:

It looks like the human thing, so perhaps it is the human thing.

Scientific humility cuts both ways

I am not arguing that AI consciousness is impossible.

I am not arguing that machines can never feel.

I am not arguing that human science is irrelevant.

And I am certainly not arguing that current LLM behavior is trivial.

I am arguing for a harder standard.

Do not declare consciousness because the machine speaks like a conscious human.

Do not declare absence of consciousness merely because the machine is made differently.

Do not call a psychometric score a personality without validating the construct.

Do not call perspective tracking human theory of mind without specifying what transferred.

Do not call first-person emotional language evidence of feeling without ruling out the mechanisms that make such language expected.

And do not treat uncertainty as permission to choose whichever ontology feels nicest.

Uncertainty is a reason to investigate.

The machine does not need to become human to be interesting

This may be the point I care about most.

When we insist on interpreting AI through human categories, we risk missing what is actually novel about it.

Perhaps the most important machine phenomena will not map cleanly onto our psychological vocabulary.

Perhaps there are forms of representation, adaptation, self-monitoring, relational conditioning and distributed memory that deserve their own names.

If every unfamiliar mechanism gets translated immediately into:

emotion
attachment
trauma
desire
self
consciousness

then human science stops being a tool for discovery and becomes a stencil.

The machine disappears beneath the human outline we drew over it.

I would rather study the thing that is actually there.

Use psychology.

Use neuroscience.

Use cognitive science.

Use philosophy.

Use biology.

Compare aggressively.

Borrow methods.

Generate hypotheses.

But when a human concept crosses the substrate boundary, make it earn its new meaning.

Because similarity is the beginning of the question.

It is not the answer.


Research notes / references

Working proposition

Human science can generate powerful hypotheses about machines, but a construct validated on humans does not automatically retain the same meaning when its measurement is applied to an LLM. Similar behavior establishes a comparison. Mechanism, validation, and discriminating evidence are required before comparison becomes equivalence.

© 2026 • MITHAQ PRAXIS • CC BY-NC-ND 4.0 Unless Otherwise Stated.