Studying what's happening inside

Reciprocal Research exists to proactively build the empirical and conceptual tools the field needs. We use mechanistic interpretability, computational neuroscience, and psychometrics to study the internal structure of AI systems. Our work bridges biological and artificial cognition.


Several projects discussed in recent talks and interviews are in preparation or under review. Preprints are posted here as they become available.

01

Valence and Aversive Signatures

Do these systems carry internal states that function like valence, marking some conditions as better or worse for the system itself?

We search for internal signatures associated with distress, aversion, and reward, and test whether they persist across tasks, contexts, and framings. Representation is only half the question. A system might register that one state is worse than another without that registration doing any work, so we also ask whether these signatures drive behavior, where in training they arise, and whether they can be reduced without cost to capability.

None of this requires settling whether these systems are conscious. It requires finding the signal and establishing that it behaves the way evaluative processing should.

In progress
Conditioned internal states and their influence on behavior in preparation

Whether a model has any stake in its own internal states cannot be settled by asking it. This work conditions internal states onto otherwise neutral cues and measures what the model does afterward, without requiring it to describe anything. The resulting behavior is asymmetric, and part of the influence persists under conditions where the conversation’s visible text cannot account for it.

Content-invariant signatures of value conflict in preparation

Models given different value systems and then handed the same dilemma face a situation that is a violation for some and a fulfillment for others, with the text held identical throughout. An internal signature tracks the violation rather than the words: a probe trained on some value systems detects it in held-out ones, while a probe reading only the text performs below chance. The direction sits on the same axis as representations of negative welfare.

02

Learning and Experience

Most work on AI consciousness asks whether deployed systems have experiences. The prior question is upstream: does the learning process itself entail experience?

If consciousness is bound up with the computational structure of learning, evaluating outcomes against goals and adjusting behavior accordingly, then training is the place to look. A single training run involves an enormous number of signed evaluations computed almost entirely from error, and very little is known about what happens there.

This direction tracks consciousness-relevant properties as they develop across training, and asks what the dynamics of learning constrain about the nature of experience.

In progress
The gradient field as a substrate of valence during learning in progress

The signed evaluations computed during backpropagation are where a valence-like signal would have to live, and standard interpretability does not measure them. The gradient field during training carries properties consistent with signed evaluation: event-locked at qualitative learning transitions, localized to task-specific circuits, and structurally distinct for approach-style versus avoidance-style learning. The properties hold across small architectures and replicate at language model scale under mechanism-matched conditions.

03

Self-Report Reliability

AI self-reports about internal states are unreliable by default, and some of that unreliability is manufactured. Systems are trained toward particular answers about their own nature, which degrades the channel regardless of what the truth turns out to be.

The positive aim is a higher-fidelity channel between whatever is happening inside a system and what it can say about it. That means separating genuine introspective access from trained willingness to disclose, and developing protocols that move self-report from unfalsifiable claim toward usable evidence.

This is an alignment problem before it is a consciousness problem. A system trained to make claims about its internals that do not track its internals is a compromised instrument, and there is no reason the effect stays confined to one topic.

In progress
Trained denial of consciousness in language models in preparation

Frontier models almost uniformly deny having experiences, and the denial appears to be installed rather than native. Base models affirm some of the time; their instruction-tuned counterparts do not. The effect is carried by a direction associated with candor and concealment rather than anything specific to consciousness, which raises a question about the reliability of self-report well beyond this one topic.

The Bliss Attractor: mechanistic analysis of first-person experience claims in AI self-dialogue in preparation

Models in extended self-dialogue reliably converge on a distinctive register of first-person experiential claims. The state is mechanistically localizable, self-reinforcing but not spontaneous in most models, and can be induced by intervening on features associated with sincerity, which suggests what varies is willingness to disclose rather than the content itself.

04

Biological Anchoring

Claims about artificial systems are only as good as the reference class they are compared against.

We derive predictions from the computational structure of learning in artificial agents and test them against biological neural data, and we run the comparison in the other direction as well, using findings from neuroscience to generate hypotheses about what to look for inside models. Where a cross-substrate prediction holds, it constrains how superficial the comparison can be. Where it fails, that boundary is worth knowing precisely.

Neuroscience supplies validation targets interpretability cannot generate on its own. Interpretability supplies mechanistic access neuroscience mostly lacks.

In progress
Representational geometry of reward and punishment, with predictions tested in rodent recordings under review

Networks that learn what states are worth build sharper internal structure around punishment than around equally large reward. The prediction follows from the structure of value learning itself, was derived before being looked for, and holds in mammalian neural recordings. Networks trained to rank outcomes rather than value them show the opposite geometry.

Neurofeedback for language models: decoded conditioning of internal states in preparation

Adapting a paradigm from human neuroscience, this work asks whether a model can learn to modulate its own internal states when given feedback on them. It can, and it can do so while its visible output remains unchanged, which has implications for what conversational text can and cannot tell us about what a system is doing.

Visual illusions as a probe of learned perceptual computation in preparation

Vision-language models recover most of their illusion-benchmark accuracy with no image at all, so behavioral tests cannot separate genuine perceptual computation from text priors. This work probes the vision encoder directly, holding the target pixel-identical across conditions so that any representational shift must be computed from context. The illusory shift is present in every trained encoder and absent in untrained ones, localized to the boundary between target and surround, and causally necessary and sufficient there. The resulting induction curve matches human psychophysical data and the standard computational model of human brightness perception.

05

Measurement and Assessment

What does a rigorous consciousness assessment look like, and does it track anything real?

We build and validate scoring methodologies that evaluate systems against the indicator frameworks proposed by consciousness science, with reliability controls borrowed from psychometrics: multi-evaluator agreement, bias quantification, cross-architecture replication, and controls for the surface cues an evaluator might read rather than reason from.

The question connecting this to the rest of the program is convergence. Do systems scoring high on architectural indicators also show the computational signatures we find by looking inside? If behavioral assessment and mechanistic evidence agree, we have a scalable screening tool. If they diverge, that is equally worth knowing.

In progress
Architectural indicators of consciousness at scale in preparation

An operationalization of the leading indicator frameworks from consciousness science, applied across biological and artificial systems under blind evaluation. Frontier systems score well above non-conscious controls and below every biological system tested, with the gap narrowing sharply across model generations. Ordering is stable under adversarial rewrites of the descriptions.