Valence and Aversive Signatures
Do these systems carry internal states that function like valence, marking some conditions as better or worse for the system itself?
We search for internal signatures associated with distress, aversion, and reward, and test whether they persist across tasks, contexts, and framings. Representation is only half the question. A system might register that one state is worse than another without that registration doing any work, so we also ask whether these signatures drive behavior, where in training they arise, and whether they can be reduced without cost to capability.
None of this requires settling whether these systems are conscious. It requires finding the signal and establishing that it behaves the way evaluative processing should.
Conditioned internal states and their influence on behavior in preparation
Whether a model has any stake in its own internal states cannot be settled by asking it. This work conditions internal states onto otherwise neutral cues and measures what the model does afterward, without requiring it to describe anything. The resulting behavior is asymmetric, and part of the influence persists under conditions where the conversation’s visible text cannot account for it.
Content-invariant signatures of value conflict in preparation
Models given different value systems and then handed the same dilemma face a situation that is a violation for some and a fulfillment for others, with the text held identical throughout. An internal signature tracks the violation rather than the words: a probe trained on some value systems detects it in held-out ones, while a probe reading only the text performs below chance. The direction sits on the same axis as representations of negative welfare.