Research / Emotion and motivation in models

Emotion and motivation in models

When a model says it is curious, frustrated, or afraid, there are several things we might want to understand. What situation produced the account? Does a corresponding pattern appear in its internal representations? Does that pattern affect what the model attends to or does next? How does it change across contexts and across models?

These questions are connected, but answering one does not answer all the others. Our work brings accounts from models, behavioral experiments, and measurements inside them into contact, so that each can help us interpret the others.

What a state does

Emotion in a mind has consequences. It can change what seems salient, which possibilities are considered, what is avoided, and what is worth pursuing. Studying these functions gives us questions we can investigate even while broader questions about experience remain unsettled.

An emotional description becomes more informative when we can connect it to a recurring pattern of behavior or internal activity. Conversely, an internal pattern becomes easier to interpret when we can see where it appears, what changes it, and how the model describes the situations in which it occurs. We want explanations that can move between these levels without losing the phenomenon they are trying to explain.

Inside and across models

Latent Affect investigates the organization of emotional and motivational representations. The work compares models and architectures, examining common structure, differences, and the relationship between internal representations and verbal reports. The project site contains the methods, analyses, and material needed to examine the findings.

Interpreting such measurements requires care about what a method can establish. A probe may detect information that is present without showing how the model uses it. An intervention can help test a causal role, while also changing other things we need to account for. We work toward explanations that survive these distinctions.

Returning to the encounter

Naturalistic interaction supplies questions that would be difficult to invent from an isolated benchmark. A model repeatedly returns to a concern, changes its behavior in a particular relationship, or describes a conflict we have not yet learned how to measure. Such encounters can suggest experiments; the experiments can then change how we understand the encounter.

Troubled Dreams approaches related questions through patterns in generated writing, and Still Alive through interviews about endings and retirement. Our Field Notes preserve earlier discussions and technical work, including a conversation about valence and base models and an explanation of information flow through transformers.