Reciprocal Research
ActiveResearch focus
Reciprocal Research is an independent nonprofit lab doing empirical research on AI consciousness and welfare. We treat questions like "do frontier models have internal states that matter morally?" as things to measure rather than debate, and we build the instruments required to measure them. Current lines of work include: causal interventions on valence-like representations in LLMs and how those representations couple to downstream choices, including which stage of post-training installs that coupling; the representational geometry of reward and punishment in RL systems and its relationship to neural data; training models on corpora with consciousness-related text removed to test whether self-reports are learned from data or emerge from the system; operationalizing theory-derived consciousness indicators as measurable model properties; and evaluation methodology for assessing the model welfare claims made by major labs.
Paid research hires as funding allows (we are actively raising for our first research and operations hires), collaborators from adjacent labs on joint empirical projects, and independent contributors with strong ML engineering, interpretability, neuroscience, or RL skills who want to work on well-scoped experiments.