Back to research teams

Theory of Interpretability

Active

Research focus

Research in the theory of interpretability relies on tracing conceptual connections in hypothesizing, clarifying underlying assumptions, and testing the fit of different theoretical frameworks for explaining model behaviors. Some key problems of interpretability can be translated into problems that have previously appeared in human cognitive sciences, which facilitates making progress on them. The aim of this project is to identify and closely examine such problems, and propose constructive directions for theoretical and empirical work.

Examples of problems to work on include but are not limited to:

What experiments can we do to test whether safety-relevant behaviors (e.g., strategic deception, hidden reasoning, world-modeling, etc.) can be reduced to other, more basic properties? Is there a way to test for "tacit representations"? What are the implications of finding such representations for interpretability in particular and AI safety more generally?

What kinds of explanation should we be looking for when studying the causal structure of models? Are there any patterns in currently available explanations?

What predictions can we make about concepts as natural kinds and phenomena of mono- and poly-semanticity, considering evidence in favor of a strong Platonic representation hypothesis?

Open to collaboration
Maybe
Looking for

volunteers and independent contributors

Join the Damaqu community to see this team's full profile.