Back to research teams

Ariadne Research Thread

Active

Research focus

We study the control and monitorability of LM-based agents, specifically focusing on low-probability but high-stakes scenarios in which models employ steganographic or encoded reasoning (chains of thought). This includes such threat models as code, research, or decision sabotage inside AI labs (alignment faking, scheming behavior).

Open to collaboration
Yes
Looking for

volunteers and independent contributors

Join the Damaqu community to see this team's full profile.