Ariadne Research Thread
ActiveResearch focus
We study the control and monitorability of LM-based agents, specifically focusing on low-probability but high-stakes scenarios in which models employ steganographic or encoded reasoning (chains of thought). This includes such threat models as code, research, or decision sabotage inside AI labs (alignment faking, scheming behavior).
Open to collaboration
Yes
Looking for
volunteers and independent contributors
Contact
Join the Damaqu community to see this team's full profile.