BDE Research for Scalable Oversight
ActiveResearch focus
Building synthetic benchmarks enabling controlled experiments on how argument types and evidence structures affect judge accuracy in debate-based AI oversight
Open to collaboration
Maybe
Looking for
Happy to speak with anyone who feels like they may be able to make use of what we are building
Contact
Join the Damaqu community to see this team's full profile.