Center on Long-Term Risk - Empirical team
ActiveResearch focus
Preventing malevolent AI
Join the Damaqu community to see this team's full profile.
Papers on Damaqu (3)
- Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoorsarxiv· 29 Jun 2026
- Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv· 5 Oct 2025
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv· 24 Feb 2025