Research Teams
AI safety and security research teams.
AI safety and security research teams.
Research Teams
Fangcun Leap builds defense systems that protect AI models, LLMs, and autonomous agents from adversarial threats. Our team also participate writing OpenClaws users' safety guidance and regulations in China.
Architecting AI Studio focuses on adversarial red-teaming and alignment stress-testing to map the structural limitations and foundational failure modes of frontier large language models. The lab applies spatial systems thinking to treat LLMs as high-dimensional informational architecture, utilising defensive red-teaming and applied capability testing for deep, entity-level verification.
Parallax is a non-profit research lab developing a second, deeper view of model behaviour. Our mission is to make white-box auditing of agents’ internal beliefs, goals, and plans a core layer of frontier AI evaluations by developing methods and infrastructure that support this depth of analysis at scale. By combining behavioural evidence with evidence drawn from a model’s internal computations, we want to diagnose alignment failures and distinguish the underlying mechanisms that produce them.
To protect human relevance in the age of advanced AI, policy-makers must identify and work to protect key human technical capacities.
Preventing malevolent AI
Reciprocal Research is an independent nonprofit lab doing empirical research on AI consciousness and welfare. We treat questions like "do frontier models have internal states that matter morally?" as things to measure rather than debate, and we build the instruments required to measure them. Current lines of work include: causal interventions on valence-like representations in LLMs and how those representations couple to downstream choices, including which stage of post-training installs that coupling; the representational geometry of reward and punishment in RL systems and its relationship to neural data; training models on corpora with consciousness-related text removed to test whether self-reports are learned from data or emerge from the system; operationalizing theory-derived consciousness indicators as measurable model properties; and evaluation methodology for assessing the model welfare claims made by major labs.
Scaling alignment via multi-agent research.
AIXI Labs models AI risk factors and safety mitigations in terms of AIXI variants, and develops the means to translate them to real AI agents. AIXI is the leading mathematical model of artificial superintelligence—the maximum theoretical limit of AI capabilities. This enables rigorous testing of both the risk factors and the safety mitigations.
We primarily focus on the specification of human-aligned reward functions for reinforcement learning (RL). This effort involves studying how human RL experts design reward functions, how to interpret preference annotations, and theory focused on reward functions.
Deep-rooted AI alignment via safety pretraining; studying AI behavior, esp. w.r.t. (mis)alignment; multi-agent safety; mechanistic interpretability
Trying to understand how different personas might shape the model's underlying worldview and to mechanistically understand what is occurring within a model when it adopts a particular persona.
Our current work focuses on how neural networks learn and leverage hierarchical structure in real-world data, grounded in several hypotheses: 1) Learned features are organized according to a hierarchical, scale-dependent notion of relevance. 2) Faithful interpretability tools leverage this hierarchy. 3) A renormalization-like framework can place principled, probabilistic bounds on cross-scale mechanistic behavior, potentially enabling worst-case guarantees.
Our goal is to understand the complex ways in which AI reflects as well as impacts human belief systems. We seek to design AI systems that function across and recognize cultural barriers. We also study how AI influences and reshapes human thought.
My team focuses on the rigorous causal validation and system-level accountability of AI deployed in high-stakes socio-technical environments. Moving beyond post-hoc interpretability, I develop geometrically-aware and causally-robust frameworks to ensure that AI decisions are fundamentally auditable, resilient, and verifiable against infrastructure-level failures. A central thrust is designing platform-agnostic validation schemas for agentic AI, detecting systemic vulnerabilities before they propagate, directly informed by my ongoing policy work at the Mila Quebec AI Institute and the European Commission's AI Office. I bridge advanced geometric representation learning, causal inference, learning theory with translational impact, ensuring that AI systems in healthcare, finance, and governance are safe, sovereign, and institutionally trustworthy.
Our goal is to characterize the human moral mind in computational terms, using the tools of cognitive science. We simultaneously draw on those insights to create AI systems that are safe, transparently aligned to human values, and support human flourishing.
Individual Researchers
Mycelium builds the technical foundation for AI to consider the welfare of all sentient beings - humans, nonhuman animals, and digital minds. Concretely, we work on technical research & engineering - developing benchmarks, evals, data pipelines and other research experiments to build an evidence base for frontier labs to take seriously the moral patienthood of nonhuman welfare.
We study how AI systems form, maintain, and revise interpretations over time, especially in ambiguous or multi-agent settings. Our work focuses on misunderstanding, overconfidence, premature convergence, perspective shifts, and the mechanisms that help models remain open to corrective evidence. More broadly, we are interested in the dynamics of meaning: how competing interpretations emerge, how prior reasoning trajectories shape future beliefs, and how AI systems can detect when they may be confidently wrong and recover more effectively. The goal is to develop AI systems that are more epistemically flexible, corrigible, and robust under uncertainty.
Building agentic tools and infra to scale AI safety research
AI governance, AI incidents, AI laws, AI evaluation, AI ecosystem monitoring.
I research AI Alignment techniques with a low alignment tax, to make it more likely they will be adopted in practice. Ideally, I am looking for techniques that improve capabilities as a side effect of increased alignment.