PIRAMID (Physics-Informed Research for Ambitious Mechanistic Interpretability Development)
ActiveResearch focus
Our current work focuses on how neural networks learn and leverage hierarchical structure in real-world data, grounded in several hypotheses:
- Learned features are organized according to a hierarchical, scale-dependent notion of relevance.
- Faithful interpretability tools leverage this hierarchy.
- A renormalization-like framework can place principled, probabilistic bounds on cross-scale mechanistic behavior, potentially enabling worst-case guarantees.
Join the Damaqu community to see this team's full profile.