Elliott Thornley
ActiveResearch focus
Using ideas from decision theory to design and train safer artificial agents.
Examples:
- POST-Agents Proposal for training shutdownable agents: https://arxiv.org/pdf/2505.20203
- Neutrality+ reward function for training shutdownable agents.
- Indifference training: https://docs.google.com/document/d/17BsrDL5Qi1aMhmELb9WhfVHx6PjPKBrtH5W7yYhZv8I/edit?usp=sharing
- Risk-averse AIs: https://drive.google.com/file/d/1KC25GrEqQMv_w_oRDF_wihUgVyyT1sBX/view?usp=sharing and https://docs.google.com/document/d/1fhQN0_auZjwac-EGTKydRvO6772HF5I-PLItp3KYyK0/edit?usp=sharing
Open to collaboration
Yes
Looking for
Volunteers and independent contributors to do ML research (and maybe some theoretical research too)
Contact
Join the Damaqu community to see this team's full profile.