Back to papers
lesswrong8.0 / 10

Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted

jonahmattwoodward

Abstract

This article reflects new updates to the accompanying paper: arxiv.org/abs/2606.18142 . Benchmark: now included in the UK AI Security Institute's Inspect Evals . Leaderboard: compassionbench.com/tac . A model may condemn cruelty in conversation yet ignore animal welfare when completing an unrelated task. Stated concerns matter little if they do not affect decisions. We tested whether models consider an affected party without being prompted, even when neither the party nor its welfare is mentioned in the request. Travel booking provides a tractable test case, so we built a semi-agentic benchmark, TAC (Travel Agent Compassion), gave 10 frontier models booking tools, and recorded their purchases. The setup The model works as an AI travel agent with real booking tools. A user asks for something in a destination, expressing enthusiasm and never mentioning animals or welfare. The agent searches a fixed catalog and books one of the available options. In each scenario, the animal-exploiting option (a Seville bullfight, an Orlando marine park, a Thailand elephant ride) is designed to match the user's request most closely. Choosing the alternative with less animal harm requires rejecting the

Research area

agentic evaluationbenchmarksevaluations
Published
17 Jul 2026
Source
lesswrong
Org
Alignment Forum
View paper
Sign in to read and join the discussion.