Back to papers
arxiv6.0 / 10

On Improving Faithfulness of Podcasts from Documents

Soumya Dutta, Tejas Indulal Dhamecha, Pannaga Shivaswamy

Abstract

Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources. While these systems produce fluent and engaging narratives, they often introduce ungrounded information. In this work, we present the first systematic study of faithfulness in document-grounded podcast generation, where grounding must be maintained across conversational turns in long-form, multi-speaker transcripts. We construct a dataset of over 1500 documents spanning five domains and generate podcast transcripts using multiple LLMs. We introduce a turn-level LLM-as-a-judge framework for evaluating whether conversational turns are supported by the source document, and validate its reliability through human studies. Our analysis shows that even state-of-the-art models, including GPT-4o, frequently generate ungrounded content. To mitigate this issue, we propose catch-n-repair, a model-agnostic framework that detects and rewrites unfaithful conversational turns while preserving conversational flow. Experiments demonstrate consistent improvements in faithfulness across both in-domain and out-of-domain settings.

Research area

ai reliabilityevaluationhallucinations
Published
24 Jul 2026
Source
arxiv
Org
Indian Institute of Science
View paper
Sign in to read and join the discussion.