Back to papers
arxiv6.0 / 10

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

Cantao Su, Menan Velayuthan, Esther Ploeger, Dong Nguyen, Anna Wegmann

Abstract

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmented: While there exist a number of tools for measuring the lexical diversity of texts, researchers lack standardized tools for quantifying diversity based on embeddings. Embedding-based diversity measures are highly flexible: They work with any embedding model and any data that can be embedded, and are thus applicable to many notions of diversity. With emb-diversity, we provide a comprehensive embedding-based diversity measurement tool, spanning a broad range of measures. We demonstrate its potential for several use cases: measuring the stylistic, semantic, language and speaker diversity of datasets. https://github.com/nlpsoc/emb-diversity/

Research area

ai biasalgorithmic fairnessrobustness
Published
22 Jul 2026
Source
arxiv
Org
Utrecht University
View paper
Sign in to read and join the discussion.