Emerging Data Practices: Data Work in the Era of Large Language Models

Data is one of the foundational aspects of making Artificial Intelligence (AI) work as intended. As large language models (LLMs) become the epicenter of AI, it is crucial to understand better how the datasets that maintain such models are created. The emergent nature of LLMs makes it critical to understand the challenges practitioners developing Gen AI technologies face to design alternatives for better responding to Gen AI's ethical issues. In this paper, we provide such understanding by reporting on 25 interviews with practitioners who handle data in three distinct development stages of different LLMs. Our contributions are (1) empirical evidence of how uncertainty, data practices, and reliance mechanisms change across LLMs' development cycle; (2) how the unique qualities of LLMs impact data practices and their implications for the future of Gen AI technologies; and (3) provide three opportunities for HCI researchers interested in supporting practitioners developing Gen AI technologies.

IBM Research, Yorktown Heights, New York, United States

IBM Research, Sao Paulo, Brazil

University of Notre Dame, South Bend, Indiana, United States

Escuela Superior Politécnica del Litoral, Guayaquil, Ecuador

10.1145/3706598.3714069

https://dl.acm.org/doi/10.1145/3706598.3714069

The ACM CHI Conference on Human Factors in Computing Systems (https://chi2025.acm.org/)

Annex Hall F206

7 件の発表

開始日時2025-04-30 20:10:00

終了日時2025-04-30 21:40:00

読み込み中…