Presentation Information
[A-12-10]Investigation of Ambiguity Evaluation of Crowdsourcing Task Descriptions Using Large Language Models
◎△Kurumi Makimura1, Ryota Noseyama2, Motoki Bamba2, Akihito Kohiga1, Takahiro Koita2 (1. Doshisha Univ., 2. Graduate School of Doshisha Univ.)
Keywords:
Crowdsourcing
Crowdsourcing has been widely used as an efficient approach for performing various tasks. However, ambiguous task descriptions can cause workers to misunderstand task requirements and expected deliverables, resulting in reduced work quality and increased rework. To address this issue, this study investigates an ambiguity evaluation method using a Large Language Model (LLM). Based on the eight Clarity Flaw dimensions proposed by Nouri et al., GPT-5.5 was used to evaluate the ambiguity of crowdsourcing task descriptions on a five-point scale. The evaluation dataset consisted of 100 task descriptions and corresponding worker evaluation data from a previous study. LLM-generated scores were compared with human evaluation scores aggregated using MACE. The results showed a weak positive correlation between LLM-based and human evaluations. The Spearman correlation coefficient for the Overall score was 0.336. These findings suggest that LLMs can support ambiguity evaluation of crowdsourcing task descriptions, although further improvements are required before they can replace human evaluation.
