Presentation Information

[4K1-GS-6a-03]Performance Evaluation of LLM Model Routing for Japanese Tasks

〇Tatsuki Ebihara1, Tomoyuki Yamamoto1, Shigeto Yoshida (1. Sharp Corporation)

Keywords:

Large Language Model,Model Routing

LLM model routing, which dynamically selects among multiple LLMs with different performance and cost profiles, has emerged as a promising approach to improving quality–cost trade-offs. However, empirical evaluations of model routing for Japanese-language tasks remain limited. This study evaluates LLM model routing on Japanese open-ended text generation tasks. We consider a binary routing setting between a stronger and a weaker model, and implement a kNN-based router using vector retrieval. We compare two labeling strategies: pairwise comparison labels and weak-model proxy labels based on the response quality of the weak model. Experiments on multiple Japanese benchmarks show that even simple proxy labels can enable practical routing under certain conditions, whereas pairwise comparison labels fail to provide consistent improvements. The results further reveal that routing effectiveness strongly depends on benchmark characteristics and performance gaps between models, and highlight the difficulty of automatically generating stable pairwise labels for open-ended Japanese tasks.