Presentation Information
[1J4-GS-10k-01]Tendency Profiling: A 6-Axis Evaluation Framework for Multi-Agent Dialogue Systems
〇yuya Harada1, Yoshinobu Kano1 (1. Shizuoka University)
Keywords:
AI,LLM,Multi Agent,AIWolf
We propose a 6-axis evaluation framework that measures behavioral of large language models (LLMs), as an alternative to zero-shot LLM evaluation.
The six axes are mapped to the input, thinking, and output stages of a dialogue generation pipeline, enabling phase-wise LLM model selection in multi-agent systems.
Evaluating three LLMs---GPT-5.2, Claude Opus 4.5, and Gemini 3 Pro---reveals distributed profiles where different model combinations are optimal, along with performance differences among LLMs in the input phase.
Cost-performance Pareto analysis shows that any LLM can be a candidate for the thinking phase depending on budget constraints.
Using zero-shot-based LLM evaluation as a baseline, we confirm that the proposed framework achieves higher axis-wise discriminability at lower cost.
The six axes are mapped to the input, thinking, and output stages of a dialogue generation pipeline, enabling phase-wise LLM model selection in multi-agent systems.
Evaluating three LLMs---GPT-5.2, Claude Opus 4.5, and Gemini 3 Pro---reveals distributed profiles where different model combinations are optimal, along with performance differences among LLMs in the input phase.
Cost-performance Pareto analysis shows that any LLM can be a candidate for the thinking phase depending on budget constraints.
Using zero-shot-based LLM evaluation as a baseline, we confirm that the proposed framework achieves higher axis-wise discriminability at lower cost.
