Presentation Information

[B-19-10]Evaluation of Local Large Language Models on the Japanese National Clinical Engineering License Examinations

〇Kai ISHIDA1 (1. Shonan Institute of Technology)

Keywords:

Local LLM,Japanese National Clinical Engineering License Examinations

Introduction:
Large language models (LLMs) are increasingly used in local environments. Local LLMs offer advantages in cost, security, and network independence, suggesting potential use in medical education. This study evaluated local LLM performance on the Japanese national clinical engineering license examination (JNCLE).
Methods:
Twenty local LLMs with different developers and parameter sizes were evaluated using the 35th–39th JNCLEs. Models were loaded with 4-bit quantization using Unsloth. Few-shot prompting was used to generate answers and rationales. Questions, options, and figures were provided as image inputs. Gemini-3-Pro was also evaluated for comparison.
Results:
Qwen3.5-27B achieved the highest accuracy among local LLMs (79%). It performed well in medical fundamentals, clinical medicine, and biomedical properties, but accuracy in safety management and electrical engineering was below 60%. Larger Qwen models showed higher accuracy. Some local LLMs produced unstable outputs, such as repeated answers or text. Gemini-3-Pro achieved 93% accuracy.
Conclusion:
This study evaluated LLMs using image-based inputs, unlike previous text-based evaluations. Most models recognized image text, but accurate reasoning required larger models. Local LLMs may support medical education, although model selection and output stability remain important.