Volume 68

Evaluating Human-Computer Interaction Capabilities of Edge LLMs in Home Energy Management: An LLM-as-a-Judge Approach Qinghu Tang, Yu Zuo, Boqun Kou, Xiao Cai, Yi Wang

https://doi.org/10.46855/energy-proceedings-12547

Abstract

Integrating Large Language Models (LLMs) into home energy management systems promises personalized, autonomous energy scheduling. However, reliance on cloud-based LLMs introduces severe privacy vulnerabilities regarding household data and critical latency bottlenecks for control loops. While deploying edge LLMs offers a localized solution, their human-computer interaction (HCI) capabilities in continuous, dynamic physical environments remain insufficiently evaluated. This work proposes a novel evaluation framework bridging this gap. We develop a context-aware scenario generator to create physically grounded HCI testbenches, and introduce an event-driven “LLM-as-a-Judge” sandbox that merges subjective semantic scoring with objective, solver-derived metrics. Case studies evaluate various sub-12B edge models, quantifying their intent comprehension and physical modeling boundaries. This research provides essential insights into the capability limits of edge intelligence, laying a foundation for their reliable deployment in privacy-sensitive energy systems.

Keywords home energy management system, large language model, edge intelligence, human-computer interaction, evaluation framework, LLM-as-a-Judge approach

Copyright ©
Energy Proceedings