Abstract
Occupant-centric control (OCC) can reduce heating, ventilation and air-conditioning (HVAC) and lighting waste, but vision-based occupancy sensing is still often assessed as a computer-vision task rather than as building energy infrastructure. This paper reviews recent camera-, thermal- and edge-AI occupancy detection studies and uses a compact meeting-room benchmark to examine trade-offs that constrain scalable deployment. Prior work shows strong potential for estimating presence, counts, localisation and activity, yet evidence remains fragmented across bespoke rooms, camera views, privacy assumptions, short tests and accuracy metrics weakly linked to control decisions. To ground the review, real-time object detectors were deployed on Raspberry Pi edge devices, using person detections as occupancy counts. Externally metered power, sustained throughput, compute use and count error were assessed together. The case study is illustrative rather than exhaustive, showing that device power was broadly similar across models while processing speed and energy per processed frame varied substantially. YOLO-family models generally gave the most reliable counts, while lighter detectors improved computational efficiency but required post-processing or higher undercounting risk. The review argues that occupancy-vision research should move beyond model ranking toward decision-oriented benchmarks covering occlusion, domain shift, privacy, thermal resilience, duty cycling and net energy benefit. Vision sensing can support intelligent buildings, but only when its energy and governance costs are explicitly considered.
Keywords Buildings, Edge AI, Vision-based occupancy sensing, Occupant-centric control
Copyright ©
Energy Proceedings