Abstract

Reliable personnel detection is critical for the safe deployment of autonomous haulage systems in underground mining, where challenging environmental conditions demand robust real-time perception. Existing research has focused primarily on fine-tuned convolutional detectors, while systematic comparisons with zero-shot vision-language models remain limited. This study presents a cross-paradigm benchmark comparing four zero-shot vision-language models (YOLO-World, Grounding DINO, OWL-ViT, and OWLv2) with four fine-tuned YOLO detectors (YOLOv8s, YOLOv9s, YOLO11s, and YOLO26s) using 31,396 real-world underground coal mine images. Detection performance was evaluated using precision, recall, F1-score, average precision, inference speed, and condition- and target scale-specific recall. The experimental results show that the fine-tuned detectors achieved F1-scores of 0.8449–0.8662 and AP50 values of 0.8791–0.9135, with YOLO26s achieving the strongest overall performance. In comparison, the zero-shot models achieved F1-scores of 0.3186–0.5484 and AP50 values of 0.2484–0.5108, with Grounding DINO performing best among the zero-shot models. Fine-tuned detectors also maintained substantially higher recall for occluded personnel and small apparent targets. These findings demonstrate a substantial performance advantage for domain-specific fine-tuning over the evaluated zero-shot approaches and establish a controlled cross-paradigm benchmark for comparing detection paradigms for underground personnel perception.

Department(s)

Mining Engineering

Publication Status

Open Access

Comments

Centers for Disease Control and Prevention, Grant U60OH012685-01-00

Keywords and Phrases

autonomous haulage; mine safety and health; object detection; personnel detection; underground mining; vision-language models; YOLO

International Standard Serial Number (ISSN)

1424-8220

Document Type

Article - Journal

Document Version

Final Version

File Type

text

Language(s)

English

Rights

© 2026 The Authors, All rights reserved.

Creative Commons Licensing

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Publication Date

01 Sep 2026

PubMed ID

42740266

Share

 
COinS