The construction industry continues to experience high rates of accidents and fatalities, underscoring the need for proactive and reliable safety management. Accurate hazard identification is essential for effective image-based monitoring and decision support on construction sites. Vision–language models (VLMs) have shown strong potential for interpreting complex visual environments; however, their deployment in safety-critical applications is limited by hallucination, where hazards are inferred without sufficient evidence. To address this limitation, this study proposes a retrieval-augmented hazard detection framework that grounds VLM outputs in visual evidence through image-to-image retrieval of visually similar hazard scenarios and structured domain knowledge within a retrieval-augmented generation (RAG) pipeline. The framework integrates image-to-image retrieval of real-world hazard scenarios, supported by a curated knowledge base of 1,040 annotated construction site images, structured hazard taxonomies, and regulation-aware verification to enforce condition-constrained hazard reasoning before compliance assessment. The proposed approach is evaluated on a separate set of 260 construction site images spanning ten construction-safety hazard categories aligned with OSHA safety domains and the construction-safety literature. At the end-to-end report level, the complete pipeline achieves an F1-score of 0.917 (precision 0.939, recall 0.896) and a hallucination rate of 6.1%, substantially lower than the VLM-only baseline. At the hazard-verification gate, adding taxonomy-constrained reasoning to RAG-1 yields a precision of 0.987 and an F1-score of 0.941, while RAG-1 alone achieves an F1-score of 0.923 with 89.8% retrieval coverage. These results indicate that visual grounding provides the largest reliability gain, with structured condition verification adding further precision and gated regulatory retrieval providing traceable compliance support after hazard confirmation.
Toward trustworthy construction safety hazard detection with visual-language models
Orooje, Mina Sadat;Re Cecconi, Fulvio;
2026-01-01
Abstract
The construction industry continues to experience high rates of accidents and fatalities, underscoring the need for proactive and reliable safety management. Accurate hazard identification is essential for effective image-based monitoring and decision support on construction sites. Vision–language models (VLMs) have shown strong potential for interpreting complex visual environments; however, their deployment in safety-critical applications is limited by hallucination, where hazards are inferred without sufficient evidence. To address this limitation, this study proposes a retrieval-augmented hazard detection framework that grounds VLM outputs in visual evidence through image-to-image retrieval of visually similar hazard scenarios and structured domain knowledge within a retrieval-augmented generation (RAG) pipeline. The framework integrates image-to-image retrieval of real-world hazard scenarios, supported by a curated knowledge base of 1,040 annotated construction site images, structured hazard taxonomies, and regulation-aware verification to enforce condition-constrained hazard reasoning before compliance assessment. The proposed approach is evaluated on a separate set of 260 construction site images spanning ten construction-safety hazard categories aligned with OSHA safety domains and the construction-safety literature. At the end-to-end report level, the complete pipeline achieves an F1-score of 0.917 (precision 0.939, recall 0.896) and a hallucination rate of 6.1%, substantially lower than the VLM-only baseline. At the hazard-verification gate, adding taxonomy-constrained reasoning to RAG-1 yields a precision of 0.987 and an F1-score of 0.941, while RAG-1 alone achieves an F1-score of 0.923 with 89.8% retrieval coverage. These results indicate that visual grounding provides the largest reliability gain, with structured condition verification adding further precision and gated regulatory retrieval providing traceable compliance support after hazard confirmation.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_45-ITcon-Orooje.pdf
accesso aperto
Descrizione: Articol
:
Publisher’s version
Dimensione
1.27 MB
Formato
Adobe PDF
|
1.27 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



