Ayoubi, M. and Arashpour, M. (2026) Early recognition of workplace hazards using data-efficient multimodal learnable prompting and parameter-efficient fine tuning. Journal of Construction Engineering and Management, 152(10): 04026165, ISSN 0733-9364
Abstract
Timely and accurate recognition of workplace hazards is critical for ensuring safety in dynamic and high-risk environments such as construction sites. However, existing video-based approaches often rely on extensive annotations and fully supervised training, and many are tailored to specific hazard types, which limits scalability to rare or diverse scenarios. This study investigates whether a data-efficient vision-language framework can support early-stage hazard recognition from short preincident video segments under limited supervision. The proposed method builds on a video-adapted contrastive language-image pre-training (CLIP)-based vision-language backbone and introduces a parameter-efficient adaptation strategy that combines learnable visual prompting with low-rank adaptation (LoRA) applied to the text encoder. Learnable visual prompts capture global, summary, and local spatiotemporal cues from short video clips, and LoRA enables lightweight semantic adaptation without fine-tuning the full backbone. This design preserves most pretrained parameters and introduces only a small number of additional trainable parameters, making it well suited to few-shot learning in data-scarce safety scenarios. To evaluate early recognition capability, this paper further examines performance under shorter temporal observation windows by reducing the amount of video available for prediction. Experiments on a curated real-world hazard video data set show that the proposed method substantially improved few-shot recognition compared with a prompt-free baseline across five hazard categories and 5-, 10-, and 15-shot supervision. Overall, the findings suggest that combining temporal visual prompting with parameter-efficient text adaptation offers a scalable engineering pathway toward assistive hazard recognition in data-scarce environments, with promising but preliminary performance.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | computer vision; construction; deep learning; hazard recognition; workplace safety |
| Index terms: | deep learning, supervision, construction site, computer vision, workplace safety, experiment, adaptation, strategy, window |
| Subjects: | artificial intelligence, user focus, data collection methods, occupational health, computer vision, management, architectural elements, control systems, work location |
| Topics: | Research Practice, Site Management, Project Management, Design Practice, Business Strategy, Health and Safety, Digital Applications |
| Descriptive scope: | 3 PCE |
N.B. Descriptive scope is a count of how many of the five facets of empirical research are indicated by the words used in title, abstract and keywords. It is not intended as a judgement on the research; merely a count of the kind of word we would expect to indicate Phenomenon, Concepts, Theoretical framing, Empirical techniques, Analytical techniques. If all five are present, then a code of “5 PCTEA” will indicate this. If you feel the coding for this record is questionable, we welcome discussion around the terms we matched or the way we categorized them. The facet you would expect may not be coded, or a facet may be coded inappropriately. This can also bear on a larger question, of which facets should be treated as defining in construction management research. Please get in touch, and we will look at it. More details here