Responsible AI in construction safety: Systematic evaluation of large language models and prompt engineering

Sammour, F; Xu, J; Wang, X; Hu, M and Zhang, Z (2026) Responsible AI in construction safety: Systematic evaluation of large language models and prompt engineering. Journal of Construction Engineering and Management, 152(1): 04025217, ISSN 0733-9364

Abstract

Responsible artificial intelligence (AI) in construction safety requires rigorous, domain-specific evaluation. Deploying large language models (LLMs) without understanding their capabilities and limitations risks generating inaccurate information and compromising worker safety. This study benchmarked three leading LLMs - GPT-3.5, GPT-4o, and Gemini 2.0 Flash - using 385 Board of Certified Safety Professionals exam questions across seven knowledge areas. All models achieved passing scores (GPT-4o: 84.6%, Gemini: 80.8%, GPT-3.5: 73.8%), demonstrating strengths in the knowledge areas of safety management systems and hazard identification. However, LLMs performed relatively poorly in science and math, emergency response, and fire prevention. Error analysis identified four key limitations impacting performance: knowledge gaps, reasoning flaws, memory issues, and calculation errors. This study also highlights the impact of four prompt engineering factors, with variations in accuracy reaching 13.5% for GPT-3.5, 7.9% for GPT-4o, and 10.4% for Gemini. However, no single prompt configuration has proven universally effective across three models and seven knowledge areas, suggesting the need for model- and task-specific optimization. These findings establish critical benchmarks for the ethical and effective deployment of AI in construction safety, identifying areas where LLMs can support safety practices and where human oversight remains essential, and offering practical insights into optimizing LLM implementation through prompt engineering.

Item Type: Article
Uncontrolled Keywords: artificial intelligence; construction safety; language models; prompt engineering
Index terms: prevention, accuracy, implementation, construction safety, reasoning, emergency response, configuration, artificial intelligence, safety management system, variation, hazard identification, large language model, science
Subjects: artificial intelligence, financial risk, health safety and environment, contractual condition, safety engineering, environmental health, specialized education, contractual arrangements, professional development, systems engineering, cognitive psychology, data science
Topics: Health and Safety, Cost Management, Procurement, Digital Applications, Information Management, Engineering Principles, Education, Contract Administration, Research Practice, Sustainability
Descriptive scope: 2 PC

N.B. Descriptive scope is a count of how many of the five facets of empirical research are indicated by the words used in title, abstract and keywords. It is not intended as a judgement on the research; merely a count of the kind of word we would expect to indicate Phenomenon, Concepts, Theoretical framing, Empirical techniques, Analytical techniques. If all five are present, then a code of “5 PCTEA” will indicate this. If you feel the coding for this record is questionable, we welcome discussion around the terms we matched or the way we categorized them. The facet you would expect may not be coded, or a facet may be coded inappropriately. This can also bear on a larger question, of which facets should be treated as defining in construction management research. Please get in touch, and we will look at it. More details here