Mun, H.; Jeong, J.; Jeong, J. and Kumi, L. (2026) Unsupervised learning approach for benchmark models to identify construction projects with high accident risk levels. Engineering, Construction and Architectural Management, 33(5), pp. 3777-3798. ISSN 0969-9988
Abstract
Purpose – The construction sector is highly prone to accidents, traditionally assessed using subjective qualitative measurements. To enhance the allocation of risk management resources and identify high-risk projects during pre-construction, an objective and quantitative approach is necessary. This study introduces a three-step clustering methodology to quantitatively evaluate accident risk levels in construction projects. Design/methodology/approach – In the first step, accident and total construction revenue by project were collected to calculate accident probabilities. In the second step, accident probabilities were calculated by project type using the data collected in the first step. After that, benchmark models were suggested using clustering methods to identify high-risk project types for risk management. Before suggesting the benchmark models, an uncertainty analysis was conducted due to the limited amount of data. In the third step, the suggested benchmark models were validated for accuracy. Findings – The results categorized risk levels for fatalities and injuries into four distinct groups. Validation through ordinal logistic regression demonstrated high explanatory power, with fatality risk levels ranging from 79.9 to 100% and injury risk levels from 90.3 to 100%. Originality/value – This benchmark model facilitates effective comparisons and analyses across various construction sectors and countries, offering a robust quantitative standard for risk management. By identifying high-risk projects such as "Dam, " this methodology enables better resource allocation during the pre-construction phase, thereby improving overall safety management in the construction industry and providing a basis for legislative applications.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | benchmark; k-means clustering; ordinal logistic regression; risk level; unsupervised learning |
| Index terms: | validation, pre-construction, clustering, resource allocation, accuracy, methodology, uncertainty analysis, fatalities, safety management, revenue, injury, risk management, construction sector, explanatory power, construction project, construction industry, logistic regression |
| Subjects: | health conditions and diseases, project delivery, data science, resource management, research methods, industry analysis, production management, environmental hazards, health risk and incident analysis, theoretical framing, statistical analysis, professional development, risk assessment, economic analysis, occupational health and safety management |
| Topics: | Sustainability, Business Strategy, Health and Safety, Risk Management, Information Management, Site Management, Digital Applications, Research Practice, Project Management |
| Descriptive scope: | 4 PCTA |
N.B. Descriptive scope is a count of how many of the five facets of empirical research are indicated by the words used in title, abstract and keywords. It is not intended as a judgement on the research; merely a count of the kind of word we would expect to indicate Phenomenon, Concepts, Theoretical framing, Empirical techniques, Analytical techniques. If all five are present, then a code of “5 PCTEA” will indicate this. If you feel the coding for this record is questionable, we welcome discussion around the terms we matched or the way we categorized them. The facet you would expect may not be coded, or a facet may be coded inappropriately. This can also bear on a larger question, of which facets should be treated as defining in construction management research. Please get in touch, and we will look at it. More details here