Cusumano, L; Olsson, N; Granath, M; Jockwer, R and Rempling, R (2025) Clustering techniques and keyword extraction with large language models for knowledge discovery in building defects data. Construction Innovation, 25(7), pp. 76-97. ISSN 1471-4175
Abstract
Purpose: The construction industry is undergoing a digital transformation and now holds large volumes of digital building defects data collected during inspections. This study aims to suggest an artificial intelligence-based method for analysing such building defects data to provide insights and knowledge faster than with traditional manual methods. Design/methodology/approach: This research explores a data set containing over 34,000 defects from hospital projects performed in Sweden from 2018 to 2021. The data mining uses keyword extraction based on both TF-IDF vectorisation and k-means clustering, the Mistral 7B model and KeyLLM. The results are compared with a content analysis using the GPT 3.5 turbo model. The analysis is performed both on an organisational and project level. Findings: The paper presents a combination of methods for analysing building defects data. The result shows that the most common problems reported during the inspections concern missing fire sealing, jointing and subceiling problems. Using k-means clustering gives fast insights into the main defect categories of the data set but requires domain knowledge. Keyword extraction using an LLM requires longer computational time but creates a deeper understanding of subcategories of defects. Finally, GPT-based content analysis is a complement to provide project-specific insights and allow user-specific requests. Research limitations/implications: The study is performed using data digitally collected in Swedish hospital projects. However, the results and methodology can be applied on other project data, such as safety inspections and warranty data. The analysis focused solely on text data. Originality/value: The method suggested in this paper uses clustering techniques and Large Language Models for analysing building defect data. The value of the proposed method is a faster process for leveraging knowledge from large amounts of unstructured text data, such as building defect reports, safety and moisture inspections and warranty issues.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | defects; inspections; knowledge generation; llm |
| Index terms: | artificial intelligence, clustering, construction industry, safety inspection, warranty, inspection, transformation, Sweden, data mining, knowledge generation, methodology, project data, content analysis, large language model, building defect |
| Subjects: | quality assurance, Geography, research methods, professional practice, knowledge management, industry analysis, performance measurement, data science, artificial intelligence, data analysis and analytics, business, data collection methods |
| Topics: | Digital Applications, Design Practice, Business Strategy, Research Practice, Quality Management, Engineering Principles, Geographical Context |
| Descriptive scope: | 5 PCTEA |
N.B. Descriptive scope is a count of how many of the five facets of empirical research are indicated by the words used in title, abstract and keywords. It is not intended as a judgement on the research; merely a count of the kind of word we would expect to indicate Phenomenon, Concepts, Theoretical framing, Empirical techniques, Analytical techniques. If all five are present, then a code of “5 PCTEA” will indicate this. If you feel the coding for this record is questionable, we welcome discussion around the terms we matched or the way we categorized them. The facet you would expect may not be coded, or a facet may be coded inappropriately. This can also bear on a larger question, of which facets should be treated as defining in construction management research. Please get in touch, and we will look at it. More details here