Liu, J; Li, W; Li, F and Zhao, X (2025) Chinese named entity recognition for bridge damage and defects based on text mining and natural language pretraining models. Journal of Construction Engineering and Management, 151(6): 04025060, ISSN 0733-9364
Abstract
Bridge inspection reports are a vital source of data for bridge management and maintenance, encompassing essential structural information indispensable for damage evaluation and decision-making. However, in the process of automatically extracting unstructured textual data and identifying damage entities, because the same type of bridge damage entity often corresponds to multiple structural components, and strong correlations along with prominent nested features exist among entities, general named entity recognition (NER) methods have limited effectiveness. To address these issues, this study introduces a novel method for NER of damage and defects in bridge inspection, leveraging text mining and pretrained natural language models. First, the study constructs a specialized corpus of bridge damage and defects from a large number of bridge inspection reports, and fine-grained entity annotations are performed on sentences describing damage and defects. Next, the study proposes an advanced bridge damage entity recognition model, which integrates pretrained natural language models with deep learning models. The model leverages the Bidirectional Encoder Representations from Transformers (BERT) pretrained model to extract vector features from Chinese characters in damage-related sentences. It then utilizes a bidirectional long short-term memory (BiLSTM) network to capture sequential patterns of multitype entity labels. Finally, it integrates conditional random fields (CRF) to enforce label constraints, generating the optimal label sequence. The model is validated through experiments using the constructed Chinese bridge inspection damage and defect named entity corpus. Experimental results demonstrate that the model proposed in this study surpasses other mainstream NER models, achieving an F1 score of 98.31% and successfully identifying seven categories of fine-grained bridge damage entities. This study not only enhances the automation of extracting information from damage-related bridge inspection text sentences but also establishes a solid foundation for building knowledge graphs in the bridge domain, advancing the development of intelligent bridge management.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | bridge damage; bridge inspection; name entity recognition; natural language processing; pretraining model |
| Index terms: | effectiveness, experiment, mining, inspection, deep learning, decision-making, automation |
| Subjects: | data collection methods, automation and robotics, performance management, decision analysis, geotechnical engineering, artificial intelligence, quality assurance |
| Topics: | Research Practice, Digital Applications, Engineering Principles, Risk Management, Quality Management |
| Descriptive scope: | 3 PCE |
N.B. Descriptive scope is a count of how many of the five facets of empirical research are indicated by the words used in title, abstract and keywords. It is not intended as a judgement on the research; merely a count of the kind of word we would expect to indicate Phenomenon, Concepts, Theoretical framing, Empirical techniques, Analytical techniques. If all five are present, then a code of “5 PCTEA” will indicate this. If you feel the coding for this record is questionable, we welcome discussion around the terms we matched or the way we categorized them. The facet you would expect may not be coded, or a facet may be coded inappropriately. This can also bear on a larger question, of which facets should be treated as defining in construction management research. Please get in touch, and we will look at it. More details here