Maco: An open-image data set with multiannotation strategies for building material counting

Huang, Q; Chen, Y; Li, Y; Liu, S; Fan, Z; Chen, X and Chen, J (2026) Maco: An open-image data set with multiannotation strategies for building material counting. Journal of Construction Engineering and Management, 152(4): 04026008, ISSN 0733-9364

Abstract

Building material management is fundamental to engineering projects. Currently, the counting of building materials is primarily conducted manually, a process that is time-consuming and prone to errors. Recent advances in computer vision and deep learning have significantly facilitated automation in the construction industry. However, the effectiveness of deep learning approaches hinges heavily on large-scale data sets with accurate, manually annotated data. Existing construction-domain image data sets are ill-suited for material counting tasks, and this lack of large-scale, publicly available data sets has become a major barrier to progress in this area. To address this gap, this study introduces and publicly releases a new large-scale image data set, called Material Counting in Construction (MACO), collected directly from construction sites. The MACO data set comprises 4,426 images and 563,595 annotated objects, covering seven common building materials and diverse real-world construction scenarios. To enhance the utility of MACO and to support more downstream tasks, three popular annotation techniques were employed: Horizontal bounding boxes (HBBs), oriented bounding boxes (OBBs), and segmentation masks. This data set is the first large-scale, multimaterial, multiannotation, and high-density open data set (with 127.3 annotations per image) among existing construction data sets. The validity of MACO is confirmed through benchmarking with two state-of-the-art one-stage object detection algorithms, achieving a maximum mean average precision at an intersection over union (IoU) threshold of 0.5 (mAP50) of 91.6%. This provides a robust benchmark for method selection in similar tasks. In addition, to compare the one-material-one-model and multimaterial-one-model paradigms, experiments were conducted on models trained on single versus multiple materials. Results indicated that the multimaterial counting model exhibits performance degradation due to cross-material feature interference, suggesting that developing a generalized detection and counting model requires further research. Overall, MACO is designed to advance intelligent construction site management, including material detection and counting, specification measurement, and robotic grasping and assembly.

Item Type: Article
Uncontrolled Keywords: building materials; computer vision; data set; deep learning; object counting
Index terms: computer vision, specification, construction site, strategy, experiment, state of the art, benchmarking, object detection, density, degradation, deep learning, validity, paradigm, building material, effectiveness, automation, open data, construction industry
Subjects: contractual condition, performance management, performance measurement, education and knowledge transfer, management, industry analysis, data management, data collection methods, material degradation and durability, artificial intelligence, work location, building materials, computer vision, automation and robotics, research dissemination and communication, evaluation and assessment methods, analytical methods
Topics: Construction Materials, Research Practice, Business Strategy, Digital Applications, Urban Studies, Design Practice, Contract Administration, Site Management, Quality Management
Descriptive scope: 5 PCTEA

N.B. Descriptive scope is a count of how many of the five facets of empirical research are indicated by the words used in title, abstract and keywords. It is not intended as a judgement on the research; merely a count of the kind of word we would expect to indicate Phenomenon, Concepts, Theoretical framing, Empirical techniques, Analytical techniques. If all five are present, then a code of “5 PCTEA” will indicate this. If you feel the coding for this record is questionable, we welcome discussion around the terms we matched or the way we categorized them. The facet you would expect may not be coded, or a facet may be coded inappropriately. This can also bear on a larger question, of which facets should be treated as defining in construction management research. Please get in touch, and we will look at it. More details here