Research Progress in YOLO-Based Road Crack Detection: A Critical Narrative Review of Architectures, Data, Evaluation and Deployment
Kaiyuan Shen *
School of Civil Engineering and Transportation, North China University of Water Resources and Electric Power, Zhengzhou, Henan 450045, China.
Fanpeng Meng
Powerchina Harbour Co., Ltd., Jinan, Shandong 250000, China.
Song Hu
Powerchina Harbour Co., Ltd., Jinan, Shandong 250000, China.
Lindong Li
Powerchina Harbour Co., Ltd., Jinan, Shandong 250000, China.
Lingchao Chen
Powerchina Harbour Co., Ltd., Jinan, Shandong 250000, China.
*Author to whom correspondence should be addressed.
Abstract
Road crack detection is a central computer-vision task in pavement inspection because cracks are visually sparse, geometrically elongated, highly variable in width and topology, and easily confused with shadows, joints, stains and road markings. The You Only Look Once (YOLO) family has become prominent in this domain because it offers an attractive accuracy-speed trade-off for vehicle-mounted, mobile and edge inspection. This critical narrative review evaluates research progress in YOLO-based road and pavement crack detection from 2016 to 26 June 2026, with emphasis on how architectural modifications, task formulation, dataset design and deployment constraints affect the strength of reported evidence. Literature was selected from accessible scholarly indexes and transport/engineering sources, supplemented by citation searching and DOI-level bibliographic verification. The evidence indicates a clear progression from straightforward transfer of generic YOLO detectors towards crack-aware designs that combine multi-scale feature fusion, attention or transformer components, lightweight convolution, revised detection heads, specialised localisation losses and detector-segmenter pipelines. Recent work increasingly addresses edge deployment, illumination variation and background interference. Nevertheless, apparent benchmark gains are difficult to compare because studies differ in datasets, crack taxonomies, train-test partitioning, image resolution, augmentation, hardware and metric definitions. Bounding-box detection also remains an imperfect representation for thin crack morphology, whereas pixel-level segmentation and geometric measurement provide richer maintenance information at higher annotation and computational cost. Cross-region generalisation, uncertainty estimation, reproducible latency measurement, calibration against maintenance-relevant severity and evaluation under adverse weather remain underdeveloped. The strongest direction is therefore not continued accumulation of isolated architectural modules, but standardised, cross-domain evaluation of integrated detection, segmentation and quantification systems under realistic operational constraints. YOLO-based crack detection is technically mature enough for increasingly credible field deployment, yet evidence for robust transfer across roads, sensors and environments remains less mature than within-dataset accuracy results suggest.
Keywords: Pavement inspection, object detection, deep learning, road damage, crack segmentation, edge computing, computer vision, infrastructure monitoring