nav emailalert searchbtn searchbox tablepage yinyongbenwen piczone journalimg journalInfo journalinfonormal searchdiv searchzone qikanlogo popupnotification paper paperNew
基于GCViT-FROS的水工混凝土裂缝精细化识别方法及应用
基金项目(Foundation): 广东省水利科技创新项目(2024-07); 江苏省自然科学基金青年项目(BK20241519); 国家自然科学基金(52409018); 云南省科技人才与平台计划:云南省数字水工程技术创新中心(No.202305AK340003)
邮箱(Email): baotf@hhu.edu.cn
DOI:
发布时间: 2026-07-24
出版时间: 2026-07-24
网络发布时间: 2026-07-24
移动端阅读
摘要:

针对水工混凝土裂缝在复杂背景下形态多变、跨度大且分布极度稀疏等特性,现有基于卷积神经网络的裂缝识别方法存在识别不连续、细节丢失等痛点,本文提出一种基于GCViT-FROS的精细化识别方法。该方法以具备层级化混合注意力机制的Global Context Vision Transformer(GCViT)为编码器,用其“局部-全局”交替感知特性,在精准捕捉边缘细节的同时构建长程依赖。在特征融合阶段,设计改进感受野模块(Modified Receptive Field Block,MRFB)并集成于特征金字塔网络,利用MRFB的膨胀卷积增强空间采样,金字塔结构提升深层信息挖掘能力。在解码阶段设计了空间自适应门控机制(Spatial Adaptive Gate,SA-Gate),融合对象上下文表示模块,利用裂缝区域语义关联及门控机制抑制背景中水渍、麻面等噪声,增强裂缝区域的像素一致性。消融实验结果验证了各核心模块对模型性能的提升作用。在对比实验中,本模型平均交并比达到了86.84%,且本模型识别裂缝的连续性与精度均优于主流模型。在溢洪道该典型水工场景中应用,验证了该方法在实际工程场景下具有良好的泛化能力与应用潜力。

Abstract:

To address the challenges posed by hydraulic concrete cracks under complex backgrounds—namely, high morphological variability, large spans, and extremely sparse distribution—existing convolutional neural network-based crack identification methods suffer from discontinuous recognition and loss of fine details. This paper proposes a refined identification method based on GCViT-FROS. The method employs the Global Context Vision Transformer (GCViT), which features a hierarchical hybrid attention mechanism, as the encoder. By leveraging its alternating local-global perception capability, the method accurately captures edge details while simultaneously establishing long-range dependencies. In the feature fusion stage, a Modified Receptive Field Block (MRFB) is designed and integrated into the feature pyramid network. The MRFB uses dilated convolutions to enhance spatial sampling, and the pyramid structure improves the ability to mine deep information. In the decoding stage, a Spatial Adaptive Gate (SA-Gate) mechanism is devised, incorporating an object context representation module. By exploiting the semantic correlations of crack regions and the gating mechanism, background noise such as water stains and rough surface textures is suppressed, thereby enhancing pixel-wise consistency within crack regions. Ablation study results validate the effectiveness of each core module in improving model performance. In comparative experiments, the proposed model achieves a mean Intersection over Union (mIoU) of 86.84%, and outperforms mainstream models in both continuity and accuracy of crack identification. Application to a typical hydraulic engineering scenario—a spillway—demonstrates the method’s strong generalization capability and application potential in real-world engineering contexts.

参考文献

[1] 任秋兵, 沈扬, 李明超, 等.水工建筑物安全监控深度分析模型及其优化研究[J].水利学报, 2021, 52(01): 71-80. (REN Qiubing, SHEN Yang, LI Mingchao, et al. Deep analysis model for safety monitoring of hydraulic structures and its optimization[J]. Journal of Hydraulic Engineering, 2021,52(01): 71-80. (in Chinese))

[2] 马春辉, 陆希, 贾冬焱, 等.复杂条件下高面板堆石坝的特殊结构体安全监测设计体系研究[J].水利学报, 2024, 55(07): 848-861+873. (MA Chunhui, LU Xi, JIA Dongyan, et al. Research on safety monitoring design system for special structures of high concrete face rockfill dams under complex conditions[J]. Journal of Hydraulic Engineering, 2024, 55(07): 848-861+873. (in Chinese))

[3] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[C]//International Conference on Learning Representations. 2021.

[4] Xie E, Wang W, Yu Z, et al. SegFormer: Simple and efficient design for semantic segmentation with transformers[C]//Advances in Neural Information Processing Systems. Curran Associates, Inc., 2021: 12077-12090.

[5] Liu Z, Lin Y, Cao Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C]//Proceedings of IEEE/CVF International Conference on Computer Vision. IEEE Computer Society, 2021: 9992-10002.

[6] 赵阳, 康飞, 万刚.基于改进CycleGAN与YOLOv8s的混凝土坝水下裂缝识别方法[J].水电能源科学, 2025, 43(04): 158-162. (ZHAO Yang, KANG Fei, WAN Gang. Underwater Crack Identification for Concrete Dams Based on Improved CycleGAN and YOLOv8s[J]. Water Resources and Power, 2025, 43(04): 158-162. (in Chinese))

[7] 高治鑫, 包腾飞, 李扬涛.基于机器学习的混凝土坝表面裂缝快速识别方法[J].水电能源科学, 2022, 40(04): 95-98. (GAO Zhixin, BAO Tengfei, LI Yangtao. Rapid Identification Method of Surface Cracks of Concrete Dams Based on Machine Learning Algorithms[J]. Water Resources and Power, 2022, 40(04): 95-98. (in Chinese))

[8] Hatamizadeh A, Yin H, Heinrich G, et al. Global context vision transformers[C]//Proceedings of the International Conference on Machine Learning. 2023.

[9] T. Lin, P. Dollár, R. Girshick, et al. Feature Pyramid Networks for Object Detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA: IEEE, 2017: 936-944.

[10] Yuan, Yuhui, Chen, Xilin, Wang, Jingdong, et al. Object-Contextual Representations for Semantic Segmentation[C]//Proceedings of the 16th European Conference on Computer Vision. ELECTR NETWORK: Springer, 2020: 173-190.

[11] 张小伟, 包腾飞.基于局部大津阈值与区域生长的坝面细小裂缝识别分割算法[J].水电能源科学, 2022, 40(2): 97-100. (ZHANG Xiaowei, BAO Tengfei. Identification and Segmentation Algorithm of Small Cracks on Dam Surface Based on Local OTSU Threshold and Regional Growth[J]. Water Resources and Power, 2022, 40(2): 97-100. (in Chinese))

[12] Liu Songtao, Huang Di, Wang Yunhong. Receptive field block net for accurate and fast object detection[C]//Computer Vision - ECCV 2018. Cham: Springer International Publishing, 2018: 404-419.

[13] 封婧仪, 梁晖, 齐智勇, 等.融合多尺度特征与注意力机制的混凝土裂缝语义分割模型[J].水力发电学报, 2025, 44(09): 114-124. (FENG Jingyi, LIANG Hui, QI Zhiyong, et al. Semantic segmentation model for concrete cracks integrating multi-scale features and attention mechanisms[J]. Journal of Hydroelectric Engineering, 2025, 44(09): 114-124. (in Chinese))

[14] Chen L C, Zhu Y, Papandreou G, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation[C]//Proceedings of the European Conference on Computer Vision. Munich, GERMANY: Springer, 2018: 801-818.

[15] Wang J, Sun K, Cheng T, et al. Deep high-resolution representation learning for visual recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 43(10): 3349-3364.

[16] Zhao H, Shi J, Qi X, et al. Pyramid scene parsing network[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, HI, USA: IEEE, 2017: 2881-2890.

基本信息:

中图分类号:TP391.41;TV544

引用信息:

[1]徐萍萍,包腾飞,龚健,等.基于GCViT-FROS的水工混凝土裂缝精细化识别方法及应用[J].水电能源科学().

基金信息:

广东省水利科技创新项目(2024-07); 江苏省自然科学基金青年项目(BK20241519); 国家自然科学基金(52409018); 云南省科技人才与平台计划:云南省数字水工程技术创新中心(No.202305AK340003)

发布时间:

2026-07-24

出版时间:

2026-07-24

网络发布时间:

2026-07-24

检 索 高级检索

引用

GB/T 7714-2015 格式引文
MLA格式引文
APA格式引文