| 时间: 2026-07-24 | 次数: |
张嘉睿, 曹哲骁, 王田,等.基于增强数据对比的自监督学习综述[J].河南理工大学学报(自然科学版),2026,45(5):48-59.
ZHANG J R, CAO Z X, WANG T,et al.Data augmentation-based contrastive self-supervised learning: A review[J].Journal of Henan Polytechnic University(Natural Science) ,2026,45(5):48-59.
基于增强数据对比的自监督学习综述
张嘉睿1, 曹哲骁1, 王田1, 王传云2, 滕婧3, 傅瑶4
1.北京航空航天大学 人工智能学院,北京 100191;2.沈阳航空航天大学 人工智能学院,辽宁 沈阳 110121;3.华北电力大学 控制与计算机工程学院,北京 102206;4.中国科学院 长春光学精密机械与物理研究所,吉林 长春 130033
摘要: 意义 传统监督学习高度依赖海量人工标注数据,面临标注成本高和易引入主观偏差等问题,限制了模型的泛化能力。同时,海量无标签数据的潜在价值亟待挖掘,自监督学习由此成为研究热点。对比学习作为判别式自监督学习的核心代表,通过构建正负样本对,在特征空间中拉近同类实例距离并推远异类实例距离,迫使网络主动学习高层语义特征。与生成式方法相比,对比学习具有更强的特征迁移能力和更低的计算成本,为通用表征学习开辟了新道路。 进展 系统梳理了对比学习的通用理论框架和核心组件,包括定义语义不变性的数据增强策略、决定特征空间判别能力的样本对构造机制、以孪生网络为基础的编码器架构,以及以InfoNCE为代表的损失函数机理。同时,总结了MoCo,SimCLR,Barlow Twins和SimSiam等典型算法的技术路径。此外,还探讨了对比学习在缓解长尾分布表征失衡与实现多模态图文语义对齐等前沿场景的应用机制,并分析了大批量训练中的假负样本干扰与跨模态异构性等现存挑战。 结论与展望对比学习有效突破了传统深度学习对标注数据的依赖瓶颈,为构建具备强大泛化能力的通用人工智能感知系统奠定了技术基础。未来,该领域仍需向纵深发展。一方面亟需设计更轻量的样本匹配算法以优化计算效率;另一方面需从数学维度增强理论解释性,揭示特征解耦与泛化的内在动因。
关键词:对比学习;自监督学习;数据增强;长尾分布;多模态学习
doi:10.16186/j.cnki.1673-9787.2026010020
基金项目:国家自然科学基金资助项目(92467108);北京市昌平区科技副区长专项项目(202504005043)
收稿日期:2026/01/13
修回日期:2026/05/23
出版日期:2026-07-24
Data augmentation-based contrastive self-supervised learning: A review
Zhang Jiarui1, Cao Zhexiao1, Wang Tian1, Wang Chuanyun2, Teng Jing3, Fu Yao4
1.School of Artificial Intelligence, Beihang University, Beijing 100191, China;2.School of Artificial Intelligence, Shenyang Aerospace University, Shenyang 110121, Liaoning, China;3.School of Control and Computer Engineering, North China Electric Power University, Beijing 102206, China;4.Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, Jilin, China
Abstract: Significance Traditional supervised learning relies heavily on large-scale manually annotated datasets, which incur high labeling costs and introduce potential human bias, thereby limiting model generalization. Meanwhile, the vast amount of unlabeled data remains underutilized. As a result, self-supervised learning has emerged as a major research focus. Among its paradigms, contrastive learning has become a representative discriminative approach. It constructs positive and negative sample pairs and learns representations by pulling similar instances closer while pushing dissimilar ones apart in the feature space, enabling models to capture high-level semantic features in a self-supervised manner. Compared with generative methods, contrastive learning demonstrates stronger transferability and lower computational costs, providing a promising direction for universal representation learning. Progress This paper systematically reviews the theoretical framework and core components of contrastive learning, including data augmentation strategies that enforce semantic invariance, sample construction mechanisms that determine feature discriminability, Siamese-network-based encoder architectures, and loss functions such as InfoNCE. Representative methods, including MoCo, SimCLR, Barlow Twins, and SimSiam, are further summarized in terms of their technical designs. In addition, emerging applications are discussed such as representation learning under long-tailed distributions and multimodal vision-language alignment. key challenges are analyzed in contrastive learning, including false negative sample issues in large-batch training and modality heterogeneity in cross-modal scenarios. Conclusions and Prospects Contrastive learning effectively alleviates the reliance on large-scale labeled data, providing a solid foundation for developing perception systems in artificial general intelligence. Further research should focus on designing lightweight and efficient sample matching mechanisms to reduce computational complexity, while also improving theoretical interpretability to better understand feature disentanglement and generalization mechanisms.
Key words:contrastive learning;self-supervised learning;data augmentation;long-tailed distribution;multimodal learning