【科研进展】课题组在鲁棒流形学习与可解释稀疏学习方向取得新进展
近日,本课题组在鲁棒流形学习与可解释稀疏学习方向取得新进展,两项成果分别发表于Neural Networks和Data Mining and Knowledge Discovery。两项成果第一作者均为本课题组毕业博士张学林,通讯作者为信息学院陈洪教授。
成果一:Bilevel Manifold Fitting
针对加性噪声与噪声维度等复杂噪声环境下流形拟合的鲁棒性与泛化性难题,本研究将元学习方案与循环一致性度量融入基于流形的生成对抗框架,提出了双层循环生成对抗网络 (Bilevel Cycle Generative Adversarial Network, BCGAN)。该网络能够自动为环境空间或隐空间数据分配掩码,筛选真正有效的信息维度,学习鲁棒的互映射,并生成新的合成样本。针对离散掩码变量难以直接学习、双层极小极大问题超梯度难以逼近等挑战,该研究设计了基于投影梯度估计的概率双层优化策略,显著降低了Hessian矩阵与Jacobian矩阵的计算开销。理论上,该文建立了随机双层极小极大算法的泛化误差上界,揭示了泛化能力与参数设置之间的关系。合成数据与真实数据上的实验验证了该方法在有效特征识别、鲁棒流形估计、流形去噪与非线性插值方面的竞争力与鲁棒性。该工作为复杂噪声环境下的鲁棒流形学习提供了新的建模范式与理论支撑。
【英文摘要】The manifold assumption states that high-dimensional ambient data possess a low-dimensional geometric structure, which has promoted many manifold learning models achieving promising performance for wide applications. However, when there exists additive noise or noisy dimensions, it is challenging to estimate the latent manifold since the noisy information tends to mislead the data-driven mapping from the ambient space to the latent space. To address this issue, we formulate a Bilevel Cycle Generative Adversarial Network (BCGAN) that comprises two generative adversarial networks and a specific manifold fitting module. This network can automatically assign masks to the ambient or latent data, learn robust mutual mappings, and generate new synthetic samples. Theoretically, we establish the upper bounds of generalization error for the stochastic bilevel minimax problems to reveal the relationship between generalization capability and parameter settings. Experiments on both synthetic and real-world datasets verify the competitiveness and robustness of the proposed approach for manifold fitting with corrupted data.

成果二:Meta Additive Model: Interpretable Sparse Learning With Auto Weighting
针对稀疏可加模型大多局限于均方误差准则下的单层学习、在非高斯扰动、离群点、噪声标签及类别不均衡等复杂噪声下性能显著退化的问题,以及传统样本重加权策略需预先指定加权函数并手动选取额外超参数的不足,基于双层优化框架提出了元可加模型(Meta Additive Model, MAM)。该方法通过MLP参数化加权函数,并利用少量元数据自动学习各损失项的数据驱动权重,实现了自动加权与稀疏变量选择的统一,可同时胜任变量选择、鲁棒回归估计与不均衡分类等多种学习任务。理论上,该文在温和条件下给出了算法的优化收敛性与泛化界保证,并证明了变量选择的一致性,将先前单层框架下的结果拓展到了双层优化设定。合成数据与真实数据上的实验表明,MAM在多种数据污染场景下均优于现有的先进可加模型。该工作是课题组前期可加模型(NeurIPS'17, ICML'20, NeurIPS'20, ICLR'22, ICML'23)研究的自然拓展。
【英文摘要】Sparse additive models have attracted much attention in high-dimensional data analysis due to their flexible representation and strong interpretability. However, most existing models are limited to single-level learning under the mean-squared error criterion, whose empirical performance can degrade significantly in the presence of complex noise, such as non-Gaussian perturbations, outliers, noisy labels, and imbalanced categories. The sample reweighting strategy is widely used to reduce the model's sensitivity to atypical data. However, it typically requires prespecifying the weighting functions and manually selecting additional hyperparameters. To address this issue, we propose a new meta additive model (MAM) based on the bilevel optimization framework, which learns data-driven weights of individual losses by parameterizing the weighting function via an MLP trained on meta data. MAM is capable of a variety of learning tasks, including variable selection, robust regression estimation, and imbalanced classification. Theoretically, MAM provides guarantees for optimization convergence, algorithmic generalization, and variable selection consistency under mild conditions. Empirically, MAM outperforms several state-of-the-art additive models on both synthetic and real-world data under various types of corruption. The implementation is available at https://github.com/zxlml/MAM.