文章摘要
农田土壤有机碳反演的变量筛选及算法组合优化——以泾河流域为例
Variable selection and algorithm combination optimization for farmland soil organic carbon inversion:a case study of the Jinghe River basin
投稿时间:2025-05-27  
DOI:10.13254/j.jare.2025.0489
中文关键词: 土壤有机碳,反演模型,机器学习,变量筛选,SHAP分析
英文关键词: soil organic carbon, inversion model, machine learning, variable screening, SHAP analysis
基金项目:国家自然科学基金项目(42271025,31961143011);陕西省重点研发计划产业链项目(2022ZDLSF06-04);陕西省创新团队项目(2021TD-52)
作者单位E-mail
张方滟 西安交通大学全球环境变化研究院, 西安 710049  
吴一平 西安交通大学全球环境变化研究院, 西安 710049
中南林业科技大学水土保持学院, 长沙 410004 
rocky.ypwu@gmail.com 
阴晓伟 西安交通大学全球环境变化研究院, 西安 710049  
孟泽昕 西安交通大学全球环境变化研究院, 西安 710049  
亚历山大罗夫·格奥尔基 俄罗斯科学院大气物理研究所, 莫斯科 119017  
李汇文 西安交通大学全球环境变化研究院, 西安 710049  
张广创 西安交通大学全球环境变化研究院, 西安 710049  
韩磊 长安大学土地工程学院, 西安 710054  
窦欣 南京信息工程大学地理科学学院, 南京 210044  
张雷 西安交通大学全球环境变化研究院, 西安 710049
陕西农业发展集团有限公司, 西安 710075 
 
王欢元 陕西农业发展集团有限公司, 西安 710075  
摘要点击次数: 715
全文下载次数: 31
中文摘要:
      为实现快速、高精度的农田土壤有机碳(SOC)反演并探讨不同变量筛选方法与机器学习算法的适宜组合方式,本研究以黄土高原典型丘陵沟壑区泾河流域农田土壤为对象,基于SOC采样数据,整合了包括MODIS多光谱数据及相关环境变量在内的126个变量,探究了全部变量集、竞争性自适应重加权采样法(CARS)筛选变量集、相关系数筛选变量集这3种变量集分别与支持向量机(SVM)、随机森林(RF)、反向传播神经网络(BPNN)和极端梯度提升树(XGBoost)这4种算法组合对模型精度的影响。结果表明,经过相关系数筛选后的11个特征变量与4种机器学习算法组合的模型精度显著优于全部变量集和CARS筛选变量集,其中与SVM算法的组合表现最优(验证集:R2=0.86,RMSE=0.95,MAE=0.76,RPD=2.66),其次为XGBoost、BPNN和RF模型。SHAP分析显示,地形变量(DEM、地形起伏度、剖面曲率)贡献度达28.41%,在地形复杂多变、耕地较为破碎的泾河流域农田SOC反演中起重要作用,其次为遥感指数(23.18%)和经纬度(20.21%)。相关系数筛选与SVM算法组合的反演结果表明,流域农田SOC范围为0.00~39.94 g·kg-1,其高值区集中在研究区域的西部边缘,低值区集中在东北部,中值区和低值区在中部和南部交错分布,平均SOC值为8.38 g·kg-1,其含量远低于全国农田SOC的平均水平。
英文摘要:
      To achieve rapid and high-precision inversion of soil organic carbon(SOC)in farmland and to explore suitable combinations of different variable screening methods and machine learning algorithms, this study focused on the farmland of the Jinghe River basin, a typical hilly-gully region of the Loess Plateau. Based on the SOC sample data, this study integrated 126 variables, including MODIS multispectral data and related environmental variables. The study explored the impacts of three variable sets:the full variable set, the competitive adaptive reweighted sampling(CARS) screened variable set, and the correlation coefficient screened variable set when combined with four machine learning algorithms:support vector machine(SVM), random forest(RF), back propagation neural network (BPNN), and extreme gradient boosting(XGBoost)on model accuracy. The results showed that models based on the 11 variables selected via correlation coefficient significantly outperformed those using the full or CARS-selected variable sets. Among these combinations, the SVM algorithm performed best(validation set:R2=0.86, RMSE=0.95, MAE=0.76, RPD=2.66), followed by the XGBoost, BPNN, and RF models. The SHAP analysis revealed that topographic variables(DEM, terrain relief, profile curvature)contributed 28.41% to the inversion of farmland SOC in the Jinghe River basin, where the topography was complex and variable and the cultivated land was relatively fragmented, followed by the remote sensing index(23.18%) and latitude/longitude(20.21%). The inversion results based on the combination of correlation coefficient selection and SVM algorithm revealed that the SOC of farmland in the basin ranged from 0.00 g·kg-1 to 39.94 g·kg-1, with high values concentrated in the western part of the basin and medium-to-low values interspersed in the central and southern areas. The average SOC content was 8.38 g·kg-1, which is lower than the national average for farmland SOC.
HTML   查看全文   查看/发表评论  下载PDF阅读器
关闭