|
| 农田土壤有机碳反演的变量筛选及算法组合优化——以泾河流域为例 |
| Variable selection and algorithm combination optimization for farmland soil organic carbon inversion:a case study of the Jinghe River basin |
| 投稿时间:2025-05-27 |
| DOI:10.13254/j.jare.2025.0489 |
| 中文关键词: 土壤有机碳,反演模型,机器学习,变量筛选,SHAP分析 |
| 英文关键词: soil organic carbon, inversion model, machine learning, variable screening, SHAP analysis |
| 基金项目:国家自然科学基金项目(42271025,31961143011);陕西省重点研发计划产业链项目(2022ZDLSF06-04);陕西省创新团队项目(2021TD-52) |
| 作者 | 单位 | E-mail | | 张方滟 | 西安交通大学全球环境变化研究院, 西安 710049 | | | 吴一平 | 西安交通大学全球环境变化研究院, 西安 710049 中南林业科技大学水土保持学院, 长沙 410004 | rocky.ypwu@gmail.com | | 阴晓伟 | 西安交通大学全球环境变化研究院, 西安 710049 | | | 孟泽昕 | 西安交通大学全球环境变化研究院, 西安 710049 | | | 亚历山大罗夫·格奥尔基 | 俄罗斯科学院大气物理研究所, 莫斯科 119017 | | | 李汇文 | 西安交通大学全球环境变化研究院, 西安 710049 | | | 张广创 | 西安交通大学全球环境变化研究院, 西安 710049 | | | 韩磊 | 长安大学土地工程学院, 西安 710054 | | | 窦欣 | 南京信息工程大学地理科学学院, 南京 210044 | | | 张雷 | 西安交通大学全球环境变化研究院, 西安 710049 陕西农业发展集团有限公司, 西安 710075 | | | 王欢元 | 陕西农业发展集团有限公司, 西安 710075 | |
|
| 摘要点击次数: 715 |
| 全文下载次数: 31 |
| 中文摘要: |
| 为实现快速、高精度的农田土壤有机碳(SOC)反演并探讨不同变量筛选方法与机器学习算法的适宜组合方式,本研究以黄土高原典型丘陵沟壑区泾河流域农田土壤为对象,基于SOC采样数据,整合了包括MODIS多光谱数据及相关环境变量在内的126个变量,探究了全部变量集、竞争性自适应重加权采样法(CARS)筛选变量集、相关系数筛选变量集这3种变量集分别与支持向量机(SVM)、随机森林(RF)、反向传播神经网络(BPNN)和极端梯度提升树(XGBoost)这4种算法组合对模型精度的影响。结果表明,经过相关系数筛选后的11个特征变量与4种机器学习算法组合的模型精度显著优于全部变量集和CARS筛选变量集,其中与SVM算法的组合表现最优(验证集:R2=0.86,RMSE=0.95,MAE=0.76,RPD=2.66),其次为XGBoost、BPNN和RF模型。SHAP分析显示,地形变量(DEM、地形起伏度、剖面曲率)贡献度达28.41%,在地形复杂多变、耕地较为破碎的泾河流域农田SOC反演中起重要作用,其次为遥感指数(23.18%)和经纬度(20.21%)。相关系数筛选与SVM算法组合的反演结果表明,流域农田SOC范围为0.00~39.94 g·kg-1,其高值区集中在研究区域的西部边缘,低值区集中在东北部,中值区和低值区在中部和南部交错分布,平均SOC值为8.38 g·kg-1,其含量远低于全国农田SOC的平均水平。 |
| 英文摘要: |
| To achieve rapid and high-precision inversion of soil organic carbon(SOC)in farmland and to explore suitable combinations of different variable screening methods and machine learning algorithms, this study focused on the farmland of the Jinghe River basin, a typical hilly-gully region of the Loess Plateau. Based on the SOC sample data, this study integrated 126 variables, including MODIS multispectral data and related environmental variables. The study explored the impacts of three variable sets:the full variable set, the competitive adaptive reweighted sampling(CARS) screened variable set, and the correlation coefficient screened variable set when combined with four machine learning algorithms:support vector machine(SVM), random forest(RF), back propagation neural network (BPNN), and extreme gradient boosting(XGBoost)on model accuracy. The results showed that models based on the 11 variables selected via correlation coefficient significantly outperformed those using the full or CARS-selected variable sets. Among these combinations, the SVM algorithm performed best(validation set:R2=0.86, RMSE=0.95, MAE=0.76, RPD=2.66), followed by the XGBoost, BPNN, and RF models. The SHAP analysis revealed that topographic variables(DEM, terrain relief, profile curvature)contributed 28.41% to the inversion of farmland SOC in the Jinghe River basin, where the topography was complex and variable and the cultivated land was relatively fragmented, followed by the remote sensing index(23.18%) and latitude/longitude(20.21%). The inversion results based on the combination of correlation coefficient selection and SVM algorithm revealed that the SOC of farmland in the basin ranged from 0.00 g·kg-1 to 39.94 g·kg-1, with high values concentrated in the western part of the basin and medium-to-low values interspersed in the central and southern areas. The average SOC content was 8.38 g·kg-1, which is lower than the national average for farmland SOC. |
| HTML
查看全文
查看/发表评论 下载PDF阅读器 |
| 关闭 |
|
|
|