An accurate estimation of soil electrical conductivity (EC) using hyperspectral techniques is of great significance for understanding the spatial distribution of solutes and soil salinization. Although spectral transformation has been widely used in data pre-processing, the performance of different pre-processing techniques (or combination methods) on different models of the same data set is still ambiguous. Moreover, extremely randomized trees (ERT) and light gradient boosting machine (LightGBM) models are new learning algorithms with good generalization performance (soil moisture and above-ground biomass), but are less studied in estimating soil salinity in the visible and near-infrared spectra. In this study, 130 soil EC data, soil measured hyperspectral data, topographic factors, conventional salinity indices such as Salinity Index 1, and two-band (2D) salinity indices such as ratio indices, were introduced. The five spectral pre-processing methods of standard normal variate (SNV), standard normal variate and detrend (SNV-DT), inverse (1/OR) (OR is original spectrum), inverse-log (Log(1/OR) and fractional order derivative (FOD) (range 0–2, with intervals of 0.25) were performed. A gradient boosting machine (GBM) was used to select sensitive spectral parameters. Models (extreme gradient boosting (XGBoost), LightGBM, random forest (RF), ERT, classification and regression tree (CART), and ridge regression (RR)) were used for inversion soil EC and model validation. The results reveal that the two-dimensional correlation coefficient highlighted EC more effectively than the one-dimensional. Under SNV and the second order derivative, the two-dimensional correlation coefficient increased by 0.286 and 0.258 compared to the one-dimension, respectively. The 13 characteristic factors of slope, NDI, SI-T, RI, profile curvature, DOA, plane curvature, SI (conventional), elevation, Int2, aspect, S1 and TWI provided 90% of the cumulative importance for EC using GBM. Among the six machine models, the ERT model performed the best for simulation (R2 = 0.98) and validation (R2 = 0.96). The ERT model showed the best performance among the EC estimation models from the reference data. The kriging map based on the ERT simulation showed a close relationship with the measured data. Our study selected the effective pre-processing methods (SNV and the 2 order derivative) using one- and two-dimensional correlation, 13 important factors and the ERT model for EC hyperspectral inversion. This provides a theoretical support for the quantitative monitoring of soil salinization on a larger scale using remote sensing techniques.
Read full abstract