南湖新闻网讯(通讯员 岳紫健) 近日,我校资源与环境学院史志华教授团队在Water Research发表了题为“Achieving causality-aware interpretable modeling of riverine suspended sediment via collaborative causal paths and multiscale temporal dependencies ”的学术论文,提出一种兼顾预测精度与过程解释的流域悬浮泥沙浓度建模框架。
一场降雨会带来多少泥沙,泥沙又将在何时进入河流?精准解答上述问题,是防治水土流失、保护水环境的重要基础。泥沙还会携带养分、重金属和病原体等物质进入水体。然而,多种环境因素相互作用,使泥沙输移呈现非线性和时间滞后,准确预测其变化并解释成因一直是流域研究的难点。
传统过程模型物理基础清晰,但对复杂环境变化的适应能力有限;机器学习预测能力较强,却往往只能发现变量之间的统计关联,难以区分真正的驱动因素与中间传递环节。模型即使预测准确,也未必能够为流域治理提供可靠依据。 针对这一问题,研究团队依据水文学原理,将降水和土地利用作为初始驱动因素,将径流、土壤含水量和植被状况作为中间环节,并考虑不同时间尺度的降水影响。团队利用赣江上游贡水流域6个国家水文站1991年至2020年的观测数据,在三种常用机器学习算法中进行了验证。结果表明,新框架在三种算法中均具有较好适用性。与常规模型相比,三项评价指标平均改善4.2%、6.2%和1.5%。模型识别出的降水、土地利用和径流等因素的重要性及其作用方向,也更符合实际水文过程。进一步分析发现,长时间滞后的降水信息和不透水地表是值得关注的误差来源。该研究实现机器学习模型 “预测准、解释清”,可为可信水文建模以及更具针对性的水土流失防治和水环境管理提供参考。

协同因果路径约束与多尺度时滞效应的流域水沙数据驱动建模框架
博士研究生岳紫健为论文第一作者,肖海兵副教授和史志华教授为论文共同通讯作者。本研究得到国家自然科学基金项目的资助。
论文链接:https://doi.org/10.1016/j.watres.2026.126756
【英文摘要】
Accurate sediment modeling is essential for effective water and environmental management, yet practitioners face a choice between two modeling paradigms. Process-based models are physically interpretable yet typically limited in generalizability and flexibility, while data-driven models are the opposite, leaving either paradigm alone insufficient to support robust management. To bridge these gaps, this study introduced a causality-aware framework embedding multitemporal response pathways within data-driven architecture. Guided by domain knowledge, we developed a daily suspended sediment concentration (SSC) prediction model with high accuracy and mechanistic transparency by encoding causal relationships as structural constraints and matching multi-temporal dynamics. Encoding the same causal structure into Shapley value-based attribution enhanced the reliability and plausibility of explainable attribution analysis for SSC variations and prediction uncertainty/errors. Validation across six subtropical watersheds in China using three algorithms (LightGBM/XGBoost/Random Forest) showed that the proposed framework outperformed conventional data-driven modeling, improving the accuracy across algorithms by 4.2% ± 3.0%, 6.2% ± 5.2%, and 1.5% ± 1.3% in NSE, KGE, and RMSE, respectively. Causality-guided Shapley values successfully disentangled root drivers (e.g., precipitation and land use) from mediators (e.g., soil water content and discharge). Consequently, feature importance shifted into physically meaningful patterns in terms of both overall causal contributions and directional relationships. Causality-guided error analysis broadly agreed with the SSC attribution, while identifying long-lagged precipitation and impervious surface factors as noteworthy contributors given their stronger influence on model uncertainty. This work highlights the practical value of integrating causal frameworks into hydro-geosciences for interpretable watershed management.