Accurate probabilistic precipitation forecasting is vital for water management and flood mitigation. Raw Ensemble Prediction Systems (EPS) often suffer from bias and under-dispersion, requiring statistical post-processing to calibrate outputs and improve reliability. This study evaluates 24-hour cumulative precipitation forecasts (24, 48, and 72-hour lead times) using Bayesian Model Averaging (BMA) and Ensemble Model Output Statistics (EMOS-CSG). Using an 8-member WRF ensemble over Iran (September 2015–February 2016), we compare two parameter estimation strategies: a “Global” approach utilizing all stations collectively, and a “Semi-local” clustering approach grouping stations with similar climatological characteristics.
Results show raw ensembles are poorly calibrated and tend to over-forecast, whereas post-processing adds substantial value, particularly at higher thresholds (>5 to >25 mm). Optimal configurations are highly model-dependent: BMA requires a Semi-local approach to mitigate under-forecasting at high accumulations, whereas EMOS-CSG achieves optimal calibration globally. Brier Skill Score (BSS) analyses indicate BMA_Global excels at the lowest threshold (>0.1 mm), while CSG_Global dominates higher thresholds. Overall, Global configurations demonstrate a consistent numerical advantage in discrimination (AUC) and accuracy (Brier Score). However, 95% confidence intervals reveal that performance differences between Global/Semi-local and BMA/EMOS-CSG approaches generally lack statistical significance.
Ultimately, CSG_Global and BMA_Semi-local emerge as the most robust frameworks. Because these findings rely on a single autumn–winter season, further evaluation using multi-season datasets is recommended to confirm generalizability for year-round operational forecasting.