Scientific Literature

Comparative evaluation of six machine learning models for multi-fuel variable compression ratio diesel engine emission prediction under leave-one-out cross-validation

Discovered On Jul 28, 2026
Primary Metric 0
The accurate prediction of exhaust emissions from compression-ignition engines fueled with alternative biofuels remains a long-standing challenge for sustainable engine development, particularly when experimental datasets are small. This study presents a systematic comparison of six machine learning (ML) models linear regression (LR), polynomial regression (degree 2), support vector regression with radial basis function kernel (SVR-RBF), random forest (RF), gradient boosting (GB), and an artificial neural network (ANN-MLP) for predicting six emission parameters (CO, HC, CO 2 , O 2 , NO x , and the air–fuel equivalence ratio λ) of a single-cylinder variable compression ratio diesel engine. The engine was operated with three neat fuels (conventional diesel, rubber seed oil biodiesel produced by two-stage acid–base transesterification, and Chlorella vulgaris microalgae biodiesel) at three compression ratios (16:1, 17:1, 18:1) and five loads (0–100% in 25% increments), forming a balanced 45-condition factorial. All models were assessed using leave-one-out cross-validation (LOOCV) with fixed literature-based hyperparameters. Gradient boosting achieved the highest mean R 2 (0.816) across the six outputs, followed by RF (0.787), Poly (0.756), LR (0.713), SVR-RBF (0.496), and ANN-MLP (− 1.758). GB topped four of six outputs (λ, O 2 , HC, CO); RF led on NO x (R 2 = 0.934), and Poly narrowly led on CO 2 (R 2 = 0.963 vs GB 0.960). The ANN-MLP yielded negative R 2 on four of six outputs, traced through architecture sensitivity sweeps and a direct LOOCV-versus-5-fold cross-validation comparison to an unfavorable parameter-to-sample ratio (~ 1.4:1) rather than to validation choice. The six emissions stratify into three predictability levels: highly predictable (λ, O 2 , CO 2 ; R 2 > 0.96), moderately predictable (NO x , HC; R 2 = 0.70–0.93), and poorly predictable (CO; R 2 < 0.40). For the present 45-sample regime, gradient boosting is recommended as the model of choice, and neural networks should be applied cautiously to small multi-fuel emission datasets.
View Raw Thread