Skip to content

Degradation model validation

Held-out validation

Nine tests the model never saw, five signals each, one parameter set
model prediction lab measurement, never fitted spread across repeat cells
Rows are the nine withheld cycling conditions, cold to hot. Columns are the five scored signals. Capacity keeps the headline honest, resistance ties the model to power and heat, and the three mode signals check the mechanisms underneath. Hover any point for values, click any panel to enlarge it below.
The five conditions the fit was allowed to see
lab measurement used in calibration
Fitted curves should sit close to their own training data, and they do. Typical fit error across these five is 1.4 points of state of health. The two 25 °C, 1C discharge panels share one lab experiment, simulated once under SOC control and once under voltage control.

Calendar storage

Ageing at rest, up to two years per condition
model prediction lab, predicted blind lab, used in calibration
One panel per storage condition, cold to warm and low to high state of charge. Storage fade depends strongly on the state of charge a cell rests at, and the model tracks that at both temperatures. Every blind storage prediction lands within 1.4 points RMSE. One blind condition, 30 percent SOC at 0 °C, is scored in the table but has no published curve data. Click any panel to enlarge it above.

Accuracy in numbers

The dot is the typical error over a whole test. The tick is the single worst moment, the number to use when sizing margins.

Capacity accuracy, every condition on one axis
predicted blind used in calibration worst single error
Blind cycling conditions average 1.8 points RMSE. The hardest condition, shallow cycling at 25 °C at 3.24 points RMSE, misses in the cautious direction, predicting more fade than the lab measured. Storage stays under 1.4 points everywhere. Hover any row for its numbers.

Full metrics per condition. Each cell reads typical error and then worst single error. Click a column header to sort.

From accuracy to confidence

A single accuracy number is not enough to plan with. What matters is how wrong the model can be, in which direction, and how often. The figure below collects every blind prediction error in the study to answer exactly that. The errors centre just below zero, so when the model misses it errs on the cautious side, and across all 358 blind checkups it never overstated remaining health by more than 3.2 points.

All 358 blind prediction errors
Model minus lab at every blind checkup, in points of state of health. Solid bars are within two points. The distribution leans left of zero, which is the safe side.

Measured coverage against a textbook bell curve:

Interval around the prediction Bell curve says Measured on 358 blind checkups
RMSE as one sigma, ±1.8 pts68.3 %75.4 %
±2 pts74.3 %80.4 %
±3 pts91.1 %91.9 %
1.96 sigma, ±3.5 pts95.0 %93.9 %
±4 pts97.7 %95.5 %

Error grows with degradation depth, not with calendar time.

The error funnel, by depth of degradation
typical error (RMSE) 95 % of errors within
The cell is rated to 80 percent state of health, so the deepest bin sits beyond its rated end of life and stresses the model past the normal service window. To pick the right uncertainty band, read the bin where the model says your duty cycle will sit at the end of the warranty window, not the deepest bin.
Version fe1f5cb5