Validation Metrics
A quantitative measure of the discrepancy between simulation and experiment, accounting for uncertainty on both sides.
Measuring the Gap
Validation requires more than a plot where the curves look close. A validation metric is a quantitative measure of the discrepancy between simulation and experiment, defined in advance and computed consistently. It turns a subjective impression of agreement into a number that can be compared, tracked, and defended.
Accounting for Both Uncertainties
A meaningful metric compares the simulation and the experiment as distributions, not points. The experiment carries measurement uncertainty; the simulation carries numerical and input uncertainty. A good metric asks whether the two overlap within their combined uncertainty, not whether their central values coincide. Two central values far apart may still agree if the uncertainties are large; two close values may disagree if the uncertainties are tight.
Kinds of Metric
- Difference metrics: the gap between simulation and experiment relative to combined uncertainty.
- Area metrics: the mismatch between the full predicted and observed distributions.
- Reliability metrics: the probability the model predictions fall within observed bounds.
Avoiding Self-Deception
The metric and the quantity of interest must be chosen before the comparison, not after seeing the data. Choosing the metric that happens to look best is a form of cheating that inflates apparent agreement. Likewise, the metric should be reported as a discrepancy with its uncertainty, not collapsed into a binary pass or fail that discards how close the result actually was.
A validation metric feeds directly into credibility assessment. It tells a design reviewer how well the model matched reality in the tested regime, and combined with the distance of the design point from that regime, it supports a judgment about how much to trust the model where no data exists. Reported honestly, it is one of the strongest pieces of evidence in a credibility case.