Design of Experiments for Surrogates
Design of experiments chooses where to sample an expensive model so a surrogate learns the most from the fewest runs.
The sampling problem
Every surrogate is only as good as the points it was trained on. Design of experiments (DoE) is the discipline of placing those points deliberately rather than randomly, so that a fixed sampling budget yields the most accurate and least biased surrogate.
Classical versus space-filling designs
Classical designs - full factorial, fractional factorial, central composite, Box-Behnken - were built for polynomial response surfaces and physical experiments with replication and noise. Computer experiments are deterministic: the same input gives the same output every time, so replication adds nothing. For deterministic simulators the goal shifts to space-filling designs that spread points evenly across the whole domain.
Space-filling families
- Latin hypercube sampling stratifies each input dimension so every level is represented once
- Maximin and minimax designs maximize the minimum distance between points to avoid clustering
- Low-discrepancy sequences (Sobol, Halton) fill space with near-uniform coverage
- Orthogonal-array-based designs control projections onto low-dimensional subspaces
Criteria for a good design
Common optimality criteria include D-optimality (minimize the volume of the coefficient confidence ellipsoid), I-optimality (minimize average prediction variance), and maximin distance. For Gaussian-process surrogates, integrated mean-squared prediction error is a natural target because it ties the design directly to predictive quality.
One-shot versus sequential
A one-shot design fixes all sample locations before any run. A sequential or adaptive design runs an initial batch, fits a surrogate, and then chooses the next point where the surrogate is most uncertain or where an optimum is likely. Sequential designs are usually more efficient because they let observed structure guide later sampling.
Practical guidance
A rough starting budget for a smooth response is roughly ten samples per input dimension, refined by validation error. High-dimensional problems demand dimension reduction or screening first, because uniform coverage becomes impossible as dimensions grow. In Kronos parameter studies of the Hyperion breeder, an initial Latin hypercube over field, current, and shaping variables seeds the surrogate, and adaptive points are added where predictive variance is largest.