Computing Library › Data Systems
Data Systems

Simulation Data Management

Simulation programs generate vast, structured outputs that must be organized, described, and pruned so results stay findable and reproducible.

The output flood

A serious simulation campaign can produce enormous volumes of output: fields on grids, time histories, checkpoints, and derived scalars, across many parameter cases. Without deliberate management this becomes an undifferentiated heap where no one can find the run behind a figure. Simulation data management is the discipline that keeps the campaign navigable and its results defensible.

Organizing a campaign

Checkpoints versus results

Checkpoints let a long run resume after interruption but are often large and transient; final fields and derived quantities are the lasting results. Distinguishing them lets a program keep results while expiring checkpoints, controlling volume without losing what matters. This is a retention policy applied with physics judgment, not a blanket rule.

Deriving before discarding

Full field output may be too large to keep indefinitely, so the durable pattern computes and stores the derived quantities a figure needs, plus enough to reproduce them, before pruning the raw fields under a documented policy. What is kept must be sufficient to regenerate every published result, so pruning decisions are tied to reproducibility requirements.

In the Kronos program

Because the Kronos machines are design and simulation rather than built hardware, simulation output is the dominant scientific data today. Frozen physics inputs define the operating points, for example the breeder Q of 3.076 and 85.0 MW, and the runs that establish them are cataloged, described, and deposited with their inputs so the published basis can be reproduced. See provenance and the open published record.