Physics-Constrained Synthetic Data Generation For Industrial Machine Learning: Enforcing Conservation And Operating-Envelope Constraints, Validating Fidelity Against Sensor Records, And Measures
DOI:
https://doi.org/10.70917/ijcisim-2026-5723Keywords:
synthetic data generation, physics-informed machine learning [9], industrial sensor data, conservation constraints, operating envelope, predictive maintenance, remaining useful life, model collapse, rare event detection, fidelity validation, train on synthetic test on realAbstract
Synthetic data generators optimise for distributional realism [4][5][6][7]. They match means, variances, and correlations, and they do not match physical law. The consequence for industrial systems is a specific and under-examined failure: a generator can produce sensor traces statistically indistinguishable from real records while violating mass and energy balance, exceeding actuator limits, or exhibiting impossible state transitions such as temperature rising with no energy input. A downstream model trained on such data learns spurious correlations rather than system dynamics, and fails silently at the rare, physically extreme operating points where reliability matters most. This paper proposes a method for generating industrial sensor data under explicit physical constraints, and specifies how it should be evaluated. It identifies four classes of constraint that an industrial generator must respect, covering conservation laws, operating-envelope bounds, monotonicity and irreversibility, and state-transition legality. It compares four enforcement mechanisms, in the loss, in the architecture, by rejection sampling, and by post-processing projection, and states what each costs, including the case against each. It specifies a fidelity validation design going beyond marginal-distribution testing to cover cross-sensor coupling, spectral content, direct measurement of constraint violation rates, and temporal dynamics. It specifies a downstream evaluation with named datasets, arms, seeds, and statistical tests. No implementation exists. No code has been written, no generator built, no data generated, and no validation performed. This is a method proposal, and the paper reports no results because there are none.