An Adaptive Rule-Driven Data Quality Model for Distributed Telemetry Systems

Authors

  • Manish Singh University Of Illinois Urbana-Champaign, USA

DOI:

https://doi.org/10.70917/ijcisim-2026-4745

Keywords:

Data quality, distributed telemetry, adaptive rules, trust scoring, behavioral profiling, semantic drift, composite rules, dataset promotion

Abstract

In distributed telemetry, machines produce a constant stream of signals as input to be used for observation, experimentation, prediction, and decision-making. Static validation rules have been widely applied for data quality assessments; however, they are unsuitable for a distributed telemetry scenario where data producers have different characteristics, traffic patterns change frequently, and the information in proximity to the data stream may be dependent not only on the data format but also on the business context. The ARDQM integrates behavior-based profiling, dynamic rule generation, multi-level rule execution, and composite and multi-action custom rules, as well as trust-aware dataset promotion in a 3-layer architecture. The ARDQM framework architecture has three main components. These are telemetry intake and behavioral profiling, adaptive rule generation and multi-level evaluation, and trust scoring and decision gating. The trust score is a weighted sum of the trust dimensions for the four validation challenges of structural, behavioral, semantic, and business validations. The composite rules represent the current data, the behavioral profiling outputs, and the historical context as conjunctions. A discrete-event simulation study over 90 simulated days for 50 heterogeneous producers, six degradation scenarios (producer dropout, schema drift, semantic drift, volume spike, correlated failure, and mild semantic drift), and 30 separate runs reveals that, when compared to static rules, ARDQM reduced false positive quality alerts by 86% under the volume-spike scenario (0.38 ± 0.04 to 0.05 ± 0.01) and improved semantic drift detection from 0.10 ± 0.03 to 0.78 ± 0.03 under the semantic-drift scenario. All of these differences to the static baseline are statistically meaningful (p < 0.001, two-sided Welch's t-test, n=30 runs per condition), with 95% CIs shown. The ablation analysis shows that composite rules drive the majority of the improvement in semantic drift detection, and that trust-based graded promotion drives the majority of the improvement in reduced false positives. Because the evaluation was done entirely in discrete-event simulation (and not in a production telemetry system), all quantitative results should be regarded as points of evidence for the method's performance in simulation; they are not validated in production. Deficiencies include cold-start performance, computational overhead, risk of rule drift, and a lack of real-world validation. This contribution is both practical and a concept, such that data quality becomes an adaptive and lifecycle-bound trust function rather than a static gate, though this is yet to be empirically confirmed on production traces.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-17

How to Cite

Manish Singh. (2026). An Adaptive Rule-Driven Data Quality Model for Distributed Telemetry Systems. International Journal of Computer Information Systems and Industrial Management Applications, 18(17s), 356–374. https://doi.org/10.70917/ijcisim-2026-4745

Issue

Section

Original Articles