Signal Generating Benchmark Health Data Using a Tabular Diffusion Transformer
Summary
This preprint proposes a two-stage cross-tabular data generation framework for producing synthetic health data from multiple heterogeneous tables that have different feature sets. Earlier approaches to synthetic tabular data generation typically only work with a single input table and have difficulty handling several heterogeneous tables whose feature sets differ from one another. In the first stage, each raw table is converted into a standardized statistical table capturing marginal distributions and pairwise correlations. In the second stage, a diffusion transformer is trained to generate synthetic statistical tables, from which synthetic raw tables are reconstructed using multivariate Gaussian sampling and an inverse probability integral transform. This two-stage framework learns a unified generative model from multiple heterogeneous tables and can generate an unlimited number of realistic synthetic tables. The authors report experiments showing high fidelity in the learned statistical representations and a favorable fidelity-diversity trade-off in the generated synthetic data.
Classification
Evidence 1
- Generating Benchmark Health Data Using a Tabular Diffusion Transformer arXiv (cs.AI) 2026-08-14 accessed 2026-08-20T05:08:17+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-8b704b9a72a9
