Observed arrival · 2026-09-04
SparkyData, a population stitched from statistics
SparkyData develops synthetic demographic datasets calibrated to public statistics, beginning with the New York metropolitan area.
Field notes
The stated workflow starts with whole households from disclosure-protected microdata, then calibrates household and person margins simultaneously rather than joining unrelated statistical tables. Additional variables are layered conditionally, and the site says records are validated against two-way joints and held-out statistics. The initial geography is the New York metropolitan area. A featured analysis contrasts 92,000 high-income renter households with alternative counts produced by gross-income and take-home-pay thresholds.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue