Publication:

Fine-Scale Structure from Coarse Observations over Irregular Domains Using Diffusion Models

Loading...
Thumbnail Image

Date

2026-06-05

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Bray, Davide. 2026. Fine-Scale Structure from Coarse Observations over Irregular Domains Using Diffusion Models. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Many scientific problems require inference at a finer scale than the scale at which data are observed. In public health, ecology, survey statistics, policy analysis, and graph-based risk modeling, responses are often available only as aggregate counts, rates, margins, or group-level labels, while the scientific target is a latent fine-scale field. This dissertation studies such problems as aggregate-supervised inverse inference: inference about fine-scale structure from observations that have passed through a known aggregation mechanism.

The central difficulty is that aggregation is many-to-one. A model may fit all observed aggregates while assigning risk, mass, labels, or intensity to the wrong fine-scale units. Fine-scale inference therefore requires more than aggregate prediction. It requires an explicit forward operator, scientifically meaningful structure, uncertainty representations that acknowledge non-identifiability, and evaluation at the scale of the latent target.

The dissertation develops this perspective through mathematical foundations, a synthesis of ecological inference, small-area estimation, weak supervision, Bayesian inverse problems, graph learning, and generative modeling, and two original methodological developments. The first, DisDiffEM, couples aggregate observations with graph-conditioned diffusion-style denoising in an approximate EM framework. The second, DiffRes, introduces a residual latent-variable formulation for graph-indexed aggregate count data. DiffRes separates predictable graph- and covariate-driven structure from lower-dimensional residual ambiguity, concentrating probabilistic and generative modeling on the part of the latent signal not already explained by structured prediction.

Empirical benchmarks distinguish aggregate agreement, fine-scale recovery, residual recovery, and uncertainty behavior. The results show that aggregate fit alone can be misleading, that the usefulness of expressive priors depends on residual dimension and sample size, and that inference-time refinement helps only when the prior and aggregate likelihood are sufficiently aligned. Overall, the dissertation argues that aggregates can support fine-scale scientific inference, but only when the model is explicit about what the data observe, what the assumptions supply, and what uncertainty remains unresolved.

Description

Other Available Sources

Research Data

Keywords

Applied mathematics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories