Zickler, ToddHan, Xinran Nicole Xinran2026-06-0920262026-05-122026Han, Xinran Nicole Xinran. 2026. Generative Models for Perceptually-Consistent Computer Vision. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.32701474https://dash.harvard.edu/handle/1/42740311Perceiving the three-dimensional shape and material of objects from images is a central challenge in both biological and computer vision. A single image conflates geometry, reflectance, and illumination, giving rise to fundamental ambiguities. For example, the same image can be exactly explained by a convex object lit from one direction or a concave object lit from another (the classical "convex/concave ambiguity''), or by a continuous three-parameter family of surfaces related by axial tilts and stretches (the "bas-relief ambiguity''). The behavior of the human visual system suggests it is aware of such ambiguities: For certain images, observers can experience spontaneous alternations between two or more competing interpretations---a phenomenon known as multistable perception. This suggests the human brain may maintain and sample from a distribution of plausible explanations of the visual input, avoiding over-commitment when the evidence is genuinely ambiguous. This thesis develops a generative perception framework that embraces, rather than suppresses, these ambiguities. The models we develop are, to our knowledge, the first neural network models that show emergent multistable perception across a variety of well-known ambiguous visual stimuli, despite being trained on modest numbers of synthetic images of ordinary, everyday objects. We develop our framework in three parts. First, we derive a curvature-based shape representation that is invariant to known shape-from-shading ambiguities, including the convex/concave and bas-relief ambiguities. We then propose a neural model that extracts this curvature statistic from small local image patches in a manner that is equivariant to image rotations and translations and stable under changes in texture and lighting. This establishes a robust bottom-up shape representation upon which generative inference can build. Second, we introduce a patch-based denoising diffusion model that samples multimodal distributions of surface normals from single shading images, guided by inter-patch consistency constraints. Despite its small size, the model produces multistable shape percepts for ambiguous stimuli while converging to veridical estimates for less ambiguous inputs, demonstrating that generative perception can capture the distributional structure of human shape perception. Third, we extend visual inference from estimating shape alone in static images to jointly estimating shape and materials from short videos. We introduce a conditional video diffusion model that generates diverse shape-and-material predictions while exploiting object motion cues to resolve ambiguities that persist in static scenes. Together, these contributions demonstrate that generative modeling enables perceptual systems that are efficient, ambiguity-aware, and aligned with human visual experience. More broadly, the generative perception paradigm---maintaining and updating distributions over scene properties as new evidence arrives---offers a principled foundation for world models and embodied agents that must plan and act under the pervasive uncertainty of the visual world.application/pdfenArtificial intelligenceComputer scienceGenerative Models for Perceptually-Consistent Computer VisionThesis or Dissertation2026-06-090000-0003-4448-330X