Publication:

Generative Models for Perceptually-Consistent Computer Vision

Loading...
Thumbnail Image

Date

2026-05-12

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Han, Xinran Nicole Xinran. 2026. Generative Models for Perceptually-Consistent Computer Vision. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Perceiving the three-dimensional shape and material of objects from images is a central challenge in both biological and computer vision. A single image conflates geometry, reflectance, and illumination, giving rise to fundamental ambiguities. For example, the same image can be exactly explained by a convex object lit from one direction or a concave object lit from another (the classical "convex/concave ambiguity''), or by a continuous three-parameter family of surfaces related by axial tilts and stretches (the "bas-relief ambiguity''). The behavior of the human visual system suggests it is aware of such ambiguities: For certain images, observers can experience spontaneous alternations between two or more competing interpretations---a phenomenon known as multistable perception. This suggests the human brain may maintain and sample from a distribution of plausible explanations of the visual input, avoiding over-commitment when the evidence is genuinely ambiguous.

This thesis develops a generative perception framework that embraces, rather than suppresses, these ambiguities. The models we develop are, to our knowledge, the first neural network models that show emergent multistable perception across a variety of well-known ambiguous visual stimuli, despite being trained on modest numbers of synthetic images of ordinary, everyday objects.

We develop our framework in three parts. First, we derive a curvature-based shape representation that is invariant to known shape-from-shading ambiguities, including the convex/concave and bas-relief ambiguities. We then propose a neural model that extracts this curvature statistic from small local image patches in a manner that is equivariant to image rotations and translations and stable under changes in texture and lighting. This establishes a robust bottom-up shape representation upon which generative inference can build.

Second, we introduce a patch-based denoising diffusion model that samples multimodal distributions of surface normals from single shading images, guided by inter-patch consistency constraints. Despite its small size, the model produces multistable shape percepts for ambiguous stimuli while converging to veridical estimates for less ambiguous inputs, demonstrating that generative perception can capture the distributional structure of human shape perception.

Third, we extend visual inference from estimating shape alone in static images to jointly estimating shape and materials from short videos. We introduce a conditional video diffusion model that generates diverse shape-and-material predictions while exploiting object motion cues to resolve ambiguities that persist in static scenes.

Together, these contributions demonstrate that generative modeling enables perceptual systems that are efficient, ambiguity-aware, and aligned with human visual experience. More broadly, the generative perception paradigm---maintaining and updating distributions over scene properties as new evidence arrives---offers a principled foundation for world models and embodied agents that must plan and act under the pervasive uncertainty of the visual world.

Description

Other Available Sources

Research Data

Keywords

Artificial intelligence, Computer science

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories