- Journal Title
- 10.1561/cgv
- Dec 8, 2025
- Foundations and Trends® in Computer Graphics and Vision
- Research Article
2
- 10.1561/0600000119
- Jan 1, 2025
- Foundations and Trends® in Computer Graphics and Vision
- Khai Nguyen
- Research Article
22
- 10.1561/0600000112
- Dec 18, 2024
- Foundations and Trends® in Computer Graphics and Vision
- Stanley Chan
The astonishing growth of generative tools in recent years has empowered many exciting applications in text-to-image generation and text-to-video generation. The underlying principle behind these generative tools is the concept of diffusion, a particular sampling mechanism that has overcome some longstanding shortcomings in previous approaches. The goal of this tutorial is to discuss the essential ideas underlying these diffusion models. The target audience of this tutorial includes undergraduate and graduate students who are interested in doing research on diffusion models or applying these tools to solve other problems.
- Research Article
95
- 10.1561/0600000110
- May 6, 2024
- Foundations and Trends® in Computer Graphics and Vision
- Chunyuan Li + 6 more
This monograph presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to generalpurpose assistants. The research landscape encompasses five core topics, categorized into two classes. (i) We start with a survey of well-established research areas: multimodal foundation models pre-trained for specific purposes, including two topics – methods of learning vision backbones for visual understanding and text-to-image generation. (ii) Then, we present recent advances in exploratory, open research areas: multimodal foundation models that aim to play the role of general-purpose assistants, including three topics – unified vision models inspired by large language models (LLMs), end-to-end training of multimodal LLMs, and chaining multimodal tools with LLMs. The target audiences of the monograph are researchers, graduate students, and professionals in computer vision and vision-language multimodal communities who are eager to learn the basics and recent advances in multimodal foundation models.
- Research Article
5
- 10.1561/0600000102
- Jan 1, 2024
- Foundations and Trends® in Computer Graphics and Vision
- Timnit Gebru + 1 more
The field of computer vision is now a multi-billion dollar enterprise, with its use in surveillance applications driving this large market share. In the last six years, computer vision researchers have started to discuss the risks and harms of some of these systems, mostly using the lens of fairness introduced in the machine learning literature to perform this analysis. While this lens is useful to uncover and mitigate a narrow segment of the harms that can be enacted through computer vision systems, it is only one of the toolkits that researchers have available to uncover and mitigate the harms of the systems they build. In this monograph, we discuss a wide range of risks and harms that can be enacted through the development and deployment of computer vision systems. We also discuss some existing technical approaches to mitigating these harms, as well as the shortcomings of these mitigation strategies. Then, we introduce computer vision researchers to harm mitigation strategies proposed by journalists, human rights activists, individuals harmed by computer vision systems, and researchers in disciplines ranging from sociology to physics. We conclude the monograph by listing principles that researchers can follow to build what we call community-rooted computer vision tools in the public interest, and give examples of such research directions. We hope that this monograph can serve as a starting point for researchers exploring the harms of current computer vision systems and attempting to steer the field into community-rooted work.
- Research Article
13
- 10.1561/0600000103
- Oct 30, 2023
- Foundations and Trends® in Computer Graphics and Vision
- Stanley H Chan + 1 more
Since the seminal work of Andrey Kolmogorov in the early 1940’s, imaging through atmospheric turbulence has grown from a pure scientific pursuit to an important subject across a multitude of civilian, space-mission, and national security applications. Fueled by the recent advancement of deep learning, the field is further experiencing a new wave of momentum of applying these learning-based techniques to the problem. However, because of the complexity of the physics of atmospheric turbulence, significant gaps remain to be filled before the power of deep learning can be fully unleashed. In particular, the goal of building the most accurate turbulence model to mimic nature is gradually shifted to designing a compromised model that can maximize the image reconstruction performance. This leads to a new field which this book is trying to explain, Computational Imaging Through Atmospheric Turbulence. The goal of this book is to present the basic concepts of turbulence physics while framing it under the theme of computational imaging. Emphasis is put on elaborating the principles of how waves propagate through atmospheric turbulence and propagation-free approaches to reproduce the effect without needing wave propagation equations. This allows for a much faster simulation while preserving the physics of turbulence, hence creating the possibility of integrating turbulence physics into the design of image reconstruction algorithms. The book is written for readers with an image processing background who are seeking to understand the physics of turbulence. Connections with deep learning are emphasized throughout the book.
- Research Article
56
- 10.1561/0600000107
- Jan 1, 2023
- Foundations and Trends® in Computer Graphics and Vision
- Yibo Yang + 2 more
Neural compression is the application of neural networks and other machine learning methods to data compression. Recent advances in statistical machine learning have opened up new possibilities for data compression, allowing compression algorithms to be learned end-to-end from data using powerful generative models such as normalizing flows, variational autoencoders, diffusion probabilistic models, and generative adversarial networks. This monograph aims to introduce this field of research to a broader machine learning audience by reviewing the necessary background in information theory (e.g., entropy coding, rate-distortion theory) and computer vision (e.g., image quality assessment, perceptual metrics), and providing a curated guide through the essential ideas and methods in the literature thus far.
- Research Article
53
- 10.1561/0600000095
- Oct 19, 2022
- Foundations and Trends® in Computer Graphics and Vision
- Gabriela Csurka + 2 more
Semantic image segmentation (SiS) plays a fundamental role in a broad variety of computer vision applications, providing key information for the global understanding of an image. This survey is an effort to summarize two decades of research in the field of SiS, where we propose a literature review of solutions starting from early historical methods followed by an overview of more recent deep learning methods including the latest trend of using transformers. We complement the review by discussing particular cases of the weak supervision and side machine learning techniques that can be used to improve the semantic segmentation such as curriculum, incremental or self-supervised learning. State-of-the-art SiS models rely on a large amount of annotated samples, which are more expensive to obtain than labels for tasks such as image classification. Since unlabeled data is instead significantly cheaper to obtain, it is not surprising that Unsupervised Domain Adaptation (UDA) reached a broad success within the semantic segmentation community. Therefore, a second core contribution of this monograph is to summarize five years of a rapidly growing field, Domain Adaptation for Semantic Image Segmentation (DASiS) which embraces the importance of semantic segmentation itself and a critical need of adapting segmentation models to new environments. In addition to providing a comprehensive survey on DASiS techniques, we unveil also newer trends such as multi-domain learning, domain generalization, domain incremental learning, test-time adaptation and source-free domain adaptation. Finally, we conclude this survey by describing datasets and benchmarks most widely used in SiS and DASiS and briefly discuss related tasks such as instance and panoptic image segmentation, as well as applications such as medical image segmentation. We hope that this monograph will provide researchers across academia and industry with a comprehensive reference guide and will help them in fostering new research directions in the field.
- Research Article
142
- 10.1561/0600000105
- Jan 1, 2022
- Foundations and Trends® in Computer Graphics and Vision
- Zhe Gan + 5 more
This monograph surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: (i) VLP for image-text tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding; (ii) VLP for core computer vision tasks, such as (open-set) image classification, object detection, and segmentation; and (iii) VLP for video-text tasks, such as video captioning, video-text retrieval, and video question answering. For each category, we present a comprehensive review of state-of-the-art methods, and discuss the progress that has been made and challenges still being faced, using specific systems and models as case studies. In addition, for each category, we discuss advanced topics being actively explored in the research community, such as big foundation models, unified modeling, in-context few-shot learning, knowledge, robustness, and computer vision in the wild, to name a few.
- Research Article
28
- 10.1561/0600000096
- Jan 1, 2021
- Foundations and Trends® in Computer Graphics and Vision
- Irene Amerini + 3 more
In the last two decades, we have witnessed an immense increase in the use of multimedia content on the internet, for multiple applications ranging from the most innocuous to very critical ones. Naturally, this emergence has given rise to many types of threats posed when this content can be manipulated/used for malicious purposes. For example, fake media can be used to drive personal opinions, ruining the image of a public figure, or for criminal activities such as terrorist propaganda and cyberbullying. The research community has of course moved to counter attack these threats by designing manipulation-detection systems based on a variety of techniques, such as signal processing, statistics, and machine learning. This research and practice activity has given rise to the field of multimedia forensics. The success of deep learning in the last decade has led to its use in multimedia forensics as well. In this survey, we look at the latest trends and deep-learning-based techniques introduced to solve three main questions investigated in the field of multimedia forensics. We begin by examining the manipulations of images and videos produced with editing tools, reporting the deep-learning approaches adopted to counter these attacks. Next, we move on to the issue of the source camera model and device identification, as well as the more recent problem of monitoring image and video sharing on social media. Finally, we look at the most recent challenge that has emerged in recent years: recognizing deepfakes, which we use to describe any content generated using artificial-intelligence techniques; we present the methods that have been introduced to show the existence of traces left in deepfake content and to detect them. For each problem, we also report the most popular metrics and datasets used today.