Class-incremental visual scene understanding for multi-stage construction via sequential knowledge distillation and exemplar replay
Class-incremental visual scene understanding for multi-stage construction via sequential knowledge distillation and exemplar replay
- Book Chapter
4
- 10.1093/acrefore/9780190236557.013.883
- Mar 22, 2023
- Oxford Research Encyclopedia of Psychology
A visual scene is a visual depiction of a real-world environment that supports human activity. The human visual system has evolved through viewing scenes, so studying scene perception allows researchers to understand how the visual system responds to the stimuli it was optimized for. Visual cognition research has established that scenes form a natural kind that is separate from both object recognition and event cognition. Visual scenes have their own category structure as well as unique neural substrates. Visual scene understanding is highly contextual, and observers use regularities among objects and between objects and scenes to leverage scene and object recognition and visual search. In this way, scene understanding forms an interface between pure visual processing and other cognitive processes. A hallmark of human scene understanding is that it can be performed incredibly rapidly—scenes can be easily understood with tens of milliseconds of exposure, even when masked, and neural correlates of scene understanding emerge before 250 ms after stimulus onset. Memory for visual scenes, though outstanding, is biased toward expanding scene representations in space and time. In all, scene understanding is a core component of human cognition that forms an interface between the mind and the world.
- Book Chapter
1
- 10.1007/978-3-319-14442-9_51
- Jan 1, 2015
Visual information understanding is known as one of the most difficult and challenging problems in the realization of machine intelligence. This paper presents research issues and overview of the current state of the art in the general flow of visual information understanding. In general, the first stage of the visual understanding starts from the object segmentation. Using the saliency map based on human visual attention model is one of the most promising methods for object segmentation. The next step is scene understanding by analyzing semantics between objects in a scene. This stage finds description of image data with a formatted text. The third step requires space understanding and context awareness using multi-view analysis. This step helps solving general occlusion problem very easily. The final stage is time series analysis of scenes and a space. After this stage, we can obtain visual information from a scene, a series of scenes, and space variations. Various technologies for visual understanding already have been tried and some of them are matured. Therefore, we need to leverage and integrate those techniques properly from the perspective of higher visual information understanding.Keywordssegmentationscene understandingspace understandingcontext awarenesstime series analysis
- Research Article
3
- 10.1007/s11263-022-01599-4
- Apr 28, 2022
- International Journal of Computer Vision
We introduce the first approach to solve the challenging problem of automatic 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant (approx 40%) improvement in semantic segmentation, reconstruction and scene flow accuracy. In addition to the evaluation on several indoor and outdoor scenes, the proposed joint 4D scene understanding framework is applied to challenging outdoor sports scenes in the wild captured with manually operated wide-baseline broadcast cameras.
- Dissertation
2
- 10.32657/10356/182101
- Jan 1, 2024
With the rapid development of industry and intelligent systems, semantic scene understanding has become essential for robotic vision in smart manufacturing. Robots have significantly advanced modern manufacturing by enabling high-quality, efficient production, extended operation durations, and work in hazardous environments. Robotic techniques have automated many processes in production lines. However, in flexible production scenarios, certain tasks cannot yet be fully handled by robots and still require human involvement. This limitation is usually caused by robots' lack of semantic understanding of the target objects in the working environment. Developing visual scene understanding techniques can enable robots to accurately recognize and localize objects or regions in visual scenes at the pixel level. These techniques greatly enhance the capability and flexibility of robots in the manufacturing industry and in various general robotic applications. Consequently, human effort in the production pipeline can be largely replaced by robots with visual understanding capabilities. This research mainly focuses on the task of 3D instance segmentation, which aims to predict both semantic and instance labels for each point in point clouds. This is a fundamental and challenging task for scene understanding, with a variety of real-world applications, such as indoor robots, autonomous driving, drones, AR/VR devices, etc. We propose five different novel methods, including one fully supervised method, two weakly supervised methods, one zero-shot method, and an augmentation method to enhance model generalization. In Chapter 3, we propose a novel proposal-free fully supervised method as Regional Purity Guide Network(RPGN). We define a novel concept of regional purity, which encodes instance-aware contextual information of the surrounding region. We also propose a pretraining pipeline for learning regional purity and design rules to generate random toy scenes by extracting samples from existing training data. Using regional purity can simultaneously prevent under-segmentation and over-segmentation problems during clustering. Although scene understanding has achieved remarkable success with deep learning techniques, it remains largely unsolved. One critical bottleneck is the significant human effort required for pixel-level labeling. To address this issue, in Chapter 4, we propose a novel weakly supervised method, RWSeg, that requires labeling only one object with a single point. Using these sparse weak labels, we introduce a unified framework with two branches to propagate semantic and instance information to unannotated regions, leveraging self-attention and random walk. Furthermore, we propose a Cross-graph Competing Random Walks (CGCRW) algorithm which encourages competition among different instance graphs to resolve ambiguities in closely positioned objects and improve the performance on instance assignment. In Chapter 5, we propose the first weakly-supervised 3D instance segmentation method that only needs categorical semantic labels as supervision, and we do not need instance-level labels. Even without having any instance-related ground-truth, we design an approach to break point clouds into raw fragments and find the most confident samples for learning instance centroids. In addition, we build a recomposed dataset to learn our defined multilevel shape-aware objectness signal. An asymmetrical object inference algorithm is followed to process core points and boundary points with different strategies, and generate high-quality pseudo instance labels to guide iterative training. In the current era dominated by large foundation models, these expansive vision models adeptly capture knowledge from vast, broad datasets, enabling them to execute zero-shot segmentation on previously unseen data. In Chapter 6, we delve into leveraging various 2D foundation models to address the challenges of 3D segmentation tasks. Our approach begins by generating initial predictions of 2D semantic masks using diverse large foundation models. These mask predictions, obtained from different frames of RGB-D video sequences, are then projected into 3D space. To produce robust 3D semantic pseudo labels, we introduce a semantic label fusion strategy that effectively combines all results through voting. Our investigation encompasses various scenarios, including zero-shot learning and limited guidance from sparse 2D point labels, allowing us to evaluate the strengths and limitations of different vision foundation models. Data augmentation is essential in deep learning for improving model generalization and robustness. While standard methods like rotations and flips have been common, they often lack high-level diversity. In Chapter 7, we explore a novel approach to automatically generate 3D labeled training data. By utilizing diffusion models and chatGPT generated text prompts, we generate diverse 2D images of single objects with various structures and appearances. Beyond texture augmentation, our method automatically alters object shapes within these images. These augmented images are then transformed into 3D objects, and virtual scenes are constructed through random composition. This approach efficiently produces a substantial amount of 3D scene data without relying on real data, offering significant advantages in addressing few-shot learning challenges and mitigating long-tailed class imbalances. Our work contributes to enhancing 3D data diversity and advancing model capabilities in scene understanding tasks.
- Research Article
2
- 10.1117/1.jei.32.3.033002
- May 3, 2023
- Journal of Electronic Imaging
With the continuous development of intelligent unmanned aerial vehicles reconnaissance, single-object detection or semantic segmentation can no longer meet the diversified requirements, while simple model stacking will cause the model to be too complex, which will seriously affect the real-time running effect. We propose a visual scene understanding algorithm based on a multitask learning network with an encoder–decoder structure. First, the efficient classification network VoVNet is selected as the feature-sharing network to obtain multiscale coding features. Second, based on the one-stage anchor-free object detector, the feature screening supplementary module implements feature interaction with segmented characters to enhance the detection capability of potential targets. Then for semantic segmentation and depth estimation pixel-level classification tasks, a cascaded chained residual pooling module is used as a parameter-sharing decoder, through the parameter-sharing mechanism to reduce the repeated decoding process to ensure the running speed. Finally, to improve the generalization ability of the model, a general-purpose dataset was constructed and network training was carried out based on the idea of knowledge distillation. Experiments on the dataset show that the performance of the multitask learning network can reach the mainstream algorithm level in semantic segmentation and depth estimation branches, whereas the object detection branch is better in multitarget recall rate, and the average time of each subtask is 39 ms, which meets the requirements of real-time performance and proves the effectiveness of the multitask network.
- Conference Article
- 10.1109/avss.2013.6636604
- Aug 1, 2013
Summary form only given. Inspired by the ability of humans to interpret and understand 3D scenes nearly effortlessly, the problem of 3D scene understanding has long been advocated as the "holy grail" of computer vision. In the early days this problem was addressed in a bottom-up fashion without enabling satisfactory or reliable results for scenes of realistic complexity. In recent years there has been considerable progress on many sub-problems of the overall 3D scene understanding problem. As the performance for these sub-tasks starts to achieve remarkable performance levels, we argue that the problem to automatically infer and understand 3D scenes should be addressed again. In this talk we will - on the one hand - highlight progress on some essential components of scene understanding such as object class recognition and articulated pose estimation and tracking. On the other hand, we will also report on our current attempt towards 3D scene understanding in the particular case of traffic scene analysis.
- Supplementary Content
6
- 10.0253/tuprints-00002237
- Jul 12, 2010
Automatic visual scene understanding is one of the ultimate goals in computer vision and has been in the field’s focus since its early beginning. Despite continuous effort over several years, applications such as autonomous driving and robotics are still unsolved and subject to active research. In recent years, improved probabilistic methods became a popular tool for current state-of-the-art computer vision algorithms. Additionally, high resolution digital imaging devices and increased computational power became available. By leveraging these methodical and technical advancements current methods obtain encouraging results in well defined environments for robust object class detection, tracking and pixel-wise semantic scene labeling and give rise to renewed hope for further progress in scene understanding for real environments. This thesis improves state-of-the-art scene understanding with monocular cameras and aims for applications on mobile platforms such as service robots or driver assistance for automotive safety. It develops and improves approaches for object class detection and semantic scene labeling and integrates those into models for global scene reasoning which exploit context at different levels. To enhance object class detection, we perform a thorough evaluation for people and pedestrian detection with the popular sliding window framework. In particular, we address pedestrian detection from a moving camera and provide new benchmark datasets for this task. As frequently used single-window metrics can fail to predict algorithm performance, we argue for application-driven image-based evaluation metrics, which allow a better system assessment. We propose and analyze features and their combination based on visual and motion cues. Detection performance is evaluated systematically for different feature-classifiers combinations which is crucial to yield best results. Our results indicate that cue combination with complementary features allow improved performance. Despite camera ego-motion, we obtain significantly better detection results for motion-enhanced pedestrian detectors. Realistic onboard applications demand real-time processing with frame rates of 10 Hz and higher. In this thesis we propose to exploit parallelism in order to achieve the required runtime performance for sliding window object detection. In a case study we employ commodity graphics hardware for the popular histograms of oriented gradients (HOG) detection approach and achieve a significant speed-up compared to a baseline CPU implementation. Furthermore, we propose an integrated dynamic conditional random field model for joint semantic scene labeling and object detection in highly dynamic scenes. Our model improves semantic context modeling and fuses low-level filter bank responses with more global object detections. Recognition performance is increased for object as well as scene classes. Integration over time needs to account for different dynamics of objects and scene classes but yields more robust results. Finally, we propose a probabilistic 3D scene model that encompasses multi-class object detection, object tracking, scene labeling, and 3D geometric relations. This integrated 3D model is able to represent complex interactions like inter-object occlusion, physical exclusion between objects, and geometric context. Inference in this model allows to recover 3D scene context and perform 3D multi-object tracking from a mobile observer, for objects of multiple categories, using only monocular video as input. Our results indicate that our joint scene tracklet model for the evidence collected over multiple frames substantially improves performance. All experiments throughout this thesis are performed on challenging real world data. We contribute several datasets that were recorded from moving cars in urban and sub-urban environments. Highly dynamic scenes are obtained while driving in normal traffic on rural roads. Our experiments support that joint models, which integrate semantic scene labeling, object detection and tracking, are well suited to improve the individual stand-alone tasks’ performance.
- Research Article
1
- 10.11591/ijai.v13.i1.pp23-34
- Mar 1, 2024
- IAES International Journal of Artificial Intelligence (IJ-AI)
<span lang="EN-US">Image captioning has been widely studied due to its ability in a visual scene understanding. Automatic visual scene understanding is useful for remote monitoring system and visually impaired people. Attention-based models, including transformer, are the current state-of-the-art architectures used in developing image captioning model. This study examines the works in the development of image captioning model, especially models that are developed based on attention mechanism. The architecture, the dataset, and the evaluation metrics analysis are done to the collected works. A general flow of image captioning model development is also presented. The literature search process carried out on Google Scholar. There are 36 literatures used in this study, including a specific image captioning development in Indonesian. It is done to take one point of view of image captioning development in a low resource language. Studies using transformer model generally achieves higher evaluation metric scores. In our finding, the highest evaluation scores on the consensus-based image description evaluation (CIDEr) c5 and c40 metrics are 138.5 and 140.5 respectively. This study gives a baseline on future development of image captioning model and brings the general concept of the image captioning development process including a picture of the development in low resource language.</span>
- Conference Article
21
- 10.1109/iros.2011.6094596
- Sep 1, 2011
We propose a novel human-robot-interaction framework for robust visual scene understanding. Without any a-priori knowledge about the objects, the task of the robot is to correctly enumerate how many of them are in the scene and segment them from the background. Our approach builds on top of state-of-the-art computer vision methods, generating object hypotheses through segmentation. This process is combined with a natural dialog system, thus including a `human in the loop' where, by exploiting the natural conversation of an advanced dialog system, the robot gains knowledge about ambiguous situations. We present an entropy-based system allowing the robot to detect the poorest object hypotheses and query the user for arbitration. Based on the information obtained from the human-robot dialog, the scene segmentation can be re-seeded and thereby improved. We present experimental results on real data that show an improved segmentation performance compared to segmentation without interaction.
- Conference Article
11
- 10.1109/iccvw54120.2021.00442
- Oct 1, 2021
We present an unsupervised adaptation approach for visual scene understanding in unstructured traffic environments. Our method is designed for unstructured real-world scenarios with dense and heterogeneous traffic consisting of cars, trucks, two-and three-wheelers, and pedestrians. We describe a new semantic segmentation technique based on unsupervised domain adaptation (DA), that can identify the class or category of each region in RGB images or videos. We also present a novel self-training algorithm for multi-source DA that improves the accuracy. Our overall approach is a deep learning-based technique and consists of an unsupervised neural network that achieves 87.18% accuracy on the challenging India Driving Dataset. Our method works well on roads that may not be well-marked or may include dirt, unidentifiable debris, potholes, etc. A key aspect of our approach is that it can also identify objects that are encountered by the model for the fist time during the testing phase. We compare our method against the state-of-the-art methods and show an improvement of 5.17% − 42.9%. Furthermore, we also conduct user studies that qualitatively validate the improvements in visual scene understanding of unstructured driving environments. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>
- Conference Article
11
- 10.1109/iccv.2019.01052
- Oct 1, 2019
We introduce the first approach to solve the challenging problem of unsupervised 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant (approx 40%) improvement in semantic segmentation, reconstruction and scene flow accuracy.
- Conference Article
1
- 10.1109/wsc57314.2022.10015401
- Dec 11, 2022
Deep neural networks (DNNs) have become a driving factor of visual scene understanding. However, the shortage of construction training images has been a major barrier to fully leverage its maximum performance potential. To address this issue, we investigate the effectiveness of synthetic images on DNN training in a common real-world scenario where only a small, biased real training image dataset is available. To this end, we synthetize numerous construction training images and conduct a DNN training experiment in real construction settings. Results show that the combined dataset-trained model always outperforms the one trained with only a small, biased real dataset. This finding indicates that an image synthetization approach has promising potential to enhance a given real training dataset in terms of data quantity and diversity. Image synthetization with automated labeling will mitigate the training image shortage, contributing to the development of more accurate and scalable DNNs for construction scene understanding.
- Research Article
11
- 10.1167/15.12.571
- Sep 1, 2015
- Journal of Vision
Research on visual scene understanding has identified a number of regions involved in processing natural scenes, but has lacked a unifying framework for understanding how these different regions are organized and interact. We propose a new organizational principle, in which scene processing relies on two distinct networks at the edge of visual cortex. The first network consists of the Transverse Occipital Sulcus (TOS, or the Occipital Place Area) and the posterior portion of the Parahippocampal Place Area (PPA). These regions have a well-defined retinotopic organization and do not show strong memory or context effects, suggesting that this network primarily processes visual features from the current view of a scene. The second network consists of the caudal Inferior Parietal Lobule (cIPL), Retrosplenial Cortex (RSC), and the anterior portion of the PPA. These regions are involved in a wide range of both visual and non-visual tasks involving episodic memory, navigation, imagination, and default mode processing, and connect information about a current scene view with a much broader temporal and spatial context. We provide evidence for this division from a diverse set of sources. Using a data-driven approach to parcellate resting-state fMRI data, we identify coherent functional regions corresponding to scene-processing areas. We then show that a network clustering analysis separates these scene-related regions into two adjacent networks, which exhibit sharp changes in connectivity properties across their narrow border. Additionally, we argue that the cIPL has been previously overlooked as a critical region for full scene understanding, based on a meta-analysis of previous functional studies as well as diffusion tractography results showing that cIPL is well-positioned to connect visual cortex with many other cortical systems. This new framework for understanding the neural substrates of scene processing bridges results from many lines of research, and makes specific predictions about functional properties of these regions. Meeting abstract presented at VSS 2015
- Conference Article
9
- 10.1109/iros40897.2019.8967576
- Nov 1, 2019
Virtual borders are an opportunity to allow users the interactive restriction of their mobile robots' workspaces, e.g. to avoid navigation errors or to exclude certain areas from working. Currently, works in this field have focused on human-robot interaction (HRI) methods to restrict the workspace. However, recent trends towards smart environments and the tremendous progress in semantic scene understanding give new opportunities to enhance the HRI-based methods. Therefore, we propose a novel learning and support system (LSS) to support users during teaching of virtual borders. Our LSS learns from user interactions employing methods from visual scene understanding and supports users through recommendations for interactions. The bidirectional interaction between the user and system is realized using augmented reality. A validation of the approach shows that the LSS robustly recognizes a limited set of typical areas for virtual borders based on previous user interactions (F1 - Score= 91.5%) while preserving the high accuracy of standard HRI-based methods with a median of Mdn= 84.6%. Moreover, this approach allows the reduction of the interaction time to a constant mean value of M = 2 seconds making it independent of the border length. This avoids a linear interaction time of standard HRI-based methods.
- Research Article
325
- 10.1109/access.2021.3090981
- Jan 1, 2021
- IEEE Access
Visual scene understanding is the core task in making any crucial decision in any computer vision system. Although popular computer vision datasets like Cityscapes, MS-COCO, PASCAL provide good benchmarks for several tasks (e.g. image classification, segmentation, object detection), these datasets are hardly suitable for post disaster damage assessments. On the other hand, existing natural disaster datasets include mainly satellite imagery which has low spatial resolution and a high revisit period. Therefore, they do not have a scope to provide quick and efficient damage assessment tasks. Unmanned Aerial Vehicle (UAV) can effortlessly access difficult places during any disaster and collect high resolution imagery that is required for aforementioned tasks of computer vision. To address these issues we present a high resolution UAV imagery, FloodNet, captured after the hurricane Harvey. This dataset demonstrates the post flooded damages of the affected areas. The images are labeled pixel-wise for semantic segmentation task and questions are produced for the task of visual question answering. FloodNet poses several challenges including detection of flooded roads and buildings and distinguishing between natural water and flooded water. With the advancement of deep learning algorithms, we can analyze the impact of any disaster which can make a precise understanding of the affected areas. In this paper, we compare and contrast the performances of baseline methods for image classification, semantic segmentation, and visual question answering on our dataset. FloodNet dataset can be downloaded from here: https://github.com/BinaLab/FloodNet-Supervised_v1.0.