Height estimation of sugarcane tip cutting position based on multimodal alignment and depth image fusion
Height estimation of sugarcane tip cutting position based on multimodal alignment and depth image fusion
- Research Article
118
- 10.1109/tie.2016.2521346
- Jun 1, 2016
- IEEE Transactions on Industrial Electronics
This paper describes a robust vision-based relative-localization approach for a moving target based on an RGB-depth (RGB-D) camera and sensor measurements from two-dimensional (2-D) light detection and ranging (LiDAR). With the proposed approach, a target’s three-dimensional (3-D) and 2-D position information is measured with an RGB-D camera and LiDAR sensor, respectively, to find the location of a target by incorporating visual-tracking algorithms, depth information of the structured light sensor, and a low-level vision-LiDAR fusion algorithm, e.g., extrinsic calibration. To produce 2-D location measurements, both visual- and depth-tracking approaches are introduced, utilizing an adaptive color-based particle filter (ACPF) (for visual tracking) and an interacting multiple-model (IMM) estimator with intermittent observations from depth-image segmentation (for depth image tracking). The 2-D LiDAR data enhance location measurements by replacing results from both visual and depth tracking; through this procedure, multiple LiDAR location measurements for a target are generated. To deal with these multiple-location measurements, we propose a modified track-to-track fusion scheme. The proposed approach shows robust localization results, even when one of the trackers fails. The proposed approach was compared to position data from a Vicon motion-capture system as the ground truth. The results of this evaluation demonstrate the superiority and robustness of the proposed approach.
- Research Article
34
- 10.3390/s16101589
- Sep 27, 2016
- Sensors (Basel, Switzerland)
RGB-D sensors (sensors with RGB camera and Depth camera) are novel sensing systems that capture RGB images along with pixel-wise depth information. Although they are widely used in various applications, RGB-D sensors have significant drawbacks including limited measurement ranges (e.g., within 3 m) and errors in depth measurement increase with distance from the sensor with respect to 3D dense mapping. In this paper, we present a novel approach to geometrically integrate the depth scene and RGB scene to enlarge the measurement distance of RGB-D sensors and enrich the details of model generated from depth images. First, precise calibration for RGB-D Sensors is introduced. In addition to the calibration of internal and external parameters for both, IR camera and RGB camera, the relative pose between RGB camera and IR camera is also calibrated. Second, to ensure poses accuracy of RGB images, a refined false features matches rejection method is introduced by combining the depth information and initial camera poses between frames of the RGB-D sensor. Then, a global optimization model is used to improve the accuracy of the camera pose, decreasing the inconsistencies between the depth frames in advance. In order to eliminate the geometric inconsistencies between RGB scene and depth scene, the scale ambiguity problem encountered during the pose estimation with RGB image sequences can be resolved by integrating the depth and visual information and a robust rigid-transformation recovery method is developed to register RGB scene to depth scene. The benefit of the proposed joint optimization method is firstly evaluated with the publicly available benchmark datasets collected with Kinect. Then, the proposed method is examined by tests with two sets of datasets collected in both outside and inside environments. The experimental results demonstrate the feasibility and robustness of the proposed method.
- Research Article
7
- 10.20965/jrm.2021.p1265
- Dec 20, 2021
- Journal of Robotics and Mechatronics
We propose a robotic forklift system for stacking multiple mesh pallets. The stacking of mesh pallets is an essential task for the shipping and storage of loads. However, stacking, the placement of pallet feet on pallet edges, is a complex problem owing to the small sizes of the feet and edges, leading to a complexity in the detection and the need for high accuracy in adjusting the pallets. To detect the pallets accurately, we utilize multiple RGB-D (RGB Depth) cameras that produce dense depth data under the limitations of the sensor position. However, the depth data contain noise. Hence, we implement a region growing-based algorithm to extract the pallet feet and edges without removing them. In addition, we design the control law based on path following control for the forklift to adjust the position and orientation of two pallets. To evaluate the performance of the proposed system, we conducted an experiment assuming a real task. The experimental results demonstrated that the proposed system can achieve a stacking operation with a real forklift and mesh pallets.
- Conference Article
11
- 10.1109/eorsa.2014.6927839
- Jun 1, 2014
Microsoft Kinect is a new and robust 3D camera, which can be used for indoor scene and 3D model reconstruction. It contains an infrared projector, an infrared camera and a RGB camera, color and depth map can be captured in one scene at the same time and run at a speed of 30 frames per second. As a handhold device, Kinect is low-cost and portable compared with lidar, but its accuracy is low. The internal parameters of both infrared camera and RGB camera, as well as their relative pose, are pre-calibrated in factory, however these average values can't meet the need of high-precision applications and parameters vary from device to device. So if we want to improve the precision of Kinect, the calibration should be done at first. In this article, we attempt to use indoor control field to calibrate Kinect sensor, and get the accurate internal parameters of both cameras and their relative pose. Different from some compute vision methods, the 3d coordinate of control points should be measured in our coordinated system and one image is enough fine to compute the internal and external parameters. As Kinect can't get the IR stream and color stream simultaneously, we get the depth and color image at first, then cover the infrared projector and get the IR image, while an infrared compensation lamp is used to make the IR map clear. We separately detect the control points in the color and IR image for getting their corresponding image points, then gain their distance values from depth image based-on IR image points. Our algorithm contains four steps: 1) Initializing the internal parameters of infrared camera and RGB camera with Zhang's method. For RGB camera, collinearity equation is applied to establish the relationship between control points and their RGB image points, then estimating external parameters and refining internal parameters based-on least square adjustment. For infrared camera, transform IR image points to Kinect 3D coordinate with their distance value and initialized internal parameters to establish the point-to-point correspondence with control points, then calculate external parameters and refining internal parameters using iterative closest points method. Furthermore, we assume a model to improve the precision of distance value, in which three additional parameters should be estimated. 2)Projecting the control points to IR and RGB image coordinate separately which consider as ideal points, and regard the corresponding image points as real points, we can estimate the distortion parameters of infrared and RGB camera. 3) Iterating 1)2) steps until it is convergence. 4) Using the external parameters of both cameras to calculate their relative pose. In our experiment, 36 control points are measured, part of them (about 20) can be seen in one image. We usually use 15 points of them to estimate parameters and others consider as check points. The results show that our method can get high precision calibration parameters which projection mean square error lower than 0.5 pixels, and the mean square error of transforming depth image points to RGB image points lower than 1.0 pixels.
- Conference Article
2
- 10.1109/iwait.2018.8369699
- Jan 1, 2018
- 2018 International Workshop on Advanced Image Technology (IWAIT)
A virtual view, or a free-viewpoint video/image, gives us a video/image of an object seen from an arbitrary viewpoint in a 3D space. One typical approach uses multiview RGB videos captured by multiple RGB cameras surrounding the object. Then, a virtual view is obtained by estimating its 3D shape from the videos. This approach has difficulty in accuracy and efficiency for estimating a 3D geometry from 2D images. Recently, an RGB-D (RGB-Depth) camera is available to capture an RGB video with a depth, which is the distance from the camera to the surface of an object, per pixel. Using an RGB-D camera, the 3D shape of an object surface can be directly obtained without estimating 3D from 2D. However, a single RGB-D camera captures only the 3D shape of the surface part that the camera faces. In this research, we propose a method to efficiently render a virtual view using multiple RGB-D cameras. In our method, the 3D shapes of different surface parts captured by the respective cameras are efficiently merged according to a virtual viewpoint. Each camera is connected to a PC, and all PCs are connected to each other for parallel processing in a PC cluster network. RGB-D data captured by the cameras have to be transferred via the network to merge. Our method effectively reduces the size of RGB-D data to transfer by view-dependent pre-rendering, in which imperfect virtual views are rendered using original RGB-D data captured by the respective cameras on their PCs in parallel. This pre-rendering greatly contributes to real-time rendering of a final virtual view.
- Conference Article
- 10.1109/icvr55215.2022.9847956
- May 26, 2022
Due to the occlusion of foreground objects in the camera's perspective, some background areas have lack of depth and color data, which are reflected as blank areas on the point cloud model. Therefore, we propose a method to repair this part. In this paper, the REALSENSE D435 RGB-D (RGB-Depth) sensor is used to capture indoor environment, and the captured color and depth images that have been calibrated and denoised are generated corresponding point cloud models. Then perform OTSU threshold segmentation on the original depth image to segment the background part of the depth image that is occluded. The next step is to use traditional image inpainting criminisi algorithm to inpaint the depth and color image with the foreground removed. Therefore, the occluded part of the background of the point cloud model has been repaired. Finally, the original point cloud and the repaired point cloud are fused, and the result of the fusion is a relatively complete point cloud model.
- Research Article
2
- 10.1049/el.2019.1095
- Sep 1, 2019
- Electronics Letters
RGB-depth (RGB-D) cameras are widely used for 3D reconstruction or human-computer interaction. To simultaneously acquire colour and depth images of an object in widely separated viewpoints, multiple RGB-D cameras are required. To fuse the colour and depth information of individual RGB-D cameras in the reference frame, the RGB-D cameras need to be fully calibrated. Among many steps of extrinsic calibration between different RGB-D cameras, bundle adjustment has been either slow or error-prone in the presence of false correspondences across the views. This Letter presents an accurate and efficient bundle adjustment method based on alternating optimisation of a robust cost function. Through experiments, the authors show that the proposed method can dramatically reduce computation time while preserving state-of-the-art accuracy.
- Conference Article
1
- 10.1117/12.2639676
- Oct 19, 2022
Our laboratory proposes a method for generating elemental images based on fusion of depth images and RGB images for integrated 3D display and has achieved certain results. However, due to the defects of commercial-grade depth cameras, accurate depth estimation of target geometry cannot be accomplished. We propose a deep learning method for estimating accurate depth data of geometry from a single RGB-D image for elemental image generation. Our proposed algorithm uses a deep convolutional network to infer surface normals, object masks and occlusion boundaries from a single RGB-D image as input. Then, based on these refined depth predictions combined with the input original depth image, the depth of all pixels, including the missing pixels in the original input depth image, is solved. Through experiments with various backbones in the proposed deep learning network structure, our resulting model has better completion on deep image inpainting.
- Research Article
5
- 10.3390/rs12071142
- Apr 3, 2020
- Remote Sensing
To provide a realistic environment for remote sensing applications, point clouds are used to realize a three-dimensional (3D) digital world for the user. Motion recognition of objects, e.g., humans, is required to provide realistic experiences in the 3D digital world. To recognize a user’s motions, 3D landmarks are provided by analyzing a 3D point cloud collected through a light detection and ranging (LiDAR) system or a red green blue (RGB) image collected visually. However, manual supervision is required to extract 3D landmarks as to whether they originate from the RGB image or the 3D point cloud. Thus, there is a need for a method for extracting 3D landmarks without manual supervision. Herein, an RGB image and a 3D point cloud are used to extract 3D landmarks. The 3D point cloud is utilized as the relative distance between a LiDAR and a user. Because it cannot contain all information the user’s entire body due to disparities, it cannot generate a dense depth image that provides the boundary of user’s body. Therefore, up-sampling is performed to increase the density of the depth image generated based on the 3D point cloud; the density depends on the 3D point cloud. This paper proposes a system for extracting 3D landmarks using 3D point clouds and RGB images without manual supervision. A depth image provides the boundary of a user’s motion and is generated by using 3D point cloud and RGB image collected by a LiDAR and an RGB camera, respectively. To extract 3D landmarks automatically, an encoder–decoder model is trained with the generated depth images, and the RGB images and 3D landmarks are extracted from these images with the trained encoder model. The method of extracting 3D landmarks using RGB depth (RGBD) images was verified experimentally, and 3D landmarks were extracted to evaluate the user’s motions with RGBD images. In this manner, landmarks could be extracted according to the user’s motions, rather than by extracting them using the RGB images. The depth images generated by the proposed method were 1.832 times denser than the up-sampling-based depth images generated with bilateral filtering.
- Research Article
- 10.1504/ijwmc.2020.10030316
- Jan 1, 2020
- International Journal of Wireless and Mobile Computing
Gesture recognition is a key research field in the human-computer interaction. At present, most of researchers focus on one-handed gesture recognition, but do not pay much attention to bimanual (two hands) gesture recognition. This paper presents a deep learning-based solution to tackle the self-occlusion and self-similarity. To solve this problem, this paper uses Kinect to collect many colour and depth images of different gestures, and each gesture contains multiple sample individuals. Colour images and depth images are used to train the recognition model of bimanual gesture respectively, and then the colour image and depth image are fused, and the bimanual gesture recognition model is trained based on colour image and depth image fusion. Then, the bimanual recognition effects of the three models are compared. The experimental results show that, regardless of the single gesture precision or the mean average precision, the bimanual gesture recognition effect of the fused model is better than the gesture recognition models based on either colour image or depth image.
- Conference Article
7
- 10.18287/1613-0073-2018-2210-300-308
- Jan 1, 2018
In this paper, we propose a new method for 3D map reconstruction using the Kinect sensor based on multiple ICP. The Kinect sensor provides RGB images as well as depth images. Since the depth and RGB color images are captured by one Kinect sensor with multiple views, each depth image should be related to the color image. After matching of the images (registration), point-to-point corresponding between two depth images is found, and they can be combined and represented in the 3D space. In order to obtain a dense 3D map of the 3D indoor environment, we design an algorithm to combine information from multiple views of the Kinect sensor. First, features extracted from color and depth images are used to localize them in a 3D scene. Next, Iterative Closest Point (ICP) algorithm is used to align all frames. As a result, a new frame is added to the dense 3D model. However, the spatial distribution and resolution of depth data affect to the performance of 3D scene reconstruction system based on ICP. In this paper we automatically divide the depth data into sub-clouds with similar resolution, to align them separately, and unify in the entire points cloud. This method is called the multiple ICP. The presented computer simulation results show an improvement in accuracy of 3D map reconstruction using real data.
- Conference Article
5
- 10.1109/iccais.2014.7020538
- Dec 1, 2014
This paper describes a robust localization approach for a moving target based on RGB-depth (RGB-D) camera and 2D light detection and ranging (LiDAR) sensor measurements. In the proposed approach, the 3D and 2D position information of a target measured by RGB-D camera and LiDAR sensor, respectively are utilized to find location of target by incorporating visual tracking algorithms, depth information of the structured light sensor and vision-LiDAR low-level fusion algorithm (e.g., extrinsic calibration). For robustness of localization, a novel approach making use of Kalman prediction and filtering with intermittent observations which are identified from depth image segmentation is proposed. The proposed depth-aided localization algorithm shows robust tracking results even if visual tracking using RGB camera fails. The experimental verification results are compared to position data from VICON motion captureas a ground truth and the results show that performance superiority and robustness of the proposed approach.
- Research Article
13
- 10.3389/fpls.2023.1094142
- May 31, 2023
- Frontiers in Plant Science
Water plays a very important role in the growth of tomato (Solanum lycopersicum L.), and how to detect the water status of tomato is the key to precise irrigation. The objective of this study is to detect the water status of tomato by fusing RGB, NIR and depth image information through deep learning. Five irrigation levels were set to cultivate tomatoes in different water states, with irrigation amounts of 150%, 125%, 100%, 75%, and 50% of reference evapotranspiration calculated by a modified Penman-Monteith equation, respectively. The water status of tomatoes was divided into five categories: severely irrigated deficit, slightly irrigated deficit, moderately irrigated, slightly over-irrigated, and severely over-irrigated. RGB images, depth images and NIR images of the upper part of the tomato plant were taken as data sets. The data sets were used to train and test the tomato water status detection models built with single-mode and multimodal deep learning networks, respectively. In the single-mode deep learning network, two CNNs, VGG-16 and Resnet-50, were trained on a single RGB image, a depth image, or a NIR image for a total of six cases. In the multimodal deep learning network, two or more of the RGB images, depth images and NIR images were trained with VGG-16 or Resnet-50, respectively, for a total of 20 combinations. Results showed that the accuracy of tomato water status detection based on single-mode deep learning ranged from 88.97% to 93.09%, while the accuracy of tomato water status detection based on multimodal deep learning ranged from 93.09% to 99.18%. The multimodal deep learning significantly outperformed the single-modal deep learning. The tomato water status detection model built using a multimodal deep learning network with ResNet-50 for RGB images and VGG-16 for depth and NIR images was optimal. This study provides a novel method for non-destructive detection of water status of tomato and gives a reference for precise irrigation management.
- Conference Article
3
- 10.23919/chicc.2018.8483387
- Jul 1, 2018
Aiming at the problem that the process of gesture recognition based on color image is greatly affected by environmental factors such as lighting, a gesture intent understanding method based on the fusion of Red-Green-Blue (RGB) data and depth data is proposed. Firstly, the gesture feature extraction based on the Speeded Up Robust Feature (SURF) method after foreground segmentation are used to get gesture information. Then, we apply Backpropagation (BP) neural network to classify and recognize gestures. The final recognition results are obtained through data fusion from recognition results based on both RGB images and depth images. We evaluated the effectiveness of the proposed method through ChaLearn Gesture Database.
- Research Article
51
- 10.1016/j.wasman.2021.12.021
- Dec 23, 2021
- Waste Management
RGB-D fusion models for construction and demolition waste detection