Paper] Development and Evaluation of Autostereoscopic 3D Live System with Real-Time Video Capture
We developed an autostereoscopic 3D live system using a minimal configuration of real-time stereo capture and a binocular autostereoscopic 3D display. We compared the displayed contents of the autostereoscopic 3D live system and conventional 2D video conferencing through both quantitative analysis and subjective evaluation. For the quantitative evaluation, we measured and assessed the fixation duration on the face area of the conversational partner using eye tracking. Our system showed longer fixation duration and a greater number of fixations on the face area compared to 2D video conferencing. For the subjective evaluation, we surveyed participants and analyzed the results. We used Wilcoxon signed-rank tests to compare our system with 2D video conferencing. The results show that our system achieves higher ratings for sense of presence and realism.
- Research Article
38
- 10.1016/j.ajodo.2016.06.037
- Jan 31, 2017
- American Journal of Orthodontics and Dentofacial Orthopedics
Role of facial attractiveness in patients with slight-to-borderline treatment need according to the Aesthetic Component of the Index of Orthodontic Treatment Need as judged by eye tracking.
- Conference Article
2
- 10.1109/icsipa.2015.7412218
- Oct 1, 2015
This paper presents the results and analysis of the objective and subjective quality evaluations of Further Reduced Resolution Depth Coding (FRRDC) method for stereoscopic 3D video. FRRDC is developed based on the Scalable Video Coding (SVC) reference software and the result are objectively evaluated using rate distortion curve and subjectively evaluated using LCD and auto-stereoscopic video displays. FRRDC uses the Down-Sampling and Up-Sampling (DSUS) method of the depth data of the stereoscopic 3D video. The emergence of numerous auto-stereoscopic displays in the market confirms the growth of 3DTV services. It is essential that the coding method of stereoscopic 3D videos produces high quality 3D videos on both stereoscopic displays and emerging auto-stereoscopic 3D video displays to ensure the interoperability and compatibility among all the different display devices. In this paper, the stereoscopic 3D videos are compressed using the H.264/SVC codec with Reduced Resolution Depth Coding (RRDC) and compared with H.264/SVC-FRRDC. The experimental results indicate good 3D depth perception of FRRDC on both stereoscopic and auto-stereoscopic display devices with lesser bit rates compared to H.264/SVC-RRDC.
- Conference Article
12
- 10.1109/iccw.2015.7247430
- Jun 1, 2015
Currently, three-dimensional (3D) video is gaining increasing popularity by providing immersive user experience. Compared against conventional 2D video, 3D video excels at bringing a “live” scene closer to the users, and/or trying to place the users in the environment of the displayed content. However, streaming 3D video sequences over the IP networks is challenging due to the impact of dynamic network conditions on user quality. Accurate objective 3D video quality assessment is critical for advanced real-time video streaming adaptation solutions. Most state-of-the-art objective 3D video quality metrics are reference-based and require access to the original 3D video sequences, which is not possible for the real-time applications. This paper proposes the extended No-reference objective 3D Video Quality Metric (eNVQM) for real time 3D video quality assessment. eNVQM establishes a correlation between network packet loss and stereoscopic 3D video quality and was tuned according to extensive subjective testing results. Performance of eNVQM is studied in comparison with two state-of-the-art objective video quality metrics: structural similarity index (SSIM) and video quality metric (VQM).
- Research Article
3
- 10.1007/s11265-013-0772-0
- Jun 18, 2013
- Journal of Signal Processing Systems
With the rapid distribution of digital video capture devices, significant videos can be captured effortlessly. The captured videos are often saved in moving pictures expert group-2 (MPEG-2) format. To prove copyright ownership, applying watermarking in MPEG-2 videos is necessary. However, little research has been devoted to the watermarking design not only for the spatial domain but also for the frequency domain and the realization of watermarking hardware. Thus, a joint data compression and watermarking system with configurable spatial and frequency domain embedding, and its very large scale integrated circuit (VLSI) architecture is presented in this paper. First, after analyzing the characteristics of videos, a novel watermarking system with two number-based keys and a shuffled image is built. It is based on the spread spectrum techniques and adaptive human visual system (AHVS). With a consideration of the cost and easiness of use, the proposed system is realized as blind detection which can dispense without the storage of the original-multimedia data. Second, the efficient VLSI architecture of our approach is designed. Various subjective and objective evaluations are performed for watermarking analysis. From the evaluation, it is realized that the system can achieve robust watermarking with high flexibility for joint data compression and low hardware complexity. Various attacks and comparisons also show the efficiency of the proposed watermarking scheme. Furthermore, the VLSI synthesis results demonstrate the high performance of the proposed architecture. Thus, the proposed system is adequate for a specific function intellectual property (IP) combined with a real-time video capture and a surveillance system.
- Conference Article
2
- 10.1109/ictai56018.2022.00133
- Oct 1, 2022
In aerospace measurement and control missions, optical equipment collects infrared images and visible images with different characteristics individually. To acquire better image quality, image fusion method is studied in this paper. Firstly, based on the idea of dense connection in image classification, a feature encoding module is designed and the resolution of the feature map is kept unchanged to preserve the position information. Secondly, a image reconstruction module based on the symmetrical reconstruction in semantic segmentation is designed to fully merge the high-level and low-level feature. Then, a lightweight encoding-feature fusion-decoding image fusion neural network with only 9 CNN layers is proposed. The experimental results demonstrated that our proposed method achieve great performance improvement whether in subjective or objective evaluation. Besides, a real-time video capture and fusion system is built based on the Black Magic video capture card Declink Duo 2 and the high performance inference toolkit tensorRT. And the built system is successfully deployed in many aerospace measurement and control missions.
- Conference Article
1
- 10.1109/cisp.2010.5646300
- Oct 1, 2010
The essential components of future e-learning systems will include real-time video capture, video transmission, and video display; a high-definition image display at the receiver in the real-time teaching depends on the promise of a highresolution video and a greater compression. This paper designed a high-performance terminal for the real-time video capture in e-learning systems; it is a terminal not only built on FPGA module and DSP parallel working system but also enhanced by improved MPEG-4 algorithm. The experiment results of this terminal demonstrated a significant effectiveness of compressing a highresolution real-time video.
- Conference Article
- 10.1109/iih-msp.2013.44
- Oct 1, 2013
This paper presents the design and implementation of a real-time 3D video capture and transfer system based on Microsoft Kinect. On the collecting layer, multiple views of texture images and depth images are acquired by 3 Kinect cameras, which are placed as a parallel array. Hole filling and noise filtering techniques are used to fill holes and smooth noise in depth image on the processing layer. On the coding layer, video stream is processed using stereo video coding or multi-view video plus depth coding methods to deal with the huge amounts of data, and then transported in network on the transport layer. At last, media stream is decoded at the decoder layer, and shown on the stereoscopic display. Experimental results show that the proposed system presented in this paper can be well effective to fill the holes and reduce the noise for depth image. Stereo video coding and multi-view video plus depth coding method are also included in this system. This system can be applied to 3D video conferences, 3D entertainment and 3D security monitoring, etc.
- Research Article
- 10.12783/dtmse/ameme2020/35520
- Apr 7, 2021
- DEStech Transactions on Materials Science and Engineering
In order to analyze the degree of people's attention to different clothing parts under the condition of clothing fit evaluation, this paper takes different fit criteria of women blouse as the research object. In this research, subjective questionnaire evaluation experiment and eye tracking experiment were conducted respectively. By analyzing the subjective evaluation results, participants have the ability to correctly evaluate the comfort of clothing and further conduct eye-tracking experimental research. Moreover, eye -tracking indexes data such as fixation duration and fixation counts of 20 areas of interest of 15 blouses with different fit degree were obtained through eye-tracking experiment. Then the mean values of eye-tracking data of different fit degree and areas of interest were compared and multivariate analysis of variance was conducted. And the heat map which can present the duration of fixation in the areas of interest was analyzed. The results showed that the total duration of fixation and the fixation counts of participants in different areas of interest were significantly different. and the degree of attention paid to the clothing parts of the women blouse was the bust, shoulder, waist and sleeve when evaluating clothing comfort with different fit criteria. This paper has scientific reference significance for improving the method of pattern making and women blouse comfort.
- Conference Article
5
- 10.1117/12.2076347
- Mar 17, 2015
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
Effective integration of 3D acquisition, reconstruction (modeling) and display technologies into a seamless systems provides augmented experience of visualizing and analyzing real objects and scenes with realistic 3D sensation. Applications can be found in medical imaging, gaming, virtual or augmented reality and hybrid simulations. Although 3D acquisition, reconstruction, and display technologies have gained significant momentum in recent years, there seems a lack of attention on synergistically combining these components into a “end-to-end” 3D visualization system. We designed, built and tested an integrated 3D visualization system that is able to capture in real-time 3D light-field images, perform 3D reconstruction to build 3D model of the objects, and display the 3D model on a large autostereoscopic screen. In this article, we will present our system architecture and component designs, hardware/software implementations, and experimental results. We will elaborate on our recent progress on sparse camera array light-field 3D acquisition, real-time dense 3D reconstruction, and autostereoscopic multi-view 3D display. A prototype is finally presented with test results to illustrate the effectiveness of our proposed integrated 3D visualization system.
- Research Article
30
- 10.1364/ao.57.00a101
- Nov 21, 2017
- Applied Optics
We study optical technologies for viewer-tracked autostereoscopic 3D display (VTA3D), which provides improved 3D image quality and extended viewing range. In particular, we utilize a technique-the so-called dynamic fusion of viewing zone (DFVZ)-for each 3D optical line to realize image quality equivalent to that achievable at optimal viewing distance, even when a viewer is moving in a depth direction. In addition, we examine quantitative properties of viewing zones provided by the VTA3D system that adopted DFVZ, revealing that the optimal viewing zone can be formed at viewer position. Last, we show that the comfort zone is extended due to DFVZ. This is demonstrated by a viewer's subjective evaluation of the 3D display system that employs both multiview autostereoscopic 3D display and DFVZ.
- Conference Article
19
- 10.1109/eurocon.2013.6624983
- Jul 1, 2013
The emerging broadband cellular technology, i.e. the Long Term Evolution (LTE), aims to support different services with high data rates and strict Quality of Service (QoS) requirements. It can thus be considered a very promising architecture for the three-dimensional (3D) video transmission. Differently from the conventional 2D video, the depth perception is the most important aspect characterizing 3D streams. It significantly influences the mobile users Quality of Experience (QoE). However, it requires the transmission of additional information as well as more bandwidth and a lower loss probability. The goal of this paper is to investigate how both 3D video formats and their average encoding rate impact on the quality experienced by the end users when the video flow is delivered through the LTE network. To this aim, objective metrics like the ratio of lost packets, Peak Signal to Noise Ratio, delay, and goodput are adopted for measuring QoS and QoE degrees. At the end of this analysis, we provide some important considerations about the LTE effectiveness for 3D video delivering.
- Research Article
16
- 10.1186/s12909-023-04851-8
- Nov 12, 2023
- BMC Medical Education
BackgroundAcquiring adequate theoretical knowledge in the field of dental radiography (DR) is essential for establishing a good foundation at the prepractical stage. Currently, nonface-to-face DR education predominantly relies on two-dimensional (2D) videos, highlighting the need for developing educational resources that address the inherent limitations of this method. We developed a virtual reality (VR) learning medium using 360° video with a prefabricated head-mounted display (pHMD) for nonface-to-face DR learning and compared it with a 2D video medium.MethodsForty-four participants were randomly assigned to a control group (n = 23; 2D video) and an experimental group (n = 21; 360° VR). DR was re-enacted by the operator and recorded using 360° video. A survey was performed to assess learning satisfaction and self-efficacy. The nonparametric statistical tests comparing the groups were conducted using SPSS statistical analysis software.ResultsLearners in the experimental group could experience VR for DR by attaching their smartphones to the pHMD. The 360° VR video with pHMD provided a step-by-step guide for DR learning from the point of view of an operator as VR. Learning satisfaction and self-efficacy were statistically significantly higher in the experimental group than the control group (p < 0.001).ConclusionsThe 360° VR videos were associated with greater learning satisfaction and self-efficacy than conventional 2D videos. However, these findings do not necessarily substantiate the educational effects of this medium, but instead suggest that it may be considered a suitable alternative for DR education in a nonface-to-face environment. However, further examination of the extent of DR knowledge gained in a nonface-to-face setting is warranted. Future research should aim to develop simulation tools based on 3D objects and also explore additional uses of 360° VR videos as prepractical learning mediums.
- Conference Article
8
- 10.1109/icip.2006.312392
- Oct 1, 2006
Three dimensional (3D) video is attracting a lot of attention as a new multimedia representation method. 3D video is a sequence of 3D models (frames) that consist of varying vertices and connectivity. In conventional 2D video compression algorithms, motion compensation (MC) using block matching algorithm is frequently employed to reduce redundancy between consecutive frames. However, there is no such technology for 3D video so far. Therefore, in this paper, we have developed an extended block matching algorithm (EMBA) to reduce temporal redundancy of geometry information of 3D video by extending the idea of 2D block matching to 3D space. In our EBMA, a cubic block is used as a matching unit and, MC is achieved efficiently by matching the mean normal vectors of the sub-blocks, which turned out to be sub-optimal by our experiments. The residual information is further transformed by discrete cosine transform (DCT) and then encoded. The extracted motion vectors are also entropy encoded. As a result of our experiments, compression ratio ranging from 10% to 18% of the original 3D video data has been achieved.
- Research Article
164
- 10.1145/1531326.1531370
- Jul 27, 2009
- ACM Transactions on Graphics
We present a set of algorithms and an associated display system capable of producing correctly rendered eye contact between a three-dimensionally transmitted remote participant and a group of observers in a 3D teleconferencing system. The participant's face is scanned in 3D at 30Hz and transmitted in real time to an autostereoscopic horizontal-parallax 3D display, displaying him or her over more than a 180° field of view observable to multiple observers. To render the geometry with correct perspective, we create a fast vertex shader based on a 6D lookup table for projecting 3D scene vertices to a range of subject angles, heights, and distances. We generalize the projection mathematics to arbitrarily shaped display surfaces, which allows us to employ a curved concave display surface to focus the high speed imagery to individual observers. To achieve two-way eye contact, we capture 2D video from a cross-polarized camera reflected to the position of the virtual participant's eyes, and display this 2D video feed on a large screen in front of the real participant, replicating the viewpoint of their virtual self. To achieve correct vertical perspective, we further leverage this image to track the position of each audience member's eyes, allowing the 3D display to render correct vertical perspective for each of the viewers around the device. The result is a one-to-many 3D teleconferencing system able to reproduce the effects of gaze, attention, and eye contact generally missing in traditional teleconferencing systems.
- Conference Article
1
- 10.1117/12.468041
- May 24, 2002
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
Auto-stereoscopic 21-inch display with eye tracking having wide viewing zone and bright image was fabricated. The image of display is projected to retinal through several optical comp onents. We calculated optical system for wider viewing zone by using Inverse-Ray Trace Method. The viewing zone of first model is 155mm (theoretical value: 161mm). We could widen viewing zone by controlling paraxial radius of curvature of spherical mirror, the distance between lenses and so on. The viewing zone of second model is 209mm. We used two spherical mirrors to obtain twice brightness. We applied eye-tracking system to the display system. Eye recogn ition is based on neural network card based on ZICS technology. We fabricated Auto-stereoscopic 21-inch display with eye tracking. We measured viewing zone based on illumination area. The viewing zone was 206mm, which was close to theoretical value. We could get twice brightness too. We could see 3D image according to position without headgear. Keywords: Autostereoscopic display, eye tracking, Optical design, Inverse ray trace, viewing zone, neural network