Tracking Any Point Methods for Markerless 3D Tissue Tracking in Endoscopic Stereo Images
Abstract Minimally invasive surgery presents challenges such as dynamic tissue motion and a limited field of view. Accurate tissue tracking has the potential to support surgical guidance, improve safety by helping avoid damage to sensitive structures, and enable context-aware robotic assistance during complex procedures. In this work, we propose a novel method for markerless 3D tissue tracking by leveraging 2D Tracking Any Point (TAP) networks. Our method combines two CoTracker models, one for temporal tracking and one for stereo matching, to estimate 3D motion from stereo endoscopic images. We evaluate the system using a clinical laparoscopic setup and a robotic arm simulating tissue motion, with experiments conducted on a synthetic 3D-printed phantom and a chicken tissue phantom. Tracking on the chicken tissue phantom yielded more reliable results, with Euclidean distance errors as low as 1.1 mm at a velocity of 10 mm/s. These findings highlight the potential of TAP-based models for accurate, markerless 3D tracking in challenging surgical scenarios.
- Research Article
45
- 10.1002/jmri.21317
- Apr 18, 2008
- Journal of Magnetic Resonance Imaging
To track three-dimensional (3D) myocardial tissue motion using slice followed cine displacement encoded imaging with stimulated echoes (DENSE). Slice following (SF) has previously been developed for 2D myocardial tagging to compensate for the effect of through-plane motion on 2D tissue tracking. By incorporating SF into a cine DENSE sequence, and applying displacement encoding in three orthogonal directions, we demonstrate the ability to track discrete elements of a slice of myocardium in 3D as the heart moves through the cardiac cycle. The SF cine DENSE tracking algorithm was validated on a moving phantom, and the effects of through-plane motion on 2D cardiac strain were investigated in six healthy subjects. A through-plane tracking accuracy of 0.46 +/- 0.32 mm was measured for a typical range of myocardial motion using a rotating phantom. In vivo 3D measurements of cardiac motion were consistent with prior myocardial tagging results. Through-plane rotation in a mid-ventricularshort-axis view was shown to decrease the magnitude of the 2D end-systolic circumferential strain by 3.91 +/- 0.43% and increase the corresponding radial strain by 6.01 +/- 1.07%. Slice followed cine DENSE provides an accurate method for 3D tissue tracking.
- Research Article
11
- 10.1016/j.ejmp.2020.10.023
- Nov 11, 2020
- Physica Medica
Simulated four-dimensional CT for markerless tumor tracking using a deep learning network with multi-task learning.
- Conference Article
1
- 10.1109/icpr.2010.620
- Aug 1, 2010
This paper presents a method for 3D reconstruction of tumors for applications in laparoscopy. This uses stereo endoscopic ultrasound images, which are simultaneously recorded. To do this, the ultrasound probe is tracked throughout the stereo endoscopic images using a particle filter and an auxiliary method based on thresholding in the HSV-space is used in order to improve the tracking. Then, the 3D pose of the ultrasound probe is calculated using conformal geometric algebra. The 2D ultrasound images have been segmented using two methods: the level sets method and morphological operators, and a comparison between their performances has been done. Finally, the processed ultrasound images are compounded into a 3D volume, using the calculated ultrasound pose.
- Conference Article
- 10.1117/12.2020924
- May 23, 2013
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
A method for 3D tracking has been developed exploiting Digital Holographic Microscopy (DHM) features. In the framework of self-consistent platform for manipulation and measurement of biological specimen we use DHM for quantitative and completely label free analysis of specimen with low amplitude contrast. Tracking capability extend the potentiality of DHM allowing to monitor the motion of appropriate probes and correlate it with sample properties. Complete 3D tracking has been obtained for the probes avoiding the issue of amplitude refocusing in traditional tracking processing. Our technique belongs to the video tracking methods that, conversely from Quadrant Photo-Diode method, opens the possibility to track multiples probes. All the common used video tracking algorithms are based on the numerical analysis of amplitude images in the focus plane and the shift of the maxima in the image plane are measured after the application of an appropriate threshold. Our approach for video tracking uses different theoretical basis. A set of interferograms is recorded and the complex wavefields are managed numerically to obtain three dimensional displacements of the probes. The procedure works properly on an higher number of probes and independently from their size. This method overcomes the traditional video tracking issues as the inability to measure the axial movement and the choice of suitable threshold mask. The novel configuration allows 3D tracking of micro-particles and simultaneously can furnish Quantitative Phase-contrast maps of tracked micro-objects by interference microscopy, without changing the configuration. In this paper, we show a new concept for a compact interferometric microscope that can ensure the multifunctionality, accomplishing accurate 3D tracking and quantitative phase-contrast analysis. Experimental results are presented and discussed for in vitro cells. Through a very simple and compact optical arrangement we show how two different functionalities can be accomplished by the same optical setup, i.e. 3D tracking of micro-object and quantitative phase contrast imaging.
- Book Chapter
- 10.1007/978-3-030-34110-7_11
- Jan 1, 2019
This paper presents a novel method for real-time 3D object detection and tracking in monocular images. The method build maps of a user-specified object from a video sequence, and stores the data for 3D object detection and tracking. The main advantage of the method lies in that it does not need existing 3D models of the objects. Instead, it first detects the target object using the state-of-the-art deep learning-based object detection method, and constructs its map using visual Simultaneous Localization and Mapping (vSLAM). The maps only need to be built once and multiple maps of different objects can be stored. A fast method is proposed to recognize the object in the map with the aid of deep learning-based detection. The method needs only one camera and is robust in cluttered environment. The mode of multiple maps allows the reuse of pre-reconstructed maps. Experimental results show that accurate, fast and robust detection and tracking are achieved.
- Research Article
- 10.1101/2025.02.13.638009
- Feb 16, 2025
- bioRxiv : the preprint server for biology
Speech movements are highly complex and require precise tuning of both spatial and timing of oral articulators to support intelligible communication. These properties also make measurement of speech movements challenging, often requiring extensive physical sensors placed around the mouth and face that are not easily tolerated by certain populations such as young children. Recent progress in machine learning-based markerless facial landmark tracking technology demonstrated its potential to provide lip tracking without the need for physical sensors, but whether such technology can provide submillimeter precision and accuracy in 3D remains unknown. Moreover, it is also unclear whether such technology can be applied to track speech movements in young children. Here, we developed a novel approach that integrates Shape Preserving Facial Landmarks with Graph Attention Networks (SPIGA), a facial landmark detector, and CoTracker, a transformer-based neural network model that jointly tracks dense points across a video sequence. We further examined and validated this novel approach by assessing its tracking precision and accuracy. The findings revealed that our approach that integrates SPIGA and CoTracker was more precise (≈ 0.15 mm in standard deviation) than SPIGA alone (≈ 0.35 mm). In addition, its 3D tracking performance was comparable to electromagnetic articulography (≈ 0.29 mm RMSE against simultaneously recorded articulograph data). Importantly, the approach performed similarly well across adults and young children (i.e., 3- and 4-year-olds). Because our framework is built upon open-source pretrained models that are fully trained, it promotes accessibility and open science while saving computing resources. Furthermore, given that this framework combines a landmark detection model (SPIGA) with a tracker model (CoTracker) to improve precision/accuracy, our novel approach serves as a proof-of-concept for enhancing the performance of a wide variety of commonly used markerless tracking applications in biology and neuroscience.
- Research Article
24
- 10.1088/0031-9155/59/17/4897
- Aug 7, 2014
- Physics in Medicine & Biology
Markerless tracking of respiration-induced tumor motion in kilo-voltage (kV) fluoroscopic image sequence is still a challenging task in real time image-guided radiation therapy (IGRT). Most of existing markerless tracking methods are based on a template matching technique or its extensions that are frequently sensitive to non-rigid tumor deformation and involve expensive computation. This paper presents a kernel-based method that is capable of tracking tumor motion in kV fluoroscopic image sequence with robust performance and low computational cost. The proposed tracking system consists of the following three steps. To enhance the contrast of kV fluoroscopic image, we firstly utilize a histogram equalization to transform the intensities of original images to a wider dynamical intensity range. A tumor target in the first frame is then represented by using a histogram-based feature vector. Subsequently, the target tracking is then formulated by maximizing a Bhattacharyya coefficient that measures the similarity between the tumor target and its candidates in the subsequent frames. The numerical solution for maximizing the Bhattacharyya coefficient is performed by a mean-shift algorithm. The proposed method was evaluated by using four clinical kV fluoroscopic image sequences. For comparison, we also implement four conventional template matching-based methods and compare their performance with our proposed method in terms of the tracking accuracy and computational cost. Experimental results demonstrated that the proposed method is superior to conventional template matching-based methods.
- Research Article
2
- 10.1117/12.2216892
- Mar 25, 2016
- Proceedings of SPIE--the International Society for Optical Engineering
Scanning-beam digital x-ray (SBDX) is an inverse geometry x-ray fluoroscopy system capable of tomosynthesis-based 3D catheter tracking. This work proposes a method of dose-reduced 3D tracking using dynamic electronic collimation (DEC) of the SBDX scanning x-ray tube. Positions in the 2D focal spot array are selectively activated to create a region-of-interest (ROI) x-ray field around the tracked catheter. The ROI position is updated for each frame based on a motion vector calculated from the two most recent 3D tracking results. The technique was evaluated with SBDX data acquired as a catheter tip inside a chest phantom was pulled along a 3D trajectory. DEC scans were retrospectively generated from the detector images stored for each focal spot position. DEC imaging of a catheter tip in a volume measuring 11.4 cm across at isocenter required 340 active focal spots per frame, versus 4473 spots in full-FOV mode. The dose-area-product (DAP) and peak skin dose (PSD) for DEC versus full field-of-view (FOV) scanning were calculated using an SBDX Monte Carlo simulation code. DAP was reduced to 7.4% to 8.4% of the full-FOV value, consistent with the relative number of active focal spots (7.6%). For image sequences with a moving catheter, PSD was 33.6% to 34.8% of the full-FOV value. The root-mean-squared-deviation between DEC-based 3D tracking coordinates and full-FOV 3D tracking coordinates was less than 0.1 mm. The 3D distance between the tracked tip and the sheath centerline averaged 0.75 mm. Dynamic electronic collimation can reduce dose with minimal change in tracking performance.
- Research Article
- 10.1016/j.mex.2026.103965
- May 23, 2026
- MethodsX
Markerless 3D hand tracking for analysis of pediatric eye-hand coordination\u2606
- Research Article
14
- 10.2352/issn.2470-1173.2016.13.iqsp-212
- Feb 14, 2016
- Electronic Imaging
Endoscopic image enhancement has become a very popular research field due to the success of minimally invasive interventions and the innovation of new technological treatment and diagnosis tools such as stereoscopic laparoscopes and the wireless capsule endoscopy. In spite of the important advances achieved in terms of image processing and enhancement, only a few techniques can be adapted to stereo endoscopic images. This can be explained by the specificities of the stereo endoscopic video acquisition process, the surgical tasks artifacts and the endoscopic domain characteristics (e.g., organ textures, edges, color distribution). In this paper we present a contrast enhancement method for stereo endoscopic images taking into consideration some of these specificities, namely those of the acquired stereo images i.e. the depth information, the binocular vision and the organs boundaries/textures. The idea is to enhance the image quality by a contrast enhancement process that exploits the local image activity, the depth information and the binocular just noticeable difference (BJND) model. The results of the conducted subjective experiment show that the proposed method produces stereo endoscopic images with sharper details of the underlying tissues and organs, without introducing any halo effect or overshooting. The observers reported as well a more depth feeling and less visual fatigue when perceiving the enhanced stereo endoscopic images.
- Research Article
2
- 10.1007/s13534-013-0098-7
- Sep 1, 2013
- Biomedical Engineering Letters
Surface to surface registration of MR and endoscopic images is the key to MR and endoscopic image fusion, which will provide the surgeon with better 3D context of the surgical site in minimally invasive procedures. However, accurate reconstruction of 3D surface from stereo endoscopic images is still a challenging task especially for the surgical site with few features. In this paper, we propose a new method to reconstruct 3D surface from stereo endoscopic images. We project a gridline light pattern onto the surgical site and then use a stereo endoscope to acquire two stereo images. The major steps in the surface reconstruction process include 1) applying an automatic method of detecting region of interest, 2) applying an image intensity correction algorithm, and 3) applying a novel automatic method to match the intersection points of the gridline pattern. We have validated our proposed technique on a liver phantom and compared our method with an existing method of similar scope. Our experiment results show that our method outperforms the existing method in terms of correct matching rate (98% vs. 47%) which is an indicator of the surface reconstruction accuracy. The proposed technique has the potential to be used in clinical practice to improve image guidance in endoscope based minimally invasive procedures. This technique may also be applied to the endoscopic procedures of other organs in the abdomen, chest cavity and pelvis such as kidneys and lungs.
- Conference Article
4
- 10.1109/nssmic.2012.6551878
- Oct 1, 2012
Motion-compensated PET of awake animals has the potential to greatly improve translational neurological investigations by enabling brain function to be studied during learning tasks and complex behaviors. Previously we have demonstrated the feasibility of performing motion-compensated brain PET on rodents, obtaining the necessary head motion data using marker-based techniques. However, markerless motion tracking would simplify animal experiments and potentially provide more accurate pose estimates over a greater range of motion. Previously we have described a markerless stereo motion tracking system and associated algorithms and validated the approach in phantoms. In this work we performed a pilot study to demonstrate motion-compensated <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">18</sup> F_FDG brain imaging in an awake, unrestrained rat using head pose measurements obtained from the markerless tracking system. Motion compensation clearly worked, resulting in easily identifiable structures in the head. However, it was also obvious that considerable residual error remained after correction. Post analysis of the motion estimates indicated that the residual error was the result of occasional spurious pose estimates, most likely caused by features on non-rigid parts of the head contributing to the pose estimation. Moreover, the line-of-response rebinning used for motion correction resulted in a large proportion of lost events, leading to noisy and inconsistent projection data. The latter is avoided by using a direct list mode reconstruction. In summary, markerless tracking continues to show promise for motion-compensated imaging of awake animals, but further optimization is required to match the accuracy and consistency of marker-based tracking.
- Abstract
- 10.1016/j.ijrobp.2020.07.647
- Oct 23, 2020
- International Journal of Radiation Oncology*Biology*Physics
Simulated Four-Dimensional CT for Markerless Tumor Tracking Using Deep Learning Network With Multi-Task Learning
- Research Article
- 10.1109/jsen.2025.3627963
- Jan 1, 2025
- IEEE Sensors Journal
Frequent intrusions of Unmanned aerial vehicles (UAVs) highlight the importance of UAV 3D tracking for security and surveillance applications. Vision-based systems for UAV perception have been widely adopted in recent years, primarily focusing on UAV 2D detection. However, these systems fall short in providing accurate 3D tracking, which is crucial for effective anti-UAV operations. To address this gap, we propose a method for UAV 3D tracking based on the fusion of vision and positioning sensors. The proposed two-stage method relies solely on the positions of cameras, eliminating the need for camera attitude information, which is typically required in traditional methods. In the 2D detection stage, UAVs are detected from monocular video using the You Only Look Once version 12 (YOLOv12) model, which generates bounding boxes for each UAV. Based on these bounding boxes, the GrabCut algorithm is applied to segment the UAV contour and compute its 2D centroid position. In the 3D tracking stage, we estimate the relative camera poses using 2D UAV’s position across views and combine them with GNSS-based camera positions to reconstruct the UAV’s 3D trajectory in the world coordinate system. In a 130 × 80 × 40 m field experiment, the proposed method achieved a median error of 4.41 m in 3D tracking, outperforming the 4.93 m error of the traditional Resection-Intersection method.
- Research Article
76
- 10.1118/1.4867859
- Mar 14, 2014
- Medical Physics
Combined magnetic resonance imaging (MRI) systems and linear accelerators for radiotherapy (MR-Linacs) are currently under development. MRI is noninvasive and nonionizing and can produce images with high soft tissue contrast. However, new tracking methods are required to obtain fast real-time spatial target localization. This study develops and evaluates a method for tracking three-dimensional (3D) respiratory liver motion in two-dimensional (2D) real-time MRI image series with high temporal and spatial resolution. The proposed method for 3D tracking in 2D real-time MRI series has three steps: (1) Recording of a 3D MRI scan and selection of a blood vessel (or tumor) structure to be tracked in subsequent 2D MRI series. (2) Generation of a library of 2D image templates oriented parallel to the 2D MRI image series by reslicing and resampling the 3D MRI scan. (3) 3D tracking of the selected structure in each real-time 2D image by finding the template and template position that yield the highest normalized cross correlation coefficient with the image. Since the tracked structure has a known 3D position relative to each template, the selection and 2D localization of a specific template translates into quantification of both the through-plane and in-plane position of the structure. As a proof of principle, 3D tracking of liver blood vessel structures was performed in five healthy volunteers in two 5.4 Hz axial, sagittal, and coronal real-time 2D MRI series of 30 s duration. In each 2D MRI series, the 3D localization was carried out twice, using nonoverlapping template libraries, which resulted in a total of 12 estimated 3D trajectories per volunteer. Validation tests carried out to support the tracking algorithm included quantification of the breathing induced 3D liver motion and liver motion directionality for the volunteers, and comparison of 2D MRI estimated positions of a structure in a watermelon with the actual positions. Axial, sagittal, and coronal 2D MRI series yielded 3D respiratory motion curves for all volunteers. The motion directionality and amplitude were very similar when measured directly as in-plane motion or estimated indirectly as through-plane motion. The mean peak-to-peak breathing amplitude was 1.6 mm (left-right), 11.0 mm (craniocaudal), and 2.5 mm (anterior-posterior). The position of the watermelon structure was estimated in 2D MRI images with a root-mean-square error of 0.52 mm (in-plane) and 0.87 mm (through-plane). A method for 3D tracking in 2D MRI series was developed and demonstrated for liver tracking in volunteers. The method would allow real-time 3D localization with integrated MR-Linac systems.