Visualization of Subjects and Their Motion in Video Using Principal Component Analysis
In recent years, videos have become an integral part of daily life and there is an enormous volume of video content. As a result, finding specific videos or scenes within this vast amount of content has become a difficult task. In order to enable quick scene recognition by visualizing both video content and relationships between frames, the authors introduce a method for visualizing the subjects and their motion by arranging video frames in a spatial layout based on their correlations. This method calculates frame correlation matrix by applying Principal Component Analysis (PCA) and arranges the video frames in a spatial layout based on the principal component scores. From the visualization results by scatter plots and 3D space, they found that the scatter plot of minimal frames and 3D space arrangement of video frames were more effective than traditional video playback for scene recognition and workload reduction.
- Research Article
20
- 10.1366/0003702884430443
- Aug 1, 1988
- Applied Spectroscopy
In NIR spectroscopy, multidimensional analyses such as Principal Component Analysis (PCA) may be applied to examine the similarity between spectra of natural products. However, such an approach is often limited by the effect of spectral interference due to water or particle size distribution of the samples. In the present work, the advantage of the elimination of such spectral interference before performing PCA was investigated. Unwanted component spectra were eliminated by a least-squares procedure. They were first orthogonalized and normalized by the Gram-Schmidt orthogonalization method. The subtraction coefficients were then assessed, similarly to principal component (PC) scores, by projection of the NIR spectra on the orthogonalized component spectra, and PCA was performed on the corrected spectra. This method was applied on an illustrative collection of wheat semolina conditioned in three levels of water content. Water was the component to be eliminated and had been previously modeled by two spectral patterns. These spectral patterns were used as the unwanted component spectra. PCA was applied independently before and after spectral correction of the collection of spectra and graphs obtained by the two procedures were compared. The squared correlation coefficient of the 3 first PC scores with water content was 0.979 before correction, with the 3 groups of water content appearing clearly on PCA graphs. After correction, the corresponding squared correlation coefficient for the 7 first PC scores was 0.016. PCA graphs obtained with corrected spectra also showed that the water effect was completely eliminated. At this moment, samples were separated according to their technological nature. The procedure developed may be useful in pattern recognition study and for automatic clustering of NIR spectra. It may also be applied in fields other than NIR spectroscopy.
- Research Article
1
- 10.4103/jfmpc.jfmpc_2077_22
- Apr 1, 2023
- Journal of Family Medicine and Primary Care
The Japanese government has promoted policies ensuring standardized medical care across the secondary medical care areas (SMCAs); however, these efforts have not been evaluated, making the current conditions unclear. Multidimensional indicators could identify these differences; thus, this study examined the regional characteristics of the medical care provision system for 21 SMCAs in Hokkaido, Japan, and the changes from 1998 to 2018. This study evaluated the characteristics of SMCAs by principal component analysis using multidimensional data related to the medical care provision system. Factor loadings and principal component scores were calculated, with the characteristics of each SMCA visually expressed using scatter plots. Additionally, data from 1998 to 2018 were analyzed to clarify the changes in SMCAs' characteristics. The primary and secondary principal components were Medical Resources and Geographical Factors, respectively. The Medical Resources components included the number of hospitals, clinics, and doctors, and an area's population of older adults, accounting for 65.28% of the total variance. The Geographical Factors components included the number of districts without doctors and the population and a land area of these districts, accounting for 23.20% of the variance. The accumulated proportion of variance was 88.47%. From 1998 to 2018, the area with the highest increase in Medical Resources was Sapporo, with numerous initial medical resources (-9.283 to -10.919). Principal component analysis summarized multidimensional indicators and evaluated SMCAs in this regional assessment. This study categorized SMCAs into four quadrants based on Medical Resources and Geographical Factors. Additionally, the difference in principal component scores between 1998 and 2018 emphasized the expanding gap in the medical care provision system among the 21 SMCAs.
- Research Article
17
- 10.1155/2020/2025072
- Jan 1, 2020
- Advances in Materials Science and Engineering
The types of crude oil for producing asphalt have a decisive influence on various performance measures (including aging resistance and durability) of asphalt. To discriminate and predict the crude oil source of different asphalt samples, a discrimination model was established using 12 greatly different infrared (IR) characteristic absorption peaks (CAPs) as predictive variables. The model was established based on diverse fingerprint recognition technologies (such as principal component analysis (PCA) and multivariate logistic regression analysis) by using attenuated total reflectance‐Fourier transform infrared spectroscopy (ATR‐FTIR). In this way, the crude oil source of different asphalt samples can be effectively discriminated. At first, by using PCA, the 12 CAPs in the IR spectra of asphalt samples were subjected to dimension reduction processing to control the variables of key factors. Moreover, the scores of various principal components in asphalt samples were calculated. Afterwards, the scores of principal components were analysed through modelling based on multivariate logistic regression analysis to discriminate and predict the crude oil source of different asphalt samples. The result showed that the logistic regression model shows a favourable goodness of fit, with the prediction accuracy reaching 93.9% for the crude oil source of asphalt samples. The method exhibits some outstanding advantages (including ease of operation and high accuracy), which is important when controlling the source and quality and improving the performance of asphalt.
- Research Article
30
- 10.3390/ijerph14091065
- Sep 1, 2017
- International Journal of Environmental Research and Public Health
Groundwater drinking water supply surveillance data were accessed to summarize water quality delivered as public and private water supplies in southern Saskatchewan as part of an exposure assessment for epidemiologic analyses of associations between water quality and type 2 diabetes or cardiovascular disease. Arsenic in drinking water has been linked to a variety of chronic diseases and previous studies have identified multiple wells with arsenic above the drinking water standard of 0.01 mg/L; therefore, arsenic concentrations were of specific interest. Principal components analysis was applied to obtain principal component (PC) scores to summarize mixtures of correlated parameters identified as health standards and those identified as aesthetic objectives in the Saskatchewan Drinking Water Quality Standards and Objective. Ordinary, universal, and empirical Bayesian kriging were used to interpolate arsenic concentrations and PC scores in southern Saskatchewan, and the results were compared. Empirical Bayesian kriging performed best across all analyses, based on having the greatest number of variables for which the root mean square error was lowest. While all of the kriging methods appeared to underestimate high values of arsenic and PC scores, empirical Bayesian kriging was chosen to summarize large scale geographic trends in groundwater-sourced drinking water quality and assess exposure to mixtures of trace metals and ions.
- Conference Article
1
- 10.1109/icicta.2012.33
- Jan 1, 2012
We have applied principal component analysis-artificial neural network (PCA-ANN) in near infrared (NIR) spectroscopy to synchronous and rapid determining the contents of polysaccharide and protein in the Coriolus versicolor Powders. Back-Propagation (BP) Networks which adopt Levenberg-Marquardt training algorithm have been developed. Via analyzing the NIR spectra matrix by principal component analysis (PCA) method, we have obtained the principal components (PC) scores. The original NIR spectra and PC scores were respectively used as input data. These developed BP Networks have been optimized by selecting suitable topologic parameters and the best numbers of training. Compare with original NIR spectra, using the PC scores as input data, the capabilities of BP networks were much better. Using these optimized BP Networks for predicting the contents of polysaccharide and protein in prediction set, the root mean square error of prediction (RMSEP) are 0.0141 and 0.0138. These results are so satisfied and NIR spectroscopy technology is convenient, rapid, no pretreatment and no pollution that this method could be popularized in the in situ measurement and the on-line quality control for fermentation.
- Conference Article
2
- 10.1109/icnc.2007.261
- Jan 1, 2007
We have applied principal component analysis -artificial neural network (PCA-ANN) in near infrared (NIR) spectroscopy to synchronous and rapid determining the contents of rifampicin (RMP), isoniazide (INH) and pyrazinamide (PZA) in compound rifampicin tablets. Back-propagation (BP) networks which adopt Levenberg-Marquardt training algorithm have been developed. Via analyzing the NIR spectra matrix by principal component analysis (PCA) method, we have obtained the principal components (PC) scores. The original NIR spectra and PC scores were respectively used as input data. These developed BP networks have been optimized by selecting suitable topologic parameters and the best numbers of training. Compare with original NIR spectra, using the PC scores as input data, the capabilities of BP networks were much better. Using these optimized BP networks for predicting the contents of RMP, INH and PZA in prediction set, the root mean square error of prediction (RMSEP) are 0.00423, 0.00320 and 0.00608. These results are so satisfied and NIR spectroscopy technology is convenient, rapid, no pretreatment and no pollution that this method could be popularized in the in situ measurement and the on-line quality control for drug production.
- Research Article
32
- 10.1016/s0169-7439(00)00111-8
- Jan 1, 2001
- Chemometrics and Intelligent Laboratory Systems
Multiple imputation and maximum likelihood principal component analysis of incomplete multivariate data from a study of the ageing of port
- Conference Article
- 10.1109/nssmic.2016.8069423
- Oct 1, 2016
The mean value of the non-displaceable binding potential (BP ND ) within a region of interest (ROI) is the traditionally-employed metric in neurological image analysis. The ability of the mean value to accurately track clinical disease progression may be limited since it does not capture the spatial pattern of tracer distribution. In this work, we employ the principal component analysis (PCA) to quantify the clinically-relevant tracer binding patterns ([11C]dihydrotetrabenazine) in high-resolution PET images of 37 Parkinson's disease subjects. The principal component (PC) scores that correspond to different binding patterns in the putamen ROIs are combined with the mean BP ND and used as the input to several linear models that aim to predict the clinical severity of the disease (disease duration). Multiple regression analysis and LASSO (least absolute shrinkage and selection operator) with cross-validation are used to evaluate the contributions of the PC scores to the accuracy of the tested models. With multiple regression analysis, the value of the adjusted R2 was 0.57 when the mean BP ND alone was used as the model input. When the PC scores were included as additional input variables, the value of the adjusted R2 increased to 0.70. The terms of the model representing the PC scores were statistically significant (p ND alone). These results demonstrate that a) the disease- and tracer-specific binding patterns can be identified in sub-cortical brain structures from high-resolution PET images, and b) such patterns may facilitate better models of the clinical disease metrics.
- Research Article
6
- 10.18421/tem122-66
- May 29, 2023
- TEM Journal
Predicting student performance in higher education based on students’ self-efficacy and learning behaviour data is challenging, because the data is changing with time. The potential of using continuous data which is collected weekly needs to be investigated to identify the effectiveness in making predictions of low-performing students. Therefore, this paper presents the analysis of continuous data using the Principal Component Analysis (PCA) and Support Vector Machine (SVM) for predicting student performance. Firstly, we proposed three patterns of the Principal Component (PC) scores to predict the trends of behaviour within a semester. Secondly, we present an analysis of using different combinations of time frames in predicting the performance using the SVM. The obtained results show that three behaviour patterns have been extracted from the Hotelling’s T² values calculated using the PC scores which were fluctuating, ascending, and descending. The use of different time frames using SVM shows different accuracy results in prediction. The use of continuous data indicates that certain data can be predicted at the early stage using multiple time frames.
- Research Article
39
- 10.1016/j.jngse.2017.01.014
- Jan 10, 2017
- Journal of Natural Gas Science and Engineering
New forecasting method for liquid rich shale gas condensate reservoirs with data driven approach using principal component analysis
- Research Article
56
- 10.1117/1.2437738
- Jan 1, 2007
- Journal of Biomedical Optics
The spectral analysis and classification for discrimination of pulsed laser-induced autofluorescence spectra of pathologically certified normal, premalignant, and malignant oral tissues recorded at a 325-nm excitation are carried out using MATLAB@R6-based principal component analysis (PCA) and k-means nearest neighbor (k-NN) analysis separately on the same set of spectral data. Six features such as mean, median, maximum intensity, energy, spectral residuals, and standard deviation are extracted from each spectrum of the 60 training samples (spectra) belonging to the normal, premalignant, and malignant groups and they are used to perform PCA on the reference database. Standard calibration models of normal, premalignant, and malignant samples are made using cluster analysis. We show that a feature vector of length 6 could be reduced to three components using the PCA technique. After performing PCA on the feature space, the first three principal component (PC) scores, which contain all the diagnostic information, are retained and the remaining scores containing only noise are discarded. The new feature space is thus constructed using three PC scores only and is used as input database for the k-NN classification. Using this transformed feature space, the centroids for normal, premalignant, and malignant samples are computed and the efficient classification for different classes of oral samples is achieved. A performance evaluation of k-NN classification results is made by calculating the statistical parameters specificity, sensitivity, and accuracy and they are found to be 100, 94.5, and 96.17%, respectively.
- Research Article
23
- 10.1002/acr.24143
- Mar 26, 2021
- Arthritis Care & Research
To determine if baseline quadriceps and hamstrings muscle activity patterns differed between those with medial-compartment knee osteoarthritis (OA) who advanced to total knee arthroplasty (TKA) and those who did not advance to TKA, and to examine associations between features extracted from principal component analysis (PCA) and discrete measures. Surface electromyograms of the vastus lateralis and medialis, rectus femoris, and lateral and medial hamstrings during walking were collected from 54 individuals with knee OA. Amplitude and temporal characteristics from PCA, co-contraction indices (CCI) for lateral and medial muscle pairs, and root mean square (RMS) amplitudes for early, mid, late, and overall stance were calculated from electromyographic waveforms. At follow-up 5 to 8 years later, 26 participants reported having undergone TKA. Analysis of variance models tested for differences in principal component (PC) scores and discrete measures between TKA and no-TKA groups (α = 0.05). Pearson's product moment correlation coefficients were calculated between PC scores and discrete variables. The TKA group had higher hamstrings activity magnitudes (PC1), prolonged activity in mid stance (PC2) for all muscles, and greater lateral CCI. TKA had higher RMS hamstrings activity for all stance phases, and higher RMS mid- and late-stance quadriceps activity. PC1 was highly correlated with RMS amplitude (highest overall and early stance). PC2 was correlated with mid- and late-stance RMS. CCIs were correlated with PC1 and PC2, with greater variance explained for PC1. Those who advanced to TKA had higher magnitudes and more prolonged agonist and antagonist activity, consistent with less joint unloading. These gait muscle activation patterns indicate a potential conservative intervention target.
- Research Article
28
- 10.1152/jn.01263.2004
- Feb 2, 2005
- Journal of Neurophysiology
In response to touches to their skin, medicinal leeches shorten their body on the side of the touch. We elicited local bends by delivering precisely controlled pressure stimuli at different locations, intensities, and durations to body-wall preparations. We video-taped the individual responses, quantifying the body-wall displacements over time using a motion-tracking algorithm based on making optic flow estimates between video frames. Using principal components analysis (PCA), we found that one to three principal components fit the behavioral data much better than did previous (cosine) measures. The amplitudes of the principal components (i.e., the principal component scores) nicely discriminated the responses to stimuli both at different locations and of different intensities. Leeches discriminated (i.e., produced distinguishable responses) between touch locations that are approximately a millimeter apart. Their ability to discriminate stimulus intensity depended on stimulus magnitude: discrimination was very acute for weak stimuli and less sensitive for stronger stimuli. In addition, increasing the stimulus duration improved the leech's ability to discriminate between stimulus intensities. Overall, the use of optic flow fields and PCA provide a powerful framework for characterizing the discrimination abilities of the leech local bend response.
- Research Article
14
- 10.1016/s1673-8527(08)60055-7
- Jun 1, 2008
- Journal of genetics and genomics = Yi chuan xue bao
An approach to incorporate linkage disequilibrium structure into genomic association analysis
- Research Article
40
- 10.1080/03610920802702535
- Apr 24, 2009
- Communications in Statistics - Theory and Methods
The monitoring of process/product profiles is presently a growing and promising area of research in statistical process control. This study is aimed at developing monitoring schemes for nonlinear profiles with random effects. We utilize the technique of principal components analysis to analyze the covariance structure of the profiles and propose monitoring schemes based on principal component (PC) scores. The number of the PC scores used in constructing control charts is crucial to the detecting power. In the Phase I analysis of historical data, due to the dependency of the PC-scores, we adopt the usual Hotelling T 2 chart to check the stability. For Phase II monitoring, we study individual PC-score control charts, a combined chart scheme that combines all the PC-score charts, and a T 2 chart. Although an individual PC-score chart may be perfect for monitoring a particular mode of variation, a chart that can detect general shifts, such as the T 2 chart and the combined chart scheme, is more feasible in practice. The performances of the schemes under study are evaluated in terms of the average run length.