Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Reconstructing Almost All of a Point Set in \(\boldsymbol{\mathbb{R}}\) d from Randomly Revealed Pairwise Distances

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Reconstructing Almost All of a Point Set in \(\boldsymbol{\mathbb{R}}\) <sup> <i>d</i> </sup> from Randomly Revealed Pairwise Distances

Similar Papers
  • Research Article
  • Cite Count Icon 4
  • 10.1093/ve/veae087
Dimensionality reduction distills complex evolutionary relationships in seasonal influenza and SARS-CoV-2.
  • Nov 14, 2024
  • Virus evolution
  • Sravani Nanduri + 3 more

Public health researchers and practitioners commonly infer phylogenies from viral genome sequences to understand transmission dynamics and identify clusters of genetically-related samples. However, viruses that reassort or recombine violate phylogenetic assumptions and require more sophisticated methods. Even when phylogenies are appropriate, they can be unnecessary or difficult to interpret without specialty knowledge. For example, pairwise distances between sequences can be enough to identify clusters of related samples or assign new samples to existing phylogenetic clusters. In this work, we tested whether dimensionality reduction methods could capture known genetic groups within two human pathogenic viruses that cause substantial human morbidity and mortality and frequently reassort or recombine, respectively: seasonal influenza A/H3N2 and SARS-CoV-2. We applied principal component analysis, multidimensional scaling (MDS), t-distributed stochastic neighbor embedding (t-SNE), and uniform manifold approximation and projection to sequences with well-defined phylogenetic clades and either reassortment (H3N2) or recombination (SARS-CoV-2). For each low-dimensional embedding of sequences, we calculated the correlation between pairwise genetic and Euclidean distances in the embedding and applied a hierarchical clustering method to identify clusters in the embedding. We measured the accuracy of clusters compared to previously defined phylogenetic clades, reassortment clusters, or recombinant lineages. We found that MDS embeddings accurately represented pairwise genetic distances including the intermediate placement of recombinant SARS-CoV-2 lineages between parental lineages. Clusters from t-SNE embeddings accurately recapitulated known phylogenetic clades, H3N2 reassortment groups, and SARS-CoV-2 recombinant lineages. We show that simple statistical methods without a biological model can accurately represent known genetic relationships for relevant human pathogenic viruses. Our open source implementation of these methods for analysis of viral genome sequences can be easily applied when phylogenetic methods are either unnecessary or inappropriate.

  • Research Article
  • Cite Count Icon 1
  • 10.2307/1547806
DNA Sequence Data in Phylogeny Reconstruction: An Overview of Techniques and Analyses
  • Oct 1, 1995
  • American Fern Journal
  • Tom A Ranker

Tom A. Ranker, DNA Sequence Data in Phylogeny Reconstruction: An Overview of Techniques and Analyses, American Fern Journal, Vol. 85, No. 4, Use of Molecular Data in Evolutionary Studies of Pteridophytes (Oct. - Dec., 1995), pp. 123-133

  • Conference Article
  • Cite Count Icon 300
  • 10.1109/infcom.2005.1497889
Mobile-assisted localization in wireless sensor networks
  • Mar 13, 2005
  • N.B Priyantha + 3 more

The localization problem is to determine an assignment of coordinates to nodes in a wireless ad-hoc or sensor network that is consistent with measured pairwise node distances. Most previously proposed solutions to this problem assume that the nodes can obtain pairwise distances to other nearby nodes using some ranging technology. However, for a variety of reasons that include obstructions and lack of reliable omnidirectional ranging, this distance information is hard to obtain in practice. Even when pairwise distances between nearby nodes are known, there may not be enough information to solve the problem uniquely. This paper describes MAL, a mobile-assisted localization method which employs a mobile user to assist in measuring distances between node pairs until these distance constraints form a globally rigid'* structure that guarantees a unique localization. We derive the required constraints on the mobile's movement and the minimum number of measurements it must collect; these constraints depend on the number of nodes visible to the mobile in a given region. We show how to guide the mobile's movement to gather a sufficient number of distance samples for node localization. We use simulations and measurements from an indoor deployment using the Cricket location system to investigate the performance of MAL, finding in real-world experiments that MAL's median pairwise distance error is less than 1.5% of the true node distance.

  • Research Article
  • Cite Count Icon 1
  • 10.1101/2024.02.07.579374
Dimensionality reduction distills complex evolutionary relationships in seasonal influenza and SARS-CoV-2.
  • Aug 29, 2024
  • bioRxiv : the preprint server for biology
  • Sravani Nanduri + 3 more

Public health researchers and practitioners commonly infer phylogenies from viral genome sequences to understand transmission dynamics and identify clusters of genetically-related samples. However, viruses that reassort or recombine violate phylogenetic assumptions and require more sophisticated methods. Even when phylogenies are appropriate, they can be unnecessary or difficult to interpret without specialty knowledge. For example, pairwise distances between sequences can be enough to identify clusters of related samples or assign new samples to existing phylogenetic clusters. In this work, we tested whether dimensionality reduction methods could capture known genetic groups within two human pathogenic viruses that cause substantial human morbidity and mortality and frequently reassort or recombine, respectively: seasonal influenza A/H3N2 and SARS-CoV-2. We applied principal component analysis (PCA), multidimensional scaling (MDS), t-distributed stochastic neighbor embedding (t-SNE), and uniform manifold approximation and projection (UMAP) to sequences with well-defined phylogenetic clades and either reassortment (H3N2) or recombination (SARS-CoV-2). For each low-dimensional embedding of sequences, we calculated the correlation between pairwise genetic and Euclidean distances in the embedding and applied a hierarchical clustering method to identify clusters in the embedding. We measured the accuracy of clusters compared to previously defined phylogenetic clades, reassortment clusters, or recombinant lineages. We found that MDS embeddings accurately represented pairwise genetic distances including the intermediate placement of recombinant SARS-CoV-2 lineages between parental lineages. Clusters from t-SNE embeddings accurately recapitulated known phylogenetic clades, H3N2 reassortment groups, and SARS-CoV-2 recombinant lineages. We show that simple statistical methods without a biological model can accurately represent known genetic relationships for relevant human pathogenic viruses. Our open source implementation of these methods for analysis of viral genome sequences can be easily applied when phylogenetic methods are either unnecessary or inappropriate.

  • Research Article
  • 10.3390/ijms27093750
The Distribution of Average Pairwise Distances Among Human Pre-miRNAs for Disease Association Analysis
  • Apr 23, 2026
  • International Journal of Molecular Sciences
  • Hsiuying Wang + 2 more

MicroRNAs (miRNAs) play essential roles in cell differentiation, development, gene regulation, and apoptosis, and have been widely implicated in numerous disease mechanisms. Owing to their regulatory importance, miRNAs are increasingly recognized as valuable disease biomarkers. Previous studies have used nucleotide sequence pairwise distances between miRNAs to explore disease associations and have derived the distribution of pairwise distances to assess the percentile rank of an observed miRNA pair. However, because a single disease may involve multiple miRNA biomarkers, evaluating the percentile rank of an average pairwise distance is often more appropriate than focusing on individual pairs. In this study, we established percentile distributions for the average pairwise nucleotide distances corresponding to different numbers of miRNAs. Applying this framework to 51 diseases and several groups of related diseases, we observed that miRNA biomarkers associated with the same disease, as well as with related diseases, often exhibit low-percentile average pairwise distances under the reference distribution. While the present study does not directly evaluate whether precursor miRNA (pre-miRNA) sequence similarity is associated with shared biological function or regulatory targets, the proposed framework provides a systematic approach for quantifying such similarity among disease-associated miRNAs and may serve as a useful foundation for future studies integrating functional and clinical validation.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 17
  • 10.1371/journal.pone.0160649
Analysis of Viral Diversity in Relation to the Recency of HIV-1C Infection in Botswana
  • Aug 23, 2016
  • PLoS ONE
  • Sikhulile Moyo + 10 more

BackgroundCross-sectional, biomarker methods to determine HIV infection recency present a promising and cost-effective alternative to the repeated testing of uninfected individuals. We evaluate a viral-based assay that uses a measure of pairwise distances (PwD) to identify HIV infection recency, and compare its performance with two serologic incidence assays, BED and LAg. In addition, we assess whether combination BED plus PwD or LAg plus PwD screening can improve predictive accuracy by reducing the likelihood of a false-recent result.MethodsThe data comes from 854 time-points and 42 participants enrolled in a primary HIV-1C infection study in Botswana. Time points after treatment initiation or with evidence of multiplicity of infection were excluded from the final analysis. PwD was calculated from quasispecies generated using single genome amplification and sequencing. We evaluated the ability of PwD to correctly classify HIV infection recency within <130, <180 and <360 days post-seroconversion using Receiver Operator Characteristics (ROC) methods. Following a secondary PwD screening, we quantified the reduction in the relative false-recency rate (rFRR) of the BED and LAg assays while maintaining a sensitivity of either 75, 80, 85 or 90%.ResultsThe final analytic sample consisted of 758 time-points from 40 participants. The PwD assay was more accurate in classifying infection recency for the 130 and 180-day cut-offs when compared with the recommended LAg and BED thresholds. A higher AUC statistic confirmed the superior predictive performance of the PwD assay for the three cut-offs. When used for combination screening, the PwD assay reduced the rFRR of the LAg assay by 52% and the BED assay by 57.8% while maintaining a 90% sensitivity for the 130 and 180-day cut-offs respectively.ConclusionPwD can accurately determine HIV infection recency. A secondary PwD screening reduces misclassification and increases the accuracy of serologic-based assays.

  • Research Article
  • Cite Count Icon 17
  • 10.5958/0976-0555.2015.00052.7
Morphological differentiation of indigenous goats in different agro-ecological zones of Vhembe district, Limpopo province, South Africa
  • Jan 1, 2015
  • Indian Journal of Animal Research
  • Tlou C Selolo + 4 more

The study was carried out to differentiate indigenous female goats in different agro-ecological zones based on their morphological traits. 551 mature female goats from semi-arid, dry sub-humid and humid agro-ecological zones were considered in this study. The morphological traits analysed were Body Weight, Body Length, Shoulder Height, Hip Height and Heart Girth. Stepwise discriminant analysis was used to check the discriminating power of the variables. Canonical discriminant procedure was applied to determine differences between indigenous goats in different zones. The analysis showed that all the variables measured have discriminating power. In canonical discriminant analysis, the first canonical variable determined was significant and accounted for 91.87% of the variation, but the second canonical variable was not significant and accounted 8.13% of the variation. The pairwise Mahalanobis distances between indigenous goats in semi-arid and dry sub-humid as well as between semi-arid and humid were significant. The pairwise distance between goats in dry sub-humid and humid was not significant. Discriminant model function correctly allocated 60.31% (semi-arid), 58.06% (humid) and 38.46% (dry subhumid) of indigenous goats into their original agro-ecological zone.

  • Research Article
  • Cite Count Icon 19
  • 10.1109/tvcg.2016.2534559
Adaptive Disentanglement based on Local Clustering in Small-World Network Visualization.
  • Feb 29, 2016
  • IEEE Transactions on Visualization and Computer Graphics
  • Arlind Nocaj + 2 more

Small-world networks have characteristically low pairwise shortest-path distances, causing distance-based layout methods to generate hairball drawings. Recent approaches thus aim at finding a sparser representation of the graph to amplify variations in pairwise distances. Since the effect of sparsification on the layout is difficult to describe analytically, the incorporated filtering parameters of these approaches typically have to be selected manually and individually for each input instance. We here propose the use of graph invariants to determine suitable parameters automatically. This allows us to perform adaptive filtering to obtain drawings in which the cluster structure is most prominent. The approach is based on an empirical relationship between input and output characteristics that is derived from real and synthetic networks.Experimental evaluation shows the effectiveness of our approach and suggests that it can be used by default to increase the robustness of force-directed layout methods.

  • Research Article
  • Cite Count Icon 58
  • 10.1099/jgv.0.001100
Baculovirus Kimura two-parameter species demarcation criterion is confirmed by the distances of 38 core gene nucleotide sequences.
  • Jul 27, 2018
  • Journal of General Virology
  • Jörg T Wennmann + 2 more

Kimura two-parameter nucleotide distance comparisons based on polyhedrin/granulin (polh/gran), late expression factor 8 (lef-8) and late expression factor 9 (lef-9) are a widely applied method for species demarcation for lepidopteran-specific baculoviruses. Baculoviruses are considered to belong to the same species when a pairwise distance threshold of 0.015 is not exceeded and are considered as possibly belonging to the same species with a distance of up to 0.050. In the present work this method was revised and extended for 172 entirely sequenced lepidopteran, hymenopteran and dipteran baculovirus genomes by applying the nucleotide sequences of all 38 known baculovirus core genes for pairwise distance calculations. On the basis of this large dataset, the previously established standard thresholds for baculovirus species demarcation were adjusted for pairwise nucleotide distances estimated from the alignments of all 38 core genes. With the newly applied thresholds for the 38 core-gene dataset, a more sophisticated Kimura two-parameter method was established, avoiding the possible influence of the chimerical polh gene of the Autographa californica multiple nucleopolyhedrovirus. Based on the new dataset, the present classification of baculovirus species was confirmed. Thereby the Kimura two-parameter method for baculovirus demarcation was extended to include the information from all 38 Baculoviridae core genes, which represent the established standard information for baculovirus phylogeny to date.

  • Research Article
  • Cite Count Icon 25
  • 10.1016/j.ympev.2010.04.036
Incorporating molecular phylogenetics with larval morphology while mitigating the effects of substitution saturation on phylogeny estimation: A new hypothesis of relationships for the flatfish family Pleuronectidae (Percomorpha: Pleuronectiformes)
  • Apr 29, 2010
  • Molecular Phylogenetics and Evolution
  • Dawn M Roje

Incorporating molecular phylogenetics with larval morphology while mitigating the effects of substitution saturation on phylogeny estimation: A new hypothesis of relationships for the flatfish family Pleuronectidae (Percomorpha: Pleuronectiformes)

  • Research Article
  • Cite Count Icon 46
  • 10.1037/0096-3445.108.1.99
Studies of the cognitive representation of spatial relations: III. A hypothetical environment.
  • Jan 1, 1979
  • Journal of Experimental Psychology: General
  • Amanda A Merrill + 1 more

This experiment investigated people's preferences for the location of facilities in an ideal town. Ten graduate students represented the relative locations of facilities (such as home, school, factory) by two methods: (a) pairwise ideal distances on a 100-point scale and (b) direct planning of locations on a Tektronix cathode ray screen. The pairwise distances were analyzed by multidimensional scaling (MDS) and the facilities were thus situated in a two-dimensional space. Subjects then expressed a preference between the direct plan and the one created by MDS. In addition, the rank order priorities of the facilities were determined for each subject. The entire procedure was repeated after 4 mo. A common central plan was evident in all cases (and rank order priorities were stable), but there was within-subject variability in the plans for different methods and test occasions. Despite such variability, subjects generally preferred their direct plan over the one created by MDS (based on pair estimates). A second group of subjects showed equal preference (on the average) for both types of town representations created by the first group. Both the pair and direct technique seem appropriate for studying cognitive representations of a hypothetical environment.

  • Research Article
  • Cite Count Icon 16
  • 10.1016/j.jviromet.2010.02.016
Evaluation of pre-screening methods for the identification of HIV-1 superinfection
  • Feb 21, 2010
  • Journal of Virological Methods
  • Andrea Rachinger + 4 more

Evaluation of pre-screening methods for the identification of HIV-1 superinfection

  • Book Chapter
  • Cite Count Icon 4
  • 10.1007/978-3-642-13818-8_17
Finding Top-k Similar Pairs of Objects Annotated with Terms from an Ontology
  • Jan 1, 2010
  • Arnab Bhattacharya + 2 more

With the growing focus on semantic searches, an increasing number of standardized ontologies are being designed to describe data. We investigate the querying of objects described by a tree-structured ontology. Specifically, we consider the case of finding the top-kbest pairs of objects that have been annotated with terms from such an ontology when the object descriptions are available only at runtime. We consider three distance measures. The first one defines the object distance as the minimum pairwise distance between the sets of terms describing them and the second one defines the distance as the average pairwise term distance. The third and most useful distance measure--earth mover's distance-- finds the best way of matching the terms and computes the distance corresponding to this best matching. We develop lower bounds that can be aggregated progressively and utilize them to speed up the search for top-kobject pairs when the earth mover's distance is used. For the minimum pairwise distance, we devise an algorithm that runs inO(D + Tklogk) time, whereDis the total information size andTis the number of terms in the ontology. We also develop a best-first search strategy for the average pairwise distance that utilizes lower bounds generated in an ordered manner. Experiments on real and synthetic datasets demonstrate the practicality and scalability of our algorithms.

  • Book Chapter
  • Cite Count Icon 8
  • 10.1007/978-3-540-77024-4_60
Relative Positions Within Small Teams of Mobile Units
  • Dec 12, 2007
  • Hongbin Li + 3 more

It is common that a small group of autonomous mobile units requires location information of each other, but it is often not applicable to build infrastructure for location service. Hence, mobile units use RF signals to determine their relative positions within the group. Several techniques have been proposed for localization, but none of them consider both units mobility and anchor unavailability. In this paper we develop a propagation scheme for spreading the signal strength information through the network, and filtering techniques are employed to process the noisy signal. We use Floyd-Warshall algorithm to generate pairwise signal distance of each pair of units. Then Multidimensional Scaling technique is used to generate relative position from pairwise distances. Due to anchor unavailability, relative positions are adjusted by certain rules to reflect the continuous mobility. We verify our methods in MICAz platform. Experimental results show that we can obtain smooth relative positions under mobility and manage moving patterns of mobile units.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 134
  • 10.1371/journal.pone.0106940
Building-up of a DNA barcode library for true bugs (insecta: hemiptera: heteroptera) of Germany reveals taxonomic uncertainties and surprises.
  • Sep 9, 2014
  • PLoS ONE
  • Michael J Raupach + 5 more

During the last few years, DNA barcoding has become an efficient method for the identification of species. In the case of insects, most published DNA barcoding studies focus on species of the Ephemeroptera, Trichoptera, Hymenoptera and especially Lepidoptera. In this study we test the efficiency of DNA barcoding for true bugs (Hemiptera: Heteroptera), an ecological and economical highly important as well as morphologically diverse insect taxon. As part of our study we analyzed DNA barcodes for 1742 specimens of 457 species, comprising 39 families of the Heteroptera. We found low nucleotide distances with a minimum pairwise K2P distance <2.2% within 21 species pairs (39 species). For ten of these species pairs (18 species), minimum pairwise distances were zero. In contrast to this, deep intraspecific sequence divergences with maximum pairwise distances >2.2% were detected for 16 traditionally recognized and valid species. With a successful identification rate of 91.5% (418 species) our study emphasizes the use of DNA barcodes for the identification of true bugs and represents an important step in building-up a comprehensive barcode library for true bugs in Germany and Central Europe as well. Our study also highlights the urgent necessity of taxonomic revisions for various taxa of the Heteroptera, with a special focus on various species of the Miridae. In this context we found evidence for on-going hybridization events within various taxonomically challenging genera (e.g. Nabis Latreille, 1802 (Nabidae), Lygus Hahn, 1833 (Miridae), Phytocoris Fallén, 1814 (Miridae)) as well as the putative existence of cryptic species (e.g. Aneurus avenius (Duffour, 1833) (Aradidae) or Orius niger (Wolff, 1811) (Anthocoridae)).

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant