Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

HipSTR-UI: A cross-platform graphical interface for accessible str genotyping from next-generation sequencing data.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Next-generation sequencing (NGS) has expanded the scope of forensic genetics by providing sequence-level resolution of Short Tandem Repeats (STRs). We developed HipSTR-UI, a cross-platform graphical interface that integrates precompiled HipSTR binaries for Windows, macOS, and Linux. The interface automates the genotyping workflow, from alignment files (BAM/CRAM) and target regions with the reference genome (FASTA) to the generation of genotypes in VCF format. HipSTR-UI includes a graphical parameter panel and exportable results. The validation was performed using Phase 3 of the 1000 Genomes Project, previously analyzed with the HipSTR command-line version. HipSTR-UI accurately reproduced the results of HipSTR, achieving 100 % concordance in genotype calls. As expected, no discrepancies were observed in concordance rates, discordant loci, or quality scores, since the interface executes the same commands as the original HipSTR tool. HipSTR-UI combines the robustness of HipSTR with a user-friendly and multilingual interface, bridging the gap between advanced sequencing technologies and routine forensic applications. Additionally, it facilitates adoption in routine forensic workflows, including human identification and kinship analysis.

Similar Papers
  • Research Article
  • Cite Count Icon 135
  • 10.1016/j.jmoldx.2012.08.001
Detection of FLT3 Internal Tandem Duplication in Targeted, Short-Read-Length, Next-Generation Sequencing Data
  • Nov 14, 2012
  • The Journal of Molecular Diagnostics
  • David H Spencer + 8 more

Detection of FLT3 Internal Tandem Duplication in Targeted, Short-Read-Length, Next-Generation Sequencing Data

  • Research Article
  • Cite Count Icon 1493
  • 10.1038/nrg2986
Genotype and SNP calling from next-generation sequencing data.
  • May 18, 2011
  • Nature Reviews Genetics
  • Rasmus Nielsen + 3 more

Meaningful analysis of next-generation sequencing (NGS) data, which are produced extensively by genetics and genomics studies, relies crucially on the accurate calling of SNPs and genotypes. Recently developed statistical methods both improve and quantify the considerable uncertainty associated with genotype calling, and will especially benefit the growing number of studies using low- to medium-coverage data. We review these methods and provide a guide for their use in NGS studies.

  • Research Article
  • Cite Count Icon 13
  • 10.1016/j.fsigen.2022.102676
Analysis and comparison of the STR genotypes called with HipSTR, STRait Razor and toaSTR by using next generation sequencing data in a Brazilian population sample
  • Feb 4, 2022
  • Forensic Science International: Genetics
  • Guilherme Valle-Silva + 6 more

Analysis and comparison of the STR genotypes called with HipSTR, STRait Razor and toaSTR by using next generation sequencing data in a Brazilian population sample

  • Book Chapter
  • Cite Count Icon 6
  • 10.1007/978-3-319-17157-9_4
Variant Calling Using NGS Data in European Aspen (Populus tremula)
  • Jan 1, 2015
  • Jing Wang + 3 more

Analysis of next-generation sequencing (NGS) data is rapidly becoming an important source of information for genetics and genomics studies. The utility of such data does, however, rely crucially on the accuracy and quality of SNP and genotype calling. Identification of genetic variants (SNPs and short indels) from NGS data is an area of active research and many recent statistical methods have been developed to both improve and quantify the large uncertainty associated with genotype calling. The detection of genetic variants from NGS data is prone to errors, due to multiple factors such as base-calling, alignment errors, and read coverage. Here we highlight some of the issues and review and exemplify some of the recent methods that have been developed for genotype calling. We also provide guidelines for their application to whole-genome re-sequencing data using a data set based on a number of European aspen (Populus tremula) individuals each sequenced to a depth of about 20× coverage per individual.

  • Research Article
  • Cite Count Icon 46
  • 10.1016/j.jmoldx.2013.10.006
Validation for Clinical Use of, and Initial Clinical Experience with, a Novel Approach to Population-Based Carrier Screening using High-Throughput, Next-Generation DNA Sequencing
  • Dec 27, 2013
  • The Journal of Molecular Diagnostics
  • Stephanie Hallam + 12 more

Validation for Clinical Use of, and Initial Clinical Experience with, a Novel Approach to Population-Based Carrier Screening using High-Throughput, Next-Generation DNA Sequencing

  • Research Article
  • Cite Count Icon 2
  • 10.1016/j.scijus.2025.101305
Discriminatory power of the Precision ID GlobalFiler™ NGS STR panel v2 in monozygotic twins for forensic applications.
  • Sep 1, 2025
  • Science & justice : journal of the Forensic Science Society
  • R I B Fonseca + 3 more

Discriminatory power of the Precision ID GlobalFiler™ NGS STR panel v2 in monozygotic twins for forensic applications.

  • Preprint Article
  • 10.7490/f1000research.1110812.1
Comparing algorithms to genotype short tandem repeats in next-generation sequencing data
  • Oct 15, 2015
  • Faculty of 1000 Research Ltd
  • Harriet Dashnow + 1 more

Short tandem repeats (STRs) are short (2-6bp) DNA sequences repeated in tandem, which make up approximately 3% of the human genome. These loci are prone to frequent mutations and high polymorphism. Dozens of neurological and developmental disorders have been attributed to STR expansions. STRs have also been implicated in a range of functions such as DNA replication and repair, chromatin organisation and regulation of gene expression. Traditionally, STR variation has been measured using capillary gel electrophoresis. This process is time-consuming and expensive, and so has tended to limit STR analysis to a handful of loci. Next-generation sequencing has the potential to address these problems. However, determining STR lengths using next-generation sequencing data is difficult. For example, many callers are limited by sequencing read lengths and polymerase slippage during PCR amplification introduces stutter noise. Recently, a small number of software tools have been developed genotype STRs in next-generation sequencing data. We have performed a general comparison of the tools published to date, identifying their application domains, assumptions and limitations. We have assessed the performance of some of the most popular STR genotyping tools on human next-generation sequencing data. When comparing STR callers we have observed drastic differences in which STR loci are identified as variant. Surprisingly, even for variant loci reported in common between tools, there is markedly low concordance between the specific genotype calls. Finally, we draw together our findings to comment on the considerations when choosing and running an STR genotyping tool, with an emphasis on applications to human disease.

  • Preprint Article
  • 10.7490/f1000research.1110901.1
Comparing algorithms to genotype short tandem repeats in next-generation sequencing data
  • Oct 29, 2015
  • F1000Research
  • Harriet Dashnow + 1 more

Short tandem repeats (STRs) are short (2-6bp) DNA sequences repeated in tandem, which make up approximately 3% of the human genome. These loci are prone to frequent mutations and high polymorphism. Dozens of neurological and developmental disorders have been attributed to STR expansions. STRs have also been implicated in a range of functions such as DNA replication and repair, chromatin organisation and regulation of gene expression. Traditionally, STR variation has been measured using capillary gel electrophoresis. This process is time-consuming and expensive, and so has tended to limit STR analysis to a handful of loci. Next-generation sequencing has the potential to address these problems. However, determining STR lengths using next-generation sequencing data is difficult. For example, many callers are limited by sequencing read lengths and polymerase slippage during PCR amplification introduces stutter noise. Recently, a small number of software tools have been developed genotype STRs in next-generation sequencing data. We have performed a general comparison of the tools published to date, identifying their application domains, assumptions and limitations. We have assessed the performance of some of the most popular STR genotyping tools on human next-generation sequencing data. When comparing STR callers we have observed drastic differences in which STR loci are identified as variant. Surprisingly, even for variant loci reported in common between tools, there is markedly low concordance between the specific genotype calls. Finally, we draw together our findings to comment on the considerations when choosing and running an STR genotyping tool, with an emphasis on applications to human disease.

  • Research Article
  • Cite Count Icon 8
  • 10.1093/bioinformatics/btt418
Joint haplotype phasing and genotype calling of multiple individuals using haplotype informative reads.
  • Aug 13, 2013
  • Bioinformatics (Oxford, England)
  • Kui Zhang + 1 more

Hidden Markov model, based on Li and Stephens model that takes into account chromosome sharing of multiple individuals, results in mainstream haplotype phasing algorithms for genotyping arrays and next-generation sequencing (NGS) data. However, existing methods based on this model assume that the allele count data are independently observed at individual sites and do not consider haplotype informative reads, i.e. reads that cover multiple heterozygous sites, which carry useful haplotype information. In our previous work, we developed a new hidden Markov model to incorporate a two-site joint emission term that captures the haplotype information across two adjacent sites. Although our model improves the accuracy of genotype calling and haplotype phasing, haplotype information in reads covering non-adjacent sites and/or more than two adjacent sites is not used because of the severe computational burden. We develop a new probabilistic model for genotype calling and haplotype phasing from NGS data that incorporates haplotype information of multiple adjacent and/or non-adjacent sites covered by a read over an arbitrary distance. We develop a new hybrid Markov Chain Monte Carlo algorithm that combines the Gibbs sampling algorithm of HapSeq and Metropolis-Hastings algorithm and is computationally feasible. We show by simulation and real data from the 1000 Genomes Project that our model offers superior performance for haplotype phasing and genotype calling for population NGS data over existing methods. HapSeq2 is available at www.ssg.uab.edu/hapseq/.

  • Research Article
  • Cite Count Icon 4
  • 10.1007/s13206-015-9210-7
Characterization of human short tandem repeats (STRs) for individual identification using the Ion Torrent
  • Jun 1, 2015
  • BioChip Journal
  • Seri Lim + 7 more

Human genomic short tandem repeats (STRs) are specific gene sequences containing base pairs that are repeatedly arranged. From the various methods available for identifying individuals, STR analysis is the method most widely used in forensic science. Conventional polymerase chain reaction (PCR) was used for STR typing, and the PCR products, consisting of amplified STR loci (amplicons) were electrophoresed with a DNA analysis device. About ten STR markers were used as standards for STR characterization and analysis of size. Extensive efforts are currently being made to explore the STR sequence diversity by analyzing multiple chromosomal loci using next generation sequencing (NGS). NGS greatly facilitates STR marker analysis for individual identification and the complete sequencing of any given sample through concurrent high-throughput sequencing of multiple loci. As a result, NGS data are more accurate and comprehensive compared to that in a conventional database. In order to overcome the limitations of the currently used size-based STR analysis method, we have typed the DNA of 13 combined DNA index system (CODIS) STR markers using Ion PGM. This kit, developed by Ion Torrent, enables the analysis of STR alleles and the sequencing of corresponding genes. We then analyzed the alleles using the HID_STR_Genotyper plugin. Through this, we determined the sequence of the allele type 15 at the D3S1358 locus in all NIST SRM 2391b samples. This allowed for the verification of the exact type of allele, which the conventional size-based STR typing methods could not resolve.

  • Research Article
  • Cite Count Icon 6
  • 10.1111/age.13043
Concordance rate in cattle and sheep between genotypes differing in Illumina GenCall quality score.
  • Feb 1, 2021
  • Animal genetics
  • D P Berry + 4 more

Proper quality control of data prior to downstream analyses is fundamental to ensure integrity of results; quality control of genomic data is no exception. While many metrics of quality control of genomic data exist, the objective of the present study was to quantify the genotype and allele concordance rate between called single nucleotide polymorphism (SNP) genotypes differing in GenCall (GC) score; the GC score is a confidence measure assigned to each Illumina genotype call. This objective was achieved using Illumina beadchip genotype data from 771 cattle (12428767 genotypes in total post-editing) and 80 sheep (1557360 SNPs genotypes in total post-editing) each genotyped in duplicate. The called genotype with the lowest associated GC score was compared to the genotype called for the same SNP in the same duplicated animal sample but with a GC score of >0.90 (assumed to represent the true genotype). The mean genotype concordance rate for a GC score of <0.300, 0.300-0.549, and ≥0.550 in the cattle (sheep in parenthesis) was 0.9467 (0.9864), 0.9707 (0.9953), and 0.9994 (0.99997) respectively; the respective allele concordance rate was 0.9730 (0.9930), 0.9849 (0.9976), and 0.9997 (0.99998). Hence, concordance eroded as the GC score of the called genotype reduced, albeit the impact was not dramatic and was not very noticeable until a GC score of <0.55. Moreover, the impact was greater and more consistent in the cattle population than in the sheep population. Furthermore, an impact of GC score on genotype concordance rate existed even for the same SNP GenTrain value; the GenTrain value is a statistical score that depicts the shape of the genotype clusters and the relative distance between the called genotype clusters.

  • Research Article
  • Cite Count Icon 55
  • 10.1053/j.gastro.2022.02.027
Comparable Results of Helicobacter pylori Antibiotic Resistance Testing of Stools vs Gastric Biopsies Using Next-Generation Sequencing
  • Feb 20, 2022
  • Gastroenterology
  • Steven F Moss + 5 more

Comparable Results of Helicobacter pylori Antibiotic Resistance Testing of Stools vs Gastric Biopsies Using Next-Generation Sequencing

  • Research Article
  • 10.1158/1538-7445.am2016-3186
Abstract 3186: Improved sensitive detection method of FLT3 (FMS-like tyrosine kinase) internal tandem duplication (ITD) mutation using next-generation sequencing technology and nested PCR
  • Jul 15, 2016
  • Cancer Research
  • Daeyoon Kim + 8 more

Sensitive detection of internal tandem duplication (ITD) mutation of FLT3 is very important in acute myeloid leukemia. To increase detection sensitivity of FLT3-ITD, we developed new detection algorithm using next generation sequencing (NGS) data. We validated results using nested polymerase chain reaction (PCR) methods. We compared results of NGS data, nested PCR and conventional PCR methods. First, using whole exome sequencing data of 83 AML patients, we applied calling algorithm for FLT3-ITD. Briefly, to detect ITDs with NGS data, the reads are aligned to a reference sequence (UCSC hg19), with BWA which is a read aligner allowing soft-clipping. Some reads can be an indication of the occurrence of ITD and BWA aligns those reads as soft-clipped. Second, we deigned two types of primer for Nested PCR. The first primer was targeted wildly for between exon14 and exon15 of FLT3 gene. Nested PCR primer was deigned to target previously reported regions which are frequently occurred ITD mutation. PCR reactions of two steps were performed using the PCR primers sequentially. In these 83 patients, FLT3-ITD was positive only in 7 patients when tested by conventional PCR methods. When NGS detection method was applied, this resulted in positive FLT3-ITD in 11 patients (11/83, 13%). When validation was performed using nested PCR, FLT3-ITD was confirmed in all of 11 patients. Nested PCR detected additional 4 patient positive for FLT3-ITD in this population. For 68 patients, FLT3-ITD was negative by both NGS and nested PCR method. Overall, NGS method improved sensitivity of FLT3-ITD detection by 57% in this population. And the concordance rate of NGS method and nested PCR was 95.2% (79/83). Then we investigated clinical significance of sensitive FLT3-ITD detection. For this, we performed nested PCR and conventional PCR at the same time in 238 AML patients to detect FLT3-ITD. Positive rate for FLT3-ITD was 20% (48/238) and 10% (24/238) by nested PCR and conventional PCR respectively. When survival analysis was performed, among patients with negative FLT3-ITD result by conventional PCR, patients who showed positive for FLT3-ITD by nested PCR had shorter overall survival compared to those who showed negative for FLT3-ITD by nested PCR. (p = 0.03). This implies that sensitive FLT3-ITD detection using nested PCR is clinically meaningful. Diagnosis of FLT3-ITD is very important genetic factor, leading a therapeutic direction for AML patient. Here we report that we have developed alternative more sensitive detection methods for FLT3-ITD based on nested PCR and NGS. Sensitive detection of FLT3-ITD was clinically meaningful, suggesting that these methods should be incorporated in a future clinical practice. Also, we want to note that, NGS method is capable of quantifying FLT3-ITD size and amount in AML patients. Citation Format: Daeyoon Kim, Yoojin Hong, Youngil Koh, Sung-Soo Yoon, Choong-Hyun Sun, Kwang-Sung Ahn, Seungmook Lee, Hongseok Yun, Suyeon Lee. Improved sensitive detection method of FLT3 (FMS-like tyrosine kinase) internal tandem duplication (ITD) mutation using next-generation sequencing technology and nested PCR. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 3186.

  • Research Article
  • 10.1158/1538-7445.am2019-1660
Abstract 1660: Identification of allelic imbalance utilizing heterozygous genotype allele frequencies and intensities
  • Jul 1, 2019
  • Cancer Research
  • Kyle Chang + 6 more

Allelic imbalance (AI) events, such as amplification, deletion or copy-neutral loss-of-heterozygosity (cn-LOH) can result in the activations of oncogenes or inactivations of tumor suppressor genes that are critical to the process of carcinogenesis and metastasis. However, tumor samples often have low tumor-cellularity and small subclones that require sensitive and robust algorithm for AI detections. To overcome this challenge, our lab has developed an original software application hapLOH (Vattathil and Scheet, 2013) that identifies AI with high sensitivity in SNP array data by utilizing heterozygous genotype allele frequencies in SNP array data using a Hidden Markov Model. Motivated by the plethora of next-generation sequencing (NGS) data, we carried this logic and functionality and developed hapLOHseq (San Lucas et al, 2016) for detection of subtle AI in NGS data using reference and alternate allele read depths and a set of haplotype estimates based on 1000 genome data. Recently, we sought to improve AI detection and provide support for different data types by developing haploh-cn. Haploh-cn extends our previous tools by supporting both SNP array and NGS data, and by incorporating intensity data and Log R ratios (LRR) into the underlying Hidden Markov Model for SNP array data, and incorporating sequencing depth for NGS data. In order to evaluate haploh-cn’s performances, we have downloaded The Cancer Genome Atlas Lung Adenocarcinoma whole exome sequencing data and corresponding Affymetrix SNP6 array data. Then, we simulated tumor cellularity by performing in-silico dilution of NGS and SNP array data. For NGS data, we down sampled tumor sequencing BAM files and mixed in matched-normal sequencing reads. For SNP6 array data, we combined the adjusted B allele frequencies of heterozygous sites in the tumor sample and matched-normal sample. Using the AI events detected in the undiluted tumor samples as the gold standard, we assessed the performance of haploh-cn in the computational diluted samples. Our results showed that haploh-cn had high sensitivity without sacrificing specificity in diluted tumor samples. In summary, haploh-cn is a robust and powerful method for profiling subtle AI in NGS and SNP array data. Citation Format: Kyle Chang, Francis A. San Lucas, Zuhal Ozcan, Smruthy Sivakumar, Yasminka A. Jakubek, Richard G. Fowler, Paul Scheet. Identification of allelic imbalance utilizing heterozygous genotype allele frequencies and intensities [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 1660.

  • Research Article
  • 10.16476/j.pibb.2016.0144
A Novel Method for Identifying Length Variations of Short Tandem Repeats Based on Next Generation Sequencing and Its Application in Human Genetic Disease Research
  • Jun 25, 2016
  • Zhangming Yan + 4 more

Next generation sequencing (NGS) technologies boosted genomic and medical research, particularly for identification of disease-causing variants. Although most types of genetic variants could be identified through NGS data analysis, there are still some limitations, such as length variations of short tandem repeats (STRs). Many genetic diseases are known to be caused by expansions of STRs, especially neurological disorders, such as Huntington disease. However, almost none of existing tools could detect STRs expanded longer than sequencing read length based on NGS. To break through the limitation, we developed a novel method for detecting length variations of STRs and estimating the length of expansions based on paired-end NGS. We applied our method in a clinical study of motor neuron disease using whole-exome sequencing and successfully identified a disease-causing expansion of STR. Our method firstly used special features of depth of read coverage at STRs to address the variant calling problem. It has widely application value in human genetic disease research and inspirational value in developing new NGS data processing tools.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant