Ensemble lemmatization with the Classical Language Toolkit
Because of the less-resourced nature of historical languages, non-standard solutions are often required for natural language processing tasks. This article introduces one such solution for historical-language lemmatization, that is the Ensemble lemmatizer for the Classical Language Toolkit, an open-source Python package that supports NLP research for historical languages. Ensemble lemmatization is the most recent development at CLTK in the repurposing and refactoring of an existing method designed for one task, specifically the backoff method as used for part-of-speech tagging, for use in a different task, namely lemmatization. This article argues for the benefits of ensemble lemmatization, specifically, flexible tool construction and the use of all available information to reach tagging decisions, and presents two use cases.
- Research Article
1
- 10.1016/j.softx.2023.101514
- Oct 16, 2023
- SoftwareX
Data that resides on the surface of a 2-sphere is common in various scientific fields, including physics, earth sciences, astronomy, and psychoacoustics. While some tools and packages exist for performing inferential statistical tests on such data and model fitting, there is currently no comprehensive open-source Python package that implements these tests. sphstat aims to fill this gap by providing an open-source Python package that implements spherical inferential tests and some model-fitting algorithms as catalogued in the authoritative reference by Fisher et al. (1993). Due to the lack of a similar open-source Python package, sphstat has the potential to be widely used in scientific and technical fields where data on the 2-sphere emerges.
- Research Article
250
- 10.1152/jn.00273.2019
- Jul 3, 2019
- Journal of Neurophysiology
Neural oscillations are widely studied using methods based on the Fourier transform, which models data as sums of sinusoids. This has successfully uncovered numerous links between oscillations and cognition or disease. However, neural data are nonsinusoidal, and these nonsinusoidal features are increasingly linked to a variety of behavioral and cognitive states, pathophysiology, and underlying neuronal circuit properties. We present a new analysis framework, one that is complementary to existing Fourier and Hilbert transform-based approaches, that quantifies oscillatory features in the time domain on a cycle-by-cycle basis. We have released this cycle-by-cycle analysis suite as "bycycle," a fully documented, open-source Python package with detailed tutorials and troubleshooting cases. This approach performs tests to assess whether an oscillation is present at any given moment and, if so, quantifies each oscillatory cycle by its amplitude, period, and waveform symmetry, the latter of which is missed with the use of conventional approaches. In a series of simulated event-related studies, we show how conventional Fourier and Hilbert transform approaches can conflate event-related changes in oscillation burst duration as increased oscillatory amplitude and as a change in the oscillation frequency, even though those features were unchanged in simulation. Our approach avoids these errors. Furthermore, we validate this approach in simulation and against experimental recordings of patients with Parkinson's disease, who are known to have nonsinusoidal beta (12-30 Hz) oscillations.NEW & NOTEWORTHY We introduce a fully documented, open-source Python package, bycycle, for analyzing neural oscillations on a cycle-by-cycle basis. This approach is complementary to traditional Fourier and Hilbert transform-based approaches but avoids specific pitfalls. First, bycycle confirms an oscillation is present, to avoid analyzing aperiodic, nonoscillatory data as oscillations. Next, it quantifies nonsinusoidal aspects of oscillations, increasingly linked to neural circuit physiology, behavioral states, and diseases. This approach is tested against simulated and real data.
- Research Article
1
- 10.3897/biss.5.75688
- Sep 27, 2021
- Biodiversity Information Science and Standards
pyOpenSci (short for Python Open Science), funded by the Alfred P. Sloan Foundation, is building a diverse community that supports well documented, open source Python software that enables open reproducible science. pyOpenSci will work with the community to openly develop best practice guidelines and open standards for scientific Python software, which will be reinforced through a community-led peer review process and training. Packages that complete the peer review process become a part of the pyOpenSci ecosystem, where maintenance can be shared to ensure longevity and stability in code. pyOpenSci packages are also eligible for a “fast tracked” acceptance to JOSS (Journal of Open Source Software). In addition, we provide review for open science tools that would be of interest to TDWG members but are not within scope for JOSS, such as API (Application Programming Interface) wrappers. pyOpenSci is built on top of the successful model of rOpenSci, founded in 2011, which has fostered the development of several useful biodiversity informatics R packages. The pyOpenSci team looks to following the lessons learned by rOpenSci, to create a similarly successful community. We invite TDWG members developing open source software tools in Python to become part of the pyOpenSci community.
- Research Article
16
- 10.1016/j.ifacol.2022.07.562
- Jan 1, 2022
- IFAC-PapersOnLine
Efficient and Simple Gaussian Process Supported Stochastic Model Predictive Control for Bioreactors using HILO-MPC
- Research Article
2
- 10.1121/10.0010552
- Apr 1, 2022
- The Journal of the Acoustical Society of America
An automated algorithm for passive acoustic detection of blue whale D-calls is developed based on established deep learning methods for image recognition via the DenseNet architecture. Koogu—an open-source Python package—was used for developing the detector. The detector was trained on annotated acoustic recordings from the Antarctic, and the performance of the detector was assessed by calculating precision and recall using a separate independent dataset also from the Antarctic. Detections from both the human analyst and automated detector were then inspected by a more experienced analyst to identify any calls missed by either approach and to adjudicate whether the apparent false-positive detections from the automated approach were actually true-positives. Lastly, an additional performance assessment was conducted using double-platform methods (via a closed-population Huggins mark recapture model) to assess the probability of detection of both the human analyst and automated detector, based on the assumption of false-positive-free and reconciled detections. According to our double-platform analysis, the automated detector performed very well with higher recall and fewer false-positives that the original human analyst.
- Abstract
- 10.1093/cdn/nzac063.018
- Jun 1, 2022
- Current Developments in Nutrition
minimod: An Open Source Python Package to Evaluate the Cost Effectiveness of Micronutrient Intervention Programs
- Research Article
24
- 10.1177/2211068214553022
- Oct 10, 2014
- SLAS Technology
PLACE: an open-source python package for laboratory automation, control, and experimentation.
- Research Article
1
- 10.1107/s1600577524005861
- Jul 30, 2024
- Journal of synchrotron radiation
In situ wavefront sensing plays a critical role in the delivery of high-quality beams for X-ray experiments. X-ray speckle-based techniques stand out among other in situ techniques for their easy experimental setup and various data acquisition modes. Although X-ray speckle-based techniques have been under development for more than a decade, there are still no user-friendly software packages for new researchers to begin with. Here, we present an open-source Python package, spexwavepy, for X-ray wavefront sensing using speckle-based techniques. This Python package covers a variety of X-ray speckle-based techniques, provides plenty of examples with real experimental data and offers detailed online documentation for users. We hope it can help new researchers learn and apply the speckle-based techniques for X-ray wavefront sensing to synchrotron radiation and X-ray free-electron laser beamlines.
- Research Article
- 10.1002/mp.18079
- Sep 1, 2025
- Medical physics
Cascaded linear models are widely used for the development and optimization of x-ray imaging systems, yet no publicly available Python implementation currently exists. We introduce CASYMIR, a flexible and open-source Python package capable of modeling direct and indirect-conversion x-ray imaging detectors under various acquisition conditions. We employed a modular software design with generalized frequency-domain expressions for each process in the detection chain, which can be implemented as serial or parallel blocks. The gain factors and other parameters are derived from the detector's characteristics, system geometry, and incident x-ray spectra, all of which can be specified by the user. The signal reaching the detector is propagated throughout the detection stages by applying these process blocks, enabling the computation of the Modulation Transfer Function (MTF) and Noise Power Spectrum (NPS) at any stage of the model. Our implementation was experimentally validated using two commercial x-ray detectors: a flat-panel a-Se detector for digital mammography and digital breast tomosynthesis, and a flat-panel scintillator (CsI) detector for dedicated breast CT. The modeled MTF had root-mean-square (RMS) percent errors below 6% for the a-Se detector, while the normalized RMS error for the NNPS was below 3%. For the CsI detector, the RMS percent error in the MTF was 5.4%, and the normalized RMS error for the NNPS was 5.8%. The CASYMIR Python package can be downloaded from https://github.com/radboud-axti/casymir_public, and it includes a standalone executable script suitable for modeling common commercial systems, along with an extensive README file and example files. CASYMIR is available as an open-source Python package under the MIT license. Given its modular and flexible structure, it can be easily modified and integrated into other simulation/virtual clinical trial pipelines where information about the detector's spatial resolution and noise performance is needed. The standalone version of CASYMIR may be particularly useful for running batch simulations with varying acquisition and system parameters, making it ideal for optimizing system design and acquisition techniques. Furthermore, given the package's modular structure, new processes can be implemented to simulate other detector and system designs.
- Research Article
- 10.1016/j.atech.2026.101874
- Mar 1, 2026
- Smart Agricultural Technology
DiscoEPG: A Python package for characterization of insect electrical penetration graph (EPG) signals
- Research Article
15
- 10.3390/s22155849
- Aug 5, 2022
- Sensors
Developing machine learning algorithms for time-series data often requires manual annotation of the data. To do so, graphical user interfaces (GUIs) are an important component. Existing Python packages for annotation and analysis of time-series data have been developed without addressing adaptability, usability, and user experience. Therefore, we developed a generic open-source Python package focusing on adaptability, usability, and user experience. The developed package, Machine Learning and Data Analytics (MaD) GUI, enables developers to rapidly create a GUI for their specific use case. Furthermore, MaD GUI enables domain experts without programming knowledge to annotate time-series data and apply algorithms to it. We conducted a small-scale study with participants from three international universities to test the adaptability of MaD GUI by developers and to test the user interface by clinicians as representatives of domain experts. MaD GUI saves up to 75% of time in contrast to using a state-of-the-art package. In line with this, subjective ratings regarding usability and user experience show that MaD GUI is preferred over a state-of-the-art package by developers and clinicians. MaD GUI reduces the effort of developers in creating GUIs for time-series analysis and offers similar usability and user experience for clinicians as a state-of-the-art package.
- Research Article
1
- 10.1016/j.ijhydene.2025.05.423
- Jul 1, 2025
- International Journal of Hydrogen Energy
Excess renewable energy can be stored by converting it to gaseous fuels such as hydrogen and methane via electrochemical reactions. Similarly, it can be stored mechanically by compressing gases such as CO 2 . The energy density of such systems is dependent on the storage density of the fluids, which is typically low under ambient conditions. Adsorption onto surfaces increases the density of fluids, avoiding the requirement for compression at high pressures and liquefaction at low temperatures. However, adsorption also alters the thermodynamics of key processes in the storage tank such as (i) refueling, (ii) discharging, (iii) dormancy, and (iv) boil-off. These effects are currently not well explored in the literature, especially for novel materials, because of the difficulties of conducting tank-scale experiments and process simulations. To alleviate this issue, we have developed an open-source python package capable of running lumped parameter dynamic simulations of sorbent-filled fluid tanks. In this paper, we explain the theory behind the model, and provide several case studies covering different applications to demonstrate the validity of our package and provide illustrative examples for potential users. Through the release of this package, we provide an easy-to-use platform for materials researchers to conduct process simulations on their novel sorbent materials, which could help ease their adoption in real industrial applications. • A python package simulating the thermodynamics of fluid storage tanks is presented. • It can simulate both conventional and sorbent-filled gas and liquid storage tanks. • The package (pytanksim) is free and open source for access by a wider audience. • Case studies are provided to validate the model and display its versatility. • These include liquid and cryo-compressed hydrogen, methane, and compressed CO 2 .
- Research Article
- 10.1016/j.softx.2026.102547
- Jun 1, 2026
- SoftwareX
Hyperspectral remote sensing captures hundreds of contiguous, narrow spectral bands in the VNIR and SWIR ranges, enabling detailed analysis of vegetation, water quality, soil properties, and other environmental variables. prismatools is an open-source Python package that facilitates processing, visualization, and analysis of PRISMA Level 2 products. It converts VNIR, SWIR and panchromatic PRISMA data into georeferenced xarray datasets, supporting seamless integration into workflows with other popular Python packages. The package also provides interactive mapping and spectral exploration leveraging the capabilities of the popular package Leafmap, along with methods for computing vegetation indices, performing PCA, extracting spectral signatures and exporting processed images.
- Research Article
108
- 10.1093/bioinformatics/btaa127
- Feb 26, 2020
- Bioinformatics
MotivationThe biological effects of human missense variants have been studied experimentally for decades but predicting their effects in clinical molecular diagnostics remains challenging. Available computational tools are usually based on the analysis of sequence conservation and structural properties of the mutant protein. We recently introduced a new machine learning method that demonstrated for the first time the significance of protein dynamics in determining the pathogenicity of missense variants.ResultsHere, we present a new interface (Rhapsody) that enables fully automated assessment of pathogenicity, incorporating both sequence coevolution data and structure- and dynamics-based features. Benchmarked against a dataset of about 20 000 annotated variants, the methodology is shown to outperform well-established and/or advanced prediction tools. We illustrate the utility of Rhapsody by in silico saturation mutagenesis studies of human H-Ras, phosphatase and tensin homolog and thiopurine S-methyltransferase.Availability and implementationThe new tool is available both as an online webserver at http://rhapsody.csb.pitt.edu and as an open-source Python package (GitHub repository: https://github.com/prody/rhapsody; PyPI package installation: pip install prody-rhapsody). Links to additional resources, tutorials and package documentation are provided in the 'Python package' section of the website.Supplementary informationSupplementary data are available at Bioinformatics online.
- Research Article
3
- 10.1177/11297298211015066
- May 10, 2021
- The Journal of Vascular Access
Dialysis vascular access, preferably an autogenous arteriovenous fistula, remains an end stage renal disease (ESRD) patient's lifeline providing a means of connecting the patient to the dialysis machine. Once an access is created, the current gold standard of care for maintenance of vascular access is angiography and angioplasty to treat stenosis. While point of care 2D ultrasound has been used to detect access problems, we sought to reproduce angiographic results comparable to the gold standard angiogram (fistulogram) using ultrasound data acquired from a conventional 2D ultrasound scanner. A 2D ultrasound probe was used to acquire a series of cross sectional images of the vascular access including arteriovenous anastomosis of a subject with a radio-cephalic fistula. These 2D B-mode images were used for 3D vessel reconstruction by binary thresholding to categorize vascular versus non-vascular structures followed by standard image segmentation to select the structure representative of dialysis vascular access and morphologic filtering. Image processing was done using open source Python Software. The open source software was able to: (1) view the gold standard fistulogram images, (2) reconstruct 2D planar images of the fistula from ultrasound data as viewed from the top, analogous to computerized tomography images, and (3) construct a 2D representation of vascular access similar to the angiogram. We present a simple approach to obtain an angiogram-like representation of the vascular access from readily available, non-proprietary 2D ultrasound data in the point of care setting. While the sono-angiogram is not intended to replace angiography, it may be useful in providing 3D imaging at the point of care in the dialysis unit, outpatient clinic, or for pre-operative planning for interventional procedures. Future work will focus on improving the robustness and quality of the imaging data while preserving the straightforward freehand approach used for ultrasound data acquisition.