StrainedSMILES2xyz: a workflow for reliable 3D structures of strained molecules from SMILES.
Accurate 3D structure generation from SMILES is essential for data-driven chemistry but often fails for strained ring systems. We introduce strainedSMILES2xyz, a Python workflow that improves conformer generation by relaxing RDKit constraints, exploring stereoisomer variants, and correcting errors using force-field refinement. Benchmarking on strained and unstrained rings shows that it outperforms existing tools, generating correct geometries in nearly all cases. The workflow is available as a Python package and Jupyter notebook.Scientific ContributionThis work identifies a critical gap in automated SMILES-to-3D structure generation for strained molecules. Many established tools, including the widely used RDKit, frequently fail or produce incorrect geometries for these systems. By explicitly targeting these failure modes, the proposed approach enables reliable 3D structure generation for chemically relevant strained molecules within fully automated workflows.
- Preprint Article
- 10.1002/essoar.10510233.1
- Jan 21, 2022
This repository creates a GUI (graphical user interface) for the BALTO (Brokered Alignment of Long-Tail Observations) project. BALTO is funded by the NSF EarthCube program. The GUI aims to provide a simplified and customizable method for users to access data sets of interest on servers that support the OpenDAP data access protocol. This interactive GUI runs within a Jupyter notebook and uses the Python packages: ipywidgets (for widget controls), ipyleaflet (for interactive maps) and pydap (an OpenDAP client). The Python source code to create the GUI and to process events is in a module called balto_gui.py that must be found in the same directory as this Jupyter notebook. Python source code for visualization of downloaded data is given in a module called balto_plot.py. This GUI consists of mulitiple panels, and supports both a tab-style and an accordion-style, which allows you to switch between GUI panels without scrolling in the notebook. You can run the notebook in a browser window without installing anything on your computer, using something called Binder. Look for the Binder icon below and a link labeled “Launch Binder”. This sets up a server in the cloud that has all the required dependencies and lets you run the notebook on that server. (Sometimes this takes a while, however.) To run this Jupyter notebook without Binder, it is recommended to install Python 3.7 from an Anaconda distribution and to then create a conda environment called balto. Simple instructions for how to create a conda environment and install the software are given in Appendix 1 of version 2 (v2) of the notebook.
- Research Article
1
- 10.1093/molbev/msag061
- Mar 9, 2026
- Molecular Biology and Evolution
piqtree (pronounced pie-cue-tree) is an easy to use, open-source Python package that provides Python script based control of IQ-TREE’s phylogenetic inference engine. piqtree builds IQ-TREE as a Python package, presenting a library of Python functions for performing many of IQ-TREE’s capabilities including phylogenetic reconstruction, ultrafast bootstrapping, branch length optimization, model selection, rapid neighbor-joining, alignment simulation, and more. As piqtree explicitly uses IQ-TREE’s phylogenetic algorithms, the computational and statistical performance of piqtree equal that of IQ-TREE. Modestly higher memory usage may be expected owing to the Python runtime and the need to load the alignment in Python. By exposing IQ-TREE’s algorithms within Python, piqtree offers users a greatly simplified experience in development of phylogenetic workflows through seamless interoperability with other Python libraries and tools mediated by the cogent3 package. It enables users to perform interactive phylogenetic analyses and visualization using, for instance, Jupyter notebooks. We present the key features available in the piqtree library and a small case study that showcases its interoperability and highlight its potential for linking a high performance phylogenetic inference engine with more user friendly interfaces. piqtree is distributed for use as a standard Python package at https://pypi.org/project/piqtree/, documentation is available at https://piqtree.readthedocs.io, user contributed solutions at https://github.com/cogent3/c3codeshare, help forums at https://github.com/iqtree/piqtree/discussions, and source code at https://github.com/iqtree/piqtree.
- Research Article
1
- 10.1093/bioinformatics/btae614
- Oct 1, 2024
- Bioinformatics (Oxford, England)
Molecular dynamics simulation is very useful but computationally demanding method of studying dynamics of biomolecular systems. Many enhanced sampling methods were developed in order to obtain the desired results in available computational time. Metadynamics and its variants are common enhanced sampling methods used for this purpose. Metadynamics simulations allow the user to gather large amounts of data, which have to be analyzed to elucidate the properties of the studied system. Here, we present metadynminer.py, a Python package that allows easy and user-friendly analysis and visualization of the results obtained from metadynamics simulations. The built-in functions automate frequent tasks and make the package easy to use for new users, while its many customization options and object-oriented nature allow for integration into specialized data analysis workflows by more advanced users. The "metadynminer.py" Python package is available under the GPL-3.0 license via PyPi and Conda. The development version is available on GitHub along with issue support (https://github.com/Jan8be/metadynminer.py). Documentation, tutorial and Jupyter Notebook (provided through the public mybinder.org service) are available at https://metadynreporter.cz.
- Preprint Article
- 10.5194/egusphere-egu25-6691
- Mar 18, 2025
In environmental data analysis, source apportionment can be an important approach to extract useful information that might otherwise be hidden within the data. The United States Environmental Protection Agency (EPA) has developed an open-source python package, the Environmental Source Apportionment Toolkit (ESAT), which enables source apportionment modeling and error estimation workflows. ESAT is intended to replace Positive Matrix Factorization v5 (PMF5) that has substantial data size limitations. ESAT is currently in alpha testing with development plans for enhanced functionality and support of large datasets, High-performance Computing (HPC) execution through a command line interface (CLI), and a standalone desktop graphical user interface (GUI). The alpha product of ESAT is publicly available and offers a complete application programming interface (API) to replicate the workflows and functionality of PMF5, with examples provided through Jupyter Notebooks. The ESAT computing module currently contains two non-negative matrix factorization (NMF) algorithms for model training, with the module designed for other algorithms to be easily added. The two algorithms currently available are the least-squares NMF (LS-NMF) and weighted-semi NMF (WS-NMF). Each algorithm offers different benefits depending on project or data requirements. The ESAT python codebase has been optimized to run in a highly parallelized manner, with most of the numerical computations implemented in Rust, a low-level language comparable in performance to C. ESAT replicates the model error estimation methods of PMF5, namely bootstrap, displacement, and a hybrid method. To facilitate experimentation and testing, ESAT contains a synthetic dataset generator and model simulator that can evaluate how well ESAT can recreate synthetic factors and contributions. Continuous development of new features are tested and added to the python package on a regular basis. One such feature is the addition of an uncertainty perturbation workflow, which will run a collection of models while slightly perturbing the uncertainty matrix, and then evaluating the impact on the solution profiles and contributions. The alpha version of the ESAT python package is available for installation from pypi at https://pypi.org/project/esat/. Further testing and development of the alpha version will proceed to a full release in late 2025. The development of a GUI desktop application is currently planned to begin after the ESAT full release.
- Research Article
8
- 10.3389/fbinf.2022.827024
- Feb 25, 2022
- Frontiers in bioinformatics
The human upper respiratory tract is the reservoir of a diverse community of commensals and potential pathogens (pathobionts), including Streptococcus pneumoniae (pneumococcus), Haemophilus influenzae, Moraxella catarrhalis, and Staphylococcus aureus, which occasionally turn into pathogens causing infectious diseases, while the contribution of many nasal microorganisms to human health remains undiscovered. To better understand the composition of the nasal microbiome community, we create a workflow of the community model, which mimics the human nasal environment. To address this challenge, constraint-based reconstruction of biochemically accurate genome-scale metabolic models (GEMs) networks of microorganisms is mandatory. Our workflow applies constraint-based modeling (CBM), simulates the metabolism between species in a given microbiome, and facilitates generating novel hypotheses on microbial interactions. Utilizing this workflow, we hope to gain a better understanding of interactions from the metabolic modeling perspective. This article presents nasal community modeling workflow (NCMW)—a python package based on GEMs of species as a starting point for understanding the composition of the nasal microbiome community. The package is constructed as a step-by-step mathematical framework for metabolic modeling and analysis of the nasal microbial community. Using constraint-based models reduces the need for culturing species in vitro, a process that is not convenient in the environment of human noses. Availability: NCMW is freely available on the Python Package Index (PIP) via pip install NCMW. The source code, documentation, and usage examples (Jupyter Notebook and example files) are available at https://github.com/manuelgloeckler/ncmw.
- Preprint Article
- 10.5194/egusphere-egu22-13193
- Mar 28, 2022
<p>As machine learning algorithms are being used more and more prominently in the meteorology and climate domains, the need for reference datasets has been identified as a priority. Moreover, boilerplate code for data handling is ubiquitous in scientific experiments. In order to focus on science, climate/meteorology/data scientists need generic and reusable domain-specific tools. To achieve these goals, we used the plugin based CliMetLab python package along with many packages listed by Pangeo.  </p><p><br>Our use case consists in providing data for machine learning algorithms in the context of the sub-seasonal to seasonal (S2S) prediction challenge 2021. The data size is about 2 Terabytes of model predictions from three different models. We experimented with providing data in multiple formats: Grib, NetCDF, and Zarr. A Pangeo recipe (using the python package pangeo_forge_recipes) was used to generate Zarr data (relying heavily on xarray and dask for parallelisation). All three versions of the S2S data have been stored on an S3 bucket located on the ECMWF European Weather Cloud (ECMWF-EWC). </p><p><br>CliMetLab aims at providing a simple interface to access climate and meteorological datasets, seamlessly downloading and caching data, converting to xarray datasets or panda dataframes, plotting data, feed them into machine learning frameworks such as tensorflow or pytorch. CliMetLab is open-source and still a Beta version (https://climetlab.readthedocs.io). The main target platform of CliMetLab is Jupyter notebooks. Additionally, a CliMetLab plugin allows shipping dataset-specific code along with a well-defined published dataset. Taking advantage of the CliMetLab tools to minimize the boilerplate code, a plugin has been developed for S2S data as a companion python package of the dataset.</p>
- Preprint Article
- 10.5194/epsc-dps2025-825
- Jul 9, 2025
Introduction: Coma images of active comets are commonly used to identify inhomogeneities that appear as coma features, which provide key insights into coma dynamics and nucleus properties [e.g., 1, 2]. We have developed a 3D Monte Carlo model to simulate such features and have successfully applied it to interpret cometary activity and nucleus characteristics in prior studies [e.g., 3, 4, 5]. We are currently preparing to release this software as a Tool to the wider astronomy community via a user-friendly web interface, Coma Factory. A companion Python package containing the underlying codes will also be made available. Our presentation will include a live demonstration of Coma Factory using both the web interface and the Python package in a Jupyter Notebook. We welcome community feedback to help guide the development of comprehensive tutorials and ensure the Tool’s broad utility.Coma Model: The 3D Monte Carlo coma model simulates the spatial distribution of particles – first-, second-, or third-generation species – emitted from a nucleus. The resulting 3D particle distribution is then projected onto a 2D skyplane corresponding to the observer’s line of sight (e.g., Figure 1), enabling direct comparison with observed coma images. The model requires input parameters that describe the comet’s activity, nucleus properties, particle-specific species characteristics, and observing geometry. It also includes the capability to simulate the effects of solar radiation pressure on all species.Coma Factory: When using Coma Factory, once the user provides the necessary input parameters, the web interface runs the Python code to generate the corresponding image in FITS format. The image header includes keywords identifying all the input parameters and hence each image created by Coma Factory maintains a digital record of the input parameters corresponding to each image. Additionally, Coma Factory may be used to produce 2D coma simulations corresponding to a range of viewing directions that enable the user to assess the 3D structure of the coma feature(s) that are useful for interpreting actual coma observations of comets.Figure 1. Top-left: Image of the active Centaur 2023 RS61 from [6]. Top-right: shown is a simulated image using Coma Factory that is consistent with the data. The bottom two panels show the same simulated image but if the data had been acquired with other facilities with different image resolutions. The same coma model was used for each of the three image simulations, where characteristics of each telescope and instrument combination were used to replicate imaging data if observations had been acquired from the specified facility. Coma Factory is capable of generating realistic comae images as illustrated by these simulations.Acknowledgments: We gratefully acknowledge the support provided by the NASA PDART Program through award 80NSSC21K0881.
- Research Article
2
- 10.5334/jors.558
- Jul 23, 2025
- Journal of Open Research Software
OGRePy is a modern, open-source Python package designed to perform symbolic tensor calculations, with a particular focus on applications in general relativity. Built on an object-oriented architecture, OGRePy encapsulates tensors, metrics, and coordinate systems as self-contained objects, automatically handling raising and lowering of indices, coordinate transformations, contractions, partial or covariant derivatives, and all tensor operations. By leveraging the capabilities of SymPy and Jupyter Notebook, OGRePy provides a robust, user-friendly environment that facilitates both research and teaching in general relativity and differential geometry. This Python package reproduces the functionality of the popular Mathematica package OGRe, while greatly improving upon it by making use of Python’s native object-oriented syntax. In this paper, we describe OGRePy’s design and implementation, and discuss its potential for reuse across research and education in mathematics and physics.
- Research Article
23
- 10.1371/journal.pone.0215137
- Apr 11, 2019
- PLoS ONE
Hybrid 3D scaffolds composed of different biomaterials with fibrous structure or enriched with different inclusions (i.e., nano- and microparticles) have already demonstrated their positive effect on cell integration and regeneration. The analysis of fibers in hybrid biomaterials, especially in a 3D space is often difficult due to their various diameters (from micro to nanoscale) and compositions. Though biomaterials processing workflows are implemented, there are no software tools for fiber analysis that can be easily integrated into such workflows. Due to the demand for reproducible science with Jupyter notebooks and the broad use of the Python programming language, we have developed the new Python package quanfima offering a complete analysis of hybrid biomaterials, that include the determination of fiber orientation, fiber and/or particle diameter and porosity. Here, we evaluate the provided tensor-based approach on a range of generated datasets under various noise conditions. Also, we show its application to the X-ray tomography datasets of polycaprolactone fibrous scaffolds pure and containing silicate-substituted hydroxyapatite microparticles, hydrogels enriched with bioglass contained strontium and alpha-tricalcium phosphate microparticles for bone tissue engineering and porous cryogel 3D scaffold for pancreatic cell culturing. The results obtained with the help of the developed package demonstrated high accuracy and performance of orientation, fibers and microparticles diameter and porosity analysis.
- Research Article
4
- 10.1167/tvst.12.2.6
- Feb 6, 2023
- Translational Vision Science & Technology
PurposeArtificial intelligence (AI) methods are changing all areas of research and have a variety of capabilities of analysis in ophthalmology, specifically in visual fields (VFs) to detect or predict vision loss progression. Whereas most of the AI algorithms are implemented in Python language, which offers numerous open-source functions and algorithms, the majority of algorithms in VF analysis are offered in the R language. This paper introduces PyVisualFields, a developed package to address this gap and make available VF analysis in the Python language.MethodsFor the first version, the R libraries for VF analysis provided by vfprogression and visualFields packages are analyzed to define the overlaps and distinct functions. Then, we defined and translated this functionality into Python with the help of the wrapper library rpy2. Besides maintaining, the subsequent versions’ milestones are established, and the third version will be R-independent.ResultsThe developed Python package is available as open-source software via the GitHub repository and is ready to be installed from PyPI. Several Jupyter notebooks are prepared to demonstrate and describe the capabilities of the PyVisualFields package in the categories of data presentation, normalization and deviation analysis, plotting, scoring, and progression analysis.ConclusionsWe developed a Python package and demonstrated its functionality for VF analysis and facilitating ophthalmic research in VF statistical analysis, illustration, and progression prediction.Translational RelevanceUsing this software package, researchers working on VF analysis can more quickly create algorithms for clinical applications using cutting-edge AI techniques.
- Research Article
- 10.1093/bioadv/vbag052
- Jan 7, 2026
- Bioinformatics advances
Illumina DNA methylation arrays have evolved rapidly, expanding genomic coverage while introducing backward incompatibilities by removing many CpG sites present in earlier versions. These changes result in systematic missing values when integrating data across array generations and substantially limiting the reuse of legacy datasets. We developed a two-stage framework for imputing missing DNA methylation values. The procedure first imputes randomly missing values using standard imputation techniques and then addresses systematic missingness using multi-output machine learning models, including support vector regression, nearest-neighbor methods, random forest models, and deep neural networks. When evaluated on real datasets with up to fifty percent induced missingness, the proposed framework consistently outperformed conventional imputation approaches. It also accurately imputes the missing CpG sites between methylation arrays and reduced representation bisulfite sequencing data, enabling robust cross-platform data integration. Analyses of large brain tumor methylation datasets demonstrate that the method restores array-specific methylation patterns while preserving biological complexity. Importantly, imputing missing methylation sites significantly improves the performance of epigenetic age prediction models. This tool is implemented in the Python package "ultra-impute," freely available at https://github.com/liguowang/ultra-impute. A code snippet demonstrating the usage of the ultra-impute package is provided in a Jupyter Notebook (https://github.com/liguowang/ultra-impute/blob/master/doc/Tutorial.ipynb).
- Research Article
184
- 10.1109/tvcg.2020.3030378
- Oct 13, 2020
- IEEE Transactions on Visualization and Computer Graphics
Natural language interfaces (NLls) have shown great promise for visual data analysis, allowing people to flexibly specify and interact with visualizations. However, developing visualization NLIs remains a challenging task, requiring low-level implementation of natural language processing (NLP) techniques as well as knowledge of visual analytic tasks and visualization design. We present NL4DV, a toolkit for natural language-driven data visualization. NL4DV is a Python package that takes as input a tabular dataset and a natural language query about that dataset. In response, the toolkit returns an analytic specification modeled as a JSON object containing data attributes, analytic tasks, and a list of Vega-Lite specifications relevant to the input query. In doing so, NL4DV aids visualization developers who may not have a background in NLP, enabling them to create new visualization NLIs or incorporate natural language input within their existing systems. We demonstrate NL4DV's usage and capabilities through four examples: 1) rendering visualizations using natural language in a Jupyter notebook, 2) developing a NLI to specify and edit Vega-Lite charts, 3) recreating data ambiguity widgets from the DataTone system, and 4) incorporating speech input to create a multimodal visualization system.
- Research Article
51
- 10.1080/15502724.2018.1518717
- May 2, 2019
- LEUKOS
ABSTRACTLuxPy is a free and open source Python package that supports several common lighting, colorimetric, color appearance, and other color science related calculations that should be useful to researchers and industry professionals. This article describes the installation of the LuxPy toolbox, provides an overview of its basic design and functionality, and gives several examples to demonstrate its basic utilities using a Jupyter notebook. LuxPy is available under a GPLv3 license at www.github.com/ksmet1977/luxpy/ or from the Python Package Index pypi.python.org/pypi/luxpy/.
- Research Article
- 10.1186/s12859-026-06474-4
- May 29, 2026
- BMC bioinformatics
The domain-linker-domain (DLD) architecture is commonly found in proteins, where flexible linkers connect consecutive domains and regulate their relative spatial positioning. Often, these linkers present partially structured elements that modulate inter-domain dynamics, directly influencing their function. From a protein design perspective, tuning the relative position and orientation of domains via the linker offers opportunities to modulate biological activity. Despite their relevance, analyzing conformational ensembles of DLD proteins remains a challenge, thereby limiting the structural insights that can be extracted. We present DL3D, a robotics-inspired visualization tool that enables intuitive analysis of the conformational space sampled by DLD proteins. DL3D discretizes the relative positions of the two domains at the linker ends and projects each conformation onto a 3D voxel map, where density is represented in grayscale to highlight the most probable configurations. In addition, quaternion-based operations allow the analysis of relative domain orientations. DL3D facilitates the structural investigation of highly flexible proteins composed of well-folded domains connected by flexible linkers. Beyond visualization, the tool supports downstream analyses such as low-dimensional conformational clustering. DL3D is implemented as a Python package and is available at: https://gitlab.laas.fr/moma/methods/analysis/dl3d/. A Jupyter notebook with usage examples is also provided.
- Preprint Article
- 10.59350/h8fm4-gm830
- Aug 17, 2020
- Front Matter
RDKit is a cheminformatics toolkit with bindings for Python. It's packed with functionality, deployed within multiple open source projects, and is widely-used in machine learning applications. RDKit can also be difficult to install. This article discusses the problem and a method for using RDKit within Jupyter notebooks. Installation Options The Python Package Index (aka PyPI, aka pip) is Python's standard package manager.