An orchestra of machine learning methods reveals landmarks in single-cell data exemplified with aging fibroblasts.

Lauritz Rasbach,Aylin Caliskan,Fatemeh Saderi,Thomas Dandekar,Tim Breitenbach

doi:10.1371/journal.pone.0302045

Abstract

In this work, a Python framework for characteristic feature extraction is developed and applied to gene expression data of human fibroblasts. Unlabeled feature selection objectively determines groups and minimal gene sets separating groups. ML explainability methods transform the features correlating with phenotypic differences into causal reasoning, supported by further pipeline and visualization tools, allowing user knowledge to boost causal reasoning. The purpose of the framework is to identify characteristic features that are causally related to phenotypic differences of single cells. The pipeline consists of several data science methods enriched with purposeful visualization of the intermediate results in order to check them systematically and infuse the domain knowledge about the investigated process. A specific focus is to extract a small but meaningful set of genes to facilitate causal reasoning for the phenotypic differences. One application could be drug target identification. For this purpose, the framework follows different steps: feature reduction (PFA), low dimensional embedding (UMAP), clustering ((H)DBSCAN), feature correlation (chi-square, mutual information), ML validation and explainability (SHAP, tree explainer). The pipeline is validated by identifying and correctly separating signature genes associated with aging in fibroblasts from single-cell gene expression measurements: PLK3, polo-like protein kinase 3; CCDC88A, Coiled-Coil Domain Containing 88A; STAT3, signal transducer and activator of transcription-3; ZNF7, Zinc Finger Protein 7; SLC24A2, solute carrier family 24 member 2 and lncRNA RP11-372K14.2. The code for the preprocessing step can be found in the GitHub repository https://github.com/AC-PHD/NoLabelPFA, along with the characteristic feature extraction https://github.com/LauritzR/characteristic-feature-extraction.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: PLOS ONE	Publication Date: Apr 17, 2024
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

An orchestra of machine learning methods reveals landmarks in single-cell data exemplified with aging fibroblasts.

Abstract

Talk to us

Similar Papers

More From: PLOS ONE

Lead the way for us

Similar Papers

Artificial Zinc Finger Fusions Targeting Sp1-binding Sites and the trans-Activator-responsive Element Potently Repress Transcription and Replication of HIV-1
Yeon-Soo Kim ... Man-Wook Hur
Journal of Biological Chemistry | VOL. 280
Yeon-Soo Kim, et. al.Yeon-Soo Kim ... Man-Wook Hur
01 Jun 2005
Journal of Biological Chemistry | VOL. 280

The Role of Zinc Finger Protein 521/Early Hematopoietic Zinc Finger Protein in Erythroid Cell Differentiation
Etsuko Matsubara ... Masaki Yasukawa
Journal of Biological Chemistry | VOL. 284
Etsuko Matsubara, et. al.Etsuko Matsubara ... Masaki Yasukawa
01 Feb 2009
Journal of Biological Chemistry | VOL. 284

A Synthetic Biology Framework for Programming Eukaryotic Transcription Functions
Ahmad S Khalil ... James J Collins
Cell | VOL. 150
Ahmad S Khalil, et. al.Ahmad S Khalil ... James J Collins
01 Aug 2012
Cell | VOL. 150

Aiolos, a lymphoid restricted transcription factor that interacts with Ikaros to regulate lymphocyte differentiation.
B Morgan
The EMBO Journal | VOL. 16
B MorganB Morgan
15 Apr 1997
The EMBO Journal | VOL. 16

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An orchestra of machine learning methods reveals landmarks in single-cell data exemplified with aging fibroblasts.

Abstract

Talk to us

Similar Papers

More From: PLOS ONE