Research Summary
Our research focuses on the development of bioinformatics and statistical approaches to address challenges posed in biological inferences from high-throughput proteomics data, and their application to biomedical problems. For this purpose, we develop algorithms for peak detection and quantification, identification of structures in multivariate data, stochastic time-course modeling to extract dynamical features, construction of protein networks and error control in the resulting inferences. In collaboration with our colleagues, experimentalists, we apply these techniques in various systems for systematic studies of post-translational modifications, proteome dynamics, signal transductions and mass informatics. Our goal is to promote identification of functional dysregulations associated with changes in the state of a biological system. An important unit of our research is a mass spectrum.

Outlined below are examples of our research work.
Protein homeostasis (proteostasis) is achieved via continuous synthesis and degradation of cellular proteins. Proteostasis is important for proper protein functioning, and it is dysregulated in many diseases. Metabolic labeling followed by LC-MS is a powerful technique to study proteostasis in a large scale (thousands of proteins).
Borzou, Ahmad, Razie Yousefi, and Rovshan G. Sadygov. Another look at matrix correlations. Bioinformatics 35.22 (2019): 4748-4753.2019, Pages 4748–4753

Metabolic labeling with heavy water followed by LC-MS is a high throughput approach to study proteostasis in vivo. Advances in mass spectrometry and sample processing have allowed consistent detection of thousands of proteins at multiple time points. However, freely available automated bioinformatics tools to analyze and extract protein decay rate constants are lacking. Here, we describe d2ome-a robust, automated software solution for in vivo protein turnover analysis. d2ome is highly scalable, uses innovative approaches to nonlinear fitting, implements Grubbs’ outlier detection and removal, uses weighted-averaging of replicates, applies a data dependent elution time windowing, and uses mass accuracy in peak detection. Here, we discuss the application of d2ome in a comparative study of protein turnover in the livers of normal vs Western diet-fed LDLR-/- mice (mouse model of nonalcoholic fatty liver disease), which contained 256 LC-MS experiments. The study revealed reduced stability of 40S ribosomal protein subunits in the Western diet-fed mice.
Sadygov, Rovshan G., et al. d2ome, software for in vivo protein turnover analysis using heavy water labeling and LC–MS, reveals alterations of hepatic proteome dynamics in a mouse model of NAFLD. Journal of proteome research 17.11 (2018): 3740-3748.

We developed a formula-based stochastic simulation strategy for TPS for in vivo studies with heavy water metabolic labeling and LC-MS. We model the rate constant (lognormal), measurement error (Laplace), peptide length (Gamma), relative abundance (RA) of the monoisotopic peak (beta regression), and the number of exchangeable hydrogens (Gamma regression). The parameters of the distributions are determined using corresponding empirical probability density functions from a large-scale dataset of murine heart proteome. The models are used in simulations of the rate constant to minimize the root-mean-squared error.
Sadygov, Vugar R., William Zhang, and Rovshan G. Sadygov. Timepoint Selection Strategy for In Vivo Proteome Dynamics from Heavy Water Metabolic Labeling and LC–MS. Journal of proteome research 19.5 (2020): 2105-2112.