Direct Maximization of Protein Identifications from Tandem Mass Spectra
M. A. Spivak, Jason Weston, Daniela M. Tomazela, Michael J. MacCoss, William Stafford Noble
- Year
- 2011
- Citations
- 32
- Access
- Open access
Abstract
The goal of many shotgun proteomics experiments is to determine the protein complement of a complex biological mixture. For many mixtures, most methodological approaches fall significantly short of this goal. Existing solutions to this problem typically subdivide the task into two stages: first identifying a collection of peptides with a low false discovery rate and then inferring from the peptides a corresponding set of proteins. In contrast, we formulate the protein identification problem as a single optimization problem, which we solve using machine learning methods. This approach is motivated by the observation that the peptide and protein level tasks are cooperative, and the solution to each can be improved by using information about the solution to the other. The resulting algorithm directly controls the relevant error rate, can incorporate a wide variety of evidence and, for complex samples, provides 18–34% more protein identifications than the current state of the art approaches. The goal of many shotgun proteomics experiments is to determine the protein complement of a complex biological mixture. For many mixtures, most methodological approaches fall significantly short of this goal. Existing solutions to this problem typically subdivide the task into two stages: first identifying a collection of peptides with a low false discovery rate and then inferring from the peptides a corresponding set of proteins. In contrast, we formulate the protein identification problem as a single optimization problem, which we solve using machine learning methods. This approach is motivated by the observation that the peptide and protein level tasks are cooperative, and the solution to each can be improved by using information about the solution to the other. The resulting algorithm directly controls the relevant error rate, can incorporate a wide variety of evidence and, for complex samples, provides 18–34% more protein identifications than the current state of the art approaches. The problem of identifying proteins from a collection of tandem mass spectra involves assigning spectra to peptides, using either a de novo or database search strategy, and then inferring the protein set from the resulting collection of peptide-spectrum matches (PSMs). 1The abbreviations used are:PSMpeptide-spectrum matchFDRfalse discovery rateHUhidden units. In practice, the goal of such an experiment is to identify as many distinct proteins as possible at a specified false discovery rate (FDR). However, most of the previous work in the context of shotgun proteomics analysis has focused on controlling error rates at the level of PSMs or peptides (1Moore R.E. Young M.K. Lee T.D. Qscore: An algorithm for evaluating Sequest database search results.J. Am. Soc. Mass Spectrom. 2002; 13: 378-386Crossref PubMed Scopus (339) Google Scholar, 2Choi H. Ghosh D. Nesvizhskii A.I. Statistical validation of peptide identifications in large-scale proteomics using target-decoy database search strategy and flexible mixture modeling.J. Proteome Res. 2008; 7: 286-292Crossref PubMed Scopus (100) Google Scholar, 3Käll L. Storey J.D. MacCoss M.J. Noble W.S. Assigning significance to peptides identified by tandem mass spectrometry using decoy databases.J. Proteome Res. 2008; 7: 29-34Crossref PubMed Scopus (449) Google Scholar, 4Choi H. Nesvizhskii A.I. False discovery rates and related statistical concepts in mass spectrometry-based proteomics.J. Proteome Res. 2008; 7: 47-50Crossref PubMed Scopus (170) Google Scholar, 5Keller A. Nesvizhskii A.I. Kolker E. Aebersold R. Empirical statistical model to estimate the accuracy of peptide identification made by MS/MS and database search.Anal. Chem. 2002; 74: 5383-5392Crossref PubMed Scopus (3928) Google Scholar, 6Fenyö D. Beavis R.C. A method for assessing the statistical significance of mass spectrometry-based protein identification using general scoring schemes.Anal. Chem. 2003; 75: 768-774Crossref PubMed Scopus (407) Google Scholar, 7Geer
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
Genetic Programming: On the Programming of Computers by Means of Natural Selection
John R. Koza
1992