The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification

Luca Narduzzi; Jan Stanstrup; Fulvio Mattivi; Pietro Franceschi

doi:10.1080/19440049.2018.1523572

The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification

Luca Narduzzi, Jan Stanstrup, Fulvio Mattivi, Pietro Franceschi

Miljøkemi

3 Citationer (Scopus)

Abstract

Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of ‘unknown’ metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called ‘Compounds Characteristics Comparison’ (CCC) emulates the experience of a plant chemist who ‘gains experience’ from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users’ confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices.

Originalsprog	Engelsk
Tidsskrift	Food Additives & Contaminants: Part A
Vol/bind	35
Udgave nummer	11
Sider (fra-til)	2145-2157
Antal sider	13
ISSN	1944-0049
DOI	https://doi.org/10.1080/19440049.2018.1523572
Status	Udgivet - 2018

Adgang til dokumentet

10.1080/19440049.2018.1523572

Citationsformater

@article{fdad35b76c35472583ecbff95a60e27e,

title = "The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification",

abstract = "Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of {\textquoteleft}unknown{\textquoteright} metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called {\textquoteleft}Compounds Characteristics Comparison{\textquoteright} (CCC) emulates the experience of a plant chemist who {\textquoteleft}gains experience{\textquoteright} from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users{\textquoteright} confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices.",

keywords = "candidate selection, grape, isotopic pattern, LC-HRMS, machine learning, metabolomics, model building, substructure recognition",

author = "Luca Narduzzi and Jan Stanstrup and Fulvio Mattivi and Pietro Franceschi",

year = "2018",

doi = "10.1080/19440049.2018.1523572",

language = "English",

volume = "35",

pages = "2145--2157",

journal = "Food Additives & Contaminants: Part A",

issn = "1944-0049",

publisher = "Taylor & Francis Online",

number = "11",

}

TY - JOUR

T1 - The Compound Characteristics Comparison (CCC) approach

T2 - a tool for improving confidence in natural compound identification

AU - Narduzzi, Luca

AU - Stanstrup, Jan

AU - Mattivi, Fulvio

AU - Franceschi, Pietro

PY - 2018

Y1 - 2018

N2 - Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of ‘unknown’ metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called ‘Compounds Characteristics Comparison’ (CCC) emulates the experience of a plant chemist who ‘gains experience’ from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users’ confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices.

AB - Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of ‘unknown’ metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called ‘Compounds Characteristics Comparison’ (CCC) emulates the experience of a plant chemist who ‘gains experience’ from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users’ confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices.

KW - candidate selection

KW - grape

KW - isotopic pattern

KW - LC-HRMS

KW - machine learning

KW - metabolomics

KW - model building

KW - substructure recognition

U2 - 10.1080/19440049.2018.1523572

DO - 10.1080/19440049.2018.1523572

M3 - Journal article

C2 - 30352003

SN - 1944-0049

VL - 35

SP - 2145

EP - 2157

JO - Food Additives & Contaminants: Part A

JF - Food Additives & Contaminants: Part A

IS - 11

ER -

The Compound Characteristics Comparison (CCC) approach: a tool for improving confidence in natural compound identification

Abstract

Adgang til dokumentet

Fingeraftryk

Citationsformater