The Human Metabolome Project (HMP) is a Canadian initiative which aims to identify and quantify all known and unknown metabolites in human tissues and biofluids. Metabolites are smaller than DNA or protein and have molecular weights < 1500 Daltons. The collection of these metabolites is known as the metabolome. The HMP was conceived, designed, and led by Dr. David Wishart at the University of Alberta. The HMP team consists of spectroscopists, organic chemists, physicians and bioinformaticians.[1][2] Since its inception in January 2005, the HMP has and continues to build databases such as the Human Metabolome Database[3] and computational resources such as MetaboAnalyst[4] that support the metabolomics community in hopes of better understanding how our body works.
Background and History
Around 2000, metabolomics, compared to sequencing the human genome, was in the developing phase. This is reflected in the number of scientific publications reporting metabolomics studies at that time. In 1998, only 1-2 metabolomics papers were published per year. However, interest in metabolomics research was mounting, with the number of publications increasing per year. However, these findings and discoveries were not readily shared, largely staying in those articles.[5] In 1999, Dr. Wishart published a paper that used machine learning to enable automated metabolite identification and quantification from nuclear magnetic resonance (NMR) spectra of a human biofluid[6]. However, he realized that there were few comprehensive resources to support this new type of work. He applied for funding from Genome Canada (GC) in 2004 for the project entitled "Building the Metabolomics Toolbox Enabling Rapid Disease Diagnosis through Metabolic Profiling" [7] and was successful. Dr. Wishart secured about $7.5 million to catalog the "chemical equivalent of the human genome" with half of the funding from GC, and the other half from Canadian Foundation for Innovation (CFI), Alberta Ingenuity Centre for Machine Learning and the University of Alberta.[2] The early HMP team had expertise in analytical chemistry, medicine, computer science and biochemistry and were largely from the University of Alberta (Wishart, Greiner, Bamforth, Marrie, Sykes, Li) and the University of Calgary (Vogel).[1] The HMP aims were to identify, quantify, index and store all known and unknown human metabolites at concentrations greater than 1 micromolar (цM), to create an electronic database and to create a physical library of compounds available for distribution so others could more easily pursue metabolomics studies.[8] Initial efforts were focused on enumerating all known human metabolites through literature surveys and experimental studies. The HMP researchers then moved towards elucidating unknown metabolites via spectral prediction or biotransformation predictions using computer science.[9]
HMDB and metabolomes. Human metabolites are challenging to characterize because they can originate from over 200 cell types and dozens of organs, their concentrations can differ by several orders of magnitude (for example, picomolar quantities for hormones to nearly molar amounts of urea[10]) and include both known and unknown compounds.[11] Over 2 years, 53 scientists combed scientific literature and performed experiments to identify and quantify 2500 metabolites in the first release of HMDB.
Legacy and Impact
Since that first $7.5 million in funding from Genome Canada and others, over $18 million[2] of subsequent funding has enabled the HMP to continue to study the known and unknown metabolites in the human body and to support the metabolomics community research efforts. As metabolites can originate from both exogenous (originating within the organism) and endogenous (from food, drugs, microbiota, cosmetics, and environment) sources, HMP researchers have created resources to comprehensively characterize these metabolites. With the fifth version of the HMDB, over 200,000 endogenous metabolites have been identified. The HMDB continues to be updated to the present day with the current version housing 253,245 metabolites.[10] Databases for the exogenous metabolites have also been developed by the HMP team. These databases contain additional details not included in the HMDB.[11] DrugBank, which compiles information about drugs and drug actions, includes details about pharmacokinetics, drug mechanism of action and absorption.[12] The Toxin Toxin Target database (T3DB), which captures information about the toxic exposome (the sum of all toxic exposures from birth to death[13]), also includes toxic threshold data (lethal dose 50 or LD50) and treatment options. FooDB contains details about food constituents[14] and the Microbial Metabolome Database has compiled data about the microbiome, their genomes and metabolites related to human health and disease.[15]
In addition to characterizing known human metabolites, the HMP has worked to identify unknown metabolites using machine learning and artificial intelligence. It has been reported that only 5% of metabolites in peaks from samples run from specialized instruments can be identified, meaning 95% are unknown.[9] The HMP team continues to develop software to help identify these unknown metabolites. To help identify unknowns from tandem mass spectra (MS/MS), Competitive Fragmentation Modeling for Metabolite Identification (CFM-ID) was created. RTpred[16] and RI pred[17] were developed to predict metabolites from liquid chromatography retention times (RT) and gas chromatography retention indices (RIs), and PROSPRE was created to predict metabolites from proton (H1) NMR spectra.[18] In addition to metabolites being generated from catabolism (the breakdown of molecules) and anabolism (synthesis of compounds), many are created from biotransformation (chemically converted for use or excretion from the body) by numerous enzymes, particularly in the liver[19]. Many of the products of biotransformations of drugs and microbial products are unknown. The HMP team has also created BioTransformer[20] to predict metabolite biotransformations to attempt to help researchers identify unknowns.
References
Wishart, D. S.; Tzur, D.; Knox, C.; Eisner, R.; Guo, A. C.; Young, N.; Cheng, D.; Jewell, K.; Arndt, D.; Sawhney, S.; Fung, C.; Nikolai, L.; Lewis, M.; Coutouly, M.-A.; Forsythe, I. (2007-01-03). "HMDB: the Human Metabolome Database". Nucleic Acids Research. 35 (Database): D521–D526. doi:10.1093/nar/gkl923. ISSN 0305-1048. PMC 1899095. PMID 17202168.
Wishart, David (2012), Suhre, Karsten (ed.), "Systems Biology Resources Arising from the Human Metabolome Project", Genetics Meets Metabolomics, New York, NY: Springer New York, pp. 157–175, doi:10.1007/978-1-4614-1689-0_11, ISBN 978-1-4614-1688-3, retrieved 2026-07-08{{citation}}: CS1 maint: work parameter with ISBN (link)
Knox, Craig; Wilson, Mike; Klinger, Christen M.; Franklin, Mark; Oler, Eponine; Wilson, Alex; Pon, Allison; Cox, Jordan; Chin, Na Eun Lucy; Strawbridge, Seth A.; Garcia-Patino, Marysol; Kruger, Ray; Sivakumaran, Aadhavya; Sanford, Selena; Doshi, Rahil (2024-01-05). "DrugBank 6.0: the DrugBank Knowledgebase for 2024". Nucleic Acids Research. 52 (D1): D1265–D1275. doi:10.1093/nar/gkad976. ISSN 1362-4962. PMC 10767804. PMID 37953279.
Wishart, David; Arndt, David; Pon, Allison; Sajed, Tanvir; Guo, An Chi; Djoumbou, Yannick; Knox, Craig; Wilson, Michael; Liang, Yongjie; Grant, Jason; Liu, Yifeng; Goldansaz, Seyed Ali; Rappaport, Stephen M. (2015). "T3DB: the toxic exposome database". Nucleic Acids Research. 43 (Database issue): D928–934. doi:10.1093/nar/gku1004. ISSN 1362-4962. PMID 25378312.
Wishart, David S.; Oler, Eponine; Peters, Harrison; Guo, AnChi; Girod, Sagan; Han, Scott; Saha, Sukanta; Lui, Vicki W.; LeVatte, Marcia; Gautam, Vasuk; Kaddurah-Daouk, Rima; Karu, Naama (2023-01-06). "MiMeDB: the Human Microbial Metabolome Database". Nucleic Acids Research. 51 (D1): D611–D620. doi:10.1093/nar/gkac868. ISSN 1362-4962. PMC 9825614. PMID 36215042.
Wishart, David S; Tian, Siyang; Allen, Dana; Oler, Eponine; Peters, Harrison; Lui, Vicki W; Gautam, Vasuk; Djoumbou-Feunang, Yannick; Greiner, Russell; Metz, Thomas O (2022-07-05). "BioTransformer 3.0—a web server for accurately predicting metabolic transformation products". Nucleic Acids Research. 50 (W1): W115–W123. doi:10.1093/nar/gkac313. ISSN 0305-1048.
LLM-generated pages with certain obvious signs of being machine generated may be deleted without notice.
These tools are prone to specific issues that violate our policies:
Instead, only summarize in your own words a range of independent, reliable, published sources that discuss the subject.
See the advice page on large language models for more information.