Skip to content

Latest commit

 

History

History
18 lines (16 loc) · 1.46 KB

20240810_chen_s_et_al.md

File metadata and controls

18 lines (16 loc) · 1.46 KB

Overview

Title: A Large-scale Reaction Dataset of Mechanistic Pathways of Organic Reactions
Authors: Shuan Chen, Ramil Babazade, Taewan Kim, Sunkyu Han, Yousung Jung
Publication Date: 2024/08/10
Publication Link: Nature Scientific Data
Alternative Publication Links: ResearchGate

Abstract

Understanding organic reaction mechanisms is crucial for interpreting the formation of products at the atomic and electronic level, but still remains as a domain of knowledgeable experts. The lack of a large-scale dataset with chemically reasonable mechanistic sequences also hinders the development of reliable machine learning models to predict organic reactions based on mechanisms as human chemists do. Here, we present a high-quality and the first large-scale reaction dataset, denoted as mech-USPTO-31K, with chemically reasonable arrow-pushing diagrams validated by synthetic chemists, encompassing a wide spectrum of polar organic reaction mechanisms. We envision this dataset curated by applying a simple and flexible method that automatically generates reaction mechanisms using autonomously extracted reaction templates and expert-coded mechanistic templates to become an invaluable tool to develop future reaction outcome prediction models and discover new reactions.