Elsevier and LG AI Research announced on September 15, 2026 that chemistry-specific AI vision technology developed by LG AI Research is now being used within Elsevier’s content extraction and curation processes for Reaxys, Elsevier’s discovery chemistry solution, making substances that appear only as images, drawings and reaction schemes in patents and scientific literature searchable.
The companies said substance information from images in patent and journal content is captured more quickly, accurately and at greater scale than was previously possible. The work enhances Elsevier’s content extraction and scientific curation of substances, reactions, bioactivities, biological targets and substance properties for Reaxys.
Why Image-Bound Chemistry Has Been Hard to Search
Much of the substance and reaction information chemists rely on is communicated through figures, drawings and reaction schemes rather than searchable text, according to the announcement. When that chemistry is not searchable, researchers can be left checking documents by hand to confirm whether a compound or reaction has already been described. Chemical drawings encode meaning through bonds, atoms, stereochemistry and spatial relationships, so a model that misreads a bond may identify the wrong compound, while one that misses a structure leaves chemists with an incomplete picture. That challenge is most acute in areas such as novelty searching, competitive intelligence and synthesis planning, and in inorganic and organometallic chemistry, where complex structures are harder to extract and index.
The MolMole Model and Its Reported Benchmarks
The technology combines molecule detection, reaction-diagram parsing and optical chemical structure recognition (OCSR) in a single model, and LG AI Research’s published benchmarking reports that it outperforms alternatives at extracting chemistry from a full document page. A May 7, 2025 post on LG AI Research’s research blog identifies the model as MolMole, developed under the group’s Deep Document Understanding program, which aims to build AI that can interpret text, graphs and tables in general documents as well as molecular structural formulas and reaction formulas in chemical papers and patents.
MolMole takes full PDF documents as input rather than requiring cropped images, and returns recognized chemical data from a document at once, according to the blog post. The model consists of three modules. ViDetect detects molecular structure regions within PDF pages and marks them with bounding boxes. ViReact identifies the positions of reactants, reaction conditions and products within reaction diagrams and classifies each region. ViMore converts recognized structures into standard chemical representations including SMILES, InChI and Mol formats, with specialized techniques for noisy, scan-based patent pages.
LG AI Research reports that ViMore achieved state-of-the-art results on three of the four standard OCSR benchmarks (CLEF, JPO, UOB and USPTO), outperforming DECIMER Image Transformer, MolScribe and MolGrapher, and that it outperformed the other models in particular on JPO, a challenging, primarily low-resolution dataset of images extracted from Japanese patent documents. On LG’s own benchmark, which evaluates 300 patent pages and 250 paper pages separately to reflect their different characteristics, the combined ViDetect and ViMore pipeline outperformed Decimer Segmentation and Image Transformer and MolDetect and MolScribe on precision and recall, and ViReact outperformed ReactionDataExtractor2.0 and RxnScribe.
The underlying paper, posted to arXiv, was first submitted on April 30, 2025 and revised on May 8, 2025. It describes MolMole as a vision-based deep learning framework that unifies molecule detection, reaction diagram parsing and OCSR into a single pipeline for extracting chemical data directly from page-level documents. Citing the lack of a standard page-level benchmark and evaluation metric, the authors also present a 550-page testset annotated with molecule bounding boxes, reaction labels and MOLfiles, along with a new evaluation metric, and report that MolMole outperforms existing toolkits on both their benchmark and public datasets.
Within Elsevier’s workflow, each extraction pipeline is validated against existing Reaxys benchmarks before it goes live, and the full pipeline underwent a testing period across Elsevier’s data and workflow tools before wider use, according to the announcement.
Executive Statements and Next Stages
Mirit Eldor, Managing Director, Life Sciences at Elsevier, said the partnership gives chemists time back by moving more chemistry out of figures and into Reaxys as curated, searchable evidence. “A structure buried in a figure should be evidence rather than a dead end,” she said.
Hwayoung Edward Lee, lead of the AI Biz Transformation Unit at LG AI Research, said the model was designed to decode complex visual chemical representations in which every bond and spatial layout holds meaning, and that the integration with Elsevier converts raw visual data into structured knowledge for researchers.
The organizations said reaction extraction is the next stage of the collaboration, extending image-based extraction beyond individual substances to broaden the reaction evidence available through Reaxys. They are also exploring further customer challenges to tackle together, pairing LG AI Research’s specialist AI capabilities with Elsevier’s chemistry content, scientific expertise and curation. The work follows Elsevier’s Responsible AI Principles and Privacy Principles, the companies said.
