Structure Search
Search patent-linked chemistry by using substructure, similarity, identical or connectivity search to find relevant records.
EMBL-EBI | Chemical Biology Resources | SureChEMBL
An open, large-scale database of compounds extracted from patent literature.
Example searches: diabet*, WO-2011043954-A1, US8129446B | Structure Search
Search patent-linked chemistry by using substructure, similarity, identical or connectivity search to find relevant records.
Connect SureChEMBL data directly through simple REST endpoints to retrieve compounds, patents and annotations.
Download bulk datasets for local analysis, reporting or integration. Use regular releases to keep tools aligned with the latest patent-mined chemistry.
Get support with searches, downloads, APIs or interpreting SureChEMBL results. Ask questions, report issues and share feedback.
Explore the SureChEMBL manuals for guidance on data, search and services. Learn how compounds, patents and annotations are organised across the resource. Follow practical examples for the web interface, downloads and REST APIs. Build clearer, more reliable workflows around SureChEMBL data.
SureChEMBL is a publicly available large-scale resource containing compounds extracted from the full text, images and attachments of patent documents. The data are extracted from the patent literature according to an automated text and image-mining pipeline on a daily basis. SureChEMBL provides access to a previously unavailable, open and timely set of annotated compound-patent associations, complemented with sophisticated combined structure and keyword-based search capabilities against the compound repository and patent document corpus. Additionally , SureChEMBL makes use of NLP (Natural Language Processing) to automatically annotate genes and diseases in the patent text. Given the wealth of knowledge hidden in patent documents, analysis of SureChEMBL data has immediate applications in drug discovery, medicinal chemistry and other commercial areas of chemical science.