Papyrus - A large scale curated dataset aimed at bioactivity predictions
DOI: 10.4121/16896406
Datacite citation style
Dataset
This repository contains the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the manuscript "Papyrus - A large scale curated dataset aimed at bioactivity predictions" (Work in Progress).
With the recent rapid growth of publicly available ligand-protein bioactivity data, there is a trove of viable data that can be used to train machine learning algorithms. However, not all data is equal in terms of size and quality, and a significant portion of researcher’s time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. As an answer to that, we have constructed the Papyrus dataset, comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with smaller datasets containing high quality data. This aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways, and also perform some rudimentary quantitative structure-activity relationship and proteochemometrics modeling. Our ambition is to create a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.
History
- 2021-10-29 first online
- 2022-04-04 published, posted
Publisher
4TU.ResearchDataFormat
g-zipped tab-separated files and g-zipped SD filesFunding
- Enhacing TRANslational SAFEty Assessment through Integrative Knowledge Management (grant code 777365) [more info...] European Commission
Organizations
Leiden Academic Centre for Drug Research (LACDR), The NetherlandsLeiden University, The Netherlands
DATA
Files (4)
- 8,743 bytesMD5:
5f92ef477ddfee6bfd5cca5d54292ab1README.txt - 2,626,964 bytesMD5:
123b42999e644e3aa5acf0a7fb29c13f05.4_combined_set_protein_targets.tsv.gz - 1,557,599,937 bytesMD5:
9f6fa88d0786b928d73d25c7c278fe2105.4_combined_set_with_stereochemistry.tsv.gz - 1,595,754,368 bytesMD5:
bf34e9f93867217b7d9e1168d5c338b205.4_combined_set_without_stereochemistry.tsv.gz -
download all files (zip)
3,155,990,012 bytes unzipped




