{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"},{"sourceId":8823311,"sourceType":"datasetVersion","datasetId":5308240}],"dockerImageVersionId":30732,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BindingDB \nis an extensive and well-regarded database dedicated to the quantitative measurement of molecular interactions, specifically focusing on protein-ligand and protein-protein binding affinities. Established to facilitate drug discovery and development, BindingDB serves as a crucial resource for researchers by providing a wealth of data that enables the understanding of how small molecules interact with their protein targets. This repository encompasses detailed information on binding constants, thermodynamic parameters, and kinetic data, collated from peer-reviewed publications and patent literature.\n\nThe database stands out for its comprehensive coverage of a wide range of biological targets, including enzymes, receptors, and transporters, which are pivotal in various therapeutic areas. By offering standardized and curated data, BindingDB supports computational and experimental efforts in pharmacology, medicinal chemistry, and molecular biology. It also integrates seamlessly with other bioinformatics tools and databases, enhancing its utility in predictive modeling and the design of new compounds with desirable biological activity.\n\nResearchers and scientists worldwide leverage BindingDB for its reliable and accessible data, contributing to advancements in drug design, identification of potential drug candidates, and the elucidation of molecular mechanisms underlying disease processes. Through continuous updates and expansions, BindingDB remains a cornerstone in the landscape of cheminformatics and bioinformatics, fostering innovation and collaboration across disciplines.","metadata":{}},{"cell_type":"markdown","source":"# OtterKnowledge \nis a platform that allows to train models from multimodal knowledge graphs and transfer the knowledge from the models to downstream tasks. Please go to https://github.com/IBM/otter-knowledge to learn more details. If you find these material useful please support us by giving us a star on our github, or cite the following paper: \n\n```\n@inproceedings{hoang2024knowledge,\n  title={Knowledge Enhanced Representation Learning for Drug Discovery},\n  author={Hoang, Thanh Lam and Sbodio, Marco Luca and Galindo, Marcos Martinez and Zayats, Mykhaylo and Fernandez-Diaz, Raul and Valls, Victor and Picco, Gabriele and Berrospi, Cesar and Lopez, Vanessa},\n  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},\n  volume={38},\n  number={9},\n  pages={10544--10552},\n  year={2024}\n}\n```","metadata":{}},{"cell_type":"markdown","source":"# Features transfered from BindingDB\nWe employed OtterKnowledge (OK) to train on BindingDB graphs, which included several million drug-target binding affinity experimental pairs. This model was then used to predict binding scores for each drug-protein pair in the Belka dataset. Additionally, we evaluated the block SMILES against the proteins. Below is an example of the output we obtained. It is important to note that OK uses MorganFingerprint and ESM as the representation of smiles and proteins. \n\nWe store the BindingDB predictions at this [location](https://www.kaggle.com/datasets/lamthuy/otterknowledge) . Due to the large size of the dataset, we have only performed predictions for 10% of the training data and the entire test data.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf_hsa = pd.read_csv(\"/kaggle/input/otterknowledge/train.csv.gz.0.1.gz.HSA.gz.train.gz_part1.csv.gz.otter_knowledge.csv/train.csv.gz.0.1.gz.HSA.gz.train.gz_part1.csv.gz.otter_knowledge.csv\")\ndf_brd4 = pd.read_csv(\"/kaggle/input/otterknowledge/train.csv.gz.0.1.gz.BRD4.gz.train.gz_part1.csv.gz.otter_knowledge.csv/train.csv.gz.0.1.gz.BRD4.gz.train.gz_part1.csv.gz.otter_knowledge.csv\")\ndf_seh = pd.read_csv(\"/kaggle/input/otterknowledge/train.csv.gz.0.1.gz.sEH.gz.train.gz_part1.csv.gz.otter_knowledge.csv/train.csv.gz.0.1.gz.sEH.gz.train.gz_part1.csv.gz.otter_knowledge.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-06-30T09:51:46.707817Z","iopub.execute_input":"2024-06-30T09:51:46.708266Z","iopub.status.idle":"2024-06-30T09:52:03.909138Z","shell.execute_reply.started":"2024-06-30T09:51:46.708233Z","shell.execute_reply":"2024-06-30T09:52:03.907837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can explore the correlation between the \"binds\" column and the IC50, Ki, Kd, and EC50 predictions below. Although the correlation is not particularly strong, there are observable trends. Given that BindingDB includes only around 1 million drugs and 9,000 proteins, it remains uncertain whether these additional features will be useful for the three proteins in this competition. If you find these features useful, please share your findings with us.","metadata":{}},{"cell_type":"code","source":"df_hsa.head()","metadata":{"execution":{"iopub.status.busy":"2024-06-30T09:52:03.911359Z","iopub.execute_input":"2024-06-30T09:52:03.911801Z","iopub.status.idle":"2024-06-30T09:52:03.937065Z","shell.execute_reply.started":"2024-06-30T09:52:03.911771Z","shell.execute_reply":"2024-06-30T09:52:03.935934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_matrix = df_hsa.corr()\ncorrelation_matrix['binds']","metadata":{"execution":{"iopub.status.busy":"2024-06-30T09:52:03.938789Z","iopub.execute_input":"2024-06-30T09:52:03.939140Z","iopub.status.idle":"2024-06-30T09:52:05.101504Z","shell.execute_reply.started":"2024-06-30T09:52:03.939110Z","shell.execute_reply":"2024-06-30T09:52:05.100231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_matrix = df_brd4.corr()\ncorrelation_matrix['binds']","metadata":{"execution":{"iopub.status.busy":"2024-06-30T09:52:05.102930Z","iopub.execute_input":"2024-06-30T09:52:05.103271Z","iopub.status.idle":"2024-06-30T09:52:06.293069Z","shell.execute_reply.started":"2024-06-30T09:52:05.103241Z","shell.execute_reply":"2024-06-30T09:52:06.291722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"correlation_matrix = df_seh.corr()\ncorrelation_matrix['binds']","metadata":{"execution":{"iopub.status.busy":"2024-06-30T09:52:06.296138Z","iopub.execute_input":"2024-06-30T09:52:06.296624Z","iopub.status.idle":"2024-06-30T09:52:07.515857Z","shell.execute_reply.started":"2024-06-30T09:52:06.296559Z","shell.execute_reply":"2024-06-30T09:52:07.514607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_brd4.IC50_smiles[df_brd4.binds==1].mean(), df_brd4.IC50_smiles[df_brd4.binds==0].mean())\nprint(df_hsa.IC50_smiles[df_hsa.binds==1].mean(), df_hsa.IC50_smiles[df_hsa.binds==0].mean())\nprint(df_seh.IC50_smiles[df_seh.binds==1].mean(), df_seh.IC50_smiles[df_seh.binds==0].mean())","metadata":{"execution":{"iopub.status.busy":"2024-06-30T09:52:07.517143Z","iopub.execute_input":"2024-06-30T09:52:07.517481Z","iopub.status.idle":"2024-06-30T09:52:07.588396Z","shell.execute_reply.started":"2024-06-30T09:52:07.517452Z","shell.execute_reply":"2024-06-30T09:52:07.587114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that the bindingdb data has a little impact on the protein brd4 more than the hsa and seh proteins. ","metadata":{}}]}