{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on April 05, 2024. By Marília Prata","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-06T21:22:58.394644Z","iopub.execute_input":"2024-04-06T21:22:58.394967Z","iopub.status.idle":"2024-04-06T21:22:59.246444Z","shell.execute_reply.started":"2024-04-06T21:22:58.394941Z","shell.execute_reply":"2024-04-06T21:22:59.245577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![](https://img1.daumcdn.net/thumb/R750x0/?scode=mtistory2&fname=https%3A%2F%2Fblog.kakaocdn.net%2Fdn%2FcGJd7r%2Fbtseqtm3y95%2FPvDLqGSXs7lELoEhDZHEgk%2Fimg.png)https://datainsider.tistory.com/142","metadata":{}},{"cell_type":"markdown","source":"\"DeepChem is a Python library for machine learning and deep learning on molecular and quantum datasets. It is built on top of PyTorch, and other popular ML frameworks. It is designed to make it easy to apply ML to new domains, and to build and benchmark new models. It is also designed to make it easy to use ML in production, by providing easy-to-use model export and deployment APIs.\"\n\nhttps://deepchem.io/about/","metadata":{}},{"cell_type":"markdown","source":"#DeepChem Citation:\n\n@manual{Intro1, \n\n title={The Basic Tools of the Deep Life Sciences}, \n organization={DeepChem},\n \n author={Ramsundar, Bharath}, \n \n howpublished = {\\url{https://github.com/deepchem/deepchem/blob/master/examples/tutorials/The_Basic_Tools_of_the_Deep_Life_Sciences.ipynb}}, \n year={2021}, \n} ","metadata":{}},{"cell_type":"code","source":"!pip install deepchem[tensorflow]","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-06T21:23:37.009116Z","iopub.execute_input":"2024-04-06T21:23:37.009405Z","iopub.status.idle":"2024-04-06T21:23:53.783550Z","shell.execute_reply.started":"2024-04-06T21:23:37.009381Z","shell.execute_reply":"2024-04-06T21:23:53.782809Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\nimport deepchem as dc","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-06T21:23:56.135853Z","iopub.execute_input":"2024-04-06T21:23:56.136182Z","iopub.status.idle":"2024-04-06T21:24:12.622279Z","shell.execute_reply.started":"2024-04-06T21:23:56.136156Z","shell.execute_reply":"2024-04-06T21:24:12.621575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\n#loading the data\ntasks, datasets, transformers = dc.molnet.load_delaney(featurizer = 'GraphConv')\ntrain_dataset, valid_dataset, test_dataset = datasets","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:24:36.726175Z","iopub.execute_input":"2024-04-06T21:24:36.726463Z","iopub.status.idle":"2024-04-06T21:24:40.462580Z","shell.execute_reply.started":"2024-04-06T21:24:36.726436Z","shell.execute_reply":"2024-04-06T21:24:40.461823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\n#create and train a model\nmodel = dc.models.GraphConvModel(n_tasks = 1, mode = 'regression', dropout = 0.2)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:24:43.893705Z","iopub.execute_input":"2024-04-06T21:24:43.894020Z","iopub.status.idle":"2024-04-06T21:24:43.942346Z","shell.execute_reply.started":"2024-04-06T21:24:43.893995Z","shell.execute_reply":"2024-04-06T21:24:43.941570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.fit(train_dataset, nb_epoch = 100)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:24:49.068944Z","iopub.execute_input":"2024-04-06T21:24:49.069802Z","iopub.status.idle":"2024-04-06T21:25:15.415788Z","shell.execute_reply.started":"2024-04-06T21:24:49.069775Z","shell.execute_reply":"2024-04-06T21:25:15.414996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\n#Model Evaluation\nmetric = dc.metrics.Metric(dc.metrics.pearson_r2_score)\nprint(\"Training set score\")\nprint(model.evaluate(train_dataset, [metric], transformers))\nprint(\"Test set score\")\nprint(model.evaluate(test_dataset, [metric], transformers))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:25:21.087280Z","iopub.execute_input":"2024-04-06T21:25:21.087613Z","iopub.status.idle":"2024-04-06T21:25:22.492084Z","shell.execute_reply.started":"2024-04-06T21:25:21.087587Z","shell.execute_reply":"2024-04-06T21:25:22.491254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#DuckDB","metadata":{}},{"cell_type":"code","source":"!pip install duckdb","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:25:27.969741Z","iopub.execute_input":"2024-04-06T21:25:27.970346Z","iopub.status.idle":"2024-04-06T21:25:38.828671Z","shell.execute_reply.started":"2024-04-06T21:25:27.970317Z","shell.execute_reply":"2024-04-06T21:25:38.827595Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By SeshuRajup  https://www.kaggle.com/code/seshurajup/eda-smiles/notebook\n\nimport duckdb\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\n\ntrain_path = '/kaggle/input/leash-BELKA/train.parquet'\ntest_path = '/kaggle/input/leash-BELKA/test.parquet'\n\ncon = duckdb.connect()","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:25:59.772931Z","iopub.execute_input":"2024-04-06T21:25:59.773277Z","iopub.status.idle":"2024-04-06T21:25:59.811921Z","shell.execute_reply.started":"2024-04-06T21:25:59.773250Z","shell.execute_reply":"2024-04-06T21:25:59.811212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By SeshuRajup  https://www.kaggle.com/code/seshurajup/eda-smiles/notebook\n\nsample = con.query(f\"\"\"(SELECT * FROM parquet_scan('{train_path}') LIMIT 10)\"\"\").df()\n\nstr(list(sample.columns))","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:26:05.979177Z","iopub.execute_input":"2024-04-06T21:26:05.979716Z","iopub.status.idle":"2024-04-06T21:26:06.075990Z","shell.execute_reply.started":"2024-04-06T21:26:05.979685Z","shell.execute_reply":"2024-04-06T21:26:06.075176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample.tail(3)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:26:11.760499Z","iopub.execute_input":"2024-04-06T21:26:11.761033Z","iopub.status.idle":"2024-04-06T21:26:11.775825Z","shell.execute_reply.started":"2024-04-06T21:26:11.761005Z","shell.execute_reply":"2024-04-06T21:26:11.774962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By SeshuRajup  https://www.kaggle.com/code/seshurajup/eda-smiles/notebook\n\nbuildingblock2_smiles_stats = con.query(f\"\"\"(SELECT buildingblock2_smiles, count(*) as buildingblock2_smiles_count FROM parquet_scan('{train_path}') GROUP BY buildingblock2_smiles ORDER BY count(*) DESC)\"\"\").df()\nbuildingblock2_smiles_stats","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:26:17.183741Z","iopub.execute_input":"2024-04-06T21:26:17.184054Z","iopub.status.idle":"2024-04-06T21:26:20.796497Z","shell.execute_reply.started":"2024-04-06T21:26:17.184031Z","shell.execute_reply":"2024-04-06T21:26:20.795599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#From the column buildingblock2_smiles:","metadata":{}},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\nsmiles = ['Cl.Nc1ccc2cccnc2c1',\n          'Cl.NCc1ccc[nH]1',\n          'CN(Cc1ccco1)Cc1ccccc1CN',\n          'Nc1cccc2cnccc12',\n          'NCc1c(F)cccc1N1CCCC1']","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:26:27.308496Z","iopub.execute_input":"2024-04-06T21:26:27.308837Z","iopub.status.idle":"2024-04-06T21:26:27.313388Z","shell.execute_reply.started":"2024-04-06T21:26:27.308816Z","shell.execute_reply":"2024-04-06T21:26:27.312340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\nfrom rdkit import Chem\nmols = [Chem.MolFromSmiles(s) for s in smiles]\nfeaturizer = dc.feat.ConvMolFeaturizer()\nx = featurizer.featurize(mols)\npredicted_solubility = model.predict_on_batch(x)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:26:33.764247Z","iopub.execute_input":"2024-04-06T21:26:33.764638Z","iopub.status.idle":"2024-04-06T21:26:33.782429Z","shell.execute_reply.started":"2024-04-06T21:26:33.764608Z","shell.execute_reply":"2024-04-06T21:26:33.781729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Salman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem\n\nfor m,s in zip(smiles, predicted_solubility):\n    print()\n    print('Molecule:', m)\n    print('Predicted solubility:', s)","metadata":{"execution":{"iopub.status.busy":"2024-04-06T21:26:38.632756Z","iopub.execute_input":"2024-04-06T21:26:38.633060Z","iopub.status.idle":"2024-04-06T21:26:38.638478Z","shell.execute_reply.started":"2024-04-06T21:26:38.633036Z","shell.execute_reply":"2024-04-06T21:26:38.637565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#This script will help in the future\n\nThis Kaggle Notebook won't help for Predictions in the present BELKA Competition . Though it may help IN the FUTURE to choose the adequate drug according to the molecular solubility. \n\nComputational prediction of drug solubility in water-based systems: Qualitative and quantitative approaches used in the current drug discovery and development setting\n\n\"This review serves to provide an update on these new approaches and how they can be used to more accurately predict solubility, and also importantly, inform us on molecular interactions and processes occurring during drug dissolution and solubilisation.\"\n\nCITATION:\n\nBergström CAS, Larsson P. Computational prediction of drug solubility in water-based systems: Qualitative and quantitative approaches used in the current drug discovery and development setting. Int J Pharm. 2018 Apr 5;540(1-2):185-193. doi: 10.1016/j.ijpharm.2018.01.044. Epub 2018 Feb 6. PMID: 29421301; PMCID: PMC5861307.\n\nhttps://www.ncbi.nlm.nih.gov/pmc/articles/PMC5861307/","metadata":{}},{"cell_type":"markdown","source":"#Drug Solubility for Medicinal Chemistry\n\n![](https://i.ytimg.com/vi/HCvOSpia_xg/sddefault.jpg)https://www.youtube.com/watch?app=desktop&v=HCvOSpia_xg","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nDeepChem https://deepchem.io/tutorials/the-basic-tools-of-the-deep-life-sciences/\n\nSeshuRajup  https://www.kaggle.com/code/seshurajup/eda-smiles/notebook\n\nSalman Ibne Eunus https://www.kaggle.com/code/salmaneunus/prediction-solubility-of-molecules-with-deepchem","metadata":{}}]}