{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Leash Bio\n##  The task is about finding small molecules binding to protein targets. It could be very helpful in searching for new drugs curing human diseases.\n### More detailed description you can find here - https://www.kaggle.com/competitions/leash-BELKA/overview","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-18T10:21:45.773248Z","iopub.execute_input":"2024-05-18T10:21:45.773886Z","iopub.status.idle":"2024-05-18T10:21:47.258865Z","shell.execute_reply.started":"2024-05-18T10:21:45.773839Z","shell.execute_reply":"2024-05-18T10:21:47.257569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Our train dataset is represented as SMILES and there is simple to present these chemicals using RDKit. Let's install this tool to create molecular properties of our molecules.** ","metadata":{}},{"cell_type":"code","source":"!pip install rdkit","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:21:47.261535Z","iopub.execute_input":"2024-05-18T10:21:47.262221Z","iopub.status.idle":"2024-05-18T10:22:07.087569Z","shell.execute_reply.started":"2024-05-18T10:21:47.262175Z","shell.execute_reply":"2024-05-18T10:22:07.086215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We will use duckdb because of the large size of the training dataset. We couldn't use all volume of features for learning our future models, so we create sample with equal parts of binding and not binding chemicals to our protein targets.**","metadata":{}},{"cell_type":"code","source":"!pip install duckdb","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:22:07.089452Z","iopub.execute_input":"2024-05-18T10:22:07.089895Z","iopub.status.idle":"2024-05-18T10:22:23.596262Z","shell.execute_reply.started":"2024-05-18T10:22:07.089859Z","shell.execute_reply":"2024-05-18T10:22:23.594460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Data Downloading**","metadata":{}},{"cell_type":"markdown","source":"*Let's create a sample from a large training dataframe using duckdb.*","metadata":{}},{"cell_type":"code","source":"import duckdb\n\ntrain_path = '/kaggle/input/leash-BELKA/train.parquet'\ntest_path = '/kaggle/input/leash-BELKA/test.parquet'\n\ncon = duckdb.connect()\n\ntrain_sample = con.query(f\"\"\"(SELECT *\n                        FROM parquet_scan('{train_path}')\n                        WHERE binds = 0\n                        ORDER BY random()\n                        LIMIT 30000)\n                        UNION ALL\n                        (SELECT *\n                        FROM parquet_scan('{train_path}')\n                        WHERE binds = 1\n                        ORDER BY random()\n                        LIMIT 30000)\"\"\").df()\n\ncon.close()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:22:23.600007Z","iopub.execute_input":"2024-05-18T10:22:23.600403Z","iopub.status.idle":"2024-05-18T10:23:18.513274Z","shell.execute_reply.started":"2024-05-18T10:22:23.600370Z","shell.execute_reply":"2024-05-18T10:23:18.511612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:23:18.515070Z","iopub.execute_input":"2024-05-18T10:23:18.515550Z","iopub.status.idle":"2024-05-18T10:23:18.549959Z","shell.execute_reply.started":"2024-05-18T10:23:18.515505Z","shell.execute_reply":"2024-05-18T10:23:18.548570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1095143%2F1901c6caa0c6c011617f4dec525d7bbe%2FKaggle%20v2%20(1).png?generation=1712179256934503&alt=media)","metadata":{}},{"cell_type":"markdown","source":"**You can find out from this illustration of DELs, that buildingblock1 is not connected with binding with targets, so we can not use this feature for learning.**","metadata":{}},{"cell_type":"markdown","source":"## **Features engineering**","metadata":{}},{"cell_type":"code","source":"from rdkit import Chem\nfrom rdkit.Chem import AllChem\nfrom rdkit.Chem import rdPartialCharges\nfrom rdkit.Chem import Descriptors\nfrom rdkit.Chem.Descriptors import ExactMolWt\nfrom rdkit.Chem.Descriptors import NumRadicalElectrons","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:23:18.551958Z","iopub.execute_input":"2024-05-18T10:23:18.552538Z","iopub.status.idle":"2024-05-18T10:23:18.807866Z","shell.execute_reply.started":"2024-05-18T10:23:18.552487Z","shell.execute_reply":"2024-05-18T10:23:18.806786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's think what properties could be useful for learning our models from the point of view of chemistry. It could be charges of active chemical groups, also is there aromatics or not in molecule, volume of chemical building blocks, electron density, number of oxigens, number of different bonds. Another idea is to find possible active surface amino acids groups, which could react with different ligands. So, there are many work we should do.**","metadata":{}},{"cell_type":"markdown","source":"## 1. Small molecule properties","metadata":{}},{"cell_type":"markdown","source":"### *First of all it is necessary to convert our SMILEs string to mol object through MolFromSmiles module.*","metadata":{}},{"cell_type":"code","source":"# Convert SMILES to RDKit molecules\ntrain_sample['full_molecule'] = train_sample['molecule_smiles'].apply(Chem.MolFromSmiles)\ntrain_sample['build_block2_mol'] = train_sample['buildingblock2_smiles'].apply(Chem.MolFromSmiles)\ntrain_sample['build_block3_mol'] = train_sample['buildingblock3_smiles'].apply(Chem.MolFromSmiles)","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:23:18.809390Z","iopub.execute_input":"2024-05-18T10:23:18.809876Z","iopub.status.idle":"2024-05-18T10:23:51.828178Z","shell.execute_reply.started":"2024-05-18T10:23:18.809834Z","shell.execute_reply":"2024-05-18T10:23:51.827020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:23:51.829650Z","iopub.execute_input":"2024-05-18T10:23:51.829980Z","iopub.status.idle":"2024-05-18T10:23:51.852967Z","shell.execute_reply.started":"2024-05-18T10:23:51.829954Z","shell.execute_reply":"2024-05-18T10:23:51.851712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" *Morgan fingerprint encodes the atom groups of a chemical into a binary vector with length and radius as its two parameters. It's quitly good property for our molecules.*","metadata":{}},{"cell_type":"code","source":"from rdkit.Chem import rdFingerprintGenerator\n\ndef generate_Morgan_Finger(molecule):\n    if molecule is None:\n        return None\n    mfpgen = rdFingerprintGenerator.GetMorganGenerator(radius=2,fpSize=2048)\n    sfp = mfpgen.GetSparseFingerprint(molecule)\n\n    return (sfp)","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:36:18.264719Z","iopub.execute_input":"2024-05-18T10:36:18.265906Z","iopub.status.idle":"2024-05-18T10:36:18.275888Z","shell.execute_reply.started":"2024-05-18T10:36:18.265859Z","shell.execute_reply":"2024-05-18T10:36:18.274340Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_MolWt(molecule):\n    if molecule is None:\n        return None\n    return (ExactMolWt(molecule))","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:36:19.553662Z","iopub.execute_input":"2024-05-18T10:36:19.554186Z","iopub.status.idle":"2024-05-18T10:36:19.560711Z","shell.execute_reply.started":"2024-05-18T10:36:19.554142Z","shell.execute_reply":"2024-05-18T10:36:19.558588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_NumRadicalEl(molecule):\n    if molecule is None:\n        return None\n    return (NumRadicalElectrons(molecule))","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:36:19.987041Z","iopub.execute_input":"2024-05-18T10:36:19.987572Z","iopub.status.idle":"2024-05-18T10:36:19.995454Z","shell.execute_reply.started":"2024-05-18T10:36:19.987538Z","shell.execute_reply":"2024-05-18T10:36:19.993951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample['full_fingerprint'] = train_sample['full_molecule'].apply(generate_Morgan_Finger)\ntrain_sample['fingerprint_2'] = train_sample['build_block2_mol'].apply(generate_Morgan_Finger)\ntrain_sample['fingerprint_3'] = train_sample['build_block3_mol'].apply(generate_Morgan_Finger)","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:36:20.942965Z","iopub.execute_input":"2024-05-18T10:36:20.943362Z","iopub.status.idle":"2024-05-18T10:36:36.132392Z","shell.execute_reply.started":"2024-05-18T10:36:20.943332Z","shell.execute_reply":"2024-05-18T10:36:36.131000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample['full_MolWt'] = train_sample['full_molecule'].apply(generate_MolWt)\ntrain_sample['MolWt_2'] = train_sample['build_block2_mol'].apply(generate_MolWt)\ntrain_sample['MolWt_3'] = train_sample['build_block3_mol'].apply(generate_MolWt)\n\ntrain_sample['full_NumRadicalEl'] = train_sample['full_molecule'].apply(generate_NumRadicalEl)\ntrain_sample['NumRadicalEl_1'] = train_sample['build_block2_mol'].apply(generate_NumRadicalEl)\ntrain_sample['NumRadicalEl_2'] = train_sample['build_block3_mol'].apply(generate_NumRadicalEl)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:36:36.134538Z","iopub.execute_input":"2024-05-18T10:36:36.134954Z","iopub.status.idle":"2024-05-18T10:36:57.200876Z","shell.execute_reply.started":"2024-05-18T10:36:36.134919Z","shell.execute_reply":"2024-05-18T10:36:57.199606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:36:57.202409Z","iopub.execute_input":"2024-05-18T10:36:57.203044Z","iopub.status.idle":"2024-05-18T10:36:57.350307Z","shell.execute_reply.started":"2024-05-18T10:36:57.203007Z","shell.execute_reply":"2024-05-18T10:36:57.348917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model creation","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import average_precision_score\nfrom sklearn.preprocessing import OneHotEncoder","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:24:31.975914Z","iopub.execute_input":"2024-05-18T10:24:31.976282Z","iopub.status.idle":"2024-05-18T10:24:32.917747Z","shell.execute_reply.started":"2024-05-18T10:24:31.976253Z","shell.execute_reply":"2024-05-18T10:24:32.916328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols = [\"protein_name\"]\ncat_cols_encoded = []\nfor col in cat_cols:\n    cat_cols_encoded += [f\"{col[0]}_{cat}\" for cat in list(train_sample[col].unique())]\n\ncat_cols_encoded\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:24:32.919503Z","iopub.execute_input":"2024-05-18T10:24:32.919926Z","iopub.status.idle":"2024-05-18T10:24:32.939289Z","shell.execute_reply.started":"2024-05-18T10:24:32.919894Z","shell.execute_reply":"2024-05-18T10:24:32.937243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# One-hot encode the protein_name\n\nohe = OneHotEncoder(sparse=False, handle_unknown='ignore')\nencoded_cols = ohe.fit_transform(train_sample[cat_cols])\ndf_prots = pd.DataFrame(encoded_cols, columns=cat_cols_encoded)\ntrain_sample = train_sample.join(df_prots)\ntrain_sample","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:24:32.941301Z","iopub.execute_input":"2024-05-18T10:24:32.941702Z","iopub.status.idle":"2024-05-18T10:24:33.173871Z","shell.execute_reply.started":"2024-05-18T10:24:32.941670Z","shell.execute_reply":"2024-05-18T10:24:33.172686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample.columns","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:24:33.175601Z","iopub.execute_input":"2024-05-18T10:24:33.176068Z","iopub.status.idle":"2024-05-18T10:24:33.184337Z","shell.execute_reply.started":"2024-05-18T10:24:33.176024Z","shell.execute_reply":"2024-05-18T10:24:33.183090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = train_sample[['full_MolWt',\n       'MolWt_2', 'MolWt_3', 'full_NumRadicalEl', 'NumRadicalEl_1',\n       'NumRadicalEl_2', 'p_HSA', 'p_BRD4', 'p_sEH']]\ny = train_sample[['binds']]","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:25:14.071762Z","iopub.execute_input":"2024-05-18T10:25:14.072271Z","iopub.status.idle":"2024-05-18T10:25:14.086667Z","shell.execute_reply.started":"2024-05-18T10:25:14.072233Z","shell.execute_reply":"2024-05-18T10:25:14.084438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:25:15.199691Z","iopub.execute_input":"2024-05-18T10:25:15.202442Z","iopub.status.idle":"2024-05-18T10:25:15.218460Z","shell.execute_reply.started":"2024-05-18T10:25:15.202400Z","shell.execute_reply":"2024-05-18T10:25:15.216945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtc_model = DecisionTreeClassifier()\ndtc_model.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:25:16.586170Z","iopub.execute_input":"2024-05-18T10:25:16.588830Z","iopub.status.idle":"2024-05-18T10:25:16.922578Z","shell.execute_reply.started":"2024-05-18T10:25:16.588770Z","shell.execute_reply":"2024-05-18T10:25:16.921246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_proba = dtc_model.predict_proba(X_test)[:, 1]  # Probability of the positive class\n\n# Calculate the mean average precision\nmap_score = average_precision_score(y_test, y_pred_proba)\nprint(f\"Mean Average Precision (mAP): {map_score:.2f}\")\n","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:25:44.309768Z","iopub.execute_input":"2024-05-18T10:25:44.310627Z","iopub.status.idle":"2024-05-18T10:25:44.338727Z","shell.execute_reply.started":"2024-05-18T10:25:44.310537Z","shell.execute_reply":"2024-05-18T10:25:44.337452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}