{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Project Summary: Predict New Medicines with BELKA\n\n# 🚀 Introduction:\nThe BELKA competition aims to harness machine learning to predict the binding affinity of small molecules to specific protein targets. This crucial step in drug development can significantly accelerate the discovery of new drugs by identifying potential candidates without extensive laboratory experiments. The dataset, provided by Leash Biosciences, includes binding interaction data between 133 million small molecules and three protein targets.\n\n# 📚 Data Preparation:\n\n**Loading Data:** The training and test data were loaded from parquet files.\n**Sampling:** Balanced the dataset by randomly sampling 30,000 binding and 30,000 non-binding interactions.\n**SMILES Conversion:** Converted SMILES strings to RDKit molecule objects for further processing.\n**Feature Generation:** Generated Extended Connectivity Fingerprints (ECFP) for each molecule, which served as molecular descriptors.\n**Protein Encoding:** One-hot encoded the protein target names to include them as features in the model.\n**Combining Features:** Combined ECFP descriptors and one-hot encoded protein names into a feature set.\n\n# 🛠️ Model Development:\n\n**Data Splitting:** Split the combined feature set and labels into training and validation sets.\n**Model Selection:** Chose the XGBoost classifier due to its effectiveness in handling structured data and binary classification tasks.\n**Training:** Trained the XGBoost model on the training set with specific parameters to optimize performance.\n\n# 🔍 Model Evaluation:\n\n**Prediction:** Predicted binding probabilities on the validation set.\n**Metrics:** Evaluated model performance using Mean Average Precision (mAP) and accuracy metrics.\n\n# Visualization:\nPlotted the distribution of binding vs. non-binding interactions.\nDisplayed the confusion matrix to visualize true vs. false positives and negatives.\n\n# 📝 Submission Preparation:\n**Test Data Processing:** Processed the test dataset similarly to the training data (SMILES conversion, ECFP generation, and protein encoding).\n**Prediction:** Predicted binding probabilities for the test set.\n**Submission File:** Created a submission file with the required format and saved it for final submission.\n\n# 🎉 Conclusion:\nThis project demonstrates a robust approach to predicting small molecule-protein interactions using machine learning. By leveraging the BELKA dataset and advanced ML techniques, we developed a model that can potentially revolutionize drug discovery, making the process faster and more efficient. Our model's performance, evaluated through mAP and accuracy metrics, highlights the feasibility of using computational methods to explore vast chemical spaces and identify promising drug candidates.","metadata":{}},{"cell_type":"code","source":"!pip install xgboost","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:35:04.292025Z","iopub.execute_input":"2024-05-17T08:35:04.293710Z","iopub.status.idle":"2024-05-17T08:35:20.295342Z","shell.execute_reply.started":"2024-05-17T08:35:04.293649Z","shell.execute_reply":"2024-05-17T08:35:20.293990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"conda install -c conda-forge rdkit\n","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:35:20.299239Z","iopub.execute_input":"2024-05-17T08:35:20.300484Z","iopub.status.idle":"2024-05-17T08:36:19.328049Z","shell.execute_reply.started":"2024-05-17T08:35:20.300438Z","shell.execute_reply":"2024-05-17T08:36:19.325009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install rdkit-pypi\n","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:36:19.331028Z","iopub.execute_input":"2024-05-17T08:36:19.331585Z","iopub.status.idle":"2024-05-17T08:36:39.955170Z","shell.execute_reply.started":"2024-05-17T08:36:19.331534Z","shell.execute_reply":"2024-05-17T08:36:39.952611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from rdkit import Chem\nprint(\"RDKit version:\", Chem.rdBase.rdkitVersion)\n","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:36:39.963537Z","iopub.execute_input":"2024-05-17T08:36:39.964718Z","iopub.status.idle":"2024-05-17T08:36:39.972527Z","shell.execute_reply.started":"2024-05-17T08:36:39.964658Z","shell.execute_reply":"2024-05-17T08:36:39.970642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install duckdb\n","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:36:39.975135Z","iopub.execute_input":"2024-05-17T08:36:39.975796Z","iopub.status.idle":"2024-05-17T08:36:56.082183Z","shell.execute_reply.started":"2024-05-17T08:36:39.975747Z","shell.execute_reply":"2024-05-17T08:36:56.080609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" conda install -c conda-forge duckdb\n","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:36:56.084506Z","iopub.execute_input":"2024-05-17T08:36:56.084930Z","iopub.status.idle":"2024-05-17T08:37:52.427804Z","shell.execute_reply.started":"2024-05-17T08:36:56.084893Z","shell.execute_reply":"2024-05-17T08:37:52.426219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import required libraries","metadata":{}},{"cell_type":"code","source":"from rdkit import Chem\nfrom rdkit.Chem import AllChem\nimport pandas as pd\nimport duckdb\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import average_precision_score, accuracy_score, confusion_matrix, ConfusionMatrixDisplay\nfrom xgboost import XGBClassifier\nfrom sklearn.preprocessing import OneHotEncoder\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport os","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:37:52.429835Z","iopub.execute_input":"2024-05-17T08:37:52.430370Z","iopub.status.idle":"2024-05-17T08:37:52.441152Z","shell.execute_reply.started":"2024-05-17T08:37:52.430324Z","shell.execute_reply":"2024-05-17T08:37:52.439777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Path to the train data","metadata":{}},{"cell_type":"code","source":"\ntrain_path = '/kaggle/input/leash-BELKA/train.parquet'","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:37:52.443162Z","iopub.execute_input":"2024-05-17T08:37:52.444035Z","iopub.status.idle":"2024-05-17T08:37:52.458937Z","shell.execute_reply.started":"2024-05-17T08:37:52.443986Z","shell.execute_reply":"2024-05-17T08:37:52.457245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Connect to DuckDB and load the data","metadata":{}},{"cell_type":"code","source":"\ncon = duckdb.connect()\nquery = f\"\"\"(SELECT *\n             FROM parquet_scan('{train_path}')\n             WHERE binds = 0\n             ORDER BY random()\n             LIMIT 30000)\n             UNION ALL\n             (SELECT *\n             FROM parquet_scan('{train_path}')\n             WHERE binds = 1\n             ORDER BY random()\n             LIMIT 30000)\"\"\"\ndf = con.execute(query).df()\ncon.close()","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:37:52.460481Z","iopub.execute_input":"2024-05-17T08:37:52.460857Z","iopub.status.idle":"2024-05-17T08:38:56.715479Z","shell.execute_reply.started":"2024-05-17T08:37:52.460827Z","shell.execute_reply":"2024-05-17T08:38:56.713991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function to convert SMILES to RDKit molecule objects","metadata":{}},{"cell_type":"code","source":"\ndf['molecule'] = df['molecule_smiles'].apply(Chem.MolFromSmiles)","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:38:56.719817Z","iopub.execute_input":"2024-05-17T08:38:56.720273Z","iopub.status.idle":"2024-05-17T08:39:16.348468Z","shell.execute_reply.started":"2024-05-17T08:38:56.720236Z","shell.execute_reply":"2024-05-17T08:39:16.347173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function to generate ECFPs (Extended Connectivity Fingerprints)","metadata":{}},{"cell_type":"code","source":"\ndef generate_ecfp(molecule, radius=2, bits=1024):\n    if molecule is None:\n        return None\n    return list(AllChem.GetMorganFingerprintAsBitVect(molecule, radius, nBits=bits))\n\ndf['ecfp'] = df['molecule'].apply(generate_ecfp)","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:39:16.350179Z","iopub.execute_input":"2024-05-17T08:39:16.350637Z","iopub.status.idle":"2024-05-17T08:40:18.776500Z","shell.execute_reply.started":"2024-05-17T08:39:16.350593Z","shell.execute_reply":"2024-05-17T08:40:18.775154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# One-hot encode the protein_name column","metadata":{}},{"cell_type":"code","source":"\nonehot_encoder = OneHotEncoder(sparse_output=False)\nprotein_onehot = onehot_encoder.fit_transform(df['protein_name'].values.reshape(-1, 1))","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:18.778373Z","iopub.execute_input":"2024-05-17T08:40:18.778728Z","iopub.status.idle":"2024-05-17T08:40:18.815816Z","shell.execute_reply.started":"2024-05-17T08:40:18.778699Z","shell.execute_reply":"2024-05-17T08:40:18.814448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Combine ECFPs and one-hot encoded protein names into the feature set","metadata":{}},{"cell_type":"code","source":"\nX = np.hstack((np.array(df['ecfp'].tolist()), protein_onehot))\ny = df['binds'].values","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:18.817363Z","iopub.execute_input":"2024-05-17T08:40:18.817762Z","iopub.status.idle":"2024-05-17T08:40:25.569278Z","shell.execute_reply.started":"2024-05-17T08:40:18.817729Z","shell.execute_reply":"2024-05-17T08:40:25.568324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split data into train and test sets","metadata":{}},{"cell_type":"code","source":"\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:25.570751Z","iopub.execute_input":"2024-05-17T08:40:25.571193Z","iopub.status.idle":"2024-05-17T08:40:25.793872Z","shell.execute_reply.started":"2024-05-17T08:40:25.571151Z","shell.execute_reply":"2024-05-17T08:40:25.792643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Initialize and train the XGBoost model","metadata":{}},{"cell_type":"code","source":"\nxgb_model = XGBClassifier(n_estimators=100, use_label_encoder=False, eval_metric='logloss', random_state=42)\nxgb_model.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:25.795588Z","iopub.execute_input":"2024-05-17T08:40:25.795946Z","iopub.status.idle":"2024-05-17T08:40:36.376017Z","shell.execute_reply.started":"2024-05-17T08:40:25.795916Z","shell.execute_reply":"2024-05-17T08:40:36.374625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predict probabilities on the test set","metadata":{}},{"cell_type":"code","source":"\ny_pred_proba = xgb_model.predict_proba(X_test)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:36.377572Z","iopub.execute_input":"2024-05-17T08:40:36.377905Z","iopub.status.idle":"2024-05-17T08:40:36.436796Z","shell.execute_reply.started":"2024-05-17T08:40:36.377876Z","shell.execute_reply":"2024-05-17T08:40:36.435667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculate the mean average precision and accuracy","metadata":{}},{"cell_type":"code","source":"\nmap_score = average_precision_score(y_test, y_pred_proba)\naccuracy = accuracy_score(y_test, xgb_model.predict(X_test))\nprint(f\"Mean Average Precision (mAP): {map_score:.2f}\")\nprint(f\"Accuracy: {accuracy:.2f}\")","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:36.438545Z","iopub.execute_input":"2024-05-17T08:40:36.439336Z","iopub.status.idle":"2024-05-17T08:40:36.504992Z","shell.execute_reply.started":"2024-05-17T08:40:36.439298Z","shell.execute_reply":"2024-05-17T08:40:36.503261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting the distribution of binding vs non-binding","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize=(8, 6))\nplt.pie(df['binds'].value_counts(), labels=['Non-binding', 'Binding'], autopct='%1.1f%%', startangle=90)\nplt.title('Distribution of Binding vs Non-binding')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:36.506948Z","iopub.execute_input":"2024-05-17T08:40:36.508238Z","iopub.status.idle":"2024-05-17T08:40:36.754824Z","shell.execute_reply.started":"2024-05-17T08:40:36.508193Z","shell.execute_reply":"2024-05-17T08:40:36.753554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting the confusion matrix","metadata":{}},{"cell_type":"code","source":"\ncm = confusion_matrix(y_test, xgb_model.predict(X_test))\ndisp = ConfusionMatrixDisplay(confusion_matrix=cm)\ndisp.plot()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-05-17T08:40:36.756884Z","iopub.execute_input":"2024-05-17T08:40:36.757775Z","iopub.status.idle":"2024-05-17T08:40:37.114932Z","shell.execute_reply.started":"2024-05-17T08:40:36.757719Z","shell.execute_reply":"2024-05-17T08:40:37.113676Z"},"trusted":true},"execution_count":null,"outputs":[]}]}