{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# References \n\nThis works is inspired by the following kernels \n\n- https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest","metadata":{}},{"cell_type":"markdown","source":"# Install Necessary libraries ","metadata":{}},{"cell_type":"markdown","source":"## DuckDB","metadata":{}},{"cell_type":"markdown","source":"We are installing duckdb here. DuckDB is an open source database management system(DBMS) which help us in fast querying and analysis of the large datasets. When we use the duckdb the data is loaded into the memory for processing , leading to faster query execution. It has SQL support and has some common SQL functionalities. It stores data in a columnar format and can be embedded with the applications. It has a very small memory usage and is easy to deploy. \n\n- **Documentation** - https://duckdb.org/docs/index\n- **Github** - https://github.com/duckdb/duckdb\n\nIn this compeitition, we are having a huge training dataset , so we an use duckdb and consider the paraquet files as the databases. ","metadata":{}},{"cell_type":"code","source":"!pip install duckdb","metadata":{"execution":{"iopub.status.busy":"2024-04-20T13:39:31.198660Z","iopub.execute_input":"2024-04-20T13:39:31.199109Z","iopub.status.idle":"2024-04-20T13:39:46.387652Z","shell.execute_reply.started":"2024-04-20T13:39:31.199073Z","shell.execute_reply":"2024-04-20T13:39:46.386346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## rdkit\n\nRdkit is an open source library for cheminformatics, which is the intersection of chemistry and computer science. It provides a wide range of functionality for working with chemical structures, molecules, and reactions.\n- It helps us to read,write and manipulate molecular structures represented in formats such as SMILES, SDF (Structure Data File), and Mol files.In this compeition , we are using the SMILES structure. \n- RDKit provides methods for generating molecular fingerprints which are binary or integer representations of molecular structures used for similarity searching and machine learning tasks.\n- It helps us in implementation of algorithms such as substructure searching, similarity searching, and property prediction.\n- Support for 2D and 3D molecular visualization.","metadata":{}},{"cell_type":"code","source":"!pip install rdkit","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:01:37.094189Z","iopub.execute_input":"2024-04-20T14:01:37.096022Z","iopub.status.idle":"2024-04-20T14:01:51.925703Z","shell.execute_reply.started":"2024-04-20T14:01:37.095966Z","shell.execute_reply":"2024-04-20T14:01:51.924493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import necessary libararies ","metadata":{}},{"cell_type":"code","source":"import duckdb\nimport pandas as pd\nfrom rdkit import Chem\nfrom rdkit.Chem import AllChem\nfrom sklearn.preprocessing import OneHotEncoder\nfrom xgboost import XGBClassifier\nfrom sklearn.model_selection import cross_val_score\nimport os","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:37:02.862726Z","iopub.execute_input":"2024-04-20T14:37:02.863195Z","iopub.status.idle":"2024-04-20T14:37:02.869409Z","shell.execute_reply.started":"2024-04-20T14:37:02.863162Z","shell.execute_reply":"2024-04-20T14:37:02.868320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Loading","metadata":{}},{"cell_type":"code","source":"train_path = '/kaggle/input/leash-BELKA/train.parquet'\ntest_path = '/kaggle/input/leash-BELKA/test.parquet'\n\ncon = duckdb.connect()\n\ndf = con.query(f\"\"\"(SELECT *\n                        FROM parquet_scan('{train_path}')\n                        WHERE binds = 0\n                        ORDER BY random()\n                        LIMIT 30000)\n                        UNION ALL\n                        (SELECT *\n                        FROM parquet_scan('{train_path}')\n                        WHERE binds = 1\n                        ORDER BY random()\n                        LIMIT 30000)\"\"\").df()\n\ncon.close()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:12:07.507893Z","iopub.execute_input":"2024-04-20T14:12:07.508319Z","iopub.status.idle":"2024-04-20T14:13:00.416685Z","shell.execute_reply.started":"2024-04-20T14:12:07.508283Z","shell.execute_reply":"2024-04-20T14:13:00.414984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:13:09.659004Z","iopub.execute_input":"2024-04-20T14:13:09.659395Z","iopub.status.idle":"2024-04-20T14:13:09.678702Z","shell.execute_reply.started":"2024-04-20T14:13:09.659363Z","shell.execute_reply":"2024-04-20T14:13:09.677683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generating ECPFS \n\nLets grab the smiles for the fully assembled molecule molecule_smiles and generate ecfps for it. We could choose different radiuses or bits, but 2 and 1024 is pretty standard.","metadata":{}},{"cell_type":"code","source":"# Convert SMILES to RDKit molecules\ndf['molecule'] = df['molecule_smiles'].apply(Chem.MolFromSmiles)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:16:24.105181Z","iopub.execute_input":"2024-04-20T14:16:24.105683Z","iopub.status.idle":"2024-04-20T14:16:45.927027Z","shell.execute_reply.started":"2024-04-20T14:16:24.105641Z","shell.execute_reply":"2024-04-20T14:16:45.925811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Generate ECFPs\ndef generate_ecfp(molecule, radius=2, bits=1024):\n    if molecule is None:\n        return None\n    return list(AllChem.GetMorganFingerprintAsBitVect(molecule, radius, nBits=bits))\n\ndf['ecfp'] = df['molecule'].apply(generate_ecfp)","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:17:30.912700Z","iopub.execute_input":"2024-04-20T14:17:30.913087Z","iopub.status.idle":"2024-04-20T14:18:33.867445Z","shell.execute_reply.started":"2024-04-20T14:17:30.913054Z","shell.execute_reply":"2024-04-20T14:18:33.866484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# One-hot encode the protein_name\nonehot_encoder = OneHotEncoder(sparse_output=False)\nprotein_onehot = onehot_encoder.fit_transform(df['protein_name'].values.reshape(-1, 1))","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:20:23.949809Z","iopub.execute_input":"2024-04-20T14:20:23.950250Z","iopub.status.idle":"2024-04-20T14:20:23.980413Z","shell.execute_reply.started":"2024-04-20T14:20:23.950197Z","shell.execute_reply":"2024-04-20T14:20:23.979412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Combine ECFPs and one-hot encoded protein_name\nX = [ecfp + protein for ecfp, protein in zip(df['ecfp'].tolist(), protein_onehot.tolist())]\ny = df['binds'].tolist()","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:20:41.800973Z","iopub.execute_input":"2024-04-20T14:20:41.801410Z","iopub.status.idle":"2024-04-20T14:20:43.042685Z","shell.execute_reply.started":"2024-04-20T14:20:41.801374Z","shell.execute_reply":"2024-04-20T14:20:43.041529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the XGBoost classifier\nxgb_model = XGBClassifier(n_estimators=100, random_state=42)\n\n# Perform cross-validation\ncv_scores = cross_val_score(xgb_model, X, y, cv=5, scoring='average_precision')\n\n# Calculate the mean average precision\nmap_score = cv_scores.mean()\nprint(f\"Mean Average Precision (mAP) with cross-validation: {map_score:.2f}\")","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:24:33.823335Z","iopub.execute_input":"2024-04-20T14:24:33.823740Z","iopub.status.idle":"2024-04-20T14:27:02.171775Z","shell.execute_reply.started":"2024-04-20T14:24:33.823709Z","shell.execute_reply":"2024-04-20T14:27:02.169329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train the XGBoost model on the entire dataset\nxgb_model.fit(X, y)\n# Process the test.parquet file chunk by chunk\ntest_file = '/kaggle/input/leash-BELKA/test.csv'\noutput_file = 'submission.csv'  # Specify the path and filename for the output file\n# Read the test.csv file into a pandas DataFrame\nfor df_test in pd.read_csv(test_file, chunksize=100000):\n    # Generate ECFPs for the molecule_smiles\n    df_test['molecule'] = df_test['molecule_smiles'].apply(Chem.MolFromSmiles)\n    df_test['ecfp'] = df_test['molecule'].apply(generate_ecfp)\n\n    # One-hot encode the protein_name\n    protein_onehot = onehot_encoder.transform(df_test['protein_name'].values.reshape(-1, 1))\n\n    # Combine ECFPs and one-hot encoded protein_name\n    X_test = [ecfp + protein for ecfp, protein in zip(df_test['ecfp'].tolist(), protein_onehot.tolist())]\n\n    # Predict the probabilities using the XGBoost model trained with cross-validation\n    probabilities = xgb_model.predict_proba(X_test)[:, 1]\n\n    # Create a DataFrame with 'id' and 'probability' columns\n    output_df = pd.DataFrame({'id': df_test['id'], 'binds': probabilities})\n\n    # Save the output DataFrame to a CSV file\n    output_df.to_csv(output_file, index=False, mode='a', header=not os.path.exists(output_file))","metadata":{"execution":{"iopub.status.busy":"2024-04-20T14:41:59.168725Z","iopub.execute_input":"2024-04-20T14:41:59.169481Z","iopub.status.idle":"2024-04-20T15:26:25.272190Z","shell.execute_reply.started":"2024-04-20T14:41:59.169407Z","shell.execute_reply":"2024-04-20T15:26:25.270852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}