{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"},{"sourceId":173728233,"sourceType":"kernelVersion"},{"sourceId":173754715,"sourceType":"kernelVersion"},{"sourceId":173757700,"sourceType":"kernelVersion"},{"sourceId":173764673,"sourceType":"kernelVersion"},{"sourceId":174671603,"sourceType":"kernelVersion"},{"sourceId":175088375,"sourceType":"kernelVersion"},{"sourceId":175115917,"sourceType":"kernelVersion"}],"dockerImageVersionId":30684,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![](https://cdn-images-1.medium.com/max/1000/1*VZNekSiJoJCcYACvIseF9Q.png)\n","metadata":{}},{"cell_type":"code","source":"from IPython.display import HTML\nimport time\n\nhandle = display(HTML(\"\"\"<marquee>👌</marquee>\"\"\"), display_id='html_marquee1')\ntime.sleep(2)\nhandle = display(HTML(\"\"\"<marquee>Note: Due to memory limitations, data separation and regression calculations were done in five notebooks.</marquee>\"\"\"), display_id='html_marquee1', update=True)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-01T17:51:02.916554Z","iopub.execute_input":"2024-05-01T17:51:02.917583Z","iopub.status.idle":"2024-05-01T17:51:04.959502Z","shell.execute_reply.started":"2024-05-01T17:51:02.917534Z","shell.execute_reply":"2024-05-01T17:51:04.958227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style=\"color:lightgray;background-color:navy;padding:1.2%;border-radius:12px 12px;font-size:1em;text-align:center\">[5]⚕️Competition Submission⚕️AutoGluon</div>\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Description</p></div>\n\n- The examples in the competition dataset are represented by a binary classification of whether a given small molecule is a binder or not to one of three protein targets.\n\n- Three protein targets were screened for this competition :\n\n>\n> **EPHX2 (sEH)**\n>\n> The first target, epoxide hydrolase 2, is encoded by the EPHX2 genetic locus, and its protein product is commonly named “soluble epoxide hydrolase”, or abbreviated to sEH.\n>\n> **BRD4**\n>\n> The second target, bromodomain 4, is encoded by the BRD4 locus and its protein product is also named BRD4. \n>\n> **ALB (HSA)**\n>\n> The third target, serum albumin, is encoded by the ALB locus and its protein product is also named ALB. \n>\n\n- The specifications of the columns of the train file and the test file are as follows :\n\n| Number | Features | Description |\n| ----------- | ----------- | ----------- |\n| 1 | **<span style=\"color: green;\">'id'** | A unique example_id that we use to identify the molecule-binding target pair. |  \n| 2 | **<span style=\"color: green;\">'buildingblock1_smiles'** | The structure, in SMILES, of the first building block | \n| 3 | **<span style=\"color: green;\">'buildingblock2_smiles'** | The structure, in SMILES, of the second building block |  \n| 4 | **<span style=\"color: green;\">'buildingblock3_smiles'** | The structure, in SMILES, of the third building block |  \n| 5 | **<span style=\"color: green;\">'molecule_smiles'** | The structure of the fully assembled molecule, in SMILES. This includes the three building blocks and the triazine core. - Note we use a [Dy] as the stand-in for the DNA linker. |  \n| 6 | **<span style=\"color: green;\">'protein_name'** | The protein target name | \n| 7 | **<span style=\"color: green;\">'binds'** | The target column. A binary class label of whether the molecule binds to the protein. Not available for the test set. |  \n ","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Details about the experiments</p></div>\n\n- DELs are libraries of small molecules with unique DNA barcodes covalently attached\n\n- DELs are manufactured by combining different building blocks\n\n![](https://cdn-images-1.medium.com/max/1000/1*4Fu8NvfopCK-1XN6p7FeUw.png)\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Insight :</p></div>\n\n>\n> The **testing data** has 1674896 rows. A large number of rows slows down the work and practically limits the use of a variety of solutions. Perhaps it would have been better if the evaluation in this competition was done with fewer samples.\n>\n> In the first notebook, we randomly separated 2,000,000 rows of **training data** by protein type. Of course, in the second, third and fourth notebooks, we used only a part of the rows to model and predict each protein. Finally, the fifth notebook is for posting challenge results.\n>\n> So if the first notebook is run several times and each time the regression calculations are done in subsequent notebooks, all these results can be used as **cross-validation**.\n>\n> In addition to the output of first notebook, you can also use the public dataset **\"BELKA_frag_1\"** at the following address :\n>\n> https://www.kaggle.com/datasets/mehrankazeminia/belka-frag-1\n>\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:navy;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>Notebooks :</p></div>\n\n| Number | Notebooks |\n| ----------- | ----------- |\n| 1 | **<span style=\"color: red;\"> ⚕️EDA & Data separation based on protein type** |\n| 2 | **<span style=\"color: navy;\">⚕️Create Models & Prediction for protein_name= sEH** |\n| 3 | **<span style=\"color: red;\"> ⚕️Create Models & Prediction for protein_name= BRD4** |\n| 4 | **<span style=\"color: navy;\">⚕️Create Models & Prediction for protein_name= HSA** |\n| 5 | **<span style=\"color: navy;\">⚕️Competition Submission** |\n","metadata":{}},{"cell_type":"markdown","source":"<p style=\"border-bottom: 10px solid navy\"></p>\n\n<p style=\"border-bottom: 10px solid navy\"></p>","metadata":{}},{"cell_type":"code","source":"import os\nimport gc\nimport glob\nimport random\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\nfrom itertools import groupby\n# ..................................\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-01T17:51:04.961322Z","iopub.execute_input":"2024-05-01T17:51:04.961672Z","iopub.status.idle":"2024-05-01T17:51:09.052834Z","shell.execute_reply.started":"2024-05-01T17:51:04.961643Z","shell.execute_reply":"2024-05-01T17:51:09.051459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls ../input/*","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:09.054190Z","iopub.execute_input":"2024-05-01T17:51:09.054900Z","iopub.status.idle":"2024-05-01T17:51:10.112005Z","shell.execute_reply.started":"2024-05-01T17:51:09.054855Z","shell.execute_reply":"2024-05-01T17:51:10.110635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:darkcyan;\">Competition Submission</span>\n\n<p style=\"border-bottom: 10px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"# protein_name ='sEH'\npred1 = pd.read_csv('/kaggle/input/2-belka-seh-autogluon-frag1/pred1.csv')\n\npred1.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:10.113917Z","iopub.execute_input":"2024-05-01T17:51:10.114966Z","iopub.status.idle":"2024-05-01T17:51:10.513617Z","shell.execute_reply.started":"2024-05-01T17:51:10.114925Z","shell.execute_reply":"2024-05-01T17:51:10.512361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# protein_name ='BRD4'\npred2 = pd.read_csv('/kaggle/input/3-belka-brd4-autogluon-frag1/pred2.csv')\n\npred2.shape","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:10.516475Z","iopub.execute_input":"2024-05-01T17:51:10.516827Z","iopub.status.idle":"2024-05-01T17:51:10.871883Z","shell.execute_reply.started":"2024-05-01T17:51:10.516799Z","shell.execute_reply":"2024-05-01T17:51:10.870789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# protein_name ='HSA'\npred3 = pd.read_csv('/kaggle/input/4-belka-hsa-autogluon-frag1/pred3.csv')\n\npred3.shape                              ","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:10.873521Z","iopub.execute_input":"2024-05-01T17:51:10.874305Z","iopub.status.idle":"2024-05-01T17:51:11.199416Z","shell.execute_reply.started":"2024-05-01T17:51:10.874248Z","shell.execute_reply":"2024-05-01T17:51:11.198268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred = pd.concat([pred1, pred2, pred3])\npred['binds'] = np.clip(pred['binds'], 0., 1.)\npred","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:11.201085Z","iopub.execute_input":"2024-05-01T17:51:11.201498Z","iopub.status.idle":"2024-05-01T17:51:11.303380Z","shell.execute_reply.started":"2024-05-01T17:51:11.201460Z","shell.execute_reply":"2024-05-01T17:51:11.302265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set()\nplt.figure(figsize=(10, 4))\nplt.hist(pred['binds'], bins=50)\n\nplt.gca().set_facecolor('lightyellow')\nplt.suptitle('Prediction Histogram', y=0.96, fontsize=16, c='darkred')\n\nround(pred['binds'].min(), 3), round(pred['binds'].max(), 3)","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:11.304855Z","iopub.execute_input":"2024-05-01T17:51:11.305322Z","iopub.status.idle":"2024-05-01T17:51:11.811648Z","shell.execute_reply.started":"2024-05-01T17:51:11.305284Z","shell.execute_reply":"2024-05-01T17:51:11.810327Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_auto_1 = pred.sort_values(by=['id'])\nsub_auto_1.to_csv('sub_auto_1.csv', index=False)\n\n# Public Score: 0.499\n!ls","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:11.815037Z","iopub.execute_input":"2024-05-01T17:51:11.816031Z","iopub.status.idle":"2024-05-01T17:51:16.769418Z","shell.execute_reply.started":"2024-05-01T17:51:11.815997Z","shell.execute_reply":"2024-05-01T17:51:16.767685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training is done with different rows of training data.\n\nsub_auto_2 = pd.read_csv('../input/5-belka-submission-autogluon-frag2/sub_auto_1.csv')\nsub_auto_3 = pd.read_csv('../input/5-belka-submission-autogluon-frag3/sub_auto_1.csv')\n\n# Similar to cross validation\n\nsub_auto = pd.read_csv('../input/leash-BELKA/sample_submission.csv')\nsub_auto['binds'] = (sub_auto_1['binds'].values + sub_auto_2['binds'].values + sub_auto_3['binds'].values) / 3\nsub_auto.to_csv('sub_auto.csv', index=False)\n# Public Score: 0.524\n!ls","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:16.771317Z","iopub.execute_input":"2024-05-01T17:51:16.771701Z","iopub.status.idle":"2024-05-01T17:51:24.229639Z","shell.execute_reply.started":"2024-05-01T17:51:16.771664Z","shell.execute_reply":"2024-05-01T17:51:24.228217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:darkcyan;\">Ensembling</span>\n\n<p style=\"border-bottom: 10px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"#Thanks to: @ricopue - Public Score: 0.546\nsub_xgb = pd.read_csv('../input/leashbio-xgb-ecfp-10m-sample-rows/submission.csv')\n\nsub_lgbm = pd.read_csv('../input/p-6-6-belka-competition-submission-lgbm/submission.csv')\n\nsubmission = pd.read_csv('../input/leash-BELKA/sample_submission.csv')\nsubmission['binds'] = (sub_auto['binds'].values* 0.35) + (sub_lgbm['binds'].values* 0.15) + (sub_xgb['binds'].values* 0.50)\nsubmission.to_csv('submission.csv', index=False)\n#Public Score: 0.553\n!ls","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:24.231748Z","iopub.execute_input":"2024-05-01T17:51:24.232115Z","iopub.status.idle":"2024-05-01T17:51:31.205498Z","shell.execute_reply.started":"2024-05-01T17:51:24.232080Z","shell.execute_reply":"2024-05-01T17:51:31.204123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set()\nplt.figure(figsize=(10, 4))\nplt.hist(submission['binds'], bins=50)\n\nplt.gca().set_facecolor('lightyellow')\nplt.suptitle('Ensembling Histogram', y=0.96, fontsize=16, c='darkred')\n\nround(submission['binds'].min(), 3), round(submission['binds'].max(), 3)","metadata":{"execution":{"iopub.status.busy":"2024-05-01T17:51:31.207686Z","iopub.execute_input":"2024-05-01T17:51:31.208174Z","iopub.status.idle":"2024-05-01T17:51:31.701001Z","shell.execute_reply.started":"2024-05-01T17:51:31.208129Z","shell.execute_reply":"2024-05-01T17:51:31.699842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<p style=\"border-bottom: 10px solid darkcyan\"></p>\n<p style=\"border-bottom: 10px solid darkcyan\"></p>","metadata":{}}]}