{
  "id": 519133,
  "title": "2nd Public/13th Private Solution",
  "url": "/competitions/leash-BELKA/writeups/loosers-2nd-public-13th-private-solution",
  "author_name": "",
  "post_date": "2024-07-09T20:42:32.577Z",
  "votes": 68,
  "comment_count": 7,
  "views": 0,
  "content": "<p>We would like to thank Kaggle for organizing such an interesting competition. We also appreciate <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> , <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , and <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> for sharing notebooks that influenced our solution. A big thanks to my teammates <a href=\"https://www.kaggle.com/yyyu54\" target=\"_blank\">@yyyu54</a> , <a href=\"https://www.kaggle.com/Ogurtsov\" target=\"_blank\">@Ogurtsov</a>, and <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> for fighting side by side until we used up all 480 submissions. We all reached the Competitions Master at the same time!</p>\n<h1>Approach Overview</h1>\n<p>We used separate approaches for molecules with shared building blocks and non-shared building blocks.</p>\n<h2>1. Shared Building Blocks</h2>\n<p>We used an ensemble of CNN, GBDT, and GNN models. </p>\n<h3>CNN models</h3>\n<p>It’s two variations of the great <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public notebook</a> by <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>. <br>\n<strong>Data:</strong> The same dataset that was used in the public notebook (<a href=\"https://www.kaggle.com/datasets/ahmedelfazouan/belka-enc-dataset\" target=\"_blank\">Link</a>). <br>\n<strong>Model architecture:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fd95ff2a9eaa0336ae9dd8ade52459721%2F2024-07-10%201.45.17.png?generation=1720553892516328&amp;alt=media\"><br>\n<strong>The major changes:</strong> We doubled the filter sizes from 32, 64, 96, to 64, 128, 192; kernel sizes were increased from 3, 3, 3, to 19, 9, 3. The ReLU activations were replaced by SiLU. Moreover, for the second model, we added a bidirectional GRU layer after the embedding and concatenated it with the convolution layers after global max pooling. <br>\n<strong>Training parameters:</strong> See the table at the end.<br>\nWe made a weighted average of the above two models trained with and without the validation set, using a total of four CNN models.</p>\n<h2>Other models:</h2>\n<h3>XGBoost (Written in R), LightGBM (Written in Python):</h3>\n<p>To achieve maximum diversity, the models were trained with different features on different subsets of the train data. All models were trained separately for each protein(<a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-gbdt-models-a-part-of-13th-place-solution\" target=\"_blank\">Code Link</a>).<br>\n<strong>Data:</strong> A sample with all binding molecules and a random sample of non-binding molecules (50M or 40M in total for GBDTs and 10M for chemprop )<br>\n<strong>Features:</strong> For one model we added predictions from chemprop (version 2.0, the output of the 3rd linear layer in the FFN) to SECFP4 (bits=1024), and for two we added BB activity features (the fraction of compounds that bind when a given BB smiles occurs at a given position) to ECFP4 (bits=1024). For LightGBM we used SECFP4 (bits=1024) and SECFP6 (bits=2048).<br>\n<strong>Training:</strong> One model was trained five times on 5 parts of the train data, each one without 20% of random BBs and others on the 50M sample (excluding validation and test subsets). <br>\n<strong>XGBoost training parameters:</strong> eta 0.05, max_depth: 25, subsample: 0.2, sampling_method: gradient_based, colsample_bytree: 0.4, min_child_weight: 4, gamma: 2, num_boost_round = 5000, early_stopping_rounds = 30.<br>\n<strong>LightGBM training parameters:</strong> max_depth: 11, bagging_fraction: 0.9, learning_rate: 0.05, colsample_bytree: 1, colsample_bynode: 0.5, lambda_l1: 1, lambda_l2: 1.5, num_leaves: 490, min_data_in_leaf': 50.</p>\n<h3>GNN:</h3>\n<p>We used this public notebook by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> with minor changes (atom types list was truncated to actual atoms in train/test sets molecules).</p>\n<p>For this part, we used weighted average to ensemble the predictions for each protein separately, using local scores (perfect correlation with LB): BRD4: 4 models, HSA: 7 models, sEH: 5 models.</p>\n<h2>2. Non-shared Building Blocks</h2>\n<p>Creating a reliable cross-validation for the molecules with nonshared BBs was difficult, so we conducted two ensemble methods based on the public score, and used them in the final submissions.</p>\n<h3>Final submission 1 (public 0.488/private 0.275): Ranking ensemble</h3>\n<p>For this solution, in order to minimize the fluctuations due to the differences between the public and private scores, we performed an ensemble of the following two ChemBERTa-based models, which gave relatively good predictions for all proteins in the non-share block.We converted the predictions of each model into ranks to account for differences in scale between models. <br>\n<strong>Model architecture:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fc20ab5e45e27faa0636a1009813c52b2%2F2024-07-10%204.45.19.png?generation=1720555352029185&amp;alt=media\"></p>\n<p><strong>Training parameters:</strong> See the table below.<br>\nFor the two models mentioned above, predictions were made for each epoch over five random folds. The average of these predictions was calculated over a total of 5 epochs × 5 folds × 2 models.</p>\n<h3>Final submission 2 (public 0.529/private 0.277): Ranking ensemble by protein</h3>\n<p>We used one XGBoost model (ECFP4 features, trained as above on 5 subsets of train data, but with validation on subsets with non-shared BBs), and seven ChemBERTa models with different classification heads and training parameters. We scored each model for each protein, and considered the scores and diversity of predictions to select the weights (BRD4: 4 models, HSA: 3 models, sEH: 4 models). It is difficult to describe all eight models in detail, so the two models with the highest weights are described in the final submission 1 (feel free to ask anything!).</p>\n<h3>Training parameters:</h3>\n<table>\n<thead>\n<tr>\n<th>Settings</th>\n<th>CNN</th>\n<th>ChemBERTa</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Optimizer</td>\n<td>Adam</td>\n<td>Adam</td>\n</tr>\n<tr>\n<td>Learning rate</td>\n<td>1e-3</td>\n<td>3e-4</td>\n</tr>\n<tr>\n<td>Optimizer momentum</td>\n<td>beta1, beta2 = 0.9, 0.999</td>\n<td>beta1, beta2 = 0.9, 0.999</td>\n</tr>\n<tr>\n<td>Optimizer weight decay</td>\n<td>0.05</td>\n<td>0.01</td>\n</tr>\n<tr>\n<td>Batch size</td>\n<td>4096</td>\n<td>1024</td>\n</tr>\n<tr>\n<td>Training epochs</td>\n<td>50</td>\n<td>5</td>\n</tr>\n<tr>\n<td>ReduceLROnPlateau</td>\n<td>patience, factor = 3, 0.05</td>\n<td>/</td>\n</tr>\n<tr>\n<td>EarlyStopping</td>\n<td>patience = 5</td>\n<td>/</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2914292",
      "postDate": "07/09/2024 20:32:38",
      "content": "<p>We would like to thank Kaggle for organizing such an interesting competition. We also appreciate <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> , <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , and <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> for sharing notebooks that influenced our solution. A big thanks to my teammates <a href=\"https://www.kaggle.com/yyyu54\" target=\"_blank\">@yyyu54</a> , <a href=\"https://www.kaggle.com/Ogurtsov\" target=\"_blank\">@Ogurtsov</a>, and <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> for fighting side by side until we used up all 480 submissions. We all reached the Competitions Master at the same time!</p>\n<h1>Approach Overview</h1>\n<p>We used separate approaches for molecules with shared building blocks and non-shared building blocks.</p>\n<h2>1. Shared Building Blocks</h2>\n<p>We used an ensemble of CNN, GBDT, and GNN models. </p>\n<h3>CNN models</h3>\n<p>It’s two variations of the great <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public notebook</a> by <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a>. <br>\n<strong>Data:</strong> The same dataset that was used in the public notebook (<a href=\"https://www.kaggle.com/datasets/ahmedelfazouan/belka-enc-dataset\" target=\"_blank\">Link</a>). <br>\n<strong>Model architecture:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fd95ff2a9eaa0336ae9dd8ade52459721%2F2024-07-10%201.45.17.png?generation=1720553892516328&amp;alt=media\"><br>\n<strong>The major changes:</strong> We doubled the filter sizes from 32, 64, 96, to 64, 128, 192; kernel sizes were increased from 3, 3, 3, to 19, 9, 3. The ReLU activations were replaced by SiLU. Moreover, for the second model, we added a bidirectional GRU layer after the embedding and concatenated it with the convolution layers after global max pooling. <br>\n<strong>Training parameters:</strong> See the table at the end.<br>\nWe made a weighted average of the above two models trained with and without the validation set, using a total of four CNN models.</p>\n<h2>Other models:</h2>\n<h3>XGBoost (Written in R), LightGBM (Written in Python):</h3>\n<p>To achieve maximum diversity, the models were trained with different features on different subsets of the train data. All models were trained separately for each protein(<a href=\"https://www.kaggle.com/code/antoninadolgorukova/belka-gbdt-models-a-part-of-13th-place-solution\" target=\"_blank\">Code Link</a>).<br>\n<strong>Data:</strong> A sample with all binding molecules and a random sample of non-binding molecules (50M or 40M in total for GBDTs and 10M for chemprop )<br>\n<strong>Features:</strong> For one model we added predictions from chemprop (version 2.0, the output of the 3rd linear layer in the FFN) to SECFP4 (bits=1024), and for two we added BB activity features (the fraction of compounds that bind when a given BB smiles occurs at a given position) to ECFP4 (bits=1024). For LightGBM we used SECFP4 (bits=1024) and SECFP6 (bits=2048).<br>\n<strong>Training:</strong> One model was trained five times on 5 parts of the train data, each one without 20% of random BBs and others on the 50M sample (excluding validation and test subsets). <br>\n<strong>XGBoost training parameters:</strong> eta 0.05, max_depth: 25, subsample: 0.2, sampling_method: gradient_based, colsample_bytree: 0.4, min_child_weight: 4, gamma: 2, num_boost_round = 5000, early_stopping_rounds = 30.<br>\n<strong>LightGBM training parameters:</strong> max_depth: 11, bagging_fraction: 0.9, learning_rate: 0.05, colsample_bytree: 1, colsample_bynode: 0.5, lambda_l1: 1, lambda_l2: 1.5, num_leaves: 490, min_data_in_leaf': 50.</p>\n<h3>GNN:</h3>\n<p>We used this public notebook by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> with minor changes (atom types list was truncated to actual atoms in train/test sets molecules).</p>\n<p>For this part, we used weighted average to ensemble the predictions for each protein separately, using local scores (perfect correlation with LB): BRD4: 4 models, HSA: 7 models, sEH: 5 models.</p>\n<h2>2. Non-shared Building Blocks</h2>\n<p>Creating a reliable cross-validation for the molecules with nonshared BBs was difficult, so we conducted two ensemble methods based on the public score, and used them in the final submissions.</p>\n<h3>Final submission 1 (public 0.488/private 0.275): Ranking ensemble</h3>\n<p>For this solution, in order to minimize the fluctuations due to the differences between the public and private scores, we performed an ensemble of the following two ChemBERTa-based models, which gave relatively good predictions for all proteins in the non-share block.We converted the predictions of each model into ranks to account for differences in scale between models. <br>\n<strong>Model architecture:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fc20ab5e45e27faa0636a1009813c52b2%2F2024-07-10%204.45.19.png?generation=1720555352029185&amp;alt=media\"></p>\n<p><strong>Training parameters:</strong> See the table below.<br>\nFor the two models mentioned above, predictions were made for each epoch over five random folds. The average of these predictions was calculated over a total of 5 epochs × 5 folds × 2 models.</p>\n<h3>Final submission 2 (public 0.529/private 0.277): Ranking ensemble by protein</h3>\n<p>We used one XGBoost model (ECFP4 features, trained as above on 5 subsets of train data, but with validation on subsets with non-shared BBs), and seven ChemBERTa models with different classification heads and training parameters. We scored each model for each protein, and considered the scores and diversity of predictions to select the weights (BRD4: 4 models, HSA: 3 models, sEH: 4 models). It is difficult to describe all eight models in detail, so the two models with the highest weights are described in the final submission 1 (feel free to ask anything!).</p>\n<h3>Training parameters:</h3>\n<table>\n<thead>\n<tr>\n<th>Settings</th>\n<th>CNN</th>\n<th>ChemBERTa</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Optimizer</td>\n<td>Adam</td>\n<td>Adam</td>\n</tr>\n<tr>\n<td>Learning rate</td>\n<td>1e-3</td>\n<td>3e-4</td>\n</tr>\n<tr>\n<td>Optimizer momentum</td>\n<td>beta1, beta2 = 0.9, 0.999</td>\n<td>beta1, beta2 = 0.9, 0.999</td>\n</tr>\n<tr>\n<td>Optimizer weight decay</td>\n<td>0.05</td>\n<td>0.01</td>\n</tr>\n<tr>\n<td>Batch size</td>\n<td>4096</td>\n<td>1024</td>\n</tr>\n<tr>\n<td>Training epochs</td>\n<td>50</td>\n<td>5</td>\n</tr>\n<tr>\n<td>ReduceLROnPlateau</td>\n<td>patience, factor = 3, 0.05</td>\n<td>/</td>\n</tr>\n<tr>\n<td>EarlyStopping</td>\n<td>patience = 5</td>\n<td>/</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "We would like to thank Kaggle for organizing such an interesting competition. We also appreciate @tetsuya3510 , @hengck23 , and @ahmedelfazouan for sharing notebooks that influenced our solution. A big thanks to my teammates @yyyu54 , @Ogurtsov, and @antoninadolgorukova for fighting side by side until we used up all 480 submissions. We all reached the Competitions Master at the same time!\n# Approach Overview\nWe used separate approaches for molecules with shared building blocks and non-shared building blocks.\n## 1. Shared Building Blocks\nWe used an ensemble of CNN, GBDT, and GNN models. \n### CNN models\nIt’s two variations of the great [public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) by @ahmedelfazouan. \n**Data:** The same dataset that was used in the public notebook ([Link](https://www.kaggle.com/datasets/ahmedelfazouan/belka-enc-dataset)). \n**Model architecture:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fd95ff2a9eaa0336ae9dd8ade52459721%2F2024-07-10%201.45.17.png?generation=1720553892516328&alt=media)\n**The major changes:** We doubled the filter sizes from 32, 64, 96, to 64, 128, 192; kernel sizes were increased from 3, 3, 3, to 19, 9, 3. The ReLU activations were replaced by SiLU. Moreover, for the second model, we added a bidirectional GRU layer after the embedding and concatenated it with the convolution layers after global max pooling. \n**Training parameters:** See the table at the end.\nWe made a weighted average of the above two models trained with and without the validation set, using a total of four CNN models.\n## Other models:\n### XGBoost (Written in R), LightGBM (Written in Python):\nTo achieve maximum diversity, the models were trained with different features on different subsets of the train data. All models were trained separately for each protein([Code Link](https://www.kaggle.com/code/antoninadolgorukova/belka-gbdt-models-a-part-of-13th-place-solution)).\n**Data:** A sample with all binding molecules and a random sample of non-binding molecules (50M or 40M in total for GBDTs and 10M for chemprop )\n**Features:** For one model we added predictions from chemprop (version 2.0, the output of the 3rd linear layer in the FFN) to SECFP4 (bits=1024), and for two we added BB activity features (the fraction of compounds that bind when a given BB smiles occurs at a given position) to ECFP4 (bits=1024). For LightGBM we used SECFP4 (bits=1024) and SECFP6 (bits=2048).\n**Training:** One model was trained five times on 5 parts of the train data, each one without 20% of random BBs and others on the 50M sample (excluding validation and test subsets). \n**XGBoost training parameters:** eta 0.05, max_depth: 25, subsample: 0.2, sampling_method: gradient_based, colsample_bytree: 0.4, min_child_weight: 4, gamma: 2, num_boost_round = 5000, early_stopping_rounds = 30.\n**LightGBM training parameters:** max_depth: 11, bagging_fraction: 0.9, learning_rate: 0.05, colsample_bytree: 1, colsample_bynode: 0.5, lambda_l1: 1, lambda_l2: 1.5, num_leaves: 490, min_data_in_leaf': 50.\n### GNN:\nWe used this public notebook by @hengck23 with minor changes (atom types list was truncated to actual atoms in train/test sets molecules).\n\n\nFor this part, we used weighted average to ensemble the predictions for each protein separately, using local scores (perfect correlation with LB): BRD4: 4 models, HSA: 7 models, sEH: 5 models.\n## 2. Non-shared Building Blocks\nCreating a reliable cross-validation for the molecules with nonshared BBs was difficult, so we conducted two ensemble methods based on the public score, and used them in the final submissions.\n### Final submission 1 (public 0.488/private 0.275): Ranking ensemble \nFor this solution, in order to minimize the fluctuations due to the differences between the public and private scores, we performed an ensemble of the following two ChemBERTa-based models, which gave relatively good predictions for all proteins in the non-share block.We converted the predictions of each model into ranks to account for differences in scale between models. \n**Model architecture:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fc20ab5e45e27faa0636a1009813c52b2%2F2024-07-10%204.45.19.png?generation=1720555352029185&alt=media)\n\n**Training parameters:** See the table below.\nFor the two models mentioned above, predictions were made for each epoch over five random folds. The average of these predictions was calculated over a total of 5 epochs × 5 folds × 2 models.\n\n### Final submission 2 (public 0.529/private 0.277): Ranking ensemble by protein\nWe used one XGBoost model (ECFP4 features, trained as above on 5 subsets of train data, but with validation on subsets with non-shared BBs), and seven ChemBERTa models with different classification heads and training parameters. We scored each model for each protein, and considered the scores and diversity of predictions to select the weights (BRD4: 4 models, HSA: 3 models, sEH: 4 models). It is difficult to describe all eight models in detail, so the two models with the highest weights are described in the final submission 1 (feel free to ask anything!).\n\n### Training parameters:\n\n| Settings               | CNN                        | ChemBERTa                 |\n| ---------------------- | -------------------------- | ------------------------- |\n| Optimizer              | Adam                       | Adam                      |\n| Learning rate          | 1e-3                       | 3e-4                      |\n| Optimizer momentum     | beta1, beta2 = 0.9, 0.999  | beta1, beta2 = 0.9, 0.999 |\n| Optimizer weight decay | 0.05                       | 0.01                      |\n| Batch size             | 4096                       | 1024                      |\n| Training epochs        | 50                         | 5                         |\n| ReduceLROnPlateau      | patience, factor = 3, 0.05 | /                         |\n| EarlyStopping          | patience = 5               | /                         |",
      "votes": null
    },
    {
      "id": "2914425",
      "postDate": "07/09/2024 23:00:34",
      "content": "<p>Thanks for sharing and congratulations!</p>",
      "rawMarkdown": "Thanks for sharing and congratulations!",
      "votes": null
    },
    {
      "id": "2914552",
      "postDate": "07/10/2024 03:04:28",
      "content": "<p>I also used ChemBerta 77-MTR as one of my models. It's good to see you used it as well. 😁</p>",
      "rawMarkdown": "I also used ChemBerta 77-MTR as one of my models. It's good to see you used it as well. 😁",
      "votes": null
    },
    {
      "id": "2915082",
      "postDate": "07/10/2024 10:19:36",
      "content": "<p>Here are the correlations between the public and private total nonshare scores (all predictions for the molecules with shared BBs = 0).<br>\nThe models included in Final submission 2 are marked in red.</p>\n<p>The two ChemBERTa models described in the solution - are №18 and №19. The XGb model (№32) has the worst score (0.122 public/0.039 private), but was only used for BRD4, for which it gives the same score (0.106 public/0.025 private) as ChemBERTa №19 (0.106 public/0.024 private).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F8b413643a37e01614bb6e58b9fa9f46e%2FScreenshot%202024-07-10%20131027.png?generation=1720606760013833&amp;alt=media\"></p>",
      "rawMarkdown": "Here are the correlations between the public and private total nonshare scores (all predictions for the molecules with shared BBs = 0).\nThe models included in Final submission 2 are marked in red.\n\nThe two ChemBERTa models described in the solution - are №18 and №19. The XGb model (№32) has the worst score (0.122 public/0.039 private), but was only used for BRD4, for which it gives the same score (0.106 public/0.025 private) as ChemBERTa №19 (0.106 public/0.024 private).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F8b413643a37e01614bb6e58b9fa9f46e%2FScreenshot%202024-07-10%20131027.png?generation=1720606760013833&alt=media)",
      "votes": null
    },
    {
      "id": "2915180",
      "postDate": "07/10/2024 11:28:56",
      "content": "<p>wow thanks for sharing your work! very impressive indeed!</p>",
      "rawMarkdown": "wow thanks for sharing your work! very impressive indeed!",
      "votes": null
    },
    {
      "id": "2915877",
      "postDate": "07/10/2024 17:22:15",
      "content": "<p>Congratulations , and thanks for sharing your work</p>",
      "rawMarkdown": "Congratulations , and thanks for sharing your work",
      "votes": null
    },
    {
      "id": "2917729",
      "postDate": "07/11/2024 18:41:08",
      "content": "<p>Thank you so much for sharing! I opted in for this competition, but it was very daunting, I only used RDKit to explore molecule structures but wasn't able to complete the whole process. Your process is so much more streamlined and also very effective 😀</p>",
      "rawMarkdown": "Thank you so much for sharing! I opted in for this competition, but it was very daunting, I only used RDKit to explore molecule structures but wasn't able to complete the whole process. Your process is so much more streamlined and also very effective 😀",
      "votes": null
    },
    {
      "id": "2988211",
      "postDate": "09/13/2024 14:25:39",
      "content": "<p>We have recently upload the <a href=\"https://github.com/statist-bhfz/leash_bio_13th_place_solution\" target=\"_blank\">train &amp; inference code for all models</a>.</p>",
      "rawMarkdown": "We have recently upload the [train & inference code for all models](https://github.com/statist-bhfz/leash_bio_13th_place_solution).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2914425,
      "author_name": "drpatrickchan",
      "author_url": "",
      "post_date": "07/09/2024 23:00:34",
      "content": "<p>Thanks for sharing and congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2914552,
      "author_name": "ahsuna123",
      "author_url": "",
      "post_date": "07/10/2024 03:04:28",
      "content": "<p>I also used ChemBerta 77-MTR as one of my models. It's good to see you used it as well. 😁</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2915082,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "07/10/2024 10:19:36",
      "content": "<p>Here are the correlations between the public and private total nonshare scores (all predictions for the molecules with shared BBs = 0).<br>\nThe models included in Final submission 2 are marked in red.</p>\n<p>The two ChemBERTa models described in the solution - are №18 and №19. The XGb model (№32) has the worst score (0.122 public/0.039 private), but was only used for BRD4, for which it gives the same score (0.106 public/0.025 private) as ChemBERTa №19 (0.106 public/0.024 private).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F8b413643a37e01614bb6e58b9fa9f46e%2FScreenshot%202024-07-10%20131027.png?generation=1720606760013833&amp;alt=media\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2915180,
      "author_name": "amulyaincorrigible",
      "author_url": "",
      "post_date": "07/10/2024 11:28:56",
      "content": "<p>wow thanks for sharing your work! very impressive indeed!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2915877,
      "author_name": "staratnyte",
      "author_url": "",
      "post_date": "07/10/2024 17:22:15",
      "content": "<p>Congratulations , and thanks for sharing your work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2917729,
      "author_name": "kaavyamahajan",
      "author_url": "",
      "post_date": "07/11/2024 18:41:08",
      "content": "<p>Thank you so much for sharing! I opted in for this competition, but it was very daunting, I only used RDKit to explore molecule structures but wasn't able to complete the whole process. Your process is so much more streamlined and also very effective 😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2988211,
      "author_name": "ogurtsov",
      "author_url": "",
      "post_date": "09/13/2024 14:25:39",
      "content": "<p>We have recently upload the <a href=\"https://github.com/statist-bhfz/leash_bio_13th_place_solution\" target=\"_blank\">train &amp; inference code for all models</a>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2914292": "We would like to thank Kaggle for organizing such an interesting competition. We also appreciate @tetsuya3510 , @hengck23 , and @ahmedelfazouan for sharing notebooks that influenced our solution. A big thanks to my teammates @yyyu54 , @Ogurtsov, and @antoninadolgorukova for fighting side by side until we used up all 480 submissions. We all reached the Competitions Master at the same time!\n# Approach Overview\nWe used separate approaches for molecules with shared building blocks and non-shared building blocks.\n## 1. Shared Building Blocks\nWe used an ensemble of CNN, GBDT, and GNN models. \n### CNN models\nIt’s two variations of the great [public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) by @ahmedelfazouan. \n**Data:** The same dataset that was used in the public notebook ([Link](https://www.kaggle.com/datasets/ahmedelfazouan/belka-enc-dataset)). \n**Model architecture:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fd95ff2a9eaa0336ae9dd8ade52459721%2F2024-07-10%201.45.17.png?generation=1720553892516328&alt=media)\n**The major changes:** We doubled the filter sizes from 32, 64, 96, to 64, 128, 192; kernel sizes were increased from 3, 3, 3, to 19, 9, 3. The ReLU activations were replaced by SiLU. Moreover, for the second model, we added a bidirectional GRU layer after the embedding and concatenated it with the convolution layers after global max pooling. \n**Training parameters:** See the table at the end.\nWe made a weighted average of the above two models trained with and without the validation set, using a total of four CNN models.\n## Other models:\n### XGBoost (Written in R), LightGBM (Written in Python):\nTo achieve maximum diversity, the models were trained with different features on different subsets of the train data. All models were trained separately for each protein([Code Link](https://www.kaggle.com/code/antoninadolgorukova/belka-gbdt-models-a-part-of-13th-place-solution)).\n**Data:** A sample with all binding molecules and a random sample of non-binding molecules (50M or 40M in total for GBDTs and 10M for chemprop )\n**Features:** For one model we added predictions from chemprop (version 2.0, the output of the 3rd linear layer in the FFN) to SECFP4 (bits=1024), and for two we added BB activity features (the fraction of compounds that bind when a given BB smiles occurs at a given position) to ECFP4 (bits=1024). For LightGBM we used SECFP4 (bits=1024) and SECFP6 (bits=2048).\n**Training:** One model was trained five times on 5 parts of the train data, each one without 20% of random BBs and others on the 50M sample (excluding validation and test subsets). \n**XGBoost training parameters:** eta 0.05, max_depth: 25, subsample: 0.2, sampling_method: gradient_based, colsample_bytree: 0.4, min_child_weight: 4, gamma: 2, num_boost_round = 5000, early_stopping_rounds = 30.\n**LightGBM training parameters:** max_depth: 11, bagging_fraction: 0.9, learning_rate: 0.05, colsample_bytree: 1, colsample_bynode: 0.5, lambda_l1: 1, lambda_l2: 1.5, num_leaves: 490, min_data_in_leaf': 50.\n### GNN:\nWe used this public notebook by @hengck23 with minor changes (atom types list was truncated to actual atoms in train/test sets molecules).\n\n\nFor this part, we used weighted average to ensemble the predictions for each protein separately, using local scores (perfect correlation with LB): BRD4: 4 models, HSA: 7 models, sEH: 5 models.\n## 2. Non-shared Building Blocks\nCreating a reliable cross-validation for the molecules with nonshared BBs was difficult, so we conducted two ensemble methods based on the public score, and used them in the final submissions.\n### Final submission 1 (public 0.488/private 0.275): Ranking ensemble \nFor this solution, in order to minimize the fluctuations due to the differences between the public and private scores, we performed an ensemble of the following two ChemBERTa-based models, which gave relatively good predictions for all proteins in the non-share block.We converted the predictions of each model into ranks to account for differences in scale between models. \n**Model architecture:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8574776%2Fc20ab5e45e27faa0636a1009813c52b2%2F2024-07-10%204.45.19.png?generation=1720555352029185&alt=media)\n\n**Training parameters:** See the table below.\nFor the two models mentioned above, predictions were made for each epoch over five random folds. The average of these predictions was calculated over a total of 5 epochs × 5 folds × 2 models.\n\n### Final submission 2 (public 0.529/private 0.277): Ranking ensemble by protein\nWe used one XGBoost model (ECFP4 features, trained as above on 5 subsets of train data, but with validation on subsets with non-shared BBs), and seven ChemBERTa models with different classification heads and training parameters. We scored each model for each protein, and considered the scores and diversity of predictions to select the weights (BRD4: 4 models, HSA: 3 models, sEH: 4 models). It is difficult to describe all eight models in detail, so the two models with the highest weights are described in the final submission 1 (feel free to ask anything!).\n\n### Training parameters:\n\n| Settings               | CNN                        | ChemBERTa                 |\n| ---------------------- | -------------------------- | ------------------------- |\n| Optimizer              | Adam                       | Adam                      |\n| Learning rate          | 1e-3                       | 3e-4                      |\n| Optimizer momentum     | beta1, beta2 = 0.9, 0.999  | beta1, beta2 = 0.9, 0.999 |\n| Optimizer weight decay | 0.05                       | 0.01                      |\n| Batch size             | 4096                       | 1024                      |\n| Training epochs        | 50                         | 5                         |\n| ReduceLROnPlateau      | patience, factor = 3, 0.05 | /                         |\n| EarlyStopping          | patience = 5               | /                         |",
    "2914425": "Thanks for sharing and congratulations!",
    "2914552": "I also used ChemBerta 77-MTR as one of my models. It's good to see you used it as well. 😁",
    "2915082": "Here are the correlations between the public and private total nonshare scores (all predictions for the molecules with shared BBs = 0).\nThe models included in Final submission 2 are marked in red.\n\nThe two ChemBERTa models described in the solution - are №18 and №19. The XGb model (№32) has the worst score (0.122 public/0.039 private), but was only used for BRD4, for which it gives the same score (0.106 public/0.025 private) as ChemBERTa №19 (0.106 public/0.024 private).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F8b413643a37e01614bb6e58b9fa9f46e%2FScreenshot%202024-07-10%20131027.png?generation=1720606760013833&alt=media)",
    "2915180": "wow thanks for sharing your work! very impressive indeed!",
    "2915877": "Congratulations , and thanks for sharing your work",
    "2917729": "Thank you so much for sharing! I opted in for this competition, but it was very daunting, I only used RDKit to explore molecule structures but wasn't able to complete the whole process. Your process is so much more streamlined and also very effective 😀",
    "2988211": "We have recently upload the [train & inference code for all models](https://github.com/statist-bhfz/leash_bio_13th_place_solution)."
  },
  "source": "meta"
}