{
  "id": 519191,
  "title": "27th Place Solution (8th in Public):  1DCNN for share and ChemBERTa for non-share",
  "url": "/competitions/leash-BELKA/discussion/519191",
  "author_name": "shina",
  "post_date": "2024-07-10T05:08:08.510000",
  "votes": 21,
  "comment_count": 0,
  "views": 0,
  "content": "<p>We adopted different strategies for shared and nonshared targets.</p>\n<h1>1. Nonshared Target Part</h1>\n<h2>CV Strategy</h2>\n<p>The split was made based on building blocks, as we saw in this helpful <a href=\"https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host\" target=\"_blank\">notebook</a>. Each building block was sorted and then systematically sampled to avoid unbalanced sampling.</p>\n<h2>Model</h2>\n<p>We used <a href=\"https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline\" target=\"_blank\">ChemBERTa-10M-MTR</a> as our base model for each protein</p>\n<h2>Training</h2>\n<h3>Fine-tuning:</h3>\n<ul>\n<li>For all proteins: We used 4M sampling for fine-tuning each model.</li>\n<li>We fine-tuned each model using 4M samples for the respective protein.</li>\n<li>For sEH Triazine core: We used a mixed dataset, with half the data from external sources and half from the regular training data.</li>\n<li>For sEH NonTriazine core: We used a dataset trained exclusively on <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">external data containing various cores</a>. This approach was chosen to potentially improve predictions for unknown cores.</li>\n<li>Overfitting Prevention: To mitigate overfitting, we froze the weights of the first three layers of the ChemBERTa model.</li>\n</ul>\n<h2>Ensemble Method</h2>\n<ul>\n<li><p>For most proteins: We tested three ensemble methods - simple average, CV-weighted average, and CV-weighted average to the power of 5. We chose the CV-weighted average to the power of 5 as it yielded the highest LB score.</p>\n<table>\n<thead>\n<tr>\n<th>Ensemble Method</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Simple Average</td>\n<td>0.486</td>\n<td>0.272</td>\n</tr>\n<tr>\n<td>CV-weighted Average</td>\n<td>0.488</td>\n<td>0.271</td>\n</tr>\n<tr>\n<td>CV-weighted Average^5</td>\n<td>0.496</td>\n<td>0.266</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>For sEH NonTriazine core: We used a simple average ensemble of the three runs with different seeds.</p></li>\n</ul>\n<p>Performance: The final Public LB score (mask share) for our nonshared target models was 0.173.</p>\n<h1>2. Shared Target Part</h1>\n<p>For shared targets, we used a 1DCNN model, improving upon the Kaggle <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public notebook</a>.</p>\n<h2>CV Strategy</h2>\n<p>We used 5-fold split generated by StratifiedKFold. All training data was used. </p>\n<h2>Model</h2>\n<ul>\n<li>Increased number of filters (32 to 64, doubling in each layer)</li>\n<li>Enhanced convolution layers (padding='same', varied kernel sizes)</li>\n<li>Added BatchNormalization</li>\n<li>Introduced shortcut connections and Attention mechanism</li>\n<li>Adjusted Dropout rate (0.1 to 0.2)</li>\n<li>Modified fully connected layers<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17343385%2F4ba074c5c07030d771f907f59c223b62%2F1DCNN_Model.png?generation=1720586518917501&amp;alt=media\"></li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>Increased number of epochs (20 to 30)</li>\n<li>Reduced patience in ReduceLROnPlateau (5 to 3)</li>\n<li>Allowed slight overfitting</li>\n</ul>\n<p>These improvements aimed to intentionally overfit the model slightly, enabling it to learn shared target data more deeply and improve scores. This strategy allowed the model to capture more detailed features of the training data, leading to improved predictive performance.</p>\n<h2>Score</h2>\n<p>We compared the LB scores of fold0 out of 15 folds to evaluate our model improvements. The results are shown in the table below:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Original notebook model</td>\n<td>0.359</td>\n<td>0.217</td>\n</tr>\n<tr>\n<td>Improved model</td>\n<td>0.387</td>\n<td>0.220</td>\n</tr>\n</tbody>\n</table>\n<p>These results demonstrate the effectiveness of our improvement strategy. By increasing model complexity and allowing deeper learning, we significantly improved our score on the public leaderboard.<br>\nAfter confirming the improvement through this single fold comparison, we proceeded with our final strategy:</p>\n<ul>\n<li>We created a new model with 5 folds, separate from the 15-fold model used for comparison.</li>\n<li>We applied our improved model architecture to all 5 folds.</li>\n<li>Finally, we created an average ensemble using these 5 folds to generate our final predictions.</li>\n</ul>\n<p>The final performance of our 5-fold ensemble 1DCNN model on the Public LB (mask nonshare) was 0.347.<br>\nThis approach allowed us to leverage the benefits of our model improvements while also utilizing ensemble techniques to further enhance our predictions.</p>\n<h1>Final Model Performance:</h1>\n<table>\n<thead>\n<tr>\n<th>Leaderboard</th>\n<th>Share(Mask NonShare)</th>\n<th>NonShare(Mask share)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public</td>\n<td>0.347</td>\n<td>0.173</td>\n</tr>\n<tr>\n<td>Private</td>\n<td>0.224</td>\n<td>0.059</td>\n</tr>\n</tbody>\n</table>\n<h1>Tried but not go well</h1>\n<ul>\n<li>Stacking approach for share and non-share (A model was constructed for each protein)<ul>\n<li>Models: 1DCNN, LGBM2, ChemBERTa for share, and ChemBERTa, 1DCNN, LGBM1, and LGBM2 for non-share (LGBM1 used morgan4+rdkit+descriptors and all data for training, LGBM2 used ECFP6 and 10M data for training, others used 10M data for training)</li></ul></li>\n<li>Procedure:</li>\n</ul>\n<ol>\n<li>By using each model, the prediction of each fold of cross-validation was carried out, and it become x. The bind values from train data the competition organizer gave became y. Catboost was used as stacking model because Catboost is good approach for category data such as bb. (but in fact, we didnot use bb as input…)</li>\n<li>To predict binds of submission, bind values of test data predicted by each model were used as x. </li>\n<li>Finally, predictions of share and non-share were merged. </li>\n</ol>\n<ul>\n<li><p>Share  LB:0.506→0.496,  PB 0.261 → 0.253</p>\n<ul>\n<li>Code：<a href=\"https://www.kaggle.com/code/kokishinbara/catboost-share-protein-loop/output\" target=\"_blank\">https://www.kaggle.com/code/kokishinbara/catboost-share-protein-loop/output</a></li>\n<li>Result: <a href=\"https://www.kaggle.com/datasets/kokishinbara/stacking-share\" target=\"_blank\">https://www.kaggle.com/datasets/kokishinbara/stacking-share</a></li></ul></li>\n<li><p>Nonshare LB0.506 → 0.497,   PB 0.261 → 0.263 </p>\n<ul>\n<li>Code：<a href=\"https://github.com/kyu999/leash-bio/blob/kyue/notebooks/stacking/nonshare_stacking.ipynb\" target=\"_blank\">https://github.com/kyu999/leash-bio/blob/kyue/notebooks/stacking/nonshare_stacking.ipynb</a></li>\n<li>Result：<a href=\"https://www.kaggle.com/datasets/kyu999/stacking-nonshare-multi-models\" target=\"_blank\">https://www.kaggle.com/datasets/kyu999/stacking-nonshare-multi-models</a></li></ul></li>\n<li><p>We tried some graph-based approatch, GCN, 3DGNN (Sphere), but did not work for us. </p></li>\n<li><p>We also tried adjusting embedding, by enumeration of the SMILES string and it did not work. </p></li>\n</ul>\n<h2>Others…</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook\" target=\"_blank\">Based on Scaffold explonation</a>, <a href=\"https://www.kaggle.com/code/kokishinbara/scaffold-exploration-num\" target=\"_blank\">Not triazine core was searched</a> and, <a href=\"https://www.kaggle.com/code/kokishinbara/without-core-leash-bio-chemberta-baseline\" target=\"_blank\">the train data without cores are trained but not go well</a>. (model = ChemBERTa, LB=0.281 PB=0.161) (Please ignore the comments in our notebook, just copied the original and refered to the code.)</li>\n<li><a href=\"https://www.kaggle.com/code/shunsukekikuchi/catboost-share-one-model/notebook\" target=\"_blank\">Catboost for stacking also used one model for proteins, but didn't go well</a> (CV of 1 fold BRD4: 0.326  HSA: 0.138  sEH: 0.680 )</li>\n<li>Experiments with Small-Scale Samples<ul>\n<li>Feature Addition: Added chemical properties as features to the LGBM model, alongside molecular fingerprints.</li>\n<li>Fingerprint Comparison: Compared the performance of multiple types of molecular fingerprints.</li>\n<li>ECFP Radius Optimization: Investigated and determined the optimal radius for ECFP.</li>\n<li>Generalization Performance: Validated the use of sEH external data to enhance the model's generalization performance.</li>\n<li>Data Filtering Impact: Examined the impact on model precision by filtering sEH external data using read count.</li>\n<li>Embedding Replacement: Replaced molecular fingerprints with Malformer embeddings to evaluate performance changes.</li>\n<li>Data Augmentation: Evaluated the effectiveness of augmented data generated using SMILES enumeration.</li>\n<li>Model Comparison: Compared the performance of the original model, ChemBERTa (10M smiles), with the updated model, ChemBERTa2 (77M smiles).</li></ul></li>\n</ul>\n<h1>Thanks</h1>\n<p>We would like to thank Kaggle for organizing such an interesting competition. We also appreciate <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> and <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> for sharing notebooks that influenced our solution. A big thanks to my teammates <a href=\"https://www.kaggle.com/kyu999\" target=\"_blank\">@kyu999</a>, <a href=\"https://www.kaggle.com/shunsukekikuchi\" target=\"_blank\">@shunsukekikuchi</a>, and <a href=\"https://www.kaggle.com/kokishinbara\" target=\"_blank\">@kokishinbara</a> for their dedication, collaboration, and tireless efforts. Without their hard work and teamwork, our achievements wouldn't have been possible.</p>",
  "messages": [
    {
      "id": 2914642,
      "postDate": "2024-07-10T05:08:08.510Z",
      "content": "<p>We adopted different strategies for shared and nonshared targets.</p>\n<h1>1. Nonshared Target Part</h1>\n<h2>CV Strategy</h2>\n<p>The split was made based on building blocks, as we saw in this helpful <a href=\"https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host\" target=\"_blank\">notebook</a>. Each building block was sorted and then systematically sampled to avoid unbalanced sampling.</p>\n<h2>Model</h2>\n<p>We used <a href=\"https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline\" target=\"_blank\">ChemBERTa-10M-MTR</a> as our base model for each protein</p>\n<h2>Training</h2>\n<h3>Fine-tuning:</h3>\n<ul>\n<li>For all proteins: We used 4M sampling for fine-tuning each model.</li>\n<li>We fine-tuned each model using 4M samples for the respective protein.</li>\n<li>For sEH Triazine core: We used a mixed dataset, with half the data from external sources and half from the regular training data.</li>\n<li>For sEH NonTriazine core: We used a dataset trained exclusively on <a href=\"https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57\" target=\"_blank\">external data containing various cores</a>. This approach was chosen to potentially improve predictions for unknown cores.</li>\n<li>Overfitting Prevention: To mitigate overfitting, we froze the weights of the first three layers of the ChemBERTa model.</li>\n</ul>\n<h2>Ensemble Method</h2>\n<ul>\n<li><p>For most proteins: We tested three ensemble methods - simple average, CV-weighted average, and CV-weighted average to the power of 5. We chose the CV-weighted average to the power of 5 as it yielded the highest LB score.</p>\n<table>\n<thead>\n<tr>\n<th>Ensemble Method</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Simple Average</td>\n<td>0.486</td>\n<td>0.272</td>\n</tr>\n<tr>\n<td>CV-weighted Average</td>\n<td>0.488</td>\n<td>0.271</td>\n</tr>\n<tr>\n<td>CV-weighted Average^5</td>\n<td>0.496</td>\n<td>0.266</td>\n</tr>\n</tbody>\n</table></li>\n<li><p>For sEH NonTriazine core: We used a simple average ensemble of the three runs with different seeds.</p></li>\n</ul>\n<p>Performance: The final Public LB score (mask share) for our nonshared target models was 0.173.</p>\n<h1>2. Shared Target Part</h1>\n<p>For shared targets, we used a 1DCNN model, improving upon the Kaggle <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public notebook</a>.</p>\n<h2>CV Strategy</h2>\n<p>We used 5-fold split generated by StratifiedKFold. All training data was used. </p>\n<h2>Model</h2>\n<ul>\n<li>Increased number of filters (32 to 64, doubling in each layer)</li>\n<li>Enhanced convolution layers (padding='same', varied kernel sizes)</li>\n<li>Added BatchNormalization</li>\n<li>Introduced shortcut connections and Attention mechanism</li>\n<li>Adjusted Dropout rate (0.1 to 0.2)</li>\n<li>Modified fully connected layers<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17343385%2F4ba074c5c07030d771f907f59c223b62%2F1DCNN_Model.png?generation=1720586518917501&amp;alt=media\"></li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>Increased number of epochs (20 to 30)</li>\n<li>Reduced patience in ReduceLROnPlateau (5 to 3)</li>\n<li>Allowed slight overfitting</li>\n</ul>\n<p>These improvements aimed to intentionally overfit the model slightly, enabling it to learn shared target data more deeply and improve scores. This strategy allowed the model to capture more detailed features of the training data, leading to improved predictive performance.</p>\n<h2>Score</h2>\n<p>We compared the LB scores of fold0 out of 15 folds to evaluate our model improvements. The results are shown in the table below:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Original notebook model</td>\n<td>0.359</td>\n<td>0.217</td>\n</tr>\n<tr>\n<td>Improved model</td>\n<td>0.387</td>\n<td>0.220</td>\n</tr>\n</tbody>\n</table>\n<p>These results demonstrate the effectiveness of our improvement strategy. By increasing model complexity and allowing deeper learning, we significantly improved our score on the public leaderboard.<br>\nAfter confirming the improvement through this single fold comparison, we proceeded with our final strategy:</p>\n<ul>\n<li>We created a new model with 5 folds, separate from the 15-fold model used for comparison.</li>\n<li>We applied our improved model architecture to all 5 folds.</li>\n<li>Finally, we created an average ensemble using these 5 folds to generate our final predictions.</li>\n</ul>\n<p>The final performance of our 5-fold ensemble 1DCNN model on the Public LB (mask nonshare) was 0.347.<br>\nThis approach allowed us to leverage the benefits of our model improvements while also utilizing ensemble techniques to further enhance our predictions.</p>\n<h1>Final Model Performance:</h1>\n<table>\n<thead>\n<tr>\n<th>Leaderboard</th>\n<th>Share(Mask NonShare)</th>\n<th>NonShare(Mask share)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Public</td>\n<td>0.347</td>\n<td>0.173</td>\n</tr>\n<tr>\n<td>Private</td>\n<td>0.224</td>\n<td>0.059</td>\n</tr>\n</tbody>\n</table>\n<h1>Tried but not go well</h1>\n<ul>\n<li>Stacking approach for share and non-share (A model was constructed for each protein)<ul>\n<li>Models: 1DCNN, LGBM2, ChemBERTa for share, and ChemBERTa, 1DCNN, LGBM1, and LGBM2 for non-share (LGBM1 used morgan4+rdkit+descriptors and all data for training, LGBM2 used ECFP6 and 10M data for training, others used 10M data for training)</li></ul></li>\n<li>Procedure:</li>\n</ul>\n<ol>\n<li>By using each model, the prediction of each fold of cross-validation was carried out, and it become x. The bind values from train data the competition organizer gave became y. Catboost was used as stacking model because Catboost is good approach for category data such as bb. (but in fact, we didnot use bb as input…)</li>\n<li>To predict binds of submission, bind values of test data predicted by each model were used as x. </li>\n<li>Finally, predictions of share and non-share were merged. </li>\n</ol>\n<ul>\n<li><p>Share  LB:0.506→0.496,  PB 0.261 → 0.253</p>\n<ul>\n<li>Code：<a href=\"https://www.kaggle.com/code/kokishinbara/catboost-share-protein-loop/output\" target=\"_blank\">https://www.kaggle.com/code/kokishinbara/catboost-share-protein-loop/output</a></li>\n<li>Result: <a href=\"https://www.kaggle.com/datasets/kokishinbara/stacking-share\" target=\"_blank\">https://www.kaggle.com/datasets/kokishinbara/stacking-share</a></li></ul></li>\n<li><p>Nonshare LB0.506 → 0.497,   PB 0.261 → 0.263 </p>\n<ul>\n<li>Code：<a href=\"https://github.com/kyu999/leash-bio/blob/kyue/notebooks/stacking/nonshare_stacking.ipynb\" target=\"_blank\">https://github.com/kyu999/leash-bio/blob/kyue/notebooks/stacking/nonshare_stacking.ipynb</a></li>\n<li>Result：<a href=\"https://www.kaggle.com/datasets/kyu999/stacking-nonshare-multi-models\" target=\"_blank\">https://www.kaggle.com/datasets/kyu999/stacking-nonshare-multi-models</a></li></ul></li>\n<li><p>We tried some graph-based approatch, GCN, 3DGNN (Sphere), but did not work for us. </p></li>\n<li><p>We also tried adjusting embedding, by enumeration of the SMILES string and it did not work. </p></li>\n</ul>\n<h2>Others…</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook\" target=\"_blank\">Based on Scaffold explonation</a>, <a href=\"https://www.kaggle.com/code/kokishinbara/scaffold-exploration-num\" target=\"_blank\">Not triazine core was searched</a> and, <a href=\"https://www.kaggle.com/code/kokishinbara/without-core-leash-bio-chemberta-baseline\" target=\"_blank\">the train data without cores are trained but not go well</a>. (model = ChemBERTa, LB=0.281 PB=0.161) (Please ignore the comments in our notebook, just copied the original and refered to the code.)</li>\n<li><a href=\"https://www.kaggle.com/code/shunsukekikuchi/catboost-share-one-model/notebook\" target=\"_blank\">Catboost for stacking also used one model for proteins, but didn't go well</a> (CV of 1 fold BRD4: 0.326  HSA: 0.138  sEH: 0.680 )</li>\n<li>Experiments with Small-Scale Samples<ul>\n<li>Feature Addition: Added chemical properties as features to the LGBM model, alongside molecular fingerprints.</li>\n<li>Fingerprint Comparison: Compared the performance of multiple types of molecular fingerprints.</li>\n<li>ECFP Radius Optimization: Investigated and determined the optimal radius for ECFP.</li>\n<li>Generalization Performance: Validated the use of sEH external data to enhance the model's generalization performance.</li>\n<li>Data Filtering Impact: Examined the impact on model precision by filtering sEH external data using read count.</li>\n<li>Embedding Replacement: Replaced molecular fingerprints with Malformer embeddings to evaluate performance changes.</li>\n<li>Data Augmentation: Evaluated the effectiveness of augmented data generated using SMILES enumeration.</li>\n<li>Model Comparison: Compared the performance of the original model, ChemBERTa (10M smiles), with the updated model, ChemBERTa2 (77M smiles).</li></ul></li>\n</ul>\n<h1>Thanks</h1>\n<p>We would like to thank Kaggle for organizing such an interesting competition. We also appreciate <a href=\"https://www.kaggle.com/ahmedelfazouan\" target=\"_blank\">@ahmedelfazouan</a> and <a href=\"https://www.kaggle.com/tetsuya3510\" target=\"_blank\">@tetsuya3510</a> for sharing notebooks that influenced our solution. A big thanks to my teammates <a href=\"https://www.kaggle.com/kyu999\" target=\"_blank\">@kyu999</a>, <a href=\"https://www.kaggle.com/shunsukekikuchi\" target=\"_blank\">@shunsukekikuchi</a>, and <a href=\"https://www.kaggle.com/kokishinbara\" target=\"_blank\">@kokishinbara</a> for their dedication, collaboration, and tireless efforts. Without their hard work and teamwork, our achievements wouldn't have been possible.</p>",
      "rawMarkdown": "We adopted different strategies for shared and nonshared targets.\n\n# 1. Nonshared Target Part\n\n## CV Strategy\nThe split was made based on building blocks, as we saw in this helpful [notebook](https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host). Each building block was sorted and then systematically sampled to avoid unbalanced sampling.\n\n## Model\nWe used [ChemBERTa-10M-MTR](https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline) as our base model for each protein\n## Training\n### Fine-tuning: \n- For all proteins: We used 4M sampling for fine-tuning each model.\n- We fine-tuned each model using 4M samples for the respective protein.\n- For sEH Triazine core: We used a mixed dataset, with half the data from external sources and half from the regular training data.\n- For sEH NonTriazine core: We used a dataset trained exclusively on [external data containing various cores](https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57). This approach was chosen to potentially improve predictions for unknown cores.\n- Overfitting Prevention: To mitigate overfitting, we froze the weights of the first three layers of the ChemBERTa model.\n\n## Ensemble Method\n- For most proteins: We tested three ensemble methods - simple average, CV-weighted average, and CV-weighted average to the power of 5. We chose the CV-weighted average to the power of 5 as it yielded the highest LB score.\n| Ensemble Method | Public | Private |\n|---|---|---|\n| Simple Average | 0.486 | 0.272 |\n| CV-weighted Average | 0.488 | 0.271 |\n| CV-weighted Average^5 | 0.496 | 0.266 |\n\n- For sEH NonTriazine core: We used a simple average ensemble of the three runs with different seeds.\n\nPerformance: The final Public LB score (mask share) for our nonshared target models was 0.173.\n\n# 2. Shared Target Part\nFor shared targets, we used a 1DCNN model, improving upon the Kaggle [public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data).\n\n## CV Strategy\nWe used 5-fold split generated by StratifiedKFold. All training data was used. \n## Model\n- Increased number of filters (32 to 64, doubling in each layer)\n- Enhanced convolution layers (padding='same', varied kernel sizes)\n- Added BatchNormalization\n- Introduced shortcut connections and Attention mechanism\n- Adjusted Dropout rate (0.1 to 0.2)\n- Modified fully connected layers\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17343385%2F4ba074c5c07030d771f907f59c223b62%2F1DCNN_Model.png?generation=1720586518917501&alt=media)\n\n## Training\n- Increased number of epochs (20 to 30)\n- Reduced patience in ReduceLROnPlateau (5 to 3)\n- Allowed slight overfitting\n\nThese improvements aimed to intentionally overfit the model slightly, enabling it to learn shared target data more deeply and improve scores. This strategy allowed the model to capture more detailed features of the training data, leading to improved predictive performance.\n\n## Score\nWe compared the LB scores of fold0 out of 15 folds to evaluate our model improvements. The results are shown in the table below:\n| Model | Public | Private |\n| --- | --- |\n| Original notebook model | 0.359 | 0.217 |\n| Improved model | 0.387 | 0.220 |\n\n\nThese results demonstrate the effectiveness of our improvement strategy. By increasing model complexity and allowing deeper learning, we significantly improved our score on the public leaderboard.\nAfter confirming the improvement through this single fold comparison, we proceeded with our final strategy:\n- We created a new model with 5 folds, separate from the 15-fold model used for comparison.\n- We applied our improved model architecture to all 5 folds.\n- Finally, we created an average ensemble using these 5 folds to generate our final predictions.\n\nThe final performance of our 5-fold ensemble 1DCNN model on the Public LB (mask nonshare) was 0.347.\nThis approach allowed us to leverage the benefits of our model improvements while also utilizing ensemble techniques to further enhance our predictions.\n\n# Final Model Performance:\n| Leaderboard | Share(Mask NonShare) | NonShare(Mask share) |\n| --- | --- | --- |\n| Public | 0.347 | 0.173 |\n| Private | 0.224 | 0.059 |\n\n# Tried but not go well\n- Stacking approach for share and non-share (A model was constructed for each protein)\n  - Models: 1DCNN, LGBM2, ChemBERTa for share, and ChemBERTa, 1DCNN, LGBM1, and LGBM2 for non-share (LGBM1 used morgan4+rdkit+descriptors and all data for training, LGBM2 used ECFP6 and 10M data for training, others used 10M data for training)\n- Procedure:\n1. By using each model, the prediction of each fold of cross-validation was carried out, and it become x. The bind values from train data the competition organizer gave became y. Catboost was used as stacking model because Catboost is good approach for category data such as bb. (but in fact, we didnot use bb as input...)\n2. To predict binds of submission, bind values of test data predicted by each model were used as x. \n3. Finally, predictions of share and non-share were merged. \n\n- Share  LB:0.506→0.496,  PB 0.261 → 0.253\n  - Code：https://www.kaggle.com/code/kokishinbara/catboost-share-protein-loop/output\n  - Result: https://www.kaggle.com/datasets/kokishinbara/stacking-share\n\n- Nonshare LB0.506 → 0.497,   PB 0.261 → 0.263 \n  - Code：https://github.com/kyu999/leash-bio/blob/kyue/notebooks/stacking/nonshare_stacking.ipynb\n  - Result：https://www.kaggle.com/datasets/kyu999/stacking-nonshare-multi-models\n\n- We tried some graph-based approatch, GCN, 3DGNN (Sphere), but did not work for us. \n- We also tried adjusting embedding, by enumeration of the SMILES string and it did not work. \n\n## Others...\n- [Based on Scaffold explonation](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook), [Not triazine core was searched](https://www.kaggle.com/code/kokishinbara/scaffold-exploration-num) and, [the train data without cores are trained but not go well](https://www.kaggle.com/code/kokishinbara/without-core-leash-bio-chemberta-baseline). (model = ChemBERTa, LB=0.281 PB=0.161) (Please ignore the comments in our notebook, just copied the original and refered to the code.)\n- [Catboost for stacking also used one model for proteins, but didn't go well](https://www.kaggle.com/code/shunsukekikuchi/catboost-share-one-model/notebook) (CV of 1 fold BRD4: 0.326  HSA: 0.138  sEH: 0.680 )\n- Experiments with Small-Scale Samples\n  - Feature Addition: Added chemical properties as features to the LGBM model, alongside molecular fingerprints.\n  - Fingerprint Comparison: Compared the performance of multiple types of molecular fingerprints.\n  - ECFP Radius Optimization: Investigated and determined the optimal radius for ECFP.\n  - Generalization Performance: Validated the use of sEH external data to enhance the model's generalization performance.\n  - Data Filtering Impact: Examined the impact on model precision by filtering sEH external data using read count.\n  - Embedding Replacement: Replaced molecular fingerprints with Malformer embeddings to evaluate performance changes.\n  - Data Augmentation: Evaluated the effectiveness of augmented data generated using SMILES enumeration.\n  - Model Comparison: Compared the performance of the original model, ChemBERTa (10M smiles), with the updated model, ChemBERTa2 (77M smiles).\n\n# Thanks\nWe would like to thank Kaggle for organizing such an interesting competition. We also appreciate @ahmedelfazouan and @tetsuya3510 for sharing notebooks that influenced our solution. A big thanks to my teammates @kyu999, @shunsukekikuchi, and @kokishinbara for their dedication, collaboration, and tireless efforts. Without their hard work and teamwork, our achievements wouldn't have been possible.",
      "votes": 21
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2914642": "We adopted different strategies for shared and nonshared targets.\n\n# 1. Nonshared Target Part\n\n## CV Strategy\nThe split was made based on building blocks, as we saw in this helpful [notebook](https://www.kaggle.com/code/thedrcat/belka-split-cv-like-the-host). Each building block was sorted and then systematically sampled to avoid unbalanced sampling.\n\n## Model\nWe used [ChemBERTa-10M-MTR](https://www.kaggle.com/code/tetsuya3510/leash-bio-chemberta-baseline) as our base model for each protein\n## Training\n### Fine-tuning: \n- For all proteins: We used 4M sampling for fine-tuning each model.\n- We fine-tuned each model using 4M samples for the respective protein.\n- For sEH Triazine core: We used a mixed dataset, with half the data from external sources and half from the regular training data.\n- For sEH NonTriazine core: We used a dataset trained exclusively on [external data containing various cores](https://chemrxiv.org/engage/chemrxiv/article-details/6438943f08c86922ffeffe57). This approach was chosen to potentially improve predictions for unknown cores.\n- Overfitting Prevention: To mitigate overfitting, we froze the weights of the first three layers of the ChemBERTa model.\n\n## Ensemble Method\n- For most proteins: We tested three ensemble methods - simple average, CV-weighted average, and CV-weighted average to the power of 5. We chose the CV-weighted average to the power of 5 as it yielded the highest LB score.\n| Ensemble Method | Public | Private |\n|---|---|---|\n| Simple Average | 0.486 | 0.272 |\n| CV-weighted Average | 0.488 | 0.271 |\n| CV-weighted Average^5 | 0.496 | 0.266 |\n\n- For sEH NonTriazine core: We used a simple average ensemble of the three runs with different seeds.\n\nPerformance: The final Public LB score (mask share) for our nonshared target models was 0.173.\n\n# 2. Shared Target Part\nFor shared targets, we used a 1DCNN model, improving upon the Kaggle [public notebook](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data).\n\n## CV Strategy\nWe used 5-fold split generated by StratifiedKFold. All training data was used. \n## Model\n- Increased number of filters (32 to 64, doubling in each layer)\n- Enhanced convolution layers (padding='same', varied kernel sizes)\n- Added BatchNormalization\n- Introduced shortcut connections and Attention mechanism\n- Adjusted Dropout rate (0.1 to 0.2)\n- Modified fully connected layers\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17343385%2F4ba074c5c07030d771f907f59c223b62%2F1DCNN_Model.png?generation=1720586518917501&alt=media)\n\n## Training\n- Increased number of epochs (20 to 30)\n- Reduced patience in ReduceLROnPlateau (5 to 3)\n- Allowed slight overfitting\n\nThese improvements aimed to intentionally overfit the model slightly, enabling it to learn shared target data more deeply and improve scores. This strategy allowed the model to capture more detailed features of the training data, leading to improved predictive performance.\n\n## Score\nWe compared the LB scores of fold0 out of 15 folds to evaluate our model improvements. The results are shown in the table below:\n| Model | Public | Private |\n| --- | --- |\n| Original notebook model | 0.359 | 0.217 |\n| Improved model | 0.387 | 0.220 |\n\n\nThese results demonstrate the effectiveness of our improvement strategy. By increasing model complexity and allowing deeper learning, we significantly improved our score on the public leaderboard.\nAfter confirming the improvement through this single fold comparison, we proceeded with our final strategy:\n- We created a new model with 5 folds, separate from the 15-fold model used for comparison.\n- We applied our improved model architecture to all 5 folds.\n- Finally, we created an average ensemble using these 5 folds to generate our final predictions.\n\nThe final performance of our 5-fold ensemble 1DCNN model on the Public LB (mask nonshare) was 0.347.\nThis approach allowed us to leverage the benefits of our model improvements while also utilizing ensemble techniques to further enhance our predictions.\n\n# Final Model Performance:\n| Leaderboard | Share(Mask NonShare) | NonShare(Mask share) |\n| --- | --- | --- |\n| Public | 0.347 | 0.173 |\n| Private | 0.224 | 0.059 |\n\n# Tried but not go well\n- Stacking approach for share and non-share (A model was constructed for each protein)\n  - Models: 1DCNN, LGBM2, ChemBERTa for share, and ChemBERTa, 1DCNN, LGBM1, and LGBM2 for non-share (LGBM1 used morgan4+rdkit+descriptors and all data for training, LGBM2 used ECFP6 and 10M data for training, others used 10M data for training)\n- Procedure:\n1. By using each model, the prediction of each fold of cross-validation was carried out, and it become x. The bind values from train data the competition organizer gave became y. Catboost was used as stacking model because Catboost is good approach for category data such as bb. (but in fact, we didnot use bb as input...)\n2. To predict binds of submission, bind values of test data predicted by each model were used as x. \n3. Finally, predictions of share and non-share were merged. \n\n- Share  LB:0.506→0.496,  PB 0.261 → 0.253\n  - Code：https://www.kaggle.com/code/kokishinbara/catboost-share-protein-loop/output\n  - Result: https://www.kaggle.com/datasets/kokishinbara/stacking-share\n\n- Nonshare LB0.506 → 0.497,   PB 0.261 → 0.263 \n  - Code：https://github.com/kyu999/leash-bio/blob/kyue/notebooks/stacking/nonshare_stacking.ipynb\n  - Result：https://www.kaggle.com/datasets/kyu999/stacking-nonshare-multi-models\n\n- We tried some graph-based approatch, GCN, 3DGNN (Sphere), but did not work for us. \n- We also tried adjusting embedding, by enumeration of the SMILES string and it did not work. \n\n## Others...\n- [Based on Scaffold explonation](https://www.kaggle.com/code/chemdatafarmer/scaffold-exploration/notebook), [Not triazine core was searched](https://www.kaggle.com/code/kokishinbara/scaffold-exploration-num) and, [the train data without cores are trained but not go well](https://www.kaggle.com/code/kokishinbara/without-core-leash-bio-chemberta-baseline). (model = ChemBERTa, LB=0.281 PB=0.161) (Please ignore the comments in our notebook, just copied the original and refered to the code.)\n- [Catboost for stacking also used one model for proteins, but didn't go well](https://www.kaggle.com/code/shunsukekikuchi/catboost-share-one-model/notebook) (CV of 1 fold BRD4: 0.326  HSA: 0.138  sEH: 0.680 )\n- Experiments with Small-Scale Samples\n  - Feature Addition: Added chemical properties as features to the LGBM model, alongside molecular fingerprints.\n  - Fingerprint Comparison: Compared the performance of multiple types of molecular fingerprints.\n  - ECFP Radius Optimization: Investigated and determined the optimal radius for ECFP.\n  - Generalization Performance: Validated the use of sEH external data to enhance the model's generalization performance.\n  - Data Filtering Impact: Examined the impact on model precision by filtering sEH external data using read count.\n  - Embedding Replacement: Replaced molecular fingerprints with Malformer embeddings to evaluate performance changes.\n  - Data Augmentation: Evaluated the effectiveness of augmented data generated using SMILES enumeration.\n  - Model Comparison: Compared the performance of the original model, ChemBERTa (10M smiles), with the updated model, ChemBERTa2 (77M smiles).\n\n# Thanks\nWe would like to thank Kaggle for organizing such an interesting competition. We also appreciate @ahmedelfazouan and @tetsuya3510 for sharing notebooks that influenced our solution. A big thanks to my teammates @kyu999, @shunsukekikuchi, and @kokishinbara for their dedication, collaboration, and tireless efforts. Without their hard work and teamwork, our achievements wouldn't have been possible."
  }
}