{
  "id": 518951,
  "title": "14th Place Solution (7th in Public): All We Need is Frequent Checkpointing",
  "url": "/competitions/leash-BELKA/writeups/gorna-14th-place-solution-7th-in-public-all-we-nee",
  "author_name": "",
  "post_date": "2024-07-09T03:47:00.360Z",
  "votes": 61,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Congratulations to all the winners and thanks to the organizers for hosting this interesting competition.<br>\nAs we all know, shared targets and nonshared targets have completely different tendencies, so we built different pipelines.</p>\n<h1>Nonshared Target Part (<a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">yu4u</a>)</h1>\n<p>The most important observation in the nonshared target part is that the validation score of the nonshared target overfits very quickly, even within just one epoch of training.<br>\nTherefore, we changed the validation interval to 0.01 epoch, and as a result, successfully obtained a checkpoint that does not overfit.</p>\n<h2>CV Strategy</h2>\n<p>CV strategy to simulate the nonshared targets is also important.<br>\nWe used a 5-fold CV strategy, where BB1, BB2, and BB3 (with the BB2 building blocks removed) building blocks were split into 5 folds. Then we removed the other data so that the training data contains only the training building blocks and the validation data contains only the validation building blocks.<br>\nWe released the notebook that generates the CV folds <a href=\"https://www.kaggle.com/code/ren4yu/leash-split-for-noshare/\" target=\"_blank\">here</a>.</p>\n<h2>Model</h2>\n<p><a href=\"https://huggingface.co/DeepChem/ChemBERTa-77M-MTR\" target=\"_blank\">ChemBERTa-77M-MTR</a></p>\n<h2>Training</h2>\n<ul>\n<li>5-fold CV.</li>\n<li>AdamW optimizer with LR=1e-3 (fixed), weight_decay=1e-5, batch_size=512.</li>\n<li>Train for one epoch and validate at each 0.01 epoch interval.</li>\n</ul>\n<h2>Selecting the Best Checkpoints</h2>\n<ul>\n<li>We found some folds have different TP distributions from whole training data, so we used only fold0 and fold2 that have similar TP distributions to the whole training data.</li>\n<li>In our training, we do not complete even one epoch of learning, so there is data that is not used in a single training run. By training with multiple seeds (11 seeds for each fold), we utilize all the data.</li>\n<li>Finally, we excluded seeds with extremely low CV scores and used an ensemble of 19 models for the final results.</li>\n<li>In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.</li>\n</ul>\n<h1>Shared Target Part (<a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">monnu</a>)</h1>\n<p>The main strategy was to improve the <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public 1DCNN model</a> by concatenating features from ECFP.<br>\nAs an option, we also trained models with additional features from ChemBERTa.</p>\n<h2>CV Strategy</h2>\n<ul>\n<li>We used 5-fold Stratified Kfold splits.</li>\n<li>We released the notebook that generates the CV folds for share <a href=\"https://www.kaggle.com/code/fuumin621/leash-split-for-share/notebook\" target=\"_blank\">here</a>.</li>\n</ul>\n<h2>Preprocess</h2>\n<ul>\n<li>ECFP: Used rdkit. r = 4, bit = 2048 or 3072</li>\n<li>1DCNN: Encoded SMILES strings into numerical values and converted them into fixed-length vectors</li>\n<li>chemberta_feature (optional): Used the ChemBERTa model to infer the SMILES strings and used the 384-dimensional output as features.</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>SMILES are passed through an embedding layer, followed by 4 layers of 1D convolution</li>\n<li>ECFP is passed through an FC layer to 128 dim</li>\n<li>The outputs above are concatenated and passed through an FC layer to output binary classification scores for 3 targets</li>\n<li>Optionally, the 384-dimensional output of ChemBERTa can be passed through an FC layer and then concatenated</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>5-fold CV.</li>\n<li>AdamW optimizer with LR=1e-3, weight_decay=0.05, batch_size=4096.</li>\n<li>num_epochs=25</li>\n</ul>\n<h2>Score</h2>\n<p>The scores of the trained models are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Model Name</th>\n<th>ECFP</th>\n<th>chemberta_feature</th>\n<th>CV</th>\n<th>PublicLB(mask noshare)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp031</td>\n<td>r=4, bit=2048</td>\n<td>No</td>\n<td>0.6589</td>\n<td>0.352</td>\n</tr>\n<tr>\n<td>exp032</td>\n<td>r=4, bit=2048</td>\n<td>Yes</td>\n<td>0.65934</td>\n<td>0.351</td>\n</tr>\n<tr>\n<td>exp039</td>\n<td>r=4, bit=3072</td>\n<td>No</td>\n<td>0.66001</td>\n<td>0.350</td>\n</tr>\n<tr>\n<td>average ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.352</td>\n</tr>\n</tbody>\n</table>\n<p>In the end, we submitted the average ensemble.</p>",
  "messages": [
    {
      "id": "2912656",
      "postDate": "07/09/2024 02:31:12",
      "content": "<p>Congratulations to all the winners and thanks to the organizers for hosting this interesting competition.<br>\nAs we all know, shared targets and nonshared targets have completely different tendencies, so we built different pipelines.</p>\n<h1>Nonshared Target Part (<a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">yu4u</a>)</h1>\n<p>The most important observation in the nonshared target part is that the validation score of the nonshared target overfits very quickly, even within just one epoch of training.<br>\nTherefore, we changed the validation interval to 0.01 epoch, and as a result, successfully obtained a checkpoint that does not overfit.</p>\n<h2>CV Strategy</h2>\n<p>CV strategy to simulate the nonshared targets is also important.<br>\nWe used a 5-fold CV strategy, where BB1, BB2, and BB3 (with the BB2 building blocks removed) building blocks were split into 5 folds. Then we removed the other data so that the training data contains only the training building blocks and the validation data contains only the validation building blocks.<br>\nWe released the notebook that generates the CV folds <a href=\"https://www.kaggle.com/code/ren4yu/leash-split-for-noshare/\" target=\"_blank\">here</a>.</p>\n<h2>Model</h2>\n<p><a href=\"https://huggingface.co/DeepChem/ChemBERTa-77M-MTR\" target=\"_blank\">ChemBERTa-77M-MTR</a></p>\n<h2>Training</h2>\n<ul>\n<li>5-fold CV.</li>\n<li>AdamW optimizer with LR=1e-3 (fixed), weight_decay=1e-5, batch_size=512.</li>\n<li>Train for one epoch and validate at each 0.01 epoch interval.</li>\n</ul>\n<h2>Selecting the Best Checkpoints</h2>\n<ul>\n<li>We found some folds have different TP distributions from whole training data, so we used only fold0 and fold2 that have similar TP distributions to the whole training data.</li>\n<li>In our training, we do not complete even one epoch of learning, so there is data that is not used in a single training run. By training with multiple seeds (11 seeds for each fold), we utilize all the data.</li>\n<li>Finally, we excluded seeds with extremely low CV scores and used an ensemble of 19 models for the final results.</li>\n<li>In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.</li>\n</ul>\n<h1>Shared Target Part (<a href=\"https://www.kaggle.com/fuumin621\" target=\"_blank\">monnu</a>)</h1>\n<p>The main strategy was to improve the <a href=\"https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data\" target=\"_blank\">public 1DCNN model</a> by concatenating features from ECFP.<br>\nAs an option, we also trained models with additional features from ChemBERTa.</p>\n<h2>CV Strategy</h2>\n<ul>\n<li>We used 5-fold Stratified Kfold splits.</li>\n<li>We released the notebook that generates the CV folds for share <a href=\"https://www.kaggle.com/code/fuumin621/leash-split-for-share/notebook\" target=\"_blank\">here</a>.</li>\n</ul>\n<h2>Preprocess</h2>\n<ul>\n<li>ECFP: Used rdkit. r = 4, bit = 2048 or 3072</li>\n<li>1DCNN: Encoded SMILES strings into numerical values and converted them into fixed-length vectors</li>\n<li>chemberta_feature (optional): Used the ChemBERTa model to infer the SMILES strings and used the 384-dimensional output as features.</li>\n</ul>\n<h2>Model</h2>\n<ul>\n<li>SMILES are passed through an embedding layer, followed by 4 layers of 1D convolution</li>\n<li>ECFP is passed through an FC layer to 128 dim</li>\n<li>The outputs above are concatenated and passed through an FC layer to output binary classification scores for 3 targets</li>\n<li>Optionally, the 384-dimensional output of ChemBERTa can be passed through an FC layer and then concatenated</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>5-fold CV.</li>\n<li>AdamW optimizer with LR=1e-3, weight_decay=0.05, batch_size=4096.</li>\n<li>num_epochs=25</li>\n</ul>\n<h2>Score</h2>\n<p>The scores of the trained models are as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Model Name</th>\n<th>ECFP</th>\n<th>chemberta_feature</th>\n<th>CV</th>\n<th>PublicLB(mask noshare)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>exp031</td>\n<td>r=4, bit=2048</td>\n<td>No</td>\n<td>0.6589</td>\n<td>0.352</td>\n</tr>\n<tr>\n<td>exp032</td>\n<td>r=4, bit=2048</td>\n<td>Yes</td>\n<td>0.65934</td>\n<td>0.351</td>\n</tr>\n<tr>\n<td>exp039</td>\n<td>r=4, bit=3072</td>\n<td>No</td>\n<td>0.66001</td>\n<td>0.350</td>\n</tr>\n<tr>\n<td>average ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.352</td>\n</tr>\n</tbody>\n</table>\n<p>In the end, we submitted the average ensemble.</p>",
      "rawMarkdown": "Congratulations to all the winners and thanks to the organizers for hosting this interesting competition.\nAs we all know, shared targets and nonshared targets have completely different tendencies, so we built different pipelines.\n\n\n# Nonshared Target Part ([yu4u](https://www.kaggle.com/ren4yu))\nThe most important observation in the nonshared target part is that the validation score of the nonshared target overfits very quickly, even within just one epoch of training.\nTherefore, we changed the validation interval to 0.01 epoch, and as a result, successfully obtained a checkpoint that does not overfit.\n\n## CV Strategy\nCV strategy to simulate the nonshared targets is also important.\nWe used a 5-fold CV strategy, where BB1, BB2, and BB3 (with the BB2 building blocks removed) building blocks were split into 5 folds. Then we removed the other data so that the training data contains only the training building blocks and the validation data contains only the validation building blocks.\nWe released the notebook that generates the CV folds [here](https://www.kaggle.com/code/ren4yu/leash-split-for-noshare/).\n\n## Model\n[ChemBERTa-77M-MTR](https://huggingface.co/DeepChem/ChemBERTa-77M-MTR)\n\n## Training\n- 5-fold CV.\n- AdamW optimizer with LR=1e-3 (fixed), weight_decay=1e-5, batch_size=512.\n- Train for one epoch and validate at each 0.01 epoch interval.\n\n## Selecting the Best Checkpoints\n- We found some folds have different TP distributions from whole training data, so we used only fold0 and fold2 that have similar TP distributions to the whole training data.\n- In our training, we do not complete even one epoch of learning, so there is data that is not used in a single training run. By training with multiple seeds (11 seeds for each fold), we utilize all the data.\n- Finally, we excluded seeds with extremely low CV scores and used an ensemble of 19 models for the final results.\n- In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.\n\n\n# Shared Target Part ([monnu](https://www.kaggle.com/fuumin621))\nThe main strategy was to improve the [public 1DCNN model](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) by concatenating features from ECFP.\nAs an option, we also trained models with additional features from ChemBERTa.\n\n## CV Strategy\n- We used 5-fold Stratified Kfold splits.\n- We released the notebook that generates the CV folds for share [here](https://www.kaggle.com/code/fuumin621/leash-split-for-share/notebook).\n\n## Preprocess\n- ECFP: Used rdkit. r = 4, bit = 2048 or 3072\n- 1DCNN: Encoded SMILES strings into numerical values and converted them into fixed-length vectors\n- chemberta_feature (optional): Used the ChemBERTa model to infer the SMILES strings and used the 384-dimensional output as features.\n\n## Model\n- SMILES are passed through an embedding layer, followed by 4 layers of 1D convolution\n- ECFP is passed through an FC layer to 128 dim\n- The outputs above are concatenated and passed through an FC layer to output binary classification scores for 3 targets\n- Optionally, the 384-dimensional output of ChemBERTa can be passed through an FC layer and then concatenated\n\n## Training\n- 5-fold CV.\n- AdamW optimizer with LR=1e-3, weight_decay=0.05, batch_size=4096.\n- num_epochs=25\n\n## Score\nThe scores of the trained models are as follows:\n| Model Name | ECFP | chemberta_feature | CV | PublicLB(mask noshare)|\n|------------|------|-------------------|----------|----------|\n| exp031     | r=4, bit=2048 | No  | 0.6589 | 0.352 |\n| exp032     | r=4, bit=2048 | Yes | 0.65934 | 0.351 |\n| exp039     | r=4, bit=3072 | No  | 0.66001 | 0.350 |\n| average ensemble     | - | - | - | 0.352 |\n\nIn the end, we submitted the average ensemble.",
      "votes": null
    },
    {
      "id": "2912673",
      "postDate": "07/09/2024 02:57:55",
      "content": "<p>Nice! Thanks for your solution</p>",
      "rawMarkdown": "Nice! Thanks for your solution",
      "votes": null
    },
    {
      "id": "2912674",
      "postDate": "07/09/2024 03:00:03",
      "content": "<p>Fascinating</p>\n<p>I have often noticed that on many proteins my first epoch is best, the idea of using a smaller validation interval is brilliant, and obvious in hindsight. Do you have any idea on why chemberta overfits so fast and other ways we could improve that? Is it just a diversity of training set issue?</p>",
      "rawMarkdown": "Fascinating\n\nI have often noticed that on many proteins my first epoch is best, the idea of using a smaller validation interval is brilliant, and obvious in hindsight. Do you have any idea on why chemberta overfits so fast and other ways we could improve that? Is it just a diversity of training set issue?",
      "votes": null
    },
    {
      "id": "2912682",
      "postDate": "07/09/2024 03:06:37",
      "content": "<p>maybe training scaffold is unique?</p>",
      "rawMarkdown": "maybe training scaffold is unique?",
      "votes": null
    },
    {
      "id": "2912709",
      "postDate": "07/09/2024 03:38:45",
      "content": "<p>Thanks. I got good informations.</p>",
      "rawMarkdown": "Thanks. I got good informations.",
      "votes": null
    },
    {
      "id": "2912735",
      "postDate": "07/09/2024 03:51:09",
      "content": "<p>I think the knowledge learned by current AI models from data cannot be generalized to the unseen protein/molecules. That could be the reason why the crazy shuffle always happened in the molecule competition! <br>\nActually, this is the third time I met the crazy shuffle in private LB (I have participated in 3 molecule competition, the shuffle happened at each time). <br>\nFor other CV/NLP/Recommendation competitions, I never meet such problems. </p>",
      "rawMarkdown": "I think the knowledge learned by current AI models from data cannot be generalized to the unseen protein/molecules. That could be the reason why the crazy shuffle always happened in the molecule competition! \nActually, this is the third time I met the crazy shuffle in private LB (I have participated in 3 molecule competition, the shuffle happened at each time). \nFor other CV/NLP/Recommendation competitions, I never meet such problems.",
      "votes": null
    },
    {
      "id": "2912744",
      "postDate": "07/09/2024 03:59:05",
      "content": "<p>I think the problem is the diversity of the training data. The training data looks big, but there are very few unique building blocks. So, there is not much diversity, and it might be easy to overfit on the noshare data.<br>\nTo avoid overfitting, the following also helped a bit:</p>\n<ul>\n<li>Stronger weight decay</li>\n<li>Freeze ChemBERTa for one epoch, then unfreeze it.</li>\n</ul>",
      "rawMarkdown": "I think the problem is the diversity of the training data. The training data looks big, but there are very few unique building blocks. So, there is not much diversity, and it might be easy to overfit on the noshare data.\nTo avoid overfitting, the following also helped a bit:\n- Stronger weight decay\n- Freeze ChemBERTa for one epoch, then unfreeze it.",
      "votes": null
    },
    {
      "id": "2912755",
      "postDate": "07/09/2024 04:08:38",
      "content": "<p>Congratulations on securing 14th place in this competition. Thanks for sharing your solution details. </p>",
      "rawMarkdown": "Congratulations on securing 14th place in this competition. Thanks for sharing your solution details.",
      "votes": null
    },
    {
      "id": "2912777",
      "postDate": "07/09/2024 04:17:26",
      "content": "<p>Our solution bore a striking resemblance to yours, but we completely overlooked the ingenious idea of using such a small validation interval. <br>\nWe're truly impressed by this innovative approach. <br>\nThank you sincerely for sharing these valuable insights.</p>",
      "rawMarkdown": "Our solution bore a striking resemblance to yours, but we completely overlooked the ingenious idea of using such a small validation interval. \nWe're truly impressed by this innovative approach. \nThank you sincerely for sharing these valuable insights.",
      "votes": null
    },
    {
      "id": "2913026",
      "postDate": "07/09/2024 07:28:49",
      "content": "<p>Nice, we also did checkpoints and scoring 10 times per epoch and found that first epoch is almost the best for predicting noshared part.</p>",
      "rawMarkdown": "Nice, we also did checkpoints and scoring 10 times per epoch and found that first epoch is almost the best for predicting noshared part.",
      "votes": null
    },
    {
      "id": "2913117",
      "postDate": "07/09/2024 09:05:14",
      "content": "<p>I forgot to mention in my write-up that we also used Exponential Moving Average (EMA) to avoid overfitting.</p>",
      "rawMarkdown": "I forgot to mention in my write-up that we also used Exponential Moving Average (EMA) to avoid overfitting.",
      "votes": null
    },
    {
      "id": "2914110",
      "postDate": "07/09/2024 18:22:33",
      "content": "<p>Nice! good job, thanks for sharing!</p>",
      "rawMarkdown": "Nice! good job, thanks for sharing!",
      "votes": null
    },
    {
      "id": "2914280",
      "postDate": "07/09/2024 20:09:38",
      "content": "<blockquote>\n  <p>In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.</p>\n</blockquote>\n<p>For 2, did you use average predictions of the three checkpoints? Or use each individual checkpoint for predicting that particular target?</p>",
      "rawMarkdown": "> In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.\n\nFor 2, did you use average predictions of the three checkpoints? Or use each individual checkpoint for predicting that particular target?",
      "votes": null
    },
    {
      "id": "2914283",
      "postDate": "07/09/2024 20:14:43",
      "content": "<p>I suspect the pretraining that the winner did could also help? Predict masked smiles to smiles, and/or predict ecfp from smiles. Predict chemically significant features. Why? With this train data, the easy path for the model is to learn \"which building block am I?\" which is exactly what you do NOT want. </p>\n<p>Target encoding alone was too effectives on shared data, so it's definitely a modeling concern in this dataset and this domain. </p>",
      "rawMarkdown": "I suspect the pretraining that the winner did could also help? Predict masked smiles to smiles, and/or predict ecfp from smiles. Predict chemically significant features. Why? With this train data, the easy path for the model is to learn \"which building block am I?\" which is exactly what you do NOT want. \n\nTarget encoding alone was too effectives on shared data, so it's definitely a modeling concern in this dataset and this domain.",
      "votes": null
    },
    {
      "id": "2914458",
      "postDate": "07/10/2024 00:15:32",
      "content": "<p>We used each individual checkpoint for predicting that particular target.</p>",
      "rawMarkdown": "We used each individual checkpoint for predicting that particular target.",
      "votes": null
    },
    {
      "id": "2918137",
      "postDate": "07/12/2024 03:47:44",
      "content": "<p>I have observed in my solution too. The first epoch performs the best. Isn't it amazing?</p>",
      "rawMarkdown": "I have observed in my solution too. The first epoch performs the best. Isn't it amazing?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2912673,
      "author_name": "horikitasaku",
      "author_url": "",
      "post_date": "07/09/2024 02:57:55",
      "content": "<p>Nice! Thanks for your solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2912674,
      "author_name": "andrewdblevins",
      "author_url": "",
      "post_date": "07/09/2024 03:00:03",
      "content": "<p>Fascinating</p>\n<p>I have often noticed that on many proteins my first epoch is best, the idea of using a smaller validation interval is brilliant, and obvious in hindsight. Do you have any idea on why chemberta overfits so fast and other ways we could improve that? Is it just a diversity of training set issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2912682,
          "author_name": "wenxuanxx",
          "author_url": "",
          "post_date": "07/09/2024 03:06:37",
          "content": "<p>maybe training scaffold is unique?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2912744,
          "author_name": "fuumin621",
          "author_url": "",
          "post_date": "07/09/2024 03:59:05",
          "content": "<p>I think the problem is the diversity of the training data. The training data looks big, but there are very few unique building blocks. So, there is not much diversity, and it might be easy to overfit on the noshare data.<br>\nTo avoid overfitting, the following also helped a bit:</p>\n<ul>\n<li>Stronger weight decay</li>\n<li>Freeze ChemBERTa for one epoch, then unfreeze it.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2913117,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "07/09/2024 09:05:14",
          "content": "<p>I forgot to mention in my write-up that we also used Exponential Moving Average (EMA) to avoid overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2914283,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/09/2024 20:14:43",
          "content": "<p>I suspect the pretraining that the winner did could also help? Predict masked smiles to smiles, and/or predict ecfp from smiles. Predict chemically significant features. Why? With this train data, the easy path for the model is to learn \"which building block am I?\" which is exactly what you do NOT want. </p>\n<p>Target encoding alone was too effectives on shared data, so it's definitely a modeling concern in this dataset and this domain. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2912709,
      "author_name": "junhanzangai",
      "author_url": "",
      "post_date": "07/09/2024 03:38:45",
      "content": "<p>Thanks. I got good informations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2912735,
      "author_name": "wudejian789",
      "author_url": "",
      "post_date": "07/09/2024 03:51:09",
      "content": "<p>I think the knowledge learned by current AI models from data cannot be generalized to the unseen protein/molecules. That could be the reason why the crazy shuffle always happened in the molecule competition! <br>\nActually, this is the third time I met the crazy shuffle in private LB (I have participated in 3 molecule competition, the shuffle happened at each time). <br>\nFor other CV/NLP/Recommendation competitions, I never meet such problems. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2912755,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/09/2024 04:08:38",
      "content": "<p>Congratulations on securing 14th place in this competition. Thanks for sharing your solution details. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2912777,
      "author_name": "shinaggle",
      "author_url": "",
      "post_date": "07/09/2024 04:17:26",
      "content": "<p>Our solution bore a striking resemblance to yours, but we completely overlooked the ingenious idea of using such a small validation interval. <br>\nWe're truly impressed by this innovative approach. <br>\nThank you sincerely for sharing these valuable insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2913026,
      "author_name": "ogurtsov",
      "author_url": "",
      "post_date": "07/09/2024 07:28:49",
      "content": "<p>Nice, we also did checkpoints and scoring 10 times per epoch and found that first epoch is almost the best for predicting noshared part.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2914110,
      "author_name": "gnomows",
      "author_url": "",
      "post_date": "07/09/2024 18:22:33",
      "content": "<p>Nice! good job, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2914280,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/09/2024 20:09:38",
      "content": "<blockquote>\n  <p>In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.</p>\n</blockquote>\n<p>For 2, did you use average predictions of the three checkpoints? Or use each individual checkpoint for predicting that particular target?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2914458,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "07/10/2024 00:15:32",
          "content": "<p>We used each individual checkpoint for predicting that particular target.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2918137,
      "author_name": "sajkazmi",
      "author_url": "",
      "post_date": "07/12/2024 03:47:44",
      "content": "<p>I have observed in my solution too. The first epoch performs the best. Isn't it amazing?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2912656": "Congratulations to all the winners and thanks to the organizers for hosting this interesting competition.\nAs we all know, shared targets and nonshared targets have completely different tendencies, so we built different pipelines.\n\n\n# Nonshared Target Part ([yu4u](https://www.kaggle.com/ren4yu))\nThe most important observation in the nonshared target part is that the validation score of the nonshared target overfits very quickly, even within just one epoch of training.\nTherefore, we changed the validation interval to 0.01 epoch, and as a result, successfully obtained a checkpoint that does not overfit.\n\n## CV Strategy\nCV strategy to simulate the nonshared targets is also important.\nWe used a 5-fold CV strategy, where BB1, BB2, and BB3 (with the BB2 building blocks removed) building blocks were split into 5 folds. Then we removed the other data so that the training data contains only the training building blocks and the validation data contains only the validation building blocks.\nWe released the notebook that generates the CV folds [here](https://www.kaggle.com/code/ren4yu/leash-split-for-noshare/).\n\n## Model\n[ChemBERTa-77M-MTR](https://huggingface.co/DeepChem/ChemBERTa-77M-MTR)\n\n## Training\n- 5-fold CV.\n- AdamW optimizer with LR=1e-3 (fixed), weight_decay=1e-5, batch_size=512.\n- Train for one epoch and validate at each 0.01 epoch interval.\n\n## Selecting the Best Checkpoints\n- We found some folds have different TP distributions from whole training data, so we used only fold0 and fold2 that have similar TP distributions to the whole training data.\n- In our training, we do not complete even one epoch of learning, so there is data that is not used in a single training run. By training with multiple seeds (11 seeds for each fold), we utilize all the data.\n- Finally, we excluded seeds with extremely low CV scores and used an ensemble of 19 models for the final results.\n- In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.\n\n\n# Shared Target Part ([monnu](https://www.kaggle.com/fuumin621))\nThe main strategy was to improve the [public 1DCNN model](https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data) by concatenating features from ECFP.\nAs an option, we also trained models with additional features from ChemBERTa.\n\n## CV Strategy\n- We used 5-fold Stratified Kfold splits.\n- We released the notebook that generates the CV folds for share [here](https://www.kaggle.com/code/fuumin621/leash-split-for-share/notebook).\n\n## Preprocess\n- ECFP: Used rdkit. r = 4, bit = 2048 or 3072\n- 1DCNN: Encoded SMILES strings into numerical values and converted them into fixed-length vectors\n- chemberta_feature (optional): Used the ChemBERTa model to infer the SMILES strings and used the 384-dimensional output as features.\n\n## Model\n- SMILES are passed through an embedding layer, followed by 4 layers of 1D convolution\n- ECFP is passed through an FC layer to 128 dim\n- The outputs above are concatenated and passed through an FC layer to output binary classification scores for 3 targets\n- Optionally, the 384-dimensional output of ChemBERTa can be passed through an FC layer and then concatenated\n\n## Training\n- 5-fold CV.\n- AdamW optimizer with LR=1e-3, weight_decay=0.05, batch_size=4096.\n- num_epochs=25\n\n## Score\nThe scores of the trained models are as follows:\n| Model Name | ECFP | chemberta_feature | CV | PublicLB(mask noshare)|\n|------------|------|-------------------|----------|----------|\n| exp031     | r=4, bit=2048 | No  | 0.6589 | 0.352 |\n| exp032     | r=4, bit=2048 | Yes | 0.65934 | 0.351 |\n| exp039     | r=4, bit=3072 | No  | 0.66001 | 0.350 |\n| average ensemble     | - | - | - | 0.352 |\n\nIn the end, we submitted the average ensemble.",
    "2912673": "Nice! Thanks for your solution",
    "2912674": "Fascinating\n\nI have often noticed that on many proteins my first epoch is best, the idea of using a smaller validation interval is brilliant, and obvious in hindsight. Do you have any idea on why chemberta overfits so fast and other ways we could improve that? Is it just a diversity of training set issue?",
    "2912682": "maybe training scaffold is unique?",
    "2912709": "Thanks. I got good informations.",
    "2912735": "I think the knowledge learned by current AI models from data cannot be generalized to the unseen protein/molecules. That could be the reason why the crazy shuffle always happened in the molecule competition! \nActually, this is the third time I met the crazy shuffle in private LB (I have participated in 3 molecule competition, the shuffle happened at each time). \nFor other CV/NLP/Recommendation competitions, I never meet such problems.",
    "2912744": "I think the problem is the diversity of the training data. The training data looks big, but there are very few unique building blocks. So, there is not much diversity, and it might be easy to overfit on the noshare data.\nTo avoid overfitting, the following also helped a bit:\n- Stronger weight decay\n- Freeze ChemBERTa for one epoch, then unfreeze it.",
    "2912755": "Congratulations on securing 14th place in this competition. Thanks for sharing your solution details.",
    "2912777": "Our solution bore a striking resemblance to yours, but we completely overlooked the ingenious idea of using such a small validation interval. \nWe're truly impressed by this innovative approach. \nThank you sincerely for sharing these valuable insights.",
    "2913026": "Nice, we also did checkpoints and scoring 10 times per epoch and found that first epoch is almost the best for predicting noshared part.",
    "2913117": "I forgot to mention in my write-up that we also used Exponential Moving Average (EMA) to avoid overfitting.",
    "2914110": "Nice! good job, thanks for sharing!",
    "2914280": "> In selecting the best checkpoints, we used two strategies: (1) a single checkpoint is selected based on the average score of the three targets, and (2) select three checkpoints based on the scores of each of the three targets. In the private LB score, the former strategy was better.\n\nFor 2, did you use average predictions of the three checkpoints? Or use each individual checkpoint for predicting that particular target?",
    "2914283": "I suspect the pretraining that the winner did could also help? Predict masked smiles to smiles, and/or predict ecfp from smiles. Predict chemically significant features. Why? With this train data, the easy path for the model is to learn \"which building block am I?\" which is exactly what you do NOT want. \n\nTarget encoding alone was too effectives on shared data, so it's definitely a modeling concern in this dataset and this domain.",
    "2914458": "We used each individual checkpoint for predicting that particular target.",
    "2918137": "I have observed in my solution too. The first epoch performs the best. Isn't it amazing?"
  },
  "source": "meta"
}