{
  "id": 663293,
  "title": "[20th place solution] AWP with Robust MLP on ESM2 Embeddings",
  "url": "/competitions/adaptive-immune-profiling-challenge-2025/discussion/663293",
  "author_name": "Hi F",
  "post_date": "2025-12-17T07:50:47.750000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, congratulations to the winners!</p>\n<p>I'd like to share a training notebook based on a simple <strong>MLP using ESM2 embeddings</strong>.\nThe main focus of this approach is the training strategy designed to improve generalization, specifically using <strong>Pseudo Labeling</strong>, <strong>AWP (Adversarial Weight Perturbation)</strong>, and <strong>robust regularization</strong>.</p>\n<h3>Key Features</h3>\n<h4>1. Model Architecture</h4>\n<ul>\n<li><strong>Input:</strong> ESM2 Embeddings (dim=640)</li>\n<li><strong>Structure:</strong> 3-layer MLP (640 → 512 → 512 → 1)</li>\n<li><strong>Layer Block:</strong> <code>Linear</code> → <code>BatchNorm</code> → <code>GELU</code> → <code>Dropout (p=0.5)</code> → <code>DropPath (p=0.2)</code></li>\n</ul>\n<blockquote>\n  <p>I used Dropout (0.5) combined with DropPath to enforce robustness in the MLP structure.</p>\n</blockquote>\n<h4>2. Training Strategy</h4>\n<ul>\n<li><strong>Pseudo Labeling:</strong>\nI incorporated pseudo-labeling to leverage additional data, helping the model learn better decision boundaries and further improve performance.</li>\n<li><strong>AWP (Adversarial Weight Perturbation):</strong>\nI implemented AWP to explicitly optimize for flat minima. By adding adversarial perturbations to the weights during training, the model becomes more robust to shifts in the test data.</li>\n</ul>\n<h4>3. Validation Strategy</h4>\n<ul>\n<li><strong>5-Seed Averaging:</strong>\nTo ensure the robustness of the CV score, I averaged the predictions over 5 different random seeds in addition to the Stratified K-Fold. This helped to reduce variance and provided a more reliable estimate of model performance.</li>\n</ul>\n<blockquote>\n  <p>train 7.  cv 0.6619 (5fold × 5seeds average)<br>\n  train 8.  cv 0.7715 (5fold × 5seeds average)</p>\n</blockquote>\n<h4>4. Limitation / Observation</h4>\n<ul>\n<li><strong>Performance on Datasets 1-6:</strong>\nOne important thing to note is that this model could not outperform the <strong>k-mer approach</strong> on Datasets 1 through 6. For those datasets, the deep learning approach did not work well in my experiments, and simple k-mer features yielded better results.</li>\n</ul>\n<p>notebook:  <a href=\"https://www.kaggle.com/code/hiroakifukuse/pseudo47-airr2025-train7-esm2-awp-25000sample-f\" target=\"_blank\">https://www.kaggle.com/code/hiroakifukuse/pseudo47-airr2025-train7-esm2-awp-25000sample-f</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2F3af61ee4bf704a365baacc261834aeaa%2Funnamed%20(2).jpg?generation=1765958438974835&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3377944,
      "postDate": "2025-12-17T07:50:47.750Z",
      "content": "<p>First of all, congratulations to the winners!</p>\n<p>I'd like to share a training notebook based on a simple <strong>MLP using ESM2 embeddings</strong>.\nThe main focus of this approach is the training strategy designed to improve generalization, specifically using <strong>Pseudo Labeling</strong>, <strong>AWP (Adversarial Weight Perturbation)</strong>, and <strong>robust regularization</strong>.</p>\n<h3>Key Features</h3>\n<h4>1. Model Architecture</h4>\n<ul>\n<li><strong>Input:</strong> ESM2 Embeddings (dim=640)</li>\n<li><strong>Structure:</strong> 3-layer MLP (640 → 512 → 512 → 1)</li>\n<li><strong>Layer Block:</strong> <code>Linear</code> → <code>BatchNorm</code> → <code>GELU</code> → <code>Dropout (p=0.5)</code> → <code>DropPath (p=0.2)</code></li>\n</ul>\n<blockquote>\n  <p>I used Dropout (0.5) combined with DropPath to enforce robustness in the MLP structure.</p>\n</blockquote>\n<h4>2. Training Strategy</h4>\n<ul>\n<li><strong>Pseudo Labeling:</strong>\nI incorporated pseudo-labeling to leverage additional data, helping the model learn better decision boundaries and further improve performance.</li>\n<li><strong>AWP (Adversarial Weight Perturbation):</strong>\nI implemented AWP to explicitly optimize for flat minima. By adding adversarial perturbations to the weights during training, the model becomes more robust to shifts in the test data.</li>\n</ul>\n<h4>3. Validation Strategy</h4>\n<ul>\n<li><strong>5-Seed Averaging:</strong>\nTo ensure the robustness of the CV score, I averaged the predictions over 5 different random seeds in addition to the Stratified K-Fold. This helped to reduce variance and provided a more reliable estimate of model performance.</li>\n</ul>\n<blockquote>\n  <p>train 7.  cv 0.6619 (5fold × 5seeds average)<br>\n  train 8.  cv 0.7715 (5fold × 5seeds average)</p>\n</blockquote>\n<h4>4. Limitation / Observation</h4>\n<ul>\n<li><strong>Performance on Datasets 1-6:</strong>\nOne important thing to note is that this model could not outperform the <strong>k-mer approach</strong> on Datasets 1 through 6. For those datasets, the deep learning approach did not work well in my experiments, and simple k-mer features yielded better results.</li>\n</ul>\n<p>notebook:  <a href=\"https://www.kaggle.com/code/hiroakifukuse/pseudo47-airr2025-train7-esm2-awp-25000sample-f\" target=\"_blank\">https://www.kaggle.com/code/hiroakifukuse/pseudo47-airr2025-train7-esm2-awp-25000sample-f</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2F3af61ee4bf704a365baacc261834aeaa%2Funnamed%20(2).jpg?generation=1765958438974835&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "First of all, congratulations to the winners!\n\nI'd like to share a training notebook based on a simple **MLP using ESM2 embeddings**.\nThe main focus of this approach is the training strategy designed to improve generalization, specifically using **Pseudo Labeling**, **AWP (Adversarial Weight Perturbation)**, and **robust regularization**.\n\n### Key Features\n\n#### 1. Model Architecture\n* **Input:** ESM2 Embeddings (dim=640)\n* **Structure:** 3-layer MLP (640 → 512 → 512 → 1)\n* **Layer Block:** `Linear` → `BatchNorm` → `GELU` → `Dropout (p=0.5)` → `DropPath (p=0.2)`\n\n> I used Dropout (0.5) combined with DropPath to enforce robustness in the MLP structure.\n\n#### 2. Training Strategy\n* **Pseudo Labeling:**\n  I incorporated pseudo-labeling to leverage additional data, helping the model learn better decision boundaries and further improve performance.\n* **AWP (Adversarial Weight Perturbation):**\n  I implemented AWP to explicitly optimize for flat minima. By adding adversarial perturbations to the weights during training, the model becomes more robust to shifts in the test data.\n\n#### 3. Validation Strategy\n* **5-Seed Averaging:**\n  To ensure the robustness of the CV score, I averaged the predictions over 5 different random seeds in addition to the Stratified K-Fold. This helped to reduce variance and provided a more reliable estimate of model performance.\n\n>train 7.  cv 0.6619 (5fold × 5seeds average)  \ntrain 8.  cv 0.7715 (5fold × 5seeds average)\n\n#### 4. Limitation / Observation\n* **Performance on Datasets 1-6:**\n  One important thing to note is that this model could not outperform the **k-mer approach** on Datasets 1 through 6. For those datasets, the deep learning approach did not work well in my experiments, and simple k-mer features yielded better results.\n\nnotebook:  https://www.kaggle.com/code/hiroakifukuse/pseudo47-airr2025-train7-esm2-awp-25000sample-f\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2F3af61ee4bf704a365baacc261834aeaa%2Funnamed%20(2).jpg?generation=1765958438974835&alt=media)",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3377944": "First of all, congratulations to the winners!\n\nI'd like to share a training notebook based on a simple **MLP using ESM2 embeddings**.\nThe main focus of this approach is the training strategy designed to improve generalization, specifically using **Pseudo Labeling**, **AWP (Adversarial Weight Perturbation)**, and **robust regularization**.\n\n### Key Features\n\n#### 1. Model Architecture\n* **Input:** ESM2 Embeddings (dim=640)\n* **Structure:** 3-layer MLP (640 → 512 → 512 → 1)\n* **Layer Block:** `Linear` → `BatchNorm` → `GELU` → `Dropout (p=0.5)` → `DropPath (p=0.2)`\n\n> I used Dropout (0.5) combined with DropPath to enforce robustness in the MLP structure.\n\n#### 2. Training Strategy\n* **Pseudo Labeling:**\n  I incorporated pseudo-labeling to leverage additional data, helping the model learn better decision boundaries and further improve performance.\n* **AWP (Adversarial Weight Perturbation):**\n  I implemented AWP to explicitly optimize for flat minima. By adding adversarial perturbations to the weights during training, the model becomes more robust to shifts in the test data.\n\n#### 3. Validation Strategy\n* **5-Seed Averaging:**\n  To ensure the robustness of the CV score, I averaged the predictions over 5 different random seeds in addition to the Stratified K-Fold. This helped to reduce variance and provided a more reliable estimate of model performance.\n\n>train 7.  cv 0.6619 (5fold × 5seeds average)  \ntrain 8.  cv 0.7715 (5fold × 5seeds average)\n\n#### 4. Limitation / Observation\n* **Performance on Datasets 1-6:**\n  One important thing to note is that this model could not outperform the **k-mer approach** on Datasets 1 through 6. For those datasets, the deep learning approach did not work well in my experiments, and simple k-mer features yielded better results.\n\nnotebook:  https://www.kaggle.com/code/hiroakifukuse/pseudo47-airr2025-train7-esm2-awp-25000sample-f\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2F3af61ee4bf704a365baacc261834aeaa%2Funnamed%20(2).jpg?generation=1765958438974835&alt=media)"
  }
}