{
  "id": 521894,
  "title": "[5th solution]: Ensemble of CNN1d, Transformer, Mamba models",
  "url": "/competitions/leash-BELKA/writeups/mamba1-one-fold-lb0-432-5th-solution-ensemble-of-c",
  "author_name": "",
  "post_date": "2024-07-26T02:27:42.600Z",
  "votes": 34,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>GitHub code:</h1>\n<p><a href=\"https://github.com/hengck23/solution-leash-BELKA\" target=\"_blank\">https://github.com/hengck23/solution-leash-BELKA</a>  </p>\n<hr>\n<h1>Brief description of the method:</h1>\n<h2>1. Tokenization</h2>\n<p>We use only character tokenization. we use CNN embedding (kernel size=3, stride=1) to learned combination of characters. We try other tokenizers like BPE, sentence piece, atom/smiles based, etc. But all perform worse than the simplest character-based tokenization.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc983825b38881f09b12733c260f0e650%2FSelection_153.png?generation=1721740500859348&amp;alt=media\" alt=\"\"></p>\n<h2>2. Network architecture</h2>\n<p>The final solution is an ensemble of 3 net architectures: cnn1d, transformer, mamba (SSM). We treat input as a sequence and the task as a 3-class ('BRD4', 'HSA', 'sEH') multi-label problem. We train with large batch sizes:  cnn1d=5000, transformer=2500, mamba=2000. For cnn1d, performance is very sensitive to BN for large batch sizes. We think it is because of:</p>\n<ul>\n<li>in-distribution and out-distribution samples have different feature values.</li>\n<li>class is imbalance (positive class is less than 1%). positive and negative samples also have different feature values.</li>\n</ul>\n<p>We use high eps=5e-3 and low momentum=0.2 for cnn1d net.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0a518b669428d59245b0f0cb51bde9ed%2FSelection_161.png?generation=1721740806722371&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1cd8da1f05892e2921dabea4c956529c%2FSelection_167.png?generation=1721777492055677&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h1>Key observation:</h1>\n<p>With 98 million training molecules, it takes quite a lot of time to train each neural net. We do not have sufficient time to train different nets for each fold. Instead, we used different folds for different nets to improve ensemble diversity.</p>\n<ul>\n<li>7 hr for cnn1d (one fold)</li>\n<li>28 hr for transformer (one fold)  </li>\n<li>36 hr for mamba (one fold)  </li>\n</ul>\n<p>This makes it difficult to compare performance for different nets. After the competition, we make some late submissions. Here are the results. It can be seen that the transformer is the most robust net.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F74090653a5768b187232ff9615694b92%2FSelection_169.png?generation=1721908106311645&amp;alt=media\" alt=\"\"></p>\n<p>Next, we compare some heatmaps generated by gradCAM:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a3fe8f4b64f8d5da65d8df136d7488e%2FSelection_190.png?generation=1721960668364141&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F084217c2f874635eacd44d98e3a21b1c%2FSelection_189.png?generation=1721960706372788&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f9feba17c936bec539bb9249c05c14f%2FSelection_187.png?generation=1721960724767966&amp;alt=media\" alt=\"\"></p>\n<p>As expected(?) activation of cnn1d is quite local. Transformer has much global activation.</p>\n<hr>\n<h2>Acknowledgement</h2>\n<h2>\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"</h2>",
  "messages": [
    {
      "id": "2933090",
      "postDate": "07/23/2024 13:21:23",
      "content": "<h1>GitHub code:</h1>\n<p><a href=\"https://github.com/hengck23/solution-leash-BELKA\" target=\"_blank\">https://github.com/hengck23/solution-leash-BELKA</a>  </p>\n<hr>\n<h1>Brief description of the method:</h1>\n<h2>1. Tokenization</h2>\n<p>We use only character tokenization. we use CNN embedding (kernel size=3, stride=1) to learned combination of characters. We try other tokenizers like BPE, sentence piece, atom/smiles based, etc. But all perform worse than the simplest character-based tokenization.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc983825b38881f09b12733c260f0e650%2FSelection_153.png?generation=1721740500859348&amp;alt=media\" alt=\"\"></p>\n<h2>2. Network architecture</h2>\n<p>The final solution is an ensemble of 3 net architectures: cnn1d, transformer, mamba (SSM). We treat input as a sequence and the task as a 3-class ('BRD4', 'HSA', 'sEH') multi-label problem. We train with large batch sizes:  cnn1d=5000, transformer=2500, mamba=2000. For cnn1d, performance is very sensitive to BN for large batch sizes. We think it is because of:</p>\n<ul>\n<li>in-distribution and out-distribution samples have different feature values.</li>\n<li>class is imbalance (positive class is less than 1%). positive and negative samples also have different feature values.</li>\n</ul>\n<p>We use high eps=5e-3 and low momentum=0.2 for cnn1d net.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0a518b669428d59245b0f0cb51bde9ed%2FSelection_161.png?generation=1721740806722371&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1cd8da1f05892e2921dabea4c956529c%2FSelection_167.png?generation=1721777492055677&amp;alt=media\" alt=\"\"></p>\n<hr>\n<h1>Key observation:</h1>\n<p>With 98 million training molecules, it takes quite a lot of time to train each neural net. We do not have sufficient time to train different nets for each fold. Instead, we used different folds for different nets to improve ensemble diversity.</p>\n<ul>\n<li>7 hr for cnn1d (one fold)</li>\n<li>28 hr for transformer (one fold)  </li>\n<li>36 hr for mamba (one fold)  </li>\n</ul>\n<p>This makes it difficult to compare performance for different nets. After the competition, we make some late submissions. Here are the results. It can be seen that the transformer is the most robust net.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F74090653a5768b187232ff9615694b92%2FSelection_169.png?generation=1721908106311645&amp;alt=media\" alt=\"\"></p>\n<p>Next, we compare some heatmaps generated by gradCAM:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a3fe8f4b64f8d5da65d8df136d7488e%2FSelection_190.png?generation=1721960668364141&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F084217c2f874635eacd44d98e3a21b1c%2FSelection_189.png?generation=1721960706372788&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f9feba17c936bec539bb9249c05c14f%2FSelection_187.png?generation=1721960724767966&amp;alt=media\" alt=\"\"></p>\n<p>As expected(?) activation of cnn1d is quite local. Transformer has much global activation.</p>\n<hr>\n<h2>Acknowledgement</h2>\n<h2>\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"</h2>",
      "rawMarkdown": "#GitHub code:\nhttps://github.com/hengck23/solution-leash-BELKA  \n\n---\n\n#Brief description of the method:\n##1. Tokenization\nWe use only character tokenization. we use CNN embedding (kernel size=3, stride=1) to learned combination of characters. We try other tokenizers like BPE, sentence piece, atom/smiles based, etc. But all perform worse than the simplest character-based tokenization.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc983825b38881f09b12733c260f0e650%2FSelection_153.png?generation=1721740500859348&alt=media)\n\n##2. Network architecture\nThe final solution is an ensemble of 3 net architectures: cnn1d, transformer, mamba (SSM). We treat input as a sequence and the task as a 3-class ('BRD4', 'HSA', 'sEH') multi-label problem. We train with large batch sizes:  cnn1d=5000, transformer=2500, mamba=2000. For cnn1d, performance is very sensitive to BN for large batch sizes. We think it is because of:\n  -  in-distribution and out-distribution samples have different feature values.\n  - class is imbalance (positive class is less than 1%). positive and negative samples also have different feature values.\n\nWe use high eps=5e-3 and low momentum=0.2 for cnn1d net.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0a518b669428d59245b0f0cb51bde9ed%2FSelection_161.png?generation=1721740806722371&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1cd8da1f05892e2921dabea4c956529c%2FSelection_167.png?generation=1721777492055677&alt=media)\n\n\n---\n#Key observation:\n\nWith 98 million training molecules, it takes quite a lot of time to train each neural net. We do not have sufficient time to train different nets for each fold. Instead, we used different folds for different nets to improve ensemble diversity.\n\n- 7 hr for cnn1d (one fold)\n- 28 hr for transformer (one fold)  \n- 36 hr for mamba (one fold)  \n\nThis makes it difficult to compare performance for different nets. After the competition, we make some late submissions. Here are the results. It can be seen that the transformer is the most robust net.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F74090653a5768b187232ff9615694b92%2FSelection_169.png?generation=1721908106311645&alt=media)\n\nNext, we compare some heatmaps generated by gradCAM:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a3fe8f4b64f8d5da65d8df136d7488e%2FSelection_190.png?generation=1721960668364141&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F084217c2f874635eacd44d98e3a21b1c%2FSelection_189.png?generation=1721960706372788&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f9feba17c936bec539bb9249c05c14f%2FSelection_187.png?generation=1721960724767966&alt=media)\n\nAs expected(?) activation of cnn1d is quite local. Transformer has much global activation.\n\n---\n\n##Acknowledgement\n##\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"",
      "votes": null
    },
    {
      "id": "2939090",
      "postDate": "07/28/2024 18:21:41",
      "content": "<p>Thanks for sharing and congratulations! Very interesting to see the model interpretation using gradcam and xsmiles. Would you mind sharing the code for that?</p>",
      "rawMarkdown": "Thanks for sharing and congratulations! Very interesting to see the model interpretation using gradcam and xsmiles. Would you mind sharing the code for that?",
      "votes": null
    },
    {
      "id": "2939224",
      "postDate": "07/28/2024 21:05:54",
      "content": "<p>please see<br>\n<a href=\"https://www.kaggle.com/code/hengck23/example-to-use-gradcam-and-xsmiles\" target=\"_blank\">https://www.kaggle.com/code/hengck23/example-to-use-gradcam-and-xsmiles</a></p>",
      "rawMarkdown": "please see\nhttps://www.kaggle.com/code/hengck23/example-to-use-gradcam-and-xsmiles",
      "votes": null
    },
    {
      "id": "2939301",
      "postDate": "07/29/2024 00:33:52",
      "content": "<p>awesome  !</p>",
      "rawMarkdown": "awesome  !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2939090,
      "author_name": "lililycai",
      "author_url": "",
      "post_date": "07/28/2024 18:21:41",
      "content": "<p>Thanks for sharing and congratulations! Very interesting to see the model interpretation using gradcam and xsmiles. Would you mind sharing the code for that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2939224,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/28/2024 21:05:54",
          "content": "<p>please see<br>\n<a href=\"https://www.kaggle.com/code/hengck23/example-to-use-gradcam-and-xsmiles\" target=\"_blank\">https://www.kaggle.com/code/hengck23/example-to-use-gradcam-and-xsmiles</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2939301,
              "author_name": "lililycai",
              "author_url": "",
              "post_date": "07/29/2024 00:33:52",
              "content": "<p>awesome  !</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2933090": "#GitHub code:\nhttps://github.com/hengck23/solution-leash-BELKA  \n\n---\n\n#Brief description of the method:\n##1. Tokenization\nWe use only character tokenization. we use CNN embedding (kernel size=3, stride=1) to learned combination of characters. We try other tokenizers like BPE, sentence piece, atom/smiles based, etc. But all perform worse than the simplest character-based tokenization.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc983825b38881f09b12733c260f0e650%2FSelection_153.png?generation=1721740500859348&alt=media)\n\n##2. Network architecture\nThe final solution is an ensemble of 3 net architectures: cnn1d, transformer, mamba (SSM). We treat input as a sequence and the task as a 3-class ('BRD4', 'HSA', 'sEH') multi-label problem. We train with large batch sizes:  cnn1d=5000, transformer=2500, mamba=2000. For cnn1d, performance is very sensitive to BN for large batch sizes. We think it is because of:\n  -  in-distribution and out-distribution samples have different feature values.\n  - class is imbalance (positive class is less than 1%). positive and negative samples also have different feature values.\n\nWe use high eps=5e-3 and low momentum=0.2 for cnn1d net.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0a518b669428d59245b0f0cb51bde9ed%2FSelection_161.png?generation=1721740806722371&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1cd8da1f05892e2921dabea4c956529c%2FSelection_167.png?generation=1721777492055677&alt=media)\n\n\n---\n#Key observation:\n\nWith 98 million training molecules, it takes quite a lot of time to train each neural net. We do not have sufficient time to train different nets for each fold. Instead, we used different folds for different nets to improve ensemble diversity.\n\n- 7 hr for cnn1d (one fold)\n- 28 hr for transformer (one fold)  \n- 36 hr for mamba (one fold)  \n\nThis makes it difficult to compare performance for different nets. After the competition, we make some late submissions. Here are the results. It can be seen that the transformer is the most robust net.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F74090653a5768b187232ff9615694b92%2FSelection_169.png?generation=1721908106311645&alt=media)\n\nNext, we compare some heatmaps generated by gradCAM:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8a3fe8f4b64f8d5da65d8df136d7488e%2FSelection_190.png?generation=1721960668364141&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F084217c2f874635eacd44d98e3a21b1c%2FSelection_189.png?generation=1721960706372788&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f9feba17c936bec539bb9249c05c14f%2FSelection_187.png?generation=1721960724767966&alt=media)\n\nAs expected(?) activation of cnn1d is quite local. Transformer has much global activation.\n\n---\n\n##Acknowledgement\n##\"We extend our thanks to HP for providing the Z8 Fury-G5 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"",
    "2939090": "Thanks for sharing and congratulations! Very interesting to see the model interpretation using gradcam and xsmiles. Would you mind sharing the code for that?",
    "2939224": "please see\nhttps://www.kaggle.com/code/hengck23/example-to-use-gradcam-and-xsmiles",
    "2939301": "awesome  !"
  },
  "source": "meta"
}