{
  "id": 402973,
  "title": "14th place solution: transformer and GRU",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/polar-geese-14th-place-solution-transformer-and-gr",
  "author_name": "",
  "post_date": "2023-04-20T14:01:28.620Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p>We would like to thank the organizers and participants of this great competition! We learned a lot and had a lot of fun. Special thanks to  <a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">ZhaounBooty</a> for his early sharing solution, <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">Robin Smits</a> for sharing further improvements of the RNN idea, and <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">Iafoss</a> for sharing the dataloader.</p>\n<p>Below is a (preliminary)  description of the two models we trained.</p>\n<p>We preprocessed the whole training dataset and saved it as a memory-mapped file. We have written an iterable dataset that allowed very fast data loading.</p>\n<h3>Transformer model:</h3>\n<p><em>Private and public score 0.983</em></p>\n<ul>\n<li>Size: 4.9M params</li>\n<li>Batch: 2000 events</li>\n<li>Hidden size: 256</li>\n<li>Activation: SiLU</li>\n<li>Layernorm</li>\n<li>Num transformer layers: 3</li>\n<li>Num heads: 4</li>\n<li>Global aggregation: [sum, mean, max, std_dev]</li>\n</ul>\n<p>Training:</p>\n<ul>\n<li>LR: 0.002</li>\n<li>Patience: 1, multiplicative decay: 0.9</li>\n<li>Run for 77 epochs</li>\n<li>32 GPUs</li>\n<li>Warmup: 2 epochs</li>\n</ul>\n<h3>Bidirectional GRU model:</h3>\n<p><em>Private and public score 0.987</em></p>\n<ul>\n<li>Size: 1.7 M parameters</li>\n<li>Batch size: 2048</li>\n<li>Input shape: (96, 6)</li>\n<li>Number of layers: 3</li>\n<li>hidden_size: 160</li>\n<li>Hidden states of the last GRU unit are passed to a fully connected layer (size: 512)</li>\n</ul>\n<p>Training</p>\n<ul>\n<li>max LR 1e-3 with a warm up and cosine annealing ( OneCycleLR, pct_start: 0.001)</li>\n<li>AdamW with weight decay 1e-4</li>\n<li>Trained for 14 epochs on the full dataset on GTX 1080 GPU for 90 hours</li>\n</ul>\n<p>In both cases, the network output was a 3-dimensional  unit vector. We then directly used the mean angular error as the loss.</p>\n<p>Interestingly, despite very limited computer resources, the GRU model scores quite well. Unfortunately, we didn't have time to scale it further on the GTX 1080.</p>",
  "messages": [
    {
      "id": "2228377",
      "postDate": "04/20/2023 13:57:06",
      "content": "<p>We would like to thank the organizers and participants of this great competition! We learned a lot and had a lot of fun. Special thanks to  <a href=\"https://www.kaggle.com/seungmoklee\" target=\"_blank\">ZhaounBooty</a> for his early sharing solution, <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">Robin Smits</a> for sharing further improvements of the RNN idea, and <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">Iafoss</a> for sharing the dataloader.</p>\n<p>Below is a (preliminary)  description of the two models we trained.</p>\n<p>We preprocessed the whole training dataset and saved it as a memory-mapped file. We have written an iterable dataset that allowed very fast data loading.</p>\n<h3>Transformer model:</h3>\n<p><em>Private and public score 0.983</em></p>\n<ul>\n<li>Size: 4.9M params</li>\n<li>Batch: 2000 events</li>\n<li>Hidden size: 256</li>\n<li>Activation: SiLU</li>\n<li>Layernorm</li>\n<li>Num transformer layers: 3</li>\n<li>Num heads: 4</li>\n<li>Global aggregation: [sum, mean, max, std_dev]</li>\n</ul>\n<p>Training:</p>\n<ul>\n<li>LR: 0.002</li>\n<li>Patience: 1, multiplicative decay: 0.9</li>\n<li>Run for 77 epochs</li>\n<li>32 GPUs</li>\n<li>Warmup: 2 epochs</li>\n</ul>\n<h3>Bidirectional GRU model:</h3>\n<p><em>Private and public score 0.987</em></p>\n<ul>\n<li>Size: 1.7 M parameters</li>\n<li>Batch size: 2048</li>\n<li>Input shape: (96, 6)</li>\n<li>Number of layers: 3</li>\n<li>hidden_size: 160</li>\n<li>Hidden states of the last GRU unit are passed to a fully connected layer (size: 512)</li>\n</ul>\n<p>Training</p>\n<ul>\n<li>max LR 1e-3 with a warm up and cosine annealing ( OneCycleLR, pct_start: 0.001)</li>\n<li>AdamW with weight decay 1e-4</li>\n<li>Trained for 14 epochs on the full dataset on GTX 1080 GPU for 90 hours</li>\n</ul>\n<p>In both cases, the network output was a 3-dimensional  unit vector. We then directly used the mean angular error as the loss.</p>\n<p>Interestingly, despite very limited computer resources, the GRU model scores quite well. Unfortunately, we didn't have time to scale it further on the GTX 1080.</p>",
      "rawMarkdown": "We would like to thank the organizers and participants of this great competition! We learned a lot and had a lot of fun. Special thanks to  [ZhaounBooty](https://www.kaggle.com/seungmoklee) for his early sharing solution, [Robin Smits](https://www.kaggle.com/rsmits) for sharing further improvements of the RNN idea, and [Iafoss](https://www.kaggle.com/iafoss) for sharing the dataloader.\n\nBelow is a (preliminary)  description of the two models we trained.\n\nWe preprocessed the whole training dataset and saved it as a memory-mapped file. We have written an iterable dataset that allowed very fast data loading.\n\n\n\n### Transformer model:\n\n*Private and public score 0.983*\n\n* Size: 4.9M params\n* Batch: 2000 events\n* Hidden size: 256\n* Activation: SiLU\n* Layernorm\n* Num transformer layers: 3\n* Num heads: 4\n* Global aggregation: [sum, mean, max, std_dev]\n\nTraining:\n* LR: 0.002\n* Patience: 1, multiplicative decay: 0.9\n* Run for 77 epochs\n* 32 GPUs\n* Warmup: 2 epochs\n\n### Bidirectional GRU model:\n\n*Private and public score 0.987*\n\n* Size: 1.7 M parameters\n* Batch size: 2048\n* Input shape: (96, 6)\n* Number of layers: 3\n* hidden_size: 160\n* Hidden states of the last GRU unit are passed to a fully connected layer (size: 512)\n\nTraining\n* max LR 1e-3 with a warm up and cosine annealing ( OneCycleLR, pct_start: 0.001)\n* AdamW with weight decay 1e-4\n* Trained for 14 epochs on the full dataset on GTX 1080 GPU for 90 hours\n\nIn both cases, the network output was a 3-dimensional  unit vector. We then directly used the mean angular error as the loss.\n\nInterestingly, despite very limited computer resources, the GRU model scores quite well. Unfortunately, we didn't have time to scale it further on the GTX 1080.",
      "votes": null
    },
    {
      "id": "2228390",
      "postDate": "04/20/2023 14:08:09",
      "content": "<p>Congratulations! Great score for single models! <br>\nIs any chance you will publish Transformer model script (architecture)? </p>",
      "rawMarkdown": "Congratulations! Great score for single models! \nIs any chance you will publish Transformer model script (architecture)?",
      "votes": null
    },
    {
      "id": "2229319",
      "postDate": "04/21/2023 09:06:01",
      "content": "<p>Hi Remek, thanks very much!</p>\n<p>Sure - here is a gist for the Transformer: <a href=\"https://gist.github.com/timinar/264d2279f5116caa8d824c9427d800a5\" target=\"_blank\">https://gist.github.com/timinar/264d2279f5116caa8d824c9427d800a5</a></p>\n<p>Would love to hear your thoughts on our implementation!</p>",
      "rawMarkdown": "Hi Remek, thanks very much!\n\nSure - here is a gist for the Transformer: https://gist.github.com/timinar/264d2279f5116caa8d824c9427d800a5\n\nWould love to hear your thoughts on our implementation!",
      "votes": null
    },
    {
      "id": "2229503",
      "postDate": "04/21/2023 12:29:04",
      "content": "<p>Very nice solution! And great to see that the GRU model is able to score into the 0.98X region. Congratulations!</p>",
      "rawMarkdown": "Very nice solution! And great to see that the GRU model is able to score into the 0.98X region. Congratulations!",
      "votes": null
    },
    {
      "id": "2229968",
      "postDate": "04/21/2023 21:41:09",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228390,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/20/2023 14:08:09",
      "content": "<p>Congratulations! Great score for single models! <br>\nIs any chance you will publish Transformer model script (architecture)? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2229319,
          "author_name": "murnanedaniel",
          "author_url": "",
          "post_date": "04/21/2023 09:06:01",
          "content": "<p>Hi Remek, thanks very much!</p>\n<p>Sure - here is a gist for the Transformer: <a href=\"https://gist.github.com/timinar/264d2279f5116caa8d824c9427d800a5\" target=\"_blank\">https://gist.github.com/timinar/264d2279f5116caa8d824c9427d800a5</a></p>\n<p>Would love to hear your thoughts on our implementation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2229503,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/21/2023 12:29:04",
      "content": "<p>Very nice solution! And great to see that the GRU model is able to score into the 0.98X region. Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2229968,
          "author_name": "inartimiryasov",
          "author_url": "",
          "post_date": "04/21/2023 21:41:09",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2228377": "We would like to thank the organizers and participants of this great competition! We learned a lot and had a lot of fun. Special thanks to  [ZhaounBooty](https://www.kaggle.com/seungmoklee) for his early sharing solution, [Robin Smits](https://www.kaggle.com/rsmits) for sharing further improvements of the RNN idea, and [Iafoss](https://www.kaggle.com/iafoss) for sharing the dataloader.\n\nBelow is a (preliminary)  description of the two models we trained.\n\nWe preprocessed the whole training dataset and saved it as a memory-mapped file. We have written an iterable dataset that allowed very fast data loading.\n\n\n\n### Transformer model:\n\n*Private and public score 0.983*\n\n* Size: 4.9M params\n* Batch: 2000 events\n* Hidden size: 256\n* Activation: SiLU\n* Layernorm\n* Num transformer layers: 3\n* Num heads: 4\n* Global aggregation: [sum, mean, max, std_dev]\n\nTraining:\n* LR: 0.002\n* Patience: 1, multiplicative decay: 0.9\n* Run for 77 epochs\n* 32 GPUs\n* Warmup: 2 epochs\n\n### Bidirectional GRU model:\n\n*Private and public score 0.987*\n\n* Size: 1.7 M parameters\n* Batch size: 2048\n* Input shape: (96, 6)\n* Number of layers: 3\n* hidden_size: 160\n* Hidden states of the last GRU unit are passed to a fully connected layer (size: 512)\n\nTraining\n* max LR 1e-3 with a warm up and cosine annealing ( OneCycleLR, pct_start: 0.001)\n* AdamW with weight decay 1e-4\n* Trained for 14 epochs on the full dataset on GTX 1080 GPU for 90 hours\n\nIn both cases, the network output was a 3-dimensional  unit vector. We then directly used the mean angular error as the loss.\n\nInterestingly, despite very limited computer resources, the GRU model scores quite well. Unfortunately, we didn't have time to scale it further on the GTX 1080.",
    "2228390": "Congratulations! Great score for single models! \nIs any chance you will publish Transformer model script (architecture)?",
    "2229319": "Hi Remek, thanks very much!\n\nSure - here is a gist for the Transformer: https://gist.github.com/timinar/264d2279f5116caa8d824c9427d800a5\n\nWould love to hear your thoughts on our implementation!",
    "2229503": "Very nice solution! And great to see that the GRU model is able to score into the 0.98X region. Congratulations!",
    "2229968": "Thank you!"
  },
  "source": "meta"
}