{
  "id": 403044,
  "title": "26th Place Solution GRU Ensemble",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/403044",
  "author_name": "",
  "post_date": "2023-04-20T22:01:00.052520700Z",
  "votes": 18,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all congratulations to all top scoring solutions and competitors. This was really a challenging but very interresting competition.</p>\n<p>The basic training and inference setup I kept using are mostly the same as the <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">training</a> and <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\" target=\"_blank\">inference</a> notebooks I published earlier in the competition.</p>\n<p>I continued optimizing those. Made them run on the TPU's in combination with TFRecords. This allowed me to continue increasing the size of the GRU models and drastically increased the training time.</p>\n<p>I also tried continuing the training of the public GNN model. With some mixed succes. Eventually my problem was the amount of compute resources available. I decided to stick with the GRU based models only. Also ensembling the GRU models seemed to have better performance than ensembling with GNN models.</p>\n<p>My final solution was a combination of 5 GRU models trained on data with a max of 128 pulses and 6 GRU models trained on data with a max of 160 pulses used for each event.</p>\n<p>My final submission can be found <a href=\"https://www.kaggle.com/code/rsmits/26th-place-solution-lstm-ensemble\" target=\"_blank\">here</a></p>\n<p><strong>What did work:</strong></p>\n<ol>\n<li>I kept increasing the number of bins to use. While performance kept increasing it started to slowly level of. Eventually I settled for using 64 bins which resulted in a head layer of 4096 dense units.</li>\n<li>Moving from the .npz files to convert everything to TFRecords. This allowed me to use almost all data and train the various notebooks efficiently on both Kaggle TPU VM and Colab Pro TPU's.</li>\n<li>Increasing the number of GRU layers.</li>\n<li>Increasing the amount of cells in each GRU layer.</li>\n<li>Depending on the model and amount of data I could train 8 to 12 epochs in a Colab Pro session of 12 hours. I used a stepwise learning rate decay. First notebook trained with around 0.0004. Second notebook continue training with 0.0003, third with 0.0002 and fourth with 0.00015.</li>\n<li>I kept using a basic batch size of 4096.</li>\n</ol>\n<p><strong>What didn't work or didn't seem to work:</strong></p>\n<ol>\n<li>Predicting - again in a classification setup - zenith and azimuth seperately. Allthough performance was good and close to just using 1 classification layer.</li>\n<li>Predicting - again in a classification setup - the x, y, z coordinates.</li>\n<li>Various learning rate schedules and larger batch sizes (beyond 4096). This didn't seem to have much effect.</li>\n<li>I did some experiments with weight decay and warm-up but it didn't seem to have much effect. This could however be related to the chosen settings.</li>\n</ol>",
  "messages": [
    {
      "id": "2228865",
      "postDate": "04/20/2023 22:01:00",
      "content": "<p>First of all congratulations to all top scoring solutions and competitors. This was really a challenging but very interresting competition.</p>\n<p>The basic training and inference setup I kept using are mostly the same as the <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">training</a> and <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference\" target=\"_blank\">inference</a> notebooks I published earlier in the competition.</p>\n<p>I continued optimizing those. Made them run on the TPU's in combination with TFRecords. This allowed me to continue increasing the size of the GRU models and drastically increased the training time.</p>\n<p>I also tried continuing the training of the public GNN model. With some mixed succes. Eventually my problem was the amount of compute resources available. I decided to stick with the GRU based models only. Also ensembling the GRU models seemed to have better performance than ensembling with GNN models.</p>\n<p>My final solution was a combination of 5 GRU models trained on data with a max of 128 pulses and 6 GRU models trained on data with a max of 160 pulses used for each event.</p>\n<p>My final submission can be found <a href=\"https://www.kaggle.com/code/rsmits/26th-place-solution-lstm-ensemble\" target=\"_blank\">here</a></p>\n<p><strong>What did work:</strong></p>\n<ol>\n<li>I kept increasing the number of bins to use. While performance kept increasing it started to slowly level of. Eventually I settled for using 64 bins which resulted in a head layer of 4096 dense units.</li>\n<li>Moving from the .npz files to convert everything to TFRecords. This allowed me to use almost all data and train the various notebooks efficiently on both Kaggle TPU VM and Colab Pro TPU's.</li>\n<li>Increasing the number of GRU layers.</li>\n<li>Increasing the amount of cells in each GRU layer.</li>\n<li>Depending on the model and amount of data I could train 8 to 12 epochs in a Colab Pro session of 12 hours. I used a stepwise learning rate decay. First notebook trained with around 0.0004. Second notebook continue training with 0.0003, third with 0.0002 and fourth with 0.00015.</li>\n<li>I kept using a basic batch size of 4096.</li>\n</ol>\n<p><strong>What didn't work or didn't seem to work:</strong></p>\n<ol>\n<li>Predicting - again in a classification setup - zenith and azimuth seperately. Allthough performance was good and close to just using 1 classification layer.</li>\n<li>Predicting - again in a classification setup - the x, y, z coordinates.</li>\n<li>Various learning rate schedules and larger batch sizes (beyond 4096). This didn't seem to have much effect.</li>\n<li>I did some experiments with weight decay and warm-up but it didn't seem to have much effect. This could however be related to the chosen settings.</li>\n</ol>",
      "rawMarkdown": "First of all congratulations to all top scoring solutions and competitors. This was really a challenging but very interresting competition.\n\nThe basic training and inference setup I kept using are mostly the same as the [training](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu) and [inference](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference) notebooks I published earlier in the competition.\n\nI continued optimizing those. Made them run on the TPU's in combination with TFRecords. This allowed me to continue increasing the size of the GRU models and drastically increased the training time.\n\nI also tried continuing the training of the public GNN model. With some mixed succes. Eventually my problem was the amount of compute resources available. I decided to stick with the GRU based models only. Also ensembling the GRU models seemed to have better performance than ensembling with GNN models.\n\nMy final solution was a combination of 5 GRU models trained on data with a max of 128 pulses and 6 GRU models trained on data with a max of 160 pulses used for each event.\n\nMy final submission can be found [here](https://www.kaggle.com/code/rsmits/26th-place-solution-lstm-ensemble)\n\n**What did work:**\n1. I kept increasing the number of bins to use. While performance kept increasing it started to slowly level of. Eventually I settled for using 64 bins which resulted in a head layer of 4096 dense units.\n2. Moving from the .npz files to convert everything to TFRecords. This allowed me to use almost all data and train the various notebooks efficiently on both Kaggle TPU VM and Colab Pro TPU's.\n3. Increasing the number of GRU layers.\n4. Increasing the amount of cells in each GRU layer.\n5. Depending on the model and amount of data I could train 8 to 12 epochs in a Colab Pro session of 12 hours. I used a stepwise learning rate decay. First notebook trained with around 0.0004. Second notebook continue training with 0.0003, third with 0.0002 and fourth with 0.00015.\n6. I kept using a basic batch size of 4096.\n\n**What didn't work or didn't seem to work:**\n1. Predicting - again in a classification setup - zenith and azimuth seperately. Allthough performance was good and close to just using 1 classification layer.\n2. Predicting - again in a classification setup - the x, y, z coordinates.\n3. Various learning rate schedules and larger batch sizes (beyond 4096). This didn't seem to have much effect.\n4. I did some experiments with weight decay and warm-up but it didn't seem to have much effect. This could however be related to the chosen settings.",
      "votes": null
    },
    {
      "id": "2229341",
      "postDate": "04/21/2023 09:30:35",
      "content": "<p>Thank you very much for publishing your versions of GRU! <br>\nInterestingly, we used a much smaller net: a Bidirectional GRU with 3 layers and a hidden size of 160, followed by one hidden layer of size 512 and a 3-dimensional output (which we normalized to one). We essentially used the angular distance score as a loss. This setup was sometimes unstable in the first couple of thousands of steps (we used a batch size of 2048 due to hardware limitations), so we had to restart the run with a different random seed. But after that, it trained well.</p>",
      "rawMarkdown": "Thank you very much for publishing your versions of GRU! \nInterestingly, we used a much smaller net: a Bidirectional GRU with 3 layers and a hidden size of 160, followed by one hidden layer of size 512 and a 3-dimensional output (which we normalized to one). We essentially used the angular distance score as a loss. This setup was sometimes unstable in the first couple of thousands of steps (we used a batch size of 2048 due to hardware limitations), so we had to restart the run with a different random seed. But after that, it trained well.",
      "votes": null
    },
    {
      "id": "2229371",
      "postDate": "04/21/2023 10:05:04",
      "content": "<p>I would like to thank you <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> for your contribution in this competition. Part of our solution was inspired by your work. </p>",
      "rawMarkdown": "I would like to thank you @rsmits for your contribution in this competition. Part of our solution was inspired by your work.",
      "votes": null
    },
    {
      "id": "2229450",
      "postDate": "04/21/2023 11:26:21",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> You're welcome :-) With an 8th place your team did great!</p>",
      "rawMarkdown": "remekkinas You're welcome :-) With an 8th place your team did great!",
      "votes": null
    },
    {
      "id": "2229502",
      "postDate": "04/21/2023 12:27:29",
      "content": "<p><a href=\"https://www.kaggle.com/inartimiryasov\" target=\"_blank\">@inartimiryasov</a> Thanks and your welcome :-) Just went through your solution. I think that the combination of using different Loss function and the fact that you used the returned states likely make the difference for improving the GRU based model to perform into the 0.98X region. Very cool!</p>",
      "rawMarkdown": "inartimiryasov Thanks and your welcome :-) Just went through your solution. I think that the combination of using different Loss function and the fact that you used the returned states likely make the difference for improving the GRU based model to perform into the 0.98X region. Very cool!",
      "votes": null
    },
    {
      "id": "2229800",
      "postDate": "04/21/2023 17:33:12",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> thanks for publishing your training and inference code. I started this competition by going through your code and it was super helpful.</p>\n<p>What was your best single model on local validation?</p>",
      "rawMarkdown": "rsmits thanks for publishing your training and inference code. I started this competition by going through your code and it was super helpful.\n\nWhat was your best single model on local validation?",
      "votes": null
    },
    {
      "id": "2229836",
      "postDate": "04/21/2023 18:37:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> Thanks and you're welcome :-)<br>\nMy best local validation was 0.99757 on training batches 1, 2 and 3. Number of pulses used was 128.</p>",
      "rawMarkdown": "Hi @crodoc Thanks and you're welcome :-)\nMy best local validation was 0.99757 on training batches 1, 2 and 3. Number of pulses used was 128.",
      "votes": null
    },
    {
      "id": "2229842",
      "postDate": "04/21/2023 18:41:22",
      "content": "<p>Nice! I wasn't able to squeeze out that much with practically the same approach.</p>",
      "rawMarkdown": "Nice! I wasn't able to squeeze out that much with practically the same approach.",
      "votes": null
    },
    {
      "id": "2229967",
      "postDate": "04/21/2023 21:40:20",
      "content": "<p>Thank you! I read what I had written in the solution, and it wasn't very clear. We also used the hidden state of the last unit (2 * hidden_size since there are hidden states for both the forward and backward GRUs). I believe this corresponds to setting return_sequences = False in TensorFlow.</p>\n<p>Apart from using a different loss function, which outperformed some other regression options that we tried, we also trained the model on the entire training set.</p>",
      "rawMarkdown": "Thank you! I read what I had written in the solution, and it wasn't very clear. We also used the hidden state of the last unit (2 * hidden_size since there are hidden states for both the forward and backward GRUs). I believe this corresponds to setting return_sequences = False in TensorFlow.\n\nApart from using a different loss function, which outperformed some other regression options that we tried, we also trained the model on the entire training set.",
      "votes": null
    },
    {
      "id": "2231052",
      "postDate": "04/23/2023 03:50:12",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> your notebooks were very much helpful for us. Thank you.</p>",
      "rawMarkdown": "rsmits your notebooks were very much helpful for us. Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2229341,
      "author_name": "inartimiryasov",
      "author_url": "",
      "post_date": "04/21/2023 09:30:35",
      "content": "<p>Thank you very much for publishing your versions of GRU! <br>\nInterestingly, we used a much smaller net: a Bidirectional GRU with 3 layers and a hidden size of 160, followed by one hidden layer of size 512 and a 3-dimensional output (which we normalized to one). We essentially used the angular distance score as a loss. This setup was sometimes unstable in the first couple of thousands of steps (we used a batch size of 2048 due to hardware limitations), so we had to restart the run with a different random seed. But after that, it trained well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2229502,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/21/2023 12:27:29",
          "content": "<p><a href=\"https://www.kaggle.com/inartimiryasov\" target=\"_blank\">@inartimiryasov</a> Thanks and your welcome :-) Just went through your solution. I think that the combination of using different Loss function and the fact that you used the returned states likely make the difference for improving the GRU based model to perform into the 0.98X region. Very cool!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2229967,
              "author_name": "inartimiryasov",
              "author_url": "",
              "post_date": "04/21/2023 21:40:20",
              "content": "<p>Thank you! I read what I had written in the solution, and it wasn't very clear. We also used the hidden state of the last unit (2 * hidden_size since there are hidden states for both the forward and backward GRUs). I believe this corresponds to setting return_sequences = False in TensorFlow.</p>\n<p>Apart from using a different loss function, which outperformed some other regression options that we tried, we also trained the model on the entire training set.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2229371,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/21/2023 10:05:04",
      "content": "<p>I would like to thank you <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> for your contribution in this competition. Part of our solution was inspired by your work. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2229450,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/21/2023 11:26:21",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> You're welcome :-) With an 8th place your team did great!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2229800,
      "author_name": "crodoc",
      "author_url": "",
      "post_date": "04/21/2023 17:33:12",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> thanks for publishing your training and inference code. I started this competition by going through your code and it was super helpful.</p>\n<p>What was your best single model on local validation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2229836,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/21/2023 18:37:12",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> Thanks and you're welcome :-)<br>\nMy best local validation was 0.99757 on training batches 1, 2 and 3. Number of pulses used was 128.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2229842,
              "author_name": "crodoc",
              "author_url": "",
              "post_date": "04/21/2023 18:41:22",
              "content": "<p>Nice! I wasn't able to squeeze out that much with practically the same approach.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2231052,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "04/23/2023 03:50:12",
      "content": "<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> your notebooks were very much helpful for us. Thank you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2228865": "First of all congratulations to all top scoring solutions and competitors. This was really a challenging but very interresting competition.\n\nThe basic training and inference setup I kept using are mostly the same as the [training](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu) and [inference](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-inference) notebooks I published earlier in the competition.\n\nI continued optimizing those. Made them run on the TPU's in combination with TFRecords. This allowed me to continue increasing the size of the GRU models and drastically increased the training time.\n\nI also tried continuing the training of the public GNN model. With some mixed succes. Eventually my problem was the amount of compute resources available. I decided to stick with the GRU based models only. Also ensembling the GRU models seemed to have better performance than ensembling with GNN models.\n\nMy final solution was a combination of 5 GRU models trained on data with a max of 128 pulses and 6 GRU models trained on data with a max of 160 pulses used for each event.\n\nMy final submission can be found [here](https://www.kaggle.com/code/rsmits/26th-place-solution-lstm-ensemble)\n\n**What did work:**\n1. I kept increasing the number of bins to use. While performance kept increasing it started to slowly level of. Eventually I settled for using 64 bins which resulted in a head layer of 4096 dense units.\n2. Moving from the .npz files to convert everything to TFRecords. This allowed me to use almost all data and train the various notebooks efficiently on both Kaggle TPU VM and Colab Pro TPU's.\n3. Increasing the number of GRU layers.\n4. Increasing the amount of cells in each GRU layer.\n5. Depending on the model and amount of data I could train 8 to 12 epochs in a Colab Pro session of 12 hours. I used a stepwise learning rate decay. First notebook trained with around 0.0004. Second notebook continue training with 0.0003, third with 0.0002 and fourth with 0.00015.\n6. I kept using a basic batch size of 4096.\n\n**What didn't work or didn't seem to work:**\n1. Predicting - again in a classification setup - zenith and azimuth seperately. Allthough performance was good and close to just using 1 classification layer.\n2. Predicting - again in a classification setup - the x, y, z coordinates.\n3. Various learning rate schedules and larger batch sizes (beyond 4096). This didn't seem to have much effect.\n4. I did some experiments with weight decay and warm-up but it didn't seem to have much effect. This could however be related to the chosen settings.",
    "2229341": "Thank you very much for publishing your versions of GRU! \nInterestingly, we used a much smaller net: a Bidirectional GRU with 3 layers and a hidden size of 160, followed by one hidden layer of size 512 and a 3-dimensional output (which we normalized to one). We essentially used the angular distance score as a loss. This setup was sometimes unstable in the first couple of thousands of steps (we used a batch size of 2048 due to hardware limitations), so we had to restart the run with a different random seed. But after that, it trained well.",
    "2229371": "I would like to thank you @rsmits for your contribution in this competition. Part of our solution was inspired by your work.",
    "2229450": "remekkinas You're welcome :-) With an 8th place your team did great!",
    "2229502": "inartimiryasov Thanks and your welcome :-) Just went through your solution. I think that the combination of using different Loss function and the fact that you used the returned states likely make the difference for improving the GRU based model to perform into the 0.98X region. Very cool!",
    "2229800": "rsmits thanks for publishing your training and inference code. I started this competition by going through your code and it was super helpful.\n\nWhat was your best single model on local validation?",
    "2229836": "Hi @crodoc Thanks and you're welcome :-)\nMy best local validation was 0.99757 on training batches 1, 2 and 3. Number of pulses used was 128.",
    "2229842": "Nice! I wasn't able to squeeze out that much with practically the same approach.",
    "2229967": "Thank you! I read what I had written in the solution, and it wasn't very clear. We also used the hidden state of the last unit (2 * hidden_size since there are hidden states for both the forward and backward GRUs). I believe this corresponds to setting return_sequences = False in TensorFlow.\n\nApart from using a different loss function, which outperformed some other regression options that we tried, we also trained the model on the entire training set.",
    "2231052": "rsmits your notebooks were very much helpful for us. Thank you."
  },
  "source": "meta"
}