{
  "id": 402969,
  "title": "10th Place Solution",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/writeups/crodoc-10th-place-solution",
  "author_name": "",
  "post_date": "2023-04-20T14:56:06.023Z",
  "votes": 36,
  "comment_count": 24,
  "views": 0,
  "content": "<p>The last few weeks have been intense, but in the end I am super happy that I've finally got my first solo gold medal. Great competition overall! It was great to see the stability between local validation and LB scores (even with using only <strong>1.5%</strong> of the data as <strong>validation</strong>).</p>\n<p>The model training process involved a train-validation split, wherein <strong>batches 11-660</strong> were utilized for <strong>training</strong>, and <strong>batches 1-10</strong> were allocated for <strong>validation</strong>. An ensemble of 8 models was formed using <strong>hill climbing</strong> or <strong>Nelder-Mead</strong> optimization techniques, with a submission time of approximately 3-3.5 hours. The inclusion of more models in the ensemble resulted in a better score; however, since I was submitting the best blend only a few hours before the end of the competition, there was not enough time to do inference on a bigger ensemble. The same model architectures were trained multiple times and blended.</p>\n<p>The ensemble's local validation score was <strong>0.97747</strong>, while the public and private LB scores were <strong>0.976</strong>.</p>\n<p>Input data for the model consisted of the following 9 features per event: <strong>sensor_x, sensor_y, sensor_z, time, charge, auxiliary, is_main_sensor, is_deep_veto, and is_deep_core</strong>.</p>\n<p>Four model architectures have been selected for the final submission.</p>\n<p><strong>Model 1</strong>:</p>\n<p>Validation score 1: <strong>1.0016</strong><br>\nValidation score 2: <strong>1.0005</strong><br>\nLoss function: <strong>CrossEntropyLoss</strong></p>\n<p>Data preprocessing steps:</p>\n<ul>\n<li>sensor_x, sensor_y, and sensor_z are divided by 600</li>\n<li>time is divided by 1000 and subtracted by its minimum value</li>\n<li>charge is divided by 300</li>\n</ul>\n<pre><code> (pl.LightningModule):\n     ():\n\n    self.bin_num = \n\n        self.gru = nn.GRU(, , num_layers=, dropout=, batch_first=, bidirectional=)\n        self.fc1 = nn.Sequential(nn.Linear(, ), nn.ReLU())\n        self.fc2 = nn.Linear(, bin_num*bin_num)\n\n     ():\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.gru(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        x = self.fc2(x)\n\n         x\n</code></pre>\n<p>Output bins were generated using <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">code</a> from <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a>. </p>\n<p><strong>Models 2-4</strong>:</p>\n<p>Loss function: <strong>VonMisesFisher3DLoss</strong></p>\n<p>Data preprocessing steps:</p>\n<ul>\n<li>sensor_x, sensor_y, and sensor_z are divided by 500</li>\n<li>time is scaled by subtracting 1.0e04 and dividing by 3.0e4</li>\n<li>charge is transformed using the logarithm base 10 and then divided by 3.0</li>\n</ul>\n<p>Validation score 3: <strong>0.9847</strong><br>\nValidation score 4: <strong>0.9859</strong></p>\n<pre><code> (pl.LightningModule):\n     ():\n\n        self.bilstm = nn.LSTM(, , num_layers=, dropout=, batch_first=, bidirectional=)\n\n        self.fc1 = nn.Sequential(nn.Linear(, ), nn.ReLU())\n        self.dropout = nn.Dropout()\n        self.fc2 = nn.Linear(, )\n\n     ():\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        x = self.dropout(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=, p=) + \n        pred_x = pred[:, ] / kappa\n        pred_y = pred[:, ] / kappa\n        pred_z = pred[:, ] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=)\n\n         pred\n</code></pre>\n<p>Validation score 5: <strong>0.9872</strong><br>\nValidation score 6: <strong>0.9887</strong></p>\n<pre><code> (pl.LightningModule):\n     ():\n        ().__init__()\n\n        self.embedding = nn.Linear(, )\n        self.bilstm = nn.LSTM(, , num_layers=, dropout=, batch_first=, bidirectional=)\n\n        self.fc1 = nn.Sequential(nn.Linear(, ), nn.ReLU())\n        self.fc2 = nn.Linear(, )\n\n     ():\n\n        x = self.embedding(x)\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=, p=) + \n        pred_x = pred[:, ] / kappa\n        pred_y = pred[:, ] / kappa\n        pred_z = pred[:, ] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=)\n\n         pred \n</code></pre>\n<p>Validation score 7: <strong>0.9842</strong><br>\nValidation score 8: <strong>0.9841</strong></p>\n<pre><code> (pl.LightningModule):\n     ():\n\n        self.embedding = nn.Linear(, )\n\n        self.bilstm = nn.LSTM(, , num_layers=, dropout=, batch_first=, bidirectional=)\n\n        self.fc1 = nn.Sequential(nn.Linear(lstm_units, ), nn.ReLU())\n        self.fc2 = nn.Linear(, )\n\n     ():\n\n        batch_sizes = batch_sizes.cpu()\n\n        x = self.embedding(x)\n\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=, p=) + \n        pred_x = pred[:, ] / kappa\n        pred_y = pred[:, ] / kappa\n        pred_z = pred[:, ] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=)\n\n         pred\n</code></pre>\n<p>Hyperparameters:<br>\nOptimizer: <strong>Adam</strong><br>\nScheduler: <strong>CosineAnnealingLR</strong><br>\nBatch size: <strong>2048</strong><br>\nMax pulses: <strong>128</strong><br>\nMax LR: <strong>1e-3</strong> or <strong>5e-4</strong><br>\nMin LR: <strong>1e-6</strong><br>\nWarmup steps: <strong>2000</strong><br>\nEpochs: <strong>10-15</strong> (possibly a few extra fine tuning epochs)</p>\n<p>The DL library used for the competition was PyTorch (Lightning), which delivered better results than TensorFlow. While blending multiple models in TensorFlow, the code produced some unanticipated errors which I wasn't able to debug. It was just simpler to switch to PyTorch.</p>\n<p>I tried a few different transformer architectures, but didn't have enough time to push it through the end. I retrained a few models similar to graphnet (<a href=\"https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite\" target=\"_blank\">code</a> from <a href=\"https://www.kaggle.com/amoshuangyc\" target=\"_blank\">@amoshuangyc</a> was very useful), but it didn't give a boost to the final ensemble.</p>",
  "messages": [
    {
      "id": "2228359",
      "postDate": "04/20/2023 13:45:12",
      "content": "<p>The last few weeks have been intense, but in the end I am super happy that I've finally got my first solo gold medal. Great competition overall! It was great to see the stability between local validation and LB scores (even with using only <strong>1.5%</strong> of the data as <strong>validation</strong>).</p>\n<p>The model training process involved a train-validation split, wherein <strong>batches 11-660</strong> were utilized for <strong>training</strong>, and <strong>batches 1-10</strong> were allocated for <strong>validation</strong>. An ensemble of 8 models was formed using <strong>hill climbing</strong> or <strong>Nelder-Mead</strong> optimization techniques, with a submission time of approximately 3-3.5 hours. The inclusion of more models in the ensemble resulted in a better score; however, since I was submitting the best blend only a few hours before the end of the competition, there was not enough time to do inference on a bigger ensemble. The same model architectures were trained multiple times and blended.</p>\n<p>The ensemble's local validation score was <strong>0.97747</strong>, while the public and private LB scores were <strong>0.976</strong>.</p>\n<p>Input data for the model consisted of the following 9 features per event: <strong>sensor_x, sensor_y, sensor_z, time, charge, auxiliary, is_main_sensor, is_deep_veto, and is_deep_core</strong>.</p>\n<p>Four model architectures have been selected for the final submission.</p>\n<p><strong>Model 1</strong>:</p>\n<p>Validation score 1: <strong>1.0016</strong><br>\nValidation score 2: <strong>1.0005</strong><br>\nLoss function: <strong>CrossEntropyLoss</strong></p>\n<p>Data preprocessing steps:</p>\n<ul>\n<li>sensor_x, sensor_y, and sensor_z are divided by 600</li>\n<li>time is divided by 1000 and subtracted by its minimum value</li>\n<li>charge is divided by 300</li>\n</ul>\n<pre><code> (pl.LightningModule):\n     ():\n\n    self.bin_num = \n\n        self.gru = nn.GRU(, , num_layers=, dropout=, batch_first=, bidirectional=)\n        self.fc1 = nn.Sequential(nn.Linear(, ), nn.ReLU())\n        self.fc2 = nn.Linear(, bin_num*bin_num)\n\n     ():\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.gru(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        x = self.fc2(x)\n\n         x\n</code></pre>\n<p>Output bins were generated using <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">code</a> from <a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a>. </p>\n<p><strong>Models 2-4</strong>:</p>\n<p>Loss function: <strong>VonMisesFisher3DLoss</strong></p>\n<p>Data preprocessing steps:</p>\n<ul>\n<li>sensor_x, sensor_y, and sensor_z are divided by 500</li>\n<li>time is scaled by subtracting 1.0e04 and dividing by 3.0e4</li>\n<li>charge is transformed using the logarithm base 10 and then divided by 3.0</li>\n</ul>\n<p>Validation score 3: <strong>0.9847</strong><br>\nValidation score 4: <strong>0.9859</strong></p>\n<pre><code> (pl.LightningModule):\n     ():\n\n        self.bilstm = nn.LSTM(, , num_layers=, dropout=, batch_first=, bidirectional=)\n\n        self.fc1 = nn.Sequential(nn.Linear(, ), nn.ReLU())\n        self.dropout = nn.Dropout()\n        self.fc2 = nn.Linear(, )\n\n     ():\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        x = self.dropout(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=, p=) + \n        pred_x = pred[:, ] / kappa\n        pred_y = pred[:, ] / kappa\n        pred_z = pred[:, ] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=)\n\n         pred\n</code></pre>\n<p>Validation score 5: <strong>0.9872</strong><br>\nValidation score 6: <strong>0.9887</strong></p>\n<pre><code> (pl.LightningModule):\n     ():\n        ().__init__()\n\n        self.embedding = nn.Linear(, )\n        self.bilstm = nn.LSTM(, , num_layers=, dropout=, batch_first=, bidirectional=)\n\n        self.fc1 = nn.Sequential(nn.Linear(, ), nn.ReLU())\n        self.fc2 = nn.Linear(, )\n\n     ():\n\n        x = self.embedding(x)\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=, p=) + \n        pred_x = pred[:, ] / kappa\n        pred_y = pred[:, ] / kappa\n        pred_z = pred[:, ] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=)\n\n         pred \n</code></pre>\n<p>Validation score 7: <strong>0.9842</strong><br>\nValidation score 8: <strong>0.9841</strong></p>\n<pre><code> (pl.LightningModule):\n     ():\n\n        self.embedding = nn.Linear(, )\n\n        self.bilstm = nn.LSTM(, , num_layers=, dropout=, batch_first=, bidirectional=)\n\n        self.fc1 = nn.Sequential(nn.Linear(lstm_units, ), nn.ReLU())\n        self.fc2 = nn.Linear(, )\n\n     ():\n\n        batch_sizes = batch_sizes.cpu()\n\n        x = self.embedding(x)\n\n        x = pack_padded_sequence(x, batch_sizes, batch_first=, enforce_sorted=)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=)\n\n        x = x.(dim=)\n        x = x.div(batch_sizes.unsqueeze(-).cuda())\n\n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=, p=) + \n        pred_x = pred[:, ] / kappa\n        pred_y = pred[:, ] / kappa\n        pred_z = pred[:, ] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=)\n\n         pred\n</code></pre>\n<p>Hyperparameters:<br>\nOptimizer: <strong>Adam</strong><br>\nScheduler: <strong>CosineAnnealingLR</strong><br>\nBatch size: <strong>2048</strong><br>\nMax pulses: <strong>128</strong><br>\nMax LR: <strong>1e-3</strong> or <strong>5e-4</strong><br>\nMin LR: <strong>1e-6</strong><br>\nWarmup steps: <strong>2000</strong><br>\nEpochs: <strong>10-15</strong> (possibly a few extra fine tuning epochs)</p>\n<p>The DL library used for the competition was PyTorch (Lightning), which delivered better results than TensorFlow. While blending multiple models in TensorFlow, the code produced some unanticipated errors which I wasn't able to debug. It was just simpler to switch to PyTorch.</p>\n<p>I tried a few different transformer architectures, but didn't have enough time to push it through the end. I retrained a few models similar to graphnet (<a href=\"https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite\" target=\"_blank\">code</a> from <a href=\"https://www.kaggle.com/amoshuangyc\" target=\"_blank\">@amoshuangyc</a> was very useful), but it didn't give a boost to the final ensemble.</p>",
      "rawMarkdown": "The last few weeks have been intense, but in the end I am super happy that I've finally got my first solo gold medal. Great competition overall! It was great to see the stability between local validation and LB scores (even with using only **1.5%** of the data as **validation**).\n\nThe model training process involved a train-validation split, wherein **batches 11-660** were utilized for **training**, and **batches 1-10** were allocated for **validation**. An ensemble of 8 models was formed using **hill climbing** or **Nelder-Mead** optimization techniques, with a submission time of approximately 3-3.5 hours. The inclusion of more models in the ensemble resulted in a better score; however, since I was submitting the best blend only a few hours before the end of the competition, there was not enough time to do inference on a bigger ensemble. The same model architectures were trained multiple times and blended.\n\nThe ensemble's local validation score was **0.97747**, while the public and private LB scores were **0.976**.\n\nInput data for the model consisted of the following 9 features per event: **sensor_x, sensor_y, sensor_z, time, charge, auxiliary, is_main_sensor, is_deep_veto, and is_deep_core**.\n\nFour model architectures have been selected for the final submission.\n\n**Model 1**:\n\nValidation score 1: **1.0016**\nValidation score 2: **1.0005**\nLoss function: **CrossEntropyLoss**\n\nData preprocessing steps:\n- sensor_x, sensor_y, and sensor_z are divided by 600\n- time is divided by 1000 and subtracted by its minimum value\n- charge is divided by 300\n\n```python\nclass Model1(pl.LightningModule):\n    def __init__(self):\n\n\tself.bin_num = 31\n\n        self.gru = nn.GRU(9, 192, num_layers=3, dropout=0.0, batch_first=True, bidirectional=True)\n        self.fc1 = nn.Sequential(nn.Linear(384, 256), nn.ReLU())\n        self.fc2 = nn.Linear(256, bin_num*bin_num)\n\n    def forward(self, x, batch_sizes):\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.gru(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        x = self.fc2(x)\n\n        return x\n```\n\nOutput bins were generated using [code](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu) from @rsmits. \n\n      \n**Models 2-4**:\n\nLoss function: **VonMisesFisher3DLoss**\n\nData preprocessing steps:\n- sensor_x, sensor_y, and sensor_z are divided by 500\n- time is scaled by subtracting 1.0e04 and dividing by 3.0e4\n- charge is transformed using the logarithm base 10 and then divided by 3.0\n    \nValidation score 3: **0.9847**\nValidation score 4: **0.9859**\n\n```python\nclass Model2(pl.LightningModule):\n    def __init__(self):\n\n        self.bilstm = nn.LSTM(9, 256, num_layers=3, dropout=0.2, batch_first=True, bidirectional=True)\n\n        self.fc1 = nn.Sequential(nn.Linear(512, 256), nn.ReLU())\n        self.dropout = nn.Dropout(0.2)\n        self.fc2 = nn.Linear(256, 3)\n\n    def forward(self, x, batch_sizes):\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        x = self.dropout(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=1, p=2) + 1e-8\n        pred_x = pred[:, 0] / kappa\n        pred_y = pred[:, 1] / kappa\n        pred_z = pred[:, 2] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=1)\n\n        return pred\n```\n     \nValidation score 5: **0.9872**\nValidation score 6: **0.9887**\n\n```python\nclass Model3(pl.LightningModule):\n    def __init__(\n        self\n    ):\n        super().__init__()\n\n        self.embedding = nn.Linear(9, 512)\n        self.bilstm = nn.LSTM(512, 256, num_layers=3, dropout=0.0, batch_first=True, bidirectional=True)\n\n        self.fc1 = nn.Sequential(nn.Linear(512, 256), nn.ReLU())\n        self.fc2 = nn.Linear(256, 3)\n\n    def forward(self, x, batch_sizes):\n\n        x = self.embedding(x)\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=1, p=2) + 1e-8\n        pred_x = pred[:, 0] / kappa\n        pred_y = pred[:, 1] / kappa\n        pred_z = pred[:, 2] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=1)\n\n        return pred \n```\n  \nValidation score 7: **0.9842**\nValidation score 8: **0.9841**\n\n```python\nclass Model4(pl.LightningModule):\n    def __init__(self):\n\n        self.embedding = nn.Linear(9, 192)\n\n        self.bilstm = nn.LSTM(192, 96, num_layers=3, dropout=0.0, batch_first=True, bidirectional=True)\n\n        self.fc1 = nn.Sequential(nn.Linear(lstm_units, 256), nn.ReLU())\n        self.fc2 = nn.Linear(256, 3)\n\n    def forward(self, x, batch_sizes):\n\n        batch_sizes = batch_sizes.cpu()\n\n        x = self.embedding(x)\n\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=1, p=2) + 1e-8\n        pred_x = pred[:, 0] / kappa\n        pred_y = pred[:, 1] / kappa\n        pred_z = pred[:, 2] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=1)\n\n        return pred\n```\n\nHyperparameters:\nOptimizer: **Adam**\nScheduler: **CosineAnnealingLR**\nBatch size: **2048**\nMax pulses: **128**\nMax LR: **1e-3** or **5e-4**\nMin LR: **1e-6**\nWarmup steps: **2000**\nEpochs: **10-15** (possibly a few extra fine tuning epochs)\n\nThe DL library used for the competition was PyTorch (Lightning), which delivered better results than TensorFlow. While blending multiple models in TensorFlow, the code produced some unanticipated errors which I wasn't able to debug. It was just simpler to switch to PyTorch.\n\nI tried a few different transformer architectures, but didn't have enough time to push it through the end. I retrained a few models similar to graphnet ([code](https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite) from @amoshuangyc was very useful), but it didn't give a boost to the final ensemble.",
      "votes": null
    },
    {
      "id": "2228424",
      "postDate": "04/20/2023 14:40:20",
      "content": "<p>double congrats, on solo gold and also becoming GM! </p>",
      "rawMarkdown": "double congrats, on solo gold and also becoming GM!",
      "votes": null
    },
    {
      "id": "2228432",
      "postDate": "04/20/2023 14:47:01",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>! Great job on the 2nd position :)</p>",
      "rawMarkdown": "Thanks @drhabib! Great job on the 2nd position :)",
      "votes": null
    },
    {
      "id": "2228461",
      "postDate": "04/20/2023 15:10:26",
      "content": "<p>Congrats on becoming a GM <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a>! Thanks for sharing the write up!</p>",
      "rawMarkdown": "Congrats on becoming a GM @crodoc! Thanks for sharing the write up!",
      "votes": null
    },
    {
      "id": "2228471",
      "postDate": "04/20/2023 15:14:38",
      "content": "<p><a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> congratulations on becoming GM!</p>\n<p>The solution is very similar to what we did with LSTMs. Were there any additional tricks with these models, that predicted a XYZ vector? Did you predict a normalized vector? For some reason this thing didn't work for us, the score of such models was always worse, than models that predicted bins.</p>",
      "rawMarkdown": "crodoc congratulations on becoming GM!\n\nThe solution is very similar to what we did with LSTMs. Were there any additional tricks with these models, that predicted a XYZ vector? Did you predict a normalized vector? For some reason this thing didn't work for us, the score of such models was always worse, than models that predicted bins.",
      "votes": null
    },
    {
      "id": "2228478",
      "postDate": "04/20/2023 15:20:31",
      "content": "<p><a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a> did you use TF or PyTorch? I was getting a worse score in TF. No idea why.</p>\n<p>For the xyz prediction, I used it practically the same as they suggested in graphnet.</p>",
      "rawMarkdown": "manwithaflower did you use TF or PyTorch? I was getting a worse score in TF. No idea why.\n\nFor the xyz prediction, I used it practically the same as they suggested in graphnet.",
      "votes": null
    },
    {
      "id": "2228481",
      "postDate": "04/20/2023 15:22:27",
      "content": "<p>Thank you 🎉</p>",
      "rawMarkdown": "Thank you 🎉",
      "votes": null
    },
    {
      "id": "2228496",
      "postDate": "04/20/2023 15:30:06",
      "content": "<p>Yes, TF. Maybe that was the problem. Very strange. But this is so funny, when I read this yours code and hyperparams, I'm just saying \"wow, literally me 😂\". Congrats again! I'll remember that I need more perseverance in future. </p>",
      "rawMarkdown": "Yes, TF. Maybe that was the problem. Very strange. But this is so funny, when I read this yours code and hyperparams, I'm just saying \"wow, literally me 😂\". Congrats again! I'll remember that I need more perseverance in future.",
      "votes": null
    },
    {
      "id": "2228501",
      "postDate": "04/20/2023 15:33:54",
      "content": "<p>I really had a negative experience with TF in this competiton. At some point even submission inference was crashing non-stop. 🤷‍♂️</p>",
      "rawMarkdown": "I really had a negative experience with TF in this competiton. At some point even submission inference was crashing non-stop. 🤷‍♂️",
      "votes": null
    },
    {
      "id": "2228548",
      "postDate": "04/20/2023 16:04:37",
      "content": "<p>Congrats on solo gold and becoming GM <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a>! Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats on solo gold and becoming GM @crodoc! Thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "2228565",
      "postDate": "04/20/2023 16:13:33",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> ! this is just insane solo GOLD </p>",
      "rawMarkdown": "Congrats @crodoc ! this is just insane solo GOLD",
      "votes": null
    },
    {
      "id": "2228580",
      "postDate": "04/20/2023 16:23:13",
      "content": "<p>Congrats with solo gold. Amazing improvement over the last week!</p>",
      "rawMarkdown": "Congrats with solo gold. Amazing improvement over the last week!",
      "votes": null
    },
    {
      "id": "2228659",
      "postDate": "04/20/2023 17:40:40",
      "content": "<p>Hearty congratulations to you <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> for the excellent performance and a special congratulations for becoming a competitions grandmaster. </p>",
      "rawMarkdown": "Hearty congratulations to you @crodoc for the excellent performance and a special congratulations for becoming a competitions grandmaster.",
      "votes": null
    },
    {
      "id": "2228679",
      "postDate": "04/20/2023 17:55:13",
      "content": "<p>Congrats on solo gold finish!  How long did training 10 epochs x 650 batches take?</p>",
      "rawMarkdown": "Congrats on solo gold finish!  How long did training 10 epochs x 650 batches take?",
      "votes": null
    },
    {
      "id": "2228783",
      "postDate": "04/20/2023 20:09:54",
      "content": "<p>I think the smallest model was 3.5h per epoch and the largest maybe 7h. It was a lot of electricity.</p>",
      "rawMarkdown": "I think the smallest model was 3.5h per epoch and the largest maybe 7h. It was a lot of electricity.",
      "votes": null
    },
    {
      "id": "2228789",
      "postDate": "04/20/2023 20:18:36",
      "content": "<p><a href=\"https://www.kaggle.com/rusg77\" target=\"_blank\">@rusg77</a> this sub was late only by a few minutes. I started to blend models a few hours before the end (I went to do a short crossfit session instead of blending right away). This sub is 0.000001 better than your LB sub 😂😂😂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6674022%2Fb40e45ac0cae28c13cb4b06c64c57dd4%2FScreenshot%20from%202023-04-20%2022-15-50.png?generation=1682021763851108&amp;alt=media\" alt=\"\"></p>\n<p>I started to get decent results quite late in the competitions. A lot of the things I tried didn't work. Everything was much better when I switched to PyTorch.</p>",
      "rawMarkdown": "rusg77 this sub was late only by a few minutes. I started to blend models a few hours before the end (I went to do a short crossfit session instead of blending right away). This sub is 0.000001 better than your LB sub 😂😂😂\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6674022%2Fb40e45ac0cae28c13cb4b06c64c57dd4%2FScreenshot%20from%202023-04-20%2022-15-50.png?generation=1682021763851108&alt=media)\n\nI started to get decent results quite late in the competitions. A lot of the things I tried didn't work. Everything was much better when I switched to PyTorch.",
      "votes": null
    },
    {
      "id": "2228810",
      "postDate": "04/20/2023 20:42:31",
      "content": "<p>Congratulations on the solo gold and becoming Grand Master!</p>\n<p>What hardware did you use to run that: Kaggle, Colab, your own GPU, …?</p>",
      "rawMarkdown": "Congratulations on the solo gold and becoming Grand Master!\n\nWhat hardware did you use to run that: Kaggle, Colab, your own GPU, ...?",
      "votes": null
    },
    {
      "id": "2228836",
      "postDate": "04/20/2023 21:20:14",
      "content": "<p>I have 6 RTX 3090s. I also rented a few A100s in the last 2 days.</p>",
      "rawMarkdown": "I have 6 RTX 3090s. I also rented a few A100s in the last 2 days.",
      "votes": null
    },
    {
      "id": "2228864",
      "postDate": "04/20/2023 21:59:41",
      "content": "<p>Congratulations! It's very weird because I also saw many cases that the performance of the TensorFlow was worse than the PyTorch. Did you use CosineAnnealing lr schedule and were all other details of the model same in the TensorFlow and the PyTorch? </p>",
      "rawMarkdown": "Congratulations! It's very weird because I also saw many cases that the performance of the TensorFlow was worse than the PyTorch. Did you use CosineAnnealing lr schedule and were all other details of the model same in the TensorFlow and the PyTorch?",
      "votes": null
    },
    {
      "id": "2228869",
      "postDate": "04/20/2023 22:04:34",
      "content": "<p><a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> This is the first time I tried TF. Not sure if I will try it again anytime soon. Wasted too much time.</p>",
      "rawMarkdown": "junseonglee11 This is the first time I tried TF. Not sure if I will try it again anytime soon. Wasted too much time.",
      "votes": null
    },
    {
      "id": "2228907",
      "postDate": "04/20/2023 23:00:36",
      "content": "<p>Are the RTX 3090s in 6 separate cases?  Or did you find a magical way to cram them together.  I am really curious 🙄</p>",
      "rawMarkdown": "Are the RTX 3090s in 6 separate cases?  Or did you find a magical way to cram them together.  I am really curious 🙄",
      "votes": null
    },
    {
      "id": "2229297",
      "postDate": "04/21/2023 08:29:26",
      "content": "<p>2 are in 2 prebuild gamer PCs. The last 4 are in the same case. <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> helped me build the machine. You can see it in this <a href=\"https://www.linkedin.com/posts/andrijamilicevic_deeplearning-nvidia-nvidiartx-activity-6969293456661704705-T4KU?utm_source=share\" target=\"_blank\">post</a>.</p>",
      "rawMarkdown": "2 are in 2 prebuild gamer PCs. The last 4 are in the same case. @init27 helped me build the machine. You can see it in this [post](https://www.linkedin.com/posts/andrijamilicevic_deeplearning-nvidia-nvidiartx-activity-6969293456661704705-T4KU?utm_source=share).",
      "votes": null
    },
    {
      "id": "2229305",
      "postDate": "04/21/2023 08:40:28",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> :) very happy to join the club :)</p>",
      "rawMarkdown": "Thanks @ravi20076 :) very happy to join the club :)",
      "votes": null
    },
    {
      "id": "2229558",
      "postDate": "04/21/2023 13:19:20",
      "content": "<p>Wow, that's a nice build.  Too bad they discontinued the 3090 Turbos.</p>",
      "rawMarkdown": "Wow, that's a nice build.  Too bad they discontinued the 3090 Turbos.",
      "votes": null
    },
    {
      "id": "2229581",
      "postDate": "04/21/2023 13:39:34",
      "content": "<p>Yea, it wasn't easy to find them. They seemed to be best buy at time. Maybe A5000 is a decent buy now?</p>",
      "rawMarkdown": "Yea, it wasn't easy to find them. They seemed to be best buy at time. Maybe A5000 is a decent buy now?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2228424,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "04/20/2023 14:40:20",
      "content": "<p>double congrats, on solo gold and also becoming GM! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2228432,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/20/2023 14:47:01",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a>! Great job on the 2nd position :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228461,
      "author_name": "ravishah1",
      "author_url": "",
      "post_date": "04/20/2023 15:10:26",
      "content": "<p>Congrats on becoming a GM <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a>! Thanks for sharing the write up!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228481,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/20/2023 15:22:27",
          "content": "<p>Thank you 🎉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228471,
      "author_name": "manwithaflower",
      "author_url": "",
      "post_date": "04/20/2023 15:14:38",
      "content": "<p><a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> congratulations on becoming GM!</p>\n<p>The solution is very similar to what we did with LSTMs. Were there any additional tricks with these models, that predicted a XYZ vector? Did you predict a normalized vector? For some reason this thing didn't work for us, the score of such models was always worse, than models that predicted bins.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228478,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/20/2023 15:20:31",
          "content": "<p><a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a> did you use TF or PyTorch? I was getting a worse score in TF. No idea why.</p>\n<p>For the xyz prediction, I used it practically the same as they suggested in graphnet.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2228496,
              "author_name": "manwithaflower",
              "author_url": "",
              "post_date": "04/20/2023 15:30:06",
              "content": "<p>Yes, TF. Maybe that was the problem. Very strange. But this is so funny, when I read this yours code and hyperparams, I'm just saying \"wow, literally me 😂\". Congrats again! I'll remember that I need more perseverance in future. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2228501,
                  "author_name": "crodoc",
                  "author_url": "",
                  "post_date": "04/20/2023 15:33:54",
                  "content": "<p>I really had a negative experience with TF in this competiton. At some point even submission inference was crashing non-stop. 🤷‍♂️</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2228548,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "04/20/2023 16:04:37",
      "content": "<p>Congrats on solo gold and becoming GM <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a>! Thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228565,
      "author_name": "vinayak121",
      "author_url": "",
      "post_date": "04/20/2023 16:13:33",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> ! this is just insane solo GOLD </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228580,
      "author_name": "rusg77",
      "author_url": "",
      "post_date": "04/20/2023 16:23:13",
      "content": "<p>Congrats with solo gold. Amazing improvement over the last week!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228789,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/20/2023 20:18:36",
          "content": "<p><a href=\"https://www.kaggle.com/rusg77\" target=\"_blank\">@rusg77</a> this sub was late only by a few minutes. I started to blend models a few hours before the end (I went to do a short crossfit session instead of blending right away). This sub is 0.000001 better than your LB sub 😂😂😂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6674022%2Fb40e45ac0cae28c13cb4b06c64c57dd4%2FScreenshot%20from%202023-04-20%2022-15-50.png?generation=1682021763851108&amp;alt=media\" alt=\"\"></p>\n<p>I started to get decent results quite late in the competitions. A lot of the things I tried didn't work. Everything was much better when I switched to PyTorch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228659,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/20/2023 17:40:40",
      "content": "<p>Hearty congratulations to you <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> for the excellent performance and a special congratulations for becoming a competitions grandmaster. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2229305,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/21/2023 08:40:28",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> :) very happy to join the club :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228679,
      "author_name": "solverworld",
      "author_url": "",
      "post_date": "04/20/2023 17:55:13",
      "content": "<p>Congrats on solo gold finish!  How long did training 10 epochs x 650 batches take?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228783,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/20/2023 20:09:54",
          "content": "<p>I think the smallest model was 3.5h per epoch and the largest maybe 7h. It was a lot of electricity.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2228810,
              "author_name": "vialactea",
              "author_url": "",
              "post_date": "04/20/2023 20:42:31",
              "content": "<p>Congratulations on the solo gold and becoming Grand Master!</p>\n<p>What hardware did you use to run that: Kaggle, Colab, your own GPU, …?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2228836,
                  "author_name": "crodoc",
                  "author_url": "",
                  "post_date": "04/20/2023 21:20:14",
                  "content": "<p>I have 6 RTX 3090s. I also rented a few A100s in the last 2 days.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2228907,
                      "author_name": "solverworld",
                      "author_url": "",
                      "post_date": "04/20/2023 23:00:36",
                      "content": "<p>Are the RTX 3090s in 6 separate cases?  Or did you find a magical way to cram them together.  I am really curious 🙄</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2229297,
                          "author_name": "crodoc",
                          "author_url": "",
                          "post_date": "04/21/2023 08:29:26",
                          "content": "<p>2 are in 2 prebuild gamer PCs. The last 4 are in the same case. <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> helped me build the machine. You can see it in this <a href=\"https://www.linkedin.com/posts/andrijamilicevic_deeplearning-nvidia-nvidiartx-activity-6969293456661704705-T4KU?utm_source=share\" target=\"_blank\">post</a>.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2229558,
                              "author_name": "solverworld",
                              "author_url": "",
                              "post_date": "04/21/2023 13:19:20",
                              "content": "<p>Wow, that's a nice build.  Too bad they discontinued the 3090 Turbos.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2229581,
                                  "author_name": "crodoc",
                                  "author_url": "",
                                  "post_date": "04/21/2023 13:39:34",
                                  "content": "<p>Yea, it wasn't easy to find them. They seemed to be best buy at time. Maybe A5000 is a decent buy now?</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2228864,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "04/20/2023 21:59:41",
      "content": "<p>Congratulations! It's very weird because I also saw many cases that the performance of the TensorFlow was worse than the PyTorch. Did you use CosineAnnealing lr schedule and were all other details of the model same in the TensorFlow and the PyTorch? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2228869,
          "author_name": "crodoc",
          "author_url": "",
          "post_date": "04/20/2023 22:04:34",
          "content": "<p><a href=\"https://www.kaggle.com/junseonglee11\" target=\"_blank\">@junseonglee11</a> This is the first time I tried TF. Not sure if I will try it again anytime soon. Wasted too much time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2228359": "The last few weeks have been intense, but in the end I am super happy that I've finally got my first solo gold medal. Great competition overall! It was great to see the stability between local validation and LB scores (even with using only **1.5%** of the data as **validation**).\n\nThe model training process involved a train-validation split, wherein **batches 11-660** were utilized for **training**, and **batches 1-10** were allocated for **validation**. An ensemble of 8 models was formed using **hill climbing** or **Nelder-Mead** optimization techniques, with a submission time of approximately 3-3.5 hours. The inclusion of more models in the ensemble resulted in a better score; however, since I was submitting the best blend only a few hours before the end of the competition, there was not enough time to do inference on a bigger ensemble. The same model architectures were trained multiple times and blended.\n\nThe ensemble's local validation score was **0.97747**, while the public and private LB scores were **0.976**.\n\nInput data for the model consisted of the following 9 features per event: **sensor_x, sensor_y, sensor_z, time, charge, auxiliary, is_main_sensor, is_deep_veto, and is_deep_core**.\n\nFour model architectures have been selected for the final submission.\n\n**Model 1**:\n\nValidation score 1: **1.0016**\nValidation score 2: **1.0005**\nLoss function: **CrossEntropyLoss**\n\nData preprocessing steps:\n- sensor_x, sensor_y, and sensor_z are divided by 600\n- time is divided by 1000 and subtracted by its minimum value\n- charge is divided by 300\n\n```python\nclass Model1(pl.LightningModule):\n    def __init__(self):\n\n\tself.bin_num = 31\n\n        self.gru = nn.GRU(9, 192, num_layers=3, dropout=0.0, batch_first=True, bidirectional=True)\n        self.fc1 = nn.Sequential(nn.Linear(384, 256), nn.ReLU())\n        self.fc2 = nn.Linear(256, bin_num*bin_num)\n\n    def forward(self, x, batch_sizes):\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.gru(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        x = self.fc2(x)\n\n        return x\n```\n\nOutput bins were generated using [code](https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu) from @rsmits. \n\n      \n**Models 2-4**:\n\nLoss function: **VonMisesFisher3DLoss**\n\nData preprocessing steps:\n- sensor_x, sensor_y, and sensor_z are divided by 500\n- time is scaled by subtracting 1.0e04 and dividing by 3.0e4\n- charge is transformed using the logarithm base 10 and then divided by 3.0\n    \nValidation score 3: **0.9847**\nValidation score 4: **0.9859**\n\n```python\nclass Model2(pl.LightningModule):\n    def __init__(self):\n\n        self.bilstm = nn.LSTM(9, 256, num_layers=3, dropout=0.2, batch_first=True, bidirectional=True)\n\n        self.fc1 = nn.Sequential(nn.Linear(512, 256), nn.ReLU())\n        self.dropout = nn.Dropout(0.2)\n        self.fc2 = nn.Linear(256, 3)\n\n    def forward(self, x, batch_sizes):\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        x = self.dropout(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=1, p=2) + 1e-8\n        pred_x = pred[:, 0] / kappa\n        pred_y = pred[:, 1] / kappa\n        pred_z = pred[:, 2] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=1)\n\n        return pred\n```\n     \nValidation score 5: **0.9872**\nValidation score 6: **0.9887**\n\n```python\nclass Model3(pl.LightningModule):\n    def __init__(\n        self\n    ):\n        super().__init__()\n\n        self.embedding = nn.Linear(9, 512)\n        self.bilstm = nn.LSTM(512, 256, num_layers=3, dropout=0.0, batch_first=True, bidirectional=True)\n\n        self.fc1 = nn.Sequential(nn.Linear(512, 256), nn.ReLU())\n        self.fc2 = nn.Linear(256, 3)\n\n    def forward(self, x, batch_sizes):\n\n        x = self.embedding(x)\n\n        batch_sizes = batch_sizes.cpu()\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=1, p=2) + 1e-8\n        pred_x = pred[:, 0] / kappa\n        pred_y = pred[:, 1] / kappa\n        pred_z = pred[:, 2] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=1)\n\n        return pred \n```\n  \nValidation score 7: **0.9842**\nValidation score 8: **0.9841**\n\n```python\nclass Model4(pl.LightningModule):\n    def __init__(self):\n\n        self.embedding = nn.Linear(9, 192)\n\n        self.bilstm = nn.LSTM(192, 96, num_layers=3, dropout=0.0, batch_first=True, bidirectional=True)\n\n        self.fc1 = nn.Sequential(nn.Linear(lstm_units, 256), nn.ReLU())\n        self.fc2 = nn.Linear(256, 3)\n\n    def forward(self, x, batch_sizes):\n\n        batch_sizes = batch_sizes.cpu()\n\n        x = self.embedding(x)\n\n        x = pack_padded_sequence(x, batch_sizes, batch_first=True, enforce_sorted=False)\n        x, _ = self.bilstm(x)\n        x, _ = pad_packed_sequence(x, batch_first=True)\n\n        x = x.sum(dim=1)\n        x = x.div(batch_sizes.unsqueeze(-1).cuda())\n        \n        x = self.fc1(x)\n        pred = self.fc2(x)\n\n        kappa = pred.norm(dim=1, p=2) + 1e-8\n        pred_x = pred[:, 0] / kappa\n        pred_y = pred[:, 1] / kappa\n        pred_z = pred[:, 2] / kappa\n        pred = torch.stack([pred_x, pred_y, pred_z, kappa], dim=1)\n\n        return pred\n```\n\nHyperparameters:\nOptimizer: **Adam**\nScheduler: **CosineAnnealingLR**\nBatch size: **2048**\nMax pulses: **128**\nMax LR: **1e-3** or **5e-4**\nMin LR: **1e-6**\nWarmup steps: **2000**\nEpochs: **10-15** (possibly a few extra fine tuning epochs)\n\nThe DL library used for the competition was PyTorch (Lightning), which delivered better results than TensorFlow. While blending multiple models in TensorFlow, the code produced some unanticipated errors which I wasn't able to debug. It was just simpler to switch to PyTorch.\n\nI tried a few different transformer architectures, but didn't have enough time to push it through the end. I retrained a few models similar to graphnet ([code](https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite) from @amoshuangyc was very useful), but it didn't give a boost to the final ensemble.",
    "2228424": "double congrats, on solo gold and also becoming GM!",
    "2228432": "Thanks @drhabib! Great job on the 2nd position :)",
    "2228461": "Congrats on becoming a GM @crodoc! Thanks for sharing the write up!",
    "2228471": "crodoc congratulations on becoming GM!\n\nThe solution is very similar to what we did with LSTMs. Were there any additional tricks with these models, that predicted a XYZ vector? Did you predict a normalized vector? For some reason this thing didn't work for us, the score of such models was always worse, than models that predicted bins.",
    "2228478": "manwithaflower did you use TF or PyTorch? I was getting a worse score in TF. No idea why.\n\nFor the xyz prediction, I used it practically the same as they suggested in graphnet.",
    "2228481": "Thank you 🎉",
    "2228496": "Yes, TF. Maybe that was the problem. Very strange. But this is so funny, when I read this yours code and hyperparams, I'm just saying \"wow, literally me 😂\". Congrats again! I'll remember that I need more perseverance in future.",
    "2228501": "I really had a negative experience with TF in this competiton. At some point even submission inference was crashing non-stop. 🤷‍♂️",
    "2228548": "Congrats on solo gold and becoming GM @crodoc! Thanks for sharing your solution.",
    "2228565": "Congrats @crodoc ! this is just insane solo GOLD",
    "2228580": "Congrats with solo gold. Amazing improvement over the last week!",
    "2228659": "Hearty congratulations to you @crodoc for the excellent performance and a special congratulations for becoming a competitions grandmaster.",
    "2228679": "Congrats on solo gold finish!  How long did training 10 epochs x 650 batches take?",
    "2228783": "I think the smallest model was 3.5h per epoch and the largest maybe 7h. It was a lot of electricity.",
    "2228789": "rusg77 this sub was late only by a few minutes. I started to blend models a few hours before the end (I went to do a short crossfit session instead of blending right away). This sub is 0.000001 better than your LB sub 😂😂😂\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6674022%2Fb40e45ac0cae28c13cb4b06c64c57dd4%2FScreenshot%20from%202023-04-20%2022-15-50.png?generation=1682021763851108&alt=media)\n\nI started to get decent results quite late in the competitions. A lot of the things I tried didn't work. Everything was much better when I switched to PyTorch.",
    "2228810": "Congratulations on the solo gold and becoming Grand Master!\n\nWhat hardware did you use to run that: Kaggle, Colab, your own GPU, ...?",
    "2228836": "I have 6 RTX 3090s. I also rented a few A100s in the last 2 days.",
    "2228864": "Congratulations! It's very weird because I also saw many cases that the performance of the TensorFlow was worse than the PyTorch. Did you use CosineAnnealing lr schedule and were all other details of the model same in the TensorFlow and the PyTorch?",
    "2228869": "junseonglee11 This is the first time I tried TF. Not sure if I will try it again anytime soon. Wasted too much time.",
    "2228907": "Are the RTX 3090s in 6 separate cases?  Or did you find a magical way to cram them together.  I am really curious 🙄",
    "2229297": "2 are in 2 prebuild gamer PCs. The last 4 are in the same case. @init27 helped me build the machine. You can see it in this [post](https://www.linkedin.com/posts/andrijamilicevic_deeplearning-nvidia-nvidiartx-activity-6969293456661704705-T4KU?utm_source=share).",
    "2229305": "Thanks @ravi20076 :) very happy to join the club :)",
    "2229558": "Wow, that's a nice build.  Too bad they discontinued the 3090 Turbos.",
    "2229581": "Yea, it wasn't easy to find them. They seemed to be best buy at time. Maybe A5000 is a decent buy now?"
  },
  "source": "meta"
}