{
  "id": 402880,
  "title": "6th place solution : GNN and Transformer part",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402880",
  "author_name": "",
  "post_date": "2023-04-20T04:05:26.762252900Z",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is a part of the 6th solution (LSTM+GNN+Transformer+group ensembling). For the main parts, please refer to the following link (TBD).</p>\n<p><strong><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153</a></strong></p>\n<p>As a newcomer to GNN, I focused on fine-tuning the baseline GraphNet solution without making significant changes.</p>\n<h1>baseline GraphNet solution finetuning (atam1231 part)</h1>\n<h2>Environment:</h2>\n<p>Ubuntu 16.04/I7-9900/64GB/NVIDIA GTX3090<br>\nPyTorch 1.13.1/CUDA 11.6</p>\n<h2>Codebase and settings:</h2>\n<ul>\n<li>Clear codebase obtained from <a href=\"https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite\" target=\"_blank\">https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7\" target=\"_blank\">https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7</a></li>\n<li>Training set: batches 1-650 except for batch 51</li>\n<li>Validation set: batch 51</li>\n<li>Optimizer: Adam with learning rate of 1e-4</li>\n<li>Scheduler: CosineAnnealingLR with a learning rate range of 1e-4 to 1e-7, t_max=4, and iteration of 25000</li>\n<li>Epoch: Interrupted training manually if the validation loss stopped decreasing (in the case of all batches, about 0 to 2 epochs).</li>\n</ul>\n<h2>Train Steps:</h2>\n<p>Step 0: Baseline solution as pretrained weights (<strong>score 1.018</strong>)<br>\nStep 1: Basic fine-tuning with 50 batches (100-150) with augmentation techniques such as random pulse and edge dropout (p=0.05) and adding noise to charge value (<strong>score 1.012</strong>).<br>\nStep 2: Basic fine-tuning with 500 batches (100-600) (<strong>score 1.004</strong>).<br>\nStep 3: Basic fine-tuning with all batches (1-650) (<strong>score 1.001</strong>).<br>\nStep 4: Gradient accumulation x20, which is essential when training from scratch (<strong>score 0.998</strong>).<br>\nStep 5: Fixing feature:pulse_len by using the original pulse length as a feature when pulse_len &gt; max_pulses (<strong>score 0.995</strong>).<br>\nStep 6: Adding MSE loss (<strong>score 0.993</strong>).<br>\nStep 7: Ensemble 7 models with minor differences in their settings (<strong>score 0.989</strong>).</p>\n<h1>transformer (outrunner part)</h1>\n<p>I trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.<br>\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).<br>\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. <br>\nEnsemble with a early checkpoint gives 0.983/0.984 (public/private).</p>\n<p><a href=\"url\" target=\"_blank\"></a></p>",
  "messages": [
    {
      "id": "2227839",
      "postDate": "04/20/2023 04:05:26",
      "content": "<p>This is a part of the 6th solution (LSTM+GNN+Transformer+group ensembling). For the main parts, please refer to the following link (TBD).</p>\n<p><strong><a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153</a></strong></p>\n<p>As a newcomer to GNN, I focused on fine-tuning the baseline GraphNet solution without making significant changes.</p>\n<h1>baseline GraphNet solution finetuning (atam1231 part)</h1>\n<h2>Environment:</h2>\n<p>Ubuntu 16.04/I7-9900/64GB/NVIDIA GTX3090<br>\nPyTorch 1.13.1/CUDA 11.6</p>\n<h2>Codebase and settings:</h2>\n<ul>\n<li>Clear codebase obtained from <a href=\"https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite\" target=\"_blank\">https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite</a></li>\n<li>inference code: <a href=\"https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7\" target=\"_blank\">https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7</a></li>\n<li>Training set: batches 1-650 except for batch 51</li>\n<li>Validation set: batch 51</li>\n<li>Optimizer: Adam with learning rate of 1e-4</li>\n<li>Scheduler: CosineAnnealingLR with a learning rate range of 1e-4 to 1e-7, t_max=4, and iteration of 25000</li>\n<li>Epoch: Interrupted training manually if the validation loss stopped decreasing (in the case of all batches, about 0 to 2 epochs).</li>\n</ul>\n<h2>Train Steps:</h2>\n<p>Step 0: Baseline solution as pretrained weights (<strong>score 1.018</strong>)<br>\nStep 1: Basic fine-tuning with 50 batches (100-150) with augmentation techniques such as random pulse and edge dropout (p=0.05) and adding noise to charge value (<strong>score 1.012</strong>).<br>\nStep 2: Basic fine-tuning with 500 batches (100-600) (<strong>score 1.004</strong>).<br>\nStep 3: Basic fine-tuning with all batches (1-650) (<strong>score 1.001</strong>).<br>\nStep 4: Gradient accumulation x20, which is essential when training from scratch (<strong>score 0.998</strong>).<br>\nStep 5: Fixing feature:pulse_len by using the original pulse length as a feature when pulse_len &gt; max_pulses (<strong>score 0.995</strong>).<br>\nStep 6: Adding MSE loss (<strong>score 0.993</strong>).<br>\nStep 7: Ensemble 7 models with minor differences in their settings (<strong>score 0.989</strong>).</p>\n<h1>transformer (outrunner part)</h1>\n<p>I trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.<br>\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).<br>\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. <br>\nEnsemble with a early checkpoint gives 0.983/0.984 (public/private).</p>\n<p><a href=\"url\" target=\"_blank\"></a></p>",
      "rawMarkdown": "This is a part of the 6th solution (LSTM+GNN+Transformer+group ensembling). For the main parts, please refer to the following link (TBD).\n\n**[https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153)**\n\nAs a newcomer to GNN, I focused on fine-tuning the baseline GraphNet solution without making significant changes.\n\n\n# baseline GraphNet solution finetuning (atam1231 part)\n\n## Environment:\n\nUbuntu 16.04/I7-9900/64GB/NVIDIA GTX3090\nPyTorch 1.13.1/CUDA 11.6\n## Codebase and settings:\n\n* Clear codebase obtained from https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite\n* inference code: [https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7](https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7)\n* Training set: batches 1-650 except for batch 51\n* Validation set: batch 51\n* Optimizer: Adam with learning rate of 1e-4\n* Scheduler: CosineAnnealingLR with a learning rate range of 1e-4 to 1e-7, t_max=4, and iteration of 25000\n* Epoch: Interrupted training manually if the validation loss stopped decreasing (in the case of all batches, about 0 to 2 epochs).\n\n## Train Steps:\n\nStep 0: Baseline solution as pretrained weights (**score 1.018**)\nStep 1: Basic fine-tuning with 50 batches (100-150) with augmentation techniques such as random pulse and edge dropout (p=0.05) and adding noise to charge value (**score 1.012**).\nStep 2: Basic fine-tuning with 500 batches (100-600) (**score 1.004**).\nStep 3: Basic fine-tuning with all batches (1-650) (**score 1.001**).\nStep 4: Gradient accumulation x20, which is essential when training from scratch (**score 0.998**).\nStep 5: Fixing feature:pulse_len by using the original pulse length as a feature when pulse_len > max_pulses (**score 0.995**).\nStep 6: Adding MSE loss (**score 0.993**).\nStep 7: Ensemble 7 models with minor differences in their settings (**score 0.989**).\n\n\n#  transformer (outrunner part)\n\nI trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. \nEnsemble with a early checkpoint gives 0.983/0.984 (public/private).\n\n\n[](url)",
      "votes": null
    },
    {
      "id": "2227949",
      "postDate": "04/20/2023 06:21:31",
      "content": "<h2>6th place solution : transformer (outrunner part)</h2>\n<p>I trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.<br>\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).<br>\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. Ensemble with a early checkpoint gives 0.983/0.984 (public/private).</p>",
      "rawMarkdown": "## 6th place solution : transformer (outrunner part)\nI trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. Ensemble with a early checkpoint gives 0.983/0.984 (public/private).",
      "votes": null
    },
    {
      "id": "2228418",
      "postDate": "04/20/2023 14:35:11",
      "content": "<p><a href=\"https://www.kaggle.com/outrunner\" target=\"_blank\">@outrunner</a> congratulations!<br>\nAre there any tricks with transformer architecture or did you use just the classic one?</p>",
      "rawMarkdown": "outrunner congratulations!\nAre there any tricks with transformer architecture or did you use just the classic one?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227949,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "04/20/2023 06:21:31",
      "content": "<h2>6th place solution : transformer (outrunner part)</h2>\n<p>I trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.<br>\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).<br>\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. Ensemble with a early checkpoint gives 0.983/0.984 (public/private).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228418,
          "author_name": "manwithaflower",
          "author_url": "",
          "post_date": "04/20/2023 14:35:11",
          "content": "<p><a href=\"https://www.kaggle.com/outrunner\" target=\"_blank\">@outrunner</a> congratulations!<br>\nAre there any tricks with transformer architecture or did you use just the classic one?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2227839": "This is a part of the 6th solution (LSTM+GNN+Transformer+group ensembling). For the main parts, please refer to the following link (TBD).\n\n**[https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/403153)**\n\nAs a newcomer to GNN, I focused on fine-tuning the baseline GraphNet solution without making significant changes.\n\n\n# baseline GraphNet solution finetuning (atam1231 part)\n\n## Environment:\n\nUbuntu 16.04/I7-9900/64GB/NVIDIA GTX3090\nPyTorch 1.13.1/CUDA 11.6\n## Codebase and settings:\n\n* Clear codebase obtained from https://www.kaggle.com/code/amoshuangyc/icecube-gnn-baseline-rewrite\n* inference code: [https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7](https://www.kaggle.com/code/atom1231/6th-graphnet-baseline-finetuning-ensemble-x7)\n* Training set: batches 1-650 except for batch 51\n* Validation set: batch 51\n* Optimizer: Adam with learning rate of 1e-4\n* Scheduler: CosineAnnealingLR with a learning rate range of 1e-4 to 1e-7, t_max=4, and iteration of 25000\n* Epoch: Interrupted training manually if the validation loss stopped decreasing (in the case of all batches, about 0 to 2 epochs).\n\n## Train Steps:\n\nStep 0: Baseline solution as pretrained weights (**score 1.018**)\nStep 1: Basic fine-tuning with 50 batches (100-150) with augmentation techniques such as random pulse and edge dropout (p=0.05) and adding noise to charge value (**score 1.012**).\nStep 2: Basic fine-tuning with 500 batches (100-600) (**score 1.004**).\nStep 3: Basic fine-tuning with all batches (1-650) (**score 1.001**).\nStep 4: Gradient accumulation x20, which is essential when training from scratch (**score 0.998**).\nStep 5: Fixing feature:pulse_len by using the original pulse length as a feature when pulse_len > max_pulses (**score 0.995**).\nStep 6: Adding MSE loss (**score 0.993**).\nStep 7: Ensemble 7 models with minor differences in their settings (**score 0.989**).\n\n\n#  transformer (outrunner part)\n\nI trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. \nEnsemble with a early checkpoint gives 0.983/0.984 (public/private).\n\n\n[](url)",
    "2227949": "## 6th place solution : transformer (outrunner part)\nI trained a transformer with 160 pulses, 4 layers, 128 dims, 4 heads. Batch size 320 on a RTX 2080 Ti.\n2 prediction heads: xyz (von Mises-Fisher) and angle (CE).\nSingle model up to 1000 pulses gives 0.985/0.985, blend with flip-z(also trained) TTA gives 0.984/0.984. Ensemble with a early checkpoint gives 0.983/0.984 (public/private).",
    "2228418": "outrunner congratulations!\nAre there any tricks with transformer architecture or did you use just the classic one?"
  },
  "source": "meta"
}