{
  "id": 383546,
  "title": "Early Sharing Prize - 1.046 Starter Kit",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/383546",
  "author_name": "",
  "post_date": "2023-02-04T07:34:27.108278300Z",
  "votes": 33,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>My 1.046 solution is now public. </p>\n<p><a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046/notebook\" target=\"_blank\">https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046/notebook</a></p>\n<p>For code competitions, I actually&nbsp;use .py files rather than notebooks as I just copy and paste relevant parts of my codebase into a single submission&nbsp;script. This means that it is very light on comments so I will summarise them here.</p>\n<ul>\n<li>I used the <a href=\"https://arxiv.org/abs/2209.03042\" target=\"_blank\">DynEdge</a> baseline from Graphnet as already shared by Rasmus <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524\" target=\"_blank\">here</a>. I added some batch norm layers and switched out the activation function</li>\n<li>I trained on 5% of the data (roughly speaking 33 batches). Validation score was 1.048.</li>\n<li>I used the ice transparency data from <a href=\"https://arxiv.org/abs/1301.5361\" target=\"_blank\">here</a></li>\n<li>I also used the following loss function: Von-Mises Fisher for azimuth, L1 for zenith, and Cosine similarity</li>\n<li>A 180-degree TTA about the z-axis also brings some gain</li>\n</ul>",
  "messages": [
    {
      "id": "2128930",
      "postDate": "02/04/2023 07:34:27",
      "content": "<p>Hi all,</p>\n<p>My 1.046 solution is now public. </p>\n<p><a href=\"https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046/notebook\" target=\"_blank\">https://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046/notebook</a></p>\n<p>For code competitions, I actually&nbsp;use .py files rather than notebooks as I just copy and paste relevant parts of my codebase into a single submission&nbsp;script. This means that it is very light on comments so I will summarise them here.</p>\n<ul>\n<li>I used the <a href=\"https://arxiv.org/abs/2209.03042\" target=\"_blank\">DynEdge</a> baseline from Graphnet as already shared by Rasmus <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524\" target=\"_blank\">here</a>. I added some batch norm layers and switched out the activation function</li>\n<li>I trained on 5% of the data (roughly speaking 33 batches). Validation score was 1.048.</li>\n<li>I used the ice transparency data from <a href=\"https://arxiv.org/abs/1301.5361\" target=\"_blank\">here</a></li>\n<li>I also used the following loss function: Von-Mises Fisher for azimuth, L1 for zenith, and Cosine similarity</li>\n<li>A 180-degree TTA about the z-axis also brings some gain</li>\n</ul>",
      "rawMarkdown": "Hi all,\n\nMy 1.046 solution is now public. \n\nhttps://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046/notebook\n\nFor code competitions, I actually use .py files rather than notebooks as I just copy and paste relevant parts of my codebase into a single submission script. This means that it is very light on comments so I will summarise them here.\n\n- I used the [DynEdge](https://arxiv.org/abs/2209.03042) baseline from Graphnet as already shared by Rasmus [here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524). I added some batch norm layers and switched out the activation function\n- I trained on 5% of the data (roughly speaking 33 batches). Validation score was 1.048.\n- I used the ice transparency data from [here](https://arxiv.org/abs/1301.5361)\n- I also used the following loss function: Von-Mises Fisher for azimuth, L1 for zenith, and Cosine similarity\n- A 180-degree TTA about the z-axis also brings some gain",
      "votes": null
    },
    {
      "id": "2129721",
      "postDate": "02/04/2023 20:15:41",
      "content": "<p>Thanks for this! Can you share your model training code?</p>",
      "rawMarkdown": "Thanks for this! Can you share your model training code?",
      "votes": null
    },
    {
      "id": "2130463",
      "postDate": "02/05/2023 13:23:12",
      "content": "<p>Very nice, will be interesting to see how a GAT or other flavours of GNN's perform :)</p>",
      "rawMarkdown": "Very nice, will be interesting to see how a GAT or other flavours of GNN's perform :)",
      "votes": null
    },
    {
      "id": "2130693",
      "postDate": "02/05/2023 16:17:00",
      "content": "<p>I use PyTorch Lightning, so you just call <code>pl.Trainer</code> with the model</p>",
      "rawMarkdown": "I use PyTorch Lightning, so you just call `pl.Trainer` with the model",
      "votes": null
    },
    {
      "id": "2130980",
      "postDate": "02/05/2023 20:12:29",
      "content": "<p>Maybe it's easier than I realize, but I was wondering precisely how to create dataset from \"roughly speaking 33 batches\"?</p>\n<p>Can you simply pass in a single python list containing 33 IceCubeSubmissionDataset objects, or something similarly trivial?</p>\n<p>For both your code and the starter notebook from competition hosts, it's not entirely trivial (to my limited understanding and familiarity with PyTorch etc) to train on larger amounts of data. If attempting everything in-memory on Kaggle, their version might run out of memory if attempting a single larger 30+ batch size conversion to a single sqlite db (I think?). Your version avoids the conversion to sqlite, but doesn't have a plug and play 'how to do a larger dataset' example code.</p>",
      "rawMarkdown": "Maybe it's easier than I realize, but I was wondering precisely how to create dataset from \"roughly speaking 33 batches\"?\n\nCan you simply pass in a single python list containing 33 IceCubeSubmissionDataset objects, or something similarly trivial?\n\nFor both your code and the starter notebook from competition hosts, it's not entirely trivial (to my limited understanding and familiarity with PyTorch etc) to train on larger amounts of data. If attempting everything in-memory on Kaggle, their version might run out of memory if attempting a single larger 30+ batch size conversion to a single sqlite db (I think?). Your version avoids the conversion to sqlite, but doesn't have a plug and play 'how to do a larger dataset' example code.",
      "votes": null
    },
    {
      "id": "2131072",
      "postDate": "02/05/2023 21:12:58",
      "content": "<p>actually, +1 to this if you like to give some light <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> <br>\nhave you used also a Parquet dataset for training?  do you recall the RAM requirements approx? </p>\n<p>BTW <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> <br>\nI have tried to modify the <code>ParquetSubmissionDataset()</code> to a suitable train dataset (I'm able to concat multiple batches but facing OOM errors for &gt;5 batches in kaggle kernels) <br>\napart from this for training there are some attributes that missing from the <code>Data</code> (torch.geometric.data) object (or I can't understand the code well yet), i.e. <br>\na) create the edges with <code>KNNGraphBuilder()</code> and b) add the labels (y)<br>\nIf I make any progress I let you know.. </p>",
      "rawMarkdown": "actually, +1 to this if you like to give some light @anjum48 \nhave you used also a Parquet dataset for training?  do you recall the RAM requirements approx? \n\nBTW @roberthatch \nI have tried to modify the `ParquetSubmissionDataset()` to a suitable train dataset (I'm able to concat multiple batches but facing OOM errors for >5 batches in kaggle kernels) \napart from this for training there are some attributes that missing from the `Data` (torch.geometric.data) object (or I can't understand the code well yet), i.e. \na) create the edges with `KNNGraphBuilder()` and b) add the labels (y)\nIf I make any progress I let you know..",
      "votes": null
    },
    {
      "id": "2131611",
      "postDate": "02/06/2023 08:29:02",
      "content": "<p>Ah yes, I forgot about this.</p>\n<p>I am training on my local machine and have created <code>.pt</code> files for each <code>Data</code> object using <code>torch.save</code> beforehand using the output of the <code>prepare_samples()</code> function - everything else is the same, other than the <code>y</code> values which are also saved in the <code>Data</code> object. I then set <code>limit_train_batches=0.05</code> in <code>Trainer</code>.</p>\n<p>It was quite tricky doing this since even with a large drive, I hit the inode limit for my partition, so I had to split the files over 2 drives. </p>\n<p>I suspect that doing it in the way the hosts have done it in the baseline code might be better though. I remember seeing some dataset related code specific to this competition recently added to the Graphnet repo</p>",
      "rawMarkdown": "Ah yes, I forgot about this.\n\nI am training on my local machine and have created `.pt` files for each `Data` object using `torch.save` beforehand using the output of the `prepare_samples()` function - everything else is the same, other than the `y` values which are also saved in the `Data` object. I then set `limit_train_batches=0.05` in `Trainer`.\n\nIt was quite tricky doing this since even with a large drive, I hit the inode limit for my partition, so I had to split the files over 2 drives. \n\nI suspect that doing it in the way the hosts have done it in the baseline code might be better though. I remember seeing some dataset related code specific to this competition recently added to the Graphnet repo",
      "votes": null
    },
    {
      "id": "2132825",
      "postDate": "02/07/2023 04:04:17",
      "content": "<p>Hmm interesting. Thanks, I guess I'll see if I can follow that and get something working. Though I'm probably gonna just make sure I have single batch training working first on each model</p>",
      "rawMarkdown": "Hmm interesting. Thanks, I guess I'll see if I can follow that and get something working. Though I'm probably gonna just make sure I have single batch training working first on each model",
      "votes": null
    },
    {
      "id": "2133278",
      "postDate": "02/07/2023 10:42:31",
      "content": "<p>Yes, this is what I'm trying now as well, to find an efficient training method (in terms of ram, hdspace, epoch time) for 1 batch and then move to a local VM.. loading parquet files is too slow so maybe we should try what anjum suggested above.  </p>",
      "rawMarkdown": "Yes, this is what I'm trying now as well, to find an efficient training method (in terms of ram, hdspace, epoch time) for 1 batch and then move to a local VM.. loading parquet files is too slow so maybe we should try what anjum suggested above.",
      "votes": null
    },
    {
      "id": "2133289",
      "postDate": "02/07/2023 10:47:29",
      "content": "<blockquote>\n  <p>files for each Data object using torch.save beforehand using the output of the prepare_samples() function</p>\n</blockquote>\n<p>prepare_samples() ? </p>\n<blockquote>\n  <p>so I had to split the files over 2 drives.</p>\n</blockquote>\n<p>so I guess we speak for &gt; 1TB ? could you also give us a rough estimate of training epoch time?</p>",
      "rawMarkdown": "> files for each Data object using torch.save beforehand using the output of the prepare_samples() function\n\nprepare_samples() ? \n\n> so I had to split the files over 2 drives.\n\nso I guess we speak for > 1TB ? could you also give us a rough estimate of training epoch time?",
      "votes": null
    },
    {
      "id": "2134937",
      "postDate": "02/08/2023 10:53:50",
      "content": "<p>Sorry I meant to say <code>prepare_sensors()</code> after joining with the event data and then saving the .pt file. Total size is just under 1TB but with the large number of small files, you may hit the inode limit before your storage limit if you are using a Linux ext partition (although it may just be the way I have my storage set up). Rough epoch time was 20 minutes on my machine using 5% of the data</p>",
      "rawMarkdown": "Sorry I meant to say `prepare_sensors()` after joining with the event data and then saving the .pt file. Total size is just under 1TB but with the large number of small files, you may hit the inode limit before your storage limit if you are using a Linux ext partition (although it may just be the way I have my storage set up). Rough epoch time was 20 minutes on my machine using 5% of the data",
      "votes": null
    },
    {
      "id": "2136838",
      "postDate": "02/09/2023 15:17:47",
      "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> The <code>graphnet_example</code> notebook trains on a single batch which ends up as ~ 2.5 GB, which is a real restriction regarding memory. Do you know if we can:</p>\n<ol>\n<li>Train on a single batch</li>\n<li>Save model weights</li>\n<li>Then load saved model weights</li>\n<li>Train on another batch</li>\n<li>Repeat</li>\n</ol>\n<p>Idk if this is called incremental training but that would at least allow us non-GPU owners train on many batches.</p>",
      "rawMarkdown": "anjum48 The `graphnet_example` notebook trains on a single batch which ends up as ~ 2.5 GB, which is a real restriction regarding memory. Do you know if we can:\n1. Train on a single batch\n2. Save model weights\n3. Then load saved model weights\n4. Train on another batch\n5. Repeat\n\nIdk if this is called incremental training but that would at least allow us non-GPU owners train on many batches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2129721,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "02/04/2023 20:15:41",
      "content": "<p>Thanks for this! Can you share your model training code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2130693,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "02/05/2023 16:17:00",
          "content": "<p>I use PyTorch Lightning, so you just call <code>pl.Trainer</code> with the model</p>",
          "votes": null,
          "replies": [
            {
              "id": 2130980,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "02/05/2023 20:12:29",
              "content": "<p>Maybe it's easier than I realize, but I was wondering precisely how to create dataset from \"roughly speaking 33 batches\"?</p>\n<p>Can you simply pass in a single python list containing 33 IceCubeSubmissionDataset objects, or something similarly trivial?</p>\n<p>For both your code and the starter notebook from competition hosts, it's not entirely trivial (to my limited understanding and familiarity with PyTorch etc) to train on larger amounts of data. If attempting everything in-memory on Kaggle, their version might run out of memory if attempting a single larger 30+ batch size conversion to a single sqlite db (I think?). Your version avoids the conversion to sqlite, but doesn't have a plug and play 'how to do a larger dataset' example code.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2131072,
                  "author_name": "imeintanis",
                  "author_url": "",
                  "post_date": "02/05/2023 21:12:58",
                  "content": "<p>actually, +1 to this if you like to give some light <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> <br>\nhave you used also a Parquet dataset for training?  do you recall the RAM requirements approx? </p>\n<p>BTW <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> <br>\nI have tried to modify the <code>ParquetSubmissionDataset()</code> to a suitable train dataset (I'm able to concat multiple batches but facing OOM errors for &gt;5 batches in kaggle kernels) <br>\napart from this for training there are some attributes that missing from the <code>Data</code> (torch.geometric.data) object (or I can't understand the code well yet), i.e. <br>\na) create the edges with <code>KNNGraphBuilder()</code> and b) add the labels (y)<br>\nIf I make any progress I let you know.. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2131611,
                      "author_name": "anjum48",
                      "author_url": "",
                      "post_date": "02/06/2023 08:29:02",
                      "content": "<p>Ah yes, I forgot about this.</p>\n<p>I am training on my local machine and have created <code>.pt</code> files for each <code>Data</code> object using <code>torch.save</code> beforehand using the output of the <code>prepare_samples()</code> function - everything else is the same, other than the <code>y</code> values which are also saved in the <code>Data</code> object. I then set <code>limit_train_batches=0.05</code> in <code>Trainer</code>.</p>\n<p>It was quite tricky doing this since even with a large drive, I hit the inode limit for my partition, so I had to split the files over 2 drives. </p>\n<p>I suspect that doing it in the way the hosts have done it in the baseline code might be better though. I remember seeing some dataset related code specific to this competition recently added to the Graphnet repo</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2132825,
                          "author_name": "roberthatch",
                          "author_url": "",
                          "post_date": "02/07/2023 04:04:17",
                          "content": "<p>Hmm interesting. Thanks, I guess I'll see if I can follow that and get something working. Though I'm probably gonna just make sure I have single batch training working first on each model</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2133278,
                              "author_name": "imeintanis",
                              "author_url": "",
                              "post_date": "02/07/2023 10:42:31",
                              "content": "<p>Yes, this is what I'm trying now as well, to find an efficient training method (in terms of ram, hdspace, epoch time) for 1 batch and then move to a local VM.. loading parquet files is too slow so maybe we should try what anjum suggested above.  </p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        },
                        {
                          "id": 2133289,
                          "author_name": "imeintanis",
                          "author_url": "",
                          "post_date": "02/07/2023 10:47:29",
                          "content": "<blockquote>\n  <p>files for each Data object using torch.save beforehand using the output of the prepare_samples() function</p>\n</blockquote>\n<p>prepare_samples() ? </p>\n<blockquote>\n  <p>so I had to split the files over 2 drives.</p>\n</blockquote>\n<p>so I guess we speak for &gt; 1TB ? could you also give us a rough estimate of training epoch time?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2134937,
                              "author_name": "anjum48",
                              "author_url": "",
                              "post_date": "02/08/2023 10:53:50",
                              "content": "<p>Sorry I meant to say <code>prepare_sensors()</code> after joining with the event data and then saving the .pt file. Total size is just under 1TB but with the large number of small files, you may hit the inode limit before your storage limit if you are using a Linux ext partition (although it may just be the way I have my storage set up). Rough epoch time was 20 minutes on my machine using 5% of the data</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2136838,
              "author_name": "alejopaullier",
              "author_url": "",
              "post_date": "02/09/2023 15:17:47",
              "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> The <code>graphnet_example</code> notebook trains on a single batch which ends up as ~ 2.5 GB, which is a real restriction regarding memory. Do you know if we can:</p>\n<ol>\n<li>Train on a single batch</li>\n<li>Save model weights</li>\n<li>Then load saved model weights</li>\n<li>Train on another batch</li>\n<li>Repeat</li>\n</ol>\n<p>Idk if this is called incremental training but that would at least allow us non-GPU owners train on many batches.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2130463,
      "author_name": "aavashsubedi",
      "author_url": "",
      "post_date": "02/05/2023 13:23:12",
      "content": "<p>Very nice, will be interesting to see how a GAT or other flavours of GNN's perform :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2128930": "Hi all,\n\nMy 1.046 solution is now public. \n\nhttps://www.kaggle.com/code/anjum48/early-sharing-prize-dynedge-1-046/notebook\n\nFor code competitions, I actually use .py files rather than notebooks as I just copy and paste relevant parts of my codebase into a single submission script. This means that it is very light on comments so I will summarise them here.\n\n- I used the [DynEdge](https://arxiv.org/abs/2209.03042) baseline from Graphnet as already shared by Rasmus [here](https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/383524). I added some batch norm layers and switched out the activation function\n- I trained on 5% of the data (roughly speaking 33 batches). Validation score was 1.048.\n- I used the ice transparency data from [here](https://arxiv.org/abs/1301.5361)\n- I also used the following loss function: Von-Mises Fisher for azimuth, L1 for zenith, and Cosine similarity\n- A 180-degree TTA about the z-axis also brings some gain",
    "2129721": "Thanks for this! Can you share your model training code?",
    "2130463": "Very nice, will be interesting to see how a GAT or other flavours of GNN's perform :)",
    "2130693": "I use PyTorch Lightning, so you just call `pl.Trainer` with the model",
    "2130980": "Maybe it's easier than I realize, but I was wondering precisely how to create dataset from \"roughly speaking 33 batches\"?\n\nCan you simply pass in a single python list containing 33 IceCubeSubmissionDataset objects, or something similarly trivial?\n\nFor both your code and the starter notebook from competition hosts, it's not entirely trivial (to my limited understanding and familiarity with PyTorch etc) to train on larger amounts of data. If attempting everything in-memory on Kaggle, their version might run out of memory if attempting a single larger 30+ batch size conversion to a single sqlite db (I think?). Your version avoids the conversion to sqlite, but doesn't have a plug and play 'how to do a larger dataset' example code.",
    "2131072": "actually, +1 to this if you like to give some light @anjum48 \nhave you used also a Parquet dataset for training?  do you recall the RAM requirements approx? \n\nBTW @roberthatch \nI have tried to modify the `ParquetSubmissionDataset()` to a suitable train dataset (I'm able to concat multiple batches but facing OOM errors for >5 batches in kaggle kernels) \napart from this for training there are some attributes that missing from the `Data` (torch.geometric.data) object (or I can't understand the code well yet), i.e. \na) create the edges with `KNNGraphBuilder()` and b) add the labels (y)\nIf I make any progress I let you know..",
    "2131611": "Ah yes, I forgot about this.\n\nI am training on my local machine and have created `.pt` files for each `Data` object using `torch.save` beforehand using the output of the `prepare_samples()` function - everything else is the same, other than the `y` values which are also saved in the `Data` object. I then set `limit_train_batches=0.05` in `Trainer`.\n\nIt was quite tricky doing this since even with a large drive, I hit the inode limit for my partition, so I had to split the files over 2 drives. \n\nI suspect that doing it in the way the hosts have done it in the baseline code might be better though. I remember seeing some dataset related code specific to this competition recently added to the Graphnet repo",
    "2132825": "Hmm interesting. Thanks, I guess I'll see if I can follow that and get something working. Though I'm probably gonna just make sure I have single batch training working first on each model",
    "2133278": "Yes, this is what I'm trying now as well, to find an efficient training method (in terms of ram, hdspace, epoch time) for 1 batch and then move to a local VM.. loading parquet files is too slow so maybe we should try what anjum suggested above.",
    "2133289": "> files for each Data object using torch.save beforehand using the output of the prepare_samples() function\n\nprepare_samples() ? \n\n> so I had to split the files over 2 drives.\n\nso I guess we speak for > 1TB ? could you also give us a rough estimate of training epoch time?",
    "2134937": "Sorry I meant to say `prepare_sensors()` after joining with the event data and then saving the .pt file. Total size is just under 1TB but with the large number of small files, you may hit the inode limit before your storage limit if you are using a Linux ext partition (although it may just be the way I have my storage set up). Rough epoch time was 20 minutes on my machine using 5% of the data",
    "2136838": "anjum48 The `graphnet_example` notebook trains on a single batch which ends up as ~ 2.5 GB, which is a real restriction regarding memory. Do you know if we can:\n1. Train on a single batch\n2. Save model weights\n3. Then load saved model weights\n4. Train on another batch\n5. Repeat\n\nIdk if this is called incremental training but that would at least allow us non-GPU owners train on many batches."
  },
  "source": "meta"
}