{
  "id": 390500,
  "title": "🌍 LSTM round-up",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/390500",
  "author_name": "",
  "post_date": "2023-02-25T21:57:36.149619700Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Starting with Zhaoun Booty's excellent set of notebooks based around an LSTM model and scoring 1.077, I tried various different things. For anybody who's interested, here's a write-up of where I'm at.</p>\n<h2>Things that worked well</h2>\n<p><strong>On-demand data loading</strong>: Using a generator for loading batches of training data allowed me to use far more batches without running out of memory.  I've been using the first 384x200k event batches (instead of 4 batches).  Training a variety of different models, they all seem to stop improving after about 200 batches.  Training in this way also means that no training sample is seen more than once, vastly removing the chance of overfitting.</p>\n<p><strong>Larger networks</strong>: With more training data came the ability to train larger networks.  I used 256 units in my LSTM and 256 units in the dense layer.</p>\n<p><strong>Bug fixes</strong>: I fixed a bug with the azimuth-shifted model, where it was shifting by a whole bin width instead of half a bin width.  This was one of the early things I did and it gave a small boost.  I didn't return to training the azimuth-shifted models later because they only help a little bit and they double the training time.  There was another bug I fixed along the way but sadly I can't remember what it was.</p>\n<p><strong>Early stopping, checkpointing, weight reloading &amp; LR scheduling</strong>: Just good practice really, but one of my runs was ruined after the model randomized late in the training.  (Probably an exploding gradients problem.)  I added an LR schedule that decreased from 1e-3 to 1e-5 over the course of the training run.</p>\n<p>With a combination of these &amp; some more minor improvements, <strong>I got the error down to 1.050</strong>.</p>\n<h2>Things that worked less well</h2>\n<p><strong>Splitting into separate azimuth &amp; zenith heads</strong>: The original notebooks bucketed the outputs into 256 bins and then did a classification task.  I noted that they did a much better job of predicting azimuth than zenith.  Therefore, I tried splitting into separate azimuth and zenith heads and also playing with the relative weights given to the two heads.  The results were very marginally positive but probably in the noise.</p>\n<p><strong>Predicting x, y &amp; z vector components</strong>: The literature suggests that it's better to predict x, y &amp; z components rather than azimuth &amp; zenith.  That didn't help for me - whether or not I forced the combination of x, y &amp; z to be a unit vector (although it was better with).  This experiment also meant that I was predicting regression targets rather than the bucketed classification targets used in the original notebooks - so it's hard to pick apart the contribution of these two things.</p>\n<h2>Promising avenues</h2>\n<p>Looking at the residual errors, the events clearly split into those that the model can predict very accurately (within a couple of degrees) and the rest where it's doing not much better than just guessing.  Interestingly, you can train another model to predict which of those two scenarios a particular sample belongs to.  Armed with that, we could train a different model for the \"hard-to-predict\" set - although I don't yet know what features that model might use that could improve its performance.</p>\n<p>Obviously the GNN baseline is still a fair bit better at 1.018.  Perhaps its time to abandon the LSTM based approach.</p>",
  "messages": [
    {
      "id": "2159623",
      "postDate": "02/25/2023 21:57:36",
      "content": "<p>Starting with Zhaoun Booty's excellent set of notebooks based around an LSTM model and scoring 1.077, I tried various different things. For anybody who's interested, here's a write-up of where I'm at.</p>\n<h2>Things that worked well</h2>\n<p><strong>On-demand data loading</strong>: Using a generator for loading batches of training data allowed me to use far more batches without running out of memory.  I've been using the first 384x200k event batches (instead of 4 batches).  Training a variety of different models, they all seem to stop improving after about 200 batches.  Training in this way also means that no training sample is seen more than once, vastly removing the chance of overfitting.</p>\n<p><strong>Larger networks</strong>: With more training data came the ability to train larger networks.  I used 256 units in my LSTM and 256 units in the dense layer.</p>\n<p><strong>Bug fixes</strong>: I fixed a bug with the azimuth-shifted model, where it was shifting by a whole bin width instead of half a bin width.  This was one of the early things I did and it gave a small boost.  I didn't return to training the azimuth-shifted models later because they only help a little bit and they double the training time.  There was another bug I fixed along the way but sadly I can't remember what it was.</p>\n<p><strong>Early stopping, checkpointing, weight reloading &amp; LR scheduling</strong>: Just good practice really, but one of my runs was ruined after the model randomized late in the training.  (Probably an exploding gradients problem.)  I added an LR schedule that decreased from 1e-3 to 1e-5 over the course of the training run.</p>\n<p>With a combination of these &amp; some more minor improvements, <strong>I got the error down to 1.050</strong>.</p>\n<h2>Things that worked less well</h2>\n<p><strong>Splitting into separate azimuth &amp; zenith heads</strong>: The original notebooks bucketed the outputs into 256 bins and then did a classification task.  I noted that they did a much better job of predicting azimuth than zenith.  Therefore, I tried splitting into separate azimuth and zenith heads and also playing with the relative weights given to the two heads.  The results were very marginally positive but probably in the noise.</p>\n<p><strong>Predicting x, y &amp; z vector components</strong>: The literature suggests that it's better to predict x, y &amp; z components rather than azimuth &amp; zenith.  That didn't help for me - whether or not I forced the combination of x, y &amp; z to be a unit vector (although it was better with).  This experiment also meant that I was predicting regression targets rather than the bucketed classification targets used in the original notebooks - so it's hard to pick apart the contribution of these two things.</p>\n<h2>Promising avenues</h2>\n<p>Looking at the residual errors, the events clearly split into those that the model can predict very accurately (within a couple of degrees) and the rest where it's doing not much better than just guessing.  Interestingly, you can train another model to predict which of those two scenarios a particular sample belongs to.  Armed with that, we could train a different model for the \"hard-to-predict\" set - although I don't yet know what features that model might use that could improve its performance.</p>\n<p>Obviously the GNN baseline is still a fair bit better at 1.018.  Perhaps its time to abandon the LSTM based approach.</p>",
      "rawMarkdown": "Starting with Zhaoun Booty's excellent set of notebooks based around an LSTM model and scoring 1.077, I tried various different things. For anybody who's interested, here's a write-up of where I'm at.\n\n## Things that worked well\n\n**On-demand data loading**: Using a generator for loading batches of training data allowed me to use far more batches without running out of memory.  I've been using the first 384x200k event batches (instead of 4 batches).  Training a variety of different models, they all seem to stop improving after about 200 batches.  Training in this way also means that no training sample is seen more than once, vastly removing the chance of overfitting.\n\n**Larger networks**: With more training data came the ability to train larger networks.  I used 256 units in my LSTM and 256 units in the dense layer.\n\n**Bug fixes**: I fixed a bug with the azimuth-shifted model, where it was shifting by a whole bin width instead of half a bin width.  This was one of the early things I did and it gave a small boost.  I didn't return to training the azimuth-shifted models later because they only help a little bit and they double the training time.  There was another bug I fixed along the way but sadly I can't remember what it was.\n\n**Early stopping, checkpointing, weight reloading & LR scheduling**: Just good practice really, but one of my runs was ruined after the model randomized late in the training.  (Probably an exploding gradients problem.)  I added an LR schedule that decreased from 1e-3 to 1e-5 over the course of the training run.\n\nWith a combination of these & some more minor improvements, **I got the error down to 1.050**.\n\n## Things that worked less well\n\n**Splitting into separate azimuth & zenith heads**: The original notebooks bucketed the outputs into 256 bins and then did a classification task.  I noted that they did a much better job of predicting azimuth than zenith.  Therefore, I tried splitting into separate azimuth and zenith heads and also playing with the relative weights given to the two heads.  The results were very marginally positive but probably in the noise.\n\n**Predicting x, y & z vector components**: The literature suggests that it's better to predict x, y & z components rather than azimuth & zenith.  That didn't help for me - whether or not I forced the combination of x, y & z to be a unit vector (although it was better with).  This experiment also meant that I was predicting regression targets rather than the bucketed classification targets used in the original notebooks - so it's hard to pick apart the contribution of these two things.\n\n## Promising avenues\n\nLooking at the residual errors, the events clearly split into those that the model can predict very accurately (within a couple of degrees) and the rest where it's doing not much better than just guessing.  Interestingly, you can train another model to predict which of those two scenarios a particular sample belongs to.  Armed with that, we could train a different model for the \"hard-to-predict\" set - although I don't yet know what features that model might use that could improve its performance.\n\nObviously the GNN baseline is still a fair bit better at 1.018.  Perhaps its time to abandon the LSTM based approach.",
      "votes": null
    },
    {
      "id": "2159839",
      "postDate": "02/26/2023 04:58:30",
      "content": "<p>All of the models seem to have the ability to predict some events very accurately - on my todo list is to compare the math models predictions, graphnet, etc. vs my LSTM.  My guess would be that the same bunch of events are good by all; and bad by all.  With the difference in performance likely just the stuff in the middle.  Several shared kernels show the interesting shape of the distribution of the angular score, and I am again betting that all the notebooks from 1.0 to 1.06 have a pretty similiar look.</p>\n<p>Been working on solving problems with data for half a century - #1 lesson learned was to look at the bad guesses and fix the root causes.  I see 3 pretty obvious areas were features exist that can help LSTM.  I find it very hard to add features to graphnet, but easy (perhaps too easy I might have too many now) to add them to LSTM.  Just finishing #1 obvious and hope to go to bed with training underway.</p>\n<p>The LSTM for sure gets better with training over a large number of batches and nice to know that I can stop adding batches when I get to 200:)   </p>\n<p>Because I find LSTM so easy to use, I plan to stay with it all the way, so I don't agree that it's time to abandon the LSTM approach.</p>",
      "rawMarkdown": "All of the models seem to have the ability to predict some events very accurately - on my todo list is to compare the math models predictions, graphnet, etc. vs my LSTM.  My guess would be that the same bunch of events are good by all; and bad by all.  With the difference in performance likely just the stuff in the middle.  Several shared kernels show the interesting shape of the distribution of the angular score, and I am again betting that all the notebooks from 1.0 to 1.06 have a pretty similiar look.\n\nBeen working on solving problems with data for half a century - #1 lesson learned was to look at the bad guesses and fix the root causes.  I see 3 pretty obvious areas were features exist that can help LSTM.  I find it very hard to add features to graphnet, but easy (perhaps too easy I might have too many now) to add them to LSTM.  Just finishing #1 obvious and hope to go to bed with training underway.\n\nThe LSTM for sure gets better with training over a large number of batches and nice to know that I can stop adding batches when I get to 200:)   \n\nBecause I find LSTM so easy to use, I plan to stay with it all the way, so I don't agree that it's time to abandon the LSTM approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2159839,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/26/2023 04:58:30",
      "content": "<p>All of the models seem to have the ability to predict some events very accurately - on my todo list is to compare the math models predictions, graphnet, etc. vs my LSTM.  My guess would be that the same bunch of events are good by all; and bad by all.  With the difference in performance likely just the stuff in the middle.  Several shared kernels show the interesting shape of the distribution of the angular score, and I am again betting that all the notebooks from 1.0 to 1.06 have a pretty similiar look.</p>\n<p>Been working on solving problems with data for half a century - #1 lesson learned was to look at the bad guesses and fix the root causes.  I see 3 pretty obvious areas were features exist that can help LSTM.  I find it very hard to add features to graphnet, but easy (perhaps too easy I might have too many now) to add them to LSTM.  Just finishing #1 obvious and hope to go to bed with training underway.</p>\n<p>The LSTM for sure gets better with training over a large number of batches and nice to know that I can stop adding batches when I get to 200:)   </p>\n<p>Because I find LSTM so easy to use, I plan to stay with it all the way, so I don't agree that it's time to abandon the LSTM approach.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2159623": "Starting with Zhaoun Booty's excellent set of notebooks based around an LSTM model and scoring 1.077, I tried various different things. For anybody who's interested, here's a write-up of where I'm at.\n\n## Things that worked well\n\n**On-demand data loading**: Using a generator for loading batches of training data allowed me to use far more batches without running out of memory.  I've been using the first 384x200k event batches (instead of 4 batches).  Training a variety of different models, they all seem to stop improving after about 200 batches.  Training in this way also means that no training sample is seen more than once, vastly removing the chance of overfitting.\n\n**Larger networks**: With more training data came the ability to train larger networks.  I used 256 units in my LSTM and 256 units in the dense layer.\n\n**Bug fixes**: I fixed a bug with the azimuth-shifted model, where it was shifting by a whole bin width instead of half a bin width.  This was one of the early things I did and it gave a small boost.  I didn't return to training the azimuth-shifted models later because they only help a little bit and they double the training time.  There was another bug I fixed along the way but sadly I can't remember what it was.\n\n**Early stopping, checkpointing, weight reloading & LR scheduling**: Just good practice really, but one of my runs was ruined after the model randomized late in the training.  (Probably an exploding gradients problem.)  I added an LR schedule that decreased from 1e-3 to 1e-5 over the course of the training run.\n\nWith a combination of these & some more minor improvements, **I got the error down to 1.050**.\n\n## Things that worked less well\n\n**Splitting into separate azimuth & zenith heads**: The original notebooks bucketed the outputs into 256 bins and then did a classification task.  I noted that they did a much better job of predicting azimuth than zenith.  Therefore, I tried splitting into separate azimuth and zenith heads and also playing with the relative weights given to the two heads.  The results were very marginally positive but probably in the noise.\n\n**Predicting x, y & z vector components**: The literature suggests that it's better to predict x, y & z components rather than azimuth & zenith.  That didn't help for me - whether or not I forced the combination of x, y & z to be a unit vector (although it was better with).  This experiment also meant that I was predicting regression targets rather than the bucketed classification targets used in the original notebooks - so it's hard to pick apart the contribution of these two things.\n\n## Promising avenues\n\nLooking at the residual errors, the events clearly split into those that the model can predict very accurately (within a couple of degrees) and the rest where it's doing not much better than just guessing.  Interestingly, you can train another model to predict which of those two scenarios a particular sample belongs to.  Armed with that, we could train a different model for the \"hard-to-predict\" set - although I don't yet know what features that model might use that could improve its performance.\n\nObviously the GNN baseline is still a fair bit better at 1.018.  Perhaps its time to abandon the LSTM based approach.",
    "2159839": "All of the models seem to have the ability to predict some events very accurately - on my todo list is to compare the math models predictions, graphnet, etc. vs my LSTM.  My guess would be that the same bunch of events are good by all; and bad by all.  With the difference in performance likely just the stuff in the middle.  Several shared kernels show the interesting shape of the distribution of the angular score, and I am again betting that all the notebooks from 1.0 to 1.06 have a pretty similiar look.\n\nBeen working on solving problems with data for half a century - #1 lesson learned was to look at the bad guesses and fix the root causes.  I see 3 pretty obvious areas were features exist that can help LSTM.  I find it very hard to add features to graphnet, but easy (perhaps too easy I might have too many now) to add them to LSTM.  Just finishing #1 obvious and hope to go to bed with training underway.\n\nThe LSTM for sure gets better with training over a large number of batches and nice to know that I can stop adding batches when I get to 200:)   \n\nBecause I find LSTM so easy to use, I plan to stay with it all the way, so I don't agree that it's time to abandon the LSTM approach."
  },
  "source": "meta"
}