{
  "id": 193143,
  "title": "LSTM insights",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/193143",
  "author_name": "",
  "post_date": "2020-10-25T14:04:02.854391600Z",
  "votes": 10,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>In this public <a href=\"https://www.kaggle.com/suryajrrafl/lstm-baseline\" target=\"_blank\">notebook</a>, I have tried using a LSTM for the given problem. The basic idea is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F42f95f0f6e5d4f2af9fdf028e90ae99f%2Flstm%20baseline%20idea.jpg?generation=1603633858768045&amp;alt=media\" alt=\"LSTM-baseline-idea\"></p>\n<p>I trained the above model on about 340k samples (LSTM input size 1024, hidden state size 64, 2 layers stacked on batch size 16, learning rate 1e-3, Adam optimiser and NLL loss function). This is the training loss curves I got.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F9f7973108712ab21ed6ad6510a24ae92%2FLSTM_baseline_train_losses_0-340k.png?generation=1603634101621918&amp;alt=media\" alt=\"Overall train loss plot\"></p>\n<p>After about 150k samples till 340k samples my training loss seemed stagnant at 240. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fee2671d00febc5bfd36463c423860421%2FLSTM_baseline_train_losses_156k-340k.png?generation=1603634198881367&amp;alt=media\" alt=\"Train_loss 150k-340k\"></p>\n<p>Is there anything I'm missing with my LSTM implementation? (Pardon me for the noobie question, I'm new to pytorch and LSTM). Meanwhile, I'm running the model to check the validation score to debug further. Feedback is most welcome. </p>",
  "messages": [
    {
      "id": "1059831",
      "postDate": "10/25/2020 14:04:02",
      "content": "<p>Hi,</p>\n<p>In this public <a href=\"https://www.kaggle.com/suryajrrafl/lstm-baseline\" target=\"_blank\">notebook</a>, I have tried using a LSTM for the given problem. The basic idea is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F42f95f0f6e5d4f2af9fdf028e90ae99f%2Flstm%20baseline%20idea.jpg?generation=1603633858768045&amp;alt=media\" alt=\"LSTM-baseline-idea\"></p>\n<p>I trained the above model on about 340k samples (LSTM input size 1024, hidden state size 64, 2 layers stacked on batch size 16, learning rate 1e-3, Adam optimiser and NLL loss function). This is the training loss curves I got.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F9f7973108712ab21ed6ad6510a24ae92%2FLSTM_baseline_train_losses_0-340k.png?generation=1603634101621918&amp;alt=media\" alt=\"Overall train loss plot\"></p>\n<p>After about 150k samples till 340k samples my training loss seemed stagnant at 240. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fee2671d00febc5bfd36463c423860421%2FLSTM_baseline_train_losses_156k-340k.png?generation=1603634198881367&amp;alt=media\" alt=\"Train_loss 150k-340k\"></p>\n<p>Is there anything I'm missing with my LSTM implementation? (Pardon me for the noobie question, I'm new to pytorch and LSTM). Meanwhile, I'm running the model to check the validation score to debug further. Feedback is most welcome. </p>",
      "rawMarkdown": "Hi,\n\nIn this public [notebook](https://www.kaggle.com/suryajrrafl/lstm-baseline), I have tried using a LSTM for the given problem. The basic idea is as follows:\n\n![LSTM-baseline-idea](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F42f95f0f6e5d4f2af9fdf028e90ae99f%2Flstm%20baseline%20idea.jpg?generation=1603633858768045&alt=media)\n\nI trained the above model on about 340k samples (LSTM input size 1024, hidden state size 64, 2 layers stacked on batch size 16, learning rate 1e-3, Adam optimiser and NLL loss function). This is the training loss curves I got.\n\n![Overall train loss plot](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F9f7973108712ab21ed6ad6510a24ae92%2FLSTM_baseline_train_losses_0-340k.png?generation=1603634101621918&alt=media)\n\nAfter about 150k samples till 340k samples my training loss seemed stagnant at 240. \n\n![Train_loss 150k-340k](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fee2671d00febc5bfd36463c423860421%2FLSTM_baseline_train_losses_156k-340k.png?generation=1603634198881367&alt=media)\n\nIs there anything I'm missing with my LSTM implementation? (Pardon me for the noobie question, I'm new to pytorch and LSTM). Meanwhile, I'm running the model to check the validation score to debug further. Feedback is most welcome.",
      "votes": null
    },
    {
      "id": "1060103",
      "postDate": "10/25/2020 19:18:38",
      "content": "<p>Whats the reason for using LSTM? </p>\n<p>Your loss is incredibly high for the amount of used samples, with resnet18 15k samples results in a train loss of 37 for me.</p>",
      "rawMarkdown": "Whats the reason for using LSTM? \n\nYour loss is incredibly high for the amount of used samples, with resnet18 15k samples results in a train loss of 37 for me.",
      "votes": null
    },
    {
      "id": "1060240",
      "postDate": "10/26/2020 00:47:31",
      "content": "<p>It could be that the learning rate is too high, or the batches may be too small / stochastic and the weight updates have become contrary between batches? <br>\nMomentum with a lower learning rate, and maybe reshuffling the samples in batches after a few epoch of training</p>",
      "rawMarkdown": "It could be that the learning rate is too high, or the batches may be too small / stochastic and the weight updates have become contrary between batches? \nMomentum with a lower learning rate, and maybe reshuffling the samples in batches after a few epoch of training",
      "votes": null
    },
    {
      "id": "1060322",
      "postDate": "10/26/2020 03:52:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>,</p>\n<p>Thanks for your feedback. </p>\n<ol>\n<li><p>Not a strict rule to use LSTMs, am new to Deep learning stuff, so trying the explore options. The nature of the problem (given sequence of past trajectories, predicting the future trajectories) made me try this. Also, I saw some research papers <a href=\"https://arxiv.org/pdf/1912.11676.pdf\" target=\"_blank\">Paper1</a>, <a href=\"https://openaccess.thecvf.com/content_WACV_2020/papers/Djuric_Uncertainty-aware_Short-term_Motion_Prediction_of_Traffic_Actors_for_Autonomous_Driving_WACV_2020_paper.pdf\" target=\"_blank\">Paper2</a> with LSTM based approaches. Actually I would like to implement seq2seq models with attention and see results. </p></li>\n<li><p>Yes, my training loss is very high for the data used (just to be explicit, by 340k samples I mean 20k batches, each of size 16). As <a href=\"https://www.kaggle.com/n3n7i\" target=\"_blank\">@n3n7i</a> pointed out, my learning rate could be a problem. Also have to check if gradient clipping (many suggested this when using RNNs to avoid vanishing gradient problem) helps.  </p></li>\n</ol>",
      "rawMarkdown": "Hi @aliabdin1,\n\nThanks for your feedback. \n\n1. Not a strict rule to use LSTMs, am new to Deep learning stuff, so trying the explore options. The nature of the problem (given sequence of past trajectories, predicting the future trajectories) made me try this. Also, I saw some research papers [Paper1](https://arxiv.org/pdf/1912.11676.pdf), [Paper2](https://openaccess.thecvf.com/content_WACV_2020/papers/Djuric_Uncertainty-aware_Short-term_Motion_Prediction_of_Traffic_Actors_for_Autonomous_Driving_WACV_2020_paper.pdf) with LSTM based approaches. Actually I would like to implement seq2seq models with attention and see results. \n\n\n2. Yes, my training loss is very high for the data used (just to be explicit, by 340k samples I mean 20k batches, each of size 16). As @n3n7i pointed out, my learning rate could be a problem. Also have to check if gradient clipping (many suggested this when using RNNs to avoid vanishing gradient problem) helps.",
      "votes": null
    },
    {
      "id": "1060327",
      "postDate": "10/26/2020 03:58:52",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/n3n7i\" target=\"_blank\">@n3n7i</a>,</p>\n<p>Thanks for your feedback. Will try Tuning the learning rate and batch size and see where it goes. I would like to iterate and compare different models and approaches. In that case, which option is generally preferred, given the huge data and limited HW resources (just using kaggle GPUs) available? </p>\n<p>Option 1 : Train on a subset of data (say 1M for 3-4 epochs)<br>\nOption 2 : Train on more data (~5M) just once </p>\n<p>My mind says option 1, but is there any rule of thumb that I'm missing?</p>",
      "rawMarkdown": "Hi @n3n7i,\n\nThanks for your feedback. Will try Tuning the learning rate and batch size and see where it goes. I would like to iterate and compare different models and approaches. In that case, which option is generally preferred, given the huge data and limited HW resources (just using kaggle GPUs) available? \n\nOption 1 : Train on a subset of data (say 1M for 3-4 epochs)\nOption 2 : Train on more data (~5M) just once \n\nMy mind says option 1, but is there any rule of thumb that I'm missing?",
      "votes": null
    },
    {
      "id": "1061341",
      "postDate": "10/27/2020 00:25:12",
      "content": "<p>That is a bit of a tricky question, i think option 1 could possibly perform a little worse on validation data, but training error should still be a reasonable indicator. Option 2 is likely to be the other way around, may have a little worse training error but probably less jump in validation error</p>",
      "rawMarkdown": "That is a bit of a tricky question, i think option 1 could possibly perform a little worse on validation data, but training error should still be a reasonable indicator. Option 2 is likely to be the other way around, may have a little worse training error but probably less jump in validation error",
      "votes": null
    },
    {
      "id": "1062010",
      "postDate": "10/27/2020 13:58:58",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a> for your feedback. I have settled on option 2 and am trying to train baseline models from scratch. I am still stuck at figuring out how to reduce my learning rate within one epoch (using only a subset of data). I suppose <strong>torch.optim.lr_scheduler</strong> functions are generally used when we train for multiple epochs. </p>\n<p>I'm trying to reduce the learning rate if mean and variance of training loss on last 'n' batches is within a defined threshold using the option described here <a href=\"https://stackoverflow.com/questions/48324152/pytorch-how-to-change-the-learning-rate-of-an-optimizer-at-any-given-moment-no\" target=\"_blank\">change learning rate</a>. Am I reinventing the wheel here or is this the way to go?</p>",
      "rawMarkdown": "Thanks @n3n77i for your feedback. I have settled on option 2 and am trying to train baseline models from scratch. I am still stuck at figuring out how to reduce my learning rate within one epoch (using only a subset of data). I suppose **torch.optim.lr_scheduler** functions are generally used when we train for multiple epochs. \n\n\nI'm trying to reduce the learning rate if mean and variance of training loss on last 'n' batches is within a defined threshold using the option described here [change learning rate](https://stackoverflow.com/questions/48324152/pytorch-how-to-change-the-learning-rate-of-an-optimizer-at-any-given-moment-no). Am I reinventing the wheel here or is this the way to go?",
      "votes": null
    },
    {
      "id": "1062594",
      "postDate": "10/28/2020 02:39:42",
      "content": "<p>ReduceLROnPlateau looks pretty close, but will still need more than one epoch with patience=0, so maybe more suited to your option 1.<br>\nYou might be able to use OneCycleLR for option 2, pct_start=0.0 i think should prevent it from increasing the learning rate. It will modify the learning rate after each batch </p>\n<p>The catch with using batch loss to update the learning rate is they are only indirectly related, some batches will have higher or lower loss somewhat regardless of the current lr. Moving averages of loss falling over some number of batches seems it would be ok (no decrease to lr), and growing variances may also work (signal to decrease lr). </p>",
      "rawMarkdown": "ReduceLROnPlateau looks pretty close, but will still need more than one epoch with patience=0, so maybe more suited to your option 1.\nYou might be able to use OneCycleLR for option 2, pct_start=0.0 i think should prevent it from increasing the learning rate. It will modify the learning rate after each batch \n\nThe catch with using batch loss to update the learning rate is they are only indirectly related, some batches will have higher or lower loss somewhat regardless of the current lr. Moving averages of loss falling over some number of batches seems it would be ok (no decrease to lr), and growing variances may also work (signal to decrease lr).",
      "votes": null
    },
    {
      "id": "1062710",
      "postDate": "10/28/2020 05:57:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a> once again for the feedback. Will look to implement it. </p>",
      "rawMarkdown": "Thanks @n3n77i once again for the feedback. Will look to implement it.",
      "votes": null
    },
    {
      "id": "1062909",
      "postDate": "10/28/2020 10:02:47",
      "content": "<p><a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a> In the image you have posted, can you please explain what those 11 frames are? Are they 11 consecutive values in the AgentDataset list where each value is the following - <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000147%2F121f495c9c75231196a512310a447d25%2FScreen%20Shot%202020-10-28%20at%203.31.12%20PM.png?generation=1603879322264805&amp;alt=media\" alt=\"\"></p>\n<p>I am basically trying to understand what the input is to the model. Thanks!</p>",
      "rawMarkdown": "suryajrrafl In the image you have posted, can you please explain what those 11 frames are? Are they 11 consecutive values in the AgentDataset list where each value is the following - ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000147%2F121f495c9c75231196a512310a447d25%2FScreen%20Shot%202020-10-28%20at%203.31.12%20PM.png?generation=1603879322264805&alt=media)\n\nI am basically trying to understand what the input is to the model. Thanks!",
      "votes": null
    },
    {
      "id": "1062930",
      "postDate": "10/28/2020 10:24:07",
      "content": "<p>Yes <a href=\"https://www.kaggle.com/arpitrf\" target=\"_blank\">@arpitrf</a> the 11 frames correspond to those in data['image'] variable</p>",
      "rawMarkdown": "Yes @arpitrf the 11 frames correspond to those in data['image'] variable",
      "votes": null
    },
    {
      "id": "1062952",
      "postDate": "10/28/2020 10:55:26",
      "content": "<p>Okay thanks for your answer <a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a>. So say we are trying to predict the 50 \"target_positions\" for agent_dataset[1]. What would the 11 frames (which correspond to data['image']) be for this?</p>",
      "rawMarkdown": "Okay thanks for your answer @suryajrrafl. So say we are trying to predict the 50 \"target_positions\" for agent_dataset[1]. What would the 11 frames (which correspond to data['image']) be for this?",
      "votes": null
    },
    {
      "id": "1062962",
      "postDate": "10/28/2020 11:10:23",
      "content": "<p>They represent the Bird's eye view image in the past 10 frames (+1 current frame). Using the past 'n' frames you are trying to predict the future 50 positions. You can increase the number in the <strong>cfg</strong> dictionary but it'll take more time to process the data. </p>",
      "rawMarkdown": "They represent the Bird's eye view image in the past 10 frames (+1 current frame). Using the past 'n' frames you are trying to predict the future 50 positions. You can increase the number in the **cfg** dictionary but it'll take more time to process the data.",
      "votes": null
    },
    {
      "id": "1062969",
      "postDate": "10/28/2020 11:25:48",
      "content": "<p>Okay. And it also needs to be ensured that the past 10 frames are in the same scene right <a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a> ?</p>",
      "rawMarkdown": "Okay. And it also needs to be ensured that the past 10 frames are in the same scene right @suryajrrafl ?",
      "votes": null
    },
    {
      "id": "1062980",
      "postDate": "10/28/2020 11:55:26",
      "content": "<p>I think that is taken care in the l5kit library function</p>",
      "rawMarkdown": "I think that is taken care in the l5kit library function",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1060103,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "10/25/2020 19:18:38",
      "content": "<p>Whats the reason for using LSTM? </p>\n<p>Your loss is incredibly high for the amount of used samples, with resnet18 15k samples results in a train loss of 37 for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1060322,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "10/26/2020 03:52:40",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>,</p>\n<p>Thanks for your feedback. </p>\n<ol>\n<li><p>Not a strict rule to use LSTMs, am new to Deep learning stuff, so trying the explore options. The nature of the problem (given sequence of past trajectories, predicting the future trajectories) made me try this. Also, I saw some research papers <a href=\"https://arxiv.org/pdf/1912.11676.pdf\" target=\"_blank\">Paper1</a>, <a href=\"https://openaccess.thecvf.com/content_WACV_2020/papers/Djuric_Uncertainty-aware_Short-term_Motion_Prediction_of_Traffic_Actors_for_Autonomous_Driving_WACV_2020_paper.pdf\" target=\"_blank\">Paper2</a> with LSTM based approaches. Actually I would like to implement seq2seq models with attention and see results. </p></li>\n<li><p>Yes, my training loss is very high for the data used (just to be explicit, by 340k samples I mean 20k batches, each of size 16). As <a href=\"https://www.kaggle.com/n3n7i\" target=\"_blank\">@n3n7i</a> pointed out, my learning rate could be a problem. Also have to check if gradient clipping (many suggested this when using RNNs to avoid vanishing gradient problem) helps.  </p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1060240,
      "author_name": "n3n77i",
      "author_url": "",
      "post_date": "10/26/2020 00:47:31",
      "content": "<p>It could be that the learning rate is too high, or the batches may be too small / stochastic and the weight updates have become contrary between batches? <br>\nMomentum with a lower learning rate, and maybe reshuffling the samples in batches after a few epoch of training</p>",
      "votes": null,
      "replies": [
        {
          "id": 1060327,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "10/26/2020 03:58:52",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/n3n7i\" target=\"_blank\">@n3n7i</a>,</p>\n<p>Thanks for your feedback. Will try Tuning the learning rate and batch size and see where it goes. I would like to iterate and compare different models and approaches. In that case, which option is generally preferred, given the huge data and limited HW resources (just using kaggle GPUs) available? </p>\n<p>Option 1 : Train on a subset of data (say 1M for 3-4 epochs)<br>\nOption 2 : Train on more data (~5M) just once </p>\n<p>My mind says option 1, but is there any rule of thumb that I'm missing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061341,
          "author_name": "n3n77i",
          "author_url": "",
          "post_date": "10/27/2020 00:25:12",
          "content": "<p>That is a bit of a tricky question, i think option 1 could possibly perform a little worse on validation data, but training error should still be a reasonable indicator. Option 2 is likely to be the other way around, may have a little worse training error but probably less jump in validation error</p>",
          "votes": null,
          "replies": [
            {
              "id": 1062010,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "10/27/2020 13:58:58",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a> for your feedback. I have settled on option 2 and am trying to train baseline models from scratch. I am still stuck at figuring out how to reduce my learning rate within one epoch (using only a subset of data). I suppose <strong>torch.optim.lr_scheduler</strong> functions are generally used when we train for multiple epochs. </p>\n<p>I'm trying to reduce the learning rate if mean and variance of training loss on last 'n' batches is within a defined threshold using the option described here <a href=\"https://stackoverflow.com/questions/48324152/pytorch-how-to-change-the-learning-rate-of-an-optimizer-at-any-given-moment-no\" target=\"_blank\">change learning rate</a>. Am I reinventing the wheel here or is this the way to go?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1062594,
              "author_name": "n3n77i",
              "author_url": "",
              "post_date": "10/28/2020 02:39:42",
              "content": "<p>ReduceLROnPlateau looks pretty close, but will still need more than one epoch with patience=0, so maybe more suited to your option 1.<br>\nYou might be able to use OneCycleLR for option 2, pct_start=0.0 i think should prevent it from increasing the learning rate. It will modify the learning rate after each batch </p>\n<p>The catch with using batch loss to update the learning rate is they are only indirectly related, some batches will have higher or lower loss somewhat regardless of the current lr. Moving averages of loss falling over some number of batches seems it would be ok (no decrease to lr), and growing variances may also work (signal to decrease lr). </p>",
              "votes": null,
              "replies": [
                {
                  "id": 1062710,
                  "author_name": "suryajrrafl",
                  "author_url": "",
                  "post_date": "10/28/2020 05:57:54",
                  "content": "<p>Thanks <a href=\"https://www.kaggle.com/n3n77i\" target=\"_blank\">@n3n77i</a> once again for the feedback. Will look to implement it. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 1062909,
      "author_name": "arpitrf",
      "author_url": "",
      "post_date": "10/28/2020 10:02:47",
      "content": "<p><a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a> In the image you have posted, can you please explain what those 11 frames are? Are they 11 consecutive values in the AgentDataset list where each value is the following - <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000147%2F121f495c9c75231196a512310a447d25%2FScreen%20Shot%202020-10-28%20at%203.31.12%20PM.png?generation=1603879322264805&amp;alt=media\" alt=\"\"></p>\n<p>I am basically trying to understand what the input is to the model. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1062930,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "10/28/2020 10:24:07",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/arpitrf\" target=\"_blank\">@arpitrf</a> the 11 frames correspond to those in data['image'] variable</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1062952,
          "author_name": "arpitrf",
          "author_url": "",
          "post_date": "10/28/2020 10:55:26",
          "content": "<p>Okay thanks for your answer <a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a>. So say we are trying to predict the 50 \"target_positions\" for agent_dataset[1]. What would the 11 frames (which correspond to data['image']) be for this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 1062962,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "10/28/2020 11:10:23",
              "content": "<p>They represent the Bird's eye view image in the past 10 frames (+1 current frame). Using the past 'n' frames you are trying to predict the future 50 positions. You can increase the number in the <strong>cfg</strong> dictionary but it'll take more time to process the data. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1062969,
              "author_name": "arpitrf",
              "author_url": "",
              "post_date": "10/28/2020 11:25:48",
              "content": "<p>Okay. And it also needs to be ensured that the past 10 frames are in the same scene right <a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a> ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 1062980,
                  "author_name": "suryajrrafl",
                  "author_url": "",
                  "post_date": "10/28/2020 11:55:26",
                  "content": "<p>I think that is taken care in the l5kit library function</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1059831": "Hi,\n\nIn this public [notebook](https://www.kaggle.com/suryajrrafl/lstm-baseline), I have tried using a LSTM for the given problem. The basic idea is as follows:\n\n![LSTM-baseline-idea](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F42f95f0f6e5d4f2af9fdf028e90ae99f%2Flstm%20baseline%20idea.jpg?generation=1603633858768045&alt=media)\n\nI trained the above model on about 340k samples (LSTM input size 1024, hidden state size 64, 2 layers stacked on batch size 16, learning rate 1e-3, Adam optimiser and NLL loss function). This is the training loss curves I got.\n\n![Overall train loss plot](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2F9f7973108712ab21ed6ad6510a24ae92%2FLSTM_baseline_train_losses_0-340k.png?generation=1603634101621918&alt=media)\n\nAfter about 150k samples till 340k samples my training loss seemed stagnant at 240. \n\n![Train_loss 150k-340k](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fee2671d00febc5bfd36463c423860421%2FLSTM_baseline_train_losses_156k-340k.png?generation=1603634198881367&alt=media)\n\nIs there anything I'm missing with my LSTM implementation? (Pardon me for the noobie question, I'm new to pytorch and LSTM). Meanwhile, I'm running the model to check the validation score to debug further. Feedback is most welcome.",
    "1060103": "Whats the reason for using LSTM? \n\nYour loss is incredibly high for the amount of used samples, with resnet18 15k samples results in a train loss of 37 for me.",
    "1060240": "It could be that the learning rate is too high, or the batches may be too small / stochastic and the weight updates have become contrary between batches? \nMomentum with a lower learning rate, and maybe reshuffling the samples in batches after a few epoch of training",
    "1060322": "Hi @aliabdin1,\n\nThanks for your feedback. \n\n1. Not a strict rule to use LSTMs, am new to Deep learning stuff, so trying the explore options. The nature of the problem (given sequence of past trajectories, predicting the future trajectories) made me try this. Also, I saw some research papers [Paper1](https://arxiv.org/pdf/1912.11676.pdf), [Paper2](https://openaccess.thecvf.com/content_WACV_2020/papers/Djuric_Uncertainty-aware_Short-term_Motion_Prediction_of_Traffic_Actors_for_Autonomous_Driving_WACV_2020_paper.pdf) with LSTM based approaches. Actually I would like to implement seq2seq models with attention and see results. \n\n\n2. Yes, my training loss is very high for the data used (just to be explicit, by 340k samples I mean 20k batches, each of size 16). As @n3n7i pointed out, my learning rate could be a problem. Also have to check if gradient clipping (many suggested this when using RNNs to avoid vanishing gradient problem) helps.",
    "1060327": "Hi @n3n7i,\n\nThanks for your feedback. Will try Tuning the learning rate and batch size and see where it goes. I would like to iterate and compare different models and approaches. In that case, which option is generally preferred, given the huge data and limited HW resources (just using kaggle GPUs) available? \n\nOption 1 : Train on a subset of data (say 1M for 3-4 epochs)\nOption 2 : Train on more data (~5M) just once \n\nMy mind says option 1, but is there any rule of thumb that I'm missing?",
    "1061341": "That is a bit of a tricky question, i think option 1 could possibly perform a little worse on validation data, but training error should still be a reasonable indicator. Option 2 is likely to be the other way around, may have a little worse training error but probably less jump in validation error",
    "1062010": "Thanks @n3n77i for your feedback. I have settled on option 2 and am trying to train baseline models from scratch. I am still stuck at figuring out how to reduce my learning rate within one epoch (using only a subset of data). I suppose **torch.optim.lr_scheduler** functions are generally used when we train for multiple epochs. \n\n\nI'm trying to reduce the learning rate if mean and variance of training loss on last 'n' batches is within a defined threshold using the option described here [change learning rate](https://stackoverflow.com/questions/48324152/pytorch-how-to-change-the-learning-rate-of-an-optimizer-at-any-given-moment-no). Am I reinventing the wheel here or is this the way to go?",
    "1062594": "ReduceLROnPlateau looks pretty close, but will still need more than one epoch with patience=0, so maybe more suited to your option 1.\nYou might be able to use OneCycleLR for option 2, pct_start=0.0 i think should prevent it from increasing the learning rate. It will modify the learning rate after each batch \n\nThe catch with using batch loss to update the learning rate is they are only indirectly related, some batches will have higher or lower loss somewhat regardless of the current lr. Moving averages of loss falling over some number of batches seems it would be ok (no decrease to lr), and growing variances may also work (signal to decrease lr).",
    "1062710": "Thanks @n3n77i once again for the feedback. Will look to implement it.",
    "1062909": "suryajrrafl In the image you have posted, can you please explain what those 11 frames are? Are they 11 consecutive values in the AgentDataset list where each value is the following - ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2000147%2F121f495c9c75231196a512310a447d25%2FScreen%20Shot%202020-10-28%20at%203.31.12%20PM.png?generation=1603879322264805&alt=media)\n\nI am basically trying to understand what the input is to the model. Thanks!",
    "1062930": "Yes @arpitrf the 11 frames correspond to those in data['image'] variable",
    "1062952": "Okay thanks for your answer @suryajrrafl. So say we are trying to predict the 50 \"target_positions\" for agent_dataset[1]. What would the 11 frames (which correspond to data['image']) be for this?",
    "1062962": "They represent the Bird's eye view image in the past 10 frames (+1 current frame). Using the past 'n' frames you are trying to predict the future 50 positions. You can increase the number in the **cfg** dictionary but it'll take more time to process the data.",
    "1062969": "Okay. And it also needs to be ensured that the past 10 frames are in the same scene right @suryajrrafl ?",
    "1062980": "I think that is taken care in the l5kit library function"
  },
  "source": "meta"
}