{
  "id": 199541,
  "title": "Things I have tried and my final solution",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/writeups/terrible-driver-things-i-have-tried-and-my-final-s",
  "author_name": "",
  "post_date": "2020-11-26T22:37:07.810Z",
  "votes": 22,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Congratulations to all the winners and thank the host for this very interesting competition. I have to admit that I did not expect a medal at all at the beginning. At some point I almost gave up, as the training was painfully slow, and I just could not get a reasonable score. This result gives me great motivation to keep trying  in the future.</p>\n<p>Forgive me if my terminology does not make sense. Any feedback would be greatly appreciated. </p>\n<p>To my understanding, this is not a problem of finding three most likely future routes, which is equivalent to finding the one most likely route, as the second most likely route will always be the most likely route shifted by one nanometre. Instead, this is a problem of finding three routes that could best represent the probability distribution. Ideally, we want to include the less likely, but nonetheless typical routes. The most damage to our score probably will be caused by those less likely but very different routes. Therefore, we want diversity in our predictions, and simple ensemble might not work. </p>\n<p>My work is based on the model shared by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">here</a>. While this simple approach to achieve multi-mode prediction is surprisingly effective, there is one thing that does not make sense to me, and I have been fighting with this problem most of the time: </p>\n<p>At every training step, the coordinates of all three predictions are all pulled towards the ground truth. The confidence of the prediction that is most close to the ground truth is increased, while the confidences for the other two predictions are decreased. Therefore, for the other two predictions that are relatively further from the ground truth, we are decreasing their confidence values (meaning now we think they are less likely) but pushing their coordinates to the ground truth (meaning making them more likely). Although at the early training stage this should not be such a big problem and the model is able to converge, I just cannot believe it can converge to an optimal point. <br>\nOther models such as classification models or NLP models would not have this problem as the target possibilities are fixed and finite. Here we are basically assigning confidence values to moving targets. </p>\n<p>I thought about several solutions: </p>\n<ul>\n<li>I constructed a “diversity” factor in the loss function. It is basically the average distance between three predictions. By adding this to the loss I was hoping I could gently push three predictions away from each other, reduce the effect of three predictions being pulled together. However, my experiments were of no success. The model either totally ignored this factor or used this factor as the only way to gain lower loss. I did not try many times because every experiment took too long. </li>\n<li>Make the coordinate space discrete and finite and assign a confidence value to every possible point in this space at each step. Then generate randomly many possible future routes. Finally do a k-means clustering to cluster the routes into three groups and take the centre of each group as the final prediction. I did not even finish the implementation of this idea as the computer power needed would be out of my reach. </li>\n<li>Very large batch size. It was until very late into the competition I suddenly realized that maybe increasing the batch size could mitigate (certainly not resolve) this problem. By letting the model see as many future possibilities as possible at each step, the model might learn to maintain the diversity of its three predictio<a href=\"url\" target=\"_blank\"></a>ns. It indeed worked, although to be honest I am not sure it was only because of the problem I mentioned above. </li>\n</ul>\n<p>So, my final solution might seem surprisingly simple to most people. I just used the good old Resnet18, a very small image setting of 150x150 with only 5 history frames, which enabled me to fit in a batch of 512 samples into my 8G VRAM. The optimizer is again good old Adam, with a learning rate starting from 0.0001 and reduce by half every 50000 steps. I trained on the full dataset, not because I think I need so much data, but to mitigate the problem of overlapping samples as discussed <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">here</a>.  </p>\n<p>I trained 400k steps, which intermittently took me more than 15 days!!! After this, my computer and I were so exhausted, so we did not try other models or optimizers. </p>\n<p>(I tried accumulating gradients to increase the effect batch size but again the training was too slow, and the early result was not fantastic, possibly because of the incompatibility with Batch Normalization.) </p>\n<p>I guess, if we could make the batch size even bigger, and image size a little larger, train longer, or maybe use a more sophisticated model, there is potential to further improve the score significantly. </p>\n<p>By the way, I have never really solved the problem of deviation between my training loss and validation loss. After removing some problematic parts of my model and setting the min future and history frames in alignment with the validation dataset, the problem was only partly solved. My training loss has reached below 10 but validation loss was never below 12.10. This is not too bad, but I know some of you get much better alignment. How did you guys get the scores aligned? </p>",
  "messages": [
    {
      "id": "1091588",
      "postDate": "11/26/2020 06:01:12",
      "content": "<p>Congratulations to all the winners and thank the host for this very interesting competition. I have to admit that I did not expect a medal at all at the beginning. At some point I almost gave up, as the training was painfully slow, and I just could not get a reasonable score. This result gives me great motivation to keep trying  in the future.</p>\n<p>Forgive me if my terminology does not make sense. Any feedback would be greatly appreciated. </p>\n<p>To my understanding, this is not a problem of finding three most likely future routes, which is equivalent to finding the one most likely route, as the second most likely route will always be the most likely route shifted by one nanometre. Instead, this is a problem of finding three routes that could best represent the probability distribution. Ideally, we want to include the less likely, but nonetheless typical routes. The most damage to our score probably will be caused by those less likely but very different routes. Therefore, we want diversity in our predictions, and simple ensemble might not work. </p>\n<p>My work is based on the model shared by <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> <a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence\" target=\"_blank\">here</a>. While this simple approach to achieve multi-mode prediction is surprisingly effective, there is one thing that does not make sense to me, and I have been fighting with this problem most of the time: </p>\n<p>At every training step, the coordinates of all three predictions are all pulled towards the ground truth. The confidence of the prediction that is most close to the ground truth is increased, while the confidences for the other two predictions are decreased. Therefore, for the other two predictions that are relatively further from the ground truth, we are decreasing their confidence values (meaning now we think they are less likely) but pushing their coordinates to the ground truth (meaning making them more likely). Although at the early training stage this should not be such a big problem and the model is able to converge, I just cannot believe it can converge to an optimal point. <br>\nOther models such as classification models or NLP models would not have this problem as the target possibilities are fixed and finite. Here we are basically assigning confidence values to moving targets. </p>\n<p>I thought about several solutions: </p>\n<ul>\n<li>I constructed a “diversity” factor in the loss function. It is basically the average distance between three predictions. By adding this to the loss I was hoping I could gently push three predictions away from each other, reduce the effect of three predictions being pulled together. However, my experiments were of no success. The model either totally ignored this factor or used this factor as the only way to gain lower loss. I did not try many times because every experiment took too long. </li>\n<li>Make the coordinate space discrete and finite and assign a confidence value to every possible point in this space at each step. Then generate randomly many possible future routes. Finally do a k-means clustering to cluster the routes into three groups and take the centre of each group as the final prediction. I did not even finish the implementation of this idea as the computer power needed would be out of my reach. </li>\n<li>Very large batch size. It was until very late into the competition I suddenly realized that maybe increasing the batch size could mitigate (certainly not resolve) this problem. By letting the model see as many future possibilities as possible at each step, the model might learn to maintain the diversity of its three predictio<a href=\"url\" target=\"_blank\"></a>ns. It indeed worked, although to be honest I am not sure it was only because of the problem I mentioned above. </li>\n</ul>\n<p>So, my final solution might seem surprisingly simple to most people. I just used the good old Resnet18, a very small image setting of 150x150 with only 5 history frames, which enabled me to fit in a batch of 512 samples into my 8G VRAM. The optimizer is again good old Adam, with a learning rate starting from 0.0001 and reduce by half every 50000 steps. I trained on the full dataset, not because I think I need so much data, but to mitigate the problem of overlapping samples as discussed <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">here</a>.  </p>\n<p>I trained 400k steps, which intermittently took me more than 15 days!!! After this, my computer and I were so exhausted, so we did not try other models or optimizers. </p>\n<p>(I tried accumulating gradients to increase the effect batch size but again the training was too slow, and the early result was not fantastic, possibly because of the incompatibility with Batch Normalization.) </p>\n<p>I guess, if we could make the batch size even bigger, and image size a little larger, train longer, or maybe use a more sophisticated model, there is potential to further improve the score significantly. </p>\n<p>By the way, I have never really solved the problem of deviation between my training loss and validation loss. After removing some problematic parts of my model and setting the min future and history frames in alignment with the validation dataset, the problem was only partly solved. My training loss has reached below 10 but validation loss was never below 12.10. This is not too bad, but I know some of you get much better alignment. How did you guys get the scores aligned? </p>",
      "rawMarkdown": "Congratulations to all the winners and thank the host for this very interesting competition. I have to admit that I did not expect a medal at all at the beginning. At some point I almost gave up, as the training was painfully slow, and I just could not get a reasonable score. This result gives me great motivation to keep trying  in the future.\n\nForgive me if my terminology does not make sense. Any feedback would be greatly appreciated. \n\nTo my understanding, this is not a problem of finding three most likely future routes, which is equivalent to finding the one most likely route, as the second most likely route will always be the most likely route shifted by one nanometre. Instead, this is a problem of finding three routes that could best represent the probability distribution. Ideally, we want to include the less likely, but nonetheless typical routes. The most damage to our score probably will be caused by those less likely but very different routes. Therefore, we want diversity in our predictions, and simple ensemble might not work. \n\nMy work is based on the model shared by @corochann [here](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence). While this simple approach to achieve multi-mode prediction is surprisingly effective, there is one thing that does not make sense to me, and I have been fighting with this problem most of the time: \n\nAt every training step, the coordinates of all three predictions are all pulled towards the ground truth. The confidence of the prediction that is most close to the ground truth is increased, while the confidences for the other two predictions are decreased. Therefore, for the other two predictions that are relatively further from the ground truth, we are decreasing their confidence values (meaning now we think they are less likely) but pushing their coordinates to the ground truth (meaning making them more likely). Although at the early training stage this should not be such a big problem and the model is able to converge, I just cannot believe it can converge to an optimal point. \nOther models such as classification models or NLP models would not have this problem as the target possibilities are fixed and finite. Here we are basically assigning confidence values to moving targets. \n\nI thought about several solutions: \n- I constructed a “diversity” factor in the loss function. It is basically the average distance between three predictions. By adding this to the loss I was hoping I could gently push three predictions away from each other, reduce the effect of three predictions being pulled together. However, my experiments were of no success. The model either totally ignored this factor or used this factor as the only way to gain lower loss. I did not try many times because every experiment took too long. \n- Make the coordinate space discrete and finite and assign a confidence value to every possible point in this space at each step. Then generate randomly many possible future routes. Finally do a k-means clustering to cluster the routes into three groups and take the centre of each group as the final prediction. I did not even finish the implementation of this idea as the computer power needed would be out of my reach. \n- Very large batch size. It was until very late into the competition I suddenly realized that maybe increasing the batch size could mitigate (certainly not resolve) this problem. By letting the model see as many future possibilities as possible at each step, the model might learn to maintain the diversity of its three predictio[](url)ns. It indeed worked, although to be honest I am not sure it was only because of the problem I mentioned above. \n\nSo, my final solution might seem surprisingly simple to most people. I just used the good old Resnet18, a very small image setting of 150x150 with only 5 history frames, which enabled me to fit in a batch of 512 samples into my 8G VRAM. The optimizer is again good old Adam, with a learning rate starting from 0.0001 and reduce by half every 50000 steps. I trained on the full dataset, not because I think I need so much data, but to mitigate the problem of overlapping samples as discussed [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762).  \n\nI trained 400k steps, which intermittently took me more than 15 days!!! After this, my computer and I were so exhausted, so we did not try other models or optimizers. \n\n(I tried accumulating gradients to increase the effect batch size but again the training was too slow, and the early result was not fantastic, possibly because of the incompatibility with Batch Normalization.) \n\nI guess, if we could make the batch size even bigger, and image size a little larger, train longer, or maybe use a more sophisticated model, there is potential to further improve the score significantly. \n\nBy the way, I have never really solved the problem of deviation between my training loss and validation loss. After removing some problematic parts of my model and setting the min future and history frames in alignment with the validation dataset, the problem was only partly solved. My training loss has reached below 10 but validation loss was never below 12.10. This is not too bad, but I know some of you get much better alignment. How did you guys get the scores aligned?",
      "votes": null
    },
    {
      "id": "1091593",
      "postDate": "11/26/2020 06:05:37",
      "content": "<p>Interesting. So you think the main thing was a large batch size that contributed to better results?</p>",
      "rawMarkdown": "Interesting. So you think the main thing was a large batch size that contributed to better results?",
      "votes": null
    },
    {
      "id": "1091598",
      "postDate": "11/26/2020 06:08:08",
      "content": "<p>That was my theory, and it did help, but to be honest I am not sure if that is the only explanation.</p>",
      "rawMarkdown": "That was my theory, and it did help, but to be honest I am not sure if that is the only explanation.",
      "votes": null
    },
    {
      "id": "1091601",
      "postDate": "11/26/2020 06:11:39",
      "content": "<p>WOW! Batch size 512! That is huge!</p>",
      "rawMarkdown": "WOW! Batch size 512! That is huge!",
      "votes": null
    },
    {
      "id": "1091603",
      "postDate": "11/26/2020 06:13:03",
      "content": "<p>Yeah, I think larger batch size should give you better result in long run. Even though it train slower (since you have less backprop per samples)</p>",
      "rawMarkdown": "Yeah, I think larger batch size should give you better result in long run. Even though it train slower (since you have less backprop per samples)",
      "votes": null
    },
    {
      "id": "1091618",
      "postDate": "11/26/2020 06:23:34",
      "content": "<p>Thanks for sharing. It seems like large batch size is the key factor.</p>",
      "rawMarkdown": "Thanks for sharing. It seems like large batch size is the key factor.",
      "votes": null
    },
    {
      "id": "1091628",
      "postDate": "11/26/2020 06:28:45",
      "content": "<p>I noticed that going from 32 to 16 bs really hurt results. I was also splitting across 2 gpus so my effective batch size was even smaller than that. I guess I should've tried even larger. </p>",
      "rawMarkdown": "I noticed that going from 32 to 16 bs really hurt results. I was also splitting across 2 gpus so my effective batch size was even smaller than that. I guess I should've tried even larger.",
      "votes": null
    },
    {
      "id": "1091668",
      "postDate": "11/26/2020 07:28:15",
      "content": "<p>Same, I tried BS 16 which affected results in a negative way. Did not try anything above 32…honestly, skimming through some recent write ups, I see that experimenting with hyperparams and backbones would have been extremely beneficial :) Most of us probably rushed into implementing complex ideas without getting some \"basics\" right</p>\n<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  Congrats on your result and thank you for sharing! </p>",
      "rawMarkdown": "Same, I tried BS 16 which affected results in a negative way. Did not try anything above 32...honestly, skimming through some recent write ups, I see that experimenting with hyperparams and backbones would have been extremely beneficial :) Most of us probably rushed into implementing complex ideas without getting some \"basics\" right\n\n@frankpanxj  Congrats on your result and thank you for sharing!",
      "votes": null
    },
    {
      "id": "1091676",
      "postDate": "11/26/2020 07:34:18",
      "content": "<p>I implemented so many papers to no avail. Finding out that batch size might have been my limiting factor is very disappointing. Lesson learned if that is the case. </p>",
      "rawMarkdown": "I implemented so many papers to no avail. Finding out that batch size might have been my limiting factor is very disappointing. Lesson learned if that is the case.",
      "votes": null
    },
    {
      "id": "1091684",
      "postDate": "11/26/2020 07:45:35",
      "content": "<p>We live and we learn….but I can say that regardless of the result, the amount we learned is priceless. Armed with new knowledge and experience, we wont make the same mistakes again in future competitions;)</p>",
      "rawMarkdown": "We live and we learn....but I can say that regardless of the result, the amount we learned is priceless. Armed with new knowledge and experience, we wont make the same mistakes again in future competitions;)",
      "votes": null
    },
    {
      "id": "1091742",
      "postDate": "11/26/2020 08:44:42",
      "content": "<p>Batch size did not matter too much for us I believe, but BS is always an interplay with LR and mostly we used BS &gt;= 64</p>",
      "rawMarkdown": "Batch size did not matter too much for us I believe, but BS is always an interplay with LR and mostly we used BS >= 64",
      "votes": null
    },
    {
      "id": "1091758",
      "postDate": "11/26/2020 08:58:38",
      "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> what's your default with BS and LR (pick your optimizer). </p>\n<p>I start with Adam and BS=32 and set LR between 5e-5 to 5e-4 depending on my assumption how valuable are imagenet features for the domain task. Changes in BS and LR lean toward directly proportional</p>",
      "rawMarkdown": "philippsinger what's your default with BS and LR (pick your optimizer). \n\nI start with Adam and BS=32 and set LR between 5e-5 to 5e-4 depending on my assumption how valuable are imagenet features for the domain task. Changes in BS and LR lean toward directly proportional",
      "votes": null
    },
    {
      "id": "1091765",
      "postDate": "11/26/2020 09:09:44",
      "content": "<p>Nevertheless congratulations is good learning. I did come late to the competition and no time to train. I implemented a solution inspired by one of the top contenders in the nuscenes challenge (<a href=\"https://arxiv.org/pdf/2005.02545.pdf)\" target=\"_blank\">https://arxiv.org/pdf/2005.02545.pdf)</a>. Still training. I got the impression that for this particular competition computational power is important. I don't think I can fit anything beyond 224x224 in my 8GB GPU.</p>",
      "rawMarkdown": "Nevertheless congratulations is good learning. I did come late to the competition and no time to train. I implemented a solution inspired by one of the top contenders in the nuscenes challenge (https://arxiv.org/pdf/2005.02545.pdf). Still training. I got the impression that for this particular competition computational power is important. I don't think I can fit anything beyond 224x224 in my 8GB GPU.",
      "votes": null
    },
    {
      "id": "1091779",
      "postDate": "11/26/2020 09:23:49",
      "content": "<p>Computational power is important and allows you to iterate through a large number of ideas, but in this case that turned out to be not 100% necessary.</p>\n<p>As we can see, <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  was able to achieve a top-tier solution without any impressive hardware, but by thinking \"outside the box\" and trying some seemingly simple things that others wouldn't even think of trying….and waiting patiently for 15 days;)<br>\nThis is what I really like about this solution.</p>",
      "rawMarkdown": "Computational power is important and allows you to iterate through a large number of ideas, but in this case that turned out to be not 100% necessary.\n\nAs we can see, @frankpanxj  was able to achieve a top-tier solution without any impressive hardware, but by thinking \"outside the box\" and trying some seemingly simple things that others wouldn't even think of trying....and waiting patiently for 15 days;)\nThis is what I really like about this solution.",
      "votes": null
    },
    {
      "id": "1091791",
      "postDate": "11/26/2020 09:35:36",
      "content": "<p><a href=\"https://www.kaggle.com/ture05\" target=\"_blank\">@ture05</a> I also implemented the joint multi head attention. Looked reasonably promising but seemed to converge to a similar result to my more vanilla models. Seems I did not train for long enough though so maybe it does eventually converge to a better result. </p>",
      "rawMarkdown": "ture05 I also implemented the joint multi head attention. Looked reasonably promising but seemed to converge to a similar result to my more vanilla models. Seems I did not train for long enough though so maybe it does eventually converge to a better result.",
      "votes": null
    },
    {
      "id": "1091794",
      "postDate": "11/26/2020 09:37:53",
      "content": "<p>It really depends on the problem, but for us a good start here was Adam BS 64 and LR 1e-4. I always use linear or cosine decay with rare exceptions.</p>",
      "rawMarkdown": "It really depends on the problem, but for us a good start here was Adam BS 64 and LR 1e-4. I always use linear or cosine decay with rare exceptions.",
      "votes": null
    },
    {
      "id": "1091804",
      "postDate": "11/26/2020 09:51:00",
      "content": "<p>Kudos to <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  definitely computational power and patience is what's needed. </p>\n<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> How far did you go? Currently 50k iter ( 64 bs) and it's outperforming the baseline (for the same iterations in training). I did add a few things, more car states and my encoding was slightly different.</p>",
      "rawMarkdown": "Kudos to @frankpanxj  definitely computational power and patience is what's needed. \n\n@ryches How far did you go? Currently 50k iter ( 64 bs) and it's outperforming the baseline (for the same iterations in training). I did add a few things, more car states and my encoding was slightly different.",
      "votes": null
    },
    {
      "id": "1091807",
      "postDate": "11/26/2020 09:54:28",
      "content": "<p>I allowed it to run for about 24 hours, not sure exactly the number of samples it was exposed to, but at that point it had converged to 17.5 validation loss. </p>",
      "rawMarkdown": "I allowed it to run for about 24 hours, not sure exactly the number of samples it was exposed to, but at that point it had converged to 17.5 validation loss.",
      "votes": null
    },
    {
      "id": "1091852",
      "postDate": "11/26/2020 10:39:49",
      "content": "<p>Thanks for sharing. And, congrats.<br>\nCould you tell us how your loss was getting down?  Even if you have log of loss, please show us.<br>\nOur team might give up training too early…</p>",
      "rawMarkdown": "Thanks for sharing. And, congrats.\nCould you tell us how your loss was getting down?  Even if you have log of loss, please show us.\nOur team might give up training too early...",
      "votes": null
    },
    {
      "id": "1091929",
      "postDate": "11/26/2020 11:41:56",
      "content": "<p>Thanks for your kind words <a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a>  😂. The only reason I could hold out that long was just because I was observing steadly decreasing loss every day, so there is no reason to stop.</p>",
      "rawMarkdown": "Thanks for your kind words @indswetrust  😂. The only reason I could hold out that long was just because I was observing steadly decreasing loss every day, so there is no reason to stop.",
      "votes": null
    },
    {
      "id": "1091933",
      "postDate": "11/26/2020 11:46:31",
      "content": "<p>iter    training_loss (avg) eval_loss<br>\n10000    23.5           27.526<br>\n20000    20.01   23.999<br>\n30000    16.69   20.397<br>\n40000    17.12   18.814<br>\n50000    15.25   17.781<br>\n60000    13.1           15.989<br>\n70000    13.93   15.538<br>\n80000    14.29   15.311<br>\n90000    12.74   15.17<br>\n100000    12.38   14.593<br>\n110000    12.54   14.021<br>\n120000    11.53   13.705<br>\n130000    12.41   13.868<br>\n140000    11.81   13.485<br>\n150000    10.79   13.321<br>\n160000    10.92   12.983<br>\n170000    10.81   12.987<br>\n180000    10.98   13.04<br>\n190000    11.83   12.893<br>\n200000    10.94   12.806<br>\n210000    10.7            12.63<br>\n220000    10.48   12.666<br>\n230000    10.6             12.527<br>\n240000   No data, the machine shut down for no reason 😂        <br>\n250000    10.43   12.555<br>\n260000    10.46   12.409<br>\n270000    10.69   12.426<br>\n280000    10.35   12.376<br>\n290000    10.55   12.347<br>\n300000    9.92            12.342<br>\n310000    11.07   12.314<br>\n320000    10.55   12.3<br>\n330000    9.94    12.317<br>\nhere I started the second round just before the competition deadline, restarted from lr of 0.00001 (drop by half every 20000 steps)<br>\n340000    10.67   12.465<br>\n350000    11.33   12.385<br>\n360000    10.1            12.352<br>\n370000    10.12   12.295<br>\n380000    10.83   12.269<br>\n390000    10.29   12.232<br>\nduring the final steps I checkpointed every 3000 steps<br>\n393000    9.98          12.2<br>\n396000    10.38   12.21<br>\n399000    10.5            12.189<br>\n402000    10.66   12.156<br>\n405000    10.25   12.241<br>\n408000    10.72   12.168</p>",
      "rawMarkdown": "iter\ttraining_loss (avg)\teval_loss\n10000\t23.5\t       27.526\n20000\t20.01\t23.999\n30000\t16.69\t20.397\n40000\t17.12\t18.814\n50000\t15.25\t17.781\n60000\t13.1\t       15.989\n70000\t13.93\t15.538\n80000\t14.29\t15.311\n90000\t12.74\t15.17\n100000\t12.38\t14.593\n110000\t12.54\t14.021\n120000\t11.53\t13.705\n130000\t12.41\t13.868\n140000\t11.81\t13.485\n150000\t10.79\t13.321\n160000\t10.92\t12.983\n170000\t10.81\t12.987\n180000\t10.98\t13.04\n190000\t11.83\t12.893\n200000\t10.94\t12.806\n210000\t10.7\t        12.63\n220000\t10.48\t12.666\n230000\t10.6\t         12.527\n240000   No data, the machine shut down for no reason 😂\t\t\n250000\t10.43\t12.555\n260000\t10.46\t12.409\n270000\t10.69\t12.426\n280000\t10.35\t12.376\n290000\t10.55\t12.347\n300000\t9.92\t        12.342\n310000\t11.07\t12.314\n320000\t10.55\t12.3\n330000\t9.94\t12.317\nhere I started the second round just before the competition deadline, restarted from lr of 0.00001 (drop by half every 20000 steps)\n340000\t10.67\t12.465\n350000\t11.33\t12.385\n360000\t10.1\t        12.352\n370000\t10.12\t12.295\n380000\t10.83\t12.269\n390000\t10.29\t12.232\nduring the final steps I checkpointed every 3000 steps\n393000\t9.98\t      12.2\n396000\t10.38\t12.21\n399000\t10.5\t        12.189\n402000\t10.66\t12.156\n405000\t10.25\t12.241\n408000\t10.72\t12.168",
      "votes": null
    },
    {
      "id": "1091999",
      "postDate": "11/26/2020 13:12:09",
      "content": "<p>Congrats and thanks for mention!<br>\nWe also discussed your 2nd thought of predicting probability distribution directly and assign 3 trajectory in later step, which may be an interesting approach for future!</p>",
      "rawMarkdown": "Congrats and thanks for mention!\nWe also discussed your 2nd thought of predicting probability distribution directly and assign 3 trajectory in later step, which may be an interesting approach for future!",
      "votes": null
    },
    {
      "id": "1092120",
      "postDate": "11/26/2020 15:01:06",
      "content": "<p>Thanks!<br>\nGreat patience you have.</p>",
      "rawMarkdown": "Thanks!\nGreat patience you have.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091593,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "11/26/2020 06:05:37",
      "content": "<p>Interesting. So you think the main thing was a large batch size that contributed to better results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091598,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "11/26/2020 06:08:08",
          "content": "<p>That was my theory, and it did help, but to be honest I am not sure if that is the only explanation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091603,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/26/2020 06:13:03",
          "content": "<p>Yeah, I think larger batch size should give you better result in long run. Even though it train slower (since you have less backprop per samples)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091628,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 06:28:45",
          "content": "<p>I noticed that going from 32 to 16 bs really hurt results. I was also splitting across 2 gpus so my effective batch size was even smaller than that. I guess I should've tried even larger. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091668,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "11/26/2020 07:28:15",
          "content": "<p>Same, I tried BS 16 which affected results in a negative way. Did not try anything above 32…honestly, skimming through some recent write ups, I see that experimenting with hyperparams and backbones would have been extremely beneficial :) Most of us probably rushed into implementing complex ideas without getting some \"basics\" right</p>\n<p><a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  Congrats on your result and thank you for sharing! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091676,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 07:34:18",
          "content": "<p>I implemented so many papers to no avail. Finding out that batch size might have been my limiting factor is very disappointing. Lesson learned if that is the case. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091684,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "11/26/2020 07:45:35",
          "content": "<p>We live and we learn….but I can say that regardless of the result, the amount we learned is priceless. Armed with new knowledge and experience, we wont make the same mistakes again in future competitions;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091742,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/26/2020 08:44:42",
          "content": "<p>Batch size did not matter too much for us I believe, but BS is always an interplay with LR and mostly we used BS &gt;= 64</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091758,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "11/26/2020 08:58:38",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> what's your default with BS and LR (pick your optimizer). </p>\n<p>I start with Adam and BS=32 and set LR between 5e-5 to 5e-4 depending on my assumption how valuable are imagenet features for the domain task. Changes in BS and LR lean toward directly proportional</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091794,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/26/2020 09:37:53",
          "content": "<p>It really depends on the problem, but for us a good start here was Adam BS 64 and LR 1e-4. I always use linear or cosine decay with rare exceptions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091601,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/26/2020 06:11:39",
      "content": "<p>WOW! Batch size 512! That is huge!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091618,
      "author_name": "sunghyunjun",
      "author_url": "",
      "post_date": "11/26/2020 06:23:34",
      "content": "<p>Thanks for sharing. It seems like large batch size is the key factor.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091765,
      "author_name": "ture05",
      "author_url": "",
      "post_date": "11/26/2020 09:09:44",
      "content": "<p>Nevertheless congratulations is good learning. I did come late to the competition and no time to train. I implemented a solution inspired by one of the top contenders in the nuscenes challenge (<a href=\"https://arxiv.org/pdf/2005.02545.pdf)\" target=\"_blank\">https://arxiv.org/pdf/2005.02545.pdf)</a>. Still training. I got the impression that for this particular competition computational power is important. I don't think I can fit anything beyond 224x224 in my 8GB GPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091779,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "11/26/2020 09:23:49",
          "content": "<p>Computational power is important and allows you to iterate through a large number of ideas, but in this case that turned out to be not 100% necessary.</p>\n<p>As we can see, <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  was able to achieve a top-tier solution without any impressive hardware, but by thinking \"outside the box\" and trying some seemingly simple things that others wouldn't even think of trying….and waiting patiently for 15 days;)<br>\nThis is what I really like about this solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091791,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 09:35:36",
          "content": "<p><a href=\"https://www.kaggle.com/ture05\" target=\"_blank\">@ture05</a> I also implemented the joint multi head attention. Looked reasonably promising but seemed to converge to a similar result to my more vanilla models. Seems I did not train for long enough though so maybe it does eventually converge to a better result. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091929,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "11/26/2020 11:41:56",
          "content": "<p>Thanks for your kind words <a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a>  😂. The only reason I could hold out that long was just because I was observing steadly decreasing loss every day, so there is no reason to stop.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091804,
      "author_name": "ture05",
      "author_url": "",
      "post_date": "11/26/2020 09:51:00",
      "content": "<p>Kudos to <a href=\"https://www.kaggle.com/frankpanxj\" target=\"_blank\">@frankpanxj</a>  definitely computational power and patience is what's needed. </p>\n<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> How far did you go? Currently 50k iter ( 64 bs) and it's outperforming the baseline (for the same iterations in training). I did add a few things, more car states and my encoding was slightly different.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091807,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "11/26/2020 09:54:28",
          "content": "<p>I allowed it to run for about 24 hours, not sure exactly the number of samples it was exposed to, but at that point it had converged to 17.5 validation loss. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091852,
      "author_name": "yasagure",
      "author_url": "",
      "post_date": "11/26/2020 10:39:49",
      "content": "<p>Thanks for sharing. And, congrats.<br>\nCould you tell us how your loss was getting down?  Even if you have log of loss, please show us.<br>\nOur team might give up training too early…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091933,
          "author_name": "frankpanxj",
          "author_url": "",
          "post_date": "11/26/2020 11:46:31",
          "content": "<p>iter    training_loss (avg) eval_loss<br>\n10000    23.5           27.526<br>\n20000    20.01   23.999<br>\n30000    16.69   20.397<br>\n40000    17.12   18.814<br>\n50000    15.25   17.781<br>\n60000    13.1           15.989<br>\n70000    13.93   15.538<br>\n80000    14.29   15.311<br>\n90000    12.74   15.17<br>\n100000    12.38   14.593<br>\n110000    12.54   14.021<br>\n120000    11.53   13.705<br>\n130000    12.41   13.868<br>\n140000    11.81   13.485<br>\n150000    10.79   13.321<br>\n160000    10.92   12.983<br>\n170000    10.81   12.987<br>\n180000    10.98   13.04<br>\n190000    11.83   12.893<br>\n200000    10.94   12.806<br>\n210000    10.7            12.63<br>\n220000    10.48   12.666<br>\n230000    10.6             12.527<br>\n240000   No data, the machine shut down for no reason 😂        <br>\n250000    10.43   12.555<br>\n260000    10.46   12.409<br>\n270000    10.69   12.426<br>\n280000    10.35   12.376<br>\n290000    10.55   12.347<br>\n300000    9.92            12.342<br>\n310000    11.07   12.314<br>\n320000    10.55   12.3<br>\n330000    9.94    12.317<br>\nhere I started the second round just before the competition deadline, restarted from lr of 0.00001 (drop by half every 20000 steps)<br>\n340000    10.67   12.465<br>\n350000    11.33   12.385<br>\n360000    10.1            12.352<br>\n370000    10.12   12.295<br>\n380000    10.83   12.269<br>\n390000    10.29   12.232<br>\nduring the final steps I checkpointed every 3000 steps<br>\n393000    9.98          12.2<br>\n396000    10.38   12.21<br>\n399000    10.5            12.189<br>\n402000    10.66   12.156<br>\n405000    10.25   12.241<br>\n408000    10.72   12.168</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1092120,
          "author_name": "yasagure",
          "author_url": "",
          "post_date": "11/26/2020 15:01:06",
          "content": "<p>Thanks!<br>\nGreat patience you have.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091999,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "11/26/2020 13:12:09",
      "content": "<p>Congrats and thanks for mention!<br>\nWe also discussed your 2nd thought of predicting probability distribution directly and assign 3 trajectory in later step, which may be an interesting approach for future!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1091588": "Congratulations to all the winners and thank the host for this very interesting competition. I have to admit that I did not expect a medal at all at the beginning. At some point I almost gave up, as the training was painfully slow, and I just could not get a reasonable score. This result gives me great motivation to keep trying  in the future.\n\nForgive me if my terminology does not make sense. Any feedback would be greatly appreciated. \n\nTo my understanding, this is not a problem of finding three most likely future routes, which is equivalent to finding the one most likely route, as the second most likely route will always be the most likely route shifted by one nanometre. Instead, this is a problem of finding three routes that could best represent the probability distribution. Ideally, we want to include the less likely, but nonetheless typical routes. The most damage to our score probably will be caused by those less likely but very different routes. Therefore, we want diversity in our predictions, and simple ensemble might not work. \n\nMy work is based on the model shared by @corochann [here](https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence). While this simple approach to achieve multi-mode prediction is surprisingly effective, there is one thing that does not make sense to me, and I have been fighting with this problem most of the time: \n\nAt every training step, the coordinates of all three predictions are all pulled towards the ground truth. The confidence of the prediction that is most close to the ground truth is increased, while the confidences for the other two predictions are decreased. Therefore, for the other two predictions that are relatively further from the ground truth, we are decreasing their confidence values (meaning now we think they are less likely) but pushing their coordinates to the ground truth (meaning making them more likely). Although at the early training stage this should not be such a big problem and the model is able to converge, I just cannot believe it can converge to an optimal point. \nOther models such as classification models or NLP models would not have this problem as the target possibilities are fixed and finite. Here we are basically assigning confidence values to moving targets. \n\nI thought about several solutions: \n- I constructed a “diversity” factor in the loss function. It is basically the average distance between three predictions. By adding this to the loss I was hoping I could gently push three predictions away from each other, reduce the effect of three predictions being pulled together. However, my experiments were of no success. The model either totally ignored this factor or used this factor as the only way to gain lower loss. I did not try many times because every experiment took too long. \n- Make the coordinate space discrete and finite and assign a confidence value to every possible point in this space at each step. Then generate randomly many possible future routes. Finally do a k-means clustering to cluster the routes into three groups and take the centre of each group as the final prediction. I did not even finish the implementation of this idea as the computer power needed would be out of my reach. \n- Very large batch size. It was until very late into the competition I suddenly realized that maybe increasing the batch size could mitigate (certainly not resolve) this problem. By letting the model see as many future possibilities as possible at each step, the model might learn to maintain the diversity of its three predictio[](url)ns. It indeed worked, although to be honest I am not sure it was only because of the problem I mentioned above. \n\nSo, my final solution might seem surprisingly simple to most people. I just used the good old Resnet18, a very small image setting of 150x150 with only 5 history frames, which enabled me to fit in a batch of 512 samples into my 8G VRAM. The optimizer is again good old Adam, with a learning rate starting from 0.0001 and reduce by half every 50000 steps. I trained on the full dataset, not because I think I need so much data, but to mitigate the problem of overlapping samples as discussed [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762).  \n\nI trained 400k steps, which intermittently took me more than 15 days!!! After this, my computer and I were so exhausted, so we did not try other models or optimizers. \n\n(I tried accumulating gradients to increase the effect batch size but again the training was too slow, and the early result was not fantastic, possibly because of the incompatibility with Batch Normalization.) \n\nI guess, if we could make the batch size even bigger, and image size a little larger, train longer, or maybe use a more sophisticated model, there is potential to further improve the score significantly. \n\nBy the way, I have never really solved the problem of deviation between my training loss and validation loss. After removing some problematic parts of my model and setting the min future and history frames in alignment with the validation dataset, the problem was only partly solved. My training loss has reached below 10 but validation loss was never below 12.10. This is not too bad, but I know some of you get much better alignment. How did you guys get the scores aligned?",
    "1091593": "Interesting. So you think the main thing was a large batch size that contributed to better results?",
    "1091598": "That was my theory, and it did help, but to be honest I am not sure if that is the only explanation.",
    "1091601": "WOW! Batch size 512! That is huge!",
    "1091603": "Yeah, I think larger batch size should give you better result in long run. Even though it train slower (since you have less backprop per samples)",
    "1091618": "Thanks for sharing. It seems like large batch size is the key factor.",
    "1091628": "I noticed that going from 32 to 16 bs really hurt results. I was also splitting across 2 gpus so my effective batch size was even smaller than that. I guess I should've tried even larger.",
    "1091668": "Same, I tried BS 16 which affected results in a negative way. Did not try anything above 32...honestly, skimming through some recent write ups, I see that experimenting with hyperparams and backbones would have been extremely beneficial :) Most of us probably rushed into implementing complex ideas without getting some \"basics\" right\n\n@frankpanxj  Congrats on your result and thank you for sharing!",
    "1091676": "I implemented so many papers to no avail. Finding out that batch size might have been my limiting factor is very disappointing. Lesson learned if that is the case.",
    "1091684": "We live and we learn....but I can say that regardless of the result, the amount we learned is priceless. Armed with new knowledge and experience, we wont make the same mistakes again in future competitions;)",
    "1091742": "Batch size did not matter too much for us I believe, but BS is always an interplay with LR and mostly we used BS >= 64",
    "1091758": "philippsinger what's your default with BS and LR (pick your optimizer). \n\nI start with Adam and BS=32 and set LR between 5e-5 to 5e-4 depending on my assumption how valuable are imagenet features for the domain task. Changes in BS and LR lean toward directly proportional",
    "1091765": "Nevertheless congratulations is good learning. I did come late to the competition and no time to train. I implemented a solution inspired by one of the top contenders in the nuscenes challenge (https://arxiv.org/pdf/2005.02545.pdf). Still training. I got the impression that for this particular competition computational power is important. I don't think I can fit anything beyond 224x224 in my 8GB GPU.",
    "1091779": "Computational power is important and allows you to iterate through a large number of ideas, but in this case that turned out to be not 100% necessary.\n\nAs we can see, @frankpanxj  was able to achieve a top-tier solution without any impressive hardware, but by thinking \"outside the box\" and trying some seemingly simple things that others wouldn't even think of trying....and waiting patiently for 15 days;)\nThis is what I really like about this solution.",
    "1091791": "ture05 I also implemented the joint multi head attention. Looked reasonably promising but seemed to converge to a similar result to my more vanilla models. Seems I did not train for long enough though so maybe it does eventually converge to a better result.",
    "1091794": "It really depends on the problem, but for us a good start here was Adam BS 64 and LR 1e-4. I always use linear or cosine decay with rare exceptions.",
    "1091804": "Kudos to @frankpanxj  definitely computational power and patience is what's needed. \n\n@ryches How far did you go? Currently 50k iter ( 64 bs) and it's outperforming the baseline (for the same iterations in training). I did add a few things, more car states and my encoding was slightly different.",
    "1091807": "I allowed it to run for about 24 hours, not sure exactly the number of samples it was exposed to, but at that point it had converged to 17.5 validation loss.",
    "1091852": "Thanks for sharing. And, congrats.\nCould you tell us how your loss was getting down?  Even if you have log of loss, please show us.\nOur team might give up training too early...",
    "1091929": "Thanks for your kind words @indswetrust  😂. The only reason I could hold out that long was just because I was observing steadly decreasing loss every day, so there is no reason to stop.",
    "1091933": "iter\ttraining_loss (avg)\teval_loss\n10000\t23.5\t       27.526\n20000\t20.01\t23.999\n30000\t16.69\t20.397\n40000\t17.12\t18.814\n50000\t15.25\t17.781\n60000\t13.1\t       15.989\n70000\t13.93\t15.538\n80000\t14.29\t15.311\n90000\t12.74\t15.17\n100000\t12.38\t14.593\n110000\t12.54\t14.021\n120000\t11.53\t13.705\n130000\t12.41\t13.868\n140000\t11.81\t13.485\n150000\t10.79\t13.321\n160000\t10.92\t12.983\n170000\t10.81\t12.987\n180000\t10.98\t13.04\n190000\t11.83\t12.893\n200000\t10.94\t12.806\n210000\t10.7\t        12.63\n220000\t10.48\t12.666\n230000\t10.6\t         12.527\n240000   No data, the machine shut down for no reason 😂\t\t\n250000\t10.43\t12.555\n260000\t10.46\t12.409\n270000\t10.69\t12.426\n280000\t10.35\t12.376\n290000\t10.55\t12.347\n300000\t9.92\t        12.342\n310000\t11.07\t12.314\n320000\t10.55\t12.3\n330000\t9.94\t12.317\nhere I started the second round just before the competition deadline, restarted from lr of 0.00001 (drop by half every 20000 steps)\n340000\t10.67\t12.465\n350000\t11.33\t12.385\n360000\t10.1\t        12.352\n370000\t10.12\t12.295\n380000\t10.83\t12.269\n390000\t10.29\t12.232\nduring the final steps I checkpointed every 3000 steps\n393000\t9.98\t      12.2\n396000\t10.38\t12.21\n399000\t10.5\t        12.189\n402000\t10.66\t12.156\n405000\t10.25\t12.241\n408000\t10.72\t12.168",
    "1091999": "Congrats and thanks for mention!\nWe also discussed your 2nd thought of predicting probability distribution directly and assign 3 trajectory in later step, which may be an interesting approach for future!",
    "1092120": "Thanks!\nGreat patience you have."
  },
  "source": "meta"
}