{
  "id": 191599,
  "title": "An incomplete summary of different approaches",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/191599",
  "author_name": "",
  "post_date": "2020-10-17T12:23:25.631597100Z",
  "votes": 50,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi, <br>\nsince I am waiting for another output of my prototyping notebook 😉, I thought it might be a good idea to summarize some topics in this challenge which might be interesting to follow up. It is currently becoming calmer and quieter, probably because everyone is testing intensively, so maybe this is a good moment to come up with some kind of a summary. Please forgive me if I don't list everything here. I have concentrated on the things that I personally found particularly exciting.</p>\n<p><strong>1. Config parameters</strong><br>\nThis is the most obvious one. Changing the params in the config can have a big impact on the score. Increasing the raster size should basically lead to better scores but is limited because of the huge amount of time needed for larger sizes. However, I experimented with the distribution of width and height quite a lot because of the findings of <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> in this thread:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323</a></p>\n<p>I tried to keep the full amount of pixels equal to benchmark different approaches against each other. If you have raster size 200<em>200, you have 40,000 pixel in sum. If you use a distribution of 55/45 for width and height instead of 50/50, you would have a raster size of 220</em>182 (-&gt; ~40k pixels). Following this idea, I have tried different combinations as:</p>\n<blockquote>\n  <p>50% – 50% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  55% – 45% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  60% – 40% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  65% – 35% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  70% – 30% width and height with 0.5, 0.3, 0.2 pixel size</p>\n</blockquote>\n<p>To be honest, I don´t see any pattern here at the moment. It seems that increasing the raster size in general improves your score, yes, but for the others parameters sometimes 0.5 is better, sometimes 0.3 and so on. I also don´t see any pattern for the distribution of width and height. But I have to admit that I use small test datasets (50k iterations with batch size 16) because of hardware limitations. </p>\n<p>I also tried resnet18, resnet 34 and resnet50 while leaving the other parameters unchanged but I didnt receive better scores with the more complex models. I didnt touch the other params yet. </p>\n<p><strong>2. Function of the rasterizer</strong><br>\nThe rasterizer is the bottleneck in this competition. We are so slow in training our little CNNs. There were some brilliant ideas to fix this issue. <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> started with the idea to pre-generate the images and save them compiled on hd but the speed gain wasnt that impressive and the images need a lot of space:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359</a></p>\n<p>Another idea was to optimize the rasterizer. You could rewrite the functions of the rasterizer in plain c-code to improve the speed. I didnt test it because I miss the required coding skills 😉.</p>\n<p><strong>3. Fixing the “confidences should sum to 1”-bug</strong><br>\nThis is not really a booster for your scores but still annoying and time consuming. Sometimes you will encounter the following error:</p>\n<blockquote>\n  <p>assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"<br>\n  AssertionError: confidences should sum to 1</p>\n</blockquote>\n<p>It seems that the gradients explode from time to time as described here:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773</a><br>\n<a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> provides also some possible solutions in the linked thread.</p>\n<p>However, I found another approach to fix this issue here:<br>\n<a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments</a><br>\n<a href=\"https://www.kaggle.com/loopdigga\" target=\"_blank\">@loopdigga</a> recommend the following approach in the commend section:</p>\n<blockquote>\n  <p>Possible solution for that could be adding some small epsilon to confidences.<br>\n  For example error = torch.log(confidences + 1e-10) - 0.5 * torch.sum(error, dim=-1)</p>\n</blockquote>\n<p><strong>4. Finding the right training strategy regarding the balance of speed / quality</strong><br>\n<a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> described some of her own strategies here:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/182787\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/182787</a><br>\nAs far as I understood she is using around of 12% of the test set which should be around 2.64 million samples.</p>\n<p>In my own tests I found that there is still a lot of variance in the mean avg loss even with 50 or 60 k iterations with a batch size of 16. So basically, we should test on larger subsets to validate our ideas.</p>\n<p>Another idea to improve the training time could be the use of weighted samples:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814</a></p>\n<p>Basically, we have a lot of scenes in our dataset in which you just need to extrapolate the coordinates straight away with the given velocity. It would therefore make sense to focus more on the samples that represent more unusual rides. An example to use weighted samples in pytorch can be found here:<br>\n<a href=\"https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\" target=\"_blank\">https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html</a></p>\n<p><strong>5. Augmentation</strong><br>\nAnother prominent topic related to CNN is augmentation. In this thread different ideas are discussed:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042</a></p>\n<p>I personally haven't tried anything in this direction yet, because I want to find out the correct parameters for the training before I dig deeper.</p>\n<p><strong>Conclusion</strong><br>\nFrom a beginner point of view this competition is quite challenging. Due to the huge amounts of data, all tests must be carefully weighed up, as there is a lack of resources at all time. Nevertheless, I find all the topics and approaches totally exciting. There is enough to do for beginners as well as experts and I am really curious which solutions the top kagglers will present in the end.</p>\n<p>Good luck to you all!</p>",
  "messages": [
    {
      "id": "1052178",
      "postDate": "10/17/2020 12:23:25",
      "content": "<p>Hi, <br>\nsince I am waiting for another output of my prototyping notebook 😉, I thought it might be a good idea to summarize some topics in this challenge which might be interesting to follow up. It is currently becoming calmer and quieter, probably because everyone is testing intensively, so maybe this is a good moment to come up with some kind of a summary. Please forgive me if I don't list everything here. I have concentrated on the things that I personally found particularly exciting.</p>\n<p><strong>1. Config parameters</strong><br>\nThis is the most obvious one. Changing the params in the config can have a big impact on the score. Increasing the raster size should basically lead to better scores but is limited because of the huge amount of time needed for larger sizes. However, I experimented with the distribution of width and height quite a lot because of the findings of <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> in this thread:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323</a></p>\n<p>I tried to keep the full amount of pixels equal to benchmark different approaches against each other. If you have raster size 200<em>200, you have 40,000 pixel in sum. If you use a distribution of 55/45 for width and height instead of 50/50, you would have a raster size of 220</em>182 (-&gt; ~40k pixels). Following this idea, I have tried different combinations as:</p>\n<blockquote>\n  <p>50% – 50% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  55% – 45% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  60% – 40% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  65% – 35% width and height with 0.5, 0.3, 0.2 pixel size<br>\n  70% – 30% width and height with 0.5, 0.3, 0.2 pixel size</p>\n</blockquote>\n<p>To be honest, I don´t see any pattern here at the moment. It seems that increasing the raster size in general improves your score, yes, but for the others parameters sometimes 0.5 is better, sometimes 0.3 and so on. I also don´t see any pattern for the distribution of width and height. But I have to admit that I use small test datasets (50k iterations with batch size 16) because of hardware limitations. </p>\n<p>I also tried resnet18, resnet 34 and resnet50 while leaving the other parameters unchanged but I didnt receive better scores with the more complex models. I didnt touch the other params yet. </p>\n<p><strong>2. Function of the rasterizer</strong><br>\nThe rasterizer is the bottleneck in this competition. We are so slow in training our little CNNs. There were some brilliant ideas to fix this issue. <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> started with the idea to pre-generate the images and save them compiled on hd but the speed gain wasnt that impressive and the images need a lot of space:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359</a></p>\n<p>Another idea was to optimize the rasterizer. You could rewrite the functions of the rasterizer in plain c-code to improve the speed. I didnt test it because I miss the required coding skills 😉.</p>\n<p><strong>3. Fixing the “confidences should sum to 1”-bug</strong><br>\nThis is not really a booster for your scores but still annoying and time consuming. Sometimes you will encounter the following error:</p>\n<blockquote>\n  <p>assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"<br>\n  AssertionError: confidences should sum to 1</p>\n</blockquote>\n<p>It seems that the gradients explode from time to time as described here:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773</a><br>\n<a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> provides also some possible solutions in the linked thread.</p>\n<p>However, I found another approach to fix this issue here:<br>\n<a href=\"https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments\" target=\"_blank\">https://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments</a><br>\n<a href=\"https://www.kaggle.com/loopdigga\" target=\"_blank\">@loopdigga</a> recommend the following approach in the commend section:</p>\n<blockquote>\n  <p>Possible solution for that could be adding some small epsilon to confidences.<br>\n  For example error = torch.log(confidences + 1e-10) - 0.5 * torch.sum(error, dim=-1)</p>\n</blockquote>\n<p><strong>4. Finding the right training strategy regarding the balance of speed / quality</strong><br>\n<a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> described some of her own strategies here:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/182787\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/182787</a><br>\nAs far as I understood she is using around of 12% of the test set which should be around 2.64 million samples.</p>\n<p>In my own tests I found that there is still a lot of variance in the mean avg loss even with 50 or 60 k iterations with a batch size of 16. So basically, we should test on larger subsets to validate our ideas.</p>\n<p>Another idea to improve the training time could be the use of weighted samples:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814</a></p>\n<p>Basically, we have a lot of scenes in our dataset in which you just need to extrapolate the coordinates straight away with the given velocity. It would therefore make sense to focus more on the samples that represent more unusual rides. An example to use weighted samples in pytorch can be found here:<br>\n<a href=\"https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\" target=\"_blank\">https://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html</a></p>\n<p><strong>5. Augmentation</strong><br>\nAnother prominent topic related to CNN is augmentation. In this thread different ideas are discussed:<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042</a></p>\n<p>I personally haven't tried anything in this direction yet, because I want to find out the correct parameters for the training before I dig deeper.</p>\n<p><strong>Conclusion</strong><br>\nFrom a beginner point of view this competition is quite challenging. Due to the huge amounts of data, all tests must be carefully weighed up, as there is a lack of resources at all time. Nevertheless, I find all the topics and approaches totally exciting. There is enough to do for beginners as well as experts and I am really curious which solutions the top kagglers will present in the end.</p>\n<p>Good luck to you all!</p>",
      "rawMarkdown": "Hi, \nsince I am waiting for another output of my prototyping notebook 😉, I thought it might be a good idea to summarize some topics in this challenge which might be interesting to follow up. It is currently becoming calmer and quieter, probably because everyone is testing intensively, so maybe this is a good moment to come up with some kind of a summary. Please forgive me if I don't list everything here. I have concentrated on the things that I personally found particularly exciting.\n\n**1. Config parameters**\nThis is the most obvious one. Changing the params in the config can have a big impact on the score. Increasing the raster size should basically lead to better scores but is limited because of the huge amount of time needed for larger sizes. However, I experimented with the distribution of width and height quite a lot because of the findings of @pestipeti in this thread:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\n\nI tried to keep the full amount of pixels equal to benchmark different approaches against each other. If you have raster size 200*200, you have 40,000 pixel in sum. If you use a distribution of 55/45 for width and height instead of 50/50, you would have a raster size of 220*182 (-> ~40k pixels). Following this idea, I have tried different combinations as:\n\n> 50% – 50% width and height with 0.5, 0.3, 0.2 pixel size\n55% – 45% width and height with 0.5, 0.3, 0.2 pixel size\n60% – 40% width and height with 0.5, 0.3, 0.2 pixel size\n65% – 35% width and height with 0.5, 0.3, 0.2 pixel size\n70% – 30% width and height with 0.5, 0.3, 0.2 pixel size\n\nTo be honest, I don´t see any pattern here at the moment. It seems that increasing the raster size in general improves your score, yes, but for the others parameters sometimes 0.5 is better, sometimes 0.3 and so on. I also don´t see any pattern for the distribution of width and height. But I have to admit that I use small test datasets (50k iterations with batch size 16) because of hardware limitations. \n\nI also tried resnet18, resnet 34 and resnet50 while leaving the other parameters unchanged but I didnt receive better scores with the more complex models. I didnt touch the other params yet. \n\n**2. Function of the rasterizer**\nThe rasterizer is the bottleneck in this competition. We are so slow in training our little CNNs. There were some brilliant ideas to fix this issue. @pestipeti started with the idea to pre-generate the images and save them compiled on hd but the speed gain wasnt that impressive and the images need a lot of space:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359\n\nAnother idea was to optimize the rasterizer. You could rewrite the functions of the rasterizer in plain c-code to improve the speed. I didnt test it because I miss the required coding skills 😉.\n\n**3. Fixing the “confidences should sum to 1”-bug**\nThis is not really a booster for your scores but still annoying and time consuming. Sometimes you will encounter the following error:\n> assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\nAssertionError: confidences should sum to 1\n\nIt seems that the gradients explode from time to time as described here:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\n@aliabdin1 provides also some possible solutions in the linked thread.\n\nHowever, I found another approach to fix this issue here:\nhttps://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments\n@loopdigga recommend the following approach in the commend section:\n> Possible solution for that could be adding some small epsilon to confidences.\nFor example error = torch.log(confidences + 1e-10) - 0.5 * torch.sum(error, dim=-1)\n\n**4. Finding the right training strategy regarding the balance of speed / quality**\n@fergusoci described some of her own strategies here:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/182787\nAs far as I understood she is using around of 12% of the test set which should be around 2.64 million samples.\n\nIn my own tests I found that there is still a lot of variance in the mean avg loss even with 50 or 60 k iterations with a batch size of 16. So basically, we should test on larger subsets to validate our ideas.\n\nAnother idea to improve the training time could be the use of weighted samples:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\n\nBasically, we have a lot of scenes in our dataset in which you just need to extrapolate the coordinates straight away with the given velocity. It would therefore make sense to focus more on the samples that represent more unusual rides. An example to use weighted samples in pytorch can be found here:\nhttps://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\n\n**5. Augmentation**\nAnother prominent topic related to CNN is augmentation. In this thread different ideas are discussed:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042\n\nI personally haven't tried anything in this direction yet, because I want to find out the correct parameters for the training before I dig deeper.\n\n**Conclusion**\nFrom a beginner point of view this competition is quite challenging. Due to the huge amounts of data, all tests must be carefully weighed up, as there is a lack of resources at all time. Nevertheless, I find all the topics and approaches totally exciting. There is enough to do for beginners as well as experts and I am really curious which solutions the top kagglers will present in the end.\n\nGood luck to you all!",
      "votes": null
    },
    {
      "id": "1052230",
      "postDate": "10/17/2020 13:39:49",
      "content": "<p>Thanks for your detailed summary!</p>\n<p>About 3. Fixing the “confidences should sum to 1”-bug<br>\nI think this is probably due to the floating point precision problem.</p>\n<blockquote>\n  <p>assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"<br>\n  AssertionError: confidences should sum to 1</p>\n</blockquote>\n<p><a href=\"https://pytorch.org/docs/stable/generated/torch.allclose.html\" target=\"_blank\">[Docs &gt; torch &gt; torch.allclose]</a><br>\n<code>torch.allclose(input, other, rtol=1e-05, atol=1e-08, equal_nan=False) → bool</code></p>\n<p>To modify 'atol' absolute tolerance might be helpful.<br>\nIn my case, I changed it from 1e-08 to 1e-06. (1e-07 sometimes make errors to my test )<br>\n<code>torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,)), atol=1e-06)</code></p>\n<p>And I try to find appropriate '<em>num_workers</em>', '<em>pin_memory</em>' of torch.utils.data.DataLoader to reduce train time.<br>\n<a href=\"https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813\" target=\"_blank\">[Discussion of pytorch.org : Guidelines for assigning num_workers to DataLoader]</a></p>",
      "rawMarkdown": "Thanks for your detailed summary!\n\nAbout 3. Fixing the “confidences should sum to 1”-bug\nI think this is probably due to the floating point precision problem.\n\n> assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\nAssertionError: confidences should sum to 1\n\n[[Docs > torch > torch.allclose]](https://pytorch.org/docs/stable/generated/torch.allclose.html)\n`torch.allclose(input, other, rtol=1e-05, atol=1e-08, equal_nan=False) → bool`\n\nTo modify 'atol' absolute tolerance might be helpful.\nIn my case, I changed it from 1e-08 to 1e-06. (1e-07 sometimes make errors to my test )\n`torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,)), atol=1e-06)`\n\n\nAnd I try to find appropriate '*num_workers*', '*pin_memory*' of torch.utils.data.DataLoader to reduce train time.\n[[Discussion of pytorch.org : Guidelines for assigning num_workers to DataLoader]](https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813)",
      "votes": null
    },
    {
      "id": "1052243",
      "postDate": "10/17/2020 13:54:03",
      "content": "<p>My intuition is that training for only 50k batches of size 16 is too little to get close to the optimum. My best model now was trained for 350k such batches, and while my training schedule is not good, I think even if I can optimize the training efficiency I wouldn't be able to squeeze it into less than 150k batches. Thank you for sharing your insights.</p>",
      "rawMarkdown": "My intuition is that training for only 50k batches of size 16 is too little to get close to the optimum. My best model now was trained for 350k such batches, and while my training schedule is not good, I think even if I can optimize the training efficiency I wouldn't be able to squeeze it into less than 150k batches. Thank you for sharing your insights.",
      "votes": null
    },
    {
      "id": "1052249",
      "postDate": "10/17/2020 14:02:16",
      "content": "<p>How long it took to train 350k <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>",
      "rawMarkdown": "How long it took to train 350k @zaharch",
      "votes": null
    },
    {
      "id": "1052255",
      "postDate": "10/17/2020 14:12:22",
      "content": "<p>It takes about 20 min per 10k, so it is about 12h. My machine is<br>\nGeForce RTX 2080-Ti, and 16 core CPU, 2 x Xeon E5-2620 v4. The bottle neck is CPU, similar to everyone. GPU utilization is 40-50%. I use 16 workers.</p>",
      "rawMarkdown": "It takes about 20 min per 10k, so it is about 12h. My machine is\nGeForce RTX 2080-Ti, and 16 core CPU, 2 x Xeon E5-2620 v4. The bottle neck is CPU, similar to everyone. GPU utilization is 40-50%. I use 16 workers.",
      "votes": null
    },
    {
      "id": "1052261",
      "postDate": "10/17/2020 14:19:46",
      "content": "<p>I agree, 50k iterations/batch size 16 is 800k samples which is quite small to even test ideas locally. </p>\n<p>For evaluation and LB submission, I found from my (limited) experiments that the largest decrease in loss occurs until about 7M samples. After that, the loss tends to flatten out (relatively), so it is not necessary to train past that point at the moment and occupy computation time.<br>\nHowever, as I mentioned, my experiments are limited since I do not have the best hardware/fastest training times…maybe people who consistently train using the full dataset may be able to provide better insight.</p>\n<p>Thank you for the write-up, <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> !</p>",
      "rawMarkdown": "I agree, 50k iterations/batch size 16 is 800k samples which is quite small to even test ideas locally. \n\nFor evaluation and LB submission, I found from my (limited) experiments that the largest decrease in loss occurs until about 7M samples. After that, the loss tends to flatten out (relatively), so it is not necessary to train past that point at the moment and occupy computation time.\nHowever, as I mentioned, my experiments are limited since I do not have the best hardware/fastest training times...maybe people who consistently train using the full dataset may be able to provide better insight.\n\nThank you for the write-up, @benbla !",
      "votes": null
    },
    {
      "id": "1060617",
      "postDate": "10/26/2020 12:00:21",
      "content": "<p>Thank you for your summary and it is a Brilliant job!</p>\n<p>I got some questions about the raster size. If we increasing the raster size and adjust the pixel size as mention in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\" target=\"_blank\">here</a>, is it necessary means that we need a bigger net for training because of the increasing information we feed in the network.</p>",
      "rawMarkdown": "Thank you for your summary and it is a Brilliant job!\n\nI got some questions about the raster size. If we increasing the raster size and adjust the pixel size as mention in [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323), is it necessary means that we need a bigger net for training because of the increasing information we feed in the network.",
      "votes": null
    },
    {
      "id": "1073994",
      "postDate": "11/10/2020 05:56:47",
      "content": "<p>hey guys! <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> <a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> <br>\ndid you trained resnet18 with this hyperparameters? or what model did you used?<br>\ni have tried resnet50 atm, but its not so good results with 30k iterations </p>",
      "rawMarkdown": "hey guys! @zaharch @indswetrust \ndid you trained resnet18 with this hyperparameters? or what model did you used?\ni have tried resnet50 atm, but its not so good results with 30k iterations",
      "votes": null
    },
    {
      "id": "1074014",
      "postDate": "11/10/2020 06:34:38",
      "content": "<p><a href=\"https://www.kaggle.com/germanjke\" target=\"_blank\">@germanjke</a> that is because 30k iterations is far too small - less than a million samples with batch size 32. As discussed by us above, that is too small to even test ideas locally, you will need to train longer to see what works and what does not work.<br>\nresnet50 - haven't tried it so can't comment on it.</p>",
      "rawMarkdown": "germanjke that is because 30k iterations is far too small - less than a million samples with batch size 32. As discussed by us above, that is too small to even test ideas locally, you will need to train longer to see what works and what does not work.\nresnet50 - haven't tried it so can't comment on it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1052230,
      "author_name": "sunghyunjun",
      "author_url": "",
      "post_date": "10/17/2020 13:39:49",
      "content": "<p>Thanks for your detailed summary!</p>\n<p>About 3. Fixing the “confidences should sum to 1”-bug<br>\nI think this is probably due to the floating point precision problem.</p>\n<blockquote>\n  <p>assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"<br>\n  AssertionError: confidences should sum to 1</p>\n</blockquote>\n<p><a href=\"https://pytorch.org/docs/stable/generated/torch.allclose.html\" target=\"_blank\">[Docs &gt; torch &gt; torch.allclose]</a><br>\n<code>torch.allclose(input, other, rtol=1e-05, atol=1e-08, equal_nan=False) → bool</code></p>\n<p>To modify 'atol' absolute tolerance might be helpful.<br>\nIn my case, I changed it from 1e-08 to 1e-06. (1e-07 sometimes make errors to my test )<br>\n<code>torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,)), atol=1e-06)</code></p>\n<p>And I try to find appropriate '<em>num_workers</em>', '<em>pin_memory</em>' of torch.utils.data.DataLoader to reduce train time.<br>\n<a href=\"https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813\" target=\"_blank\">[Discussion of pytorch.org : Guidelines for assigning num_workers to DataLoader]</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1052243,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "10/17/2020 13:54:03",
      "content": "<p>My intuition is that training for only 50k batches of size 16 is too little to get close to the optimum. My best model now was trained for 350k such batches, and while my training schedule is not good, I think even if I can optimize the training efficiency I wouldn't be able to squeeze it into less than 150k batches. Thank you for sharing your insights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1052249,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "10/17/2020 14:02:16",
          "content": "<p>How long it took to train 350k <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1052255,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "10/17/2020 14:12:22",
          "content": "<p>It takes about 20 min per 10k, so it is about 12h. My machine is<br>\nGeForce RTX 2080-Ti, and 16 core CPU, 2 x Xeon E5-2620 v4. The bottle neck is CPU, similar to everyone. GPU utilization is 40-50%. I use 16 workers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1052261,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "10/17/2020 14:19:46",
          "content": "<p>I agree, 50k iterations/batch size 16 is 800k samples which is quite small to even test ideas locally. </p>\n<p>For evaluation and LB submission, I found from my (limited) experiments that the largest decrease in loss occurs until about 7M samples. After that, the loss tends to flatten out (relatively), so it is not necessary to train past that point at the moment and occupy computation time.<br>\nHowever, as I mentioned, my experiments are limited since I do not have the best hardware/fastest training times…maybe people who consistently train using the full dataset may be able to provide better insight.</p>\n<p>Thank you for the write-up, <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1073994,
          "author_name": "germanjke",
          "author_url": "",
          "post_date": "11/10/2020 05:56:47",
          "content": "<p>hey guys! <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> <a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> <br>\ndid you trained resnet18 with this hyperparameters? or what model did you used?<br>\ni have tried resnet50 atm, but its not so good results with 30k iterations </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1074014,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "11/10/2020 06:34:38",
          "content": "<p><a href=\"https://www.kaggle.com/germanjke\" target=\"_blank\">@germanjke</a> that is because 30k iterations is far too small - less than a million samples with batch size 32. As discussed by us above, that is too small to even test ideas locally, you will need to train longer to see what works and what does not work.<br>\nresnet50 - haven't tried it so can't comment on it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1060617,
      "author_name": "spicychicken38",
      "author_url": "",
      "post_date": "10/26/2020 12:00:21",
      "content": "<p>Thank you for your summary and it is a Brilliant job!</p>\n<p>I got some questions about the raster size. If we increasing the raster size and adjust the pixel size as mention in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\" target=\"_blank\">here</a>, is it necessary means that we need a bigger net for training because of the increasing information we feed in the network.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1052178": "Hi, \nsince I am waiting for another output of my prototyping notebook 😉, I thought it might be a good idea to summarize some topics in this challenge which might be interesting to follow up. It is currently becoming calmer and quieter, probably because everyone is testing intensively, so maybe this is a good moment to come up with some kind of a summary. Please forgive me if I don't list everything here. I have concentrated on the things that I personally found particularly exciting.\n\n**1. Config parameters**\nThis is the most obvious one. Changing the params in the config can have a big impact on the score. Increasing the raster size should basically lead to better scores but is limited because of the huge amount of time needed for larger sizes. However, I experimented with the distribution of width and height quite a lot because of the findings of @pestipeti in this thread:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323\n\nI tried to keep the full amount of pixels equal to benchmark different approaches against each other. If you have raster size 200*200, you have 40,000 pixel in sum. If you use a distribution of 55/45 for width and height instead of 50/50, you would have a raster size of 220*182 (-> ~40k pixels). Following this idea, I have tried different combinations as:\n\n> 50% – 50% width and height with 0.5, 0.3, 0.2 pixel size\n55% – 45% width and height with 0.5, 0.3, 0.2 pixel size\n60% – 40% width and height with 0.5, 0.3, 0.2 pixel size\n65% – 35% width and height with 0.5, 0.3, 0.2 pixel size\n70% – 30% width and height with 0.5, 0.3, 0.2 pixel size\n\nTo be honest, I don´t see any pattern here at the moment. It seems that increasing the raster size in general improves your score, yes, but for the others parameters sometimes 0.5 is better, sometimes 0.3 and so on. I also don´t see any pattern for the distribution of width and height. But I have to admit that I use small test datasets (50k iterations with batch size 16) because of hardware limitations. \n\nI also tried resnet18, resnet 34 and resnet50 while leaving the other parameters unchanged but I didnt receive better scores with the more complex models. I didnt touch the other params yet. \n\n**2. Function of the rasterizer**\nThe rasterizer is the bottleneck in this competition. We are so slow in training our little CNNs. There were some brilliant ideas to fix this issue. @pestipeti started with the idea to pre-generate the images and save them compiled on hd but the speed gain wasnt that impressive and the images need a lot of space:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180359\n\nAnother idea was to optimize the rasterizer. You could rewrite the functions of the rasterizer in plain c-code to improve the speed. I didnt test it because I miss the required coding skills 😉.\n\n**3. Fixing the “confidences should sum to 1”-bug**\nThis is not really a booster for your scores but still annoying and time consuming. Sometimes you will encounter the following error:\n> assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\nAssertionError: confidences should sum to 1\n\nIt seems that the gradients explode from time to time as described here:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/187773\n@aliabdin1 provides also some possible solutions in the linked thread.\n\nHowever, I found another approach to fix this issue here:\nhttps://www.kaggle.com/corochann/lyft-training-with-multi-mode-confidence/comments\n@loopdigga recommend the following approach in the commend section:\n> Possible solution for that could be adding some small epsilon to confidences.\nFor example error = torch.log(confidences + 1e-10) - 0.5 * torch.sum(error, dim=-1)\n\n**4. Finding the right training strategy regarding the balance of speed / quality**\n@fergusoci described some of her own strategies here:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/182787\nAs far as I understood she is using around of 12% of the test set which should be around 2.64 million samples.\n\nIn my own tests I found that there is still a lot of variance in the mean avg loss even with 50 or 60 k iterations with a batch size of 16. So basically, we should test on larger subsets to validate our ideas.\n\nAnother idea to improve the training time could be the use of weighted samples:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/183814\n\nBasically, we have a lot of scenes in our dataset in which you just need to extrapolate the coordinates straight away with the given velocity. It would therefore make sense to focus more on the samples that represent more unusual rides. An example to use weighted samples in pytorch can be found here:\nhttps://pytorch.org/docs/1.1.0/_modules/torch/utils/data/sampler.html\n\n**5. Augmentation**\nAnother prominent topic related to CNN is augmentation. In this thread different ideas are discussed:\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/191042\n\nI personally haven't tried anything in this direction yet, because I want to find out the correct parameters for the training before I dig deeper.\n\n**Conclusion**\nFrom a beginner point of view this competition is quite challenging. Due to the huge amounts of data, all tests must be carefully weighed up, as there is a lack of resources at all time. Nevertheless, I find all the topics and approaches totally exciting. There is enough to do for beginners as well as experts and I am really curious which solutions the top kagglers will present in the end.\n\nGood luck to you all!",
    "1052230": "Thanks for your detailed summary!\n\nAbout 3. Fixing the “confidences should sum to 1”-bug\nI think this is probably due to the floating point precision problem.\n\n> assert torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,))), \"confidences should sum to 1\"\nAssertionError: confidences should sum to 1\n\n[[Docs > torch > torch.allclose]](https://pytorch.org/docs/stable/generated/torch.allclose.html)\n`torch.allclose(input, other, rtol=1e-05, atol=1e-08, equal_nan=False) → bool`\n\nTo modify 'atol' absolute tolerance might be helpful.\nIn my case, I changed it from 1e-08 to 1e-06. (1e-07 sometimes make errors to my test )\n`torch.allclose(torch.sum(confidences, dim=1), confidences.new_ones((batch_size,)), atol=1e-06)`\n\n\nAnd I try to find appropriate '*num_workers*', '*pin_memory*' of torch.utils.data.DataLoader to reduce train time.\n[[Discussion of pytorch.org : Guidelines for assigning num_workers to DataLoader]](https://discuss.pytorch.org/t/guidelines-for-assigning-num-workers-to-dataloader/813)",
    "1052243": "My intuition is that training for only 50k batches of size 16 is too little to get close to the optimum. My best model now was trained for 350k such batches, and while my training schedule is not good, I think even if I can optimize the training efficiency I wouldn't be able to squeeze it into less than 150k batches. Thank you for sharing your insights.",
    "1052249": "How long it took to train 350k @zaharch",
    "1052255": "It takes about 20 min per 10k, so it is about 12h. My machine is\nGeForce RTX 2080-Ti, and 16 core CPU, 2 x Xeon E5-2620 v4. The bottle neck is CPU, similar to everyone. GPU utilization is 40-50%. I use 16 workers.",
    "1052261": "I agree, 50k iterations/batch size 16 is 800k samples which is quite small to even test ideas locally. \n\nFor evaluation and LB submission, I found from my (limited) experiments that the largest decrease in loss occurs until about 7M samples. After that, the loss tends to flatten out (relatively), so it is not necessary to train past that point at the moment and occupy computation time.\nHowever, as I mentioned, my experiments are limited since I do not have the best hardware/fastest training times...maybe people who consistently train using the full dataset may be able to provide better insight.\n\nThank you for the write-up, @benbla !",
    "1060617": "Thank you for your summary and it is a Brilliant job!\n\nI got some questions about the raster size. If we increasing the raster size and adjust the pixel size as mention in [here](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/178323), is it necessary means that we need a bigger net for training because of the increasing information we feed in the network.",
    "1073994": "hey guys! @zaharch @indswetrust \ndid you trained resnet18 with this hyperparameters? or what model did you used?\ni have tried resnet50 atm, but its not so good results with 30k iterations",
    "1074014": "germanjke that is because 30k iterations is far too small - less than a million samples with batch size 32. As discussed by us above, that is too small to even test ideas locally, you will need to train longer to see what works and what does not work.\nresnet50 - haven't tried it so can't comment on it."
  },
  "source": "meta"
}