{
  "id": 199497,
  "title": "The baseline model seems doing pretty well",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/199497",
  "author_name": "",
  "post_date": "2020-11-26T00:33:37.597081200Z",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First, Congratulations to all the winners!<br>\nI am really looking forward to all the winner's strategies, but at the same time, just want to share the baseline method at around 17.x. </p>\n<p>It is kind of embarrassed that the best model I got is the very same simple original Resnet34 model on the public notebook (and probably the same as that in the original lyft tutorial notebook). None of my other change to the network really works well.</p>\n<p>The only difference that brought me from the original 23.6xx to 17.x is by switching to <code>batch_sizse=64</code>, lowering the learning rate to 1e-4, using the one cycle policy lr scheduler with SGD instead of Adam. And just be patient not to kill the training before it even finish one epoch ;)<br>\nThe total train time is about 35 hours on my machine since on my Windows machine, I can only afford 6 workers, or the memory will explode despite having 64GB of ram. (Oh well, Windows… Time to install Linux)</p>\n<p>Other trick I found is that using <code>pin_memory=True</code> with <code>prefetch_factor=32</code> on the DataLoder makes thing train faster. During inference, I set <code>pin_memory=False</code>, but <code>batch_size=128</code>, <code>prefetch_factor=4</code> allows it to finish inference in like 7 mins.</p>\n<p>I think I have learned a lot of pytorch and numpy on this competition! Thanks Lyft for having such nice competition and dataset!</p>",
  "messages": [
    {
      "id": "1091346",
      "postDate": "11/26/2020 00:33:37",
      "content": "<p>First, Congratulations to all the winners!<br>\nI am really looking forward to all the winner's strategies, but at the same time, just want to share the baseline method at around 17.x. </p>\n<p>It is kind of embarrassed that the best model I got is the very same simple original Resnet34 model on the public notebook (and probably the same as that in the original lyft tutorial notebook). None of my other change to the network really works well.</p>\n<p>The only difference that brought me from the original 23.6xx to 17.x is by switching to <code>batch_sizse=64</code>, lowering the learning rate to 1e-4, using the one cycle policy lr scheduler with SGD instead of Adam. And just be patient not to kill the training before it even finish one epoch ;)<br>\nThe total train time is about 35 hours on my machine since on my Windows machine, I can only afford 6 workers, or the memory will explode despite having 64GB of ram. (Oh well, Windows… Time to install Linux)</p>\n<p>Other trick I found is that using <code>pin_memory=True</code> with <code>prefetch_factor=32</code> on the DataLoder makes thing train faster. During inference, I set <code>pin_memory=False</code>, but <code>batch_size=128</code>, <code>prefetch_factor=4</code> allows it to finish inference in like 7 mins.</p>\n<p>I think I have learned a lot of pytorch and numpy on this competition! Thanks Lyft for having such nice competition and dataset!</p>",
      "rawMarkdown": "First, Congratulations to all the winners!\nI am really looking forward to all the winner's strategies, but at the same time, just want to share the baseline method at around 17.x. \n\nIt is kind of embarrassed that the best model I got is the very same simple original Resnet34 model on the public notebook (and probably the same as that in the original lyft tutorial notebook). None of my other change to the network really works well.\n\nThe only difference that brought me from the original 23.6xx to 17.x is by switching to `batch_sizse=64`, lowering the learning rate to 1e-4, using the one cycle policy lr scheduler with SGD instead of Adam. And just be patient not to kill the training before it even finish one epoch ;)\nThe total train time is about 35 hours on my machine since on my Windows machine, I can only afford 6 workers, or the memory will explode despite having 64GB of ram. (Oh well, Windows... Time to install Linux)\n\nOther trick I found is that using `pin_memory=True` with `prefetch_factor=32` on the DataLoder makes thing train faster. During inference, I set `pin_memory=False`, but `batch_size=128`, `prefetch_factor=4` allows it to finish inference in like 7 mins.\n\nI think I have learned a lot of pytorch and numpy on this competition! Thanks Lyft for having such nice competition and dataset!",
      "votes": null
    },
    {
      "id": "1091354",
      "postDate": "11/26/2020 00:40:32",
      "content": "<p>set history=2 and image size to 256x256.<br>\nthe baseline CNN resnet34 open kernel can achieve this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcf89cdee217bff0f0da6e6e2120c8f53%2FSelection_031.png?generation=1606351230804902&amp;alt=media\" alt=\"\"></p>\n<p>this is fast to train because only 2 history frame needs to be rendered</p>",
      "rawMarkdown": "set history=2 and image size to 256x256.\nthe baseline CNN resnet34 open kernel can achieve this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcf89cdee217bff0f0da6e6e2120c8f53%2FSelection_031.png?generation=1606351230804902&alt=media)\n\nthis is fast to train because only 2 history frame needs to be rendered",
      "votes": null
    },
    {
      "id": "1091356",
      "postDate": "11/26/2020 00:43:06",
      "content": "<p>I see. I think I tried that but I wasn't patient to wait until it finish one epoch since the validation score was doing very bad on the first 2/5 of epoch. I guess I should have try the entire epoch.<br>\nAre you using <code>batch_size=256</code>?</p>",
      "rawMarkdown": "I see. I think I tried that but I wasn't patient to wait until it finish one epoch since the validation score was doing very bad on the first 2/5 of epoch. I guess I should have try the entire epoch.\nAre you using `batch_size=256`?",
      "votes": null
    },
    {
      "id": "1091369",
      "postDate": "11/26/2020 00:58:15",
      "content": "<p>i am using 128</p>",
      "rawMarkdown": "i am using 128",
      "votes": null
    },
    {
      "id": "1091413",
      "postDate": "11/26/2020 02:08:32",
      "content": "<p>I also utilized public baselines. But overfit problem soon surfaces. In my case the dropout technique help solve that problem by A LOT. I used dropout rate of 0.3 and got here. Looking forward to see others solutions.</p>",
      "rawMarkdown": "I also utilized public baselines. But overfit problem soon surfaces. In my case the dropout technique help solve that problem by A LOT. I used dropout rate of 0.3 and got here. Looking forward to see others solutions.",
      "votes": null
    },
    {
      "id": "1091443",
      "postDate": "11/26/2020 03:05:28",
      "content": "<p>Impressive, is this with 1 epoch training on train zarr or full train zarr or something else?</p>",
      "rawMarkdown": "Impressive, is this with 1 epoch training on train zarr or full train zarr or something else?",
      "votes": null
    },
    {
      "id": "1091449",
      "postDate": "11/26/2020 03:11:36",
      "content": "<p>Congrats, a medal is a medal no matter what your solution is ;) What training data you used? Any sampling? Also how many epochs have you trained for?</p>",
      "rawMarkdown": "Congrats, a medal is a medal no matter what your solution is ;) What training data you used? Any sampling? Also how many epochs have you trained for?",
      "votes": null
    },
    {
      "id": "1091460",
      "postDate": "11/26/2020 03:28:26",
      "content": "<p>Thanks! I just used the simple train set and train for about 1 epoch (24M samples). I didn't have the time to try the train_full.</p>",
      "rawMarkdown": "Thanks! I just used the simple train set and train for about 1 epoch (24M samples). I didn't have the time to try the train_full.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091354,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/26/2020 00:40:32",
      "content": "<p>set history=2 and image size to 256x256.<br>\nthe baseline CNN resnet34 open kernel can achieve this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcf89cdee217bff0f0da6e6e2120c8f53%2FSelection_031.png?generation=1606351230804902&amp;alt=media\" alt=\"\"></p>\n<p>this is fast to train because only 2 history frame needs to be rendered</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091356,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/26/2020 00:43:06",
          "content": "<p>I see. I think I tried that but I wasn't patient to wait until it finish one epoch since the validation score was doing very bad on the first 2/5 of epoch. I guess I should have try the entire epoch.<br>\nAre you using <code>batch_size=256</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091369,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/26/2020 00:58:15",
          "content": "<p>i am using 128</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091443,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "11/26/2020 03:05:28",
          "content": "<p>Impressive, is this with 1 epoch training on train zarr or full train zarr or something else?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091413,
      "author_name": "ryanchun",
      "author_url": "",
      "post_date": "11/26/2020 02:08:32",
      "content": "<p>I also utilized public baselines. But overfit problem soon surfaces. In my case the dropout technique help solve that problem by A LOT. I used dropout rate of 0.3 and got here. Looking forward to see others solutions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1091449,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "11/26/2020 03:11:36",
      "content": "<p>Congrats, a medal is a medal no matter what your solution is ;) What training data you used? Any sampling? Also how many epochs have you trained for?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091460,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/26/2020 03:28:26",
          "content": "<p>Thanks! I just used the simple train set and train for about 1 epoch (24M samples). I didn't have the time to try the train_full.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1091346": "First, Congratulations to all the winners!\nI am really looking forward to all the winner's strategies, but at the same time, just want to share the baseline method at around 17.x. \n\nIt is kind of embarrassed that the best model I got is the very same simple original Resnet34 model on the public notebook (and probably the same as that in the original lyft tutorial notebook). None of my other change to the network really works well.\n\nThe only difference that brought me from the original 23.6xx to 17.x is by switching to `batch_sizse=64`, lowering the learning rate to 1e-4, using the one cycle policy lr scheduler with SGD instead of Adam. And just be patient not to kill the training before it even finish one epoch ;)\nThe total train time is about 35 hours on my machine since on my Windows machine, I can only afford 6 workers, or the memory will explode despite having 64GB of ram. (Oh well, Windows... Time to install Linux)\n\nOther trick I found is that using `pin_memory=True` with `prefetch_factor=32` on the DataLoder makes thing train faster. During inference, I set `pin_memory=False`, but `batch_size=128`, `prefetch_factor=4` allows it to finish inference in like 7 mins.\n\nI think I have learned a lot of pytorch and numpy on this competition! Thanks Lyft for having such nice competition and dataset!",
    "1091354": "set history=2 and image size to 256x256.\nthe baseline CNN resnet34 open kernel can achieve this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fcf89cdee217bff0f0da6e6e2120c8f53%2FSelection_031.png?generation=1606351230804902&alt=media)\n\nthis is fast to train because only 2 history frame needs to be rendered",
    "1091356": "I see. I think I tried that but I wasn't patient to wait until it finish one epoch since the validation score was doing very bad on the first 2/5 of epoch. I guess I should have try the entire epoch.\nAre you using `batch_size=256`?",
    "1091369": "i am using 128",
    "1091413": "I also utilized public baselines. But overfit problem soon surfaces. In my case the dropout technique help solve that problem by A LOT. I used dropout rate of 0.3 and got here. Looking forward to see others solutions.",
    "1091443": "Impressive, is this with 1 epoch training on train zarr or full train zarr or something else?",
    "1091449": "Congrats, a medal is a medal no matter what your solution is ;) What training data you used? Any sampling? Also how many epochs have you trained for?",
    "1091460": "Thanks! I just used the simple train set and train for about 1 epoch (24M samples). I didn't have the time to try the train_full."
  },
  "source": "meta"
}