{
  "id": 195781,
  "title": "Question: How many samples to train on to get competitive results?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/195781",
  "author_name": "",
  "post_date": "2020-11-07T12:08:42.126716700Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi I have the basic ResNet34 model as in the popular notebooks in this challenge by <a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> and <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and I trained it on 100,000 iterations (batch_size = 16) and got an LB score of 78.xx. I wanted to ask -</p>\n<ol>\n<li>Does this score seem justifiable keeping in mind that I trained on 1,600,000 samples (out of  the ~22.5M samples in AgentDataset)?</li>\n<li>How many samples does one usually train on (let's assume the model to be this resnet34 only) in this competition to get close to the 23.xx score or better? </li>\n</ol>",
  "messages": [
    {
      "id": "1071788",
      "postDate": "11/07/2020 12:08:42",
      "content": "<p>Hi I have the basic ResNet34 model as in the popular notebooks in this challenge by <a href=\"https://www.kaggle.com/huanvo\" target=\"_blank\">@huanvo</a> and <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> and I trained it on 100,000 iterations (batch_size = 16) and got an LB score of 78.xx. I wanted to ask -</p>\n<ol>\n<li>Does this score seem justifiable keeping in mind that I trained on 1,600,000 samples (out of  the ~22.5M samples in AgentDataset)?</li>\n<li>How many samples does one usually train on (let's assume the model to be this resnet34 only) in this competition to get close to the 23.xx score or better? </li>\n</ol>",
      "rawMarkdown": "Hi I have the basic ResNet34 model as in the popular notebooks in this challenge by @huanvo and @corochann and I trained it on 100,000 iterations (batch_size = 16) and got an LB score of 78.xx. I wanted to ask -\n1. Does this score seem justifiable keeping in mind that I trained on 1,600,000 samples (out of  the ~22.5M samples in AgentDataset)?\n2. How many samples does one usually train on (let's assume the model to be this resnet34 only) in this competition to get close to the 23.xx score or better?",
      "votes": null
    },
    {
      "id": "1075672",
      "postDate": "11/11/2020 21:42:23",
      "content": "<p>200k for me</p>",
      "rawMarkdown": "200k for me",
      "votes": null
    },
    {
      "id": "1075677",
      "postDate": "11/11/2020 21:47:27",
      "content": "<p>Do you mean 200k samples or 200k iterations? And if you are talking about iterations which batch size did you chose?</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Do you mean 200k samples or 200k iterations? And if you are talking about iterations which batch size did you chose?\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "1075696",
      "postDate": "11/11/2020 22:27:36",
      "content": "<p>batch_size=40 200k iterations</p>",
      "rawMarkdown": "batch_size=40 200k iterations",
      "votes": null
    },
    {
      "id": "1075697",
      "postDate": "11/11/2020 22:28:32",
      "content": "<p>i think how much iteration is not importance, how you processing your data is more focus</p>",
      "rawMarkdown": "i think how much iteration is not importance, how you processing your data is more focus",
      "votes": null
    },
    {
      "id": "1075699",
      "postDate": "11/11/2020 22:35:49",
      "content": "<p><a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> Your current LB (14.808) is from 200k (8M sample)?</p>",
      "rawMarkdown": "doanquanvietnamca Your current LB (14.808) is from 200k (8M sample)?",
      "votes": null
    },
    {
      "id": "1075715",
      "postDate": "11/11/2020 23:08:38",
      "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  yeah. </p>",
      "rawMarkdown": "pestipeti  yeah.",
      "votes": null
    },
    {
      "id": "1075878",
      "postDate": "11/12/2020 03:50:47",
      "content": "<p>I have tried a 300k iterations (batch_size 16) on the original ResNet34, I got LB score of 36.878.<br>\nSecond try 1.2M iterations (batch_size 16) gave me LB score 23.xxx.<br>\nFinal try 24M samples gave me LB 17.549.</p>",
      "rawMarkdown": "I have tried a 300k iterations (batch_size 16) on the original ResNet34, I got LB score of 36.878.\nSecond try 1.2M iterations (batch_size 16) gave me LB score 23.xxx.\nFinal try 24M samples gave me LB 17.549.",
      "votes": null
    },
    {
      "id": "1082134",
      "postDate": "11/17/2020 16:19:12",
      "content": "<p><a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> your LB is great wrt amount of samples you trained! Can you give some insights on the data processing?… some example maybe…</p>\n<blockquote>\n  <p>how you processing your data is more focus</p>\n</blockquote>",
      "rawMarkdown": "doanquanvietnamca your LB is great wrt amount of samples you trained! Can you give some insights on the data processing?... some example maybe...\n>how you processing your data is more focus",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1075672,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "11/11/2020 21:42:23",
      "content": "<p>200k for me</p>",
      "votes": null,
      "replies": [
        {
          "id": 1075677,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "11/11/2020 21:47:27",
          "content": "<p>Do you mean 200k samples or 200k iterations? And if you are talking about iterations which batch size did you chose?</p>\n<p>Thanks in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075696,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "11/11/2020 22:27:36",
          "content": "<p>batch_size=40 200k iterations</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075697,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "11/11/2020 22:28:32",
          "content": "<p>i think how much iteration is not importance, how you processing your data is more focus</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075699,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/11/2020 22:35:49",
          "content": "<p><a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> Your current LB (14.808) is from 200k (8M sample)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1075715,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "11/11/2020 23:08:38",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>  yeah. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1082134,
          "author_name": "pawankumarsahu",
          "author_url": "",
          "post_date": "11/17/2020 16:19:12",
          "content": "<p><a href=\"https://www.kaggle.com/doanquanvietnamca\" target=\"_blank\">@doanquanvietnamca</a> your LB is great wrt amount of samples you trained! Can you give some insights on the data processing?… some example maybe…</p>\n<blockquote>\n  <p>how you processing your data is more focus</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1075878,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "11/12/2020 03:50:47",
      "content": "<p>I have tried a 300k iterations (batch_size 16) on the original ResNet34, I got LB score of 36.878.<br>\nSecond try 1.2M iterations (batch_size 16) gave me LB score 23.xxx.<br>\nFinal try 24M samples gave me LB 17.549.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1071788": "Hi I have the basic ResNet34 model as in the popular notebooks in this challenge by @huanvo and @corochann and I trained it on 100,000 iterations (batch_size = 16) and got an LB score of 78.xx. I wanted to ask -\n1. Does this score seem justifiable keeping in mind that I trained on 1,600,000 samples (out of  the ~22.5M samples in AgentDataset)?\n2. How many samples does one usually train on (let's assume the model to be this resnet34 only) in this competition to get close to the 23.xx score or better?",
    "1075672": "200k for me",
    "1075677": "Do you mean 200k samples or 200k iterations? And if you are talking about iterations which batch size did you chose?\n\nThanks in advance!",
    "1075696": "batch_size=40 200k iterations",
    "1075697": "i think how much iteration is not importance, how you processing your data is more focus",
    "1075699": "doanquanvietnamca Your current LB (14.808) is from 200k (8M sample)?",
    "1075715": "pestipeti  yeah.",
    "1075878": "I have tried a 300k iterations (batch_size 16) on the original ResNet34, I got LB score of 36.878.\nSecond try 1.2M iterations (batch_size 16) gave me LB score 23.xxx.\nFinal try 24M samples gave me LB 17.549.",
    "1082134": "doanquanvietnamca your LB is great wrt amount of samples you trained! Can you give some insights on the data processing?... some example maybe...\n>how you processing your data is more focus"
  },
  "source": "meta"
}