{
  "id": 186921,
  "title": "How are you Validating?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/186921",
  "author_name": "",
  "post_date": "2020-09-26T15:02:28.074914800Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Given the bottleneck due to rasterization with limited number of cpu cores, how are you validating your model? Running validation on complete validation set provided will consume a lot of time. For every n step I am validating on a small sample of validation set but there could be better ways to do it.</p>\n<p>Even after validating on small sample I still use LB score as a reference to evaluate my model which I don't think is a very good idea. Any ideas on how we could build a decent validation pipeline?</p>",
  "messages": [
    {
      "id": "1028050",
      "postDate": "09/26/2020 15:02:28",
      "content": "<p>Given the bottleneck due to rasterization with limited number of cpu cores, how are you validating your model? Running validation on complete validation set provided will consume a lot of time. For every n step I am validating on a small sample of validation set but there could be better ways to do it.</p>\n<p>Even after validating on small sample I still use LB score as a reference to evaluate my model which I don't think is a very good idea. Any ideas on how we could build a decent validation pipeline?</p>",
      "rawMarkdown": "Given the bottleneck due to rasterization with limited number of cpu cores, how are you validating your model? Running validation on complete validation set provided will consume a lot of time. For every n step I am validating on a small sample of validation set but there could be better ways to do it.\n\nEven after validating on small sample I still use LB score as a reference to evaluate my model which I don't think is a very good idea. Any ideas on how we could build a decent validation pipeline?",
      "votes": null
    },
    {
      "id": "1028211",
      "postDate": "09/26/2020 17:10:22",
      "content": "<p>Use as much of the validation set as you want to spend time?</p>\n<p>Small portion for a first indication, more if the model seems to be good. And use all in the end when it comes to choosing your two final submissions?</p>\n<p>I personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.</p>",
      "rawMarkdown": "Use as much of the validation set as you want to spend time?\n\nSmall portion for a first indication, more if the model seems to be good. And use all in the end when it comes to choosing your two final submissions?\n\nI personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.",
      "votes": null
    },
    {
      "id": "1028404",
      "postDate": "09/26/2020 20:27:02",
      "content": "<blockquote>\n  <p>For every n step I am validating on a small sample of validation set but there could be better ways to do it.</p>\n</blockquote>\n<p>Im exactly doing the same thing. A higher usage of the validation set would reflect the performance of the model more accurate but due to the rasterization, I dont see any different strategies than this. </p>\n<p>How many frames do you use for evaluation?</p>",
      "rawMarkdown": "> For every n step I am validating on a small sample of validation set but there could be better ways to do it.\n\nIm exactly doing the same thing. A higher usage of the validation set would reflect the performance of the model more accurate but due to the rasterization, I dont see any different strategies than this. \n\nHow many frames do you use for evaluation?",
      "votes": null
    },
    {
      "id": "1028618",
      "postDate": "09/27/2020 03:46:02",
      "content": "<p>I guess, what I am looking for is a way to sampling validation set to get a good representation of our dataset which we can use to verify our training results. </p>\n<p><code>I personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.</code><br>\nI agree with this, but I am just looking for a good way to get a small sample from validation set which I could use to speed up my experiment validation.</p>",
      "rawMarkdown": "I guess, what I am looking for is a way to sampling validation set to get a good representation of our dataset which we can use to verify our training results. \n\n`I personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.`\nI agree with this, but I am just looking for a good way to get a small sample from validation set which I could use to speed up my experiment validation.",
      "votes": null
    },
    {
      "id": "1060509",
      "postDate": "10/26/2020 09:09:56",
      "content": "<p>Is there any way we can save the rasterized data instead of doing rasterization on the fly?</p>",
      "rawMarkdown": "Is there any way we can save the rasterized data instead of doing rasterization on the fly?",
      "votes": null
    },
    {
      "id": "1060539",
      "postDate": "10/26/2020 09:39:58",
      "content": "<p><a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a> <br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">Take a look at this post</a></p>",
      "rawMarkdown": "louis925 \n[Take a look at this post](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1028211,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "09/26/2020 17:10:22",
      "content": "<p>Use as much of the validation set as you want to spend time?</p>\n<p>Small portion for a first indication, more if the model seems to be good. And use all in the end when it comes to choosing your two final submissions?</p>\n<p>I personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1028618,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "09/27/2020 03:46:02",
          "content": "<p>I guess, what I am looking for is a way to sampling validation set to get a good representation of our dataset which we can use to verify our training results. </p>\n<p><code>I personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.</code><br>\nI agree with this, but I am just looking for a good way to get a small sample from validation set which I could use to speed up my experiment validation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1028404,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "09/26/2020 20:27:02",
      "content": "<blockquote>\n  <p>For every n step I am validating on a small sample of validation set but there could be better ways to do it.</p>\n</blockquote>\n<p>Im exactly doing the same thing. A higher usage of the validation set would reflect the performance of the model more accurate but due to the rasterization, I dont see any different strategies than this. </p>\n<p>How many frames do you use for evaluation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1060509,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "10/26/2020 09:09:56",
      "content": "<p>Is there any way we can save the rasterized data instead of doing rasterization on the fly?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1060539,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "10/26/2020 09:39:58",
          "content": "<p><a href=\"https://www.kaggle.com/louis925\" target=\"_blank\">@louis925</a> <br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637\" target=\"_blank\">Take a look at this post</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1028050": "Given the bottleneck due to rasterization with limited number of cpu cores, how are you validating your model? Running validation on complete validation set provided will consume a lot of time. For every n step I am validating on a small sample of validation set but there could be better ways to do it.\n\nEven after validating on small sample I still use LB score as a reference to evaluate my model which I don't think is a very good idea. Any ideas on how we could build a decent validation pipeline?",
    "1028211": "Use as much of the validation set as you want to spend time?\n\nSmall portion for a first indication, more if the model seems to be good. And use all in the end when it comes to choosing your two final submissions?\n\nI personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.",
    "1028404": "> For every n step I am validating on a small sample of validation set but there could be better ways to do it.\n\nIm exactly doing the same thing. A higher usage of the validation set would reflect the performance of the model more accurate but due to the rasterization, I dont see any different strategies than this. \n\nHow many frames do you use for evaluation?",
    "1028618": "I guess, what I am looking for is a way to sampling validation set to get a good representation of our dataset which we can use to verify our training results. \n\n`I personally love the giant dataset. Less prone to outliers, overfitting, lucky seeds etc.`\nI agree with this, but I am just looking for a good way to get a small sample from validation set which I could use to speed up my experiment validation.",
    "1060509": "Is there any way we can save the rasterized data instead of doing rasterization on the fly?",
    "1060539": "louis925 \n[Take a look at this post](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/177637)"
  },
  "source": "meta"
}