{
  "id": 187780,
  "title": "Training masks for faster prototyping",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/187780",
  "author_name": "",
  "post_date": "2020-09-30T08:58:37.124507200Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all, </p>\n<p>I created a mask for the training set that some of you might find useful, you can find it here: <a href=\"https://www.kaggle.com/fnands/training-mask-100-200-10\" target=\"_blank\">link to dataset</a>. </p>\n<p>Additionally, if you want to change the parameters and make your own mask the notebook that created it can be found here: <a href=\"https://www.kaggle.com/fnands/makesubmasks\" target=\"_blank\">notebook to make masks</a></p>\n<h2>Rationale:</h2>\n<p>As the data has a sampling rate of 10Hz, the world doesn't change much from one frame to the other. This means that there are many near duplicates in the normal dataset.     <br>\nSeeing as most of us are still prototyping our methods, this means we are spending a lot of time running over data that is very similar, and with such a large dataset as we have here that can severely slow down our testing.  </p>\n<p>This mask picks out a subset of the training data that looks similar to our test set and that is lower in near-duplicates than the full training set.   </p>\n<p><strong>Let me be clear on this point</strong>: I think this is a useful tool for prototyping, as it will pick out a selection of high quality samples from the training dataset, but I won't be using it for my final training (and neither should you)</p>",
  "messages": [
    {
      "id": "1032513",
      "postDate": "09/30/2020 08:58:37",
      "content": "<p>Hi all, </p>\n<p>I created a mask for the training set that some of you might find useful, you can find it here: <a href=\"https://www.kaggle.com/fnands/training-mask-100-200-10\" target=\"_blank\">link to dataset</a>. </p>\n<p>Additionally, if you want to change the parameters and make your own mask the notebook that created it can be found here: <a href=\"https://www.kaggle.com/fnands/makesubmasks\" target=\"_blank\">notebook to make masks</a></p>\n<h2>Rationale:</h2>\n<p>As the data has a sampling rate of 10Hz, the world doesn't change much from one frame to the other. This means that there are many near duplicates in the normal dataset.     <br>\nSeeing as most of us are still prototyping our methods, this means we are spending a lot of time running over data that is very similar, and with such a large dataset as we have here that can severely slow down our testing.  </p>\n<p>This mask picks out a subset of the training data that looks similar to our test set and that is lower in near-duplicates than the full training set.   </p>\n<p><strong>Let me be clear on this point</strong>: I think this is a useful tool for prototyping, as it will pick out a selection of high quality samples from the training dataset, but I won't be using it for my final training (and neither should you)</p>",
      "rawMarkdown": "Hi all, \n\nI created a mask for the training set that some of you might find useful, you can find it here: [link to dataset](https://www.kaggle.com/fnands/training-mask-100-200-10). \n\nAdditionally, if you want to change the parameters and make your own mask the notebook that created it can be found here: [notebook to make masks](https://www.kaggle.com/fnands/makesubmasks)\n\n## Rationale: \nAs the data has a sampling rate of 10Hz, the world doesn't change much from one frame to the other. This means that there are many near duplicates in the normal dataset.     \nSeeing as most of us are still prototyping our methods, this means we are spending a lot of time running over data that is very similar, and with such a large dataset as we have here that can severely slow down our testing.  \n\nThis mask picks out a subset of the training data that looks similar to our test set and that is lower in near-duplicates than the full training set.   \n\n**Let me be clear on this point**: I think this is a useful tool for prototyping, as it will pick out a selection of high quality samples from the training dataset, but I won't be using it for my final training (and neither should you)",
      "votes": null
    },
    {
      "id": "1032616",
      "postDate": "09/30/2020 10:14:58",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> . The training on full_zarr seem never ending. Will experiment with parameters in your book and tell you how it goes. </p>",
      "rawMarkdown": "thanks @fnands . The training on full_zarr seem never ending. Will experiment with parameters in your book and tell you how it goes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1032616,
      "author_name": "deepakrajpurushothaman",
      "author_url": "",
      "post_date": "09/30/2020 10:14:58",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> . The training on full_zarr seem never ending. Will experiment with parameters in your book and tell you how it goes. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1032513": "Hi all, \n\nI created a mask for the training set that some of you might find useful, you can find it here: [link to dataset](https://www.kaggle.com/fnands/training-mask-100-200-10). \n\nAdditionally, if you want to change the parameters and make your own mask the notebook that created it can be found here: [notebook to make masks](https://www.kaggle.com/fnands/makesubmasks)\n\n## Rationale: \nAs the data has a sampling rate of 10Hz, the world doesn't change much from one frame to the other. This means that there are many near duplicates in the normal dataset.     \nSeeing as most of us are still prototyping our methods, this means we are spending a lot of time running over data that is very similar, and with such a large dataset as we have here that can severely slow down our testing.  \n\nThis mask picks out a subset of the training data that looks similar to our test set and that is lower in near-duplicates than the full training set.   \n\n**Let me be clear on this point**: I think this is a useful tool for prototyping, as it will pick out a selection of high quality samples from the training dataset, but I won't be using it for my final training (and neither should you)",
    "1032616": "thanks @fnands . The training on full_zarr seem never ending. Will experiment with parameters in your book and tell you how it goes."
  },
  "source": "meta"
}