{
  "id": 22050,
  "title": "Resized images (test data), data augmentation and LB score",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22050",
  "author_name": "",
  "post_date": "2016-07-05T13:21:50.217Z",
  "votes": null,
  "comment_count": 2,
  "views": 566,
  "content": "<p>I am trying to figure out the meaning of this text: &quot;To discourage hand labeling, we have supplemented the test dataset with some images that are resized. These processed images are ignored and don't count towards your score&quot;. Why would rescaled images discourage hand labelling? What do they mean by &quot;your score&quot;? Is it the current LB score (30% of the test data) or the final score? If the current LB score is computed based on resized images, data augmentation (e.g. using zoom on training data) could improve the LB score, but this gain would not account for the final score, since these processed images will be ignored. In this sense, data augmentation could even prejudice the final score.</p>",
  "messages": [
    {
      "id": "126006",
      "postDate": "07/05/2016 13:21:50",
      "content": "<p>I am trying to figure out the meaning of this text: &quot;To discourage hand labeling, we have supplemented the test dataset with some images that are resized. These processed images are ignored and don't count towards your score&quot;. Why would rescaled images discourage hand labelling? What do they mean by &quot;your score&quot;? Is it the current LB score (30% of the test data) or the final score? If the current LB score is computed based on resized images, data augmentation (e.g. using zoom on training data) could improve the LB score, but this gain would not account for the final score, since these processed images will be ignored. In this sense, data augmentation could even prejudice the final score.</p>",
      "rawMarkdown": "I am trying to figure out the meaning of this text: \"To discourage hand labeling, we have supplemented the test dataset with some images that are resized. These processed images are ignored and don't count towards your score\". Why would rescaled images discourage hand labelling? What do they mean by \"your score\"? Is it the current LB score (30% of the test data) or the final score? If the current LB score is computed based on resized images, data augmentation (e.g. using zoom on training data) could improve the LB score, but this gain would not account for the final score, since these processed images will be ignored. In this sense, data augmentation could even prejudice the final score.",
      "votes": null
    },
    {
      "id": "126030",
      "postDate": "07/05/2016 17:40:32",
      "content": "<p>Let's say there were 1,000 test images. It might be reasonable to hand label all of these. And what I mean by hand label is predicting the class, not labeling features in the image. But if you create another 9,000 fake images that aren't scored, and mix them in with the test images, it would be much harder to finish hand labeling in the allotted time of the contest.</p>\n\n<p>These extra images aren't part of the score calculation (either public LB or private LB). They are filtered out before a score is calculated. You can literally put any class in those rows (if you knew which were the extras) and it wouldn't change your LB score.</p>",
      "rawMarkdown": "Let's say there were 1,000 test images. It might be reasonable to hand label all of these. And what I mean by hand label is predicting the class, not labeling features in the image. But if you create another 9,000 fake images that aren't scored, and mix them in with the test images, it would be much harder to finish hand labeling in the allotted time of the contest.\r\n\r\nThese extra images aren't part of the score calculation (either public LB or private LB). They are filtered out before a score is calculated. You can literally put any class in those rows (if you knew which were the extras) and it wouldn't change your LB score.",
      "votes": null
    },
    {
      "id": "126039",
      "postDate": "07/05/2016 18:56:15",
      "content": "<p>Thanks for this information. In this case, the gains on the public LB resulting from the use of data augmentation tend to reflect the gains on the private LB.</p>",
      "rawMarkdown": "Thanks for this information. In this case, the gains on the public LB resulting from the use of data augmentation tend to reflect the gains on the private LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 126030,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "07/05/2016 17:40:32",
      "content": "<p>Let's say there were 1,000 test images. It might be reasonable to hand label all of these. And what I mean by hand label is predicting the class, not labeling features in the image. But if you create another 9,000 fake images that aren't scored, and mix them in with the test images, it would be much harder to finish hand labeling in the allotted time of the contest.</p>\n\n<p>These extra images aren't part of the score calculation (either public LB or private LB). They are filtered out before a score is calculated. You can literally put any class in those rows (if you knew which were the extras) and it wouldn't change your LB score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126039,
      "author_name": "oswaldoludwig",
      "author_url": "",
      "post_date": "07/05/2016 18:56:15",
      "content": "<p>Thanks for this information. In this case, the gains on the public LB resulting from the use of data augmentation tend to reflect the gains on the private LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "126006": "I am trying to figure out the meaning of this text: \"To discourage hand labeling, we have supplemented the test dataset with some images that are resized. These processed images are ignored and don't count towards your score\". Why would rescaled images discourage hand labelling? What do they mean by \"your score\"? Is it the current LB score (30% of the test data) or the final score? If the current LB score is computed based on resized images, data augmentation (e.g. using zoom on training data) could improve the LB score, but this gain would not account for the final score, since these processed images will be ignored. In this sense, data augmentation could even prejudice the final score.",
    "126030": "Let's say there were 1,000 test images. It might be reasonable to hand label all of these. And what I mean by hand label is predicting the class, not labeling features in the image. But if you create another 9,000 fake images that aren't scored, and mix them in with the test images, it would be much harder to finish hand labeling in the allotted time of the contest.\r\n\r\nThese extra images aren't part of the score calculation (either public LB or private LB). They are filtered out before a score is calculated. You can literally put any class in those rows (if you knew which were the extras) and it wouldn't change your LB score.",
    "126039": "Thanks for this information. In this case, the gains on the public LB resulting from the use of data augmentation tend to reflect the gains on the private LB."
  },
  "source": "meta"
}