{
  "id": 26600,
  "title": "Small amount of training data?",
  "url": "/competitions/dstl-satellite-imagery-feature-detection/discussion/26600",
  "author_name": "",
  "post_date": "2016-12-17T01:29:49.417Z",
  "votes": 2,
  "comment_count": 6,
  "views": 521,
  "content": "<p>Looking at the data, there are more than 20Gb of images. Grid_sizes has 450 unique imageIDs, however train_wkt has only 22. Why do we have all of these images in the dataset if we only have labelled training data for less than 5%?</p>\n\n<p>Am I missing something here?</p>",
  "messages": [
    {
      "id": "150844",
      "postDate": "12/17/2016 01:29:49",
      "content": "<p>Looking at the data, there are more than 20Gb of images. Grid_sizes has 450 unique imageIDs, however train_wkt has only 22. Why do we have all of these images in the dataset if we only have labelled training data for less than 5%?</p>\n\n<p>Am I missing something here?</p>",
      "rawMarkdown": "Looking at the data, there are more than 20Gb of images. Grid_sizes has 450 unique imageIDs, however train_wkt has only 22. Why do we have all of these images in the dataset if we only have labelled training data for less than 5%?\r\n\r\nAm I missing something here?",
      "votes": null
    },
    {
      "id": "150860",
      "postDate": "12/17/2016 03:58:53",
      "content": "<p>The remaining images are the test set:</p>\n\n<ul>\n<li>sample_submission.csv has the 429 unique ImageIDs.</li>\n<li>train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).</li>\n<li>grid_sizes.csv has 450 ImageIDs.</li>\n<li>three_band/ has total 450 images.</li>\n<li>sixteen_band/ has 1350 images (3 images for each of the 450 ImageIDs = 1350 ).</li>\n</ul>\n\n<p>Everything seems to sum up correctly.</p>",
      "rawMarkdown": "The remaining images are the test set:\r\n\r\n - sample_submission.csv has the 429 unique ImageIDs.\r\n - train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).\r\n - grid_sizes.csv has 450 ImageIDs.\r\n - three_band/ has total 450 images.\r\n - sixteen_band/ has 1350 images (3 images for each of the 450 ImageIDs = 1350 ).\r\n\r\nEverything seems to sum up correctly.",
      "votes": null
    },
    {
      "id": "150862",
      "postDate": "12/17/2016 04:04:35",
      "content": "<p>No, the updated v2 train_wkt.csv has 22 imageIDs, check the data page again for the new file.</p>\n\n<p>The only thing that sums up is that 3*450 = 1350. It doesn't change the fact that there are only labels for 22 images, the remaining 428 are unlabelled images which makes them a part of the test set. It just seems odd to me that the test set is more than 20 times larger than the training set.</p>\n\n<p>I guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?</p>",
      "rawMarkdown": "No, the updated v2 train_wkt.csv has 22 imageIDs, check the data page again for the new file.\r\n\r\nThe only thing that sums up is that 3*450 = 1350. It doesn't change the fact that there are only labels for 22 images, the remaining 428 are unlabelled images which makes them a part of the test set. It just seems odd to me that the test set is more than 20 times larger than the training set.\r\n\r\nI guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?",
      "votes": null
    },
    {
      "id": "150863",
      "postDate": "12/17/2016 04:07:21",
      "content": "<blockquote>\n  <p>train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).</p>\n</blockquote>\n\n<p>You were correct.  It was announced in another thread that another image was added to the train set -- the train_wkt_v2.csv does indeed have 22 ImageIDs. Sorry for the confusion.</p>",
      "rawMarkdown": "> train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).\r\n\r\nYou were correct.  It was announced in another thread that another image was added to the train set -- the train_wkt_v2.csv does indeed have 22 ImageIDs. Sorry for the confusion.",
      "votes": null
    },
    {
      "id": "150954",
      "postDate": "12/17/2016 20:27:27",
      "content": "<p>[quote=grantbey;150862]</p>\n\n<p>I guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?</p>\n\n<p>[/quote]</p>\n\n<p>@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it <em>everywhere</em> to save the most human resources.</p>",
      "rawMarkdown": "[quote=grantbey;150862]\r\n\r\nI guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?\r\n\r\n[/quote]\r\n\r\n@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it *everywhere* to save the most human resources.",
      "votes": null
    },
    {
      "id": "150956",
      "postDate": "12/17/2016 20:46:00",
      "content": "<p>[quote=zero zero;150954]</p>\n\n<p>@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it <em>everywhere</em> to save the most human resources.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks, I guess I just thought it was unusual given other ML applications I've played with. I was mostly confirming that what I was seeing was correct and that I wasn't missing some crucial point and thus ignoring tons of training data.</p>\n\n<p>Thanks for your response!</p>",
      "rawMarkdown": "[quote=zero zero;150954]\r\n\r\n@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it *everywhere* to save the most human resources.\r\n\r\n[/quote]\r\n\r\nThanks, I guess I just thought it was unusual given other ML applications I've played with. I was mostly confirming that what I was seeing was correct and that I wasn't missing some crucial point and thus ignoring tons of training data.\r\n\r\nThanks for your response!",
      "votes": null
    },
    {
      "id": "387758",
      "postDate": "09/15/2018 16:06:27",
      "content": "<p>Manually tagging the data is not more a hurdle. Check this out,</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=tYqnsp-OcLQ\">https://www.youtube.com/watch?v=tYqnsp-OcLQ</a></p>\n\n<p><a href=\"https://www.youtube.com/watch?v=xkSEnDIlvhI\">https://www.youtube.com/watch?v=xkSEnDIlvhI</a></p>\n\n<p>you can quickly create your own image and video segmentation data in no time!!</p>\n\n<p>Neurala’s Brain Builder Tool: <a href=\"https://hubs.ly/H0dkvRK0\">https://hubs.ly/H0dkvRK0</a></p>",
      "rawMarkdown": "Manually tagging the data is not more a hurdle. Check this out,\n\nhttps://www.youtube.com/watch?v=tYqnsp-OcLQ\n\nhttps://www.youtube.com/watch?v=xkSEnDIlvhI\n\nyou can quickly create your own image and video segmentation data in no time!!\n\nNeurala’s Brain Builder Tool: https://hubs.ly/H0dkvRK0",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 150860,
      "author_name": "amanbh",
      "author_url": "",
      "post_date": "12/17/2016 03:58:53",
      "content": "<p>The remaining images are the test set:</p>\n\n<ul>\n<li>sample_submission.csv has the 429 unique ImageIDs.</li>\n<li>train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).</li>\n<li>grid_sizes.csv has 450 ImageIDs.</li>\n<li>three_band/ has total 450 images.</li>\n<li>sixteen_band/ has 1350 images (3 images for each of the 450 ImageIDs = 1350 ).</li>\n</ul>\n\n<p>Everything seems to sum up correctly.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 150862,
      "author_name": "grantbey",
      "author_url": "",
      "post_date": "12/17/2016 04:04:35",
      "content": "<p>No, the updated v2 train_wkt.csv has 22 imageIDs, check the data page again for the new file.</p>\n\n<p>The only thing that sums up is that 3*450 = 1350. It doesn't change the fact that there are only labels for 22 images, the remaining 428 are unlabelled images which makes them a part of the test set. It just seems odd to me that the test set is more than 20 times larger than the training set.</p>\n\n<p>I guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?</p>",
      "votes": null,
      "replies": [
        {
          "id": 150954,
          "author_name": "zerozero",
          "author_url": "",
          "post_date": "12/17/2016 20:27:27",
          "content": "<p>[quote=grantbey;150862]</p>\n\n<p>I guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?</p>\n\n<p>[/quote]</p>\n\n<p>@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it <em>everywhere</em> to save the most human resources.</p>",
          "votes": null,
          "replies": [
            {
              "id": 150956,
              "author_name": "grantbey",
              "author_url": "",
              "post_date": "12/17/2016 20:46:00",
              "content": "<p>[quote=zero zero;150954]</p>\n\n<p>@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it <em>everywhere</em> to save the most human resources.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks, I guess I just thought it was unusual given other ML applications I've played with. I was mostly confirming that what I was seeing was correct and that I wasn't missing some crucial point and thus ignoring tons of training data.</p>\n\n<p>Thanks for your response!</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 387758,
          "author_name": "ajagetia",
          "author_url": "",
          "post_date": "09/15/2018 16:06:27",
          "content": "<p>Manually tagging the data is not more a hurdle. Check this out,</p>\n\n<p><a href=\"https://www.youtube.com/watch?v=tYqnsp-OcLQ\">https://www.youtube.com/watch?v=tYqnsp-OcLQ</a></p>\n\n<p><a href=\"https://www.youtube.com/watch?v=xkSEnDIlvhI\">https://www.youtube.com/watch?v=xkSEnDIlvhI</a></p>\n\n<p>you can quickly create your own image and video segmentation data in no time!!</p>\n\n<p>Neurala’s Brain Builder Tool: <a href=\"https://hubs.ly/H0dkvRK0\">https://hubs.ly/H0dkvRK0</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 150863,
      "author_name": "amanbh",
      "author_url": "",
      "post_date": "12/17/2016 04:07:21",
      "content": "<blockquote>\n  <p>train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).</p>\n</blockquote>\n\n<p>You were correct.  It was announced in another thread that another image was added to the train set -- the train_wkt_v2.csv does indeed have 22 ImageIDs. Sorry for the confusion.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "150844": "Looking at the data, there are more than 20Gb of images. Grid_sizes has 450 unique imageIDs, however train_wkt has only 22. Why do we have all of these images in the dataset if we only have labelled training data for less than 5%?\r\n\r\nAm I missing something here?",
    "150860": "The remaining images are the test set:\r\n\r\n - sample_submission.csv has the 429 unique ImageIDs.\r\n - train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).\r\n - grid_sizes.csv has 450 ImageIDs.\r\n - three_band/ has total 450 images.\r\n - sixteen_band/ has 1350 images (3 images for each of the 450 ImageIDs = 1350 ).\r\n\r\nEverything seems to sum up correctly.",
    "150862": "No, the updated v2 train_wkt.csv has 22 imageIDs, check the data page again for the new file.\r\n\r\nThe only thing that sums up is that 3*450 = 1350. It doesn't change the fact that there are only labels for 22 images, the remaining 428 are unlabelled images which makes them a part of the test set. It just seems odd to me that the test set is more than 20 times larger than the training set.\r\n\r\nI guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?",
    "150863": "> train_wkt.csv has 21 unique ImageIDs (you incorrectly said 22).\r\n\r\nYou were correct.  It was announced in another thread that another image was added to the train set -- the train_wkt_v2.csv does indeed have 22 ImageIDs. Sorry for the confusion.",
    "150954": "[quote=grantbey;150862]\r\n\r\nI guess my question was: does this seem strange to anyone else? Or am I missing something crucial here?\r\n\r\n[/quote]\r\n\r\n@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it *everywhere* to save the most human resources.",
    "150956": "[quote=zero zero;150954]\r\n\r\n@grantbey, not strange at all that there's a large test set.  It makes it harder for any hand labeling during the competition and final judgement.  It's also more real-worldish as in practice one would want to develop a good model and then apply it *everywhere* to save the most human resources.\r\n\r\n[/quote]\r\n\r\nThanks, I guess I just thought it was unusual given other ML applications I've played with. I was mostly confirming that what I was seeing was correct and that I wasn't missing some crucial point and thus ignoring tons of training data.\r\n\r\nThanks for your response!",
    "387758": "Manually tagging the data is not more a hurdle. Check this out,\n\nhttps://www.youtube.com/watch?v=tYqnsp-OcLQ\n\nhttps://www.youtube.com/watch?v=xkSEnDIlvhI\n\nyou can quickly create your own image and video segmentation data in no time!!\n\nNeurala’s Brain Builder Tool: https://hubs.ly/H0dkvRK0"
  },
  "source": "meta"
}