{
  "id": 37800,
  "title": "The real test set has the same size as the train set!",
  "url": "/competitions/carvana-image-masking-challenge/discussion/37800",
  "author_name": "Guillermo Barbadillo",
  "post_date": "2017-08-09T15:56:16.118000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have made 25 random submissions to estimate the real size of the test set. As it is said in the challenge description:</p>\n\n<blockquote>\n  <p>To deter hand labeling, we have supplemented the test set with car images that are ignored in scoring.</p>\n</blockquote>\n\n<p>So I think knowing the size of the real test set is important because the bigger the test set the most trust-able will be the public score.</p>\n\n<p>Combining the results of the random submissions with some simulations I have come to the conclusion that <strong>the test set has the same size as the train set</strong>. If you want to know more details about it check the link below. <br>\n<a href=\"https://www.kaggle.com/ironbar/real-test-set-size-estimation\">https://www.kaggle.com/ironbar/real-test-set-size-estimation</a></p>\n\n<p><img src=\"https://www.kaggle.io/svf/1416485/e129e2fe1090f91e45d14127dfa2e4b0/__results___files/__results___11_1.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>This means that for computing the public score 1696 images are being used, and that the final private score will be computed using 3392 images.</p>\n\n<p>I leave this open questions to the community:  </p>\n\n<ul>\n<li>How is the division between public and test set been made? Do images of the same car belong to both categories or just one?</li>\n<li>Could the big difference in maker distribution between train and test set be caused by unused cars?</li>\n</ul>",
  "messages": [
    {
      "id": 211633,
      "postDate": "2017-08-09T15:56:16.120Z",
      "content": "<p>I have made 25 random submissions to estimate the real size of the test set. As it is said in the challenge description:</p>\n\n<blockquote>\n  <p>To deter hand labeling, we have supplemented the test set with car images that are ignored in scoring.</p>\n</blockquote>\n\n<p>So I think knowing the size of the real test set is important because the bigger the test set the most trust-able will be the public score.</p>\n\n<p>Combining the results of the random submissions with some simulations I have come to the conclusion that <strong>the test set has the same size as the train set</strong>. If you want to know more details about it check the link below. <br>\n<a href=\"https://www.kaggle.com/ironbar/real-test-set-size-estimation\">https://www.kaggle.com/ironbar/real-test-set-size-estimation</a></p>\n\n<p><img src=\"https://www.kaggle.io/svf/1416485/e129e2fe1090f91e45d14127dfa2e4b0/__results___files/__results___11_1.png\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>This means that for computing the public score 1696 images are being used, and that the final private score will be computed using 3392 images.</p>\n\n<p>I leave this open questions to the community:  </p>\n\n<ul>\n<li>How is the division between public and test set been made? Do images of the same car belong to both categories or just one?</li>\n<li>Could the big difference in maker distribution between train and test set be caused by unused cars?</li>\n</ul>",
      "rawMarkdown": "I have made 25 random submissions to estimate the real size of the test set. As it is said in the challenge description:\n\n&gt; To deter hand labeling, we have supplemented the test set with car images that are ignored in scoring.\n\nSo I think knowing the size of the real test set is important because the bigger the test set the most trust-able will be the public score.\n\nCombining the results of the random submissions with some simulations I have come to the conclusion that **the test set has the same size as the train set**. If you want to know more details about it check the link below.  \nhttps://www.kaggle.com/ironbar/real-test-set-size-estimation\n\n![enter image description here][1]\n\n\nThis means that for computing the public score 1696 images are being used, and that the final private score will be computed using 3392 images.\n\nI leave this open questions to the community:  \n\n* How is the division between public and test set been made? Do images of the same car belong to both categories or just one?\n* Could the big difference in maker distribution between train and test set be caused by unused cars?\n\n\n  [1]: https://www.kaggle.io/svf/1416485/e129e2fe1090f91e45d14127dfa2e4b0/__results___files/__results___11_1.png",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "211633": "I have made 25 random submissions to estimate the real size of the test set. As it is said in the challenge description:\n\n&gt; To deter hand labeling, we have supplemented the test set with car images that are ignored in scoring.\n\nSo I think knowing the size of the real test set is important because the bigger the test set the most trust-able will be the public score.\n\nCombining the results of the random submissions with some simulations I have come to the conclusion that **the test set has the same size as the train set**. If you want to know more details about it check the link below.  \nhttps://www.kaggle.com/ironbar/real-test-set-size-estimation\n\n![enter image description here][1]\n\n\nThis means that for computing the public score 1696 images are being used, and that the final private score will be computed using 3392 images.\n\nI leave this open questions to the community:  \n\n* How is the division between public and test set been made? Do images of the same car belong to both categories or just one?\n* Could the big difference in maker distribution between train and test set be caused by unused cars?\n\n\n  [1]: https://www.kaggle.io/svf/1416485/e129e2fe1090f91e45d14127dfa2e4b0/__results___files/__results___11_1.png"
  }
}