{
  "id": 34207,
  "title": "Image sizes information",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/34207",
  "author_name": "",
  "post_date": "2017-06-05T19:13:01.402373800Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Currently, the image sizes seems to contain valuable information about the Type. For instance, for Type_1 and Type_2 34% and 41% of pics are of size 3264 x 2448, however for the Type_3 it's a 66%. The difference seems too big to be a mere noise due to sampling. </p>\n\n<p>Any thoughts on this ? The annoying thing with this is that it can go both ways:</p>\n\n<ul>\n<li>if the test set contains the same kind of images, training the model that uses this information is valuable</li>\n<li>if this was due to the sampling method, it can kill your model</li>\n</ul>\n\n<p>In any case, this is not, or should not be, the goal of the competition, so it may be useful if the admins specify what's the distribution in the test set.</p>",
  "messages": [
    {
      "id": "189319",
      "postDate": "06/05/2017 19:13:01",
      "content": "<p>Currently, the image sizes seems to contain valuable information about the Type. For instance, for Type_1 and Type_2 34% and 41% of pics are of size 3264 x 2448, however for the Type_3 it's a 66%. The difference seems too big to be a mere noise due to sampling. </p>\n\n<p>Any thoughts on this ? The annoying thing with this is that it can go both ways:</p>\n\n<ul>\n<li>if the test set contains the same kind of images, training the model that uses this information is valuable</li>\n<li>if this was due to the sampling method, it can kill your model</li>\n</ul>\n\n<p>In any case, this is not, or should not be, the goal of the competition, so it may be useful if the admins specify what's the distribution in the test set.</p>",
      "rawMarkdown": "Currently, the image sizes seems to contain valuable information about the Type. For instance, for Type_1 and Type_2 34% and 41% of pics are of size 3264 x 2448, however for the Type_3 it's a 66%. The difference seems too big to be a mere noise due to sampling. \n\nAny thoughts on this ? The annoying thing with this is that it can go both ways:\n\n- if the test set contains the same kind of images, training the model that uses this information is valuable\n- if this was due to the sampling method, it can kill your model\n\nIn any case, this is not, or should not be, the goal of the competition, so it may be useful if the admins specify what's the distribution in the test set.",
      "votes": null
    },
    {
      "id": "189333",
      "postDate": "06/05/2017 20:10:06",
      "content": "<p>Admins could ensure this does not affect the competition outcome by resizing the images in the stage 2 test set. Better to remove the possibility that it impacts the results IMO.</p>",
      "rawMarkdown": "Admins could ensure this does not affect the competition outcome by resizing the images in the stage 2 test set. Better to remove the possibility that it impacts the results IMO.",
      "votes": null
    },
    {
      "id": "189710",
      "postDate": "06/06/2017 17:58:45",
      "content": "<p>Thanks for sharing this. Please do not rely on the image size as a feature as the second stage images will be resized. </p>",
      "rawMarkdown": "Thanks for sharing this. Please do not rely on the image size as a feature as the second stage images will be resized.",
      "votes": null
    },
    {
      "id": "190647",
      "postDate": "06/08/2017 05:51:14",
      "content": "<p>That's pure overfitting in its classical definition of describing noise in the dataset. Could be helpful to win Kaggle competitions but never in the real world.</p>",
      "rawMarkdown": "That's pure overfitting in its classical definition of describing noise in the dataset. Could be helpful to win Kaggle competitions but never in the real world.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 189333,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "06/05/2017 20:10:06",
      "content": "<p>Admins could ensure this does not affect the competition outcome by resizing the images in the stage 2 test set. Better to remove the possibility that it impacts the results IMO.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 189710,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "06/06/2017 17:58:45",
      "content": "<p>Thanks for sharing this. Please do not rely on the image size as a feature as the second stage images will be resized. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 190647,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "06/08/2017 05:51:14",
      "content": "<p>That's pure overfitting in its classical definition of describing noise in the dataset. Could be helpful to win Kaggle competitions but never in the real world.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "189319": "Currently, the image sizes seems to contain valuable information about the Type. For instance, for Type_1 and Type_2 34% and 41% of pics are of size 3264 x 2448, however for the Type_3 it's a 66%. The difference seems too big to be a mere noise due to sampling. \n\nAny thoughts on this ? The annoying thing with this is that it can go both ways:\n\n- if the test set contains the same kind of images, training the model that uses this information is valuable\n- if this was due to the sampling method, it can kill your model\n\nIn any case, this is not, or should not be, the goal of the competition, so it may be useful if the admins specify what's the distribution in the test set.",
    "189333": "Admins could ensure this does not affect the competition outcome by resizing the images in the stage 2 test set. Better to remove the possibility that it impacts the results IMO.",
    "189710": "Thanks for sharing this. Please do not rely on the image size as a feature as the second stage images will be resized.",
    "190647": "That's pure overfitting in its classical definition of describing noise in the dataset. Could be helpful to win Kaggle competitions but never in the real world."
  },
  "source": "meta"
}