{
  "id": 35114,
  "title": "Why make things harder then needed?",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/35114",
  "author_name": "",
  "post_date": "2017-06-22T14:00:02.217212200Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is a bit of a rant. Two latest competitions where I participated (Fisheries and Intel's) were unnecessarily complicated by the absence of ways to identify subsequent images in the former and some sort of anonymized patient ID in the latter (to use additional images), making correct CV (kudos to raddar for sharing his strategy) complex and unreliable. It seems to me that the absence of some sort of patient ID is artificial and does not reflect the reality. Having multiple images from the same patient would have helped everyone's models and patient IDs would have helped to avoid leaks from additional images while not introducing unreasonable assumptions (as if hospitals don't have this information). It's a win-win for all :)</p>",
  "messages": [
    {
      "id": "194992",
      "postDate": "06/22/2017 14:00:02",
      "content": "<p>This is a bit of a rant. Two latest competitions where I participated (Fisheries and Intel's) were unnecessarily complicated by the absence of ways to identify subsequent images in the former and some sort of anonymized patient ID in the latter (to use additional images), making correct CV (kudos to raddar for sharing his strategy) complex and unreliable. It seems to me that the absence of some sort of patient ID is artificial and does not reflect the reality. Having multiple images from the same patient would have helped everyone's models and patient IDs would have helped to avoid leaks from additional images while not introducing unreasonable assumptions (as if hospitals don't have this information). It's a win-win for all :)</p>",
      "rawMarkdown": "This is a bit of a rant. Two latest competitions where I participated (Fisheries and Intel's) were unnecessarily complicated by the absence of ways to identify subsequent images in the former and some sort of anonymized patient ID in the latter (to use additional images), making correct CV (kudos to raddar for sharing his strategy) complex and unreliable. It seems to me that the absence of some sort of patient ID is artificial and does not reflect the reality. Having multiple images from the same patient would have helped everyone's models and patient IDs would have helped to avoid leaks from additional images while not introducing unreasonable assumptions (as if hospitals don't have this information). It's a win-win for all :)",
      "votes": null
    },
    {
      "id": "194997",
      "postDate": "06/22/2017 14:31:25",
      "content": "<p>I totally agree bad taste in my mouth!</p>",
      "rawMarkdown": "I totally agree bad taste in my mouth!",
      "votes": null
    },
    {
      "id": "195013",
      "postDate": "06/22/2017 15:44:40",
      "content": "<p>I agree too.</p>\n\n<p>Data should be better prepared and verified - changing labels in 2/3 of the competition isn't very good idea, especially when it's not the first problem with the data. ID's would be very helpful and providing them wouldn't hurt, unless the competition is more about solving unnecessary data imperfections than the problem it's designed for.</p>\n\n<p>LB mining should be addressed too, I understand that the labels will be released in the end and we can retrain but really, considering the dataset size, those additional 512 images can make quite a difference during model validation. Let's not create an unfair advantage for those willing to mine, this should either be forbidden or when somone mines them, labels should be posted obligatorily. </p>",
      "rawMarkdown": "I agree too.\n\n Data should be better prepared and verified - changing labels in 2/3 of the competition isn't very good idea, especially when it's not the first problem with the data. ID's would be very helpful and providing them wouldn't hurt, unless the competition is more about solving unnecessary data imperfections than the problem it's designed for.\n\nLB mining should be addressed too, I understand that the labels will be released in the end and we can retrain but really, considering the dataset size, those additional 512 images can make quite a difference during model validation. Let's not create an unfair advantage for those willing to mine, this should either be forbidden or when somone mines them, labels should be posted obligatorily.",
      "votes": null
    },
    {
      "id": "195015",
      "postDate": "06/22/2017 15:54:03",
      "content": "<p>Oh yeah! Forgot about LB mining. This is another issue, which I'm not sure how to solve. But it should definitely be banned. While this will not prevent those who really want to mine using some fake accounts it will at least make things more complicated for them.</p>",
      "rawMarkdown": "Oh yeah! Forgot about LB mining. This is another issue, which I'm not sure how to solve. But it should definitely be banned. While this will not prevent those who really want to mine using some fake accounts it will at least make things more complicated for them.",
      "votes": null
    },
    {
      "id": "195063",
      "postDate": "06/22/2017 17:39:14",
      "content": "<p>For the LB mining, what about a format similar to the Stage 2? Where the Stage 1 is simply for entry(labels already provided) and Stage 2 is just a new dataset being released. This would more or less be a blind competition since the Stage 1 LB would mean nothing, but then again the Stage 2 public LB meant nothing...</p>",
      "rawMarkdown": "For the LB mining, what about a format similar to the Stage 2? Where the Stage 1 is simply for entry(labels already provided) and Stage 2 is just a new dataset being released. This would more or less be a blind competition since the Stage 1 LB would mean nothing, but then again the Stage 2 public LB meant nothing...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 194997,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "06/22/2017 14:31:25",
      "content": "<p>I totally agree bad taste in my mouth!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195013,
      "author_name": "wrosinski",
      "author_url": "",
      "post_date": "06/22/2017 15:44:40",
      "content": "<p>I agree too.</p>\n\n<p>Data should be better prepared and verified - changing labels in 2/3 of the competition isn't very good idea, especially when it's not the first problem with the data. ID's would be very helpful and providing them wouldn't hurt, unless the competition is more about solving unnecessary data imperfections than the problem it's designed for.</p>\n\n<p>LB mining should be addressed too, I understand that the labels will be released in the end and we can retrain but really, considering the dataset size, those additional 512 images can make quite a difference during model validation. Let's not create an unfair advantage for those willing to mine, this should either be forbidden or when somone mines them, labels should be posted obligatorily. </p>",
      "votes": null,
      "replies": [
        {
          "id": 195015,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "06/22/2017 15:54:03",
          "content": "<p>Oh yeah! Forgot about LB mining. This is another issue, which I'm not sure how to solve. But it should definitely be banned. While this will not prevent those who really want to mine using some fake accounts it will at least make things more complicated for them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195063,
          "author_name": "timothyrimnac",
          "author_url": "",
          "post_date": "06/22/2017 17:39:14",
          "content": "<p>For the LB mining, what about a format similar to the Stage 2? Where the Stage 1 is simply for entry(labels already provided) and Stage 2 is just a new dataset being released. This would more or less be a blind competition since the Stage 1 LB would mean nothing, but then again the Stage 2 public LB meant nothing...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "194992": "This is a bit of a rant. Two latest competitions where I participated (Fisheries and Intel's) were unnecessarily complicated by the absence of ways to identify subsequent images in the former and some sort of anonymized patient ID in the latter (to use additional images), making correct CV (kudos to raddar for sharing his strategy) complex and unreliable. It seems to me that the absence of some sort of patient ID is artificial and does not reflect the reality. Having multiple images from the same patient would have helped everyone's models and patient IDs would have helped to avoid leaks from additional images while not introducing unreasonable assumptions (as if hospitals don't have this information). It's a win-win for all :)",
    "194997": "I totally agree bad taste in my mouth!",
    "195013": "I agree too.\n\n Data should be better prepared and verified - changing labels in 2/3 of the competition isn't very good idea, especially when it's not the first problem with the data. ID's would be very helpful and providing them wouldn't hurt, unless the competition is more about solving unnecessary data imperfections than the problem it's designed for.\n\nLB mining should be addressed too, I understand that the labels will be released in the end and we can retrain but really, considering the dataset size, those additional 512 images can make quite a difference during model validation. Let's not create an unfair advantage for those willing to mine, this should either be forbidden or when somone mines them, labels should be posted obligatorily.",
    "195015": "Oh yeah! Forgot about LB mining. This is another issue, which I'm not sure how to solve. But it should definitely be banned. While this will not prevent those who really want to mine using some fake accounts it will at least make things more complicated for them.",
    "195063": "For the LB mining, what about a format similar to the Stage 2? Where the Stage 1 is simply for entry(labels already provided) and Stage 2 is just a new dataset being released. This would more or less be a blind competition since the Stage 1 LB would mean nothing, but then again the Stage 2 public LB meant nothing..."
  },
  "source": "meta"
}