{
  "id": 176085,
  "title": "SIIM 2020 vs SIIM 2019",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/176085",
  "author_name": "",
  "post_date": "2020-08-20T12:44:15.464023100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Last year, around the same season, there was a competition <br>\n\"SIIM-ACR Pneumothorax Segmentation\"<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation</a> </p>\n<p>At first it seems that these two competitions have a lot in common:  same organizer, medical image analysis problems, and at the end of both competitions there was a shake-up of a similar size</p>\n<p>Anybody who took part also in \"SIIM-ACR Pneumothorax Segmentation\", what do you think about both of them?</p>",
  "messages": [
    {
      "id": "978836",
      "postDate": "08/20/2020 12:44:15",
      "content": "<p>Last year, around the same season, there was a competition <br>\n\"SIIM-ACR Pneumothorax Segmentation\"<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation</a> </p>\n<p>At first it seems that these two competitions have a lot in common:  same organizer, medical image analysis problems, and at the end of both competitions there was a shake-up of a similar size</p>\n<p>Anybody who took part also in \"SIIM-ACR Pneumothorax Segmentation\", what do you think about both of them?</p>",
      "rawMarkdown": "Last year, around the same season, there was a competition \n\"SIIM-ACR Pneumothorax Segmentation\"\n[https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation](url) \n\nAt first it seems that these two competitions have a lot in common:  same organizer, medical image analysis problems, and at the end of both competitions there was a shake-up of a similar size\n\nAnybody who took part also in \"SIIM-ACR Pneumothorax Segmentation\", what do you think about both of them?",
      "votes": null
    },
    {
      "id": "979432",
      "postDate": "08/20/2020 20:47:36",
      "content": "<p>Large shakeups tend to happen because of the following two reasons:</p>\n<ol>\n<li>A lot of people are - mindlessly - stack, blend, blend of the blends for gains on the public LB, completely ignoring CV. </li>\n<li>In the medical and anomaly detection fields (in general) positive samples are very scarce. Either because of the cost of annotation or they don't have more samples at all, like rare disease etc. </li>\n</ol>\n<p>With a very small amount of positive samples, you can either:<br>\na, spread all of the different types of positives across train and test<br>\nb, leave a few types of cases to test only</p>\n<p>With option A, you will have a better model for the types of cases you have, but you will have no indication how well your model can generalize. You can only measure the general performance with option B, but it can cause lose CV-public-private correlations. </p>",
      "rawMarkdown": "Large shakeups tend to happen because of the following two reasons:\n1. A lot of people are - mindlessly - stack, blend, blend of the blends for gains on the public LB, completely ignoring CV. \n2. In the medical and anomaly detection fields (in general) positive samples are very scarce. Either because of the cost of annotation or they don't have more samples at all, like rare disease etc. \n\nWith a very small amount of positive samples, you can either:\na, spread all of the different types of positives across train and test\nb, leave a few types of cases to test only\n\nWith option A, you will have a better model for the types of cases you have, but you will have no indication how well your model can generalize. You can only measure the general performance with option B, but it can cause lose CV-public-private correlations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 979432,
      "author_name": "doncalculator",
      "author_url": "",
      "post_date": "08/20/2020 20:47:36",
      "content": "<p>Large shakeups tend to happen because of the following two reasons:</p>\n<ol>\n<li>A lot of people are - mindlessly - stack, blend, blend of the blends for gains on the public LB, completely ignoring CV. </li>\n<li>In the medical and anomaly detection fields (in general) positive samples are very scarce. Either because of the cost of annotation or they don't have more samples at all, like rare disease etc. </li>\n</ol>\n<p>With a very small amount of positive samples, you can either:<br>\na, spread all of the different types of positives across train and test<br>\nb, leave a few types of cases to test only</p>\n<p>With option A, you will have a better model for the types of cases you have, but you will have no indication how well your model can generalize. You can only measure the general performance with option B, but it can cause lose CV-public-private correlations. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "978836": "Last year, around the same season, there was a competition \n\"SIIM-ACR Pneumothorax Segmentation\"\n[https://www.kaggle.com/c/siim-acr-pneumothorax-segmentation](url) \n\nAt first it seems that these two competitions have a lot in common:  same organizer, medical image analysis problems, and at the end of both competitions there was a shake-up of a similar size\n\nAnybody who took part also in \"SIIM-ACR Pneumothorax Segmentation\", what do you think about both of them?",
    "979432": "Large shakeups tend to happen because of the following two reasons:\n1. A lot of people are - mindlessly - stack, blend, blend of the blends for gains on the public LB, completely ignoring CV. \n2. In the medical and anomaly detection fields (in general) positive samples are very scarce. Either because of the cost of annotation or they don't have more samples at all, like rare disease etc. \n\nWith a very small amount of positive samples, you can either:\na, spread all of the different types of positives across train and test\nb, leave a few types of cases to test only\n\nWith option A, you will have a better model for the types of cases you have, but you will have no indication how well your model can generalize. You can only measure the general performance with option B, but it can cause lose CV-public-private correlations."
  },
  "source": "meta"
}