{
  "id": 175595,
  "title": "What was the point?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175595",
  "author_name": "kvigly",
  "post_date": "2020-08-18T18:11:05.483000",
  "votes": 31,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Indeed. What was the point of this competition? </p>\n<ol>\n<li>Why it was done in a .csv submission format, which allowed to just blend N top kernels and get a medal? </li>\n<li>Why was it just a binary segmentation? previous years ISIC had a (better) classification into 4 classes - isn't it more clinically relevant? </li>\n<li>Top solutions - can they really be converted into something that can be further used outside of the competition? </li>\n<li>Terrible CV\\LB correlation - again, suggests that top-LB solutions may not be the best, in fact.a</li>\n</ol>\n<p>I understand this competition is a great point to maybe get started in DL (thus so many people participating) - nice metrics, pictures, where you can understand, what is going on, decent amount of data - basically, everything you need to learning things about the DL, but was that the aim of the organizing team? :)  </p>",
  "messages": [
    {
      "id": 976214,
      "postDate": "2020-08-18T18:11:05.483Z",
      "content": "<p>Indeed. What was the point of this competition? </p>\n<ol>\n<li>Why it was done in a .csv submission format, which allowed to just blend N top kernels and get a medal? </li>\n<li>Why was it just a binary segmentation? previous years ISIC had a (better) classification into 4 classes - isn't it more clinically relevant? </li>\n<li>Top solutions - can they really be converted into something that can be further used outside of the competition? </li>\n<li>Terrible CV\\LB correlation - again, suggests that top-LB solutions may not be the best, in fact.a</li>\n</ol>\n<p>I understand this competition is a great point to maybe get started in DL (thus so many people participating) - nice metrics, pictures, where you can understand, what is going on, decent amount of data - basically, everything you need to learning things about the DL, but was that the aim of the organizing team? :)  </p>",
      "rawMarkdown": "Indeed. What was the point of this competition? \n1. Why it was done in a .csv submission format, which allowed to just blend N top kernels and get a medal? \n2. Why was it just a binary segmentation? previous years ISIC had a (better) classification into 4 classes - isn't it more clinically relevant? \n3. Top solutions - can they really be converted into something that can be further used outside of the competition? \n4. Terrible CV\\LB correlation - again, suggests that top-LB solutions may not be the best, in fact.a\n\nI understand this competition is a great point to maybe get started in DL (thus so many people participating) - nice metrics, pictures, where you can understand, what is going on, decent amount of data - basically, everything you need to learning things about the DL, but was that the aim of the organizing team? :)  \n",
      "votes": 31
    },
    {
      "id": 976273,
      "postDate": "2020-08-18T18:53:54.743Z",
      "content": "<p>What is the issue?</p>\n<p>When host allows arbitrary models it means the host is interested to see how far the limit can be pushed.</p>\n<p>Then you can evaluate how a single model compare to this limit.</p>\n<p>It is a very common misconception that solutions have to be used as is to be of value.</p>",
      "rawMarkdown": "What is the issue?\n\nWhen host allows arbitrary models it means the host is interested to see how far the limit can be pushed.\n\nThen you can evaluate how a single model compare to this limit.\n\nIt is a very common misconception that solutions have to be used as is to be of value.",
      "votes": 21,
      "replies": [
        {
          "id": 976319,
          "postDate": "2020-08-18T19:34:07.933Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> makes a very important point here.</p>",
          "rawMarkdown": "@cpmpml makes a very important point here.",
          "votes": 1
        }
      ]
    },
    {
      "id": 976281,
      "postDate": "2020-08-18T19:04:45.740Z",
      "content": "<p>You are correct in your questions. The real issue with this competition is a very small number of positive images. As a result, if we take the same model and run it on different parts of the test data we get very different AUC ROC(even without overfitting to the public LB). <br>\nThis means, there is a low correlation between the rank of the model in the private LB and how good it will perform on a much larger dataset.<br>\nThe low number of positive examples, also made the use of all of the patient's images for prediction, almost insignificant, which was one of the organizers' objective.    </p>",
      "rawMarkdown": "You are correct in your questions. The real issue with this competition is a very small number of positive images. As a result, if we take the same model and run it on different parts of the test data we get very different AUC ROC(even without overfitting to the public LB). \nThis means, there is a low correlation between the rank of the model in the private LB and how good it will perform on a much larger dataset.\nThe low number of positive examples, also made the use of all of the patient's images for prediction, almost insignificant, which was one of the organizers' objective.    ",
      "votes": 8
    },
    {
      "id": 976257,
      "postDate": "2020-08-18T18:43:22.763Z",
      "content": "<p>For #4 I think it's actually a good thing. From the participants' perspective, it teaches us how dangerous overfitting LB can be. From the host's perspective, it helps them pick the ones who are building robust models with good generalization ability on unseen data. If the CV-LB correlation was high (BTW some teams like the 1st place seemed to have found a reliable CV that correlates well with LB), it could be that LB-Private Board is also highly correlated, then it would be easy to do the modeling in an overfitting way but still get good results, which would make the resulting models less reliable/favorable in the real world. So I guess the key is to find a reliable CV that can generalize well on unseen data. </p>",
      "rawMarkdown": "For #4 I think it's actually a good thing. From the participants' perspective, it teaches us how dangerous overfitting LB can be. From the host's perspective, it helps them pick the ones who are building robust models with good generalization ability on unseen data. If the CV-LB correlation was high (BTW some teams like the 1st place seemed to have found a reliable CV that correlates well with LB), it could be that LB-Private Board is also highly correlated, then it would be easy to do the modeling in an overfitting way but still get good results, which would make the resulting models less reliable/favorable in the real world. So I guess the key is to find a reliable CV that can generalize well on unseen data. ",
      "votes": 5,
      "replies": [
        {
          "id": 976261,
          "postDate": "2020-08-18T18:46:22.997Z",
          "content": "<p>I dont agree here fully, I think the public LB should at least somehow be representative. We also dont know how well private LB correlates with CV, you only know it from some of the success stories. </p>",
          "rawMarkdown": "I dont agree here fully, I think the public LB should at least somehow be representative. We also dont know how well private LB correlates with CV, you only know it from some of the success stories. ",
          "votes": 6
        },
        {
          "id": 976283,
          "postDate": "2020-08-18T19:06:59.567Z",
          "content": "<p>I might be wrong, just my feeling (and the result of my experiments) is that LB is somewhat representative, although not a strict \"CV increase and LB increase with the same amount\" that kind of correlation, it's still correlated. Also, I think although this might be tricky for a competition, it seems to be more like a real-world messy setup where you don't know how your model performs until it goes live, and your best shot is to find something you can trust. Maybe CV, maybe LB, maybe a combo of both, as long as you have a strong enough logic to support it. IMO, if private LB does not correlate with our CV, it's our job to find a better CV set up. LOL</p>",
          "rawMarkdown": "I might be wrong, just my feeling (and the result of my experiments) is that LB is somewhat representative, although not a strict \"CV increase and LB increase with the same amount\" that kind of correlation, it's still correlated. Also, I think although this might be tricky for a competition, it seems to be more like a real-world messy setup where you don't know how your model performs until it goes live, and your best shot is to find something you can trust. Maybe CV, maybe LB, maybe a combo of both, as long as you have a strong enough logic to support it. IMO, if private LB does not correlate with our CV, it's our job to find a better CV set up. LOL",
          "votes": 2
        },
        {
          "id": 976298,
          "postDate": "2020-08-18T19:18:40.547Z",
          "content": "<p>The problem is when public LB correlates with CV, but private LB does not. And this can easily happen here as there is still a large random range in private LB scores.</p>",
          "rawMarkdown": "The problem is when public LB correlates with CV, but private LB does not. And this can easily happen here as there is still a large random range in private LB scores.",
          "votes": 8
        }
      ]
    },
    {
      "id": 978272,
      "postDate": "2020-08-20T04:26:11.697Z",
      "content": "<p>Should have been a two-stage code only competition </p>",
      "rawMarkdown": "Should have been a two-stage code only competition ",
      "votes": 1
    },
    {
      "id": 977727,
      "postDate": "2020-08-19T17:09:08.750Z",
      "content": "<p>Regarding your 4th point on CV/LB correlation. What I would suggest for Kaggle is to work more carefully with the organizers on the data partitioning in the future. Potential points to consider might be:</p>\n<ol>\n<li>Make sure that the public test set is at least <code>N</code> observations, otherwise use a <code>50/50</code> split between public and private test sets. Check that the class distribution is similar if there is a high imbalance. </li>\n<li>Run a couple of distribution consistency checks. In the case of image competitions, it could be image sizes, mean colors, etc. In the case of tabular data, it could be different descriptive statistics. Try to arrange the partitioning such that the differences between public/private/train are not too big.</li>\n<li>Train a couple of baseline models to check if the correlation between performance measures on different subsets looks decent.</li>\n</ol>\n<p>This would of course require additional investment from the platform/organizer side. And even with a perfect split, people are still going to overfit the public LB. But a good split would reduce the impact of pure luck, which is beneficial for all parties.</p>",
      "rawMarkdown": "Regarding your 4th point on CV/LB correlation. What I would suggest for Kaggle is to work more carefully with the organizers on the data partitioning in the future. Potential points to consider might be:\n1. Make sure that the public test set is at least `N` observations, otherwise use a `50/50` split between public and private test sets. Check that the class distribution is similar if there is a high imbalance. \n2. Run a couple of distribution consistency checks. In the case of image competitions, it could be image sizes, mean colors, etc. In the case of tabular data, it could be different descriptive statistics. Try to arrange the partitioning such that the differences between public/private/train are not too big.\n3. Train a couple of baseline models to check if the correlation between performance measures on different subsets looks decent.\n\nThis would of course require additional investment from the platform/organizer side. And even with a perfect split, people are still going to overfit the public LB. But a good split would reduce the impact of pure luck, which is beneficial for all parties.",
      "votes": 1
    },
    {
      "id": 977721,
      "postDate": "2020-08-19T17:02:18.450Z",
      "content": "<p>4) Looking at how imbalance the dataset was, I wonder if it is easy to make a CV/LB correlation. I suspect it is not easy (excepted if you have tons of data, the company probably doesn't have them)</p>",
      "rawMarkdown": "4) Looking at how imbalance the dataset was, I wonder if it is easy to make a CV/LB correlation. I suspect it is not easy (excepted if you have tons of data, the company probably doesn't have them)",
      "votes": 1
    },
    {
      "id": 976294,
      "postDate": "2020-08-18T19:17:00.600Z",
      "content": "<p>Just by hosting the competition you get people talking and raise awareness of the issues in this space. The value of the competition is beyond the produced models.</p>",
      "rawMarkdown": "Just by hosting the competition you get people talking and raise awareness of the issues in this space. The value of the competition is beyond the produced models.",
      "votes": 2
    },
    {
      "id": 976270,
      "postDate": "2020-08-18T18:51:14.303Z",
      "content": "<p>True! From what I have seen, most of the people(including me) resorted to transfer learning + metadata + extensive ensembling of their models. I am not sure if there is a novel solution offered, although the winner's training strategy is quite interesting and unique. But if anyone tried out of the box, I would love to hear new ideas even though they didn't top, because they deserve more attention and who knows, probably they would perform even better if they are provided with good resources (storage and multiple GPUs). 😃<br>\nAlso, the main aim of the competition is to use contextual information, and see if that could improve performance. I am curious to know if anyone succeeded in coming to a conclusion about whether contextual information is useful or not.</p>",
      "rawMarkdown": "True! From what I have seen, most of the people(including me) resorted to transfer learning + metadata + extensive ensembling of their models. I am not sure if there is a novel solution offered, although the winner's training strategy is quite interesting and unique. But if anyone tried out of the box, I would love to hear new ideas even though they didn't top, because they deserve more attention and who knows, probably they would perform even better if they are provided with good resources (storage and multiple GPUs). 😃\nAlso, the main aim of the competition is to use contextual information, and see if that could improve performance. I am curious to know if anyone succeeded in coming to a conclusion about whether contextual information is useful or not.",
      "votes": 2,
      "replies": [
        {
          "id": 979305,
          "postDate": "2020-08-20T18:50:28.190Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 976305,
      "postDate": "2020-08-18T19:22:54.517Z",
      "content": "<p>The first point is major to me, why was it CSV submission, for sure there will be people who copy submissions or even private share it. In my eyes this is definetly Kaggles fault, they should have made it a Kernel Submission only for more safety. The clean out now would not have been necessary.</p>",
      "rawMarkdown": "The first point is major to me, why was it CSV submission, for sure there will be people who copy submissions or even private share it. In my eyes this is definetly Kaggles fault, they should have made it a Kernel Submission only for more safety. The clean out now would not have been necessary."
    }
  ],
  "comments": [
    {
      "id": 976273,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-18T18:53:54.743000",
      "content": "<p>What is the issue?</p>\n<p>When host allows arbitrary models it means the host is interested to see how far the limit can be pushed.</p>\n<p>Then you can evaluate how a single model compare to this limit.</p>\n<p>It is a very common misconception that solutions have to be used as is to be of value.</p>",
      "votes": 21,
      "replies": [
        {
          "id": 976319,
          "author_name": "jsyphil",
          "author_url": "",
          "post_date": "2020-08-18T19:34:07.933000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> makes a very important point here.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 976281,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-08-18T19:04:45.740000",
      "content": "<p>You are correct in your questions. The real issue with this competition is a very small number of positive images. As a result, if we take the same model and run it on different parts of the test data we get very different AUC ROC(even without overfitting to the public LB). <br>\nThis means, there is a low correlation between the rank of the model in the private LB and how good it will perform on a much larger dataset.<br>\nThe low number of positive examples, also made the use of all of the patient's images for prediction, almost insignificant, which was one of the organizers' objective.    </p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 976257,
      "author_name": "DHZM",
      "author_url": "",
      "post_date": "2020-08-18T18:43:22.763000",
      "content": "<p>For #4 I think it's actually a good thing. From the participants' perspective, it teaches us how dangerous overfitting LB can be. From the host's perspective, it helps them pick the ones who are building robust models with good generalization ability on unseen data. If the CV-LB correlation was high (BTW some teams like the 1st place seemed to have found a reliable CV that correlates well with LB), it could be that LB-Private Board is also highly correlated, then it would be easy to do the modeling in an overfitting way but still get good results, which would make the resulting models less reliable/favorable in the real world. So I guess the key is to find a reliable CV that can generalize well on unseen data. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 976261,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-18T18:46:22.997000",
          "content": "<p>I dont agree here fully, I think the public LB should at least somehow be representative. We also dont know how well private LB correlates with CV, you only know it from some of the success stories. </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 976283,
          "author_name": "DHZM",
          "author_url": "",
          "post_date": "2020-08-18T19:06:59.567000",
          "content": "<p>I might be wrong, just my feeling (and the result of my experiments) is that LB is somewhat representative, although not a strict \"CV increase and LB increase with the same amount\" that kind of correlation, it's still correlated. Also, I think although this might be tricky for a competition, it seems to be more like a real-world messy setup where you don't know how your model performs until it goes live, and your best shot is to find something you can trust. Maybe CV, maybe LB, maybe a combo of both, as long as you have a strong enough logic to support it. IMO, if private LB does not correlate with our CV, it's our job to find a better CV set up. LOL</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 976298,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-18T19:18:40.547000",
          "content": "<p>The problem is when public LB correlates with CV, but private LB does not. And this can easily happen here as there is still a large random range in private LB scores.</p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 978272,
      "author_name": "Utkarsh Chandra Srivastava",
      "author_url": "",
      "post_date": "2020-08-20T04:26:11.697000",
      "content": "<p>Should have been a two-stage code only competition </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 977727,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2020-08-19T17:09:08.750000",
      "content": "<p>Regarding your 4th point on CV/LB correlation. What I would suggest for Kaggle is to work more carefully with the organizers on the data partitioning in the future. Potential points to consider might be:</p>\n<ol>\n<li>Make sure that the public test set is at least <code>N</code> observations, otherwise use a <code>50/50</code> split between public and private test sets. Check that the class distribution is similar if there is a high imbalance. </li>\n<li>Run a couple of distribution consistency checks. In the case of image competitions, it could be image sizes, mean colors, etc. In the case of tabular data, it could be different descriptive statistics. Try to arrange the partitioning such that the differences between public/private/train are not too big.</li>\n<li>Train a couple of baseline models to check if the correlation between performance measures on different subsets looks decent.</li>\n</ol>\n<p>This would of course require additional investment from the platform/organizer side. And even with a perfect split, people are still going to overfit the public LB. But a good split would reduce the impact of pure luck, which is beneficial for all parties.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 977721,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-08-19T17:02:18.450000",
      "content": "<p>4) Looking at how imbalance the dataset was, I wonder if it is easy to make a CV/LB correlation. I suspect it is not easy (excepted if you have tons of data, the company probably doesn't have them)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976294,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2020-08-18T19:17:00.600000",
      "content": "<p>Just by hosting the competition you get people talking and raise awareness of the issues in this space. The value of the competition is beyond the produced models.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 976270,
      "author_name": "AkundiPrathyusha",
      "author_url": "",
      "post_date": "2020-08-18T18:51:14.303000",
      "content": "<p>True! From what I have seen, most of the people(including me) resorted to transfer learning + metadata + extensive ensembling of their models. I am not sure if there is a novel solution offered, although the winner's training strategy is quite interesting and unique. But if anyone tried out of the box, I would love to hear new ideas even though they didn't top, because they deserve more attention and who knows, probably they would perform even better if they are provided with good resources (storage and multiple GPUs). 😃<br>\nAlso, the main aim of the competition is to use contextual information, and see if that could improve performance. I am curious to know if anyone succeeded in coming to a conclusion about whether contextual information is useful or not.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 979305,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-20T18:50:28.190000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 976305,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-08-18T19:22:54.517000",
      "content": "<p>The first point is major to me, why was it CSV submission, for sure there will be people who copy submissions or even private share it. In my eyes this is definetly Kaggles fault, they should have made it a Kernel Submission only for more safety. The clean out now would not have been necessary.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "976214": "Indeed. What was the point of this competition? \n1. Why it was done in a .csv submission format, which allowed to just blend N top kernels and get a medal? \n2. Why was it just a binary segmentation? previous years ISIC had a (better) classification into 4 classes - isn't it more clinically relevant? \n3. Top solutions - can they really be converted into something that can be further used outside of the competition? \n4. Terrible CV\\LB correlation - again, suggests that top-LB solutions may not be the best, in fact.a\n\nI understand this competition is a great point to maybe get started in DL (thus so many people participating) - nice metrics, pictures, where you can understand, what is going on, decent amount of data - basically, everything you need to learning things about the DL, but was that the aim of the organizing team? :)  \n",
    "976273": "What is the issue?\n\nWhen host allows arbitrary models it means the host is interested to see how far the limit can be pushed.\n\nThen you can evaluate how a single model compare to this limit.\n\nIt is a very common misconception that solutions have to be used as is to be of value.",
    "976281": "You are correct in your questions. The real issue with this competition is a very small number of positive images. As a result, if we take the same model and run it on different parts of the test data we get very different AUC ROC(even without overfitting to the public LB). \nThis means, there is a low correlation between the rank of the model in the private LB and how good it will perform on a much larger dataset.\nThe low number of positive examples, also made the use of all of the patient's images for prediction, almost insignificant, which was one of the organizers' objective.    ",
    "976257": "For #4 I think it's actually a good thing. From the participants' perspective, it teaches us how dangerous overfitting LB can be. From the host's perspective, it helps them pick the ones who are building robust models with good generalization ability on unseen data. If the CV-LB correlation was high (BTW some teams like the 1st place seemed to have found a reliable CV that correlates well with LB), it could be that LB-Private Board is also highly correlated, then it would be easy to do the modeling in an overfitting way but still get good results, which would make the resulting models less reliable/favorable in the real world. So I guess the key is to find a reliable CV that can generalize well on unseen data. ",
    "978272": "Should have been a two-stage code only competition ",
    "977727": "Regarding your 4th point on CV/LB correlation. What I would suggest for Kaggle is to work more carefully with the organizers on the data partitioning in the future. Potential points to consider might be:\n1. Make sure that the public test set is at least `N` observations, otherwise use a `50/50` split between public and private test sets. Check that the class distribution is similar if there is a high imbalance. \n2. Run a couple of distribution consistency checks. In the case of image competitions, it could be image sizes, mean colors, etc. In the case of tabular data, it could be different descriptive statistics. Try to arrange the partitioning such that the differences between public/private/train are not too big.\n3. Train a couple of baseline models to check if the correlation between performance measures on different subsets looks decent.\n\nThis would of course require additional investment from the platform/organizer side. And even with a perfect split, people are still going to overfit the public LB. But a good split would reduce the impact of pure luck, which is beneficial for all parties.",
    "977721": "4) Looking at how imbalance the dataset was, I wonder if it is easy to make a CV/LB correlation. I suspect it is not easy (excepted if you have tons of data, the company probably doesn't have them)",
    "976294": "Just by hosting the competition you get people talking and raise awareness of the issues in this space. The value of the competition is beyond the produced models.",
    "976270": "True! From what I have seen, most of the people(including me) resorted to transfer learning + metadata + extensive ensembling of their models. I am not sure if there is a novel solution offered, although the winner's training strategy is quite interesting and unique. But if anyone tried out of the box, I would love to hear new ideas even though they didn't top, because they deserve more attention and who knows, probably they would perform even better if they are provided with good resources (storage and multiple GPUs). 😃\nAlso, the main aim of the competition is to use contextual information, and see if that could improve performance. I am curious to know if anyone succeeded in coming to a conclusion about whether contextual information is useful or not.",
    "976305": "The first point is major to me, why was it CSV submission, for sure there will be people who copy submissions or even private share it. In my eyes this is definetly Kaggles fault, they should have made it a Kernel Submission only for more safety. The clean out now would not have been necessary."
  }
}