{
  "id": 175379,
  "title": "still interested to know : 0.99 on public LB, how is it done?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175379",
  "author_name": "",
  "post_date": "2020-08-18T03:50:46.798487300Z",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>i saw a couple of teams have 0.99 on public LB before the competition ends. I wonder how it is done?</p>\n<p>(even if it is overfitting, i love to know about the details)<br>\nThanks!</p>",
  "messages": [
    {
      "id": "974783",
      "postDate": "08/18/2020 03:50:46",
      "content": "<p>i saw a couple of teams have 0.99 on public LB before the competition ends. I wonder how it is done?</p>\n<p>(even if it is overfitting, i love to know about the details)<br>\nThanks!</p>",
      "rawMarkdown": "i saw a couple of teams have 0.99 on public LB before the competition ends. I wonder how it is done?\n\n(even if it is overfitting, i love to know about the details)\nThanks!",
      "votes": null
    },
    {
      "id": "974805",
      "postDate": "08/18/2020 04:00:22",
      "content": "<p>I think you are looking for <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111\" target=\"_blank\">this</a>.</p>",
      "rawMarkdown": "I think you are looking for [this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111).",
      "votes": null
    },
    {
      "id": "974846",
      "postDate": "08/18/2020 04:27:23",
      "content": "<p>there are some general discussion in the link. but i hope to see how details. e.g. actual submission csv, scores or number of probes, etc</p>",
      "rawMarkdown": "there are some general discussion in the link. but i hope to see how details. e.g. actual submission csv, scores or number of probes, etc",
      "votes": null
    },
    {
      "id": "974878",
      "postDate": "08/18/2020 04:39:04",
      "content": "<p>I not sure but I think you can download data from <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\" target=\"_blank\">ISIC archive </a> and overfit on it, I think it might contained some test samples. But I don't have any evidence it just a hypothesis.</p>",
      "rawMarkdown": "I not sure but I think you can download data from [ISIC archive ](https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery) and overfit on it, I think it might contained some test samples. But I don't have any evidence it just a hypothesis.",
      "votes": null
    },
    {
      "id": "975020",
      "postDate": "08/18/2020 05:47:58",
      "content": "<p>Same curious here! It's amazing. </p>",
      "rawMarkdown": "Same curious here! It's amazing.",
      "votes": null
    },
    {
      "id": "975064",
      "postDate": "08/18/2020 06:10:23",
      "content": "<p>It is not fresh as far as I know. Please see the original <a href=\"http://blog.mrtz.org/2015/03/09/competition.html\" target=\"_blank\">post </a> and  kaggle <a href=\"https://www.kaggle.com/olegtrott/the-perfect-score-script\" target=\"_blank\">kernel</a> with related <a href=\"https://www.kaggle.com/c/restaurant-revenue-prediction/discussion/13950\" target=\"_blank\">discussion</a>. The difference of applying in this comp could be probing top N records (for example top 3000) rather than probing all 10982. Meanwhile, I also find a recent <a href=\"https://github.com/google-research/google-research/tree/master/grouptesting\" target=\"_blank\">work </a> (<a href=\"https://arxiv.org/abs/2004.12508\" target=\"_blank\">paper</a>) done by Google using sequential Monte carlo samplers for COVID-19 group testing that has the potiental to further optimize the method.</p>",
      "rawMarkdown": "It is not fresh as far as I know. Please see the original [post ](http://blog.mrtz.org/2015/03/09/competition.html) and  kaggle [kernel](https://www.kaggle.com/olegtrott/the-perfect-score-script) with related [discussion](https://www.kaggle.com/c/restaurant-revenue-prediction/discussion/13950). The difference of applying in this comp could be probing top N records (for example top 3000) rather than probing all 10982. Meanwhile, I also find a recent [work ](https://github.com/google-research/google-research/tree/master/grouptesting) ([paper](https://arxiv.org/abs/2004.12508)) done by Google using sequential Monte carlo samplers for COVID-19 group testing that has the potiental to further optimize the method.",
      "votes": null
    },
    {
      "id": "975240",
      "postDate": "08/18/2020 08:01:52",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> do you plan to make a write up?<br>\n<a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> ?</p>\n<p>I would be interested in three things personally:</p>\n<ul>\n<li>what was your approach for optimal LB probing?</li>\n<li>why would you do that? especially <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> who did this very early in the competition. Having all this light being at the top while everyone knows you achieved an impossible score… all this in a medical research competition…</li>\n<li>what did you learn from the experiments and what did you do with labels you got?</li>\n</ul>",
      "rawMarkdown": "zaharch do you plan to make a write up?\n@sirishks ?\n\nI would be interested in three things personally:\n- what was your approach for optimal LB probing?\n- why would you do that? especially @sirishks who did this very early in the competition. Having all this light being at the top while everyone knows you achieved an impossible score... all this in a medical research competition...\n- what did you learn from the experiments and what did you do with labels you got?",
      "votes": null
    },
    {
      "id": "975433",
      "postDate": "08/18/2020 09:47:19",
      "content": "<p>I am also waiting for <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> writeup, especially the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111\" target=\"_blank\">Gurobi MIP part</a>.</p>\n<p>When I started working on this competition, I kept getting bad results and was frustrated because of 2 reasons:</p>\n<ol>\n<li>It was impossible to <strong>visually</strong> identify whether a lesion is a MM or not, especially the Amelanotic kind.</li>\n<li>Many images looked like MMs and I wanted to figure out the <strong>real count</strong> i.e. distribution of test data.</li>\n</ol>\n<p>So, I kept making more and more submissions. As others have pointed out, probing Public LB is not very helpful and wastes precious submissions. I only did this experiment to check the accuracy of my model, whether the model is learning what is useful or not.</p>\n<p>In this competition I figured out how to climb <strong>Public</strong> LB, may be next one I will try to <strong>stay</strong> in a higher rank and avoid a shakeup 😪</p>\n<p>Regarding the actual technique, I have already posted the <strong>secret recipe</strong> in multiple posts - see:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497#901185\" target=\"_blank\">ignoring bottom rows</a>,</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#930723\" target=\"_blank\">estimating number of MMs in test data</a><ul>\n<li>(<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#932194\" target=\"_blank\">Optimo confirmed the calculations</a>),</li></ul></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437\" target=\"_blank\">checking duplicates</a><ul>\n<li>(<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956491\" target=\"_blank\">Chris confirmed the math calculations</a>),</li></ul></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">power averaging concepts</a></li>\n</ul>",
      "rawMarkdown": "I am also waiting for [@zaharch](https://www.kaggle.com/zaharch) writeup, especially the [Gurobi MIP part](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111).\n\nWhen I started working on this competition, I kept getting bad results and was frustrated because of 2 reasons:\n1. It was impossible to **visually** identify whether a lesion is a MM or not, especially the Amelanotic kind.\n2. Many images looked like MMs and I wanted to figure out the **real count** i.e. distribution of test data.\n\nSo, I kept making more and more submissions. As others have pointed out, probing Public LB is not very helpful and wastes precious submissions. I only did this experiment to check the accuracy of my model, whether the model is learning what is useful or not.\n\nIn this competition I figured out how to climb **Public** LB, may be next one I will try to **stay** in a higher rank and avoid a shakeup 😪\n\nRegarding the actual technique, I have already posted the **secret recipe** in multiple posts - see:\n- [ignoring bottom rows](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497#901185),\n- [estimating number of MMs in test data](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#930723)\n   - ([Optimo confirmed the calculations](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#932194)),\n- [checking duplicates](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437)\n   - ([Chris confirmed the math calculations](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956491)),\n- [power averaging concepts](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653)",
      "votes": null
    },
    {
      "id": "976007",
      "postDate": "08/18/2020 15:22:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> , <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> , I published my story in a different post, please take a look:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175559\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175559</a></p>",
      "rawMarkdown": "Hi @hengck23, @optimo , @sirishks , I published my story in a different post, please take a look:\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175559",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974805,
      "author_name": "wuliaokaola",
      "author_url": "",
      "post_date": "08/18/2020 04:00:22",
      "content": "<p>I think you are looking for <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111\" target=\"_blank\">this</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 974846,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/18/2020 04:27:23",
          "content": "<p>there are some general discussion in the link. but i hope to see how details. e.g. actual submission csv, scores or number of probes, etc</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974878,
      "author_name": "ivanvoid",
      "author_url": "",
      "post_date": "08/18/2020 04:39:04",
      "content": "<p>I not sure but I think you can download data from <a href=\"https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery\" target=\"_blank\">ISIC archive </a> and overfit on it, I think it might contained some test samples. But I don't have any evidence it just a hypothesis.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975020,
      "author_name": "yuanlin08",
      "author_url": "",
      "post_date": "08/18/2020 05:47:58",
      "content": "<p>Same curious here! It's amazing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975064,
      "author_name": "dxchen",
      "author_url": "",
      "post_date": "08/18/2020 06:10:23",
      "content": "<p>It is not fresh as far as I know. Please see the original <a href=\"http://blog.mrtz.org/2015/03/09/competition.html\" target=\"_blank\">post </a> and  kaggle <a href=\"https://www.kaggle.com/olegtrott/the-perfect-score-script\" target=\"_blank\">kernel</a> with related <a href=\"https://www.kaggle.com/c/restaurant-revenue-prediction/discussion/13950\" target=\"_blank\">discussion</a>. The difference of applying in this comp could be probing top N records (for example top 3000) rather than probing all 10982. Meanwhile, I also find a recent <a href=\"https://github.com/google-research/google-research/tree/master/grouptesting\" target=\"_blank\">work </a> (<a href=\"https://arxiv.org/abs/2004.12508\" target=\"_blank\">paper</a>) done by Google using sequential Monte carlo samplers for COVID-19 group testing that has the potiental to further optimize the method.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975240,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/18/2020 08:01:52",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> do you plan to make a write up?<br>\n<a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> ?</p>\n<p>I would be interested in three things personally:</p>\n<ul>\n<li>what was your approach for optimal LB probing?</li>\n<li>why would you do that? especially <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> who did this very early in the competition. Having all this light being at the top while everyone knows you achieved an impossible score… all this in a medical research competition…</li>\n<li>what did you learn from the experiments and what did you do with labels you got?</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975433,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "08/18/2020 09:47:19",
      "content": "<p>I am also waiting for <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> writeup, especially the <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111\" target=\"_blank\">Gurobi MIP part</a>.</p>\n<p>When I started working on this competition, I kept getting bad results and was frustrated because of 2 reasons:</p>\n<ol>\n<li>It was impossible to <strong>visually</strong> identify whether a lesion is a MM or not, especially the Amelanotic kind.</li>\n<li>Many images looked like MMs and I wanted to figure out the <strong>real count</strong> i.e. distribution of test data.</li>\n</ol>\n<p>So, I kept making more and more submissions. As others have pointed out, probing Public LB is not very helpful and wastes precious submissions. I only did this experiment to check the accuracy of my model, whether the model is learning what is useful or not.</p>\n<p>In this competition I figured out how to climb <strong>Public</strong> LB, may be next one I will try to <strong>stay</strong> in a higher rank and avoid a shakeup 😪</p>\n<p>Regarding the actual technique, I have already posted the <strong>secret recipe</strong> in multiple posts - see:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497#901185\" target=\"_blank\">ignoring bottom rows</a>,</li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#930723\" target=\"_blank\">estimating number of MMs in test data</a><ul>\n<li>(<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#932194\" target=\"_blank\">Optimo confirmed the calculations</a>),</li></ul></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437\" target=\"_blank\">checking duplicates</a><ul>\n<li>(<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956491\" target=\"_blank\">Chris confirmed the math calculations</a>),</li></ul></li>\n<li><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\" target=\"_blank\">power averaging concepts</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976007,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/18/2020 15:22:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> , <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> , I published my story in a different post, please take a look:</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175559\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175559</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974783": "i saw a couple of teams have 0.99 on public LB before the competition ends. I wonder how it is done?\n\n(even if it is overfitting, i love to know about the details)\nThanks!",
    "974805": "I think you are looking for [this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111).",
    "974846": "there are some general discussion in the link. but i hope to see how details. e.g. actual submission csv, scores or number of probes, etc",
    "974878": "I not sure but I think you can download data from [ISIC archive ](https://www.isic-archive.com/#!/topWithHeader/onlyHeaderTop/gallery) and overfit on it, I think it might contained some test samples. But I don't have any evidence it just a hypothesis.",
    "975020": "Same curious here! It's amazing.",
    "975064": "It is not fresh as far as I know. Please see the original [post ](http://blog.mrtz.org/2015/03/09/competition.html) and  kaggle [kernel](https://www.kaggle.com/olegtrott/the-perfect-score-script) with related [discussion](https://www.kaggle.com/c/restaurant-revenue-prediction/discussion/13950). The difference of applying in this comp could be probing top N records (for example top 3000) rather than probing all 10982. Meanwhile, I also find a recent [work ](https://github.com/google-research/google-research/tree/master/grouptesting) ([paper](https://arxiv.org/abs/2004.12508)) done by Google using sequential Monte carlo samplers for COVID-19 group testing that has the potiental to further optimize the method.",
    "975240": "zaharch do you plan to make a write up?\n@sirishks ?\n\nI would be interested in three things personally:\n- what was your approach for optimal LB probing?\n- why would you do that? especially @sirishks who did this very early in the competition. Having all this light being at the top while everyone knows you achieved an impossible score... all this in a medical research competition...\n- what did you learn from the experiments and what did you do with labels you got?",
    "975433": "I am also waiting for [@zaharch](https://www.kaggle.com/zaharch) writeup, especially the [Gurobi MIP part](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/174515#970111).\n\nWhen I started working on this competition, I kept getting bad results and was frustrated because of 2 reasons:\n1. It was impossible to **visually** identify whether a lesion is a MM or not, especially the Amelanotic kind.\n2. Many images looked like MMs and I wanted to figure out the **real count** i.e. distribution of test data.\n\nSo, I kept making more and more submissions. As others have pointed out, probing Public LB is not very helpful and wastes precious submissions. I only did this experiment to check the accuracy of my model, whether the model is learning what is useful or not.\n\nIn this competition I figured out how to climb **Public** LB, may be next one I will try to **stay** in a higher rank and avoid a shakeup 😪\n\nRegarding the actual technique, I have already posted the **secret recipe** in multiple posts - see:\n- [ignoring bottom rows](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497#901185),\n- [estimating number of MMs in test data](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#930723)\n   - ([Optimo confirmed the calculations](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167215#932194)),\n- [checking duplicates](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956437)\n   - ([Chris confirmed the math calculations](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/171930#956491)),\n- [power averaging concepts](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653)",
    "976007": "Hi @hengck23, @optimo , @sirishks , I published my story in a different post, please take a look:\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175559"
  },
  "source": "meta"
}