{
  "id": 175418,
  "title": "Pseudo labelling: What worked and what didn't?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175418",
  "author_name": "",
  "post_date": "2020-08-18T06:09:34.178928800Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Pseudo labelling was potentially very useful (while being very risky) in this competition due to two reasons:</p>\n<ol>\n<li>Heavily skewed dataset (only ~2.3% positive samples)</li>\n<li>'Mystery' images not present in training but present in (public as well as private) test.</li>\n</ol>\n<p><strong>How was your experience while using pseudo labelled images? Did it work for you?</strong></p>\n<p>We were able to score up to 0.9445 pvt LB (0.9525 public LB) using a blend of 11 pseudo labelled single models each trained on a single fold (though this was not one of our final submissions). Details of each model are as below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F763007%2Ff6a4cde9cea78bc74594b2d253951b47%2FScreenshot%202020-08-18%20at%2010.34.27%20AM.png?generation=1597728232400582&amp;alt=media\" alt=\"\"></p>\n<p><strong>Links to Pseudo-labelled dataset created by us:</strong><br>\n<a href=\"https://www.kaggle.com/rohitagarwal/pseudo\" target=\"_blank\">Dataset1</a><br>\n<a href=\"https://www.kaggle.com/rohitagarwal/pseudoall\" target=\"_blank\">Dataset2</a><br>\n<a href=\"https://www.kaggle.com/rohitagarwal/pseudolabeling\" target=\"_blank\">Dataset3</a></p>\n<h2>How did we perform the Pseudo-labelling?</h2>\n<ul>\n<li>We considered all submissions that scored high on public LB. This includes private as well as public submissions.</li>\n<li>We identified images that are consistently ranked high across all these submissions and marked those as positive.</li>\n<li>We identified images that are consistently ranked low across all these submissions and marked those as negative.</li>\n<li>We maintained the ratio of positive / negative to be around 2.2%-2.4% in line with the distribution of the rest of the data.</li>\n<li>We created a TFRecord for each of these set of pseudo-labelled images.</li>\n<li>We include this fold for training (and never for validation, where's the fun otherwise?).</li>\n</ul>\n<h2>What did we experiment?</h2>\n<ul>\n<li>Number of submissions to consider for this exercise</li>\n<li>Thresholds for considering positive images</li>\n<li>Whether we should consider all images or only the 'mystery' images (1920x1080).</li>\n</ul>\n<h2>What worked for us?</h2>\n<ul>\n<li>We considered ~20-30 submissions (all that scored more than 0.95xx on the public LB). The lower the overlap of the images in the top ranks across all the submissions, the better for us because it was more likely that those were actually positive and in the public LB</li>\n<li>We considered multiple thresholds: top 260 images, top 130 images, top 65 images, etc. - no clear winner here, however generally the tighter the threshold the better it is</li>\n<li>We finally considered only the 'mystery' images pseudo-labelled. These images were otherwise not present in the training data and had a distribution which was significantly different from all the other images.</li>\n</ul>\n<p>Now that the private LB is released, I will soon probe it to identify how much of our labelling was actually accurate.</p>",
  "messages": [
    {
      "id": "975062",
      "postDate": "08/18/2020 06:09:34",
      "content": "<p>Pseudo labelling was potentially very useful (while being very risky) in this competition due to two reasons:</p>\n<ol>\n<li>Heavily skewed dataset (only ~2.3% positive samples)</li>\n<li>'Mystery' images not present in training but present in (public as well as private) test.</li>\n</ol>\n<p><strong>How was your experience while using pseudo labelled images? Did it work for you?</strong></p>\n<p>We were able to score up to 0.9445 pvt LB (0.9525 public LB) using a blend of 11 pseudo labelled single models each trained on a single fold (though this was not one of our final submissions). Details of each model are as below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F763007%2Ff6a4cde9cea78bc74594b2d253951b47%2FScreenshot%202020-08-18%20at%2010.34.27%20AM.png?generation=1597728232400582&amp;alt=media\" alt=\"\"></p>\n<p><strong>Links to Pseudo-labelled dataset created by us:</strong><br>\n<a href=\"https://www.kaggle.com/rohitagarwal/pseudo\" target=\"_blank\">Dataset1</a><br>\n<a href=\"https://www.kaggle.com/rohitagarwal/pseudoall\" target=\"_blank\">Dataset2</a><br>\n<a href=\"https://www.kaggle.com/rohitagarwal/pseudolabeling\" target=\"_blank\">Dataset3</a></p>\n<h2>How did we perform the Pseudo-labelling?</h2>\n<ul>\n<li>We considered all submissions that scored high on public LB. This includes private as well as public submissions.</li>\n<li>We identified images that are consistently ranked high across all these submissions and marked those as positive.</li>\n<li>We identified images that are consistently ranked low across all these submissions and marked those as negative.</li>\n<li>We maintained the ratio of positive / negative to be around 2.2%-2.4% in line with the distribution of the rest of the data.</li>\n<li>We created a TFRecord for each of these set of pseudo-labelled images.</li>\n<li>We include this fold for training (and never for validation, where's the fun otherwise?).</li>\n</ul>\n<h2>What did we experiment?</h2>\n<ul>\n<li>Number of submissions to consider for this exercise</li>\n<li>Thresholds for considering positive images</li>\n<li>Whether we should consider all images or only the 'mystery' images (1920x1080).</li>\n</ul>\n<h2>What worked for us?</h2>\n<ul>\n<li>We considered ~20-30 submissions (all that scored more than 0.95xx on the public LB). The lower the overlap of the images in the top ranks across all the submissions, the better for us because it was more likely that those were actually positive and in the public LB</li>\n<li>We considered multiple thresholds: top 260 images, top 130 images, top 65 images, etc. - no clear winner here, however generally the tighter the threshold the better it is</li>\n<li>We finally considered only the 'mystery' images pseudo-labelled. These images were otherwise not present in the training data and had a distribution which was significantly different from all the other images.</li>\n</ul>\n<p>Now that the private LB is released, I will soon probe it to identify how much of our labelling was actually accurate.</p>",
      "rawMarkdown": "Pseudo labelling was potentially very useful (while being very risky) in this competition due to two reasons:\n\n1. Heavily skewed dataset (only ~2.3% positive samples)\n2. 'Mystery' images not present in training but present in (public as well as private) test.\n\n**How was your experience while using pseudo labelled images? Did it work for you?**\n\nWe were able to score up to 0.9445 pvt LB (0.9525 public LB) using a blend of 11 pseudo labelled single models each trained on a single fold (though this was not one of our final submissions). Details of each model are as below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F763007%2Ff6a4cde9cea78bc74594b2d253951b47%2FScreenshot%202020-08-18%20at%2010.34.27%20AM.png?generation=1597728232400582&alt=media)\n\n**Links to Pseudo-labelled dataset created by us:**\n[Dataset1](https://www.kaggle.com/rohitagarwal/pseudo)\n[Dataset2](https://www.kaggle.com/rohitagarwal/pseudoall)\n[Dataset3](https://www.kaggle.com/rohitagarwal/pseudolabeling)\n\n## How did we perform the Pseudo-labelling?\n\n- We considered all submissions that scored high on public LB. This includes private as well as public submissions.\n- We identified images that are consistently ranked high across all these submissions and marked those as positive.\n- We identified images that are consistently ranked low across all these submissions and marked those as negative.\n- We maintained the ratio of positive / negative to be around 2.2%-2.4% in line with the distribution of the rest of the data.\n- We created a TFRecord for each of these set of pseudo-labelled images.\n- We include this fold for training (and never for validation, where's the fun otherwise?).\n\n## What did we experiment?\n\n- Number of submissions to consider for this exercise\n- Thresholds for considering positive images\n- Whether we should consider all images or only the 'mystery' images (1920x1080).\n\n## What worked for us?\n- We considered ~20-30 submissions (all that scored more than 0.95xx on the public LB). The lower the overlap of the images in the top ranks across all the submissions, the better for us because it was more likely that those were actually positive and in the public LB\n- We considered multiple thresholds: top 260 images, top 130 images, top 65 images, etc. - no clear winner here, however generally the tighter the threshold the better it is\n- We finally considered only the 'mystery' images pseudo-labelled. These images were otherwise not present in the training data and had a distribution which was significantly different from all the other images.\n\nNow that the private LB is released, I will soon probe it to identify how much of our labelling was actually accurate.",
      "votes": null
    },
    {
      "id": "975276",
      "postDate": "08/18/2020 08:18:28",
      "content": "<p>Nice - I was wondering if someone tried PSeudo Labelling. Good that it apparently worked for you.<br>\n I didn't have time to try certain things and left this out. </p>",
      "rawMarkdown": "Nice - I was wondering if someone tried PSeudo Labelling. Good that it apparently worked for you.\n I didn't have time to try certain things and left this out.",
      "votes": null
    },
    {
      "id": "975351",
      "postDate": "08/18/2020 08:57:19",
      "content": "<p>Hi, \nplease which method of Deep Learning fits with the dynamic environment big data size, as like massive data connectivity in wireless communication, and we have milimeterwave beamforming massive multi-input-multi-output devices in a dynamic environment.\nWhich one of these methods will be fit to solve high computational complexity, to get maximum policy value, perfect beamforming vector, or perfect decision-maker?\nCan use GAN, DDPG, Policy Search method?\nwhich of DL methods GAN, DDPG, Policy Search method fit with the Dynamic environment?\nThank you in advanced</p>\n\n<p>Alaaedi</p>",
      "rawMarkdown": "Hi, \nplease which method of Deep Learning fits with the dynamic environment big data size, as like massive data connectivity in wireless communication, and we have milimeterwave beamforming massive multi-input-multi-output devices in a dynamic environment.\nWhich one of these methods will be fit to solve high computational complexity, to get maximum policy value, perfect beamforming vector, or perfect decision-maker?\nCan use GAN, DDPG, Policy Search method?\nwhich of DL methods GAN, DDPG, Policy Search method fit with the Dynamic environment?\nThank you in advanced\n\nAlaaedi",
      "votes": null
    },
    {
      "id": "978302",
      "postDate": "08/20/2020 05:07:18",
      "content": "<p>We also tried it but could not get good result. As we enter very late in this comp ,almost 20 days before it ends so could not do much experiment on pseudo labeling.<br>\nOur approach was to use only positive class images and a little conservative positive/negative ratio we took was of 1.5% to 2.2%. </p>",
      "rawMarkdown": "We also tried it but could not get good result. As we enter very late in this comp ,almost 20 days before it ends so could not do much experiment on pseudo labeling.\nOur approach was to use only positive class images and a little conservative positive/negative ratio we took was of 1.5% to 2.2%.",
      "votes": null
    },
    {
      "id": "978479",
      "postDate": "08/20/2020 07:53:08",
      "content": "<p>Nice experiments! Have you already checked on the private LB how accurate was the pseudo-labeling? I was always wondering about the lower accuracy bound, at which pseudo-labeling is able to bring improvements.</p>\n<p>We did some experiments with pseudo-labeling of the top-100 images predicted to be malignant by our best submission, which was a blend of multiple models. The results on performance gains were not very conclusive: it helped a bit with some models but mostly introduced too much noise.</p>",
      "rawMarkdown": "Nice experiments! Have you already checked on the private LB how accurate was the pseudo-labeling? I was always wondering about the lower accuracy bound, at which pseudo-labeling is able to bring improvements.\n\nWe did some experiments with pseudo-labeling of the top-100 images predicted to be malignant by our best submission, which was a blend of multiple models. The results on performance gains were not very conclusive: it helped a bit with some models but mostly introduced too much noise.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 975276,
      "author_name": "vishram6",
      "author_url": "",
      "post_date": "08/18/2020 08:18:28",
      "content": "<p>Nice - I was wondering if someone tried PSeudo Labelling. Good that it apparently worked for you.<br>\n I didn't have time to try certain things and left this out. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 978302,
      "author_name": "rajnishe",
      "author_url": "",
      "post_date": "08/20/2020 05:07:18",
      "content": "<p>We also tried it but could not get good result. As we enter very late in this comp ,almost 20 days before it ends so could not do much experiment on pseudo labeling.<br>\nOur approach was to use only positive class images and a little conservative positive/negative ratio we took was of 1.5% to 2.2%. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 978479,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "08/20/2020 07:53:08",
      "content": "<p>Nice experiments! Have you already checked on the private LB how accurate was the pseudo-labeling? I was always wondering about the lower accuracy bound, at which pseudo-labeling is able to bring improvements.</p>\n<p>We did some experiments with pseudo-labeling of the top-100 images predicted to be malignant by our best submission, which was a blend of multiple models. The results on performance gains were not very conclusive: it helped a bit with some models but mostly introduced too much noise.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975351,
      "author_name": "halayed",
      "author_url": "",
      "post_date": "08/18/2020 08:57:19",
      "content": "<p>Hi, \nplease which method of Deep Learning fits with the dynamic environment big data size, as like massive data connectivity in wireless communication, and we have milimeterwave beamforming massive multi-input-multi-output devices in a dynamic environment.\nWhich one of these methods will be fit to solve high computational complexity, to get maximum policy value, perfect beamforming vector, or perfect decision-maker?\nCan use GAN, DDPG, Policy Search method?\nwhich of DL methods GAN, DDPG, Policy Search method fit with the Dynamic environment?\nThank you in advanced</p>\n\n<p>Alaaedi</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "975062": "Pseudo labelling was potentially very useful (while being very risky) in this competition due to two reasons:\n\n1. Heavily skewed dataset (only ~2.3% positive samples)\n2. 'Mystery' images not present in training but present in (public as well as private) test.\n\n**How was your experience while using pseudo labelled images? Did it work for you?**\n\nWe were able to score up to 0.9445 pvt LB (0.9525 public LB) using a blend of 11 pseudo labelled single models each trained on a single fold (though this was not one of our final submissions). Details of each model are as below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F763007%2Ff6a4cde9cea78bc74594b2d253951b47%2FScreenshot%202020-08-18%20at%2010.34.27%20AM.png?generation=1597728232400582&alt=media)\n\n**Links to Pseudo-labelled dataset created by us:**\n[Dataset1](https://www.kaggle.com/rohitagarwal/pseudo)\n[Dataset2](https://www.kaggle.com/rohitagarwal/pseudoall)\n[Dataset3](https://www.kaggle.com/rohitagarwal/pseudolabeling)\n\n## How did we perform the Pseudo-labelling?\n\n- We considered all submissions that scored high on public LB. This includes private as well as public submissions.\n- We identified images that are consistently ranked high across all these submissions and marked those as positive.\n- We identified images that are consistently ranked low across all these submissions and marked those as negative.\n- We maintained the ratio of positive / negative to be around 2.2%-2.4% in line with the distribution of the rest of the data.\n- We created a TFRecord for each of these set of pseudo-labelled images.\n- We include this fold for training (and never for validation, where's the fun otherwise?).\n\n## What did we experiment?\n\n- Number of submissions to consider for this exercise\n- Thresholds for considering positive images\n- Whether we should consider all images or only the 'mystery' images (1920x1080).\n\n## What worked for us?\n- We considered ~20-30 submissions (all that scored more than 0.95xx on the public LB). The lower the overlap of the images in the top ranks across all the submissions, the better for us because it was more likely that those were actually positive and in the public LB\n- We considered multiple thresholds: top 260 images, top 130 images, top 65 images, etc. - no clear winner here, however generally the tighter the threshold the better it is\n- We finally considered only the 'mystery' images pseudo-labelled. These images were otherwise not present in the training data and had a distribution which was significantly different from all the other images.\n\nNow that the private LB is released, I will soon probe it to identify how much of our labelling was actually accurate.",
    "975276": "Nice - I was wondering if someone tried PSeudo Labelling. Good that it apparently worked for you.\n I didn't have time to try certain things and left this out.",
    "975351": "Hi, \nplease which method of Deep Learning fits with the dynamic environment big data size, as like massive data connectivity in wireless communication, and we have milimeterwave beamforming massive multi-input-multi-output devices in a dynamic environment.\nWhich one of these methods will be fit to solve high computational complexity, to get maximum policy value, perfect beamforming vector, or perfect decision-maker?\nCan use GAN, DDPG, Policy Search method?\nwhich of DL methods GAN, DDPG, Policy Search method fit with the Dynamic environment?\nThank you in advanced\n\nAlaaedi",
    "978302": "We also tried it but could not get good result. As we enter very late in this comp ,almost 20 days before it ends so could not do much experiment on pseudo labeling.\nOur approach was to use only positive class images and a little conservative positive/negative ratio we took was of 1.5% to 2.2%.",
    "978479": "Nice experiments! Have you already checked on the private LB how accurate was the pseudo-labeling? I was always wondering about the lower accuracy bound, at which pseudo-labeling is able to bring improvements.\n\nWe did some experiments with pseudo-labeling of the top-100 images predicted to be malignant by our best submission, which was a blend of multiple models. The results on performance gains were not very conclusive: it helped a bit with some models but mostly introduced too much noise."
  },
  "source": "meta"
}