{
  "id": 159847,
  "title": "why is nobody talking about pseudo labeling here?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/159847",
  "author_name": "",
  "post_date": "2020-06-18T23:35:01.355112700Z",
  "votes": 5,
  "comment_count": 12,
  "views": 0,
  "content": "<p>really wondering. any educated guesses ;) ?</p>",
  "messages": [
    {
      "id": "892493",
      "postDate": "06/18/2020 23:35:01",
      "content": "<p>really wondering. any educated guesses ;) ?</p>",
      "rawMarkdown": "really wondering. any educated guesses ;) ?",
      "votes": null
    },
    {
      "id": "892547",
      "postDate": "06/19/2020 02:04:30",
      "content": "<p>Saving it for later maybe?</p>",
      "rawMarkdown": "Saving it for later maybe?",
      "votes": null
    },
    {
      "id": "892569",
      "postDate": "06/19/2020 02:33:31",
      "content": "<p>because we dont have unlabelled data?</p>",
      "rawMarkdown": "because we dont have unlabelled data?",
      "votes": null
    },
    {
      "id": "892578",
      "postDate": "06/19/2020 02:48:37",
      "content": "<p>I think the idea of pseudo labeling is to augment the train set with a subset of the test data + predicted labels. The subset only includes instances on which our learning algorithm makes predictions with a high level of confidence. So, our unlabeled data come from the test set.</p>",
      "rawMarkdown": "I think the idea of pseudo labeling is to augment the train set with a subset of the test data + predicted labels. The subset only includes instances on which our learning algorithm makes predictions with a high level of confidence. So, our unlabeled data come from the test set.",
      "votes": null
    },
    {
      "id": "892588",
      "postDate": "06/19/2020 03:03:14",
      "content": "<p>It's too early. We need good enough models to generate pseudo labels.</p>",
      "rawMarkdown": "It's too early. We need good enough models to generate pseudo labels.",
      "votes": null
    },
    {
      "id": "892738",
      "postDate": "06/19/2020 05:58:54",
      "content": "<p>😂 </p>",
      "rawMarkdown": "😂",
      "votes": null
    },
    {
      "id": "892976",
      "postDate": "06/19/2020 09:44:01",
      "content": "<p><a href=\"/zzy990106\">@zzy990106</a>  <a href=\"/graf10a\">@graf10a</a> \n probably yes - but as far as I know (not much) its best to take only the very likely examples and add them to the training data - so there should not be much difference between the 0.94 models now and the 0.96 models later - hmm</p>",
      "rawMarkdown": "zzy990106  @graf10a \n probably yes - but as far as I know (not much) its best to take only the very likely examples and add them to the training data - so there should not be much difference between the 0.94 models now and the 0.96 models later - hmm",
      "votes": null
    },
    {
      "id": "893016",
      "postDate": "06/19/2020 10:21:50",
      "content": "<p>Since there is such a huge amount of external labeled data available, it seems risky to be using uncertain labels inferred from the test set IMO. Nevertheless, the images taken from the test set definitely are more similar to the test set than external images :)...</p>",
      "rawMarkdown": "Since there is such a huge amount of external labeled data available, it seems risky to be using uncertain labels inferred from the test set IMO. Nevertheless, the images taken from the test set definitely are more similar to the test set than external images :)...",
      "votes": null
    },
    {
      "id": "894266",
      "postDate": "06/20/2020 09:27:12",
      "content": "<p>I tried but not help. May be need to do for more models</p>",
      "rawMarkdown": "I tried but not help. May be need to do for more models",
      "votes": null
    },
    {
      "id": "896941",
      "postDate": "06/22/2020 14:32:08",
      "content": "<p>Usually it is a bit early to start using pseudo-labelling. You can use pseudo-labelling when we are sure your best-performing model is not overfitting while performing quite high on LB (my honest guess is around 0.95). Also if you want to do pseudo-labelling anyway, make sure to use very confident predictions usually for positive targets AUC &gt; 0.99 and for negative ones AUC &lt; 0.02, otherwise you will introduce a strong confirmation bias.</p>",
      "rawMarkdown": "Usually it is a bit early to start using pseudo-labelling. You can use pseudo-labelling when we are sure your best-performing model is not overfitting while performing quite high on LB (my honest guess is around 0.95). Also if you want to do pseudo-labelling anyway, make sure to use very confident predictions usually for positive targets AUC &gt; 0.99 and for negative ones AUC &lt; 0.02, otherwise you will introduce a strong confirmation bias.",
      "votes": null
    },
    {
      "id": "900125",
      "postDate": "06/24/2020 16:42:10",
      "content": "<p>Where are you planning to gather those unlabeled images from, Test data ? </p>",
      "rawMarkdown": "Where are you planning to gather those unlabeled images from, Test data ?",
      "votes": null
    },
    {
      "id": "907358",
      "postDate": "06/29/2020 22:17:12",
      "content": "<p>Yeah I tried that, there is not much improvement.</p>",
      "rawMarkdown": "Yeah I tried that, there is not much improvement.",
      "votes": null
    },
    {
      "id": "907395",
      "postDate": "06/29/2020 23:54:11",
      "content": "<p>can you specify the amount of improvement and can you tell the score of the model you used to predict the target values?</p>",
      "rawMarkdown": "can you specify the amount of improvement and can you tell the score of the model you used to predict the target values?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 892547,
      "author_name": "graf10a",
      "author_url": "",
      "post_date": "06/19/2020 02:04:30",
      "content": "<p>Saving it for later maybe?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 892569,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "06/19/2020 02:33:31",
      "content": "<p>because we dont have unlabelled data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 892578,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "06/19/2020 02:48:37",
          "content": "<p>I think the idea of pseudo labeling is to augment the train set with a subset of the test data + predicted labels. The subset only includes instances on which our learning algorithm makes predictions with a high level of confidence. So, our unlabeled data come from the test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 892738,
          "author_name": "changewow",
          "author_url": "",
          "post_date": "06/19/2020 05:58:54",
          "content": "<p>😂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 892588,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "06/19/2020 03:03:14",
      "content": "<p>It's too early. We need good enough models to generate pseudo labels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 892976,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "06/19/2020 09:44:01",
          "content": "<p><a href=\"/zzy990106\">@zzy990106</a>  <a href=\"/graf10a\">@graf10a</a> \n probably yes - but as far as I know (not much) its best to take only the very likely examples and add them to the training data - so there should not be much difference between the 0.94 models now and the 0.96 models later - hmm</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893016,
      "author_name": "group16",
      "author_url": "",
      "post_date": "06/19/2020 10:21:50",
      "content": "<p>Since there is such a huge amount of external labeled data available, it seems risky to be using uncertain labels inferred from the test set IMO. Nevertheless, the images taken from the test set definitely are more similar to the test set than external images :)...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 894266,
      "author_name": "dinhthilan",
      "author_url": "",
      "post_date": "06/20/2020 09:27:12",
      "content": "<p>I tried but not help. May be need to do for more models</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 896941,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/22/2020 14:32:08",
      "content": "<p>Usually it is a bit early to start using pseudo-labelling. You can use pseudo-labelling when we are sure your best-performing model is not overfitting while performing quite high on LB (my honest guess is around 0.95). Also if you want to do pseudo-labelling anyway, make sure to use very confident predictions usually for positive targets AUC &gt; 0.99 and for negative ones AUC &lt; 0.02, otherwise you will introduce a strong confirmation bias.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 900125,
      "author_name": "niteshx2",
      "author_url": "",
      "post_date": "06/24/2020 16:42:10",
      "content": "<p>Where are you planning to gather those unlabeled images from, Test data ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 907358,
      "author_name": "abhishek60",
      "author_url": "",
      "post_date": "06/29/2020 22:17:12",
      "content": "<p>Yeah I tried that, there is not much improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 907395,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "06/29/2020 23:54:11",
          "content": "<p>can you specify the amount of improvement and can you tell the score of the model you used to predict the target values?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "892493": "really wondering. any educated guesses ;) ?",
    "892547": "Saving it for later maybe?",
    "892569": "because we dont have unlabelled data?",
    "892578": "I think the idea of pseudo labeling is to augment the train set with a subset of the test data + predicted labels. The subset only includes instances on which our learning algorithm makes predictions with a high level of confidence. So, our unlabeled data come from the test set.",
    "892588": "It's too early. We need good enough models to generate pseudo labels.",
    "892738": "😂",
    "892976": "zzy990106  @graf10a \n probably yes - but as far as I know (not much) its best to take only the very likely examples and add them to the training data - so there should not be much difference between the 0.94 models now and the 0.96 models later - hmm",
    "893016": "Since there is such a huge amount of external labeled data available, it seems risky to be using uncertain labels inferred from the test set IMO. Nevertheless, the images taken from the test set definitely are more similar to the test set than external images :)...",
    "894266": "I tried but not help. May be need to do for more models",
    "896941": "Usually it is a bit early to start using pseudo-labelling. You can use pseudo-labelling when we are sure your best-performing model is not overfitting while performing quite high on LB (my honest guess is around 0.95). Also if you want to do pseudo-labelling anyway, make sure to use very confident predictions usually for positive targets AUC &gt; 0.99 and for negative ones AUC &lt; 0.02, otherwise you will introduce a strong confirmation bias.",
    "900125": "Where are you planning to gather those unlabeled images from, Test data ?",
    "907358": "Yeah I tried that, there is not much improvement.",
    "907395": "can you specify the amount of improvement and can you tell the score of the model you used to predict the target values?"
  },
  "source": "meta"
}