{
  "id": 242778,
  "title": "PseudoLabelling",
  "url": "/competitions/bms-molecular-translation/discussion/242778",
  "author_name": "",
  "post_date": "2021-05-30T18:11:44.434152Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Has anyone found pseudolabelling to help? Also, when pseudolabelling, should you retrain the model on the new labels + training set, or finetune the model on the extra data?</p>",
  "messages": [
    {
      "id": "1328984",
      "postDate": "05/30/2021 18:11:44",
      "content": "<p>Has anyone found pseudolabelling to help? Also, when pseudolabelling, should you retrain the model on the new labels + training set, or finetune the model on the extra data?</p>",
      "rawMarkdown": "Has anyone found pseudolabelling to help? Also, when pseudolabelling, should you retrain the model on the new labels + training set, or finetune the model on the extra data?",
      "votes": null
    },
    {
      "id": "1331533",
      "postDate": "06/01/2021 14:22:57",
      "content": "<p>Good question. I'm wondering that too. I'd normally go with post-training on both the training data and pseudo-labelled data as that decreases the effect of overfitting to your own targets.</p>",
      "rawMarkdown": "Good question. I'm wondering that too. I'd normally go with post-training on both the training data and pseudo-labelled data as that decreases the effect of overfitting to your own targets.",
      "votes": null
    },
    {
      "id": "1331535",
      "postDate": "06/01/2021 14:24:09",
      "content": "<p>It's curious to me that this isn't being talked about more. Either it's not useful, or really useful lol</p>",
      "rawMarkdown": "It's curious to me that this isn't being talked about more. Either it's not useful, or really useful lol",
      "votes": null
    },
    {
      "id": "1331541",
      "postDate": "06/01/2021 14:28:34",
      "content": "<p>Well at least when you have an ensemble you can take the test-predictions where all models give the same prediction ==&gt; Pseudo-labels which are very likely correct. While those are the easier ones, it should improve performance. <br>\nAlthough maybe not use it over the whole training. But about that I can only guess, I have no experience.<br>\n(Why is pseudo-labeling even allowed? It clearly biases against test set)</p>",
      "rawMarkdown": "Well at least when you have an ensemble you can take the test-predictions where all models give the same prediction ==> Pseudo-labels which are very likely correct. While those are the easier ones, it should improve performance. \nAlthough maybe not use it over the whole training. But about that I can only guess, I have no experience.\n(Why is pseudo-labeling even allowed? It clearly biases against test set)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1331533,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "06/01/2021 14:22:57",
      "content": "<p>Good question. I'm wondering that too. I'd normally go with post-training on both the training data and pseudo-labelled data as that decreases the effect of overfitting to your own targets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1331535,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "06/01/2021 14:24:09",
          "content": "<p>It's curious to me that this isn't being talked about more. Either it's not useful, or really useful lol</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1331541,
          "author_name": "cepheidq",
          "author_url": "",
          "post_date": "06/01/2021 14:28:34",
          "content": "<p>Well at least when you have an ensemble you can take the test-predictions where all models give the same prediction ==&gt; Pseudo-labels which are very likely correct. While those are the easier ones, it should improve performance. <br>\nAlthough maybe not use it over the whole training. But about that I can only guess, I have no experience.<br>\n(Why is pseudo-labeling even allowed? It clearly biases against test set)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1328984": "Has anyone found pseudolabelling to help? Also, when pseudolabelling, should you retrain the model on the new labels + training set, or finetune the model on the extra data?",
    "1331533": "Good question. I'm wondering that too. I'd normally go with post-training on both the training data and pseudo-labelled data as that decreases the effect of overfitting to your own targets.",
    "1331535": "It's curious to me that this isn't being talked about more. Either it's not useful, or really useful lol",
    "1331541": "Well at least when you have an ensemble you can take the test-predictions where all models give the same prediction ==> Pseudo-labels which are very likely correct. While those are the easier ones, it should improve performance. \nAlthough maybe not use it over the whole training. But about that I can only guess, I have no experience.\n(Why is pseudo-labeling even allowed? It clearly biases against test set)"
  },
  "source": "meta"
}