{
  "id": 321303,
  "title": "About the pseudo label of this competition",
  "url": "/competitions/happy-whale-and-dolphin/discussion/321303",
  "author_name": "",
  "post_date": "2022-04-26T06:45:56.865813400Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It's been a week since the competition ended, but I'm still trying to figure out how to catch up with the top few scores. I tried a lot of pseudo-labels but the score barely improved, so I would like to ask you any good suggestions for the handling of pseudo-labels.</p>",
  "messages": [
    {
      "id": "1768312",
      "postDate": "04/26/2022 06:45:56",
      "content": "<p>It's been a week since the competition ended, but I'm still trying to figure out how to catch up with the top few scores. I tried a lot of pseudo-labels but the score barely improved, so I would like to ask you any good suggestions for the handling of pseudo-labels.</p>",
      "rawMarkdown": "It's been a week since the competition ended, but I'm still trying to figure out how to catch up with the top few scores. I tried a lot of pseudo-labels but the score barely improved, so I would like to ask you any good suggestions for the handling of pseudo-labels.",
      "votes": null
    },
    {
      "id": "1770036",
      "postDate": "04/27/2022 22:22:36",
      "content": "<p>Pseudo labeling is very powerful in this competition. Our team did not use pseudo label during the competition, but we just tried pseudo labels now after the competition and posted our results in an UPDATE section in our solution <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298\" target=\"_blank\">here</a></p>\n<p>The simpliest way to use pseudo is to take your <code>submission.csv</code> file and identify your model's 66% most confident predictions. If you use embeddings, then use best cosine similarity. If you use a voting classifier ensemble, then use predictions with most counts. Either way make a new dataset with those 66% best test prediction individual id.</p>\n<p>Next train your best models with train plus the new datasets created from test data with pseudo labels. That single model should perform similar to your ensemble LB score. Make a new ensemble with 50% new single model and 50% old ensemble. That should boost your LB.</p>\n<p>You can then repeat this process making pseudo labels from your new ensemble which should achieve a new best LB</p>",
      "rawMarkdown": "Pseudo labeling is very powerful in this competition. Our team did not use pseudo label during the competition, but we just tried pseudo labels now after the competition and posted our results in an UPDATE section in our solution [here][1]\n\nThe simpliest way to use pseudo is to take your `submission.csv` file and identify your model's 66% most confident predictions. If you use embeddings, then use best cosine similarity. If you use a voting classifier ensemble, then use predictions with most counts. Either way make a new dataset with those 66% best test prediction individual id.\n\nNext train your best models with train plus the new datasets created from test data with pseudo labels. That single model should perform similar to your ensemble LB score. Make a new ensemble with 50% new single model and 50% old ensemble. That should boost your LB.\n\nYou can then repeat this process making pseudo labels from your new ensemble which should achieve a new best LB\n\n[1]: https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298",
      "votes": null
    },
    {
      "id": "1770652",
      "postDate": "04/28/2022 13:23:00",
      "content": "<p>Thanks for your reply, your detailed process helped me a lot.  :=)</p>",
      "rawMarkdown": "Thanks for your reply, your detailed process helped me a lot.  :=)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1770036,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/27/2022 22:22:36",
      "content": "<p>Pseudo labeling is very powerful in this competition. Our team did not use pseudo label during the competition, but we just tried pseudo labels now after the competition and posted our results in an UPDATE section in our solution <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298\" target=\"_blank\">here</a></p>\n<p>The simpliest way to use pseudo is to take your <code>submission.csv</code> file and identify your model's 66% most confident predictions. If you use embeddings, then use best cosine similarity. If you use a voting classifier ensemble, then use predictions with most counts. Either way make a new dataset with those 66% best test prediction individual id.</p>\n<p>Next train your best models with train plus the new datasets created from test data with pseudo labels. That single model should perform similar to your ensemble LB score. Make a new ensemble with 50% new single model and 50% old ensemble. That should boost your LB.</p>\n<p>You can then repeat this process making pseudo labels from your new ensemble which should achieve a new best LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 1770652,
          "author_name": "yangranran",
          "author_url": "",
          "post_date": "04/28/2022 13:23:00",
          "content": "<p>Thanks for your reply, your detailed process helped me a lot.  :=)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1768312": "It's been a week since the competition ended, but I'm still trying to figure out how to catch up with the top few scores. I tried a lot of pseudo-labels but the score barely improved, so I would like to ask you any good suggestions for the handling of pseudo-labels.",
    "1770036": "Pseudo labeling is very powerful in this competition. Our team did not use pseudo label during the competition, but we just tried pseudo labels now after the competition and posted our results in an UPDATE section in our solution [here][1]\n\nThe simpliest way to use pseudo is to take your `submission.csv` file and identify your model's 66% most confident predictions. If you use embeddings, then use best cosine similarity. If you use a voting classifier ensemble, then use predictions with most counts. Either way make a new dataset with those 66% best test prediction individual id.\n\nNext train your best models with train plus the new datasets created from test data with pseudo labels. That single model should perform similar to your ensemble LB score. Make a new ensemble with 50% new single model and 50% old ensemble. That should boost your LB.\n\nYou can then repeat this process making pseudo labels from your new ensemble which should achieve a new best LB\n\n[1]: https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/320298",
    "1770652": "Thanks for your reply, your detailed process helped me a lot.  :=)"
  },
  "source": "meta"
}