{
  "id": 417379,
  "title": "💪PSEUDO LABELLING-\"Unlocking Unlabeled Data\" 💪🚀",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/417379",
  "author_name": "",
  "post_date": "2023-06-15T13:07:40.298390800Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<h2>-&gt; <strong>Dataset 3 includes tiles extracted from an additional 9 WSIs</strong></h2>\n<h2>-&gt; <strong>These images are not annotated, i.e. it is not labelled.</strong></h2>\n<h2>-&gt; <strong>Hence we can train the model on dataset1 and dataset2 to predict the output of the dataset3.</strong></h2>\n<h2>-&gt; <strong>Then using the output of the dataset3(CAN TAKE ONLY THE CONFIDENT PREDICTIONS) as the label of it, we can train a model with all the 3 datasets now to get a better model</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11971090%2Fc315ab508925dec53176f3f187cff942%2FPseudo-Labeling-Learning-Architecture.png?generation=1686834604096016&amp;alt=media\" alt=\"Pseudo labelling image\"></p>",
  "messages": [
    {
      "id": "2303759",
      "postDate": "06/15/2023 13:07:40",
      "content": "<h2>-&gt; <strong>Dataset 3 includes tiles extracted from an additional 9 WSIs</strong></h2>\n<h2>-&gt; <strong>These images are not annotated, i.e. it is not labelled.</strong></h2>\n<h2>-&gt; <strong>Hence we can train the model on dataset1 and dataset2 to predict the output of the dataset3.</strong></h2>\n<h2>-&gt; <strong>Then using the output of the dataset3(CAN TAKE ONLY THE CONFIDENT PREDICTIONS) as the label of it, we can train a model with all the 3 datasets now to get a better model</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11971090%2Fc315ab508925dec53176f3f187cff942%2FPseudo-Labeling-Learning-Architecture.png?generation=1686834604096016&amp;alt=media\" alt=\"Pseudo labelling image\"></p>",
      "rawMarkdown": "## -> **Dataset 3 includes tiles extracted from an additional 9 WSIs**\n\n## -> **These images are not annotated, i.e. it is not labelled.**\n\n## -> **Hence we can train the model on dataset1 and dataset2 to predict the output of the dataset3.** \n\n## -> **Then using the output of the dataset3(CAN TAKE ONLY THE CONFIDENT PREDICTIONS) as the label of it, we can train a model with all the 3 datasets now to get a better model**\n\n![Pseudo labelling image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11971090%2Fc315ab508925dec53176f3f187cff942%2FPseudo-Labeling-Learning-Architecture.png?generation=1686834604096016&alt=media)",
      "votes": null
    },
    {
      "id": "2303800",
      "postDate": "06/15/2023 13:33:27",
      "content": "<p>I am experimenting with this approach. I've kept val constant across the old and the new model and experimented with predicting on unlabelled training data with confidence of 0.95, 0.9, 0.85 (i.e high confidence). So far, I've had little success. The new model has lower mAPs than the old model though I haven't tried the new model on the lb test set. Note, lower the confidence, lower the mAP. I would be interested to hear if anyone else has managed to make this approach work or not work.</p>",
      "rawMarkdown": "I am experimenting with this approach. I've kept val constant across the old and the new model and experimented with predicting on unlabelled training data with confidence of 0.95, 0.9, 0.85 (i.e high confidence). So far, I've had little success. The new model has lower mAPs than the old model though I haven't tried the new model on the lb test set. Note, lower the confidence, lower the mAP. I would be interested to hear if anyone else has managed to make this approach work or not work.",
      "votes": null
    },
    {
      "id": "2303924",
      "postDate": "06/15/2023 14:56:21",
      "content": "<p>use dataset 1 for the predictions of the dataset3 then you will be getting more better labels as dataset 1 is the expert verified dataset.</p>",
      "rawMarkdown": "use dataset 1 for the predictions of the dataset3 then you will be getting more better labels as dataset 1 is the expert verified dataset.",
      "votes": null
    },
    {
      "id": "2304278",
      "postDate": "06/15/2023 21:43:22",
      "content": "<p>The problem is that training just on dataset1 results in sub-optimal mAP because the dataset1 is quite small. Dataset1 has 422 images, dataset2 has 1211 images, dataset3 has 5400 images. </p>",
      "rawMarkdown": "The problem is that training just on dataset1 results in sub-optimal mAP because the dataset1 is quite small. Dataset1 has 422 images, dataset2 has 1211 images, dataset3 has 5400 images.",
      "votes": null
    },
    {
      "id": "2314066",
      "postDate": "06/23/2023 06:18:00",
      "content": "<p>I just post a <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418908\" target=\"_blank\">discussion</a>, my training precision has a sudden decrese, is it possible caused by the different dataset?</p>",
      "rawMarkdown": "I just post a [discussion](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418908), my training precision has a sudden decrese, is it possible caused by the different dataset?",
      "votes": null
    },
    {
      "id": "2314108",
      "postDate": "06/23/2023 06:54:39",
      "content": "<p>what you can do is train on the dataset 2 and keep dataset 1 for validation.<br>\nThen after training this model..<br>\nThen use this one to predict the dataset 3 </p>",
      "rawMarkdown": "what you can do is train on the dataset 2 and keep dataset 1 for validation.\nThen after training this model..\nThen use this one to predict the dataset 3",
      "votes": null
    },
    {
      "id": "2372265",
      "postDate": "08/03/2023 15:31:14",
      "content": "<p>Hi, how did you determinate THE CONFIDENT PREDICTIONS?.</p>\n<p>Thanks in advance</p>",
      "rawMarkdown": "Hi, how did you determinate THE CONFIDENT PREDICTIONS?.\n\nThanks in advance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303800,
      "author_name": "achandra1",
      "author_url": "",
      "post_date": "06/15/2023 13:33:27",
      "content": "<p>I am experimenting with this approach. I've kept val constant across the old and the new model and experimented with predicting on unlabelled training data with confidence of 0.95, 0.9, 0.85 (i.e high confidence). So far, I've had little success. The new model has lower mAPs than the old model though I haven't tried the new model on the lb test set. Note, lower the confidence, lower the mAP. I would be interested to hear if anyone else has managed to make this approach work or not work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303924,
          "author_name": "vishakkbhat",
          "author_url": "",
          "post_date": "06/15/2023 14:56:21",
          "content": "<p>use dataset 1 for the predictions of the dataset3 then you will be getting more better labels as dataset 1 is the expert verified dataset.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2304278,
              "author_name": "achandra1",
              "author_url": "",
              "post_date": "06/15/2023 21:43:22",
              "content": "<p>The problem is that training just on dataset1 results in sub-optimal mAP because the dataset1 is quite small. Dataset1 has 422 images, dataset2 has 1211 images, dataset3 has 5400 images. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2314066,
                  "author_name": "chg0901",
                  "author_url": "",
                  "post_date": "06/23/2023 06:18:00",
                  "content": "<p>I just post a <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418908\" target=\"_blank\">discussion</a>, my training precision has a sudden decrese, is it possible caused by the different dataset?</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2314108,
                  "author_name": "vishakkbhat",
                  "author_url": "",
                  "post_date": "06/23/2023 06:54:39",
                  "content": "<p>what you can do is train on the dataset 2 and keep dataset 1 for validation.<br>\nThen after training this model..<br>\nThen use this one to predict the dataset 3 </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2372265,
      "author_name": "pablolarrosa",
      "author_url": "",
      "post_date": "08/03/2023 15:31:14",
      "content": "<p>Hi, how did you determinate THE CONFIDENT PREDICTIONS?.</p>\n<p>Thanks in advance</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2303759": "## -> **Dataset 3 includes tiles extracted from an additional 9 WSIs**\n\n## -> **These images are not annotated, i.e. it is not labelled.**\n\n## -> **Hence we can train the model on dataset1 and dataset2 to predict the output of the dataset3.** \n\n## -> **Then using the output of the dataset3(CAN TAKE ONLY THE CONFIDENT PREDICTIONS) as the label of it, we can train a model with all the 3 datasets now to get a better model**\n\n![Pseudo labelling image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11971090%2Fc315ab508925dec53176f3f187cff942%2FPseudo-Labeling-Learning-Architecture.png?generation=1686834604096016&alt=media)",
    "2303800": "I am experimenting with this approach. I've kept val constant across the old and the new model and experimented with predicting on unlabelled training data with confidence of 0.95, 0.9, 0.85 (i.e high confidence). So far, I've had little success. The new model has lower mAPs than the old model though I haven't tried the new model on the lb test set. Note, lower the confidence, lower the mAP. I would be interested to hear if anyone else has managed to make this approach work or not work.",
    "2303924": "use dataset 1 for the predictions of the dataset3 then you will be getting more better labels as dataset 1 is the expert verified dataset.",
    "2304278": "The problem is that training just on dataset1 results in sub-optimal mAP because the dataset1 is quite small. Dataset1 has 422 images, dataset2 has 1211 images, dataset3 has 5400 images.",
    "2314066": "I just post a [discussion](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/418908), my training precision has a sudden decrese, is it possible caused by the different dataset?",
    "2314108": "what you can do is train on the dataset 2 and keep dataset 1 for validation.\nThen after training this model..\nThen use this one to predict the dataset 3",
    "2372265": "Hi, how did you determinate THE CONFIDENT PREDICTIONS?.\n\nThanks in advance"
  },
  "source": "meta"
}