{
  "id": 544423,
  "title": "Pseudo Labels It may Work?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/544423",
  "author_name": "",
  "post_date": "2024-11-05T04:16:04.371580200Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>some target (sii) are missing, so we can predicting through training a model on available (sii) , we retrain the model and we compare the performance of the model before and after leveraging the missing labels(semi-supervised)</p>",
  "messages": [
    {
      "id": "3036929",
      "postDate": "11/05/2024 04:16:04",
      "content": "<p>some target (sii) are missing, so we can predicting through training a model on available (sii) , we retrain the model and we compare the performance of the model before and after leveraging the missing labels(semi-supervised)</p>",
      "rawMarkdown": "some target (sii) are missing, so we can predicting through training a model on available (sii) , we retrain the model and we compare the performance of the model before and after leveraging the missing labels(semi-supervised)",
      "votes": null
    },
    {
      "id": "3037299",
      "postDate": "11/05/2024 15:15:18",
      "content": "<p>You want to predict the labels and then take these as a ground truth for training a final model - it's a BIG no-no, because you are lying to your final model. It's simpler just to use SMOTE to get a lot of synthetic data, but not to serve some predictions as true labels.</p>",
      "rawMarkdown": "You want to predict the labels and then take these as a ground truth for training a final model - it's a BIG no-no, because you are lying to your final model. It's simpler just to use SMOTE to get a lot of synthetic data, but not to serve some predictions as true labels.",
      "votes": null
    },
    {
      "id": "3037633",
      "postDate": "11/05/2024 22:31:47",
      "content": "<p>Once you start using Pseudo labels, your Train score &amp; cv score will start to improve, I guess because we're filling the missed sii with the obvious signals. And I guess the cv score is not reliable anymore<br>\nHere's one of my many examples, it doesn't get worse, but it doesn't improve.</p>\n<p>Mean Train QWK --&gt; 0.6671<br>\nMean Validation QWK ---&gt; 0.5828<br>\nOptimized thresholds: [0.5198199  1.2157459  2.55957693]<br>\n----&gt; || Optimized QWK SCORE ::  0.625<br>\nPublic score 0.453</p>",
      "rawMarkdown": "Once you start using Pseudo labels, your Train score & cv score will start to improve, I guess because we're filling the missed sii with the obvious signals. And I guess the cv score is not reliable anymore\nHere's one of my many examples, it doesn't get worse, but it doesn't improve.\n\n\nMean Train QWK --> 0.6671\nMean Validation QWK ---> 0.5828\nOptimized thresholds: [0.5198199  1.2157459  2.55957693]\n----> || Optimized QWK SCORE ::  0.625\nPublic score 0.453",
      "votes": null
    },
    {
      "id": "3037922",
      "postDate": "11/06/2024 11:18:59",
      "content": "<p>Of course, pseudolabeling may work, but it needs a lot of work to see an improvement. </p>",
      "rawMarkdown": "Of course, pseudolabeling may work, but it needs a lot of work to see an improvement.",
      "votes": null
    },
    {
      "id": "3037955",
      "postDate": "11/06/2024 12:08:29",
      "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> , i think you saying that thre final model may detect some pattern that related to the pretrained model are irrelevant to real pattern between the features and true labels</p>",
      "rawMarkdown": "eu1234 , i think you saying that thre final model may detect some pattern that related to the pretrained model are irrelevant to real pattern between the features and true labels",
      "votes": null
    },
    {
      "id": "3037957",
      "postDate": "11/06/2024 12:12:54",
      "content": "<p><a href=\"https://www.kaggle.com/jankowalski2000\" target=\"_blank\">@jankowalski2000</a> , the problem is you can rely on the pseudolabeling results , such as <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> it may not reflect the real pattern in the data, in other word , like to train the model to increase score in the train data not to detect pattern into data (may lead to overfitting)</p>",
      "rawMarkdown": "jankowalski2000 , the problem is you can rely on the pseudolabeling results , such as @eu1234 it may not reflect the real pattern in the data, in other word , like to train the model to increase score in the train data not to detect pattern into data (may lead to overfitting)",
      "votes": null
    },
    {
      "id": "3037959",
      "postDate": "11/06/2024 12:14:32",
      "content": "<p>Yes, training on wrongly predicted labels and taking them as truth you are misleading your final model. </p>",
      "rawMarkdown": "Yes, training on wrongly predicted labels and taking them as truth you are misleading your final model.",
      "votes": null
    },
    {
      "id": "3037960",
      "postDate": "11/06/2024 12:15:22",
      "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> , honetely i don't have experiments with pseudo labels approach, but in Overview of the competition , unsupervsed learning approach mentioned that why i think semi-supervised may be beneficial</p>",
      "rawMarkdown": "tomyuen , honetely i don't have experiments with pseudo labels approach, but in Overview of the competition , unsupervsed learning approach mentioned that why i think semi-supervised may be beneficial",
      "votes": null
    },
    {
      "id": "3037967",
      "postDate": "11/06/2024 12:26:04",
      "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> absolutely, ithink 6.7G of tabular data i atleast mid-size data so in necessary we don't need to augment data risking to misleading our model</p>",
      "rawMarkdown": "eu1234 absolutely, ithink 6.7G of tabular data i atleast mid-size data so in necessary we don't need to augment data risking to misleading our model",
      "votes": null
    },
    {
      "id": "3045773",
      "postDate": "11/14/2024 19:13:21",
      "content": "<p>Good job, Tom 😃</p>\n<p>Totally agree with <a href=\"https://www.kaggle.com/saidkoussi\" target=\"_blank\">@saidkoussi</a>, being semi-supervised will help the final rusult.</p>\n<p>Perhaps you could try giving different weights? Specifically, the pseudo-labels can have less influence on the model than the model with the truth labels.</p>\n<p>Looking forward to your exciting results!</p>",
      "rawMarkdown": "Good job, Tom 😃\n\nTotally agree with @saidkoussi, being semi-supervised will help the final rusult.\n\nPerhaps you could try giving different weights? Specifically, the pseudo-labels can have less influence on the model than the model with the truth labels.\n\nLooking forward to your exciting results!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3037299,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "11/05/2024 15:15:18",
      "content": "<p>You want to predict the labels and then take these as a ground truth for training a final model - it's a BIG no-no, because you are lying to your final model. It's simpler just to use SMOTE to get a lot of synthetic data, but not to serve some predictions as true labels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3037955,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "11/06/2024 12:08:29",
          "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> , i think you saying that thre final model may detect some pattern that related to the pretrained model are irrelevant to real pattern between the features and true labels</p>",
          "votes": null,
          "replies": [
            {
              "id": 3037959,
              "author_name": "eu1234",
              "author_url": "",
              "post_date": "11/06/2024 12:14:32",
              "content": "<p>Yes, training on wrongly predicted labels and taking them as truth you are misleading your final model. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3037967,
                  "author_name": "saidkoussi",
                  "author_url": "",
                  "post_date": "11/06/2024 12:26:04",
                  "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> absolutely, ithink 6.7G of tabular data i atleast mid-size data so in necessary we don't need to augment data risking to misleading our model</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3037633,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "11/05/2024 22:31:47",
      "content": "<p>Once you start using Pseudo labels, your Train score &amp; cv score will start to improve, I guess because we're filling the missed sii with the obvious signals. And I guess the cv score is not reliable anymore<br>\nHere's one of my many examples, it doesn't get worse, but it doesn't improve.</p>\n<p>Mean Train QWK --&gt; 0.6671<br>\nMean Validation QWK ---&gt; 0.5828<br>\nOptimized thresholds: [0.5198199  1.2157459  2.55957693]<br>\n----&gt; || Optimized QWK SCORE ::  0.625<br>\nPublic score 0.453</p>",
      "votes": null,
      "replies": [
        {
          "id": 3037960,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "11/06/2024 12:15:22",
          "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> , honetely i don't have experiments with pseudo labels approach, but in Overview of the competition , unsupervsed learning approach mentioned that why i think semi-supervised may be beneficial</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3045773,
          "author_name": "fangzitao",
          "author_url": "",
          "post_date": "11/14/2024 19:13:21",
          "content": "<p>Good job, Tom 😃</p>\n<p>Totally agree with <a href=\"https://www.kaggle.com/saidkoussi\" target=\"_blank\">@saidkoussi</a>, being semi-supervised will help the final rusult.</p>\n<p>Perhaps you could try giving different weights? Specifically, the pseudo-labels can have less influence on the model than the model with the truth labels.</p>\n<p>Looking forward to your exciting results!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3037922,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "11/06/2024 11:18:59",
      "content": "<p>Of course, pseudolabeling may work, but it needs a lot of work to see an improvement. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3037957,
          "author_name": "saidkoussi",
          "author_url": "",
          "post_date": "11/06/2024 12:12:54",
          "content": "<p><a href=\"https://www.kaggle.com/jankowalski2000\" target=\"_blank\">@jankowalski2000</a> , the problem is you can rely on the pseudolabeling results , such as <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> it may not reflect the real pattern in the data, in other word , like to train the model to increase score in the train data not to detect pattern into data (may lead to overfitting)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3036929": "some target (sii) are missing, so we can predicting through training a model on available (sii) , we retrain the model and we compare the performance of the model before and after leveraging the missing labels(semi-supervised)",
    "3037299": "You want to predict the labels and then take these as a ground truth for training a final model - it's a BIG no-no, because you are lying to your final model. It's simpler just to use SMOTE to get a lot of synthetic data, but not to serve some predictions as true labels.",
    "3037633": "Once you start using Pseudo labels, your Train score & cv score will start to improve, I guess because we're filling the missed sii with the obvious signals. And I guess the cv score is not reliable anymore\nHere's one of my many examples, it doesn't get worse, but it doesn't improve.\n\n\nMean Train QWK --> 0.6671\nMean Validation QWK ---> 0.5828\nOptimized thresholds: [0.5198199  1.2157459  2.55957693]\n----> || Optimized QWK SCORE ::  0.625\nPublic score 0.453",
    "3037922": "Of course, pseudolabeling may work, but it needs a lot of work to see an improvement.",
    "3037955": "eu1234 , i think you saying that thre final model may detect some pattern that related to the pretrained model are irrelevant to real pattern between the features and true labels",
    "3037957": "jankowalski2000 , the problem is you can rely on the pseudolabeling results , such as @eu1234 it may not reflect the real pattern in the data, in other word , like to train the model to increase score in the train data not to detect pattern into data (may lead to overfitting)",
    "3037959": "Yes, training on wrongly predicted labels and taking them as truth you are misleading your final model.",
    "3037960": "tomyuen , honetely i don't have experiments with pseudo labels approach, but in Overview of the competition , unsupervsed learning approach mentioned that why i think semi-supervised may be beneficial",
    "3037967": "eu1234 absolutely, ithink 6.7G of tabular data i atleast mid-size data so in necessary we don't need to augment data risking to misleading our model",
    "3045773": "Good job, Tom 😃\n\nTotally agree with @saidkoussi, being semi-supervised will help the final rusult.\n\nPerhaps you could try giving different weights? Specifically, the pseudo-labels can have less influence on the model than the model with the truth labels.\n\nLooking forward to your exciting results!"
  },
  "source": "meta"
}