{
  "id": 535023,
  "title": "Modeling Approach: Semi-Supervised Learning",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535023",
  "author_name": "Andreas Bisiadis",
  "post_date": "2024-09-19T20:35:20.260000",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<blockquote>\n  <p>The majority of measures are missing for most participants. In particular, the target sii is missing for about 25% of participants in the training set. You may wish to apply non-supervised learning techniques to this data. </p>\n</blockquote>\n<p>An simple yet effective approach, that focuses on the partial absence of the target column, would be:</p>\n<ol>\n<li>Train a LightGBM model on the labeled samples.</li>\n<li>Use the previous model to generate predictions on the unlabaled samples.</li>\n<li>Use the predictions as (pseudo) target and add this target to the initial training data.</li>\n<li>Having a dataset with both <code>labeled</code> and <code>pseudo-labeled</code> samples, we train once again a (LightGBM) model. This model will be used for inference on the test data.</li>\n</ol>\n<p>For more information, visit <a href=\"https://www.ibm.com/topics/semi-supervised-learning#:~:text=Contributors%3A%20Dave%20Bergmann-,What%20is%20semi%2Dsupervised%20learning%3F,for%20classification%20and%20regression%20tasks.\" target=\"_blank\">this</a> page.</p>",
  "messages": [
    {
      "id": 2993465,
      "postDate": "2024-09-19T20:35:20.260Z",
      "content": "<blockquote>\n  <p>The majority of measures are missing for most participants. In particular, the target sii is missing for about 25% of participants in the training set. You may wish to apply non-supervised learning techniques to this data. </p>\n</blockquote>\n<p>An simple yet effective approach, that focuses on the partial absence of the target column, would be:</p>\n<ol>\n<li>Train a LightGBM model on the labeled samples.</li>\n<li>Use the previous model to generate predictions on the unlabaled samples.</li>\n<li>Use the predictions as (pseudo) target and add this target to the initial training data.</li>\n<li>Having a dataset with both <code>labeled</code> and <code>pseudo-labeled</code> samples, we train once again a (LightGBM) model. This model will be used for inference on the test data.</li>\n</ol>\n<p>For more information, visit <a href=\"https://www.ibm.com/topics/semi-supervised-learning#:~:text=Contributors%3A%20Dave%20Bergmann-,What%20is%20semi%2Dsupervised%20learning%3F,for%20classification%20and%20regression%20tasks.\" target=\"_blank\">this</a> page.</p>",
      "rawMarkdown": ">  The majority of measures are missing for most participants. In particular, the target sii is missing for about 25% of participants in the training set. You may wish to apply non-supervised learning techniques to this data. \n\nAn simple yet effective approach, that focuses on the partial absence of the target column, would be:\n\n1. Train a LightGBM model on the labeled samples.\n2. Use the previous model to generate predictions on the unlabaled samples.\n3. Use the predictions as (pseudo) target and add this target to the initial training data.\n4. Having a dataset with both `labeled` and `pseudo-labeled` samples, we train once again a (LightGBM) model. This model will be used for inference on the test data.\n\nFor more information, visit [this](https://www.ibm.com/topics/semi-supervised-learning#:~:text=Contributors%3A%20Dave%20Bergmann-,What%20is%20semi%2Dsupervised%20learning%3F,for%20classification%20and%20regression%20tasks.) page.",
      "votes": 26
    },
    {
      "id": 3017281,
      "postDate": "2024-10-14T18:01:49.043Z",
      "content": "<p>Andrew, did you have a go at implementing pseudo-labelling? I was thinking of giving it a try, but when I looked at the rows with NAN's in the SII I found a lot of other data was NAN, which made me a little wary of this approach.</p>\n<p>I've already used imputation to 'patch' up the NANs which improved my results when training with the SII NAN rows missing, but the number or rows data where SII is NAN is only 1224. That compares to 2736 where SII has a value.</p>\n<p>So its not like are there are huge amounts of data that are unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).</p>\n<p>…. A few more details … the \"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data … but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:<br>\nPhysical-Season                            509<br>\nPhysical-BMI                               729<br>\nPhysical-Height                            727<br>\nPhysical-Weight                            720<br>\nPhysical-Waist_Circumference               809<br>\nPhysical-Diastolic_BP                      748<br>\nPhysical-HeartRate                         743<br>\nPhysical-Systolic_BP                       748<br>\nPreInt_EduHx-computerinternet_hoursday     577</p>\n<p>The physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.</p>",
      "rawMarkdown": "Andrew, did you have a go at implementing pseudo-labelling? I was thinking of giving it a try, but when I looked at the rows with NAN's in the SII I found a lot of other data was NAN, which made me a little wary of this approach.\n\nI've already used imputation to 'patch' up the NANs which improved my results when training with the SII NAN rows missing, but the number or rows data where SII is NAN is only 1224. That compares to 2736 where SII has a value.\n\nSo its not like are there are huge amounts of data that are unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).\n\n.... A few more details ... the \"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data ... but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:\nPhysical-Season                            509\nPhysical-BMI                               729\nPhysical-Height                            727\nPhysical-Weight                            720\nPhysical-Waist_Circumference               809\nPhysical-Diastolic_BP                      748\nPhysical-HeartRate                         743\nPhysical-Systolic_BP                       748\nPreInt_EduHx-computerinternet_hoursday     577\n\nThe physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.",
      "votes": 3
    },
    {
      "id": 2993514,
      "postDate": "2024-09-19T21:45:47.100Z",
      "content": "<p>Perhaps it would be beneficial for the final model to be different from the one generating the pseudo labels, so it is not just enforcing what the previous model learned. I could see a confidence measure derived from the standard deviation of CV predictions for the pseudo labels across different models being useful as well. </p>\n<p>I also wonder if training a binary classifier to detect unlabelled training data would be of any use.</p>\n<p>I'm new to kaggle so these are just some ideas I'm throwing out here.</p>",
      "rawMarkdown": "Perhaps it would be beneficial for the final model to be different from the one generating the pseudo labels, so it is not just enforcing what the previous model learned. I could see a confidence measure derived from the standard deviation of CV predictions for the pseudo labels across different models being useful as well. \n\nI also wonder if training a binary classifier to detect unlabelled training data would be of any use.\n\nI'm new to kaggle so these are just some ideas I'm throwing out here.",
      "votes": 1,
      "replies": [
        {
          "id": 2993650,
          "postDate": "2024-09-20T04:40:30.237Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2995978,
      "postDate": "2024-09-22T20:46:26.253Z",
      "content": "<p>Have you tried any imputation for missing values first? I think the rows with missing sii might still be useful for imputing any values for nan in other fields. </p>",
      "rawMarkdown": "Have you tried any imputation for missing values first? I think the rows with missing sii might still be useful for imputing any values for nan in other fields. ",
      "votes": 2
    },
    {
      "id": 3523108,
      "postDate": "2026-09-10T15:50:10.107Z",
      "content": "<p>This dataset raises an intriguing question: <strong>what can ML learn before every example is labeled?</strong> 🤔</p>\n<p><a href=\"https://mlguidance.blogspot.com/2026/09/what-are-models-of-semi-supervised.html\" target=\"_blank\">https://mlguidance.blogspot.com/2026/09/what-are-models-of-semi-supervised.html</a></p>",
      "rawMarkdown": "This dataset raises an intriguing question: **what can ML learn before every example is labeled?** 🤔\n\nhttps://mlguidance.blogspot.com/2026/09/what-are-models-of-semi-supervised.html"
    },
    {
      "id": 3013037,
      "postDate": "2024-10-09T15:45:19.877Z",
      "content": "<p>I almost missed this discussion and it's <a href=\"https://www.kaggle.com/wti200\" target=\"_blank\">@wti200</a> who points me here! Thanks both.<br>\nThe IBM graph is fab, showing how the semi can help.</p>",
      "rawMarkdown": "I almost missed this discussion and it's @wti200 who points me here! Thanks both.\nThe IBM graph is fab, showing how the semi can help."
    },
    {
      "id": 2993814,
      "postDate": "2024-09-20T09:16:23.247Z",
      "content": "<p>this seems like an obvious first choice, and that usually means there's a better way</p>",
      "rawMarkdown": "this seems like an obvious first choice, and that usually means there's a better way"
    }
  ],
  "comments": [
    {
      "id": 3017281,
      "author_name": "stuart houghton",
      "author_url": "",
      "post_date": "2024-10-14T18:01:49.043000",
      "content": "<p>Andrew, did you have a go at implementing pseudo-labelling? I was thinking of giving it a try, but when I looked at the rows with NAN's in the SII I found a lot of other data was NAN, which made me a little wary of this approach.</p>\n<p>I've already used imputation to 'patch' up the NANs which improved my results when training with the SII NAN rows missing, but the number or rows data where SII is NAN is only 1224. That compares to 2736 where SII has a value.</p>\n<p>So its not like are there are huge amounts of data that are unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).</p>\n<p>…. A few more details … the \"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data … but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:<br>\nPhysical-Season                            509<br>\nPhysical-BMI                               729<br>\nPhysical-Height                            727<br>\nPhysical-Weight                            720<br>\nPhysical-Waist_Circumference               809<br>\nPhysical-Diastolic_BP                      748<br>\nPhysical-HeartRate                         743<br>\nPhysical-Systolic_BP                       748<br>\nPreInt_EduHx-computerinternet_hoursday     577</p>\n<p>The physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2993514,
      "author_name": "MXMLNV",
      "author_url": "",
      "post_date": "2024-09-19T21:45:47.100000",
      "content": "<p>Perhaps it would be beneficial for the final model to be different from the one generating the pseudo labels, so it is not just enforcing what the previous model learned. I could see a confidence measure derived from the standard deviation of CV predictions for the pseudo labels across different models being useful as well. </p>\n<p>I also wonder if training a binary classifier to detect unlabelled training data would be of any use.</p>\n<p>I'm new to kaggle so these are just some ideas I'm throwing out here.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2993650,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-09-20T04:40:30.237000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2995978,
      "author_name": "Chandan Nayak",
      "author_url": "",
      "post_date": "2024-09-22T20:46:26.253000",
      "content": "<p>Have you tried any imputation for missing values first? I think the rows with missing sii might still be useful for imputing any values for nan in other fields. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3523108,
      "author_name": "Kaustubh Verma",
      "author_url": "",
      "post_date": "2026-09-10T15:50:10.107000",
      "content": "<p>This dataset raises an intriguing question: <strong>what can ML learn before every example is labeled?</strong> 🤔</p>\n<p><a href=\"https://mlguidance.blogspot.com/2026/09/what-are-models-of-semi-supervised.html\" target=\"_blank\">https://mlguidance.blogspot.com/2026/09/what-are-models-of-semi-supervised.html</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3013037,
      "author_name": "Tom Yuen",
      "author_url": "",
      "post_date": "2024-10-09T15:45:19.877000",
      "content": "<p>I almost missed this discussion and it's <a href=\"https://www.kaggle.com/wti200\" target=\"_blank\">@wti200</a> who points me here! Thanks both.<br>\nThe IBM graph is fab, showing how the semi can help.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2993814,
      "author_name": "g john rao",
      "author_url": "",
      "post_date": "2024-09-20T09:16:23.247000",
      "content": "<p>this seems like an obvious first choice, and that usually means there's a better way</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2993465": ">  The majority of measures are missing for most participants. In particular, the target sii is missing for about 25% of participants in the training set. You may wish to apply non-supervised learning techniques to this data. \n\nAn simple yet effective approach, that focuses on the partial absence of the target column, would be:\n\n1. Train a LightGBM model on the labeled samples.\n2. Use the previous model to generate predictions on the unlabaled samples.\n3. Use the predictions as (pseudo) target and add this target to the initial training data.\n4. Having a dataset with both `labeled` and `pseudo-labeled` samples, we train once again a (LightGBM) model. This model will be used for inference on the test data.\n\nFor more information, visit [this](https://www.ibm.com/topics/semi-supervised-learning#:~:text=Contributors%3A%20Dave%20Bergmann-,What%20is%20semi%2Dsupervised%20learning%3F,for%20classification%20and%20regression%20tasks.) page.",
    "3017281": "Andrew, did you have a go at implementing pseudo-labelling? I was thinking of giving it a try, but when I looked at the rows with NAN's in the SII I found a lot of other data was NAN, which made me a little wary of this approach.\n\nI've already used imputation to 'patch' up the NANs which improved my results when training with the SII NAN rows missing, but the number or rows data where SII is NAN is only 1224. That compares to 2736 where SII has a value.\n\nSo its not like are there are huge amounts of data that are unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).\n\n.... A few more details ... the \"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data ... but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:\nPhysical-Season                            509\nPhysical-BMI                               729\nPhysical-Height                            727\nPhysical-Weight                            720\nPhysical-Waist_Circumference               809\nPhysical-Diastolic_BP                      748\nPhysical-HeartRate                         743\nPhysical-Systolic_BP                       748\nPreInt_EduHx-computerinternet_hoursday     577\n\nThe physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.",
    "2993514": "Perhaps it would be beneficial for the final model to be different from the one generating the pseudo labels, so it is not just enforcing what the previous model learned. I could see a confidence measure derived from the standard deviation of CV predictions for the pseudo labels across different models being useful as well. \n\nI also wonder if training a binary classifier to detect unlabelled training data would be of any use.\n\nI'm new to kaggle so these are just some ideas I'm throwing out here.",
    "2995978": "Have you tried any imputation for missing values first? I think the rows with missing sii might still be useful for imputing any values for nan in other fields. ",
    "3523108": "This dataset raises an intriguing question: **what can ML learn before every example is labeled?** 🤔\n\nhttps://mlguidance.blogspot.com/2026/09/what-are-models-of-semi-supervised.html",
    "3013037": "I almost missed this discussion and it's @wti200 who points me here! Thanks both.\nThe IBM graph is fab, showing how the semi can help.",
    "2993814": "this seems like an obvious first choice, and that usually means there's a better way"
  }
}