{
  "id": 540476,
  "title": "Approaches to unlabelled data and sparsity of features ...",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/540476",
  "author_name": "stuart houghton",
  "post_date": "2024-10-14T18:19:16.744000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Sparsity of data see,s like a real problem here. I was thinking about implementing pseudo-labelling, but when I looked at the unlabeled rows, I found a lot of other data was other features were NAN, which made me a little wary of this approach.</p>\n<p>I've already used imputation to 'patch' up the features NANs which improved my results when training with the unlabeled rows removed, but the number or unlabeled rows is only 1224. That compares to 2736 where SII has a value.</p>\n<p>So its not like are there are huge amounts of data that is unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).</p>\n<p>\"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data … but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:<br>\nPhysical-Season                            509<br>\nPhysical-BMI                               729<br>\nPhysical-Height                            727<br>\nPhysical-Weight                            720<br>\nPhysical-Waist_Circumference               809<br>\nPhysical-Diastolic_BP                      748<br>\nPhysical-HeartRate                         743<br>\nPhysical-Systolic_BP                       748<br>\nPreInt_EduHx-computerinternet_hoursday     577</p>\n<p>The physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.</p>\n<p>Any thoughts on other approaches? Imputation seems to have yielded the best results for me so far.</p>",
  "messages": [
    {
      "id": 3017294,
      "postDate": "2024-10-14T18:19:16.743Z",
      "content": "<p>Sparsity of data see,s like a real problem here. I was thinking about implementing pseudo-labelling, but when I looked at the unlabeled rows, I found a lot of other data was other features were NAN, which made me a little wary of this approach.</p>\n<p>I've already used imputation to 'patch' up the features NANs which improved my results when training with the unlabeled rows removed, but the number or unlabeled rows is only 1224. That compares to 2736 where SII has a value.</p>\n<p>So its not like are there are huge amounts of data that is unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).</p>\n<p>\"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data … but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:<br>\nPhysical-Season                            509<br>\nPhysical-BMI                               729<br>\nPhysical-Height                            727<br>\nPhysical-Weight                            720<br>\nPhysical-Waist_Circumference               809<br>\nPhysical-Diastolic_BP                      748<br>\nPhysical-HeartRate                         743<br>\nPhysical-Systolic_BP                       748<br>\nPreInt_EduHx-computerinternet_hoursday     577</p>\n<p>The physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.</p>\n<p>Any thoughts on other approaches? Imputation seems to have yielded the best results for me so far.</p>",
      "rawMarkdown": "Sparsity of data see,s like a real problem here. I was thinking about implementing pseudo-labelling, but when I looked at the unlabeled rows, I found a lot of other data was other features were NAN, which made me a little wary of this approach.\n\nI've already used imputation to 'patch' up the features NANs which improved my results when training with the unlabeled rows removed, but the number or unlabeled rows is only 1224. That compares to 2736 where SII has a value.\n\nSo its not like are there are huge amounts of data that is unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).\n\n\"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data ... but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:\nPhysical-Season                            509\nPhysical-BMI                               729\nPhysical-Height                            727\nPhysical-Weight                            720\nPhysical-Waist_Circumference               809\nPhysical-Diastolic_BP                      748\nPhysical-HeartRate                         743\nPhysical-Systolic_BP                       748\nPreInt_EduHx-computerinternet_hoursday     577\n\nThe physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.\n\nAny thoughts on other approaches? Imputation seems to have yielded the best results for me so far.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3017294": "Sparsity of data see,s like a real problem here. I was thinking about implementing pseudo-labelling, but when I looked at the unlabeled rows, I found a lot of other data was other features were NAN, which made me a little wary of this approach.\n\nI've already used imputation to 'patch' up the features NANs which improved my results when training with the unlabeled rows removed, but the number or unlabeled rows is only 1224. That compares to 2736 where SII has a value.\n\nSo its not like are there are huge amounts of data that is unlabelled and the quality of the data avaiable in the unlabelled data is poor too (lots of other empty features).\n\n\"Physical\" and \"PreInt_EduHx\" columns values are most abundant with very few missing values in the labeled data ... but when it comes to unlabeled rows (1224 rows), these columns have the following number of missing values:\nPhysical-Season                            509\nPhysical-BMI                               729\nPhysical-Height                            727\nPhysical-Weight                            720\nPhysical-Waist_Circumference               809\nPhysical-Diastolic_BP                      748\nPhysical-HeartRate                         743\nPhysical-Systolic_BP                       748\nPreInt_EduHx-computerinternet_hoursday     577\n\nThe physical columns are also the most correlated to SII (apart from the PCIAT columns which are basically how SII is calcuated) but maybe that only due to their relative abundance.\n\nAny thoughts on other approaches? Imputation seems to have yielded the best results for me so far."
  }
}