{
  "id": 537514,
  "title": "How to predict missing TARGET in train.csv？",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/537514",
  "author_name": "",
  "post_date": "2024-10-03T15:16:16.229880200Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>In 3960 rows of train.csv, there are 1224 rows with missing sii. What a big portion!</p>\n<p>【In Competition Datasets Description】<br>\nThe target sii is missing for a portion of the participants in the training set. You may wish to apply non-supervised learning techniques to this data. </p>\n<p>【In Shared Baseline Model】<br>\nI found that dropping 1224 rows directly in training set can also get not bad score in public Leaderboard.</p>\n<p>【Methods I tried to fill missing sii】<br>\n(1) Lightgbm：all the predicted sii is 0.0 ,no matter what parameter I use, I don't know why.<br>\n(2) KNN: the distribution of predicted sii vary greatly from different n_neighbors.</p>",
  "messages": [
    {
      "id": "3005969",
      "postDate": "10/03/2024 15:16:16",
      "content": "<p>In 3960 rows of train.csv, there are 1224 rows with missing sii. What a big portion!</p>\n<p>【In Competition Datasets Description】<br>\nThe target sii is missing for a portion of the participants in the training set. You may wish to apply non-supervised learning techniques to this data. </p>\n<p>【In Shared Baseline Model】<br>\nI found that dropping 1224 rows directly in training set can also get not bad score in public Leaderboard.</p>\n<p>【Methods I tried to fill missing sii】<br>\n(1) Lightgbm：all the predicted sii is 0.0 ,no matter what parameter I use, I don't know why.<br>\n(2) KNN: the distribution of predicted sii vary greatly from different n_neighbors.</p>",
      "rawMarkdown": "In 3960 rows of train.csv, there are 1224 rows with missing sii. What a big portion!\n\n【In Competition Datasets Description】\nThe target sii is missing for a portion of the participants in the training set. You may wish to apply non-supervised learning techniques to this data. \n\n【In Shared Baseline Model】\nI found that dropping 1224 rows directly in training set can also get not bad score in public Leaderboard.\n\n【Methods I tried to fill missing sii】\n(1) Lightgbm：all the predicted sii is 0.0 ,no matter what parameter I use, I don't know why.\n(2) KNN: the distribution of predicted sii vary greatly from different n_neighbors.",
      "votes": null
    },
    {
      "id": "3006218",
      "postDate": "10/03/2024 21:23:59",
      "content": "<p>You can use your normal code, split the train set into train_df and train1_df, train1_df having missing sii. then you use train_df to train train1_df (instead of test_df in normal case), merge, and go ahead with what you usually do.</p>\n<p>I tried many times, I do get a few sii= 2 or 3, but final performance drop. May be the result I feed into train1_df is rubbish. 🤣</p>",
      "rawMarkdown": "You can use your normal code, split the train set into train_df and train1_df, train1_df having missing sii. then you use train_df to train train1_df (instead of test_df in normal case), merge, and go ahead with what you usually do.\n\nI tried many times, I do get a few sii= 2 or 3, but final performance drop. May be the result I feed into train1_df is rubbish. 🤣",
      "votes": null
    },
    {
      "id": "3006282",
      "postDate": "10/04/2024 02:08:57",
      "content": "<p>Thanks.</p>\n<p>I only got few sii =3 when KNN n_neighbors&lt;=3, but I think it is not a reliable method to fill the missing sii.<br>\nBecause the features also have many missing values, I need to fill the missing values in features first, and I also not sure whether the method to fill the missing values in features is reliable.😅</p>",
      "rawMarkdown": "Thanks.\n\nI only got few sii =3 when KNN n_neighbors<=3, but I think it is not a reliable method to fill the missing sii.\nBecause the features also have many missing values, I need to fill the missing values in features first, and I also not sure whether the method to fill the missing values in features is reliable.😅",
      "votes": null
    },
    {
      "id": "3006673",
      "postDate": "10/04/2024 12:17:40",
      "content": "<p>全量样本才5000左右，训练样本3960，有目标样本2736，有额外数据的996个。<br>\n总体的样本数本身已经很小了，而无目标和不带有额外数据的样本数占比很大，意味着你不论是做丢弃和填充，都会严重的影响数据分布，进而影响真实结果。所以实际上没有特别好的方法解决，只能看这个数据集实验下来，丢弃和填充哪种效果更好</p>",
      "rawMarkdown": "全量样本才5000左右，训练样本3960，有目标样本2736，有额外数据的996个。\n总体的样本数本身已经很小了，而无目标和不带有额外数据的样本数占比很大，意味着你不论是做丢弃和填充，都会严重的影响数据分布，进而影响真实结果。所以实际上没有特别好的方法解决，只能看这个数据集实验下来，丢弃和填充哪种效果更好",
      "votes": null
    },
    {
      "id": "3006693",
      "postDate": "10/04/2024 12:39:09",
      "content": "<p>Thanks. I will try both.</p>\n<p>In the rows with known target, the distribution of different sii is :<br>\n0.0    687<br>\n1.0    386<br>\n2.0    142<br>\n3.0      9<br>\nI think that it is possible the same distribution on the unknown data set.</p>",
      "rawMarkdown": "Thanks. I will try both.\n\nIn the rows with known target, the distribution of different sii is :\n0.0    687\n1.0    386\n2.0    142\n3.0      9\nI think that it is possible the same distribution on the unknown data set.",
      "votes": null
    },
    {
      "id": "3006851",
      "postDate": "10/04/2024 15:54:54",
      "content": "<p>很好的想法，但是有个问题是，你不论怎样对无目标的结果做预测，你的结果是作为最后的ground truth来做训练。必然无法保证这1000多个数据的ground truth预测的百分百正确，同时会给最后的那个模型带来很大的干扰。我们最终的目标还是将最后的模型调整的很好，所以对于无target的数据是很难做处理的</p>",
      "rawMarkdown": "很好的想法，但是有个问题是，你不论怎样对无目标的结果做预测，你的结果是作为最后的ground truth来做训练。必然无法保证这1000多个数据的ground truth预测的百分百正确，同时会给最后的那个模型带来很大的干扰。我们最终的目标还是将最后的模型调整的很好，所以对于无target的数据是很难做处理的",
      "votes": null
    },
    {
      "id": "3006955",
      "postDate": "10/04/2024 18:09:54",
      "content": "<p>I am trying to use class_weight and some balancing methods such as smote. My performance is slightly better balancing the classes</p>",
      "rawMarkdown": "I am trying to use class_weight and some balancing methods such as smote. My performance is slightly better balancing the classes",
      "votes": null
    },
    {
      "id": "3007296",
      "postDate": "10/05/2024 07:18:31",
      "content": "<p>Thanks. I will try.</p>\n<p>I tried KNN with metric='wminkowski', and set features weight as pearson correlation with sii.<br>\nBut the performance seems not better than other methods.😭</p>",
      "rawMarkdown": "Thanks. I will try.\n\nI tried KNN with metric='wminkowski', and set features weight as pearson correlation with sii.\nBut the performance seems not better than other methods.😭",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3006218,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/03/2024 21:23:59",
      "content": "<p>You can use your normal code, split the train set into train_df and train1_df, train1_df having missing sii. then you use train_df to train train1_df (instead of test_df in normal case), merge, and go ahead with what you usually do.</p>\n<p>I tried many times, I do get a few sii= 2 or 3, but final performance drop. May be the result I feed into train1_df is rubbish. 🤣</p>",
      "votes": null,
      "replies": [
        {
          "id": 3006282,
          "author_name": "ggggpeushmy",
          "author_url": "",
          "post_date": "10/04/2024 02:08:57",
          "content": "<p>Thanks.</p>\n<p>I only got few sii =3 when KNN n_neighbors&lt;=3, but I think it is not a reliable method to fill the missing sii.<br>\nBecause the features also have many missing values, I need to fill the missing values in features first, and I also not sure whether the method to fill the missing values in features is reliable.😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3006673,
      "author_name": "handsomeevy",
      "author_url": "",
      "post_date": "10/04/2024 12:17:40",
      "content": "<p>全量样本才5000左右，训练样本3960，有目标样本2736，有额外数据的996个。<br>\n总体的样本数本身已经很小了，而无目标和不带有额外数据的样本数占比很大，意味着你不论是做丢弃和填充，都会严重的影响数据分布，进而影响真实结果。所以实际上没有特别好的方法解决，只能看这个数据集实验下来，丢弃和填充哪种效果更好</p>",
      "votes": null,
      "replies": [
        {
          "id": 3006693,
          "author_name": "ggggpeushmy",
          "author_url": "",
          "post_date": "10/04/2024 12:39:09",
          "content": "<p>Thanks. I will try both.</p>\n<p>In the rows with known target, the distribution of different sii is :<br>\n0.0    687<br>\n1.0    386<br>\n2.0    142<br>\n3.0      9<br>\nI think that it is possible the same distribution on the unknown data set.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3006851,
              "author_name": "handsomeevy",
              "author_url": "",
              "post_date": "10/04/2024 15:54:54",
              "content": "<p>很好的想法，但是有个问题是，你不论怎样对无目标的结果做预测，你的结果是作为最后的ground truth来做训练。必然无法保证这1000多个数据的ground truth预测的百分百正确，同时会给最后的那个模型带来很大的干扰。我们最终的目标还是将最后的模型调整的很好，所以对于无target的数据是很难做处理的</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3006955,
      "author_name": "alyssonamaral",
      "author_url": "",
      "post_date": "10/04/2024 18:09:54",
      "content": "<p>I am trying to use class_weight and some balancing methods such as smote. My performance is slightly better balancing the classes</p>",
      "votes": null,
      "replies": [
        {
          "id": 3007296,
          "author_name": "ggggpeushmy",
          "author_url": "",
          "post_date": "10/05/2024 07:18:31",
          "content": "<p>Thanks. I will try.</p>\n<p>I tried KNN with metric='wminkowski', and set features weight as pearson correlation with sii.<br>\nBut the performance seems not better than other methods.😭</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3005969": "In 3960 rows of train.csv, there are 1224 rows with missing sii. What a big portion!\n\n【In Competition Datasets Description】\nThe target sii is missing for a portion of the participants in the training set. You may wish to apply non-supervised learning techniques to this data. \n\n【In Shared Baseline Model】\nI found that dropping 1224 rows directly in training set can also get not bad score in public Leaderboard.\n\n【Methods I tried to fill missing sii】\n(1) Lightgbm：all the predicted sii is 0.0 ,no matter what parameter I use, I don't know why.\n(2) KNN: the distribution of predicted sii vary greatly from different n_neighbors.",
    "3006218": "You can use your normal code, split the train set into train_df and train1_df, train1_df having missing sii. then you use train_df to train train1_df (instead of test_df in normal case), merge, and go ahead with what you usually do.\n\nI tried many times, I do get a few sii= 2 or 3, but final performance drop. May be the result I feed into train1_df is rubbish. 🤣",
    "3006282": "Thanks.\n\nI only got few sii =3 when KNN n_neighbors<=3, but I think it is not a reliable method to fill the missing sii.\nBecause the features also have many missing values, I need to fill the missing values in features first, and I also not sure whether the method to fill the missing values in features is reliable.😅",
    "3006673": "全量样本才5000左右，训练样本3960，有目标样本2736，有额外数据的996个。\n总体的样本数本身已经很小了，而无目标和不带有额外数据的样本数占比很大，意味着你不论是做丢弃和填充，都会严重的影响数据分布，进而影响真实结果。所以实际上没有特别好的方法解决，只能看这个数据集实验下来，丢弃和填充哪种效果更好",
    "3006693": "Thanks. I will try both.\n\nIn the rows with known target, the distribution of different sii is :\n0.0    687\n1.0    386\n2.0    142\n3.0      9\nI think that it is possible the same distribution on the unknown data set.",
    "3006851": "很好的想法，但是有个问题是，你不论怎样对无目标的结果做预测，你的结果是作为最后的ground truth来做训练。必然无法保证这1000多个数据的ground truth预测的百分百正确，同时会给最后的那个模型带来很大的干扰。我们最终的目标还是将最后的模型调整的很好，所以对于无target的数据是很难做处理的",
    "3006955": "I am trying to use class_weight and some balancing methods such as smote. My performance is slightly better balancing the classes",
    "3007296": "Thanks. I will try.\n\nI tried KNN with metric='wminkowski', and set features weight as pearson correlation with sii.\nBut the performance seems not better than other methods.😭"
  },
  "source": "meta"
}