{
  "id": 539536,
  "title": "Ideas for samples with missing 'sii'",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/539536",
  "author_name": "",
  "post_date": "2024-10-09T12:24:26.549696400Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I was asking myself what's the point of all the samples with missing 'sii' in the training set? Do I miss something or did you find any use of them? The host suggest to apply non-supervised learning on them, but I don't see the point as we can predict on them with supervised model and even so these will not be as 'ground truth' to take them into training set further. Your thoughts champions?</p>",
  "messages": [
    {
      "id": "3012856",
      "postDate": "10/09/2024 12:24:26",
      "content": "<p>I was asking myself what's the point of all the samples with missing 'sii' in the training set? Do I miss something or did you find any use of them? The host suggest to apply non-supervised learning on them, but I don't see the point as we can predict on them with supervised model and even so these will not be as 'ground truth' to take them into training set further. Your thoughts champions?</p>",
      "rawMarkdown": "I was asking myself what's the point of all the samples with missing 'sii' in the training set? Do I miss something or did you find any use of them? The host suggest to apply non-supervised learning on them, but I don't see the point as we can predict on them with supervised model and even so these will not be as 'ground truth' to take them into training set further. Your thoughts champions?",
      "votes": null
    },
    {
      "id": "3012994",
      "postDate": "10/09/2024 14:58:56",
      "content": "<p>This <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535023\" target=\"_blank\">discussion topic</a> touches on your question. I hope it helps you.</p>",
      "rawMarkdown": "This [discussion topic](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535023) touches on your question. I hope it helps you.",
      "votes": null
    },
    {
      "id": "3012998",
      "postDate": "10/09/2024 15:04:38",
      "content": "<p>In that post he proposes what I mentioned in my post - that we can predict them, but this is not ground truth and training on them is just introducing noise to the actual true data. We can use SMOTE to generate more data if needed, but to train on samples with predicted targets doesn't sound very promising.</p>",
      "rawMarkdown": "In that post he proposes what I mentioned in my post - that we can predict them, but this is not ground truth and training on them is just introducing noise to the actual true data. We can use SMOTE to generate more data if needed, but to train on samples with predicted targets doesn't sound very promising.",
      "votes": null
    },
    {
      "id": "3013014",
      "postDate": "10/09/2024 15:18:37",
      "content": "<p>His post comes down to generating pseudo labels and that by itself increases data availability. And this just might improve the model performance (there are no guaranties of course). And yes this is no ground truth.</p>",
      "rawMarkdown": "His post comes down to generating pseudo labels and that by itself increases data availability. And this just might improve the model performance (there are no guaranties of course). And yes this is no ground truth.",
      "votes": null
    },
    {
      "id": "3013039",
      "postDate": "10/09/2024 15:49:11",
      "content": "<p>filling in the sii into the missed cells will not help. It's just noise. I tried it many times 🤣.<br>\nThose data is offering something, for example, average BMI in the larger population, but they're not critical.<br>\nI just read that IBM link in that discussion topic, worth reading.</p>",
      "rawMarkdown": "filling in the sii into the missed cells will not help. It's just noise. I tried it many times 🤣.\nThose data is offering something, for example, average BMI in the larger population, but they're not critical.\nI just read that IBM link in that discussion topic, worth reading.",
      "votes": null
    },
    {
      "id": "3014873",
      "postDate": "10/11/2024 17:05:52",
      "content": "<p>Good point - we can use them to calculate a broader range of values in case we impute missing values or other type of analysis</p>",
      "rawMarkdown": "Good point - we can use them to calculate a broader range of values in case we impute missing values or other type of analysis",
      "votes": null
    },
    {
      "id": "3016021",
      "postDate": "10/13/2024 09:11:51",
      "content": "<p>While it may seem like there are some gaps, there’s actually a lot of potential here! You can get creative, especially when you have time-series data for the missing SII datapoints. Plus, using broader datapoints can help fill in those gaps. There’s plenty of opportunity to make this work!\"</p>",
      "rawMarkdown": "While it may seem like there are some gaps, there’s actually a lot of potential here! You can get creative, especially when you have time-series data for the missing SII datapoints. Plus, using broader datapoints can help fill in those gaps. There’s plenty of opportunity to make this work!\"",
      "votes": null
    },
    {
      "id": "3016064",
      "postDate": "10/13/2024 09:50:11",
      "content": "<p>As far as I checked there are no missing 'sii' for the available id series (if we drop that 2-3 samples with several missing PCIAT). But about using broader datapoints for imputation I agree that these can be useful.</p>",
      "rawMarkdown": "As far as I checked there are no missing 'sii' for the available id series (if we drop that 2-3 samples with several missing PCIAT). But about using broader datapoints for imputation I agree that these can be useful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3012994,
      "author_name": "wti200",
      "author_url": "",
      "post_date": "10/09/2024 14:58:56",
      "content": "<p>This <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535023\" target=\"_blank\">discussion topic</a> touches on your question. I hope it helps you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3012998,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/09/2024 15:04:38",
          "content": "<p>In that post he proposes what I mentioned in my post - that we can predict them, but this is not ground truth and training on them is just introducing noise to the actual true data. We can use SMOTE to generate more data if needed, but to train on samples with predicted targets doesn't sound very promising.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3013014,
              "author_name": "wti200",
              "author_url": "",
              "post_date": "10/09/2024 15:18:37",
              "content": "<p>His post comes down to generating pseudo labels and that by itself increases data availability. And this just might improve the model performance (there are no guaranties of course). And yes this is no ground truth.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3013039,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/09/2024 15:49:11",
      "content": "<p>filling in the sii into the missed cells will not help. It's just noise. I tried it many times 🤣.<br>\nThose data is offering something, for example, average BMI in the larger population, but they're not critical.<br>\nI just read that IBM link in that discussion topic, worth reading.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3014873,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/11/2024 17:05:52",
          "content": "<p>Good point - we can use them to calculate a broader range of values in case we impute missing values or other type of analysis</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3016021,
      "author_name": "chanpreetsingh07",
      "author_url": "",
      "post_date": "10/13/2024 09:11:51",
      "content": "<p>While it may seem like there are some gaps, there’s actually a lot of potential here! You can get creative, especially when you have time-series data for the missing SII datapoints. Plus, using broader datapoints can help fill in those gaps. There’s plenty of opportunity to make this work!\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 3016064,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/13/2024 09:50:11",
          "content": "<p>As far as I checked there are no missing 'sii' for the available id series (if we drop that 2-3 samples with several missing PCIAT). But about using broader datapoints for imputation I agree that these can be useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3012856": "I was asking myself what's the point of all the samples with missing 'sii' in the training set? Do I miss something or did you find any use of them? The host suggest to apply non-supervised learning on them, but I don't see the point as we can predict on them with supervised model and even so these will not be as 'ground truth' to take them into training set further. Your thoughts champions?",
    "3012994": "This [discussion topic](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535023) touches on your question. I hope it helps you.",
    "3012998": "In that post he proposes what I mentioned in my post - that we can predict them, but this is not ground truth and training on them is just introducing noise to the actual true data. We can use SMOTE to generate more data if needed, but to train on samples with predicted targets doesn't sound very promising.",
    "3013014": "His post comes down to generating pseudo labels and that by itself increases data availability. And this just might improve the model performance (there are no guaranties of course). And yes this is no ground truth.",
    "3013039": "filling in the sii into the missed cells will not help. It's just noise. I tried it many times 🤣.\nThose data is offering something, for example, average BMI in the larger population, but they're not critical.\nI just read that IBM link in that discussion topic, worth reading.",
    "3014873": "Good point - we can use them to calculate a broader range of values in case we impute missing values or other type of analysis",
    "3016021": "While it may seem like there are some gaps, there’s actually a lot of potential here! You can get creative, especially when you have time-series data for the missing SII datapoints. Plus, using broader datapoints can help fill in those gaps. There’s plenty of opportunity to make this work!\"",
    "3016064": "As far as I checked there are no missing 'sii' for the available id series (if we drop that 2-3 samples with several missing PCIAT). But about using broader datapoints for imputation I agree that these can be useful."
  },
  "source": "meta"
}