{
  "id": 542246,
  "title": "Questions about Actigraphy Data Usage and Potential Data Leakage from Imputation in Kaggle Child Mind Institute — Problematic Internet Use Competition",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/542246",
  "author_name": "",
  "post_date": "2024-10-23T18:39:21.257059800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have a few questions regarding the Kaggle Child Mind Institute — Problematic Internet Use Competition</p>\n<pre><code>What is  actual purpose   data  parquet  (actigraphy)? Given that only  portion   observations   training  have these data,    test , there are only  observations  actigraphy data, is  possible  extract meaningful information  this?\n\nWhy does  competition description suggest imputing  missing sii data  clustering ( maybe I misinterpreted this)? Is  recommended  impute  sii data   available features even  we apply  unsupervised method (e.g., clustering patients)? Would this lead  data leakage  we plan  use supervised learning later ,    ?\n</code></pre>",
  "messages": [
    {
      "id": "3026396",
      "postDate": "10/23/2024 18:39:21",
      "content": "<p>I have a few questions regarding the Kaggle Child Mind Institute — Problematic Internet Use Competition</p>\n<pre><code>What is  actual purpose   data  parquet  (actigraphy)? Given that only  portion   observations   training  have these data,    test , there are only  observations  actigraphy data, is  possible  extract meaningful information  this?\n\nWhy does  competition description suggest imputing  missing sii data  clustering ( maybe I misinterpreted this)? Is  recommended  impute  sii data   available features even  we apply  unsupervised method (e.g., clustering patients)? Would this lead  data leakage  we plan  use supervised learning later ,    ?\n</code></pre>",
      "rawMarkdown": "I have a few questions regarding the Kaggle Child Mind Institute — Problematic Internet Use Competition\n\n    What is the actual purpose of the data in parquet format (actigraphy)? Given that only a portion of the observations in the training set have these data, and in the test set, there are only two observations with actigraphy data, is it possible to extract meaningful information from this?\n\n    Why does the competition description suggest imputing the missing sii data using clustering (or maybe I misinterpreted this)? Is it recommended to impute the sii data using the available features even if we apply an unsupervised method (e.g., clustering patients)? Would this lead to data leakage if we plan to use supervised learning later on, or am I mistaken?",
      "votes": null
    },
    {
      "id": "3029387",
      "postDate": "10/27/2024 08:38:24",
      "content": "<p>As you mentioned the actigraphy series are available just for a part of the samples and this is the case for the test set too, BUT the test set you see is not the real one, it's just a pseudo set that we can verify if our code runs right. The real test set we don't see, so there are a lot more samples in the main test set and a lot more actigraphy series (the exact number we don't know). <br>\nAbout imputing missing 'sii' - predicting them and then using them as ground truth for training is not the good practice. You can use the samples with missing 'sii' for a broader range analysis for imputing, group separation or other feature engineering techniques.</p>",
      "rawMarkdown": "As you mentioned the actigraphy series are available just for a part of the samples and this is the case for the test set too, BUT the test set you see is not the real one, it's just a pseudo set that we can verify if our code runs right. The real test set we don't see, so there are a lot more samples in the main test set and a lot more actigraphy series (the exact number we don't know). \nAbout imputing missing 'sii' - predicting them and then using them as ground truth for training is not the good practice. You can use the samples with missing 'sii' for a broader range analysis for imputing, group separation or other feature engineering techniques.",
      "votes": null
    },
    {
      "id": "3033785",
      "postDate": "11/01/2024 14:09:50",
      "content": "<p>this was useful! Thanks!</p>",
      "rawMarkdown": "this was useful! Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3029387,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "10/27/2024 08:38:24",
      "content": "<p>As you mentioned the actigraphy series are available just for a part of the samples and this is the case for the test set too, BUT the test set you see is not the real one, it's just a pseudo set that we can verify if our code runs right. The real test set we don't see, so there are a lot more samples in the main test set and a lot more actigraphy series (the exact number we don't know). <br>\nAbout imputing missing 'sii' - predicting them and then using them as ground truth for training is not the good practice. You can use the samples with missing 'sii' for a broader range analysis for imputing, group separation or other feature engineering techniques.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3033785,
          "author_name": "adriamb3",
          "author_url": "",
          "post_date": "11/01/2024 14:09:50",
          "content": "<p>this was useful! Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3026396": "I have a few questions regarding the Kaggle Child Mind Institute — Problematic Internet Use Competition\n\n    What is the actual purpose of the data in parquet format (actigraphy)? Given that only a portion of the observations in the training set have these data, and in the test set, there are only two observations with actigraphy data, is it possible to extract meaningful information from this?\n\n    Why does the competition description suggest imputing the missing sii data using clustering (or maybe I misinterpreted this)? Is it recommended to impute the sii data using the available features even if we apply an unsupervised method (e.g., clustering patients)? Would this lead to data leakage if we plan to use supervised learning later on, or am I mistaken?",
    "3029387": "As you mentioned the actigraphy series are available just for a part of the samples and this is the case for the test set too, BUT the test set you see is not the real one, it's just a pseudo set that we can verify if our code runs right. The real test set we don't see, so there are a lot more samples in the main test set and a lot more actigraphy series (the exact number we don't know). \nAbout imputing missing 'sii' - predicting them and then using them as ground truth for training is not the good practice. You can use the samples with missing 'sii' for a broader range analysis for imputing, group separation or other feature engineering techniques.",
    "3033785": "this was useful! Thanks!"
  },
  "source": "meta"
}