{
  "id": 551063,
  "title": "In using KNNImputer for missing value imputation, does increasing k lead to more leakage, or does decreasing k result in more leakage?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551063",
  "author_name": "",
  "post_date": "2024-12-11T03:50:13.078348800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In using KNNImputer for missing value imputation, does increasing k lead to more leakage, or does decreasing k result in more leakage?</p>",
  "messages": [
    {
      "id": "3069085",
      "postDate": "12/11/2024 03:50:13",
      "content": "<p>In using KNNImputer for missing value imputation, does increasing k lead to more leakage, or does decreasing k result in more leakage?</p>",
      "rawMarkdown": "In using KNNImputer for missing value imputation, does increasing k lead to more leakage, or does decreasing k result in more leakage?",
      "votes": null
    },
    {
      "id": "3069221",
      "postDate": "12/11/2024 07:29:39",
      "content": "<p>The problem is not k, the problem is to use an imputer in this competition and the way it is used : using an imputer when more than a half of values are missing in so many columns is crazy, using an imputer to impute missing values of targets is even more crazy.</p>",
      "rawMarkdown": "The problem is not k, the problem is to use an imputer in this competition and the way it is used : using an imputer when more than a half of values are missing in so many columns is crazy, using an imputer to impute missing values of targets is even more crazy.",
      "votes": null
    },
    {
      "id": "3069243",
      "postDate": "12/11/2024 08:15:31",
      "content": "<p>k has nothing to do with data leakage. It is the amount of 'closeby' datapoints the model considers for your datapoint you want to impute. For example k=3, it checks the 3 closest datapoints and averages the value of them (or whatever weight is setup) for your imputation.</p>\n<p>Dataleakage happens if you impute all the columns of your df (including the target value) and train/test split afterwards. That means you imputed the target value on your test df -&gt; data leakage</p>",
      "rawMarkdown": "k has nothing to do with data leakage. It is the amount of 'closeby' datapoints the model considers for your datapoint you want to impute. For example k=3, it checks the 3 closest datapoints and averages the value of them (or whatever weight is setup) for your imputation.\n\nDataleakage happens if you impute all the columns of your df (including the target value) and train/test split afterwards. That means you imputed the target value on your test df -> data leakage",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3069221,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/11/2024 07:29:39",
      "content": "<p>The problem is not k, the problem is to use an imputer in this competition and the way it is used : using an imputer when more than a half of values are missing in so many columns is crazy, using an imputer to impute missing values of targets is even more crazy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3069243,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "12/11/2024 08:15:31",
      "content": "<p>k has nothing to do with data leakage. It is the amount of 'closeby' datapoints the model considers for your datapoint you want to impute. For example k=3, it checks the 3 closest datapoints and averages the value of them (or whatever weight is setup) for your imputation.</p>\n<p>Dataleakage happens if you impute all the columns of your df (including the target value) and train/test split afterwards. That means you imputed the target value on your test df -&gt; data leakage</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3069085": "In using KNNImputer for missing value imputation, does increasing k lead to more leakage, or does decreasing k result in more leakage?",
    "3069221": "The problem is not k, the problem is to use an imputer in this competition and the way it is used : using an imputer when more than a half of values are missing in so many columns is crazy, using an imputer to impute missing values of targets is even more crazy.",
    "3069243": "k has nothing to do with data leakage. It is the amount of 'closeby' datapoints the model considers for your datapoint you want to impute. For example k=3, it checks the 3 closest datapoints and averages the value of them (or whatever weight is setup) for your imputation.\n\nDataleakage happens if you impute all the columns of your df (including the target value) and train/test split afterwards. That means you imputed the target value on your test df -> data leakage"
  },
  "source": "meta"
}