{
  "id": 200690,
  "title": "What to do first when Doing K Fold Cv , Folds creation or Feature Engineering ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/200690",
  "author_name": "",
  "post_date": "2020-12-01T12:58:42.049718300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , would be glad if you and other Kagglers would help me clarify this . Suppose in this competition we are doing stratfication of users based on their occurences in train data . Now what we do first . </p>\n<ol>\n<li><p>Do we first do feature engineering based on content_id and user_id mean features and after that do fold creation . But doing so will cause leakage across validation sets . </p></li>\n<li><p>We first create folds based on some stratification strategy and then for each fold we calculate this mean features Separately .</p></li>\n</ol>\n<p>so is method 2 a Good choice ?</p>",
  "messages": [
    {
      "id": "1098121",
      "postDate": "12/01/2020 12:58:42",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , would be glad if you and other Kagglers would help me clarify this . Suppose in this competition we are doing stratfication of users based on their occurences in train data . Now what we do first . </p>\n<ol>\n<li><p>Do we first do feature engineering based on content_id and user_id mean features and after that do fold creation . But doing so will cause leakage across validation sets . </p></li>\n<li><p>We first create folds based on some stratification strategy and then for each fold we calculate this mean features Separately .</p></li>\n</ol>\n<p>so is method 2 a Good choice ?</p>",
      "rawMarkdown": "rohanrao @cdeotte , would be glad if you and other Kagglers would help me clarify this . Suppose in this competition we are doing stratfication of users based on their occurences in train data . Now what we do first . \n1. Do we first do feature engineering based on content_id and user_id mean features and after that do fold creation . But doing so will cause leakage across validation sets . \n\n2. We first create folds based on some stratification strategy and then for each fold we calculate this mean features Separately .\n\nso is method 2 a Good choice ?",
      "votes": null
    },
    {
      "id": "1098287",
      "postDate": "12/01/2020 14:48:27",
      "content": "<p>It's method 2. If you have k folds and you're computing features on fold i, remove temporarily the targets of fold i, then compute the features and you'll be sure that you're not leaking the target in the features for fold i.</p>",
      "rawMarkdown": "It's method 2. If you have k folds and you're computing features on fold i, remove temporarily the targets of fold i, then compute the features and you'll be sure that you're not leaking the target in the features for fold i.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1098287,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "12/01/2020 14:48:27",
      "content": "<p>It's method 2. If you have k folds and you're computing features on fold i, remove temporarily the targets of fold i, then compute the features and you'll be sure that you're not leaking the target in the features for fold i.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1098121": "rohanrao @cdeotte , would be glad if you and other Kagglers would help me clarify this . Suppose in this competition we are doing stratfication of users based on their occurences in train data . Now what we do first . \n1. Do we first do feature engineering based on content_id and user_id mean features and after that do fold creation . But doing so will cause leakage across validation sets . \n\n2. We first create folds based on some stratification strategy and then for each fold we calculate this mean features Separately .\n\nso is method 2 a Good choice ?",
    "1098287": "It's method 2. If you have k folds and you're computing features on fold i, remove temporarily the targets of fold i, then compute the features and you'll be sure that you're not leaking the target in the features for fold i."
  },
  "source": "meta"
}