{
  "id": 388682,
  "title": "Something strange about CV and LB",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388682",
  "author_name": "",
  "post_date": "2023-02-19T01:07:18.561762600Z",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>For each group_level part, I use corresponding data to make feature engineering  and train models(without historical and future informations or any encoding). What I am confused is that lesser features I use, higher CV it is, but lower LB score as result. Some examples are below:</p>\n<table>\n<thead>\n<tr>\n<th>origin features(each level)</th>\n<th>CV</th>\n<th>most important features(each level)</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>300,500,600</td>\n<td>0.6933</td>\n<td>not reduced</td>\n<td></td>\n<td>0.693</td>\n</tr>\n<tr>\n<td>300,500,600</td>\n<td>0.6935</td>\n<td>200,300,400</td>\n<td>0.6943</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>500,800,1000</td>\n<td>0.6935</td>\n<td>100,110,120</td>\n<td>0.6958</td>\n<td>0.691</td>\n</tr>\n</tbody>\n</table>\n<p>How could this be?</p>",
  "messages": [
    {
      "id": "2150121",
      "postDate": "02/19/2023 01:07:18",
      "content": "<p>For each group_level part, I use corresponding data to make feature engineering  and train models(without historical and future informations or any encoding). What I am confused is that lesser features I use, higher CV it is, but lower LB score as result. Some examples are below:</p>\n<table>\n<thead>\n<tr>\n<th>origin features(each level)</th>\n<th>CV</th>\n<th>most important features(each level)</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>300,500,600</td>\n<td>0.6933</td>\n<td>not reduced</td>\n<td></td>\n<td>0.693</td>\n</tr>\n<tr>\n<td>300,500,600</td>\n<td>0.6935</td>\n<td>200,300,400</td>\n<td>0.6943</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>500,800,1000</td>\n<td>0.6935</td>\n<td>100,110,120</td>\n<td>0.6958</td>\n<td>0.691</td>\n</tr>\n</tbody>\n</table>\n<p>How could this be?</p>",
      "rawMarkdown": "For each group_level part, I use corresponding data to make feature engineering  and train models(without historical and future informations or any encoding). What I am confused is that lesser features I use, higher CV it is, but lower LB score as result. Some examples are below:\n| origin features(each level) | CV | most important features(each level) | CV |  LB |\n| --- | --- | --- | --- | --- |\n| 300,500,600 | 0.6933 | not reduced |  | 0.693 |\n| 300,500,600 | 0.6935 | 200,300,400 | 0.6943 | 0.692 |\n| 500,800,1000 | 0.6935 | 100,110,120 | 0.6958 | 0.691 |\n\nHow could this be?",
      "votes": null
    },
    {
      "id": "2150127",
      "postDate": "02/19/2023 01:10:34",
      "content": "<p>It could be a CV leak. How do you choose your features? If you use information from all folds to select features for fold 1, then you are leaking the OOF targets of fold 1 by using feature selection information from folds 2,3,4,5</p>",
      "rawMarkdown": "It could be a CV leak. How do you choose your features? If you use information from all folds to select features for fold 1, then you are leaking the OOF targets of fold 1 by using feature selection information from folds 2,3,4,5",
      "votes": null
    },
    {
      "id": "2150134",
      "postDate": "02/19/2023 01:16:36",
      "content": "<p>Hi Chris, I don't create features about targets, only about something like the elapsed times and index numbers between two events.</p>",
      "rawMarkdown": "Hi Chris, I don't create features about targets, only about something like the elapsed times and index numbers between two events.",
      "votes": null
    },
    {
      "id": "2150140",
      "postDate": "02/19/2023 01:20:13",
      "content": "<p>Yes, but how do you pick the top 100 features?</p>",
      "rawMarkdown": "Yes, but how do you pick the top 100 features?",
      "votes": null
    },
    {
      "id": "2150141",
      "postDate": "02/19/2023 01:20:35",
      "content": "<p>Do you use feature importance from folds 2,3,4,5 to help select features for fold 1? If so, that is a leak.</p>",
      "rawMarkdown": "Do you use feature importance from folds 2,3,4,5 to help select features for fold 1? If so, that is a leak.",
      "votes": null
    },
    {
      "id": "2150148",
      "postDate": "02/19/2023 01:27:02",
      "content": "<p>Oh it is. I used the average feature importance values to pick 'top' 100 features.<br>\nin each fold:</p>\n<pre><code>fold_importance_df = pd.DataFrame()\nfold_importance_df[\"feature\"] = FEATURES\nfold_importance_df[\"importance\"] = clf.feature_importances_\nfold_importance_df[\"fold\"] = i + 1\nfeature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n</code></pre>\n<p>then:</p>\n<pre><code>feature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)\ndisplay(feature_importance_df.head(10))\n</code></pre>",
      "rawMarkdown": "Oh it is. I used the average feature importance values to pick 'top' 100 features.\nin each fold:\n```\nfold_importance_df = pd.DataFrame()\nfold_importance_df[\"feature\"] = FEATURES\nfold_importance_df[\"importance\"] = clf.feature_importances_\nfold_importance_df[\"fold\"] = i + 1\nfeature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n```\nthen:\n```\nfeature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)\ndisplay(feature_importance_df.head(10))\n```",
      "votes": null
    },
    {
      "id": "2150149",
      "postDate": "02/19/2023 01:27:55",
      "content": "<p>Feature selection has lots of opportunities for introducing leaks. For example if we do permutation importance, that is using the targets to pick features. So it results in an optimistic CV score. With all methods we must be careful that we aren't using too much information from targets.</p>",
      "rawMarkdown": "Feature selection has lots of opportunities for introducing leaks. For example if we do permutation importance, that is using the targets to pick features. So it results in an optimistic CV score. With all methods we must be careful that we aren't using too much information from targets.",
      "votes": null
    },
    {
      "id": "2180935",
      "postDate": "03/14/2023 07:35:14",
      "content": "<p>I faced the same situation. Thanks so much, Chris and Joseph!! I learn so much from your discussion.</p>",
      "rawMarkdown": "I faced the same situation. Thanks so much, Chris and Joseph!! I learn so much from your discussion.",
      "votes": null
    },
    {
      "id": "2193197",
      "postDate": "03/23/2023 06:14:36",
      "content": "<p>hi Chris, if the above method leads to cv leakage as you said, then how to select features in cv that will not lead to cv leakage, or is there any other better method? Thank you very much for your reply</p>",
      "rawMarkdown": "hi Chris, if the above method leads to cv leakage as you said, then how to select features in cv that will not lead to cv leakage, or is there any other better method? Thank you very much for your reply",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2150127,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/19/2023 01:10:34",
      "content": "<p>It could be a CV leak. How do you choose your features? If you use information from all folds to select features for fold 1, then you are leaking the OOF targets of fold 1 by using feature selection information from folds 2,3,4,5</p>",
      "votes": null,
      "replies": [
        {
          "id": 2150134,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "02/19/2023 01:16:36",
          "content": "<p>Hi Chris, I don't create features about targets, only about something like the elapsed times and index numbers between two events.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2150140,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/19/2023 01:20:13",
              "content": "<p>Yes, but how do you pick the top 100 features?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2150141,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/19/2023 01:20:35",
              "content": "<p>Do you use feature importance from folds 2,3,4,5 to help select features for fold 1? If so, that is a leak.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2150148,
                  "author_name": "takanashihumbert",
                  "author_url": "",
                  "post_date": "02/19/2023 01:27:02",
                  "content": "<p>Oh it is. I used the average feature importance values to pick 'top' 100 features.<br>\nin each fold:</p>\n<pre><code>fold_importance_df = pd.DataFrame()\nfold_importance_df[\"feature\"] = FEATURES\nfold_importance_df[\"importance\"] = clf.feature_importances_\nfold_importance_df[\"fold\"] = i + 1\nfeature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n</code></pre>\n<p>then:</p>\n<pre><code>feature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)\ndisplay(feature_importance_df.head(10))\n</code></pre>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2150149,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/19/2023 01:27:55",
              "content": "<p>Feature selection has lots of opportunities for introducing leaks. For example if we do permutation importance, that is using the targets to pick features. So it results in an optimistic CV score. With all methods we must be careful that we aren't using too much information from targets.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2193197,
          "author_name": "shiyusha",
          "author_url": "",
          "post_date": "03/23/2023 06:14:36",
          "content": "<p>hi Chris, if the above method leads to cv leakage as you said, then how to select features in cv that will not lead to cv leakage, or is there any other better method? Thank you very much for your reply</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2180935,
      "author_name": "mengvision",
      "author_url": "",
      "post_date": "03/14/2023 07:35:14",
      "content": "<p>I faced the same situation. Thanks so much, Chris and Joseph!! I learn so much from your discussion.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2150121": "For each group_level part, I use corresponding data to make feature engineering  and train models(without historical and future informations or any encoding). What I am confused is that lesser features I use, higher CV it is, but lower LB score as result. Some examples are below:\n| origin features(each level) | CV | most important features(each level) | CV |  LB |\n| --- | --- | --- | --- | --- |\n| 300,500,600 | 0.6933 | not reduced |  | 0.693 |\n| 300,500,600 | 0.6935 | 200,300,400 | 0.6943 | 0.692 |\n| 500,800,1000 | 0.6935 | 100,110,120 | 0.6958 | 0.691 |\n\nHow could this be?",
    "2150127": "It could be a CV leak. How do you choose your features? If you use information from all folds to select features for fold 1, then you are leaking the OOF targets of fold 1 by using feature selection information from folds 2,3,4,5",
    "2150134": "Hi Chris, I don't create features about targets, only about something like the elapsed times and index numbers between two events.",
    "2150140": "Yes, but how do you pick the top 100 features?",
    "2150141": "Do you use feature importance from folds 2,3,4,5 to help select features for fold 1? If so, that is a leak.",
    "2150148": "Oh it is. I used the average feature importance values to pick 'top' 100 features.\nin each fold:\n```\nfold_importance_df = pd.DataFrame()\nfold_importance_df[\"feature\"] = FEATURES\nfold_importance_df[\"importance\"] = clf.feature_importances_\nfold_importance_df[\"fold\"] = i + 1\nfeature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n```\nthen:\n```\nfeature_importance_df = feature_importance_df.groupby(['feature'])['importance'].agg(['mean']).sort_values(by='mean', ascending=False)\ndisplay(feature_importance_df.head(10))\n```",
    "2150149": "Feature selection has lots of opportunities for introducing leaks. For example if we do permutation importance, that is using the targets to pick features. So it results in an optimistic CV score. With all methods we must be careful that we aren't using too much information from targets.",
    "2180935": "I faced the same situation. Thanks so much, Chris and Joseph!! I learn so much from your discussion.",
    "2193197": "hi Chris, if the above method leads to cv leakage as you said, then how to select features in cv that will not lead to cv leakage, or is there any other better method? Thank you very much for your reply"
  },
  "source": "meta"
}