{
  "id": 396872,
  "title": "How to ensure that cv does not leak？",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/396872",
  "author_name": "",
  "post_date": "2023-03-23T07:13:12.319941800Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>i stuck on a problem about cv leak.I first created a lot of features, then increased my cv score from 0.65(all features) to 0.67 by cv pick features (adding and ranking the importance of each feature under each problem), but my LB score was only 0.62, which I realized was causing a cv leak. Now I am very confused about how to filter the features without leading to leakage. For example, if I filter the features with the filter method on all the training sets and then go to cv, does it also lead to information leakage? So how do we do feature select?</p>",
  "messages": [
    {
      "id": "2193254",
      "postDate": "03/23/2023 07:13:12",
      "content": "<p>i stuck on a problem about cv leak.I first created a lot of features, then increased my cv score from 0.65(all features) to 0.67 by cv pick features (adding and ranking the importance of each feature under each problem), but my LB score was only 0.62, which I realized was causing a cv leak. Now I am very confused about how to filter the features without leading to leakage. For example, if I filter the features with the filter method on all the training sets and then go to cv, does it also lead to information leakage? So how do we do feature select?</p>",
      "rawMarkdown": "i stuck on a problem about cv leak.I first created a lot of features, then increased my cv score from 0.65(all features) to 0.67 by cv pick features (adding and ranking the importance of each feature under each problem), but my LB score was only 0.62, which I realized was causing a cv leak. Now I am very confused about how to filter the features without leading to leakage. For example, if I filter the features with the filter method on all the training sets and then go to cv, does it also lead to information leakage? So how do we do feature select?",
      "votes": null
    },
    {
      "id": "2193268",
      "postDate": "03/23/2023 07:20:52",
      "content": "<p>use <code>groupby(['session_id', 'level_group']).xxxx</code> to create features.  I guess this would be no data leakage.</p>",
      "rawMarkdown": "use ```groupby(['session_id', 'level_group']).xxxx``` to create features.  I guess this would be no data leakage.",
      "votes": null
    },
    {
      "id": "2194734",
      "postDate": "03/24/2023 05:55:35",
      "content": "<p>Hi, I think his problem lies in the feature selection process based on feature importance. I also faced the same problem.</p>\n<p>My understanding is that the calculation of feature importance involves test data information. If you do a 5-fold cross validation and choose those features that are more important, those chosen features better fits the data you use, which causes overfitting.</p>\n<p>So, could you suggest any feasible way for this feature selection process? Thanks in advance.</p>",
      "rawMarkdown": "Hi, I think his problem lies in the feature selection process based on feature importance. I also faced the same problem.\n\nMy understanding is that the calculation of feature importance involves test data information. If you do a 5-fold cross validation and choose those features that are more important, those chosen features better fits the data you use, which causes overfitting.\n\nSo, could you suggest any feasible way for this feature selection process? Thanks in advance.",
      "votes": null
    },
    {
      "id": "2194749",
      "postDate": "03/24/2023 06:14:01",
      "content": "<p>I also faced this problem. I tried some feature selection methods, but it always give me better cv but worse Lb. So I didn't do feature selection now.</p>",
      "rawMarkdown": "I also faced this problem. I tried some feature selection methods, but it always give me better cv but worse Lb. So I didn't do feature selection now.",
      "votes": null
    },
    {
      "id": "2194754",
      "postDate": "03/24/2023 06:19:16",
      "content": "<p>Thank you for your immediate reply. 加油:)</p>",
      "rawMarkdown": "Thank you for your immediate reply. 加油:)",
      "votes": null
    },
    {
      "id": "2194763",
      "postDate": "03/24/2023 06:24:17",
      "content": "<p>Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like <a href=\"https://www.kaggle.com/code/dansbecker/permutation-importance\" target=\"_blank\">permutation importance</a>. I have not tried it yet.</p>",
      "rawMarkdown": "Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like [permutation importance](https://www.kaggle.com/code/dansbecker/permutation-importance). I have not tried it yet.",
      "votes": null
    },
    {
      "id": "2195907",
      "postDate": "03/25/2023 01:17:19",
      "content": "<p><a href=\"https://www.kaggle.com/yunqicao\" target=\"_blank\">@yunqicao</a> <a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a>  Thank you for your discussion,I still have questions, if you don't mind. If you don't do feature selection, what is the approximate number of features you have to maintain?</p>",
      "rawMarkdown": "yunqicao @hookman  Thank you for your discussion,I still have questions, if you don't mind. If you don't do feature selection, what is the approximate number of features you have to maintain?",
      "votes": null
    },
    {
      "id": "2195956",
      "postDate": "03/25/2023 02:16:40",
      "content": "<p>About 1000 to 3000 features. Actually, I'm working on it, trying to reduce the number of features, but still make little progress.</p>",
      "rawMarkdown": "About 1000 to 3000 features. Actually, I'm working on it, trying to reduce the number of features, but still make little progress.",
      "votes": null
    },
    {
      "id": "2196073",
      "postDate": "03/25/2023 06:05:00",
      "content": "<p>about 1500.</p>",
      "rawMarkdown": "about 1500.",
      "votes": null
    },
    {
      "id": "2196314",
      "postDate": "03/25/2023 09:54:16",
      "content": "<blockquote>\n  <p>Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like <a href=\"https://www.kaggle.com/code/dansbecker/permutation-importance\" target=\"_blank\">permutation importance</a>. I have not tried it yet.</p>\n</blockquote>\n<p>interesting, can I ask you why you recommend methods which don't involve feature importance of gradient boost tree? <a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a> </p>",
      "rawMarkdown": "> Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like [permutation importance](https://www.kaggle.com/code/dansbecker/permutation-importance). I have not tried it yet.\n\ninteresting, can I ask you why you recommend methods which don't involve feature importance of gradient boost tree? @hookman",
      "votes": null
    },
    {
      "id": "2196574",
      "postDate": "03/25/2023 14:00:59",
      "content": "<p>Sure. maybe you can see this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">discussion</a> here. Use the mean of 5 folds feature importance could be a kind of leakage and it give better cv but worse lb. So I guess we can use some methods which don't involve feature importance of gradient boost tree to avoid this leakage.</p>",
      "rawMarkdown": "Sure. maybe you can see this [discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682) here. Use the mean of 5 folds feature importance could be a kind of leakage and it give better cv but worse lb. So I guess we can use some methods which don't involve feature importance of gradient boost tree to avoid this leakage.",
      "votes": null
    },
    {
      "id": "2198384",
      "postDate": "03/27/2023 00:55:43",
      "content": "<p>Thanks! I will read it :)</p>",
      "rawMarkdown": "Thanks! I will read it :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2193268,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "03/23/2023 07:20:52",
      "content": "<p>use <code>groupby(['session_id', 'level_group']).xxxx</code> to create features.  I guess this would be no data leakage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2194734,
          "author_name": "yunqicao",
          "author_url": "",
          "post_date": "03/24/2023 05:55:35",
          "content": "<p>Hi, I think his problem lies in the feature selection process based on feature importance. I also faced the same problem.</p>\n<p>My understanding is that the calculation of feature importance involves test data information. If you do a 5-fold cross validation and choose those features that are more important, those chosen features better fits the data you use, which causes overfitting.</p>\n<p>So, could you suggest any feasible way for this feature selection process? Thanks in advance.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2194749,
              "author_name": "hookman",
              "author_url": "",
              "post_date": "03/24/2023 06:14:01",
              "content": "<p>I also faced this problem. I tried some feature selection methods, but it always give me better cv but worse Lb. So I didn't do feature selection now.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2194754,
                  "author_name": "yunqicao",
                  "author_url": "",
                  "post_date": "03/24/2023 06:19:16",
                  "content": "<p>Thank you for your immediate reply. 加油:)</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2194763,
                      "author_name": "hookman",
                      "author_url": "",
                      "post_date": "03/24/2023 06:24:17",
                      "content": "<p>Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like <a href=\"https://www.kaggle.com/code/dansbecker/permutation-importance\" target=\"_blank\">permutation importance</a>. I have not tried it yet.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2195907,
                          "author_name": "shiyusha",
                          "author_url": "",
                          "post_date": "03/25/2023 01:17:19",
                          "content": "<p><a href=\"https://www.kaggle.com/yunqicao\" target=\"_blank\">@yunqicao</a> <a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a>  Thank you for your discussion,I still have questions, if you don't mind. If you don't do feature selection, what is the approximate number of features you have to maintain?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2195956,
                              "author_name": "yunqicao",
                              "author_url": "",
                              "post_date": "03/25/2023 02:16:40",
                              "content": "<p>About 1000 to 3000 features. Actually, I'm working on it, trying to reduce the number of features, but still make little progress.</p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 2196073,
                              "author_name": "hookman",
                              "author_url": "",
                              "post_date": "03/25/2023 06:05:00",
                              "content": "<p>about 1500.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        },
                        {
                          "id": 2196314,
                          "author_name": "econjt",
                          "author_url": "",
                          "post_date": "03/25/2023 09:54:16",
                          "content": "<blockquote>\n  <p>Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like <a href=\"https://www.kaggle.com/code/dansbecker/permutation-importance\" target=\"_blank\">permutation importance</a>. I have not tried it yet.</p>\n</blockquote>\n<p>interesting, can I ask you why you recommend methods which don't involve feature importance of gradient boost tree? <a href=\"https://www.kaggle.com/hookman\" target=\"_blank\">@hookman</a> </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2196574,
                              "author_name": "hookman",
                              "author_url": "",
                              "post_date": "03/25/2023 14:00:59",
                              "content": "<p>Sure. maybe you can see this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">discussion</a> here. Use the mean of 5 folds feature importance could be a kind of leakage and it give better cv but worse lb. So I guess we can use some methods which don't involve feature importance of gradient boost tree to avoid this leakage.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2198384,
                                  "author_name": "econjt",
                                  "author_url": "",
                                  "post_date": "03/27/2023 00:55:43",
                                  "content": "<p>Thanks! I will read it :)</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2193254": "i stuck on a problem about cv leak.I first created a lot of features, then increased my cv score from 0.65(all features) to 0.67 by cv pick features (adding and ranking the importance of each feature under each problem), but my LB score was only 0.62, which I realized was causing a cv leak. Now I am very confused about how to filter the features without leading to leakage. For example, if I filter the features with the filter method on all the training sets and then go to cv, does it also lead to information leakage? So how do we do feature select?",
    "2193268": "use ```groupby(['session_id', 'level_group']).xxxx``` to create features.  I guess this would be no data leakage.",
    "2194734": "Hi, I think his problem lies in the feature selection process based on feature importance. I also faced the same problem.\n\nMy understanding is that the calculation of feature importance involves test data information. If you do a 5-fold cross validation and choose those features that are more important, those chosen features better fits the data you use, which causes overfitting.\n\nSo, could you suggest any feasible way for this feature selection process? Thanks in advance.",
    "2194749": "I also faced this problem. I tried some feature selection methods, but it always give me better cv but worse Lb. So I didn't do feature selection now.",
    "2194754": "Thank you for your immediate reply. 加油:)",
    "2194763": "Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like [permutation importance](https://www.kaggle.com/code/dansbecker/permutation-importance). I have not tried it yet.",
    "2195907": "yunqicao @hookman  Thank you for your discussion,I still have questions, if you don't mind. If you don't do feature selection, what is the approximate number of features you have to maintain?",
    "2195956": "About 1000 to 3000 features. Actually, I'm working on it, trying to reduce the number of features, but still make little progress.",
    "2196073": "about 1500.",
    "2196314": "> Thanks. Maybe we can try some method which don't involve feature importance of gradient boost tree, like [permutation importance](https://www.kaggle.com/code/dansbecker/permutation-importance). I have not tried it yet.\n\ninteresting, can I ask you why you recommend methods which don't involve feature importance of gradient boost tree? @hookman",
    "2196574": "Sure. maybe you can see this [discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682) here. Use the mean of 5 folds feature importance could be a kind of leakage and it give better cv but worse lb. So I guess we can use some methods which don't involve feature importance of gradient boost tree to avoid this leakage.",
    "2198384": "Thanks! I will read it :)"
  },
  "source": "meta"
}