{
  "id": 399802,
  "title": "Classic CV LB Thread",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/399802",
  "author_name": "Gaurav Rawat",
  "post_date": "2023-04-05T15:22:25.678000",
  "votes": 20,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Hi I was curious about how the CV and LB correlates for other folks . Posting our team results best below </p>\n<ul>\n<li><code>CV 70082 LB 0.701</code></li>\n<li><code>CV 70053 LB 0.703</code></li>\n<li><code>CV 70099 LB 0.705</code></li>\n<li><code>CV 70107  LB 0.703</code></li>\n<li><code>CV 70107  LB 0.705</code> Latest submission😑</li>\n</ul>\n<p>As mentioned in some posts there is a 0.001-0.002 gap in our case . FE helps the best but Hyper-param Tuning/Threshold in our case does improve CV in some cases but we don't see same in LB which could be hyper-params overfit but would be good to know what you guys are relying on .</p>",
  "messages": [
    {
      "id": 2210723,
      "postDate": "2023-04-05T15:22:25.680Z",
      "content": "<p>Hi I was curious about how the CV and LB correlates for other folks . Posting our team results best below </p>\n<ul>\n<li><code>CV 70082 LB 0.701</code></li>\n<li><code>CV 70053 LB 0.703</code></li>\n<li><code>CV 70099 LB 0.705</code></li>\n<li><code>CV 70107  LB 0.703</code></li>\n<li><code>CV 70107  LB 0.705</code> Latest submission😑</li>\n</ul>\n<p>As mentioned in some posts there is a 0.001-0.002 gap in our case . FE helps the best but Hyper-param Tuning/Threshold in our case does improve CV in some cases but we don't see same in LB which could be hyper-params overfit but would be good to know what you guys are relying on .</p>",
      "rawMarkdown": "Hi I was curious about how the CV and LB correlates for other folks . Posting our team results best below \n\n- `CV 70082 LB 0.701`\n- `CV 70053 LB 0.703`\n- `CV 70099 LB 0.705`\n- `CV 70107  LB 0.703`\n- `CV 70107  LB 0.705` Latest submission😑\n\nAs mentioned in some posts there is a 0.001-0.002 gap in our case . FE helps the best but Hyper-param Tuning/Threshold in our case does improve CV in some cases but we don't see same in LB which could be hyper-params overfit but would be good to know what you guys are relying on .",
      "votes": 20
    },
    {
      "id": 2259488,
      "postDate": "2023-05-15T03:16:22.913Z",
      "content": "<p>CV:0.6997368774742736 LB:0.702 <br>\nCV:0.700170167261253 LB:0.703 <br>\nCV:0.7004214709822321 LB:0.705 <br>\nCV:0.7002594561184703 LB:0.702</p>\n<p>CV does not seem to correlate with LB well.</p>",
      "rawMarkdown": "CV:0.6997368774742736 LB:0.702 \nCV:0.700170167261253 LB:0.703 \nCV:0.7004214709822321 LB:0.705 \nCV:0.7002594561184703 LB:0.702\n\nCV does not seem to correlate with LB well.",
      "votes": 3,
      "replies": [
        {
          "id": 2259493,
          "postDate": "2023-05-15T03:23:58.817Z",
          "content": "<p>True observing the same very confusing</p>",
          "rawMarkdown": "True observing the same very confusing",
          "votes": 1
        },
        {
          "id": 2260253,
          "postDate": "2023-05-15T14:59:07.447Z",
          "content": "<p>Yes, i have the same problem. It seems that after cv became 0.700xx (just FE) CV-LB correlation decreased. </p>\n<p>I tried using the old test set as a validation set and found that the data is quite noisy. For example, a small change in cv (~ 0.00002-0.00005) can change the test score by 0.001-0.003.</p>",
          "rawMarkdown": "Yes, i have the same problem. It seems that after cv became 0.700xx (just FE) CV-LB correlation decreased. \n\nI tried using the old test set as a validation set and found that the data is quite noisy. For example, a small change in cv (~ 0.00002-0.00005) can change the test score by 0.001-0.003.\n\n",
          "replies": [
            {
              "id": 2260971,
              "postDate": "2023-05-16T03:01:31.147Z",
              "content": "<p>Also are you guys <a href=\"https://www.kaggle.com/myppka\" target=\"_blank\">@myppka</a>  and <a href=\"https://www.kaggle.com/yurimaeda\" target=\"_blank\">@yurimaeda</a> having issues with selecting threshold :/ , is a bigger threshold more that gives confidence and overfits less or one that gives or optimizes on CV . some thresholds seem perform better on LB than cv 😧</p>",
              "rawMarkdown": "Also are you guys @myppka  and @yurimaeda having issues with selecting threshold :/ , is a bigger threshold more that gives confidence and overfits less or one that gives or optimizes on CV . some thresholds seem perform better on LB than cv 😧",
              "votes": 1
            },
            {
              "id": 2261173,
              "postDate": "2023-05-16T06:45:31.360Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2261174,
              "postDate": "2023-05-16T06:46:06.457Z",
              "content": "<p>In my case, only at a threshold of 0.62 was it possible to achieve a difference between CV and LB greater than 0.003</p>",
              "rawMarkdown": "In my case, only at a threshold of 0.62 was it possible to achieve a difference between CV and LB greater than 0.003",
              "votes": 1
            },
            {
              "id": 2261630,
              "postDate": "2023-05-16T13:21:45.693Z",
              "content": "<p>Do you guys submit using a single model, or an ensemble of different cv folds? I haven't tried ensembling yet, but single models from different folds for me give drastically different LB scores, disregarding the validation score of that fold.</p>",
              "rawMarkdown": "Do you guys submit using a single model, or an ensemble of different cv folds? I haven't tried ensembling yet, but single models from different folds for me give drastically different LB scores, disregarding the validation score of that fold."
            },
            {
              "id": 2261657,
              "postDate": "2023-05-16T13:41:07.170Z",
              "content": "<p>Single model for now…</p>",
              "rawMarkdown": "Single model for now…",
              "votes": 1
            },
            {
              "id": 2261734,
              "postDate": "2023-05-16T14:14:11.917Z",
              "content": "<p>Also single model</p>",
              "rawMarkdown": "Also single model",
              "votes": 1
            },
            {
              "id": 2261782,
              "postDate": "2023-05-16T14:46:35.857Z",
              "content": "<p>Oh, so when you guys say the CV is something, do you mean the macro F1 of all folds or score of a single validation fold?</p>",
              "rawMarkdown": "Oh, so when you guys say the CV is something, do you mean the macro F1 of all folds or score of a single validation fold?"
            },
            {
              "id": 2261916,
              "postDate": "2023-05-16T16:13:47.183Z",
              "content": "<p>I use macro F1 of all folds.<br>\nYou can <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218#2193827\" target=\"_blank\">retrain</a> one model per question using 100% of data. I think the results should become more stable</p>",
              "rawMarkdown": "I use macro F1 of all folds.\nYou can [retrain](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218#2193827) one model per question using 100% of data. I think the results should become more stable\n",
              "votes": 1
            },
            {
              "id": 2261926,
              "postDate": "2023-05-16T16:16:50.277Z",
              "content": "<p>Thanks for the advice, I did this at the start of the competition, but at times it felt a bit unstable, so I stopped doing that, maybe it's time to explore that option again :)</p>",
              "rawMarkdown": "Thanks for the advice, I did this at the start of the competition, but at times it felt a bit unstable, so I stopped doing that, maybe it's time to explore that option again :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2211456,
      "postDate": "2023-04-06T04:04:15.090Z",
      "content": "<p>My current best result</p>\n<p>CV  69775      LB  0.700<br>\nCV  69823      LB  0.698</p>\n<p>Some thoughts about the correlation between CV and LB:</p>\n<ul>\n<li><p>Correlation became a mess as soon as I tried to do feature selection (as <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> mentioned deleting features that have 0 importance in every fold, which I believed shouldn't cause any leakage ). Maybe there is something wrong with my feature selection strategy, I have currently abandoned the idea of doing any feature selection.</p></li>\n<li><p>Correlation worked ok for most cases if you just add new features, but not all cases. My best LB was achieved by some blind attempt that added a feature that cause a 0.0002 drop in CV, but it boosted LB from 0.698 to 0.700.   </p></li>\n<li><p>Correlation worked much better for time-based features than coordinate-based features. </p></li>\n</ul>",
      "rawMarkdown": "My current best result\n\nCV  69775      LB  0.700\nCV  69823      LB  0.698\n\nSome thoughts about the correlation between CV and LB:\n\n- Correlation became a mess as soon as I tried to do feature selection (as @gauravbrills mentioned deleting features that have 0 importance in every fold, which I believed shouldn't cause any leakage ). Maybe there is something wrong with my feature selection strategy, I have currently abandoned the idea of doing any feature selection.\n\n- Correlation worked ok for most cases if you just add new features, but not all cases. My best LB was achieved by some blind attempt that added a feature that cause a 0.0002 drop in CV, but it boosted LB from 0.698 to 0.700.   \n\n- Correlation worked much better for time-based features than coordinate-based features. ",
      "votes": 3,
      "replies": [
        {
          "id": 2212015,
          "postDate": "2023-04-06T12:54:34.623Z",
          "content": "<p>For seeing <strong><em>“Correlation became a mess as soon as I tried to do feature selection”</em></strong> i remember something.</p>\n<p>Actually i have tried to do feature selecton with feat-importance.</p>\n<p>For example, i begin with 1000 featres and regularly train with xgboost, after training i drop 5(for example) least important feature.</p>\n<p>After about 100 iters , i found that only first 2-4 iter will boost CV weakly, then CV will descend quickly，so i give up to use feature selection.</p>\n<p>Is there any relation between my experiment and your mess correlation？</p>\n<p>It seems that feature selection dosen't work well on the dataset.</p>",
          "rawMarkdown": "For seeing ***“Correlation became a mess as soon as I tried to do feature selection”*** i remember something.\n\nActually i have tried to do feature selecton with feat-importance.\n\nFor example, i begin with 1000 featres and regularly train with xgboost, after training i drop 5(for example) least important feature.\n\nAfter about 100 iters , i found that only first 2-4 iter will boost CV weakly, then CV will descend quickly，so i give up to use feature selection.\n\nIs there any relation between my experiment and your mess correlation？\n\nIt seems that feature selection dosen't work well on the dataset.",
          "votes": 1,
          "replies": [
            {
              "id": 2212032,
              "postDate": "2023-04-06T13:14:22.723Z",
              "content": "<p>Sorry, i rechecked my experiment i find <strong><em>\"then CV will descend quickly\"</em></strong> is wrong, actually CV float around(-0.0005,+0.0003), it seems that with feature num reduced, it make no sense for the model.</p>",
              "rawMarkdown": "Sorry, i rechecked my experiment i find ***\"then CV will descend quickly\"*** is wrong, actually CV float around(-0.0005,+0.0003), it seems that with feature num reduced, it make no sense for the model."
            },
            {
              "id": 2212067,
              "postDate": "2023-04-06T13:58:16.057Z",
              "content": "<p>I am also finding similar issue with Feature Selection specially removing zero importance features . Need to maybe check permutation importance but will be too time consuming :| </p>",
              "rawMarkdown": "I am also finding similar issue with Feature Selection specially removing zero importance features . Need to maybe check permutation importance but will be too time consuming :| ",
              "votes": 1
            },
            {
              "id": 2214242,
              "postDate": "2023-04-08T09:15:04.620Z",
              "content": "<p>I took the same strategy as <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> mentioned, deleting features that have zero importance in every fold(to avoid leakage) at first. It showed some minimal boost like 0.0005 for CV which I considered as noise, then LB drop sharply like 0.02.</p>",
              "rawMarkdown": "I took the same strategy as @gauravbrills mentioned, deleting features that have zero importance in every fold(to avoid leakage) at first. It showed some minimal boost like 0.0005 for CV which I considered as noise, then LB drop sharply like 0.02.",
              "votes": 1
            },
            {
              "id": 2214259,
              "postDate": "2023-04-08T09:46:30.817Z",
              "content": "<p>Maybe the reduce of features caused overfit？</p>\n<p>May i ask how much proportion have you guys reduced?</p>",
              "rawMarkdown": "Maybe the reduce of features caused overfit？\n\nMay i ask how much proportion have you guys reduced?"
            },
            {
              "id": 2214337,
              "postDate": "2023-04-08T11:39:35.170Z",
              "content": "<p>It was not supposed to cause overfitting, since I only eliminated features that have 0 importance for every model in that specific level_group. There are more 0-importance features for each fold/question, but only a small portion that has 0 importance for every fold/question, like 2% of features. I did not expect a 0.02 drop for LB since 2% seems insignificant to me. </p>\n<p>Also when I tried to eliminate 4%, 6% 8%… of features, CV went back and forth but had some improvements like 0.005 overtime, LB also improved a little back but never met the optimal score with 100% features. </p>\n<p>Just realized maybe I misunderstood how <code>XGB.feature_importace</code> works, I should recheck that.</p>",
              "rawMarkdown": "It was not supposed to cause overfitting, since I only eliminated features that have 0 importance for every model in that specific level_group. There are more 0-importance features for each fold/question, but only a small portion that has 0 importance for every fold/question, like 2% of features. I did not expect a 0.02 drop for LB since 2% seems insignificant to me. \n\nAlso when I tried to eliminate 4%, 6% 8%... of features, CV went back and forth but had some improvements like 0.005 overtime, LB also improved a little back but never met the optimal score with 100% features. \n\nJust realized maybe I misunderstood how `XGB.feature_importace` works, I should recheck that."
            },
            {
              "id": 2214399,
              "postDate": "2023-04-08T12:45:53.153Z",
              "content": "<p>Eliminating features that have 0 importance should not change anything indeed. </p>\n<p>However, could the <strong>col_sample</strong> parameter (or any random sampling of features) be impacted by this ? I have also come across surprising behavior when changing the order of features in my dataframe. I was not expecting to see any change of performance and yet, the order of features seems to make a difference. </p>\n<p>I have always concluded that it showed how unstable my model was, the same way when changing the random seed drastically changes performance. </p>\n<p>On my side, I haven't reached the same performance on the LB as you guys, but I am performing feature selection based on both gain, weight and cover by taking the intersection of the best (top 40% (arbitrarily chosen)) features of each importance type. As a result, I end up with a model that has way less features and better performance ( +0.003) and lower generalization gap between train/validation. </p>",
              "rawMarkdown": "Eliminating features that have 0 importance should not change anything indeed. \n\nHowever, could the **col_sample** parameter (or any random sampling of features) be impacted by this ? I have also come across surprising behavior when changing the order of features in my dataframe. I was not expecting to see any change of performance and yet, the order of features seems to make a difference. \n\nI have always concluded that it showed how unstable my model was, the same way when changing the random seed drastically changes performance. \n\nOn my side, I haven't reached the same performance on the LB as you guys, but I am performing feature selection based on both gain, weight and cover by taking the intersection of the best (top 40% (arbitrarily chosen)) features of each importance type. As a result, I end up with a model that has way less features and better performance ( +0.003) and lower generalization gap between train/validation. ",
              "votes": 3
            },
            {
              "id": 2215002,
              "postDate": "2023-04-09T01:36:05.323Z",
              "content": "<p>A great feature selection method,may I ask if both cv and lb increased simultaneously in the end？</p>",
              "rawMarkdown": "A great feature selection method,may I ask if both cv and lb increased simultaneously in the end？"
            },
            {
              "id": 2215459,
              "postDate": "2023-04-09T10:12:12.360Z",
              "content": "<p>Yes, in my case both cv and lb increased simultaneously, but I am facing the same issue as everyone : noticing some fluctuation on the LB. At least, by reducing the number of features, it's faster to train models and try new features/ideas.</p>",
              "rawMarkdown": "Yes, in my case both cv and lb increased simultaneously, but I am facing the same issue as everyone : noticing some fluctuation on the LB. At least, by reducing the number of features, it's faster to train models and try new features/ideas.",
              "votes": 1
            },
            {
              "id": 2215869,
              "postDate": "2023-04-09T16:30:05.073Z",
              "content": "<p>Also did anyone also observe this in our case tinkering hyperparams like tree depth etc does improve cv but gives worse in lb , not sure it overfits a bit 😐</p>",
              "rawMarkdown": "Also did anyone also observe this in our case tinkering hyperparams like tree depth etc does improve cv but gives worse in lb , not sure it overfits a bit 😐"
            }
          ]
        }
      ]
    },
    {
      "id": 2258965,
      "postDate": "2023-05-14T15:23:29.080Z",
      "content": "<p>How's everyone's CV LB correlations stands after the update? <br>\nFor me CV 0.6997 - LB 0.701</p>",
      "rawMarkdown": "How's everyone's CV LB correlations stands after the update? \nFor me CV 0.6997 - LB 0.701",
      "votes": 1
    },
    {
      "id": 2217368,
      "postDate": "2023-04-10T20:15:35.233Z",
      "content": "<p>It seems the distribution of train doesn't resemble public LB's so much.</p>\n<blockquote>\n  <p>CV: 0.69952 LB: 0.702<br>\n  CV: 0.70025 LB: 0.703<br>\n  CV: 0.70033 LB: 0.704</p>\n</blockquote>",
      "rawMarkdown": "It seems the distribution of train doesn't resemble public LB's so much.\n>CV: 0.69952 LB: 0.702\n>CV: 0.70025 LB: 0.703\n>CV: 0.70033 LB: 0.704",
      "votes": 1,
      "replies": [
        {
          "id": 2217388,
          "postDate": "2023-04-10T21:08:39.557Z",
          "content": "<p>You LB is bit better seems than cv ,but correlation good . we are somewhat not getting this right now :| </p>",
          "rawMarkdown": "You LB is bit better seems than cv ,but correlation good . we are somewhat not getting this right now :| "
        }
      ]
    },
    {
      "id": 2241558,
      "postDate": "2023-05-01T15:41:58.133Z",
      "content": "<p>For me, cv is stable. </p>",
      "rawMarkdown": "For me, cv is stable. "
    },
    {
      "id": 2218062,
      "postDate": "2023-04-11T12:01:13.457Z",
      "content": "<p>My LB scores are consistently lower by 0.005-0.010 than my CV scores. This is not in line with what most of you are reporting in this and other threads. I'm not doing anything spectacular, using XGBClassifier with plenty of features, most of those are copied from public notebooks. I have no idea what I'm doing wrong, I'd appreciate if someone could help me understand this. Thanks!</p>",
      "rawMarkdown": "My LB scores are consistently lower by 0.005-0.010 than my CV scores. This is not in line with what most of you are reporting in this and other threads. I'm not doing anything spectacular, using XGBClassifier with plenty of features, most of those are copied from public notebooks. I have no idea what I'm doing wrong, I'd appreciate if someone could help me understand this. Thanks!"
    },
    {
      "id": 2216772,
      "postDate": "2023-04-10T10:37:14.797Z",
      "content": "<p>My current results:</p>\n<ul>\n<li>Best CV: CV 0.69951 LB 0.699 </li>\n<li>Best LB: CV 0.69914 LB 0.703</li>\n</ul>\n<p>After changing the test set, the correlation between CV and LB decreased(</p>",
      "rawMarkdown": "My current results:\n- Best CV: CV 0.69951 LB 0.699 \n- Best LB: CV 0.69914 LB 0.703\n\nAfter changing the test set, the correlation between CV and LB decreased(\n"
    },
    {
      "id": 2216386,
      "postDate": "2023-04-10T04:01:15.633Z",
      "content": "<p>0.628 LB 0.578</p>",
      "rawMarkdown": "0.628 LB 0.578"
    },
    {
      "id": 2216115,
      "postDate": "2023-04-09T19:35:10.880Z",
      "content": "<p>CV 0.6992 LB 0.701 but it is unstable.<br>\n There is probably 0.002 noise/overfitting as my LB drops to 0.699 with small changes that improve CV<br>\nthreshold optimization may be overfitting too?</p>",
      "rawMarkdown": "CV 0.6992 LB 0.701 but it is unstable.\n There is probably 0.002 noise/overfitting as my LB drops to 0.699 with small changes that improve CV\nthreshold optimization may be overfitting too?",
      "replies": [
        {
          "id": 2216192,
          "postDate": "2023-04-09T21:06:35.900Z",
          "content": "<p>we getting same behaviour😀 guess think still the magic is missing to cross this</p>",
          "rawMarkdown": "we getting same behaviour😀 guess think still the magic is missing to cross this"
        }
      ]
    },
    {
      "id": 2215809,
      "postDate": "2023-04-09T15:48:23.727Z",
      "content": "<p>CV: 0.6995 LB: 0.698,</p>\n<p>I wonder if I'm overfitting, seeing you guys all have higher LB scores than CV scores.</p>",
      "rawMarkdown": "CV: 0.6995 LB: 0.698,\n\nI wonder if I'm overfitting, seeing you guys all have higher LB scores than CV scores.",
      "replies": [
        {
          "id": 2216118,
          "postDate": "2023-04-09T19:37:25.157Z",
          "content": "<p>did you select the features with high importance ? that may overfit.</p>",
          "rawMarkdown": "did you select the features with high importance ? that may overfit.",
          "replies": [
            {
              "id": 2216153,
              "postDate": "2023-04-09T20:17:30.927Z",
              "content": "<p>No, the only place I can think of that can overfit the validation set is that I do early stopping based on the validation set.</p>",
              "rawMarkdown": "No, the only place I can think of that can overfit the validation set is that I do early stopping based on the validation set."
            }
          ]
        }
      ]
    },
    {
      "id": 2215648,
      "postDate": "2023-04-09T13:36:01.523Z",
      "content": "<p>MY CV 0.6983 LB: 0.701</p>",
      "rawMarkdown": "MY CV 0.6983 LB: 0.701"
    },
    {
      "id": 2215037,
      "postDate": "2023-04-09T03:40:29.730Z",
      "content": "<p>My current best result</p>\n<p>CV 70096 LB 0.703</p>",
      "rawMarkdown": "My current best result\n\nCV 70096 LB 0.703",
      "replies": [
        {
          "id": 2215646,
          "postDate": "2023-04-09T13:34:46.913Z",
          "content": "<p>single model ?</p>",
          "rawMarkdown": "single model ?",
          "replies": [
            {
              "id": 2215876,
              "postDate": "2023-04-09T16:32:10.003Z",
              "content": "<p>Were you Able to ensemble different models our per question model seems also very slow due to too many features?</p>",
              "rawMarkdown": "Were you Able to ensemble different models our per question model seems also very slow due to too many features?"
            },
            {
              "id": 2215893,
              "postDate": "2023-04-09T16:41:21.407Z",
              "content": "<p>How much total time it's taking to inference.</p>",
              "rawMarkdown": "How much total time it's taking to inference."
            },
            {
              "id": 2215964,
              "postDate": "2023-04-09T17:20:52.233Z",
              "content": "<p>Mine very slow around 6-7 hours :( need to optimize some calculations </p>",
              "rawMarkdown": "Mine very slow around 6-7 hours :( need to optimize some calculations "
            },
            {
              "id": 2216338,
              "postDate": "2023-04-10T02:49:13.550Z",
              "content": "<p>my features might be less than you, but I can only ensemble two models for each questions </p>",
              "rawMarkdown": "my features might be less than you, but I can only ensemble two models for each questions ",
              "votes": 1
            }
          ]
        },
        {
          "id": 2215867,
          "postDate": "2023-04-09T16:28:20.770Z",
          "content": "<p>Nice we got similar cv but not that lb</p>",
          "rawMarkdown": "Nice we got similar cv but not that lb"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2259488,
      "author_name": "theta",
      "author_url": "",
      "post_date": "2023-05-15T03:16:22.913000",
      "content": "<p>CV:0.6997368774742736 LB:0.702 <br>\nCV:0.700170167261253 LB:0.703 <br>\nCV:0.7004214709822321 LB:0.705 <br>\nCV:0.7002594561184703 LB:0.702</p>\n<p>CV does not seem to correlate with LB well.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2259493,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2023-05-15T03:23:58.817000",
          "content": "<p>True observing the same very confusing</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2260253,
          "author_name": "superslon",
          "author_url": "",
          "post_date": "2023-05-15T14:59:07.447000",
          "content": "<p>Yes, i have the same problem. It seems that after cv became 0.700xx (just FE) CV-LB correlation decreased. </p>\n<p>I tried using the old test set as a validation set and found that the data is quite noisy. For example, a small change in cv (~ 0.00002-0.00005) can change the test score by 0.001-0.003.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2260971,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-05-16T03:01:31.147000",
              "content": "<p>Also are you guys <a href=\"https://www.kaggle.com/myppka\" target=\"_blank\">@myppka</a>  and <a href=\"https://www.kaggle.com/yurimaeda\" target=\"_blank\">@yurimaeda</a> having issues with selecting threshold :/ , is a bigger threshold more that gives confidence and overfits less or one that gives or optimizes on CV . some thresholds seem perform better on LB than cv 😧</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2261173,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-16T06:45:31.360000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261174,
              "author_name": "superslon",
              "author_url": "",
              "post_date": "2023-05-16T06:46:06.457000",
              "content": "<p>In my case, only at a threshold of 0.62 was it possible to achieve a difference between CV and LB greater than 0.003</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2261630,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-05-16T13:21:45.693000",
              "content": "<p>Do you guys submit using a single model, or an ensemble of different cv folds? I haven't tried ensembling yet, but single models from different folds for me give drastically different LB scores, disregarding the validation score of that fold.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261657,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-05-16T13:41:07.170000",
              "content": "<p>Single model for now…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2261734,
              "author_name": "superslon",
              "author_url": "",
              "post_date": "2023-05-16T14:14:11.917000",
              "content": "<p>Also single model</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2261782,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-05-16T14:46:35.857000",
              "content": "<p>Oh, so when you guys say the CV is something, do you mean the macro F1 of all folds or score of a single validation fold?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261916,
              "author_name": "superslon",
              "author_url": "",
              "post_date": "2023-05-16T16:13:47.183000",
              "content": "<p>I use macro F1 of all folds.<br>\nYou can <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218#2193827\" target=\"_blank\">retrain</a> one model per question using 100% of data. I think the results should become more stable</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2261926,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-05-16T16:16:50.277000",
              "content": "<p>Thanks for the advice, I did this at the start of the competition, but at times it felt a bit unstable, so I stopped doing that, maybe it's time to explore that option again :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2211456,
      "author_name": "Ya Xu",
      "author_url": "",
      "post_date": "2023-04-06T04:04:15.090000",
      "content": "<p>My current best result</p>\n<p>CV  69775      LB  0.700<br>\nCV  69823      LB  0.698</p>\n<p>Some thoughts about the correlation between CV and LB:</p>\n<ul>\n<li><p>Correlation became a mess as soon as I tried to do feature selection (as <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> mentioned deleting features that have 0 importance in every fold, which I believed shouldn't cause any leakage ). Maybe there is something wrong with my feature selection strategy, I have currently abandoned the idea of doing any feature selection.</p></li>\n<li><p>Correlation worked ok for most cases if you just add new features, but not all cases. My best LB was achieved by some blind attempt that added a feature that cause a 0.0002 drop in CV, but it boosted LB from 0.698 to 0.700.   </p></li>\n<li><p>Correlation worked much better for time-based features than coordinate-based features. </p></li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 2212015,
          "author_name": "LeLeCHAA",
          "author_url": "",
          "post_date": "2023-04-06T12:54:34.623000",
          "content": "<p>For seeing <strong><em>“Correlation became a mess as soon as I tried to do feature selection”</em></strong> i remember something.</p>\n<p>Actually i have tried to do feature selecton with feat-importance.</p>\n<p>For example, i begin with 1000 featres and regularly train with xgboost, after training i drop 5(for example) least important feature.</p>\n<p>After about 100 iters , i found that only first 2-4 iter will boost CV weakly, then CV will descend quickly，so i give up to use feature selection.</p>\n<p>Is there any relation between my experiment and your mess correlation？</p>\n<p>It seems that feature selection dosen't work well on the dataset.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2212032,
              "author_name": "LeLeCHAA",
              "author_url": "",
              "post_date": "2023-04-06T13:14:22.723000",
              "content": "<p>Sorry, i rechecked my experiment i find <strong><em>\"then CV will descend quickly\"</em></strong> is wrong, actually CV float around(-0.0005,+0.0003), it seems that with feature num reduced, it make no sense for the model.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2212067,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-04-06T13:58:16.057000",
              "content": "<p>I am also finding similar issue with Feature Selection specially removing zero importance features . Need to maybe check permutation importance but will be too time consuming :| </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2214242,
              "author_name": "Ya Xu",
              "author_url": "",
              "post_date": "2023-04-08T09:15:04.620000",
              "content": "<p>I took the same strategy as <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> mentioned, deleting features that have zero importance in every fold(to avoid leakage) at first. It showed some minimal boost like 0.0005 for CV which I considered as noise, then LB drop sharply like 0.02.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2214259,
              "author_name": "LeLeCHAA",
              "author_url": "",
              "post_date": "2023-04-08T09:46:30.817000",
              "content": "<p>Maybe the reduce of features caused overfit？</p>\n<p>May i ask how much proportion have you guys reduced?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2214337,
              "author_name": "Ya Xu",
              "author_url": "",
              "post_date": "2023-04-08T11:39:35.170000",
              "content": "<p>It was not supposed to cause overfitting, since I only eliminated features that have 0 importance for every model in that specific level_group. There are more 0-importance features for each fold/question, but only a small portion that has 0 importance for every fold/question, like 2% of features. I did not expect a 0.02 drop for LB since 2% seems insignificant to me. </p>\n<p>Also when I tried to eliminate 4%, 6% 8%… of features, CV went back and forth but had some improvements like 0.005 overtime, LB also improved a little back but never met the optimal score with 100% features. </p>\n<p>Just realized maybe I misunderstood how <code>XGB.feature_importace</code> works, I should recheck that.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2214399,
              "author_name": "Fisch",
              "author_url": "",
              "post_date": "2023-04-08T12:45:53.153000",
              "content": "<p>Eliminating features that have 0 importance should not change anything indeed. </p>\n<p>However, could the <strong>col_sample</strong> parameter (or any random sampling of features) be impacted by this ? I have also come across surprising behavior when changing the order of features in my dataframe. I was not expecting to see any change of performance and yet, the order of features seems to make a difference. </p>\n<p>I have always concluded that it showed how unstable my model was, the same way when changing the random seed drastically changes performance. </p>\n<p>On my side, I haven't reached the same performance on the LB as you guys, but I am performing feature selection based on both gain, weight and cover by taking the intersection of the best (top 40% (arbitrarily chosen)) features of each importance type. As a result, I end up with a model that has way less features and better performance ( +0.003) and lower generalization gap between train/validation. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2215002,
              "author_name": "cjz",
              "author_url": "",
              "post_date": "2023-04-09T01:36:05.323000",
              "content": "<p>A great feature selection method,may I ask if both cv and lb increased simultaneously in the end？</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2215459,
              "author_name": "Fisch",
              "author_url": "",
              "post_date": "2023-04-09T10:12:12.360000",
              "content": "<p>Yes, in my case both cv and lb increased simultaneously, but I am facing the same issue as everyone : noticing some fluctuation on the LB. At least, by reducing the number of features, it's faster to train models and try new features/ideas.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2215869,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-04-09T16:30:05.073000",
              "content": "<p>Also did anyone also observe this in our case tinkering hyperparams like tree depth etc does improve cv but gives worse in lb , not sure it overfits a bit 😐</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2258965,
      "author_name": "Woprime",
      "author_url": "",
      "post_date": "2023-05-14T15:23:29.080000",
      "content": "<p>How's everyone's CV LB correlations stands after the update? <br>\nFor me CV 0.6997 - LB 0.701</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2217368,
      "author_name": "Joseph Zhou",
      "author_url": "",
      "post_date": "2023-04-10T20:15:35.233000",
      "content": "<p>It seems the distribution of train doesn't resemble public LB's so much.</p>\n<blockquote>\n  <p>CV: 0.69952 LB: 0.702<br>\n  CV: 0.70025 LB: 0.703<br>\n  CV: 0.70033 LB: 0.704</p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 2217388,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2023-04-10T21:08:39.557000",
          "content": "<p>You LB is bit better seems than cv ,but correlation good . we are somewhat not getting this right now :| </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2241558,
      "author_name": "Angelina Tseng",
      "author_url": "",
      "post_date": "2023-05-01T15:41:58.133000",
      "content": "<p>For me, cv is stable. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2218062,
      "author_name": "GBalázs",
      "author_url": "",
      "post_date": "2023-04-11T12:01:13.457000",
      "content": "<p>My LB scores are consistently lower by 0.005-0.010 than my CV scores. This is not in line with what most of you are reporting in this and other threads. I'm not doing anything spectacular, using XGBClassifier with plenty of features, most of those are copied from public notebooks. I have no idea what I'm doing wrong, I'd appreciate if someone could help me understand this. Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2216772,
      "author_name": "superslon",
      "author_url": "",
      "post_date": "2023-04-10T10:37:14.797000",
      "content": "<p>My current results:</p>\n<ul>\n<li>Best CV: CV 0.69951 LB 0.699 </li>\n<li>Best LB: CV 0.69914 LB 0.703</li>\n</ul>\n<p>After changing the test set, the correlation between CV and LB decreased(</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2216386,
      "author_name": "Kefan Xu",
      "author_url": "",
      "post_date": "2023-04-10T04:01:15.633000",
      "content": "<p>0.628 LB 0.578</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2216115,
      "author_name": "Elias",
      "author_url": "",
      "post_date": "2023-04-09T19:35:10.880000",
      "content": "<p>CV 0.6992 LB 0.701 but it is unstable.<br>\n There is probably 0.002 noise/overfitting as my LB drops to 0.699 with small changes that improve CV<br>\nthreshold optimization may be overfitting too?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2216192,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2023-04-09T21:06:35.900000",
          "content": "<p>we getting same behaviour😀 guess think still the magic is missing to cross this</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2215809,
      "author_name": "Woprime",
      "author_url": "",
      "post_date": "2023-04-09T15:48:23.727000",
      "content": "<p>CV: 0.6995 LB: 0.698,</p>\n<p>I wonder if I'm overfitting, seeing you guys all have higher LB scores than CV scores.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2216118,
          "author_name": "Elias",
          "author_url": "",
          "post_date": "2023-04-09T19:37:25.157000",
          "content": "<p>did you select the features with high importance ? that may overfit.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2216153,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-04-09T20:17:30.927000",
              "content": "<p>No, the only place I can think of that can overfit the validation set is that I do early stopping based on the validation set.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2215648,
      "author_name": "HZM",
      "author_url": "",
      "post_date": "2023-04-09T13:36:01.523000",
      "content": "<p>MY CV 0.6983 LB: 0.701</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2215037,
      "author_name": "老肥",
      "author_url": "",
      "post_date": "2023-04-09T03:40:29.730000",
      "content": "<p>My current best result</p>\n<p>CV 70096 LB 0.703</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2215646,
          "author_name": "HZM",
          "author_url": "",
          "post_date": "2023-04-09T13:34:46.913000",
          "content": "<p>single model ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2215876,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-04-09T16:32:10.003000",
              "content": "<p>Were you Able to ensemble different models our per question model seems also very slow due to too many features?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2215893,
              "author_name": "Ashish Kumar Singh 🇮🇳",
              "author_url": "",
              "post_date": "2023-04-09T16:41:21.407000",
              "content": "<p>How much total time it's taking to inference.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2215964,
              "author_name": "Gaurav Rawat",
              "author_url": "",
              "post_date": "2023-04-09T17:20:52.233000",
              "content": "<p>Mine very slow around 6-7 hours :( need to optimize some calculations </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2216338,
              "author_name": "HZM",
              "author_url": "",
              "post_date": "2023-04-10T02:49:13.550000",
              "content": "<p>my features might be less than you, but I can only ensemble two models for each questions </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2215867,
          "author_name": "Gaurav Rawat",
          "author_url": "",
          "post_date": "2023-04-09T16:28:20.770000",
          "content": "<p>Nice we got similar cv but not that lb</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2210723": "Hi I was curious about how the CV and LB correlates for other folks . Posting our team results best below \n\n- `CV 70082 LB 0.701`\n- `CV 70053 LB 0.703`\n- `CV 70099 LB 0.705`\n- `CV 70107  LB 0.703`\n- `CV 70107  LB 0.705` Latest submission😑\n\nAs mentioned in some posts there is a 0.001-0.002 gap in our case . FE helps the best but Hyper-param Tuning/Threshold in our case does improve CV in some cases but we don't see same in LB which could be hyper-params overfit but would be good to know what you guys are relying on .",
    "2259488": "CV:0.6997368774742736 LB:0.702 \nCV:0.700170167261253 LB:0.703 \nCV:0.7004214709822321 LB:0.705 \nCV:0.7002594561184703 LB:0.702\n\nCV does not seem to correlate with LB well.",
    "2211456": "My current best result\n\nCV  69775      LB  0.700\nCV  69823      LB  0.698\n\nSome thoughts about the correlation between CV and LB:\n\n- Correlation became a mess as soon as I tried to do feature selection (as @gauravbrills mentioned deleting features that have 0 importance in every fold, which I believed shouldn't cause any leakage ). Maybe there is something wrong with my feature selection strategy, I have currently abandoned the idea of doing any feature selection.\n\n- Correlation worked ok for most cases if you just add new features, but not all cases. My best LB was achieved by some blind attempt that added a feature that cause a 0.0002 drop in CV, but it boosted LB from 0.698 to 0.700.   \n\n- Correlation worked much better for time-based features than coordinate-based features. ",
    "2258965": "How's everyone's CV LB correlations stands after the update? \nFor me CV 0.6997 - LB 0.701",
    "2217368": "It seems the distribution of train doesn't resemble public LB's so much.\n>CV: 0.69952 LB: 0.702\n>CV: 0.70025 LB: 0.703\n>CV: 0.70033 LB: 0.704",
    "2241558": "For me, cv is stable. ",
    "2218062": "My LB scores are consistently lower by 0.005-0.010 than my CV scores. This is not in line with what most of you are reporting in this and other threads. I'm not doing anything spectacular, using XGBClassifier with plenty of features, most of those are copied from public notebooks. I have no idea what I'm doing wrong, I'd appreciate if someone could help me understand this. Thanks!",
    "2216772": "My current results:\n- Best CV: CV 0.69951 LB 0.699 \n- Best LB: CV 0.69914 LB 0.703\n\nAfter changing the test set, the correlation between CV and LB decreased(\n",
    "2216386": "0.628 LB 0.578",
    "2216115": "CV 0.6992 LB 0.701 but it is unstable.\n There is probably 0.002 noise/overfitting as my LB drops to 0.699 with small changes that improve CV\nthreshold optimization may be overfitting too?",
    "2215809": "CV: 0.6995 LB: 0.698,\n\nI wonder if I'm overfitting, seeing you guys all have higher LB scores than CV scores.",
    "2215648": "MY CV 0.6983 LB: 0.701",
    "2215037": "My current best result\n\nCV 70096 LB 0.703"
  }
}