{
  "id": 91602,
  "title": "Is your CV-publicLB correlation stable?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91602",
  "author_name": "",
  "post_date": "2019-05-07T01:42:29.761874600Z",
  "votes": 15,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I create this topic to let people share their CV-publicLB correlation. </p>\n\n<p>Does your CV improvement translate to public LB improvement? And if so, how stable is it? (Always, sometimes, rarely...?). </p>\n\n<p>In my case, it is only “relatively so-so”, which means 70% of the times my better CV gets better LB. However, there are some extreme cases which contradict the correlation (very good CV but very bad LB, or vice versa). Those cannot make me sleep well. </p>",
  "messages": [
    {
      "id": "528055",
      "postDate": "05/07/2019 01:42:29",
      "content": "<p>I create this topic to let people share their CV-publicLB correlation. </p>\n\n<p>Does your CV improvement translate to public LB improvement? And if so, how stable is it? (Always, sometimes, rarely...?). </p>\n\n<p>In my case, it is only “relatively so-so”, which means 70% of the times my better CV gets better LB. However, there are some extreme cases which contradict the correlation (very good CV but very bad LB, or vice versa). Those cannot make me sleep well. </p>",
      "rawMarkdown": "I create this topic to let people share their CV-publicLB correlation. \n\nDoes your CV improvement translate to public LB improvement? And if so, how stable is it? (Always, sometimes, rarely...?). \n\nIn my case, it is only “relatively so-so”, which means 70% of the times my better CV gets better LB. However, there are some extreme cases which contradict the correlation (very good CV but very bad LB, or vice versa). Those cannot make me sleep well.",
      "votes": null
    },
    {
      "id": "528099",
      "postDate": "05/07/2019 04:37:17",
      "content": "<p>I’ve noticed the same thing with my CV technique. It could be entirely dependent on the way CV is computed though. Shuffled, not shuffled, quake-wise, number of folds- all could play a factor in the CV to LB relationship. </p>",
      "rawMarkdown": "I’ve noticed the same thing with my CV technique. It could be entirely dependent on the way CV is computed though. Shuffled, not shuffled, quake-wise, number of folds- all could play a factor in the CV to LB relationship.",
      "votes": null
    },
    {
      "id": "528126",
      "postDate": "05/07/2019 05:41:28",
      "content": "<p>Yes. So I create this topic to see if there is anyone in this world having a stable CV-LB correlation. At least we need to believe it’s possible or not. </p>",
      "rawMarkdown": "Yes. So I create this topic to see if there is anyone in this world having a stable CV-LB correlation. At least we need to believe it’s possible or not.",
      "votes": null
    },
    {
      "id": "528204",
      "postDate": "05/07/2019 08:56:40",
      "content": "<p>My belief is that no one has a stable CV LB relationship.</p>",
      "rawMarkdown": "My belief is that no one has a stable CV LB relationship.",
      "votes": null
    },
    {
      "id": "528218",
      "postDate": "05/07/2019 09:20:37",
      "content": "<p>Well. Everything is possible IMHO. </p>",
      "rawMarkdown": "Well. Everything is possible IMHO.",
      "votes": null
    },
    {
      "id": "528255",
      "postDate": "05/07/2019 11:17:59",
      "content": "<p>I would not worry too much if it is not stable as public LB is useless.  I know, I repeat myself endlessly ;)</p>",
      "rawMarkdown": "I would not worry too much if it is not stable as public LB is useless.  I know, I repeat myself endlessly ;)",
      "votes": null
    },
    {
      "id": "528256",
      "postDate": "05/07/2019 11:20:40",
      "content": "<p>My CV-LB correlation is stable within the same model, but not between models. \nSo I can pretty much monitor the effect of adding/removing features or modifying the preprocessing as long as I'm using the same CV, same seed and same model. But when it comes to comparing between different models, it stops making sense. I think it is mainly because different models learn different baselines with similar CV score. Because of that, I am still using models without tuning (mostly default parameters).</p>",
      "rawMarkdown": "My CV-LB correlation is stable within the same model, but not between models. \nSo I can pretty much monitor the effect of adding/removing features or modifying the preprocessing as long as I'm using the same CV, same seed and same model. But when it comes to comparing between different models, it stops making sense. I think it is mainly because different models learn different baselines with similar CV score. Because of that, I am still using models without tuning (mostly default parameters).",
      "votes": null
    },
    {
      "id": "528260",
      "postDate": "05/07/2019 11:26:20",
      "content": "<p>Not stable of course, but to what extent? Do you have any extreme cases? (say, improved CV, but unexpectedly and surprisingly much much worse LB, or vice versa)</p>",
      "rawMarkdown": "Not stable of course, but to what extent? Do you have any extreme cases? (say, improved CV, but unexpectedly and surprisingly much much worse LB, or vice versa)",
      "votes": null
    },
    {
      "id": "528289",
      "postDate": "05/07/2019 12:28:31",
      "content": "<p>Said differently: I am not using LB score in any way to guide me.</p>",
      "rawMarkdown": "Said differently: I am not using LB score in any way to guide me.",
      "votes": null
    },
    {
      "id": "528340",
      "postDate": "05/07/2019 14:29:01",
      "content": "<p>my CV, LB relationship is pretty stable. ofc not the third decimal, but first and usually second decimal are in line. \njust using standard 3KFold without shuffling so far to test features</p>",
      "rawMarkdown": "my CV, LB relationship is pretty stable. ofc not the third decimal, but first and usually second decimal are in line. \njust using standard 3KFold without shuffling so far to test features",
      "votes": null
    },
    {
      "id": "528368",
      "postDate": "05/07/2019 15:44:36",
      "content": "<p>I agree with <a href=\"/cpmpml\">@cpmpml</a>. With only 340 samples, public lb should be considered one more fold in cross-validation imho.</p>",
      "rawMarkdown": "I agree with @cpmpml. With only 340 samples, public lb should be considered one more fold in cross-validation imho.",
      "votes": null
    },
    {
      "id": "528378",
      "postDate": "05/07/2019 16:07:43",
      "content": "<p>Right, saying it is totally useless is too strong.  Looking at it as one fold is more accurate.</p>",
      "rawMarkdown": "Right, saying it is totally useless is too strong.  Looking at it as one fold is more accurate.",
      "votes": null
    },
    {
      "id": "528415",
      "postDate": "05/07/2019 18:33:15",
      "content": "<p>I trust the LB a little more when I shut one eye ;)</p>",
      "rawMarkdown": "I trust the LB a little more when I shut one eye ;)",
      "votes": null
    },
    {
      "id": "528429",
      "postDate": "05/07/2019 19:34:07",
      "content": "<p>As someone performing terribly on the leaderboard, I lend my support to this viewpoint!</p>",
      "rawMarkdown": "As someone performing terribly on the leaderboard, I lend my support to this viewpoint!",
      "votes": null
    },
    {
      "id": "528436",
      "postDate": "05/07/2019 20:10:39",
      "content": "<blockquote>\n  <p>Right, saying it is totally useless is too strong. Looking at it as one fold is more accurate.</p>\n</blockquote>\n\n<p>This sums it together very well!</p>",
      "rawMarkdown": "&gt; Right, saying it is totally useless is too strong. Looking at it as one fold is more accurate.\n\nThis sums it together very well!",
      "votes": null
    },
    {
      "id": "529086",
      "postDate": "05/09/2019 06:56:00",
      "content": "<p>I think the correlation CV-LB is strongly influenced by the high variance of training data (some subsets have more correlation than others). I tested a NN using different shuffling points (via random_state option) and for some shuffling it seems to have higher correlation as if LB is better simulated by just some subsets of training data. Still looking into the issue.  </p>",
      "rawMarkdown": "I think the correlation CV-LB is strongly influenced by the high variance of training data (some subsets have more correlation than others). I tested a NN using different shuffling points (via random_state option) and for some shuffling it seems to have higher correlation as if LB is better simulated by just some subsets of training data. Still looking into the issue.",
      "votes": null
    },
    {
      "id": "529350",
      "postDate": "05/09/2019 17:36:54",
      "content": "<p>I think cv and LB has no relation too.\nBut the LB must right.\nAnd can not sleep well.</p>",
      "rawMarkdown": "I think cv and LB has no relation too.\nBut the LB must right.\nAnd can not sleep well.",
      "votes": null
    },
    {
      "id": "529605",
      "postDate": "05/10/2019 10:14:24",
      "content": "<p>Our blending CV-LB correlation is (75%) stable, but single model is not.</p>",
      "rawMarkdown": "Our blending CV-LB correlation is (75%) stable, but single model is not.",
      "votes": null
    },
    {
      "id": "529637",
      "postDate": "05/10/2019 11:50:05",
      "content": "<p>From my view it is stable and dependent on each other also little bit related due to variances in training set and influenced also on LB score. </p>",
      "rawMarkdown": "From my view it is stable and dependent on each other also little bit related due to variances in training set and influenced also on LB score.",
      "votes": null
    },
    {
      "id": "540721",
      "postDate": "06/01/2019 00:12:36",
      "content": "<p>Every time I look at the correlation between my CV and LB, it like tells me \"You know nothing, John Snow\"..</p>",
      "rawMarkdown": "Every time I look at the correlation between my CV and LB, it like tells me \"You know nothing, John Snow\"..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 528099,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "05/07/2019 04:37:17",
      "content": "<p>I’ve noticed the same thing with my CV technique. It could be entirely dependent on the way CV is computed though. Shuffled, not shuffled, quake-wise, number of folds- all could play a factor in the CV to LB relationship. </p>",
      "votes": null,
      "replies": [
        {
          "id": 528126,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/07/2019 05:41:28",
          "content": "<p>Yes. So I create this topic to see if there is anyone in this world having a stable CV-LB correlation. At least we need to believe it’s possible or not. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528204,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "05/07/2019 08:56:40",
      "content": "<p>My belief is that no one has a stable CV LB relationship.</p>",
      "votes": null,
      "replies": [
        {
          "id": 528218,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/07/2019 09:20:37",
          "content": "<p>Well. Everything is possible IMHO. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528340,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/07/2019 14:29:01",
          "content": "<p>my CV, LB relationship is pretty stable. ofc not the third decimal, but first and usually second decimal are in line. \njust using standard 3KFold without shuffling so far to test features</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528255,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/07/2019 11:17:59",
      "content": "<p>I would not worry too much if it is not stable as public LB is useless.  I know, I repeat myself endlessly ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 528260,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/07/2019 11:26:20",
          "content": "<p>Not stable of course, but to what extent? Do you have any extreme cases? (say, improved CV, but unexpectedly and surprisingly much much worse LB, or vice versa)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528289,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 12:28:31",
          "content": "<p>Said differently: I am not using LB score in any way to guide me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528368,
          "author_name": "jsaguiar",
          "author_url": "",
          "post_date": "05/07/2019 15:44:36",
          "content": "<p>I agree with <a href=\"/cpmpml\">@cpmpml</a>. With only 340 samples, public lb should be considered one more fold in cross-validation imho.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528378,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/07/2019 16:07:43",
          "content": "<p>Right, saying it is totally useless is too strong.  Looking at it as one fold is more accurate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528415,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "05/07/2019 18:33:15",
          "content": "<p>I trust the LB a little more when I shut one eye ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528429,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/07/2019 19:34:07",
          "content": "<p>As someone performing terribly on the leaderboard, I lend my support to this viewpoint!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 528436,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "05/07/2019 20:10:39",
          "content": "<blockquote>\n  <p>Right, saying it is totally useless is too strong. Looking at it as one fold is more accurate.</p>\n</blockquote>\n\n<p>This sums it together very well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 528256,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "05/07/2019 11:20:40",
      "content": "<p>My CV-LB correlation is stable within the same model, but not between models. \nSo I can pretty much monitor the effect of adding/removing features or modifying the preprocessing as long as I'm using the same CV, same seed and same model. But when it comes to comparing between different models, it stops making sense. I think it is mainly because different models learn different baselines with similar CV score. Because of that, I am still using models without tuning (mostly default parameters).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529086,
      "author_name": "gianfrancobarone",
      "author_url": "",
      "post_date": "05/09/2019 06:56:00",
      "content": "<p>I think the correlation CV-LB is strongly influenced by the high variance of training data (some subsets have more correlation than others). I tested a NN using different shuffling points (via random_state option) and for some shuffling it seems to have higher correlation as if LB is better simulated by just some subsets of training data. Still looking into the issue.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529350,
      "author_name": "zjp2origin",
      "author_url": "",
      "post_date": "05/09/2019 17:36:54",
      "content": "<p>I think cv and LB has no relation too.\nBut the LB must right.\nAnd can not sleep well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529605,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "05/10/2019 10:14:24",
      "content": "<p>Our blending CV-LB correlation is (75%) stable, but single model is not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529637,
      "author_name": "himusoni",
      "author_url": "",
      "post_date": "05/10/2019 11:50:05",
      "content": "<p>From my view it is stable and dependent on each other also little bit related due to variances in training set and influenced also on LB score. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 540721,
      "author_name": "elvenmonk",
      "author_url": "",
      "post_date": "06/01/2019 00:12:36",
      "content": "<p>Every time I look at the correlation between my CV and LB, it like tells me \"You know nothing, John Snow\"..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "528055": "I create this topic to let people share their CV-publicLB correlation. \n\nDoes your CV improvement translate to public LB improvement? And if so, how stable is it? (Always, sometimes, rarely...?). \n\nIn my case, it is only “relatively so-so”, which means 70% of the times my better CV gets better LB. However, there are some extreme cases which contradict the correlation (very good CV but very bad LB, or vice versa). Those cannot make me sleep well.",
    "528099": "I’ve noticed the same thing with my CV technique. It could be entirely dependent on the way CV is computed though. Shuffled, not shuffled, quake-wise, number of folds- all could play a factor in the CV to LB relationship.",
    "528126": "Yes. So I create this topic to see if there is anyone in this world having a stable CV-LB correlation. At least we need to believe it’s possible or not.",
    "528204": "My belief is that no one has a stable CV LB relationship.",
    "528218": "Well. Everything is possible IMHO.",
    "528255": "I would not worry too much if it is not stable as public LB is useless.  I know, I repeat myself endlessly ;)",
    "528256": "My CV-LB correlation is stable within the same model, but not between models. \nSo I can pretty much monitor the effect of adding/removing features or modifying the preprocessing as long as I'm using the same CV, same seed and same model. But when it comes to comparing between different models, it stops making sense. I think it is mainly because different models learn different baselines with similar CV score. Because of that, I am still using models without tuning (mostly default parameters).",
    "528260": "Not stable of course, but to what extent? Do you have any extreme cases? (say, improved CV, but unexpectedly and surprisingly much much worse LB, or vice versa)",
    "528289": "Said differently: I am not using LB score in any way to guide me.",
    "528340": "my CV, LB relationship is pretty stable. ofc not the third decimal, but first and usually second decimal are in line. \njust using standard 3KFold without shuffling so far to test features",
    "528368": "I agree with @cpmpml. With only 340 samples, public lb should be considered one more fold in cross-validation imho.",
    "528378": "Right, saying it is totally useless is too strong.  Looking at it as one fold is more accurate.",
    "528415": "I trust the LB a little more when I shut one eye ;)",
    "528429": "As someone performing terribly on the leaderboard, I lend my support to this viewpoint!",
    "528436": "&gt; Right, saying it is totally useless is too strong. Looking at it as one fold is more accurate.\n\nThis sums it together very well!",
    "529086": "I think the correlation CV-LB is strongly influenced by the high variance of training data (some subsets have more correlation than others). I tested a NN using different shuffling points (via random_state option) and for some shuffling it seems to have higher correlation as if LB is better simulated by just some subsets of training data. Still looking into the issue.",
    "529350": "I think cv and LB has no relation too.\nBut the LB must right.\nAnd can not sleep well.",
    "529605": "Our blending CV-LB correlation is (75%) stable, but single model is not.",
    "529637": "From my view it is stable and dependent on each other also little bit related due to variances in training set and influenced also on LB score.",
    "540721": "Every time I look at the correlation between my CV and LB, it like tells me \"You know nothing, John Snow\".."
  },
  "source": "meta"
}