{
  "id": 75747,
  "title": "Correlation between CV and LB",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/75747",
  "author_name": "",
  "post_date": "2018-12-26T00:06:57.547222400Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've seen this problem arise in several other competitions, but in this one the discrepancy seems particularly egregious, or perhaps I'm missing something. </p>\n\n<p>I used Leo's <a href=\"https://www.kaggle.com/bluexleoxgreen/simple-feature-lightgbm-baseline\">kernel</a>, added a few features, the CV score went from 0.83 to 0.86, but the LB came down from 0.183 to 0.167. </p>\n\n<p>Still early days, and I'll report on whatever new findings I get. </p>",
  "messages": [
    {
      "id": "445206",
      "postDate": "12/26/2018 00:06:57",
      "content": "<p>I've seen this problem arise in several other competitions, but in this one the discrepancy seems particularly egregious, or perhaps I'm missing something. </p>\n\n<p>I used Leo's <a href=\"https://www.kaggle.com/bluexleoxgreen/simple-feature-lightgbm-baseline\">kernel</a>, added a few features, the CV score went from 0.83 to 0.86, but the LB came down from 0.183 to 0.167. </p>\n\n<p>Still early days, and I'll report on whatever new findings I get. </p>",
      "rawMarkdown": "I've seen this problem arise in several other competitions, but in this one the discrepancy seems particularly egregious, or perhaps I'm missing something. \n\nI used Leo's [kernel][1], added a few features, the CV score went from 0.83 to 0.86, but the LB came down from 0.183 to 0.167. \n\nStill early days, and I'll report on whatever new findings I get. \n\n  [1]: https://www.kaggle.com/bluexleoxgreen/simple-feature-lightgbm-baseline",
      "votes": null
    },
    {
      "id": "445213",
      "postDate": "12/26/2018 00:49:13",
      "content": "<p>If adding features makes CV go up but LB go down then I think you are overfitting to the training data and need to improve CV. </p>",
      "rawMarkdown": "If adding features makes CV go up but LB go down then I think you are overfitting to the training data and need to improve CV.",
      "votes": null
    },
    {
      "id": "445214",
      "postDate": "12/26/2018 00:51:09",
      "content": "<p>That was my impression too. But in a situation like this, how does one judge performance? I cant tell how my model is performing based on a higher CV... </p>",
      "rawMarkdown": "That was my impression too. But in a situation like this, how does one judge performance? I cant tell how my model is performing based on a higher CV...",
      "votes": null
    },
    {
      "id": "445230",
      "postDate": "12/26/2018 02:40:41",
      "content": "<p>Once CV and LB start improving together you can trust CV but I think the model you based your model on is overfitting a lot to begin with. I think using AUC metric rather than MCC might be the problem but I’m not really sure at this time. </p>",
      "rawMarkdown": "Once CV and LB start improving together you can trust CV but I think the model you based your model on is overfitting a lot to begin with. I think using AUC metric rather than MCC might be the problem but I’m not really sure at this time.",
      "votes": null
    },
    {
      "id": "445231",
      "postDate": "12/26/2018 02:43:31",
      "content": "<p>That's certainly possible. Currently trying a LightGBM with a custom feval corresponding to MCC... let's see how that goes. </p>",
      "rawMarkdown": "That's certainly possible. Currently trying a LightGBM with a custom feval corresponding to MCC... let's see how that goes.",
      "votes": null
    },
    {
      "id": "445236",
      "postDate": "12/26/2018 02:54:56",
      "content": "<p>Hope it helps! I am interested to see what effect it will have on CV and LB scores. </p>",
      "rawMarkdown": "Hope it helps! I am interested to see what effect it will have on CV and LB scores.",
      "votes": null
    },
    {
      "id": "445397",
      "postDate": "12/26/2018 11:28:15",
      "content": "<p>One severe issue with the kernel is that 'id_measurement' is used as a feature. Additionally in most cases the target is the same for all phases within measurement, but here a single measurement can be split across folds. You should drop 'id_measurement' as well as 'phase' from the features for training and prediction. The use of those features lead to overfitting. </p>\n\n<p>Additionally you seem to be comparing AUC locally with MCC on LB.</p>",
      "rawMarkdown": "One severe issue with the kernel is that 'id\\_measurement' is used as a feature. Additionally in most cases the target is the same for all phases within measurement, but here a single measurement can be split across folds. You should drop 'id\\_measurement' as well as 'phase' from the features for training and prediction. The use of those features lead to overfitting. \n\nAdditionally you seem to be comparing AUC locally with MCC on LB.",
      "votes": null
    },
    {
      "id": "445465",
      "postDate": "12/26/2018 14:25:00",
      "content": "<p>Id_measurement definitely. But why phase? And yes, as Jack and I discussed, I think the metric is a but issue here </p>",
      "rawMarkdown": "Id_measurement definitely. But why phase? And yes, as Jack and I discussed, I think the metric is a but issue here",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 445213,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/26/2018 00:49:13",
      "content": "<p>If adding features makes CV go up but LB go down then I think you are overfitting to the training data and need to improve CV. </p>",
      "votes": null,
      "replies": [
        {
          "id": 445214,
          "author_name": "delayedkarma",
          "author_url": "",
          "post_date": "12/26/2018 00:51:09",
          "content": "<p>That was my impression too. But in a situation like this, how does one judge performance? I cant tell how my model is performing based on a higher CV... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445230,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "12/26/2018 02:40:41",
          "content": "<p>Once CV and LB start improving together you can trust CV but I think the model you based your model on is overfitting a lot to begin with. I think using AUC metric rather than MCC might be the problem but I’m not really sure at this time. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445231,
          "author_name": "delayedkarma",
          "author_url": "",
          "post_date": "12/26/2018 02:43:31",
          "content": "<p>That's certainly possible. Currently trying a LightGBM with a custom feval corresponding to MCC... let's see how that goes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445236,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "12/26/2018 02:54:56",
          "content": "<p>Hope it helps! I am interested to see what effect it will have on CV and LB scores. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 445397,
      "author_name": "helgith",
      "author_url": "",
      "post_date": "12/26/2018 11:28:15",
      "content": "<p>One severe issue with the kernel is that 'id_measurement' is used as a feature. Additionally in most cases the target is the same for all phases within measurement, but here a single measurement can be split across folds. You should drop 'id_measurement' as well as 'phase' from the features for training and prediction. The use of those features lead to overfitting. </p>\n\n<p>Additionally you seem to be comparing AUC locally with MCC on LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 445465,
          "author_name": "delayedkarma",
          "author_url": "",
          "post_date": "12/26/2018 14:25:00",
          "content": "<p>Id_measurement definitely. But why phase? And yes, as Jack and I discussed, I think the metric is a but issue here </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "445206": "I've seen this problem arise in several other competitions, but in this one the discrepancy seems particularly egregious, or perhaps I'm missing something. \n\nI used Leo's [kernel][1], added a few features, the CV score went from 0.83 to 0.86, but the LB came down from 0.183 to 0.167. \n\nStill early days, and I'll report on whatever new findings I get. \n\n  [1]: https://www.kaggle.com/bluexleoxgreen/simple-feature-lightgbm-baseline",
    "445213": "If adding features makes CV go up but LB go down then I think you are overfitting to the training data and need to improve CV.",
    "445214": "That was my impression too. But in a situation like this, how does one judge performance? I cant tell how my model is performing based on a higher CV...",
    "445230": "Once CV and LB start improving together you can trust CV but I think the model you based your model on is overfitting a lot to begin with. I think using AUC metric rather than MCC might be the problem but I’m not really sure at this time.",
    "445231": "That's certainly possible. Currently trying a LightGBM with a custom feval corresponding to MCC... let's see how that goes.",
    "445236": "Hope it helps! I am interested to see what effect it will have on CV and LB scores.",
    "445397": "One severe issue with the kernel is that 'id\\_measurement' is used as a feature. Additionally in most cases the target is the same for all phases within measurement, but here a single measurement can be split across folds. You should drop 'id\\_measurement' as well as 'phase' from the features for training and prediction. The use of those features lead to overfitting. \n\nAdditionally you seem to be comparing AUC locally with MCC on LB.",
    "445465": "Id_measurement definitely. But why phase? And yes, as Jack and I discussed, I think the metric is a but issue here"
  },
  "source": "meta"
}