{
  "id": 82462,
  "title": "CV/LB score",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/82462",
  "author_name": "",
  "post_date": "2019-03-01T13:35:56.847141400Z",
  "votes": 9,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi everyone! I'm opening this topic because the score on cv  with respect to lb score appears quite strange.  My best CV is 0.54 on training and 0.63 on test and 1.74 on LB.  My best LB score instead is 1.55 that has 1.89 on training and 2.19 on test. Are you, Dear Competitors, willing to share your own observations about this difference? \nThanks.</p>",
  "messages": [
    {
      "id": "481498",
      "postDate": "03/01/2019 13:35:56",
      "content": "<p>Hi everyone! I'm opening this topic because the score on cv  with respect to lb score appears quite strange.  My best CV is 0.54 on training and 0.63 on test and 1.74 on LB.  My best LB score instead is 1.55 that has 1.89 on training and 2.19 on test. Are you, Dear Competitors, willing to share your own observations about this difference? \nThanks.</p>",
      "rawMarkdown": "Hi everyone! I'm opening this topic because the score on cv  with respect to lb score appears quite strange.  My best CV is 0.54 on training and 0.63 on test and 1.74 on LB.  My best LB score instead is 1.55 that has 1.89 on training and 2.19 on test. Are you, Dear Competitors, willing to share your own observations about this difference? \nThanks.",
      "votes": null
    },
    {
      "id": "481515",
      "postDate": "03/01/2019 13:59:56",
      "content": "<p>There is another topic on this subject. Usually LB scores are 0.5 points below CV scores.\nA CV at 0.54 has a leakage.\nI am at around 2.0 CV and 1.4 to 1.6 LB (CV score is quite stable but LB score looks quite random).</p>",
      "rawMarkdown": "There is another topic on this subject. Usually LB scores are 0.5 points below CV scores.\nA CV at 0.54 has a leakage.\nI am at around 2.0 CV and 1.4 to 1.6 LB (CV score is quite stable but LB score looks quite random).",
      "votes": null
    },
    {
      "id": "481838",
      "postDate": "03/01/2019 22:45:11",
      "content": "<p>1.94 to 2.0CV with 1.432 to 1.6 LB</p>",
      "rawMarkdown": "1.94 to 2.0CV with 1.432 to 1.6 LB",
      "votes": null
    },
    {
      "id": "482007",
      "postDate": "03/02/2019 07:26:42",
      "content": "<p>I'm new to Kaggle, but had a similar experience in the ELO competition.\nI suspect this has a lot to do with the amount of data that is used to evaluate the LB vs your own CV.\nIn this competition, only 13% is used.</p>",
      "rawMarkdown": "I'm new to Kaggle, but had a similar experience in the ELO competition.\nI suspect this has a lot to do with the amount of data that is used to evaluate the LB vs your own CV.\nIn this competition, only 13% is used.",
      "votes": null
    },
    {
      "id": "482019",
      "postDate": "03/02/2019 07:45:54",
      "content": "<p>This competition is strange. I'd carefully checked and ensure that my model does not leak (i.e., train set (including early stopping set) does not overlap with validation set) during the stacking model later on. I got around 1.9x to 2.0x CV for the 1st stage models, and got below 1.8x CV and 1.6x on public LB for the stacked model. Does stacking work? I don't know. If we trust CV until the end we must stay very very low in the public LB :) I'm not sure who will have that courage :)</p>",
      "rawMarkdown": "This competition is strange. I'd carefully checked and ensure that my model does not leak (i.e., train set (including early stopping set) does not overlap with validation set) during the stacking model later on. I got around 1.9x to 2.0x CV for the 1st stage models, and got below 1.8x CV and 1.6x on public LB for the stacked model. Does stacking work? I don't know. If we trust CV until the end we must stay very very low in the public LB :) I'm not sure who will have that courage :)",
      "votes": null
    },
    {
      "id": "482035",
      "postDate": "03/02/2019 08:16:28",
      "content": "<p>Thanks Kha, I'll take this into consideration.</p>",
      "rawMarkdown": "Thanks Kha, I'll take this into consideration.",
      "votes": null
    },
    {
      "id": "486545",
      "postDate": "03/08/2019 23:55:19",
      "content": "<p>my CV is about 2.0 but my LB is 1.42 to 1.8, a little strange to me.</p>",
      "rawMarkdown": "my CV is about 2.0 but my LB is 1.42 to 1.8, a little strange to me.",
      "votes": null
    },
    {
      "id": "486919",
      "postDate": "03/09/2019 17:00:36",
      "content": "<p>I think the reason is very simple - training set consists of 17 recordings and the test - 2600, so it's one of the big chellenges of this conmpetion. All the statictical feature normalization techiques migth be effective. </p>",
      "rawMarkdown": "I think the reason is very simple - training set consists of 17 recordings and the test - 2600, so it's one of the big chellenges of this conmpetion. All the statictical feature normalization techiques migth be effective.",
      "votes": null
    },
    {
      "id": "486951",
      "postDate": "03/09/2019 18:49:36",
      "content": "<p>By any chance, are you sampling with replacement when choosing training segments from the data stream? Segment overlap could be one cause for leakage.</p>",
      "rawMarkdown": "By any chance, are you sampling with replacement when choosing training segments from the data stream? Segment overlap could be one cause for leakage.",
      "votes": null
    },
    {
      "id": "487672",
      "postDate": "03/11/2019 10:09:07",
      "content": "<p>Training sets consists of only a few quakes, but is about 4000+ bins of measurements, Test set is about 2600 bins. Don't confuse number of quakes with number of measured bins.</p>",
      "rawMarkdown": "Training sets consists of only a few quakes, but is about 4000+ bins of measurements, Test set is about 2600 bins. Don't confuse number of quakes with number of measured bins.",
      "votes": null
    },
    {
      "id": "487673",
      "postDate": "03/11/2019 10:10:30",
      "content": "<p>Segment overlap, you mean when quakes occur? I know this would be an issue, but should't be too huge as there are only about 15 corrupted bins that way (of the 4000+)</p>",
      "rawMarkdown": "Segment overlap, you mean when quakes occur? I know this would be an issue, but should't be too huge as there are only about 15 corrupted bins that way (of the 4000+)",
      "votes": null
    },
    {
      "id": "487906",
      "postDate": "03/11/2019 16:19:52",
      "content": "<p>Yes, but the organizers claimed that 2600 test frames were taken from different recordings that are not overlap with training.. So it's a pure statistical challenge.</p>",
      "rawMarkdown": "Yes, but the organizers claimed that 2600 test frames were taken from different recordings that are not overlap with training.. So it's a pure statistical challenge.",
      "votes": null
    },
    {
      "id": "492273",
      "postDate": "03/17/2019 02:24:43",
      "content": "<p>I think by segment overlap, Matei is referring to how the training set is split up. For example, if there are 629,000,000 data points, you can split this up into 4193 non-overlapping files, each containing 150,000 data points, i.e. no file contains the same data point.</p>\n\n<p>However, you can also split up the 629,000,000 data points, into overlapping files. For example, my first file can be from index 0 to index 149,999, my second file can be from 10000 to 159,999. In doing so you will end up with 62,000+ files, with a lot of overlapping data.</p>\n\n<p>You can try this out yourself - you can achieve a very low CV. I can get down to 0.9 CV, but my LB score goes down to 1.73, compared to CV of 2.0 and LB of 1.57 with non-overlapping training data.</p>",
      "rawMarkdown": "I think by segment overlap, Matei is referring to how the training set is split up. For example, if there are 629,000,000 data points, you can split this up into 4193 non-overlapping files, each containing 150,000 data points, i.e. no file contains the same data point.\n\nHowever, you can also split up the 629,000,000 data points, into overlapping files. For example, my first file can be from index 0 to index 149,999, my second file can be from 10000 to 159,999. In doing so you will end up with 62,000+ files, with a lot of overlapping data.\n\nYou can try this out yourself - you can achieve a very low CV. I can get down to 0.9 CV, but my LB score goes down to 1.73, compared to CV of 2.0 and LB of 1.57 with non-overlapping training data.",
      "votes": null
    },
    {
      "id": "492433",
      "postDate": "03/17/2019 09:28:05",
      "content": "<p>Oh yes, that doesn't sound good!\nWould there be any benefits to overlapping these? Or is it always an error to do so?</p>",
      "rawMarkdown": "Oh yes, that doesn't sound good!\nWould there be any benefits to overlapping these? Or is it always an error to do so?",
      "votes": null
    },
    {
      "id": "492463",
      "postDate": "03/17/2019 10:15:22",
      "content": "<p>Not that I can think of. This is what I think would be considered leakage. You would basically be training your model with the duplicate data, whereas the test data you are predicting is highly varied in comparison.</p>\n\n<p>Now imagine using this model to predict the data, and getting feedback on 13% of your predictions. \nI believe this is why there is such a discrepancy between CV and LB, if your training segments overlap.</p>",
      "rawMarkdown": "Not that I can think of. This is what I think would be considered leakage. You would basically be training your model with the duplicate data, whereas the test data you are predicting is highly varied in comparison.\n\nNow imagine using this model to predict the data, and getting feedback on 13% of your predictions. \nI believe this is why there is such a discrepancy between CV and LB, if your training segments overlap.",
      "votes": null
    },
    {
      "id": "492504",
      "postDate": "03/17/2019 11:41:28",
      "content": "<p>It would be a leakage only if you have overlapping segments in different folds (likely to happen if you use shuffle). </p>",
      "rawMarkdown": "It would be a leakage only if you have overlapping segments in different folds (likely to happen if you use shuffle).",
      "votes": null
    },
    {
      "id": "495024",
      "postDate": "03/20/2019 14:51:04",
      "content": "<p>it is really important to trust your CV, because the public data is 13% only. So there are high chances of overfitting.</p>",
      "rawMarkdown": "it is really important to trust your CV, because the public data is 13% only. So there are high chances of overfitting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 481515,
      "author_name": "zidmie",
      "author_url": "",
      "post_date": "03/01/2019 13:59:56",
      "content": "<p>There is another topic on this subject. Usually LB scores are 0.5 points below CV scores.\nA CV at 0.54 has a leakage.\nI am at around 2.0 CV and 1.4 to 1.6 LB (CV score is quite stable but LB score looks quite random).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481838,
      "author_name": "andyatkinson",
      "author_url": "",
      "post_date": "03/01/2019 22:45:11",
      "content": "<p>1.94 to 2.0CV with 1.432 to 1.6 LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 482007,
      "author_name": "vocm88",
      "author_url": "",
      "post_date": "03/02/2019 07:26:42",
      "content": "<p>I'm new to Kaggle, but had a similar experience in the ELO competition.\nI suspect this has a lot to do with the amount of data that is used to evaluate the LB vs your own CV.\nIn this competition, only 13% is used.</p>",
      "votes": null,
      "replies": [
        {
          "id": 482019,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/02/2019 07:45:54",
          "content": "<p>This competition is strange. I'd carefully checked and ensure that my model does not leak (i.e., train set (including early stopping set) does not overlap with validation set) during the stacking model later on. I got around 1.9x to 2.0x CV for the 1st stage models, and got below 1.8x CV and 1.6x on public LB for the stacked model. Does stacking work? I don't know. If we trust CV until the end we must stay very very low in the public LB :) I'm not sure who will have that courage :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 482035,
          "author_name": "vocm88",
          "author_url": "",
          "post_date": "03/02/2019 08:16:28",
          "content": "<p>Thanks Kha, I'll take this into consideration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495024,
          "author_name": "roydatascience",
          "author_url": "",
          "post_date": "03/20/2019 14:51:04",
          "content": "<p>it is really important to trust your CV, because the public data is 13% only. So there are high chances of overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 486545,
      "author_name": "asterisk",
      "author_url": "",
      "post_date": "03/08/2019 23:55:19",
      "content": "<p>my CV is about 2.0 but my LB is 1.42 to 1.8, a little strange to me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 486919,
      "author_name": "",
      "author_url": "",
      "post_date": "03/09/2019 17:00:36",
      "content": "<p>I think the reason is very simple - training set consists of 17 recordings and the test - 2600, so it's one of the big chellenges of this conmpetion. All the statictical feature normalization techiques migth be effective. </p>",
      "votes": null,
      "replies": [
        {
          "id": 487672,
          "author_name": "theupgrade",
          "author_url": "",
          "post_date": "03/11/2019 10:09:07",
          "content": "<p>Training sets consists of only a few quakes, but is about 4000+ bins of measurements, Test set is about 2600 bins. Don't confuse number of quakes with number of measured bins.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 487906,
          "author_name": "",
          "author_url": "",
          "post_date": "03/11/2019 16:19:52",
          "content": "<p>Yes, but the organizers claimed that 2600 test frames were taken from different recordings that are not overlap with training.. So it's a pure statistical challenge.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 486951,
      "author_name": "mateiionita",
      "author_url": "",
      "post_date": "03/09/2019 18:49:36",
      "content": "<p>By any chance, are you sampling with replacement when choosing training segments from the data stream? Segment overlap could be one cause for leakage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 487673,
          "author_name": "theupgrade",
          "author_url": "",
          "post_date": "03/11/2019 10:10:30",
          "content": "<p>Segment overlap, you mean when quakes occur? I know this would be an issue, but should't be too huge as there are only about 15 corrupted bins that way (of the 4000+)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 492273,
          "author_name": "vocm88",
          "author_url": "",
          "post_date": "03/17/2019 02:24:43",
          "content": "<p>I think by segment overlap, Matei is referring to how the training set is split up. For example, if there are 629,000,000 data points, you can split this up into 4193 non-overlapping files, each containing 150,000 data points, i.e. no file contains the same data point.</p>\n\n<p>However, you can also split up the 629,000,000 data points, into overlapping files. For example, my first file can be from index 0 to index 149,999, my second file can be from 10000 to 159,999. In doing so you will end up with 62,000+ files, with a lot of overlapping data.</p>\n\n<p>You can try this out yourself - you can achieve a very low CV. I can get down to 0.9 CV, but my LB score goes down to 1.73, compared to CV of 2.0 and LB of 1.57 with non-overlapping training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 492433,
          "author_name": "theupgrade",
          "author_url": "",
          "post_date": "03/17/2019 09:28:05",
          "content": "<p>Oh yes, that doesn't sound good!\nWould there be any benefits to overlapping these? Or is it always an error to do so?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 492463,
          "author_name": "vocm88",
          "author_url": "",
          "post_date": "03/17/2019 10:15:22",
          "content": "<p>Not that I can think of. This is what I think would be considered leakage. You would basically be training your model with the duplicate data, whereas the test data you are predicting is highly varied in comparison.</p>\n\n<p>Now imagine using this model to predict the data, and getting feedback on 13% of your predictions. \nI believe this is why there is such a discrepancy between CV and LB, if your training segments overlap.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 492504,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "03/17/2019 11:41:28",
          "content": "<p>It would be a leakage only if you have overlapping segments in different folds (likely to happen if you use shuffle). </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "481498": "Hi everyone! I'm opening this topic because the score on cv  with respect to lb score appears quite strange.  My best CV is 0.54 on training and 0.63 on test and 1.74 on LB.  My best LB score instead is 1.55 that has 1.89 on training and 2.19 on test. Are you, Dear Competitors, willing to share your own observations about this difference? \nThanks.",
    "481515": "There is another topic on this subject. Usually LB scores are 0.5 points below CV scores.\nA CV at 0.54 has a leakage.\nI am at around 2.0 CV and 1.4 to 1.6 LB (CV score is quite stable but LB score looks quite random).",
    "481838": "1.94 to 2.0CV with 1.432 to 1.6 LB",
    "482007": "I'm new to Kaggle, but had a similar experience in the ELO competition.\nI suspect this has a lot to do with the amount of data that is used to evaluate the LB vs your own CV.\nIn this competition, only 13% is used.",
    "482019": "This competition is strange. I'd carefully checked and ensure that my model does not leak (i.e., train set (including early stopping set) does not overlap with validation set) during the stacking model later on. I got around 1.9x to 2.0x CV for the 1st stage models, and got below 1.8x CV and 1.6x on public LB for the stacked model. Does stacking work? I don't know. If we trust CV until the end we must stay very very low in the public LB :) I'm not sure who will have that courage :)",
    "482035": "Thanks Kha, I'll take this into consideration.",
    "486545": "my CV is about 2.0 but my LB is 1.42 to 1.8, a little strange to me.",
    "486919": "I think the reason is very simple - training set consists of 17 recordings and the test - 2600, so it's one of the big chellenges of this conmpetion. All the statictical feature normalization techiques migth be effective.",
    "486951": "By any chance, are you sampling with replacement when choosing training segments from the data stream? Segment overlap could be one cause for leakage.",
    "487672": "Training sets consists of only a few quakes, but is about 4000+ bins of measurements, Test set is about 2600 bins. Don't confuse number of quakes with number of measured bins.",
    "487673": "Segment overlap, you mean when quakes occur? I know this would be an issue, but should't be too huge as there are only about 15 corrupted bins that way (of the 4000+)",
    "487906": "Yes, but the organizers claimed that 2600 test frames were taken from different recordings that are not overlap with training.. So it's a pure statistical challenge.",
    "492273": "I think by segment overlap, Matei is referring to how the training set is split up. For example, if there are 629,000,000 data points, you can split this up into 4193 non-overlapping files, each containing 150,000 data points, i.e. no file contains the same data point.\n\nHowever, you can also split up the 629,000,000 data points, into overlapping files. For example, my first file can be from index 0 to index 149,999, my second file can be from 10000 to 159,999. In doing so you will end up with 62,000+ files, with a lot of overlapping data.\n\nYou can try this out yourself - you can achieve a very low CV. I can get down to 0.9 CV, but my LB score goes down to 1.73, compared to CV of 2.0 and LB of 1.57 with non-overlapping training data.",
    "492433": "Oh yes, that doesn't sound good!\nWould there be any benefits to overlapping these? Or is it always an error to do so?",
    "492463": "Not that I can think of. This is what I think would be considered leakage. You would basically be training your model with the duplicate data, whereas the test data you are predicting is highly varied in comparison.\n\nNow imagine using this model to predict the data, and getting feedback on 13% of your predictions. \nI believe this is why there is such a discrepancy between CV and LB, if your training segments overlap.",
    "492504": "It would be a leakage only if you have overlapping segments in different folds (likely to happen if you use shuffle).",
    "495024": "it is really important to trust your CV, because the public data is 13% only. So there are high chances of overfitting."
  },
  "source": "meta"
}