{
  "id": 87157,
  "title": "Can you share your cv and lb? and lb & cv uncoincident",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/87157",
  "author_name": "",
  "post_date": "2019-03-29T06:24:31.873222Z",
  "votes": 7,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi all:\nCurrently:\nmy single model result:\ncv 2.08, lb 1.504 (with very few features)\ncv 2.02, lb 1.535 (with many features from public)\nBTW, i can say almost sure there is no data leak in the cv 2.02 version.\nCan you share your result and any idea about why this cv and lb uncoincident? thank you very much!</p>",
  "messages": [
    {
      "id": "502853",
      "postDate": "03/29/2019 06:24:31",
      "content": "<p>Hi all:\nCurrently:\nmy single model result:\ncv 2.08, lb 1.504 (with very few features)\ncv 2.02, lb 1.535 (with many features from public)\nBTW, i can say almost sure there is no data leak in the cv 2.02 version.\nCan you share your result and any idea about why this cv and lb uncoincident? thank you very much!</p>",
      "rawMarkdown": "Hi all:\nCurrently:\nmy single model result:\ncv 2.08, lb 1.504 (with very few features)\ncv 2.02, lb 1.535 (with many features from public)\nBTW, i can say almost sure there is no data leak in the cv 2.02 version.\nCan you share your result and any idea about why this cv and lb uncoincident? thank you very much!",
      "votes": null
    },
    {
      "id": "502925",
      "postDate": "03/29/2019 09:03:59",
      "content": "<p>As it was shown in the discussions, the PL is calculated with an \"easy\" part of the test data, so the lb is better than cv. I would say trust your cv as the LB is calculated on only 13% of the test data.\nmy best model has 1.73 cv 1.623 lb, it's also confusing for me as I see most people having ~2.0+ cv and 1.4-1.5 lb</p>",
      "rawMarkdown": "As it was shown in the discussions, the PL is calculated with an \"easy\" part of the test data, so the lb is better than cv. I would say trust your cv as the LB is calculated on only 13% of the test data.\nmy best model has 1.73 cv 1.623 lb, it's also confusing for me as I see most people having ~2.0+ cv and 1.4-1.5 lb",
      "votes": null
    },
    {
      "id": "503025",
      "postDate": "03/29/2019 11:53:46",
      "content": "<p>1.73 CV is amazing, I'd be much happier with that than a high LB!  Are you validating by earthquake, or with random 150,000-row segments?</p>",
      "rawMarkdown": "1.73 CV is amazing, I'd be much happier with that than a high LB!  Are you validating by earthquake, or with random 150,000-row segments?",
      "votes": null
    },
    {
      "id": "503029",
      "postDate": "03/29/2019 12:02:49",
      "content": "<p>Thank you , Yep, we should trust cv. Your 1.73 cv is  amazing!</p>",
      "rawMarkdown": "Thank you , Yep, we should trust cv. Your 1.73 cv is  amazing!",
      "votes": null
    },
    {
      "id": "503042",
      "postDate": "03/29/2019 12:18:10",
      "content": "<p>earthquake by earthquake</p>",
      "rawMarkdown": "earthquake by earthquake",
      "votes": null
    },
    {
      "id": "503340",
      "postDate": "03/29/2019 20:47:04",
      "content": "<p>As I believe others have mentioned, the LB is based off of earthquakes with relatively low TTFs, which most models have, at least from my experience, been able to predict pretty well. It's the earthquakes with long TTFs that the models tended to have struggled. </p>\n\n<p>So your CV is taking into account both scenarios whereas the LB, at least from what it seems, is not. </p>",
      "rawMarkdown": "As I believe others have mentioned, the LB is based off of earthquakes with relatively low TTFs, which most models have, at least from my experience, been able to predict pretty well. It's the earthquakes with long TTFs that the models tended to have struggled. \n\nSo your CV is taking into account both scenarios whereas the LB, at least from what it seems, is not.",
      "votes": null
    },
    {
      "id": "503456",
      "postDate": "03/30/2019 02:34:53",
      "content": "<p>You mean the lb data have relatively low TTFs, but trian data have both high TTFs and low TTFs, Am i right? If so do you have ideas about  just public lb has relatively low FFTs or both public lb and test lb have relatively low FFTs, Thank you very much!</p>",
      "rawMarkdown": "You mean the lb data have relatively low TTFs, but trian data have both high TTFs and low TTFs, Am i right? If so do you have ideas about  just public lb has relatively low FFTs or both public lb and test lb have relatively low FFTs, Thank you very much!",
      "votes": null
    },
    {
      "id": "503462",
      "postDate": "03/30/2019 02:42:32",
      "content": "<p>Yes, the train data has a mixture of high and low TTFs. </p>\n\n<p>What some others have discovered (by outputting essentially noise) in other discussion threads is that the public leaderboard has a significantly higher proportion of low TTFs than our training data.</p>\n\n<p>So for example, suppose that 20% of our training data has low TTFs (let's say TTF &lt; 3). The public leaderboard is more around 40-50%.</p>\n\n<p>However, we don't know if this also applies to the private leaderboard. I mean, how could we? It could be just as generous with low TTFs, it could be much harsher. </p>\n\n<p>That's why others have advised using the model with the best CV score, as it (should) take into account performance on both low and high TTFs. </p>",
      "rawMarkdown": "Yes, the train data has a mixture of high and low TTFs. \n\nWhat some others have discovered (by outputting essentially noise) in other discussion threads is that the public leaderboard has a significantly higher proportion of low TTFs than our training data.\n\nSo for example, suppose that 20% of our training data has low TTFs (let's say TTF &lt; 3). The public leaderboard is more around 40-50%.\n\nHowever, we don't know if this also applies to the private leaderboard. I mean, how could we? It could be just as generous with low TTFs, it could be much harsher. \n\nThat's why others have advised using the model with the best CV score, as it (should) take into account performance on both low and high TTFs.",
      "votes": null
    },
    {
      "id": "503467",
      "postDate": "03/30/2019 03:02:47",
      "content": "<p>Yes , we should trust cv, thank you very much for this usefull information for me(i joined this competition 2 days ago)</p>",
      "rawMarkdown": "Yes , we should trust cv, thank you very much for this usefull information for me(i joined this competition 2 days ago)",
      "votes": null
    },
    {
      "id": "504081",
      "postDate": "03/31/2019 00:02:12",
      "content": "<p>I got 1.801 cv and 1.713 lb....so frustrating</p>",
      "rawMarkdown": "I got 1.801 cv and 1.713 lb....so frustrating",
      "votes": null
    },
    {
      "id": "504555",
      "postDate": "03/31/2019 20:08:23",
      "content": "<p>Here's my take on it:</p>\n\n<p>According to Bertrand RL's \"Additional Info\" post, \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time\".</p>\n\n<p>So presumably the test set, taken as a whole, has the same TTF distribution as the training set.  The 87% of the test set that's used for the public LB seems to have lower TTF's than the training set, so the 13% that will be used for the private leaderboard has to have higher TTF's on average.</p>\n\n<p>I'm finding that the distributions of my features are pretty nearly identical between the training set and the test set, which I take as support of this hypothesis.</p>",
      "rawMarkdown": "Here's my take on it:\n\nAccording to Bertrand RL's \"Additional Info\" post, \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time\".\n\nSo presumably the test set, taken as a whole, has the same TTF distribution as the training set.  The 87% of the test set that's used for the public LB seems to have lower TTF's than the training set, so the 13% that will be used for the private leaderboard has to have higher TTF's on average.\n\n I'm finding that the distributions of my features are pretty nearly identical between the training set and the test set, which I take as support of this hypothesis.",
      "votes": null
    },
    {
      "id": "504762",
      "postDate": "04/01/2019 06:01:24",
      "content": "<p>Good!</p>",
      "rawMarkdown": "Good!",
      "votes": null
    },
    {
      "id": "505091",
      "postDate": "04/01/2019 14:48:41",
      "content": "<p>Running adversarial validation using basic features I get a binary logloss of ~0.47, indicating there isn't much separating train and test apart from TTF. This shouldn't come as much of a surprise since train and test both come from the same experiment but it's nice to get confirmation. </p>",
      "rawMarkdown": "Running adversarial validation using basic features I get a binary logloss of ~0.47, indicating there isn't much separating train and test apart from TTF. This shouldn't come as much of a surprise since train and test both come from the same experiment but it's nice to get confirmation.",
      "votes": null
    },
    {
      "id": "505707",
      "postDate": "04/02/2019 12:43:45",
      "content": "<p>Actually, it seems very interesting. How did you manage to have so small gap between cv and lb?</p>",
      "rawMarkdown": "Actually, it seems very interesting. How did you manage to have so small gap between cv and lb?",
      "votes": null
    },
    {
      "id": "505716",
      "postDate": "04/02/2019 13:03:57",
      "content": "<p>The weird thing is the more features I have, the higher (worse) LB score I get. When I look at the min/max value of the submission, the range is narrower for the model with more features. All I can say is something is wrong. </p>",
      "rawMarkdown": "The weird thing is the more features I have, the higher (worse) LB score I get. When I look at the min/max value of the submission, the range is narrower for the model with more features. All I can say is something is wrong.",
      "votes": null
    },
    {
      "id": "505763",
      "postDate": "04/02/2019 13:56:41",
      "content": "<p>Ordinarily I'd say that is just overfitting. But apparently, the public test set has less variance in time_to_failure than the train set, so perhaps your model is generalising better but is less accurate for simpler cases as a result. I'm wondering if the best approach for this competition is to have multiple models for segments that appear to have low-TTF, and segments that appear to have high-TTF.</p>",
      "rawMarkdown": "Ordinarily I'd say that is just overfitting. But apparently, the public test set has less variance in time_to_failure than the train set, so perhaps your model is generalising better but is less accurate for simpler cases as a result. I'm wondering if the best approach for this competition is to have multiple models for segments that appear to have low-TTF, and segments that appear to have high-TTF.",
      "votes": null
    },
    {
      "id": "505819",
      "postDate": "04/02/2019 15:14:40",
      "content": "<p>I would say until the result from private LB is out, I can't be sure whether it's overfitting or generalizing well. Thanks for your feedback though. </p>",
      "rawMarkdown": "I would say until the result from private LB is out, I can't be sure whether it's overfitting or generalizing well. Thanks for your feedback though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 502925,
      "author_name": "kerzer",
      "author_url": "",
      "post_date": "03/29/2019 09:03:59",
      "content": "<p>As it was shown in the discussions, the PL is calculated with an \"easy\" part of the test data, so the lb is better than cv. I would say trust your cv as the LB is calculated on only 13% of the test data.\nmy best model has 1.73 cv 1.623 lb, it's also confusing for me as I see most people having ~2.0+ cv and 1.4-1.5 lb</p>",
      "votes": null,
      "replies": [
        {
          "id": 503025,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "03/29/2019 11:53:46",
          "content": "<p>1.73 CV is amazing, I'd be much happier with that than a high LB!  Are you validating by earthquake, or with random 150,000-row segments?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503029,
          "author_name": "codlife",
          "author_url": "",
          "post_date": "03/29/2019 12:02:49",
          "content": "<p>Thank you , Yep, we should trust cv. Your 1.73 cv is  amazing!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503042,
          "author_name": "kerzer",
          "author_url": "",
          "post_date": "03/29/2019 12:18:10",
          "content": "<p>earthquake by earthquake</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 503340,
      "author_name": "conormcnamara",
      "author_url": "",
      "post_date": "03/29/2019 20:47:04",
      "content": "<p>As I believe others have mentioned, the LB is based off of earthquakes with relatively low TTFs, which most models have, at least from my experience, been able to predict pretty well. It's the earthquakes with long TTFs that the models tended to have struggled. </p>\n\n<p>So your CV is taking into account both scenarios whereas the LB, at least from what it seems, is not. </p>",
      "votes": null,
      "replies": [
        {
          "id": 503456,
          "author_name": "codlife",
          "author_url": "",
          "post_date": "03/30/2019 02:34:53",
          "content": "<p>You mean the lb data have relatively low TTFs, but trian data have both high TTFs and low TTFs, Am i right? If so do you have ideas about  just public lb has relatively low FFTs or both public lb and test lb have relatively low FFTs, Thank you very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503462,
          "author_name": "conormcnamara",
          "author_url": "",
          "post_date": "03/30/2019 02:42:32",
          "content": "<p>Yes, the train data has a mixture of high and low TTFs. </p>\n\n<p>What some others have discovered (by outputting essentially noise) in other discussion threads is that the public leaderboard has a significantly higher proportion of low TTFs than our training data.</p>\n\n<p>So for example, suppose that 20% of our training data has low TTFs (let's say TTF &lt; 3). The public leaderboard is more around 40-50%.</p>\n\n<p>However, we don't know if this also applies to the private leaderboard. I mean, how could we? It could be just as generous with low TTFs, it could be much harsher. </p>\n\n<p>That's why others have advised using the model with the best CV score, as it (should) take into account performance on both low and high TTFs. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503467,
          "author_name": "codlife",
          "author_url": "",
          "post_date": "03/30/2019 03:02:47",
          "content": "<p>Yes , we should trust cv, thank you very much for this usefull information for me(i joined this competition 2 days ago)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 504081,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "03/31/2019 00:02:12",
      "content": "<p>I got 1.801 cv and 1.713 lb....so frustrating</p>",
      "votes": null,
      "replies": [
        {
          "id": 504762,
          "author_name": "codlife",
          "author_url": "",
          "post_date": "04/01/2019 06:01:24",
          "content": "<p>Good!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505707,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "04/02/2019 12:43:45",
          "content": "<p>Actually, it seems very interesting. How did you manage to have so small gap between cv and lb?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505716,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/02/2019 13:03:57",
          "content": "<p>The weird thing is the more features I have, the higher (worse) LB score I get. When I look at the min/max value of the submission, the range is narrower for the model with more features. All I can say is something is wrong. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505763,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/02/2019 13:56:41",
          "content": "<p>Ordinarily I'd say that is just overfitting. But apparently, the public test set has less variance in time_to_failure than the train set, so perhaps your model is generalising better but is less accurate for simpler cases as a result. I'm wondering if the best approach for this competition is to have multiple models for segments that appear to have low-TTF, and segments that appear to have high-TTF.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 505819,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/02/2019 15:14:40",
          "content": "<p>I would say until the result from private LB is out, I can't be sure whether it's overfitting or generalizing well. Thanks for your feedback though. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 504555,
      "author_name": "vgates",
      "author_url": "",
      "post_date": "03/31/2019 20:08:23",
      "content": "<p>Here's my take on it:</p>\n\n<p>According to Bertrand RL's \"Additional Info\" post, \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time\".</p>\n\n<p>So presumably the test set, taken as a whole, has the same TTF distribution as the training set.  The 87% of the test set that's used for the public LB seems to have lower TTF's than the training set, so the 13% that will be used for the private leaderboard has to have higher TTF's on average.</p>\n\n<p>I'm finding that the distributions of my features are pretty nearly identical between the training set and the test set, which I take as support of this hypothesis.</p>",
      "votes": null,
      "replies": [
        {
          "id": 505091,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/01/2019 14:48:41",
          "content": "<p>Running adversarial validation using basic features I get a binary logloss of ~0.47, indicating there isn't much separating train and test apart from TTF. This shouldn't come as much of a surprise since train and test both come from the same experiment but it's nice to get confirmation. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "502853": "Hi all:\nCurrently:\nmy single model result:\ncv 2.08, lb 1.504 (with very few features)\ncv 2.02, lb 1.535 (with many features from public)\nBTW, i can say almost sure there is no data leak in the cv 2.02 version.\nCan you share your result and any idea about why this cv and lb uncoincident? thank you very much!",
    "502925": "As it was shown in the discussions, the PL is calculated with an \"easy\" part of the test data, so the lb is better than cv. I would say trust your cv as the LB is calculated on only 13% of the test data.\nmy best model has 1.73 cv 1.623 lb, it's also confusing for me as I see most people having ~2.0+ cv and 1.4-1.5 lb",
    "503025": "1.73 CV is amazing, I'd be much happier with that than a high LB!  Are you validating by earthquake, or with random 150,000-row segments?",
    "503029": "Thank you , Yep, we should trust cv. Your 1.73 cv is  amazing!",
    "503042": "earthquake by earthquake",
    "503340": "As I believe others have mentioned, the LB is based off of earthquakes with relatively low TTFs, which most models have, at least from my experience, been able to predict pretty well. It's the earthquakes with long TTFs that the models tended to have struggled. \n\nSo your CV is taking into account both scenarios whereas the LB, at least from what it seems, is not.",
    "503456": "You mean the lb data have relatively low TTFs, but trian data have both high TTFs and low TTFs, Am i right? If so do you have ideas about  just public lb has relatively low FFTs or both public lb and test lb have relatively low FFTs, Thank you very much!",
    "503462": "Yes, the train data has a mixture of high and low TTFs. \n\nWhat some others have discovered (by outputting essentially noise) in other discussion threads is that the public leaderboard has a significantly higher proportion of low TTFs than our training data.\n\nSo for example, suppose that 20% of our training data has low TTFs (let's say TTF &lt; 3). The public leaderboard is more around 40-50%.\n\nHowever, we don't know if this also applies to the private leaderboard. I mean, how could we? It could be just as generous with low TTFs, it could be much harsher. \n\nThat's why others have advised using the model with the best CV score, as it (should) take into account performance on both low and high TTFs.",
    "503467": "Yes , we should trust cv, thank you very much for this usefull information for me(i joined this competition 2 days ago)",
    "504081": "I got 1.801 cv and 1.713 lb....so frustrating",
    "504555": "Here's my take on it:\n\nAccording to Bertrand RL's \"Additional Info\" post, \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time\".\n\nSo presumably the test set, taken as a whole, has the same TTF distribution as the training set.  The 87% of the test set that's used for the public LB seems to have lower TTF's than the training set, so the 13% that will be used for the private leaderboard has to have higher TTF's on average.\n\n I'm finding that the distributions of my features are pretty nearly identical between the training set and the test set, which I take as support of this hypothesis.",
    "504762": "Good!",
    "505091": "Running adversarial validation using basic features I get a binary logloss of ~0.47, indicating there isn't much separating train and test apart from TTF. This shouldn't come as much of a surprise since train and test both come from the same experiment but it's nice to get confirmation.",
    "505707": "Actually, it seems very interesting. How did you manage to have so small gap between cv and lb?",
    "505716": "The weird thing is the more features I have, the higher (worse) LB score I get. When I look at the min/max value of the submission, the range is narrower for the model with more features. All I can say is something is wrong.",
    "505763": "Ordinarily I'd say that is just overfitting. But apparently, the public test set has less variance in time_to_failure than the train set, so perhaps your model is generalising better but is less accurate for simpler cases as a result. I'm wondering if the best approach for this competition is to have multiple models for segments that appear to have low-TTF, and segments that appear to have high-TTF.",
    "505819": "I would say until the result from private LB is out, I can't be sure whether it's overfitting or generalizing well. Thanks for your feedback though."
  },
  "source": "meta"
}