{
  "id": 89496,
  "title": "What metrics would you consider to choose your final submission?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89496",
  "author_name": "",
  "post_date": "2019-04-15T01:54:36.062782900Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Kagglers, as in many other posts here, CV score doesn't correlate to LB score. I guess there is not much we can do in terms of validation scheme - normal 5 fold, stratified 5 fold, or even CV by each earthquake group. </p>\n\n<p>For me, LB score stays at around 1.5 even though my local CV reaches 1.7 after adding new features without using stacking and without any parameter tuning. So I am totally lost as to how I should proceed. So what's your opinion on choosing submission?</p>\n\n<p>In this regard, the following two posts are particularly inspiring.</p>\n\n<p><a href=\"/amjad85\">@amjad85</a> (<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975\">link</a>):</p>\n\n<p><code>\nFrom my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.\n</code></p>\n\n<p><a href=\"/eplistical\">@eplistical</a> (<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/86530#latest-515572\">Link</a>)\n```\nI submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time_to_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time_to_failure &lt; 6.0 (1618 points in total) to be 10000.0</p>\n\n<p>I got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time_to_failure distribution in the test data:</p>\n\n<p>For public data, there are about 83 points (24%) whose time_to_failure is large and 258 points (76%) whose time_to_faliure is small. \nFor private data, there are about 923 points (40%) whose time_to_failure is large and 1360 points (60%) whose time_to_faliure is small.</p>\n\n<p>Since most models have higher accuracy for those data points with small time_to_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time_to_failure data points in the public test data.\n```</p>",
  "messages": [
    {
      "id": "516798",
      "postDate": "04/15/2019 01:54:36",
      "content": "<p>Hi Kagglers, as in many other posts here, CV score doesn't correlate to LB score. I guess there is not much we can do in terms of validation scheme - normal 5 fold, stratified 5 fold, or even CV by each earthquake group. </p>\n\n<p>For me, LB score stays at around 1.5 even though my local CV reaches 1.7 after adding new features without using stacking and without any parameter tuning. So I am totally lost as to how I should proceed. So what's your opinion on choosing submission?</p>\n\n<p>In this regard, the following two posts are particularly inspiring.</p>\n\n<p><a href=\"/amjad85\">@amjad85</a> (<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975\">link</a>):</p>\n\n<p><code>\nFrom my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.\n</code></p>\n\n<p><a href=\"/eplistical\">@eplistical</a> (<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/86530#latest-515572\">Link</a>)\n```\nI submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time_to_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time_to_failure &lt; 6.0 (1618 points in total) to be 10000.0</p>\n\n<p>I got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time_to_failure distribution in the test data:</p>\n\n<p>For public data, there are about 83 points (24%) whose time_to_failure is large and 258 points (76%) whose time_to_faliure is small. \nFor private data, there are about 923 points (40%) whose time_to_failure is large and 1360 points (60%) whose time_to_faliure is small.</p>\n\n<p>Since most models have higher accuracy for those data points with small time_to_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time_to_failure data points in the public test data.\n```</p>",
      "rawMarkdown": "Hi Kagglers, as in many other posts here, CV score doesn't correlate to LB score. I guess there is not much we can do in terms of validation scheme - normal 5 fold, stratified 5 fold, or even CV by each earthquake group. \n\nFor me, LB score stays at around 1.5 even though my local CV reaches 1.7 after adding new features without using stacking and without any parameter tuning. So I am totally lost as to how I should proceed. So what's your opinion on choosing submission?\n\nIn this regard, the following two posts are particularly inspiring.\n\n @amjad85 ([link](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975)):\n\n```\nFrom my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.\n```\n\n@eplistical ([Link](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/86530#latest-515572))\n```\nI submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time_to_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time_to_failure &lt; 6.0 (1618 points in total) to be 10000.0\n\nI got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time_to_failure distribution in the test data:\n\nFor public data, there are about 83 points (24%) whose time_to_failure is large and 258 points (76%) whose time_to_faliure is small. \nFor private data, there are about 923 points (40%) whose time_to_failure is large and 1360 points (60%) whose time_to_faliure is small.\n\nSince most models have higher accuracy for those data points with small time_to_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time_to_failure data points in the public test data.\n```",
      "votes": null
    },
    {
      "id": "516904",
      "postDate": "04/15/2019 07:14:40",
      "content": "<p>Actually, i am new on this competition, and what puzzles me is the validation setup. With a robust validation setup we can choose the best CV submissions.\nConcerning the validation, i cannot figure out if it must be done with simple Kfold or with a GroupKFold on earthquake ID.  It would be nice i could know the number of earthquakes of the test set</p>",
      "rawMarkdown": "Actually, i am new on this competition, and what puzzles me is the validation setup. With a robust validation setup we can choose the best CV submissions.\nConcerning the validation, i cannot figure out if it must be done with simple Kfold or with a GroupKFold on earthquake ID.  It would be nice i could know the number of earthquakes of the test set",
      "votes": null
    },
    {
      "id": "517011",
      "postDate": "04/15/2019 11:20:37",
      "content": "<p>Why choose wisely? Just select one sub with high LB, and other with high CV :) </p>",
      "rawMarkdown": "Why choose wisely? Just select one sub with high LB, and other with high CV :)",
      "votes": null
    },
    {
      "id": "518005",
      "postDate": "04/16/2019 18:47:25",
      "content": "<p>My CV and LB scores have been negatively inverted so far</p>",
      "rawMarkdown": "My CV and LB scores have been negatively inverted so far",
      "votes": null
    },
    {
      "id": "518122",
      "postDate": "04/16/2019 21:53:43",
      "content": "<p>You mean you got a lower LB score for a higher CV score (or vice versa)?</p>",
      "rawMarkdown": "You mean you got a lower LB score for a higher CV score (or vice versa)?",
      "votes": null
    },
    {
      "id": "518149",
      "postDate": "04/16/2019 23:03:45",
      "content": "<p>Thanks for the mention. I'm glad you found what I wrote inspiring. \nI somehow missed the rule about team merging and number of submissions, so I wasn't conservative in submitting solutions. After so many submissions, I think the CV and LB score are correlated, so I would use both CV and LB scores in deciding which submissions to choose.  I wouldn't be too confident about a prediction with low CV but high LB score, since it may indicate lack of generalization to new data with different distribution.</p>",
      "rawMarkdown": "Thanks for the mention. I'm glad you found what I wrote inspiring. \nI somehow missed the rule about team merging and number of submissions, so I wasn't conservative in submitting solutions. After so many submissions, I think the CV and LB score are correlated, so I would use both CV and LB scores in deciding which submissions to choose.  I wouldn't be too confident about a prediction with low CV but high LB score, since it may indicate lack of generalization to new data with different distribution.",
      "votes": null
    },
    {
      "id": "518224",
      "postDate": "04/17/2019 01:41:21",
      "content": "<p><a href=\"/amjad85\">@amjad85</a>, I started to realize in order to make CV and LB in line and trustworthy, you have to make sure your cross-validation scheme is free from mistakes, errors, and leakages.</p>",
      "rawMarkdown": "amjad85, I started to realize in order to make CV and LB in line and trustworthy, you have to make sure your cross-validation scheme is free from mistakes, errors, and leakages.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 516904,
      "author_name": "dkaraflos",
      "author_url": "",
      "post_date": "04/15/2019 07:14:40",
      "content": "<p>Actually, i am new on this competition, and what puzzles me is the validation setup. With a robust validation setup we can choose the best CV submissions.\nConcerning the validation, i cannot figure out if it must be done with simple Kfold or with a GroupKFold on earthquake ID.  It would be nice i could know the number of earthquakes of the test set</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517011,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "04/15/2019 11:20:37",
      "content": "<p>Why choose wisely? Just select one sub with high LB, and other with high CV :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518005,
      "author_name": "halldalton94",
      "author_url": "",
      "post_date": "04/16/2019 18:47:25",
      "content": "<p>My CV and LB scores have been negatively inverted so far</p>",
      "votes": null,
      "replies": [
        {
          "id": 518122,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/16/2019 21:53:43",
          "content": "<p>You mean you got a lower LB score for a higher CV score (or vice versa)?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 518149,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "04/16/2019 23:03:45",
      "content": "<p>Thanks for the mention. I'm glad you found what I wrote inspiring. \nI somehow missed the rule about team merging and number of submissions, so I wasn't conservative in submitting solutions. After so many submissions, I think the CV and LB score are correlated, so I would use both CV and LB scores in deciding which submissions to choose.  I wouldn't be too confident about a prediction with low CV but high LB score, since it may indicate lack of generalization to new data with different distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 518224,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/17/2019 01:41:21",
          "content": "<p><a href=\"/amjad85\">@amjad85</a>, I started to realize in order to make CV and LB in line and trustworthy, you have to make sure your cross-validation scheme is free from mistakes, errors, and leakages.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "516798": "Hi Kagglers, as in many other posts here, CV score doesn't correlate to LB score. I guess there is not much we can do in terms of validation scheme - normal 5 fold, stratified 5 fold, or even CV by each earthquake group. \n\nFor me, LB score stays at around 1.5 even though my local CV reaches 1.7 after adding new features without using stacking and without any parameter tuning. So I am totally lost as to how I should proceed. So what's your opinion on choosing submission?\n\nIn this regard, the following two posts are particularly inspiring.\n\n @amjad85 ([link](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89301#515975)):\n\n```\nFrom my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.\n```\n\n@eplistical ([Link](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/86530#latest-515572))\n```\nI submitted a simple model (~1.52 LB) with the following modifications:\n1) Force all predicted data points with time_to_failure &gt; 6.0 (1006 points in total) to be 10000.0\n2) Force all predicted data points with time_to_failure &lt; 6.0 (1618 points in total) to be 10000.0\n\nI got MAE ~2442 for (1) and ~7556 for (2).\nFrom these results we can approximate the time_to_failure distribution in the test data:\n\nFor public data, there are about 83 points (24%) whose time_to_failure is large and 258 points (76%) whose time_to_faliure is small. \nFor private data, there are about 923 points (40%) whose time_to_failure is large and 1360 points (60%) whose time_to_faliure is small.\n\nSince most models have higher accuracy for those data points with small time_to_failure (as I observed), the above distribution partially explains why people get lower LB error than their CV: There are more small time_to_failure data points in the public test data.\n```",
    "516904": "Actually, i am new on this competition, and what puzzles me is the validation setup. With a robust validation setup we can choose the best CV submissions.\nConcerning the validation, i cannot figure out if it must be done with simple Kfold or with a GroupKFold on earthquake ID.  It would be nice i could know the number of earthquakes of the test set",
    "517011": "Why choose wisely? Just select one sub with high LB, and other with high CV :)",
    "518005": "My CV and LB scores have been negatively inverted so far",
    "518122": "You mean you got a lower LB score for a higher CV score (or vice versa)?",
    "518149": "Thanks for the mention. I'm glad you found what I wrote inspiring. \nI somehow missed the rule about team merging and number of submissions, so I wasn't conservative in submitting solutions. After so many submissions, I think the CV and LB score are correlated, so I would use both CV and LB scores in deciding which submissions to choose.  I wouldn't be too confident about a prediction with low CV but high LB score, since it may indicate lack of generalization to new data with different distribution.",
    "518224": "amjad85, I started to realize in order to make CV and LB in line and trustworthy, you have to make sure your cross-validation scheme is free from mistakes, errors, and leakages."
  },
  "source": "meta"
}