{
  "id": 89301,
  "title": "Different folds result in very different LB scores",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89301",
  "author_name": "",
  "post_date": "2019-04-12T19:46:34.337819Z",
  "votes": 21,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Lots of people have posted on the pitfalls of leaderboard in this competition, but just wanted to share my experience to illustrate just how unreliable the leaderboard is. I split the data into 5 folds. My model takes a while to train and since I'm impatient I've been submitting results trained on a single fold. The first fold resulted in an LB score of 1.351 (yay!). However, training the exact same model on a different fold results in an LB score of 1.452.</p>\n\n<p>There have been some other posts about how the subset used to calculate the LB consists of a disproportionate number of samples with low time to failure. I calculated the mean TTF predicted by my model on the test set:</p>\n\n<p>fold 1: 5.031828\nfold 2: 5.233134</p>\n\n<p>The first fold seems to be biased towards a shorter TTF for whatever reason. I don't have any reason to believe fold 1 will be any better when evaluated on the entire test set even though it's a lot better on the LB. In case anyone is wondering I got oof scores of ~1.86 and ~1.97. Anyway hope this helps make some sense of the LB scores.</p>",
  "messages": [
    {
      "id": "515585",
      "postDate": "04/12/2019 19:46:34",
      "content": "<p>Lots of people have posted on the pitfalls of leaderboard in this competition, but just wanted to share my experience to illustrate just how unreliable the leaderboard is. I split the data into 5 folds. My model takes a while to train and since I'm impatient I've been submitting results trained on a single fold. The first fold resulted in an LB score of 1.351 (yay!). However, training the exact same model on a different fold results in an LB score of 1.452.</p>\n\n<p>There have been some other posts about how the subset used to calculate the LB consists of a disproportionate number of samples with low time to failure. I calculated the mean TTF predicted by my model on the test set:</p>\n\n<p>fold 1: 5.031828\nfold 2: 5.233134</p>\n\n<p>The first fold seems to be biased towards a shorter TTF for whatever reason. I don't have any reason to believe fold 1 will be any better when evaluated on the entire test set even though it's a lot better on the LB. In case anyone is wondering I got oof scores of ~1.86 and ~1.97. Anyway hope this helps make some sense of the LB scores.</p>",
      "rawMarkdown": "Lots of people have posted on the pitfalls of leaderboard in this competition, but just wanted to share my experience to illustrate just how unreliable the leaderboard is. I split the data into 5 folds. My model takes a while to train and since I'm impatient I've been submitting results trained on a single fold. The first fold resulted in an LB score of 1.351 (yay!). However, training the exact same model on a different fold results in an LB score of 1.452.\n\nThere have been some other posts about how the subset used to calculate the LB consists of a disproportionate number of samples with low time to failure. I calculated the mean TTF predicted by my model on the test set:\n\nfold 1: 5.031828\nfold 2: 5.233134\n\nThe first fold seems to be biased towards a shorter TTF for whatever reason. I don't have any reason to believe fold 1 will be any better when evaluated on the entire test set even though it's a lot better on the LB. In case anyone is wondering I got oof scores of ~1.86 and ~1.97. Anyway hope this helps make some sense of the LB scores.",
      "votes": null
    },
    {
      "id": "515637",
      "postDate": "04/12/2019 21:48:33",
      "content": "<p>How are you splitting your folds?</p>",
      "rawMarkdown": "How are you splitting your folds?",
      "votes": null
    },
    {
      "id": "515770",
      "postDate": "04/13/2019 05:25:09",
      "content": "<p>I follow.\nI am using quake-based split and I am nowhere near MAE ~ 2 🤔</p>",
      "rawMarkdown": "I follow.\nI am using quake-based split and I am nowhere near MAE ~ 2 🤔",
      "votes": null
    },
    {
      "id": "515975",
      "postDate": "04/13/2019 13:10:37",
      "content": "<p>Thank you for sharing this info.  The organizers have kept the design of data splitting a secret, but the most likely scenario is that the public LB is based only on the first 1-2 cycles/earthquakes in the test data, and the private LB comprises completely different cycles/earthquakes. So if your fold has a majority samples from similar cycles to public LB, your score will be better.</p>\n\n<p>Many have pointed out that the models tend to have worse prediction on segments with high TTF values, but I have a different explanation. \nAt the beginning of each cycle, the signal properties of all cycles are the same. However, in some cycles, the signal escalates/intensifies quickly leading to a short cycle. In other cycles, the signal escalates in a slower pace, or encounter mini-quakes that significantly decrease the intensity, ultimately leading to a long cycle. If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%.  There seem to be no features in the signal that inform which behavior the cycle will take except the \"mean\" of the signal. But so far, I'm not sure if the mean of the signal is physically important feature or batch-effect noise that should be eliminated by centering or high-pass filtering.</p>\n\n<p>From my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.</p>",
      "rawMarkdown": "Thank you for sharing this info.  The organizers have kept the design of data splitting a secret, but the most likely scenario is that the public LB is based only on the first 1-2 cycles/earthquakes in the test data, and the private LB comprises completely different cycles/earthquakes. So if your fold has a majority samples from similar cycles to public LB, your score will be better.\n\nMany have pointed out that the models tend to have worse prediction on segments with high TTF values, but I have a different explanation. \nAt the beginning of each cycle, the signal properties of all cycles are the same. However, in some cycles, the signal escalates/intensifies quickly leading to a short cycle. In other cycles, the signal escalates in a slower pace, or encounter mini-quakes that significantly decrease the intensity, ultimately leading to a long cycle. If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%.  There seem to be no features in the signal that inform which behavior the cycle will take except the \"mean\" of the signal. But so far, I'm not sure if the mean of the signal is physically important feature or batch-effect noise that should be eliminated by centering or high-pass filtering.\n\nFrom my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.",
      "votes": null
    },
    {
      "id": "517373",
      "postDate": "04/16/2019 00:53:27",
      "content": "<p>Some of my submissions which score around 1.410 have mean around 5.57-5.59. Others which score around 1.440 have mean around 5.35-5.50. My current best single model CV has (test) mean 5.610 which scores 1.413.</p>\n\n<p>To me, higher mean of predictions means that the model leans towards higher ttf. Given 2 submissions with same CV, same LB, the one with larger mean may perform better on private LB.</p>\n\n<p>Anyone with this kind of information, please share so all of us can have some insights?</p>",
      "rawMarkdown": "Some of my submissions which score around 1.410 have mean around 5.57-5.59. Others which score around 1.440 have mean around 5.35-5.50. My current best single model CV has (test) mean 5.610 which scores 1.413.\n\nTo me, higher mean of predictions means that the model leans towards higher ttf. Given 2 submissions with same CV, same LB, the one with larger mean may perform better on private LB.\n\nAnyone with this kind of information, please share so all of us can have some insights?",
      "votes": null
    },
    {
      "id": "522103",
      "postDate": "04/23/2019 21:31:17",
      "content": "<p>This seems very high to me. Mine is under 5.2</p>",
      "rawMarkdown": "This seems very high to me. Mine is under 5.2",
      "votes": null
    },
    {
      "id": "523929",
      "postDate": "04/27/2019 12:47:46",
      "content": "<p>what is your quake-based split means? </p>",
      "rawMarkdown": "what is your quake-based split means?",
      "votes": null
    },
    {
      "id": "525626",
      "postDate": "05/01/2019 12:22:45",
      "content": "<p><a href=\"/amjad85\">@amjad85</a> How did you come up with that estimate? </p>\n\n<p>\"If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%\"</p>",
      "rawMarkdown": "amjad85 How did you come up with that estimate? \n\n\"If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%\"",
      "votes": null
    },
    {
      "id": "525635",
      "postDate": "05/01/2019 12:45:30",
      "content": "<p>I tried it. Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement, but obviously not so useful for the test set. </p>",
      "rawMarkdown": "I tried it. Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement, but obviously not so useful for the test set.",
      "votes": null
    },
    {
      "id": "525638",
      "postDate": "05/01/2019 12:46:37",
      "content": "<p>I see. Thanks! </p>",
      "rawMarkdown": "I see. Thanks!",
      "votes": null
    },
    {
      "id": "526090",
      "postDate": "05/02/2019 10:27:10",
      "content": "<blockquote>\n  <p>Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement,</p>\n</blockquote>\n\n<p>If it does improve your cv score then your cv setting is flawed as it fails to detect this over fitting.</p>",
      "rawMarkdown": "&gt; Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement,\n\nIf it does improve your cv score then your cv setting is flawed as it fails to detect this over fitting.",
      "votes": null
    },
    {
      "id": "526176",
      "postDate": "05/02/2019 14:06:37",
      "content": "<p>Yes I was shuffling when I tried it.</p>",
      "rawMarkdown": "Yes I was shuffling when I tried it.",
      "votes": null
    },
    {
      "id": "529056",
      "postDate": "05/09/2019 05:15:07",
      "content": "<p>Hi there, when you guys talk about \"mean value\", are you referring to mean of predicted TTF of all test samples?</p>",
      "rawMarkdown": "Hi there, when you guys talk about \"mean value\", are you referring to mean of predicted TTF of all test samples?",
      "votes": null
    },
    {
      "id": "529057",
      "postDate": "05/09/2019 05:15:41",
      "content": "<p><a href=\"/khahuras\">@khahuras</a> <a href=\"/areveillon\">@areveillon</a>, forgot to @ you two</p>",
      "rawMarkdown": "khahuras @areveillon, forgot to @ you two",
      "votes": null
    },
    {
      "id": "529067",
      "postDate": "05/09/2019 05:33:46",
      "content": "<p>Yes it is the mean value of predicted test samples.</p>",
      "rawMarkdown": "Yes it is the mean value of predicted test samples.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 515637,
      "author_name": "halldalton94",
      "author_url": "",
      "post_date": "04/12/2019 21:48:33",
      "content": "<p>How are you splitting your folds?</p>",
      "votes": null,
      "replies": [
        {
          "id": 515770,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "04/13/2019 05:25:09",
          "content": "<p>I follow.\nI am using quake-based split and I am nowhere near MAE ~ 2 🤔</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523929,
          "author_name": "supportvectordevin",
          "author_url": "",
          "post_date": "04/27/2019 12:47:46",
          "content": "<p>what is your quake-based split means? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 515975,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "04/13/2019 13:10:37",
      "content": "<p>Thank you for sharing this info.  The organizers have kept the design of data splitting a secret, but the most likely scenario is that the public LB is based only on the first 1-2 cycles/earthquakes in the test data, and the private LB comprises completely different cycles/earthquakes. So if your fold has a majority samples from similar cycles to public LB, your score will be better.</p>\n\n<p>Many have pointed out that the models tend to have worse prediction on segments with high TTF values, but I have a different explanation. \nAt the beginning of each cycle, the signal properties of all cycles are the same. However, in some cycles, the signal escalates/intensifies quickly leading to a short cycle. In other cycles, the signal escalates in a slower pace, or encounter mini-quakes that significantly decrease the intensity, ultimately leading to a long cycle. If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%.  There seem to be no features in the signal that inform which behavior the cycle will take except the \"mean\" of the signal. But so far, I'm not sure if the mean of the signal is physically important feature or batch-effect noise that should be eliminated by centering or high-pass filtering.</p>\n\n<p>From my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525626,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "05/01/2019 12:22:45",
          "content": "<p><a href=\"/amjad85\">@amjad85</a> How did you come up with that estimate? </p>\n\n<p>\"If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525635,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/01/2019 12:45:30",
          "content": "<p>I tried it. Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement, but obviously not so useful for the test set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525638,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "05/01/2019 12:46:37",
          "content": "<p>I see. Thanks! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526090,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/02/2019 10:27:10",
          "content": "<blockquote>\n  <p>Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement,</p>\n</blockquote>\n\n<p>If it does improve your cv score then your cv setting is flawed as it fails to detect this over fitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526176,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "05/02/2019 14:06:37",
          "content": "<p>Yes I was shuffling when I tried it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517373,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "04/16/2019 00:53:27",
      "content": "<p>Some of my submissions which score around 1.410 have mean around 5.57-5.59. Others which score around 1.440 have mean around 5.35-5.50. My current best single model CV has (test) mean 5.610 which scores 1.413.</p>\n\n<p>To me, higher mean of predictions means that the model leans towards higher ttf. Given 2 submissions with same CV, same LB, the one with larger mean may perform better on private LB.</p>\n\n<p>Anyone with this kind of information, please share so all of us can have some insights?</p>",
      "votes": null,
      "replies": [
        {
          "id": 522103,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "04/23/2019 21:31:17",
          "content": "<p>This seems very high to me. Mine is under 5.2</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529056,
          "author_name": "",
          "author_url": "",
          "post_date": "05/09/2019 05:15:07",
          "content": "<p>Hi there, when you guys talk about \"mean value\", are you referring to mean of predicted TTF of all test samples?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529057,
          "author_name": "",
          "author_url": "",
          "post_date": "05/09/2019 05:15:41",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> <a href=\"/areveillon\">@areveillon</a>, forgot to @ you two</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529067,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "05/09/2019 05:33:46",
          "content": "<p>Yes it is the mean value of predicted test samples.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "515585": "Lots of people have posted on the pitfalls of leaderboard in this competition, but just wanted to share my experience to illustrate just how unreliable the leaderboard is. I split the data into 5 folds. My model takes a while to train and since I'm impatient I've been submitting results trained on a single fold. The first fold resulted in an LB score of 1.351 (yay!). However, training the exact same model on a different fold results in an LB score of 1.452.\n\nThere have been some other posts about how the subset used to calculate the LB consists of a disproportionate number of samples with low time to failure. I calculated the mean TTF predicted by my model on the test set:\n\nfold 1: 5.031828\nfold 2: 5.233134\n\nThe first fold seems to be biased towards a shorter TTF for whatever reason. I don't have any reason to believe fold 1 will be any better when evaluated on the entire test set even though it's a lot better on the LB. In case anyone is wondering I got oof scores of ~1.86 and ~1.97. Anyway hope this helps make some sense of the LB scores.",
    "515637": "How are you splitting your folds?",
    "515770": "I follow.\nI am using quake-based split and I am nowhere near MAE ~ 2 🤔",
    "515975": "Thank you for sharing this info.  The organizers have kept the design of data splitting a secret, but the most likely scenario is that the public LB is based only on the first 1-2 cycles/earthquakes in the test data, and the private LB comprises completely different cycles/earthquakes. So if your fold has a majority samples from similar cycles to public LB, your score will be better.\n\nMany have pointed out that the models tend to have worse prediction on segments with high TTF values, but I have a different explanation. \nAt the beginning of each cycle, the signal properties of all cycles are the same. However, in some cycles, the signal escalates/intensifies quickly leading to a short cycle. In other cycles, the signal escalates in a slower pace, or encounter mini-quakes that significantly decrease the intensity, ultimately leading to a long cycle. If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%.  There seem to be no features in the signal that inform which behavior the cycle will take except the \"mean\" of the signal. But so far, I'm not sure if the mean of the signal is physically important feature or batch-effect noise that should be eliminated by centering or high-pass filtering.\n\nFrom my observation, all models have a trade-off between the two types of cycles, the better the model on one type, the worse it behaves on the other. Unfortunately, the final results will depend on the distribution of the private LB test data, and the best balanced model may not win.",
    "517373": "Some of my submissions which score around 1.410 have mean around 5.57-5.59. Others which score around 1.440 have mean around 5.35-5.50. My current best single model CV has (test) mean 5.610 which scores 1.413.\n\nTo me, higher mean of predictions means that the model leans towards higher ttf. Given 2 submissions with same CV, same LB, the one with larger mean may perform better on private LB.\n\nAnyone with this kind of information, please share so all of us can have some insights?",
    "522103": "This seems very high to me. Mine is under 5.2",
    "523929": "what is your quake-based split means?",
    "525626": "amjad85 How did you come up with that estimate? \n\n\"If we can distinguish between the two types (say you use the ID of the cycle in training), the MAE will drop about 40-50%\"",
    "525635": "I tried it. Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement, but obviously not so useful for the test set.",
    "525638": "I see. Thanks!",
    "526090": "&gt; Add the quake ID as a categorical feature to your model and see. It will lead to significant improvement,\n\nIf it does improve your cv score then your cv setting is flawed as it fails to detect this over fitting.",
    "526176": "Yes I was shuffling when I tried it.",
    "529056": "Hi there, when you guys talk about \"mean value\", are you referring to mean of predicted TTF of all test samples?",
    "529057": "khahuras @areveillon, forgot to @ you two",
    "529067": "Yes it is the mean value of predicted test samples."
  },
  "source": "meta"
}