{
  "id": 94389,
  "title": "Private Leaderboard",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94389",
  "author_name": "",
  "post_date": "2019-06-04T08:48:59.214418900Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everybody, \nPlease someone has an, mathematical, explanation for what happen on the change between Private Leaderboard and Public Leaderboard ?\nMany thanks in advance</p>",
  "messages": [
    {
      "id": "542977",
      "postDate": "06/04/2019 08:48:59",
      "content": "<p>Hello everybody, \nPlease someone has an, mathematical, explanation for what happen on the change between Private Leaderboard and Public Leaderboard ?\nMany thanks in advance</p>",
      "rawMarkdown": "Hello everybody, \nPlease someone has an, mathematical, explanation for what happen on the change between Private Leaderboard and Public Leaderboard ?\nMany thanks in advance",
      "votes": null
    },
    {
      "id": "543244",
      "postDate": "06/04/2019 12:07:23",
      "content": "<p>Suppose we have only 8 points in the private data, 1 point in public and 15 points in the training data -- these are the maxima of the time to failure​, or the lengths of the segments between earthquakes. The solution minimising the Mean Absolute Error (MAE) in the training data is the median of 15 points. This is not the optimal solution for the public or the private. The final score should be aiming for the meadia​n of the private, but many people optimise for the training. This is my (over) simplified view. (I didn't optimise toward the public, so don't know if or how other people did)</p>",
      "rawMarkdown": "Suppose we have only 8 points in the private data, 1 point in public and 15 points in the training data -- these are the maxima of the time to failure​, or the lengths of the segments between earthquakes. The solution minimising the Mean Absolute Error (MAE) in the training data is the median of 15 points. This is not the optimal solution for the public or the private. The final score should be aiming for the meadia​n of the private, but many people optimise for the training. This is my (over) simplified view. (I didn't optimise toward the public, so don't know if or how other people did)",
      "votes": null
    },
    {
      "id": "543312",
      "postDate": "06/04/2019 13:14:25",
      "content": "<p>Hi Jun, Thanks a lot for your answer and congratulation for your 2end place :) \nI understand as well that the score expected for private data is the score who should be approch. However, what i found in this completion that  public and train data can't be a good candidates to score well...  and i see a lot of kagglers ares disappointed to work hard in order to score well on public and the difference with private is too large... \nhave you detected that difference during the competition? </p>",
      "rawMarkdown": "Hi Jun, Thanks a lot for your answer and congratulation for your 2end place :) \nI understand as well that the score expected for private data is the score who should be approch. However, what i found in this completion that  public and train data can't be a good candidates to score well...  and i see a lot of kagglers ares disappointed to work hard in order to score well on public and the difference with private is too large... \nhave you detected that difference during the competition?",
      "votes": null
    },
    {
      "id": "543317",
      "postDate": "06/04/2019 13:16:51",
      "content": "<p>Many people have detected that difference during the competition. It has been discussed at length.</p>",
      "rawMarkdown": "Many people have detected that difference during the competition. It has been discussed at length.",
      "votes": null
    },
    {
      "id": "543328",
      "postDate": "06/04/2019 13:25:24",
      "content": "<p>sorry but that difference is abused. If i take my example: i crashed from 291 place with score 1.391 to 1501 place with score of 2.57746. And i have no explanation for what happen... </p>",
      "rawMarkdown": "sorry but that difference is abused. If i take my example: i crashed from 291 place with score 1.391 to 1501 place with score of 2.57746. And i have no explanation for what happen...",
      "votes": null
    },
    {
      "id": "543380",
      "postDate": "06/04/2019 14:01:58",
      "content": "<p>The public - train difference seems obvious from the sample_submission.csv score (e.g. \"Public set time to failure distribution\" discussion). The train - test (mostly private) difference is not that clear. I realised that difference by plotting the histogram of standard deviation of the train vs test, and later confirmed it with the p4677 figure. It is always a good idea to check the train - test data difference by looking the distribution of the features. The 1st place solution does this very systematically.</p>",
      "rawMarkdown": "The public - train difference seems obvious from the sample_submission.csv score (e.g. \"Public set time to failure distribution\" discussion). The train - test (mostly private) difference is not that clear. I realised that difference by plotting the histogram of standard deviation of the train vs test, and later confirmed it with the p4677 figure. It is always a good idea to check the train - test data difference by looking the distribution of the features. The 1st place solution does this very systematically.",
      "votes": null
    },
    {
      "id": "543399",
      "postDate": "06/04/2019 14:09:34",
      "content": "<p>Of course it is abused, <a href=\"/aminepy\">@aminepy</a> \nThis makes kaggle somewhat different from real world problems.</p>\n\n<p>But at kaggle we have the option to look at the test set in many competitions. \nHave a look at my teammember's explanation of how we tackled that task with train and test being very different: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543389\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543389</a></p>\n\n<p>As Jun said, to us this was no surprise and we analysed the data in detail to get to our final submissions. </p>",
      "rawMarkdown": "Of course it is abused, @aminepy \nThis makes kaggle somewhat different from real world problems.\n\nBut at kaggle we have the option to look at the test set in many competitions. \nHave a look at my teammember's explanation of how we tackled that task with train and test being very different: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543389\n\nAs Jun said, to us this was no surprise and we analysed the data in detail to get to our final submissions.",
      "votes": null
    },
    {
      "id": "543416",
      "postDate": "06/04/2019 14:19:14",
      "content": "<p>I see, probably something more than what I explained is happening in your case. That is probably some kind of overfit. (1) Do your \"local CV\" scores improve as your public score improve? If you don't double check your imporovemen​​t in public LB with your local data, there is a danger of overfitting to public. (2) Do your features depend on the mean of the data? The distribution of mean is different between train and test (there was a discussion on that) and it could cause not generalizing to majo​rit​y of the test data (Maybe the public is right after train and relatively similar to train, but private is more different. I am not confident about this though). This is a special case of checking the distribution of your features between train and test.</p>",
      "rawMarkdown": "I see, probably something more than what I explained is happening in your case. That is probably some kind of overfit. (1) Do your \"local CV\" scores improve as your public score improve? If you don't double check your imporovemen​​t in public LB with your local data, there is a danger of overfitting to public. (2) Do your features depend on the mean of the data? The distribution of mean is different between train and test (there was a discussion on that) and it could cause not generalizing to majo​rit​y of the test data (Maybe the public is right after train and relatively similar to train, but private is more different. I am not confident about this though). This is a special case of checking the distribution of your features between train and test.",
      "votes": null
    },
    {
      "id": "543460",
      "postDate": "06/04/2019 14:37:42",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> , <a href=\"/junkoda\">@junkoda</a>  Many thanks guys for your explanations!  I have taken note of your remarks :) \nAgain congratulations for your high scores, glade to be in touch with you :)</p>",
      "rawMarkdown": "ilu000 , @junkoda  Many thanks guys for your explanations!  I have taken note of your remarks :) \nAgain congratulations for your high scores, glade to be in touch with you :)",
      "votes": null
    },
    {
      "id": "543561",
      "postDate": "06/04/2019 15:41:57",
      "content": "<p>This is very simple. The metric is MAE and the mean(ttf) of Train, TestPublic and TestPrivate are:  5.66, 4 and 6.4</p>\n\n<p>So if you overfit your model to perform well on TestPublic your mean(prediction) would be close to 4. This is 2.4 away from the mean of TestPrivate and hurts a lot the MAE metric. </p>",
      "rawMarkdown": "This is very simple. The metric is MAE and the mean(ttf) of Train, TestPublic and TestPrivate are:  5.66, 4 and 6.4\n\nSo if you overfit your model to perform well on TestPublic your mean(prediction) would be close to 4. This is 2.4 away from the mean of TestPrivate and hurts a lot the MAE metric.",
      "votes": null
    },
    {
      "id": "544145",
      "postDate": "06/05/2019 07:41:14",
      "content": "<p><a href=\"/titericz\">@titericz</a> : Thanks a lot for your explanation :) \nPlease, what \"mean(ttf)\" means? </p>",
      "rawMarkdown": "titericz : Thanks a lot for your explanation :) \nPlease, what \"mean(ttf)\" means?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543244,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "06/04/2019 12:07:23",
      "content": "<p>Suppose we have only 8 points in the private data, 1 point in public and 15 points in the training data -- these are the maxima of the time to failure​, or the lengths of the segments between earthquakes. The solution minimising the Mean Absolute Error (MAE) in the training data is the median of 15 points. This is not the optimal solution for the public or the private. The final score should be aiming for the meadia​n of the private, but many people optimise for the training. This is my (over) simplified view. (I didn't optimise toward the public, so don't know if or how other people did)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543312,
      "author_name": "aminepy",
      "author_url": "",
      "post_date": "06/04/2019 13:14:25",
      "content": "<p>Hi Jun, Thanks a lot for your answer and congratulation for your 2end place :) \nI understand as well that the score expected for private data is the score who should be approch. However, what i found in this completion that  public and train data can't be a good candidates to score well...  and i see a lot of kagglers ares disappointed to work hard in order to score well on public and the difference with private is too large... \nhave you detected that difference during the competition? </p>",
      "votes": null,
      "replies": [
        {
          "id": 543317,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "06/04/2019 13:16:51",
          "content": "<p>Many people have detected that difference during the competition. It has been discussed at length.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543328,
          "author_name": "aminepy",
          "author_url": "",
          "post_date": "06/04/2019 13:25:24",
          "content": "<p>sorry but that difference is abused. If i take my example: i crashed from 291 place with score 1.391 to 1501 place with score of 2.57746. And i have no explanation for what happen... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543380,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "06/04/2019 14:01:58",
          "content": "<p>The public - train difference seems obvious from the sample_submission.csv score (e.g. \"Public set time to failure distribution\" discussion). The train - test (mostly private) difference is not that clear. I realised that difference by plotting the histogram of standard deviation of the train vs test, and later confirmed it with the p4677 figure. It is always a good idea to check the train - test data difference by looking the distribution of the features. The 1st place solution does this very systematically.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543399,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "06/04/2019 14:09:34",
          "content": "<p>Of course it is abused, <a href=\"/aminepy\">@aminepy</a> \nThis makes kaggle somewhat different from real world problems.</p>\n\n<p>But at kaggle we have the option to look at the test set in many competitions. \nHave a look at my teammember's explanation of how we tackled that task with train and test being very different: <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543389\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543389</a></p>\n\n<p>As Jun said, to us this was no surprise and we analysed the data in detail to get to our final submissions. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543416,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "06/04/2019 14:19:14",
          "content": "<p>I see, probably something more than what I explained is happening in your case. That is probably some kind of overfit. (1) Do your \"local CV\" scores improve as your public score improve? If you don't double check your imporovemen​​t in public LB with your local data, there is a danger of overfitting to public. (2) Do your features depend on the mean of the data? The distribution of mean is different between train and test (there was a discussion on that) and it could cause not generalizing to majo​rit​y of the test data (Maybe the public is right after train and relatively similar to train, but private is more different. I am not confident about this though). This is a special case of checking the distribution of your features between train and test.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543460,
      "author_name": "aminepy",
      "author_url": "",
      "post_date": "06/04/2019 14:37:42",
      "content": "<p><a href=\"/ilu000\">@ilu000</a> , <a href=\"/junkoda\">@junkoda</a>  Many thanks guys for your explanations!  I have taken note of your remarks :) \nAgain congratulations for your high scores, glade to be in touch with you :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543561,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:41:57",
      "content": "<p>This is very simple. The metric is MAE and the mean(ttf) of Train, TestPublic and TestPrivate are:  5.66, 4 and 6.4</p>\n\n<p>So if you overfit your model to perform well on TestPublic your mean(prediction) would be close to 4. This is 2.4 away from the mean of TestPrivate and hurts a lot the MAE metric. </p>",
      "votes": null,
      "replies": [
        {
          "id": 544145,
          "author_name": "aminepy",
          "author_url": "",
          "post_date": "06/05/2019 07:41:14",
          "content": "<p><a href=\"/titericz\">@titericz</a> : Thanks a lot for your explanation :) \nPlease, what \"mean(ttf)\" means? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "542977": "Hello everybody, \nPlease someone has an, mathematical, explanation for what happen on the change between Private Leaderboard and Public Leaderboard ?\nMany thanks in advance",
    "543244": "Suppose we have only 8 points in the private data, 1 point in public and 15 points in the training data -- these are the maxima of the time to failure​, or the lengths of the segments between earthquakes. The solution minimising the Mean Absolute Error (MAE) in the training data is the median of 15 points. This is not the optimal solution for the public or the private. The final score should be aiming for the meadia​n of the private, but many people optimise for the training. This is my (over) simplified view. (I didn't optimise toward the public, so don't know if or how other people did)",
    "543312": "Hi Jun, Thanks a lot for your answer and congratulation for your 2end place :) \nI understand as well that the score expected for private data is the score who should be approch. However, what i found in this completion that  public and train data can't be a good candidates to score well...  and i see a lot of kagglers ares disappointed to work hard in order to score well on public and the difference with private is too large... \nhave you detected that difference during the competition?",
    "543317": "Many people have detected that difference during the competition. It has been discussed at length.",
    "543328": "sorry but that difference is abused. If i take my example: i crashed from 291 place with score 1.391 to 1501 place with score of 2.57746. And i have no explanation for what happen...",
    "543380": "The public - train difference seems obvious from the sample_submission.csv score (e.g. \"Public set time to failure distribution\" discussion). The train - test (mostly private) difference is not that clear. I realised that difference by plotting the histogram of standard deviation of the train vs test, and later confirmed it with the p4677 figure. It is always a good idea to check the train - test data difference by looking the distribution of the features. The 1st place solution does this very systematically.",
    "543399": "Of course it is abused, @aminepy \nThis makes kaggle somewhat different from real world problems.\n\nBut at kaggle we have the option to look at the test set in many competitions. \nHave a look at my teammember's explanation of how we tackled that task with train and test being very different: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543389\n\nAs Jun said, to us this was no surprise and we analysed the data in detail to get to our final submissions.",
    "543416": "I see, probably something more than what I explained is happening in your case. That is probably some kind of overfit. (1) Do your \"local CV\" scores improve as your public score improve? If you don't double check your imporovemen​​t in public LB with your local data, there is a danger of overfitting to public. (2) Do your features depend on the mean of the data? The distribution of mean is different between train and test (there was a discussion on that) and it could cause not generalizing to majo​rit​y of the test data (Maybe the public is right after train and relatively similar to train, but private is more different. I am not confident about this though). This is a special case of checking the distribution of your features between train and test.",
    "543460": "ilu000 , @junkoda  Many thanks guys for your explanations!  I have taken note of your remarks :) \nAgain congratulations for your high scores, glade to be in touch with you :)",
    "543561": "This is very simple. The metric is MAE and the mean(ttf) of Train, TestPublic and TestPrivate are:  5.66, 4 and 6.4\n\nSo if you overfit your model to perform well on TestPublic your mean(prediction) would be close to 4. This is 2.4 away from the mean of TestPrivate and hurts a lot the MAE metric.",
    "544145": "titericz : Thanks a lot for your explanation :) \nPlease, what \"mean(ttf)\" means?"
  },
  "source": "meta"
}