{
  "id": 94355,
  "title": "Three Keys to this Competition",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94355",
  "author_name": "",
  "post_date": "2019-06-04T03:45:59.058923100Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>My shoulda-coulda-woulda submission of 2.34085 would have landed me in 11th place. I think there were several keys to this competition, one of which I did not give enough weight ... literally.</p>\n\n<p><strong>1. Do not shuffle time series data.</strong>\nSimilarly do not train with data combined from before and after the point(s) you are predicting. Given some of the discussions about this, I am sure that some participants will disagree with me. However, when data comes from a continuous source such as a geophysical or biological process, the data is frequently nonstationary. Since we were told that the public/private test data came after the train data in time, a cross validation method should respect the temporal nature of the train data. I found the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662\">discussion</a>/kernel by @felipefonte99 to be particularly revealing. Also, playing around with the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90838#latest-536337\">p4581 data set</a>, I found it to be rather different from the data set of this competition to the point that I had trouble leveraging it. Also, I observed a substantial shift across time in some features derived from the p4581 data set that were very apparent to the naked eye. I ended up using the first half of the competition data as training and the second half of the data as validation. I hated having to use so much data in the validation fold, but it was the best split that allowed for similar distributions between the two folds.</p>\n\n<p><strong>2. Predictions that fit peaks well are probably a sign of overfitting</strong>\nI saw no evidence that it is possible to predict both low and high peaks in the same fold without overfitting. As I thought about this more, it made sense. When a trial begins, there is likely nothing in this presumably quiescent state that would distinguish it from longer or shorter quiescent states. The maximum values of the predictions for each earthquake should be about the same.</p>\n\n<p><strong>3. Scale the predictions</strong>\nThis is where I messed up. At one point I started to apply a multiplicative factor to my predictions, but then decided it was too much of a gamble, hence, my unsubmitted 2.34085 score. The concept behind scaling up the predictions is based on the idea from #2 that the maximum values for each quake should be about the same and there is an optimal scaling factor to make them best match the mean of the peaks of the public/private test data. This scaling factor can be estimated from the data provided in a paper authored by the host that is mentioned <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086#latest-542678\">here</a> by @robikscube.</p>",
  "messages": [
    {
      "id": "542704",
      "postDate": "06/04/2019 03:45:59",
      "content": "<p>My shoulda-coulda-woulda submission of 2.34085 would have landed me in 11th place. I think there were several keys to this competition, one of which I did not give enough weight ... literally.</p>\n\n<p><strong>1. Do not shuffle time series data.</strong>\nSimilarly do not train with data combined from before and after the point(s) you are predicting. Given some of the discussions about this, I am sure that some participants will disagree with me. However, when data comes from a continuous source such as a geophysical or biological process, the data is frequently nonstationary. Since we were told that the public/private test data came after the train data in time, a cross validation method should respect the temporal nature of the train data. I found the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662\">discussion</a>/kernel by @felipefonte99 to be particularly revealing. Also, playing around with the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90838#latest-536337\">p4581 data set</a>, I found it to be rather different from the data set of this competition to the point that I had trouble leveraging it. Also, I observed a substantial shift across time in some features derived from the p4581 data set that were very apparent to the naked eye. I ended up using the first half of the competition data as training and the second half of the data as validation. I hated having to use so much data in the validation fold, but it was the best split that allowed for similar distributions between the two folds.</p>\n\n<p><strong>2. Predictions that fit peaks well are probably a sign of overfitting</strong>\nI saw no evidence that it is possible to predict both low and high peaks in the same fold without overfitting. As I thought about this more, it made sense. When a trial begins, there is likely nothing in this presumably quiescent state that would distinguish it from longer or shorter quiescent states. The maximum values of the predictions for each earthquake should be about the same.</p>\n\n<p><strong>3. Scale the predictions</strong>\nThis is where I messed up. At one point I started to apply a multiplicative factor to my predictions, but then decided it was too much of a gamble, hence, my unsubmitted 2.34085 score. The concept behind scaling up the predictions is based on the idea from #2 that the maximum values for each quake should be about the same and there is an optimal scaling factor to make them best match the mean of the peaks of the public/private test data. This scaling factor can be estimated from the data provided in a paper authored by the host that is mentioned <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086#latest-542678\">here</a> by @robikscube.</p>",
      "rawMarkdown": "My shoulda-coulda-woulda submission of 2.34085 would have landed me in 11th place. I think there were several keys to this competition, one of which I did not give enough weight ... literally.\n\n**1. Do not shuffle time series data.**\nSimilarly do not train with data combined from before and after the point(s) you are predicting. Given some of the discussions about this, I am sure that some participants will disagree with me. However, when data comes from a continuous source such as a geophysical or biological process, the data is frequently nonstationary. Since we were told that the public/private test data came after the train data in time, a cross validation method should respect the temporal nature of the train data. I found the [discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662)/kernel by @felipefonte99 to be particularly revealing. Also, playing around with the [p4581 data set](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90838#latest-536337), I found it to be rather different from the data set of this competition to the point that I had trouble leveraging it. Also, I observed a substantial shift across time in some features derived from the p4581 data set that were very apparent to the naked eye. I ended up using the first half of the competition data as training and the second half of the data as validation. I hated having to use so much data in the validation fold, but it was the best split that allowed for similar distributions between the two folds.\n\n**2. Predictions that fit peaks well are probably a sign of overfitting**\nI saw no evidence that it is possible to predict both low and high peaks in the same fold without overfitting. As I thought about this more, it made sense. When a trial begins, there is likely nothing in this presumably quiescent state that would distinguish it from longer or shorter quiescent states. The maximum values of the predictions for each earthquake should be about the same.\n\n**3. Scale the predictions**\nThis is where I messed up. At one point I started to apply a multiplicative factor to my predictions, but then decided it was too much of a gamble, hence, my unsubmitted 2.34085 score. The concept behind scaling up the predictions is based on the idea from #2 that the maximum values for each quake should be about the same and there is an optimal scaling factor to make them best match the mean of the peaks of the public/private test data. This scaling factor can be estimated from the data provided in a paper authored by the host that is mentioned [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086#latest-542678) by @robikscube.",
      "votes": null
    },
    {
      "id": "542721",
      "postDate": "06/04/2019 04:02:13",
      "content": "<p>I actually shuffled it. I did it because I thought that the small size of the chunks made each one not so related to the sorrounding ones. It is the information of just an instant.\nI agree that scaling the predictions is important. I transformed the features separately for train and test to standard normal distribution.</p>",
      "rawMarkdown": "I actually shuffled it. I did it because I thought that the small size of the chunks made each one not so related to the sorrounding ones. It is the information of just an instant.\nI agree that scaling the predictions is important. I transformed the features separately for train and test to standard normal distribution.",
      "votes": null
    },
    {
      "id": "542738",
      "postDate": "06/04/2019 04:17:54",
      "content": "<p>If only i could think this way after reading from discussion that the possibility mean shifting in private test data (as information leak from the paper ).  congratulation for your gold medals. Nice adaptation</p>",
      "rawMarkdown": "If only i could think this way after reading from discussion that the possibility mean shifting in private test data (as information leak from the paper ).  congratulation for your gold medals. Nice adaptation",
      "votes": null
    },
    {
      "id": "542745",
      "postDate": "06/04/2019 04:24:15",
      "content": "<p>Thank you. Just to clarify, I didn't transformed because what I saw in the paper. I did it because I saw a shift in the distribution of features in test set compared to the training set.</p>",
      "rawMarkdown": "Thank you. Just to clarify, I didn't transformed because what I saw in the paper. I did it because I saw a shift in the distribution of features in test set compared to the training set.",
      "votes": null
    },
    {
      "id": "542746",
      "postDate": "06/04/2019 04:26:06",
      "content": "<p>Well played, even without looking at the paper, you able to avoid the illusion :). This is very input for next competitions for me. Thanks so much</p>",
      "rawMarkdown": "Well played, even without looking at the paper, you able to avoid the illusion :). This is very input for next competitions for me. Thanks so much",
      "votes": null
    },
    {
      "id": "542817",
      "postDate": "06/04/2019 05:47:58",
      "content": "<p>I tried to build some different model to classify sequence of segments and it didn't worked at all. Maybe that explains why different segment inside the same quake period are not correlated and Kfold can be used normally. In my solution I used nested CV strategy.</p>",
      "rawMarkdown": "I tried to build some different model to classify sequence of segments and it didn't worked at all. Maybe that explains why different segment inside the same quake period are not correlated and Kfold can be used normally. In my solution I used nested CV strategy.",
      "votes": null
    },
    {
      "id": "543347",
      "postDate": "06/04/2019 13:40:35",
      "content": "<p>That's an interesting observation. I'm wondering though why leave-one-out majorly overfit (see <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662\">here</a>). Perhaps there are micro correlations but not so much macro correlations (except if a piece of the laboratory apparatus falls off/degrades of course)?</p>",
      "rawMarkdown": "That's an interesting observation. I'm wondering though why leave-one-out majorly overfit (see [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662)). Perhaps there are micro correlations but not so much macro correlations (except if a piece of the laboratory apparatus falls off/degrades of course)?",
      "votes": null
    },
    {
      "id": "543356",
      "postDate": "06/04/2019 13:43:55",
      "content": "<p>CarlosPK, thank you for sharing that you shuffled the data and still performed very well. I hope you'll consider posting your solution. I'm curious to see how you tackled this competition.</p>",
      "rawMarkdown": "CarlosPK, thank you for sharing that you shuffled the data and still performed very well. I hope you'll consider posting your solution. I'm curious to see how you tackled this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 542721,
      "author_name": "carlospk",
      "author_url": "",
      "post_date": "06/04/2019 04:02:13",
      "content": "<p>I actually shuffled it. I did it because I thought that the small size of the chunks made each one not so related to the sorrounding ones. It is the information of just an instant.\nI agree that scaling the predictions is important. I transformed the features separately for train and test to standard normal distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 542738,
          "author_name": "arisukma",
          "author_url": "",
          "post_date": "06/04/2019 04:17:54",
          "content": "<p>If only i could think this way after reading from discussion that the possibility mean shifting in private test data (as information leak from the paper ).  congratulation for your gold medals. Nice adaptation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 542745,
          "author_name": "carlospk",
          "author_url": "",
          "post_date": "06/04/2019 04:24:15",
          "content": "<p>Thank you. Just to clarify, I didn't transformed because what I saw in the paper. I did it because I saw a shift in the distribution of features in test set compared to the training set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 542746,
          "author_name": "arisukma",
          "author_url": "",
          "post_date": "06/04/2019 04:26:06",
          "content": "<p>Well played, even without looking at the paper, you able to avoid the illusion :). This is very input for next competitions for me. Thanks so much</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543356,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "06/04/2019 13:43:55",
          "content": "<p>CarlosPK, thank you for sharing that you shuffled the data and still performed very well. I hope you'll consider posting your solution. I'm curious to see how you tackled this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542817,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 05:47:58",
      "content": "<p>I tried to build some different model to classify sequence of segments and it didn't worked at all. Maybe that explains why different segment inside the same quake period are not correlated and Kfold can be used normally. In my solution I used nested CV strategy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543347,
          "author_name": "trentb",
          "author_url": "",
          "post_date": "06/04/2019 13:40:35",
          "content": "<p>That's an interesting observation. I'm wondering though why leave-one-out majorly overfit (see <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662\">here</a>). Perhaps there are micro correlations but not so much macro correlations (except if a piece of the laboratory apparatus falls off/degrades of course)?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "542704": "My shoulda-coulda-woulda submission of 2.34085 would have landed me in 11th place. I think there were several keys to this competition, one of which I did not give enough weight ... literally.\n\n**1. Do not shuffle time series data.**\nSimilarly do not train with data combined from before and after the point(s) you are predicting. Given some of the discussions about this, I am sure that some participants will disagree with me. However, when data comes from a continuous source such as a geophysical or biological process, the data is frequently nonstationary. Since we were told that the public/private test data came after the train data in time, a cross validation method should respect the temporal nature of the train data. I found the [discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662)/kernel by @felipefonte99 to be particularly revealing. Also, playing around with the [p4581 data set](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90838#latest-536337), I found it to be rather different from the data set of this competition to the point that I had trouble leveraging it. Also, I observed a substantial shift across time in some features derived from the p4581 data set that were very apparent to the naked eye. I ended up using the first half of the competition data as training and the second half of the data as validation. I hated having to use so much data in the validation fold, but it was the best split that allowed for similar distributions between the two folds.\n\n**2. Predictions that fit peaks well are probably a sign of overfitting**\nI saw no evidence that it is possible to predict both low and high peaks in the same fold without overfitting. As I thought about this more, it made sense. When a trial begins, there is likely nothing in this presumably quiescent state that would distinguish it from longer or shorter quiescent states. The maximum values of the predictions for each earthquake should be about the same.\n\n**3. Scale the predictions**\nThis is where I messed up. At one point I started to apply a multiplicative factor to my predictions, but then decided it was too much of a gamble, hence, my unsubmitted 2.34085 score. The concept behind scaling up the predictions is based on the idea from #2 that the maximum values for each quake should be about the same and there is an optimal scaling factor to make them best match the mean of the peaks of the public/private test data. This scaling factor can be estimated from the data provided in a paper authored by the host that is mentioned [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94086#latest-542678) by @robikscube.",
    "542721": "I actually shuffled it. I did it because I thought that the small size of the chunks made each one not so related to the sorrounding ones. It is the information of just an instant.\nI agree that scaling the predictions is important. I transformed the features separately for train and test to standard normal distribution.",
    "542738": "If only i could think this way after reading from discussion that the possibility mean shifting in private test data (as information leak from the paper ).  congratulation for your gold medals. Nice adaptation",
    "542745": "Thank you. Just to clarify, I didn't transformed because what I saw in the paper. I did it because I saw a shift in the distribution of features in test set compared to the training set.",
    "542746": "Well played, even without looking at the paper, you able to avoid the illusion :). This is very input for next competitions for me. Thanks so much",
    "542817": "I tried to build some different model to classify sequence of segments and it didn't worked at all. Maybe that explains why different segment inside the same quake period are not correlated and Kfold can be used normally. In my solution I used nested CV strategy.",
    "543347": "That's an interesting observation. I'm wondering though why leave-one-out majorly overfit (see [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93325#latest-537662)). Perhaps there are micro correlations but not so much macro correlations (except if a piece of the laboratory apparatus falls off/degrades of course)?",
    "543356": "CarlosPK, thank you for sharing that you shuffled the data and still performed very well. I hope you'll consider posting your solution. I'm curious to see how you tackled this competition."
  },
  "source": "meta"
}