{
  "id": 94396,
  "title": "Unconventional gold (16th place solution)",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/miguel-perez-unconventional-gold-16th-place-soluti",
  "author_name": "",
  "post_date": "2019-06-04T10:00:52.947Z",
  "votes": 30,
  "comment_count": 12,
  "views": 0,
  "content": "<p>To begin with, I want to congratulate all competitors regardless of leaderboard ranking for what has sure been a lot of work this months. In many cases it may feel to be unrewarded. Don't feel despair, just make sure to learn as much as possible including final wrapups and  it will have been worth it. </p>\n\n<p>So, here are my hints, hoping that will be useful for someone:</p>\n\n<h3>GENERAL IDEAS:</h3>\n\n<ul>\n<li><p>Assessing the amount of noise in a competition is the key to generalization. (And it is never easy)</p></li>\n<li><p>You can win without stacking, and even without a too complex model, as long as you frame the problem correctly. </p></li>\n<li><p>Don't be trapped by the Public Leaderboard race, just treat it as another fold of your validation or -sometimes- ignore it altogether</p></li>\n</ul>\n\n<h3>MY -CONDENSED- PATH TO THIS GOLD:</h3>\n\n<ul>\n<li><p>First month:  lost time trying to make a RNN work in a thousand of creative ways. After lots of work had to bite the bullet:  RNN was not the way to go. <strong>Difficult decision</strong>: Start from scratch.</p></li>\n<li><p>Second month: The fight to find a proper validation scheme, the realization of just how overfittable the dataset was. Lots of feature engineering, most of it the kind already seen in forums, only \"original\" ones some computationally expensive Matrix Profile stats that turned out not to be too useful. <strong>Difficult decision</strong>:  Use random forest. (Less tunning, less sparsity of feature importance than boosting, important to make also feature selection of the less overfitting ones).  </p></li>\n<li><p>Last three days: As <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">this post</a>  disclosed there was a potential leak that showed train-test relationship. Not a too easy to leverage leak, and risky to use though. Seeing how important it has been in the end some kind of systematic processing of the image would have been extremely effective, but  I didn't do that. Instead I simply gained a couple of insights: </p></li>\n</ul>\n\n<p>First insight: outliers removal. Outliers are not just out of range values but more in general train data that is unusual vs. what is to be expected. I realized that the Earthques 5 and 6 were actually a \"synthetic\" split of what in the experiment was considered a single earhtquake, i.e. the stress theshold was different . <strong>Difficult decision</strong>: to remove those two earthquakes data of an already small dataset.</p>\n\n<p>Second insight: I had already observed that bimodal distribution of predictions was very different between train and test showing  that shorter earquakes in test were way longer than shorter earthquakes in train. That was an ugly warning, and fitted quite well with what at a glance could be seen in the experiment. I visually estimated the proportion and lowered the weight of \"small\" earthquakes in training set by 1/5, no more fancy adjustment than that,  noise was too big for more fine grained tunning anyway. <strong>Last unconfortable decision</strong>: select a submission that placed me like 2300th in public leaderboard cause it made sense.</p>\n\n<h3>ONE LAST WARNING</h3>\n\n<p>So, that was it. Hope it was of some interest. One last warning, don't take any of this decissions to rigidly (like using Random Forest, for example, I never choose a RF without a reason to do it). <strong>In the end, it is all about generalization and every competition is different</strong>. </p>\n\n<p>As I said in the beginning, just make sure to keep on learning and to have fun :-)</p>",
  "messages": [
    {
      "id": "543012",
      "postDate": "06/04/2019 09:14:39",
      "content": "<p>To begin with, I want to congratulate all competitors regardless of leaderboard ranking for what has sure been a lot of work this months. In many cases it may feel to be unrewarded. Don't feel despair, just make sure to learn as much as possible including final wrapups and  it will have been worth it. </p>\n\n<p>So, here are my hints, hoping that will be useful for someone:</p>\n\n<h3>GENERAL IDEAS:</h3>\n\n<ul>\n<li><p>Assessing the amount of noise in a competition is the key to generalization. (And it is never easy)</p></li>\n<li><p>You can win without stacking, and even without a too complex model, as long as you frame the problem correctly. </p></li>\n<li><p>Don't be trapped by the Public Leaderboard race, just treat it as another fold of your validation or -sometimes- ignore it altogether</p></li>\n</ul>\n\n<h3>MY -CONDENSED- PATH TO THIS GOLD:</h3>\n\n<ul>\n<li><p>First month:  lost time trying to make a RNN work in a thousand of creative ways. After lots of work had to bite the bullet:  RNN was not the way to go. <strong>Difficult decision</strong>: Start from scratch.</p></li>\n<li><p>Second month: The fight to find a proper validation scheme, the realization of just how overfittable the dataset was. Lots of feature engineering, most of it the kind already seen in forums, only \"original\" ones some computationally expensive Matrix Profile stats that turned out not to be too useful. <strong>Difficult decision</strong>:  Use random forest. (Less tunning, less sparsity of feature importance than boosting, important to make also feature selection of the less overfitting ones).  </p></li>\n<li><p>Last three days: As <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">this post</a>  disclosed there was a potential leak that showed train-test relationship. Not a too easy to leverage leak, and risky to use though. Seeing how important it has been in the end some kind of systematic processing of the image would have been extremely effective, but  I didn't do that. Instead I simply gained a couple of insights: </p></li>\n</ul>\n\n<p>First insight: outliers removal. Outliers are not just out of range values but more in general train data that is unusual vs. what is to be expected. I realized that the Earthques 5 and 6 were actually a \"synthetic\" split of what in the experiment was considered a single earhtquake, i.e. the stress theshold was different . <strong>Difficult decision</strong>: to remove those two earthquakes data of an already small dataset.</p>\n\n<p>Second insight: I had already observed that bimodal distribution of predictions was very different between train and test showing  that shorter earquakes in test were way longer than shorter earthquakes in train. That was an ugly warning, and fitted quite well with what at a glance could be seen in the experiment. I visually estimated the proportion and lowered the weight of \"small\" earthquakes in training set by 1/5, no more fancy adjustment than that,  noise was too big for more fine grained tunning anyway. <strong>Last unconfortable decision</strong>: select a submission that placed me like 2300th in public leaderboard cause it made sense.</p>\n\n<h3>ONE LAST WARNING</h3>\n\n<p>So, that was it. Hope it was of some interest. One last warning, don't take any of this decissions to rigidly (like using Random Forest, for example, I never choose a RF without a reason to do it). <strong>In the end, it is all about generalization and every competition is different</strong>. </p>\n\n<p>As I said in the beginning, just make sure to keep on learning and to have fun :-)</p>",
      "rawMarkdown": "To begin with, I want to congratulate all competitors regardless of leaderboard ranking for what has sure been a lot of work this months. In many cases it may feel to be unrewarded. Don't feel despair, just make sure to learn as much as possible including final wrapups and  it will have been worth it. \n\nSo, here are my hints, hoping that will be useful for someone:\n\n### GENERAL IDEAS:\n\n-  Assessing the amount of noise in a competition is the key to generalization. (And it is never easy)\n\n- You can win without stacking, and even without a too complex model, as long as you frame the problem correctly. \n \n- Don't be trapped by the Public Leaderboard race, just treat it as another fold of your validation or -sometimes- ignore it altogether\n\n### MY -CONDENSED- PATH TO THIS GOLD:\n\n- First month:  lost time trying to make a RNN work in a thousand of creative ways. After lots of work had to bite the bullet:  RNN was not the way to go. **Difficult decision**: Start from scratch.\n\n- Second month: The fight to find a proper validation scheme, the realization of just how overfittable the dataset was. Lots of feature engineering, most of it the kind already seen in forums, only \"original\" ones some computationally expensive Matrix Profile stats that turned out not to be too useful. **Difficult decision**:  Use random forest. (Less tunning, less sparsity of feature importance than boosting, important to make also feature selection of the less overfitting ones).  \n\n- Last three days: As [this post](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844)  disclosed there was a potential leak that showed train-test relationship. Not a too easy to leverage leak, and risky to use though. Seeing how important it has been in the end some kind of systematic processing of the image would have been extremely effective, but  I didn't do that. Instead I simply gained a couple of insights: \n\nFirst insight: outliers removal. Outliers are not just out of range values but more in general train data that is unusual vs. what is to be expected. I realized that the Earthques 5 and 6 were actually a \"synthetic\" split of what in the experiment was considered a single earhtquake, i.e. the stress theshold was different . **Difficult decision**: to remove those two earthquakes data of an already small dataset.\n\nSecond insight: I had already observed that bimodal distribution of predictions was very different between train and test showing  that shorter earquakes in test were way longer than shorter earthquakes in train. That was an ugly warning, and fitted quite well with what at a glance could be seen in the experiment. I visually estimated the proportion and lowered the weight of \"small\" earthquakes in training set by 1/5, no more fancy adjustment than that,  noise was too big for more fine grained tunning anyway. **Last unconfortable decision**: select a submission that placed me like 2300th in public leaderboard cause it made sense.\n\n### ONE LAST WARNING\n\nSo, that was it. Hope it was of some interest. One last warning, don't take any of this decissions to rigidly (like using Random Forest, for example, I never choose a RF without a reason to do it). **In the end, it is all about generalization and every competition is different**. \n\nAs I said in the beginning, just make sure to keep on learning and to have fun :-)",
      "votes": null
    },
    {
      "id": "543030",
      "postDate": "06/04/2019 09:28:21",
      "content": "<p>Thanks for sharing <a href=\"/miguelpm\">@miguelpm</a> </p>",
      "rawMarkdown": "Thanks for sharing @miguelpm",
      "votes": null
    },
    {
      "id": "543046",
      "postDate": "06/04/2019 09:48:47",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats! Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "543078",
      "postDate": "06/04/2019 10:14:18",
      "content": "<p>Thanks for sharing and congrats for gold!  I'm not surprised RF works well, I even had a KNN that worked quite well ;)</p>",
      "rawMarkdown": "Thanks for sharing and congrats for gold!  I'm not surprised RF works well, I even had a KNN that worked quite well ;)",
      "votes": null
    },
    {
      "id": "543272",
      "postDate": "06/04/2019 12:33:39",
      "content": "<p>Congratulations! :)</p>",
      "rawMarkdown": "Congratulations! :)",
      "votes": null
    },
    {
      "id": "543576",
      "postDate": "06/04/2019 15:48:05",
      "content": "<p>Congrats <a href=\"/miguelpm\">@miguelpm</a> . And thanks for sharing. </p>",
      "rawMarkdown": "Congrats @miguelpm . And thanks for sharing.",
      "votes": null
    },
    {
      "id": "543733",
      "postDate": "06/04/2019 18:23:49",
      "content": "<p>Great !!!!!!!</p>",
      "rawMarkdown": "Great !!!!!!!",
      "votes": null
    },
    {
      "id": "543907",
      "postDate": "06/04/2019 23:36:00",
      "content": "<p>Congrats! I've learned a lot from this post, because there are many what you did with what you thought. Thank you for sharing!</p>",
      "rawMarkdown": "Congrats! I've learned a lot from this post, because there are many what you did with what you thought. Thank you for sharing!",
      "votes": null
    },
    {
      "id": "546494",
      "postDate": "06/06/2019 16:10:25",
      "content": "<p>Interesting, thanks for  sharing</p>",
      "rawMarkdown": "Interesting, thanks for  sharing",
      "votes": null
    },
    {
      "id": "573614",
      "postDate": "07/12/2019 14:03:53",
      "content": "<p>Thanks for sharing <a href=\"/miguelpm\">@miguelpm</a> , may I know how you use Matrix Profile to create new features as you mentioned. Link to some docs should be enough. Thanks.</p>",
      "rawMarkdown": "Thanks for sharing @miguelpm , may I know how you use Matrix Profile to create new features as you mentioned. Link to some docs should be enough. Thanks.",
      "votes": null
    },
    {
      "id": "573824",
      "postDate": "07/12/2019 20:02:13",
      "content": "<p>Sure, it's a very interesting tool, here it's all you need to learn about it <a href=\"https://www.cs.ucr.edu/~eamonn/MatrixProfile.html\">https://www.cs.ucr.edu/~eamonn/MatrixProfile.html</a></p>",
      "rawMarkdown": "Sure, it's a very interesting tool, here it's all you need to learn about it https://www.cs.ucr.edu/~eamonn/MatrixProfile.html",
      "votes": null
    },
    {
      "id": "585725",
      "postDate": "07/28/2019 00:03:53",
      "content": "<p>Interesting and smart </p>",
      "rawMarkdown": "Interesting and smart",
      "votes": null
    },
    {
      "id": "589434",
      "postDate": "07/31/2019 22:13:47",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543030,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/04/2019 09:28:21",
      "content": "<p>Thanks for sharing <a href=\"/miguelpm\">@miguelpm</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543046,
      "author_name": "prashanththangavel",
      "author_url": "",
      "post_date": "06/04/2019 09:48:47",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543078,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2019 10:14:18",
      "content": "<p>Thanks for sharing and congrats for gold!  I'm not surprised RF works well, I even had a KNN that worked quite well ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543272,
      "author_name": "timon88",
      "author_url": "",
      "post_date": "06/04/2019 12:33:39",
      "content": "<p>Congratulations! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543576,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 15:48:05",
      "content": "<p>Congrats <a href=\"/miguelpm\">@miguelpm</a> . And thanks for sharing. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543733,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "06/04/2019 18:23:49",
      "content": "<p>Great !!!!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 543907,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "06/04/2019 23:36:00",
      "content": "<p>Congrats! I've learned a lot from this post, because there are many what you did with what you thought. Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546494,
      "author_name": "rachidabida",
      "author_url": "",
      "post_date": "06/06/2019 16:10:25",
      "content": "<p>Interesting, thanks for  sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 573614,
      "author_name": "jaiwei98",
      "author_url": "",
      "post_date": "07/12/2019 14:03:53",
      "content": "<p>Thanks for sharing <a href=\"/miguelpm\">@miguelpm</a> , may I know how you use Matrix Profile to create new features as you mentioned. Link to some docs should be enough. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 573824,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "07/12/2019 20:02:13",
          "content": "<p>Sure, it's a very interesting tool, here it's all you need to learn about it <a href=\"https://www.cs.ucr.edu/~eamonn/MatrixProfile.html\">https://www.cs.ucr.edu/~eamonn/MatrixProfile.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 585725,
      "author_name": "keenborder",
      "author_url": "",
      "post_date": "07/28/2019 00:03:53",
      "content": "<p>Interesting and smart </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 589434,
      "author_name": "kerzer",
      "author_url": "",
      "post_date": "07/31/2019 22:13:47",
      "content": "<p>nice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543012": "To begin with, I want to congratulate all competitors regardless of leaderboard ranking for what has sure been a lot of work this months. In many cases it may feel to be unrewarded. Don't feel despair, just make sure to learn as much as possible including final wrapups and  it will have been worth it. \n\nSo, here are my hints, hoping that will be useful for someone:\n\n### GENERAL IDEAS:\n\n-  Assessing the amount of noise in a competition is the key to generalization. (And it is never easy)\n\n- You can win without stacking, and even without a too complex model, as long as you frame the problem correctly. \n \n- Don't be trapped by the Public Leaderboard race, just treat it as another fold of your validation or -sometimes- ignore it altogether\n\n### MY -CONDENSED- PATH TO THIS GOLD:\n\n- First month:  lost time trying to make a RNN work in a thousand of creative ways. After lots of work had to bite the bullet:  RNN was not the way to go. **Difficult decision**: Start from scratch.\n\n- Second month: The fight to find a proper validation scheme, the realization of just how overfittable the dataset was. Lots of feature engineering, most of it the kind already seen in forums, only \"original\" ones some computationally expensive Matrix Profile stats that turned out not to be too useful. **Difficult decision**:  Use random forest. (Less tunning, less sparsity of feature importance than boosting, important to make also feature selection of the less overfitting ones).  \n\n- Last three days: As [this post](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844)  disclosed there was a potential leak that showed train-test relationship. Not a too easy to leverage leak, and risky to use though. Seeing how important it has been in the end some kind of systematic processing of the image would have been extremely effective, but  I didn't do that. Instead I simply gained a couple of insights: \n\nFirst insight: outliers removal. Outliers are not just out of range values but more in general train data that is unusual vs. what is to be expected. I realized that the Earthques 5 and 6 were actually a \"synthetic\" split of what in the experiment was considered a single earhtquake, i.e. the stress theshold was different . **Difficult decision**: to remove those two earthquakes data of an already small dataset.\n\nSecond insight: I had already observed that bimodal distribution of predictions was very different between train and test showing  that shorter earquakes in test were way longer than shorter earthquakes in train. That was an ugly warning, and fitted quite well with what at a glance could be seen in the experiment. I visually estimated the proportion and lowered the weight of \"small\" earthquakes in training set by 1/5, no more fancy adjustment than that,  noise was too big for more fine grained tunning anyway. **Last unconfortable decision**: select a submission that placed me like 2300th in public leaderboard cause it made sense.\n\n### ONE LAST WARNING\n\nSo, that was it. Hope it was of some interest. One last warning, don't take any of this decissions to rigidly (like using Random Forest, for example, I never choose a RF without a reason to do it). **In the end, it is all about generalization and every competition is different**. \n\nAs I said in the beginning, just make sure to keep on learning and to have fun :-)",
    "543030": "Thanks for sharing @miguelpm",
    "543046": "Congrats! Thanks for sharing :)",
    "543078": "Thanks for sharing and congrats for gold!  I'm not surprised RF works well, I even had a KNN that worked quite well ;)",
    "543272": "Congratulations! :)",
    "543576": "Congrats @miguelpm . And thanks for sharing.",
    "543733": "Great !!!!!!!",
    "543907": "Congrats! I've learned a lot from this post, because there are many what you did with what you thought. Thank you for sharing!",
    "546494": "Interesting, thanks for  sharing",
    "573614": "Thanks for sharing @miguelpm , may I know how you use Matrix Profile to create new features as you mentioned. Link to some docs should be enough. Thanks.",
    "573824": "Sure, it's a very interesting tool, here it's all you need to learn about it https://www.cs.ucr.edu/~eamonn/MatrixProfile.html",
    "585725": "Interesting and smart",
    "589434": "nice"
  },
  "source": "meta"
}