{
  "id": 94594,
  "title": "19th place solution (GBDT + post-processing)",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/yahatan-19th-place-solution-gbdt-post-processing",
  "author_name": "",
  "post_date": "2019-06-05T14:19:56.567Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my brilliant teammates.</p>\n\n<p>In this thread, I want to share the detail of our submission finally ranked 19th place. Our another submission is <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94450#latest-544035\">described by my teammate</a>. Please check it.</p>\n\n<h2>Model</h2>\n\n<p>The prediction is made by averaging of LightGBM, XGBoost and CatBoost. Each model is averaged by 25 different seeds. Hyperparameters are borrowed from kernels by <a href=\"https://www.kaggle.com/artgor/seismic-data-eda-and-baseline\"></a><a href=\"/artgor\">@artgor</a> and <a href=\"https://www.kaggle.com/gideonvos/earthquake-prediction-with-xgboost-lb1-496\"></a><a href=\"/gideonvos\">@gideonvos</a></p>\n\n<p>We used quake based 5-fold split. The distribution of target values can differs a lot between the folds. Early stopping in this setting easily leads overfitting to validation set. So we use\n - The same number of rounds of GBDT through folds (like lightgbm.cv)\n - boost_from_mean = True (or its equivalent) in all models (discussed in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91500#latest-528287\">this thread</a>)\n - Splits manually picked from all possible splits of 15 quakes into 5 folds. We picked splits which estimated to have low between-fold variance and mean of validation score.</p>\n\n<p>We also trained models predict TSF (time since failure). We used both TTF and TSF in post-processing.</p>\n\n<h2>Features</h2>\n\n<p>We used feature from <a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">Masters Final Project</a>. Some of them have very different distribution between train and test. So we trained GBDTs which use a single feature, did adversarial validation, and discarded features achieve higher AUC than 0.55.</p>\n\n<p>Then, we did a forward feature selection, which iteratively adds a feature with the best gain. It converges after selecting 15 features.</p>\n\n<h2>Post-processing</h2>\n\n<p>As already discussed, the test TTF distribution can be estimated using the figure shown in the organizer's paper. The figure is a <a href=\"https://en.wikipedia.org/wiki/Vector_graphics\">vector graphics</a>. So it can be expanded arbitrarily without loosing image quality. Using this, we measured peak to peak distance of sheer stress manually and dug out TTFs after train set as the below image. (Of course, it may have some small errors)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13411/estimated_test.png\" alt=\"estimated_test\"></p>\n\n<p>Blue line shows estimated TTFs. Then, we estimated where the test set is. Orange line shows MAEs between public LB scores for the constant values 0-10 (shown in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268\">this discussion</a>) and scores for them calculated in public LB size segments.  (These inferences are based on the assumption that the test set is made from one large segments and each instance is adjacent to another instance. We confirmed the test set has no overlap by brute-force matching) Red lines show estimated starting point of Public LB, starting and ending point of Private LB. Surprisingly, it almost matched the TTF expected by <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268\"></a><a href=\"/mykper\">@mykper</a>.</p>\n\n<p>Finally, we calculated the median of peak values weighted by the length of the quakes in test set and make the prediction as follows.</p>\n\n<p><code>\npred = np.where(pred_ttf &amp;lt; 4, pred_ttf, weithed_median - pred_tsf)\n</code></p>\n\n<p>We think predicting large part of TSF is difficult and predicting small part of TTF is relatively easy. So we use raw TTF if it is smaller than 4 and median - TSF otherwise. </p>\n\n<p>Post-processed oof prediction to training set:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13412/postprocessed_valid.png\" alt=\"postprocessed_valid\"></p>\n\n<p>What I regret is that it was late to find this and it did not come up with a clever use. I am so surprised by the interesting use of it in the solution of other participants.</p>\n\n<p>Our final submission is scored 2.10809 in Public LB (2.40159 in Private LB). I want to thank my teammates for agreeing to choose such a risky submission.</p>\n\n<p>That's all. Thank you for everyone.</p>",
  "messages": [
    {
      "id": "544429",
      "postDate": "06/05/2019 13:54:55",
      "content": "<p>First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my brilliant teammates.</p>\n\n<p>In this thread, I want to share the detail of our submission finally ranked 19th place. Our another submission is <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94450#latest-544035\">described by my teammate</a>. Please check it.</p>\n\n<h2>Model</h2>\n\n<p>The prediction is made by averaging of LightGBM, XGBoost and CatBoost. Each model is averaged by 25 different seeds. Hyperparameters are borrowed from kernels by <a href=\"https://www.kaggle.com/artgor/seismic-data-eda-and-baseline\"></a><a href=\"/artgor\">@artgor</a> and <a href=\"https://www.kaggle.com/gideonvos/earthquake-prediction-with-xgboost-lb1-496\"></a><a href=\"/gideonvos\">@gideonvos</a></p>\n\n<p>We used quake based 5-fold split. The distribution of target values can differs a lot between the folds. Early stopping in this setting easily leads overfitting to validation set. So we use\n - The same number of rounds of GBDT through folds (like lightgbm.cv)\n - boost_from_mean = True (or its equivalent) in all models (discussed in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91500#latest-528287\">this thread</a>)\n - Splits manually picked from all possible splits of 15 quakes into 5 folds. We picked splits which estimated to have low between-fold variance and mean of validation score.</p>\n\n<p>We also trained models predict TSF (time since failure). We used both TTF and TSF in post-processing.</p>\n\n<h2>Features</h2>\n\n<p>We used feature from <a href=\"https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392\">Masters Final Project</a>. Some of them have very different distribution between train and test. So we trained GBDTs which use a single feature, did adversarial validation, and discarded features achieve higher AUC than 0.55.</p>\n\n<p>Then, we did a forward feature selection, which iteratively adds a feature with the best gain. It converges after selecting 15 features.</p>\n\n<h2>Post-processing</h2>\n\n<p>As already discussed, the test TTF distribution can be estimated using the figure shown in the organizer's paper. The figure is a <a href=\"https://en.wikipedia.org/wiki/Vector_graphics\">vector graphics</a>. So it can be expanded arbitrarily without loosing image quality. Using this, we measured peak to peak distance of sheer stress manually and dug out TTFs after train set as the below image. (Of course, it may have some small errors)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13411/estimated_test.png\" alt=\"estimated_test\"></p>\n\n<p>Blue line shows estimated TTFs. Then, we estimated where the test set is. Orange line shows MAEs between public LB scores for the constant values 0-10 (shown in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268\">this discussion</a>) and scores for them calculated in public LB size segments.  (These inferences are based on the assumption that the test set is made from one large segments and each instance is adjacent to another instance. We confirmed the test set has no overlap by brute-force matching) Red lines show estimated starting point of Public LB, starting and ending point of Private LB. Surprisingly, it almost matched the TTF expected by <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268\"></a><a href=\"/mykper\">@mykper</a>.</p>\n\n<p>Finally, we calculated the median of peak values weighted by the length of the quakes in test set and make the prediction as follows.</p>\n\n<p><code>\npred = np.where(pred_ttf &amp;lt; 4, pred_ttf, weithed_median - pred_tsf)\n</code></p>\n\n<p>We think predicting large part of TSF is difficult and predicting small part of TTF is relatively easy. So we use raw TTF if it is smaller than 4 and median - TSF otherwise. </p>\n\n<p>Post-processed oof prediction to training set:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13412/postprocessed_valid.png\" alt=\"postprocessed_valid\"></p>\n\n<p>What I regret is that it was late to find this and it did not come up with a clever use. I am so surprised by the interesting use of it in the solution of other participants.</p>\n\n<p>Our final submission is scored 2.10809 in Public LB (2.40159 in Private LB). I want to thank my teammates for agreeing to choose such a risky submission.</p>\n\n<p>That's all. Thank you for everyone.</p>",
      "rawMarkdown": "First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my brilliant teammates.\n\nIn this thread, I want to share the detail of our submission finally ranked 19th place. Our another submission is [described by my teammate](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94450#latest-544035). Please check it.\n\nModel\n---\n\nThe prediction is made by averaging of LightGBM, XGBoost and CatBoost. Each model is averaged by 25 different seeds. Hyperparameters are borrowed from kernels by [@artgor](https://www.kaggle.com/artgor/seismic-data-eda-and-baseline) and [@gideonvos](https://www.kaggle.com/gideonvos/earthquake-prediction-with-xgboost-lb1-496)\n\nWe used quake based 5-fold split. The distribution of target values can differs a lot between the folds. Early stopping in this setting easily leads overfitting to validation set. So we use\n - The same number of rounds of GBDT through folds (like lightgbm.cv)\n - boost\\_from\\_mean = True (or its equivalent) in all models (discussed in [this thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91500#latest-528287))\n - Splits manually picked from all possible splits of 15 quakes into 5 folds. We picked splits which estimated to have low between-fold variance and mean of validation score.\n\nWe also trained models predict TSF (time since failure). We used both TTF and TSF in post-processing.\n\nFeatures\n---\n\nWe used feature from [Masters Final Project](https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392). Some of them have very different distribution between train and test. So we trained GBDTs which use a single feature, did adversarial validation, and discarded features achieve higher AUC than 0.55.\n\nThen, we did a forward feature selection, which iteratively adds a feature with the best gain. It converges after selecting 15 features.\n\nPost-processing\n---\nAs already discussed, the test TTF distribution can be estimated using the figure shown in the organizer's paper. The figure is a [vector graphics](https://en.wikipedia.org/wiki/Vector_graphics). So it can be expanded arbitrarily without loosing image quality. Using this, we measured peak to peak distance of sheer stress manually and dug out TTFs after train set as the below image. (Of course, it may have some small errors)\n\n![estimated_test](https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13411/estimated_test.png)\n\nBlue line shows estimated TTFs. Then, we estimated where the test set is. Orange line shows MAEs between public LB scores for the constant values 0-10 (shown in [this discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268)) and scores for them calculated in public LB size segments.  (These inferences are based on the assumption that the test set is made from one large segments and each instance is adjacent to another instance. We confirmed the test set has no overlap by brute-force matching) Red lines show estimated starting point of Public LB, starting and ending point of Private LB. Surprisingly, it almost matched the TTF expected by [@mykper](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268).\n\nFinally, we calculated the median of peak values weighted by the length of the quakes in test set and make the prediction as follows.\n\n```\npred = np.where(pred_ttf &lt; 4, pred_ttf, weithed_median - pred_tsf)\n```\n\nWe think predicting large part of TSF is difficult and predicting small part of TTF is relatively easy. So we use raw TTF if it is smaller than 4 and median - TSF otherwise. \n\nPost-processed oof prediction to training set:\n![postprocessed_valid](https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13412/postprocessed_valid.png)\n\nWhat I regret is that it was late to find this and it did not come up with a clever use. I am so surprised by the interesting use of it in the solution of other participants.\n\nOur final submission is scored 2.10809 in Public LB (2.40159 in Private LB). I want to thank my teammates for agreeing to choose such a risky submission.\n\nThat's all. Thank you for everyone.",
      "votes": null
    },
    {
      "id": "544446",
      "postDate": "06/05/2019 14:16:44",
      "content": "<p>Thanks for sharing and congrats on the result!</p>",
      "rawMarkdown": "Thanks for sharing and congrats on the result!",
      "votes": null
    },
    {
      "id": "544454",
      "postDate": "06/05/2019 14:23:35",
      "content": "<p>Thanks. I learned a lot from your post in discussion.</p>",
      "rawMarkdown": "Thanks. I learned a lot from your post in discussion.",
      "votes": null
    },
    {
      "id": "544678",
      "postDate": "06/05/2019 19:00:40",
      "content": "<p>Congrats and thanks for sharing <a href=\"/zaburo\">@zaburo</a> !</p>",
      "rawMarkdown": "Congrats and thanks for sharing @zaburo !",
      "votes": null
    },
    {
      "id": "544755",
      "postDate": "06/05/2019 21:44:59",
      "content": "<p>Congrats and glad my parameters could help. Unfortunately I dropped to bronze after making the huge mistake of one last submission which ended up overfitting in private. Had I stuck to my original I would have landed silver too. You live and learn I guess!</p>",
      "rawMarkdown": "Congrats and glad my parameters could help. Unfortunately I dropped to bronze after making the huge mistake of one last submission which ended up overfitting in private. Had I stuck to my original I would have landed silver too. You live and learn I guess!",
      "votes": null
    },
    {
      "id": "546108",
      "postDate": "06/06/2019 09:05:12",
      "content": "<p>Congrats!  And Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congrats!  And Thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "547714",
      "postDate": "06/08/2019 06:57:01",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 544446,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2019 14:16:44",
      "content": "<p>Thanks for sharing and congrats on the result!</p>",
      "votes": null,
      "replies": [
        {
          "id": 544454,
          "author_name": "zaburo",
          "author_url": "",
          "post_date": "06/05/2019 14:23:35",
          "content": "<p>Thanks. I learned a lot from your post in discussion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544678,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/05/2019 19:00:40",
      "content": "<p>Congrats and thanks for sharing <a href=\"/zaburo\">@zaburo</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544755,
      "author_name": "gideonvos",
      "author_url": "",
      "post_date": "06/05/2019 21:44:59",
      "content": "<p>Congrats and glad my parameters could help. Unfortunately I dropped to bronze after making the huge mistake of one last submission which ended up overfitting in private. Had I stuck to my original I would have landed silver too. You live and learn I guess!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546108,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "06/06/2019 09:05:12",
      "content": "<p>Congrats!  And Thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547714,
      "author_name": "prashantkikani",
      "author_url": "",
      "post_date": "06/08/2019 06:57:01",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "544429": "First of all, thanks to the organizers and all participants in this competition! And thanks a lot to my brilliant teammates.\n\nIn this thread, I want to share the detail of our submission finally ranked 19th place. Our another submission is [described by my teammate](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94450#latest-544035). Please check it.\n\nModel\n---\n\nThe prediction is made by averaging of LightGBM, XGBoost and CatBoost. Each model is averaged by 25 different seeds. Hyperparameters are borrowed from kernels by [@artgor](https://www.kaggle.com/artgor/seismic-data-eda-and-baseline) and [@gideonvos](https://www.kaggle.com/gideonvos/earthquake-prediction-with-xgboost-lb1-496)\n\nWe used quake based 5-fold split. The distribution of target values can differs a lot between the folds. Early stopping in this setting easily leads overfitting to validation set. So we use\n - The same number of rounds of GBDT through folds (like lightgbm.cv)\n - boost\\_from\\_mean = True (or its equivalent) in all models (discussed in [this thread](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91500#latest-528287))\n - Splits manually picked from all possible splits of 15 quakes into 5 folds. We picked splits which estimated to have low between-fold variance and mean of validation score.\n\nWe also trained models predict TSF (time since failure). We used both TTF and TSF in post-processing.\n\nFeatures\n---\n\nWe used feature from [Masters Final Project](https://www.kaggle.com/vettejeep/masters-final-project-model-lb-1-392). Some of them have very different distribution between train and test. So we trained GBDTs which use a single feature, did adversarial validation, and discarded features achieve higher AUC than 0.55.\n\nThen, we did a forward feature selection, which iteratively adds a feature with the best gain. It converges after selecting 15 features.\n\nPost-processing\n---\nAs already discussed, the test TTF distribution can be estimated using the figure shown in the organizer's paper. The figure is a [vector graphics](https://en.wikipedia.org/wiki/Vector_graphics). So it can be expanded arbitrarily without loosing image quality. Using this, we measured peak to peak distance of sheer stress manually and dug out TTFs after train set as the below image. (Of course, it may have some small errors)\n\n![estimated_test](https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13411/estimated_test.png)\n\nBlue line shows estimated TTFs. Then, we estimated where the test set is. Orange line shows MAEs between public LB scores for the constant values 0-10 (shown in [this discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268)) and scores for them calculated in public LB size segments.  (These inferences are based on the assumption that the test set is made from one large segments and each instance is adjacent to another instance. We confirmed the test set has no overlap by brute-force matching) Red lines show estimated starting point of Public LB, starting and ending point of Private LB. Surprisingly, it almost matched the TTF expected by [@mykper](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583#latest-535268).\n\nFinally, we calculated the median of peak values weighted by the length of the quakes in test set and make the prediction as follows.\n\n```\npred = np.where(pred_ttf &lt; 4, pred_ttf, weithed_median - pred_tsf)\n```\n\nWe think predicting large part of TSF is difficult and predicting small part of TTF is relatively easy. So we use raw TTF if it is smaller than 4 and median - TSF otherwise. \n\nPost-processed oof prediction to training set:\n![postprocessed_valid](https://storage.googleapis.com/kaggle-forum-message-attachments/544429/13412/postprocessed_valid.png)\n\nWhat I regret is that it was late to find this and it did not come up with a clever use. I am so surprised by the interesting use of it in the solution of other participants.\n\nOur final submission is scored 2.10809 in Public LB (2.40159 in Private LB). I want to thank my teammates for agreeing to choose such a risky submission.\n\nThat's all. Thank you for everyone.",
    "544446": "Thanks for sharing and congrats on the result!",
    "544454": "Thanks. I learned a lot from your post in discussion.",
    "544678": "Congrats and thanks for sharing @zaburo !",
    "544755": "Congrats and glad my parameters could help. Unfortunately I dropped to bronze after making the huge mistake of one last submission which ended up overfitting in private. Had I stuck to my original I would have landed silver too. You live and learn I guess!",
    "546108": "Congrats!  And Thanks for sharing your solution.",
    "547714": "Thanks!"
  },
  "source": "meta"
}