{
  "id": 94570,
  "title": "87th place private and 1st place public solution",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/kopeyka-87th-place-private-and-1st-place-public-so",
  "author_name": "",
  "post_date": "2019-06-05T10:10:30.623Z",
  "votes": 23,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I want to thank Kaggle and the great community of this competition participants for making this journey possible! This is a write-up of our work with my teammate <a href=\"/idog90\">@idog90</a>. </p>\n\n<p>First, please check out <a href=\"https://medium.com/@zaharchikishev/how-to-lb-probe-on-kaggle-c0aa21458bfe?source=friends_link&amp;sk=543b4ebcafee24697979b69efdb35adf\">the LB probing strategy</a> that gave us the public leader-board first place. In this post I will describe the other components of our solution, focusing on the submission that gave us 87th place.</p>\n\n<p><strong>features and models</strong>\nWe generated 270 features, mostly taken from the public kernels with some modifications. By using recursive feature elimination with permutation importance by LGBM we reduced the feature set to 90 features. Regarding the segment mean, we subtracted the mean from the signal before calculating the features, but left it as an additional feature because it improved both CV and LB. Leaving it was possibly a critical mistake. \nModels that participated are LGBM, XGB, CatBoost, SVR and shallow NN by fastai. Additionally we added a median of classification output of LGBM, by dividing the targets space into 11 classes. \nThe 6 models oof predictions were blended together by linear stacking with MAE objective function and non-negativity constraints. Using 6 models and stacking gave more or less similar scores to the best single model scores (XGB or CatBoost), so this setup is probably redundant but no harm.</p>\n\n<p><strong>CV strategy</strong>\nWe used 25k stride to generate train segments, and then assigned every 35 consecutive segment into a group of its own. Then we removed the segments with data overlapping into the next group, and used stratified group 5-fold CV strategy with the groups as defined above. I think it worked well for us. </p>\n\n<p><strong>secret weapon</strong>\nWe used two extra elements that gave us the medal. The second final submission didn't include it and scored somewhere around 500 on the private.</p>\n\n<p>a. First, we merged short quakes 5 and 6 together, considering them as a single long quake. This is of course to make the train distribution to be more similar to the test. This is the change that made the difference.</p>\n\n<p>b. Second, we added public test segments to the train with their best estimated scores from the LB probing strategy linked above. I don't think it had much impact, it was less than 2% of the train data, and noisy. </p>\n\n<p><strong>using MIP solution</strong>\nAs also mentioned in the medium article linked above, our mixed integer programming (MIP) solution for the LB probing had a nice ability to calculate public LB scores of potential submissions with great precision (0.02). In a sense, the MIP model learnt very well the subspace of all \"normal\" submissions and therefore gave excellent estimations for anything in that subspace. We used it to estimate public LB scores for our final submissions candidates on the last day.</p>\n\n<p><strong>what didn't work</strong>\nHuge amount of things that we tried didn't work well. To name the major ones:</p>\n\n<p>a. <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94435#latest-543574\">batch start identification</a> didn't help\nb. transformer NNs on raw data gave bad CV and LB scores\nc. spectograms with CNN was OK but worse than our best models and we didn't include it at the end.\ne. reweighting of the train set into the known test distribution, fitting histograms of the classification models to match the test distribution, - didn't work well on test, still an enigma for me.\nf. genetic programming, - but we are not experts, tried only a simple configuration\ng. more complex stacking, - seems to overfit</p>\n\n<p>Thanks for reading that far, and hopefully see you in the future competitions!</p>",
  "messages": [
    {
      "id": "544217",
      "postDate": "06/05/2019 09:23:45",
      "content": "<p>I want to thank Kaggle and the great community of this competition participants for making this journey possible! This is a write-up of our work with my teammate <a href=\"/idog90\">@idog90</a>. </p>\n\n<p>First, please check out <a href=\"https://medium.com/@zaharchikishev/how-to-lb-probe-on-kaggle-c0aa21458bfe?source=friends_link&amp;sk=543b4ebcafee24697979b69efdb35adf\">the LB probing strategy</a> that gave us the public leader-board first place. In this post I will describe the other components of our solution, focusing on the submission that gave us 87th place.</p>\n\n<p><strong>features and models</strong>\nWe generated 270 features, mostly taken from the public kernels with some modifications. By using recursive feature elimination with permutation importance by LGBM we reduced the feature set to 90 features. Regarding the segment mean, we subtracted the mean from the signal before calculating the features, but left it as an additional feature because it improved both CV and LB. Leaving it was possibly a critical mistake. \nModels that participated are LGBM, XGB, CatBoost, SVR and shallow NN by fastai. Additionally we added a median of classification output of LGBM, by dividing the targets space into 11 classes. \nThe 6 models oof predictions were blended together by linear stacking with MAE objective function and non-negativity constraints. Using 6 models and stacking gave more or less similar scores to the best single model scores (XGB or CatBoost), so this setup is probably redundant but no harm.</p>\n\n<p><strong>CV strategy</strong>\nWe used 25k stride to generate train segments, and then assigned every 35 consecutive segment into a group of its own. Then we removed the segments with data overlapping into the next group, and used stratified group 5-fold CV strategy with the groups as defined above. I think it worked well for us. </p>\n\n<p><strong>secret weapon</strong>\nWe used two extra elements that gave us the medal. The second final submission didn't include it and scored somewhere around 500 on the private.</p>\n\n<p>a. First, we merged short quakes 5 and 6 together, considering them as a single long quake. This is of course to make the train distribution to be more similar to the test. This is the change that made the difference.</p>\n\n<p>b. Second, we added public test segments to the train with their best estimated scores from the LB probing strategy linked above. I don't think it had much impact, it was less than 2% of the train data, and noisy. </p>\n\n<p><strong>using MIP solution</strong>\nAs also mentioned in the medium article linked above, our mixed integer programming (MIP) solution for the LB probing had a nice ability to calculate public LB scores of potential submissions with great precision (0.02). In a sense, the MIP model learnt very well the subspace of all \"normal\" submissions and therefore gave excellent estimations for anything in that subspace. We used it to estimate public LB scores for our final submissions candidates on the last day.</p>\n\n<p><strong>what didn't work</strong>\nHuge amount of things that we tried didn't work well. To name the major ones:</p>\n\n<p>a. <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94435#latest-543574\">batch start identification</a> didn't help\nb. transformer NNs on raw data gave bad CV and LB scores\nc. spectograms with CNN was OK but worse than our best models and we didn't include it at the end.\ne. reweighting of the train set into the known test distribution, fitting histograms of the classification models to match the test distribution, - didn't work well on test, still an enigma for me.\nf. genetic programming, - but we are not experts, tried only a simple configuration\ng. more complex stacking, - seems to overfit</p>\n\n<p>Thanks for reading that far, and hopefully see you in the future competitions!</p>",
      "rawMarkdown": "I want to thank Kaggle and the great community of this competition participants for making this journey possible! This is a write-up of our work with my teammate @idog90. \n\nFirst, please check out [the LB probing strategy](https://medium.com/@zaharchikishev/how-to-lb-probe-on-kaggle-c0aa21458bfe?source=friends_link&amp;sk=543b4ebcafee24697979b69efdb35adf) that gave us the public leader-board first place. In this post I will describe the other components of our solution, focusing on the submission that gave us 87th place.\n\n**features and models**\nWe generated 270 features, mostly taken from the public kernels with some modifications. By using recursive feature elimination with permutation importance by LGBM we reduced the feature set to 90 features. Regarding the segment mean, we subtracted the mean from the signal before calculating the features, but left it as an additional feature because it improved both CV and LB. Leaving it was possibly a critical mistake. \nModels that participated are LGBM, XGB, CatBoost, SVR and shallow NN by fastai. Additionally we added a median of classification output of LGBM, by dividing the targets space into 11 classes. \nThe 6 models oof predictions were blended together by linear stacking with MAE objective function and non-negativity constraints. Using 6 models and stacking gave more or less similar scores to the best single model scores (XGB or CatBoost), so this setup is probably redundant but no harm.\n\n**CV strategy**\nWe used 25k stride to generate train segments, and then assigned every 35 consecutive segment into a group of its own. Then we removed the segments with data overlapping into the next group, and used stratified group 5-fold CV strategy with the groups as defined above. I think it worked well for us. \n\n**secret weapon**\nWe used two extra elements that gave us the medal. The second final submission didn't include it and scored somewhere around 500 on the private.\n\na. First, we merged short quakes 5 and 6 together, considering them as a single long quake. This is of course to make the train distribution to be more similar to the test. This is the change that made the difference.\n\nb. Second, we added public test segments to the train with their best estimated scores from the LB probing strategy linked above. I don't think it had much impact, it was less than 2% of the train data, and noisy. \n\n**using MIP solution**\nAs also mentioned in the medium article linked above, our mixed integer programming (MIP) solution for the LB probing had a nice ability to calculate public LB scores of potential submissions with great precision (0.02). In a sense, the MIP model learnt very well the subspace of all \"normal\" submissions and therefore gave excellent estimations for anything in that subspace. We used it to estimate public LB scores for our final submissions candidates on the last day.\n\n**what didn't work**\nHuge amount of things that we tried didn't work well. To name the major ones:\n\na. [batch start identification](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94435#latest-543574) didn't help\nb. transformer NNs on raw data gave bad CV and LB scores\nc. spectograms with CNN was OK but worse than our best models and we didn't include it at the end.\ne. reweighting of the train set into the known test distribution, fitting histograms of the classification models to match the test distribution, - didn't work well on test, still an enigma for me.\nf. genetic programming, - but we are not experts, tried only a simple configuration\ng. more complex stacking, - seems to overfit\n\n\nThanks for reading that far, and hopefully see you in the future competitions!",
      "votes": null
    },
    {
      "id": "544248",
      "postDate": "06/05/2019 09:57:17",
      "content": "<p>Congrats on the result, even if you may feel disappointed.  I love the use of MIP given my history with CPLEX ;)  Very inventive way of using mathematical optimizaiton to attack a ML problem.</p>\n\n<p>And I am glad someone actually tried this:</p>\n\n<blockquote>\n  <p>we merged short quakes 5 and 6 together</p>\n</blockquote>\n\n<p>It was blatant in some of the papers, yet I am not sure many teams did it.  We did not, and this may be the single most effective change to make to improve our models significantly.  Well done!</p>",
      "rawMarkdown": "Congrats on the result, even if you may feel disappointed.  I love the use of MIP given my history with CPLEX ;)  Very inventive way of using mathematical optimizaiton to attack a ML problem.\n\nAnd I am glad someone actually tried this:\n\n&gt; we merged short quakes 5 and 6 together\n\nIt was blatant in some of the papers, yet I am not sure many teams did it.  We did not, and this may be the single most effective change to make to improve our models significantly.  Well done!",
      "votes": null
    },
    {
      "id": "544258",
      "postDate": "06/05/2019 10:16:00",
      "content": "<p>thank you! We did quite well, after all. It was so easy to drop into no-medals zone. </p>\n\n<p>I used CPLEX for 2 years while working at IBM research, and love it. My major is also in optimization, so I love to use these tools.</p>\n\n<p>Congratulations on your results, hopefully we will compete once again. </p>",
      "rawMarkdown": "thank you! We did quite well, after all. It was so easy to drop into no-medals zone. \n\nI used CPLEX for 2 years while working at IBM research, and love it. My major is also in optimization, so I love to use these tools.\n\nCongratulations on your results, hopefully we will compete once again.",
      "votes": null
    },
    {
      "id": "544269",
      "postDate": "06/05/2019 10:41:11",
      "content": "<p>Merging the quakes makes the mean higher and thus made train data closer to test indeed.</p>",
      "rawMarkdown": "Merging the quakes makes the mean higher and thus made train data closer to test indeed.",
      "votes": null
    },
    {
      "id": "544362",
      "postDate": "06/05/2019 12:40:57",
      "content": "<p>This is really  interesting, including that MIP solution. </p>\n\n<p>To merge earthquakes 5 and 6 was clever. I considered it but in the end decided to treat them as potentially harmful anomalies and remove them both from train. Anyway those two earthquakes were something to take care of so... clever move!</p>\n\n<p>Thanks for sharing, and congratulations!</p>",
      "rawMarkdown": "This is really  interesting, including that MIP solution. \n\nTo merge earthquakes 5 and 6 was clever. I considered it but in the end decided to treat them as potentially harmful anomalies and remove them both from train. Anyway those two earthquakes were something to take care of so... clever move!\n\nThanks for sharing, and congratulations!",
      "votes": null
    },
    {
      "id": "545885",
      "postDate": "06/06/2019 02:19:00",
      "content": "<p>Regarding the merge of earthquakes, I thought my eyes were seeing things in the papers. I kept going back and forth between the images in the papers and my own plots and was wondering if there was a mistake somewhere.</p>",
      "rawMarkdown": "Regarding the merge of earthquakes, I thought my eyes were seeing things in the papers. I kept going back and forth between the images in the papers and my own plots and was wondering if there was a mistake somewhere.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 544248,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2019 09:57:17",
      "content": "<p>Congrats on the result, even if you may feel disappointed.  I love the use of MIP given my history with CPLEX ;)  Very inventive way of using mathematical optimizaiton to attack a ML problem.</p>\n\n<p>And I am glad someone actually tried this:</p>\n\n<blockquote>\n  <p>we merged short quakes 5 and 6 together</p>\n</blockquote>\n\n<p>It was blatant in some of the papers, yet I am not sure many teams did it.  We did not, and this may be the single most effective change to make to improve our models significantly.  Well done!</p>",
      "votes": null,
      "replies": [
        {
          "id": 544258,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "06/05/2019 10:16:00",
          "content": "<p>thank you! We did quite well, after all. It was so easy to drop into no-medals zone. </p>\n\n<p>I used CPLEX for 2 years while working at IBM research, and love it. My major is also in optimization, so I love to use these tools.</p>\n\n<p>Congratulations on your results, hopefully we will compete once again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544269,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/05/2019 10:41:11",
          "content": "<p>Merging the quakes makes the mean higher and thus made train data closer to test indeed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544362,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "06/05/2019 12:40:57",
      "content": "<p>This is really  interesting, including that MIP solution. </p>\n\n<p>To merge earthquakes 5 and 6 was clever. I considered it but in the end decided to treat them as potentially harmful anomalies and remove them both from train. Anyway those two earthquakes were something to take care of so... clever move!</p>\n\n<p>Thanks for sharing, and congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 545885,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "06/06/2019 02:19:00",
      "content": "<p>Regarding the merge of earthquakes, I thought my eyes were seeing things in the papers. I kept going back and forth between the images in the papers and my own plots and was wondering if there was a mistake somewhere.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "544217": "I want to thank Kaggle and the great community of this competition participants for making this journey possible! This is a write-up of our work with my teammate @idog90. \n\nFirst, please check out [the LB probing strategy](https://medium.com/@zaharchikishev/how-to-lb-probe-on-kaggle-c0aa21458bfe?source=friends_link&amp;sk=543b4ebcafee24697979b69efdb35adf) that gave us the public leader-board first place. In this post I will describe the other components of our solution, focusing on the submission that gave us 87th place.\n\n**features and models**\nWe generated 270 features, mostly taken from the public kernels with some modifications. By using recursive feature elimination with permutation importance by LGBM we reduced the feature set to 90 features. Regarding the segment mean, we subtracted the mean from the signal before calculating the features, but left it as an additional feature because it improved both CV and LB. Leaving it was possibly a critical mistake. \nModels that participated are LGBM, XGB, CatBoost, SVR and shallow NN by fastai. Additionally we added a median of classification output of LGBM, by dividing the targets space into 11 classes. \nThe 6 models oof predictions were blended together by linear stacking with MAE objective function and non-negativity constraints. Using 6 models and stacking gave more or less similar scores to the best single model scores (XGB or CatBoost), so this setup is probably redundant but no harm.\n\n**CV strategy**\nWe used 25k stride to generate train segments, and then assigned every 35 consecutive segment into a group of its own. Then we removed the segments with data overlapping into the next group, and used stratified group 5-fold CV strategy with the groups as defined above. I think it worked well for us. \n\n**secret weapon**\nWe used two extra elements that gave us the medal. The second final submission didn't include it and scored somewhere around 500 on the private.\n\na. First, we merged short quakes 5 and 6 together, considering them as a single long quake. This is of course to make the train distribution to be more similar to the test. This is the change that made the difference.\n\nb. Second, we added public test segments to the train with their best estimated scores from the LB probing strategy linked above. I don't think it had much impact, it was less than 2% of the train data, and noisy. \n\n**using MIP solution**\nAs also mentioned in the medium article linked above, our mixed integer programming (MIP) solution for the LB probing had a nice ability to calculate public LB scores of potential submissions with great precision (0.02). In a sense, the MIP model learnt very well the subspace of all \"normal\" submissions and therefore gave excellent estimations for anything in that subspace. We used it to estimate public LB scores for our final submissions candidates on the last day.\n\n**what didn't work**\nHuge amount of things that we tried didn't work well. To name the major ones:\n\na. [batch start identification](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94435#latest-543574) didn't help\nb. transformer NNs on raw data gave bad CV and LB scores\nc. spectograms with CNN was OK but worse than our best models and we didn't include it at the end.\ne. reweighting of the train set into the known test distribution, fitting histograms of the classification models to match the test distribution, - didn't work well on test, still an enigma for me.\nf. genetic programming, - but we are not experts, tried only a simple configuration\ng. more complex stacking, - seems to overfit\n\n\nThanks for reading that far, and hopefully see you in the future competitions!",
    "544248": "Congrats on the result, even if you may feel disappointed.  I love the use of MIP given my history with CPLEX ;)  Very inventive way of using mathematical optimizaiton to attack a ML problem.\n\nAnd I am glad someone actually tried this:\n\n&gt; we merged short quakes 5 and 6 together\n\nIt was blatant in some of the papers, yet I am not sure many teams did it.  We did not, and this may be the single most effective change to make to improve our models significantly.  Well done!",
    "544258": "thank you! We did quite well, after all. It was so easy to drop into no-medals zone. \n\nI used CPLEX for 2 years while working at IBM research, and love it. My major is also in optimization, so I love to use these tools.\n\nCongratulations on your results, hopefully we will compete once again.",
    "544269": "Merging the quakes makes the mean higher and thus made train data closer to test indeed.",
    "544362": "This is really  interesting, including that MIP solution. \n\nTo merge earthquakes 5 and 6 was clever. I considered it but in the end decided to treat them as potentially harmful anomalies and remove them both from train. Anyway those two earthquakes were something to take care of so... clever move!\n\nThanks for sharing, and congratulations!",
    "545885": "Regarding the merge of earthquakes, I thought my eyes were seeing things in the papers. I kept going back and forth between the images in the papers and my own plots and was wondering if there was a mistake somewhere."
  },
  "source": "meta"
}