{
  "id": 94484,
  "title": "5th place solution",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/glory-or-death-5th-place-solution",
  "author_name": "",
  "post_date": "2019-06-05T00:44:55.977Z",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello kaggers!</p>\n\n<p>There have been many good solutions posted already, but I haven't yet seen some of the things I did, so I'd like to share some of my insights in this competition.</p>\n\n<ul>\n<li><p>First thing I realized is that no public kernel or my own features can distinguish between high and low quake intervals. Models usually settle for something that can be described as the mean quake interval multiplied by a factor between 0 and 1 with 1 being far from the next quake and 0 being right before the quake.</p></li>\n<li><p>Knowing this, I decided to eliminate the influence of the time between quakes on the target and instead predict that value between 0 and 1. I called it time fraction and it's <code>time_to_failure/time_between_quakes</code>. Then I just multiply the predicted values by the mean time between quakes (actually, a slightly lower value worked better, so I treated it as a hyper-parameter). This allowed my models to run longer before early stopping, converging to a better solution.</p></li>\n</ul>\n\n<p>*<em>Feature engineering *</em></p>\n\n<ul>\n<li><p>I used ~20 features in the end. Here is the picture of features' permutation importances:\n<img src=\"https://i.ibb.co/LYZGPF4/fig1.png\" alt=\"permutation importances\"></p></li>\n<li><p>Most of them involve some sort of signal filtering followed by some feature extraction. To filter the signal, I used Butterworth filters (<code>filtered_f1-f2_feature_name</code>), wavelet decompositions of various levels (<code>wavelet_name_decomposition_level_feature_name</code>), or removing large peaks. My most common features were histogram-based. Basically, it's the summed values at the peak or the tail of the histogram.  Some other features are peak-based or from public kernels.</p></li>\n<li><p>I did extensive feature selection, which improved my result significantly. I basically tried all possible wavelets in the pywt library with histogram-based features and did a stage-wise forward feature selection on them. To speed up feature selection, I used a simple SVR with a gaussian kernel instead of my final model.</p>\n\n<p><strong>Model</strong> \nI used a simple feed-forward neural net in pytorch, I found that it works much better with my features than LGBM. I added some bells and whistles to it, like adjusting learning rate on plateaus, batch normalization, adding 0 mean, 0.05 std gaussian noise to the training batches, training several models on each fold and picking the best one to avoid unlucky random weights initialization.</p>\n\n<p><strong>Cross-validation</strong> \nI loaded training data on a per-quake basis, so there was no overlap between quakes. I didn't test anything else, it just made sense to me to do it that way. I did 15-fold quake-based CV uniting the first and the last bits to the neighboring quake. So, my folds were [0,1],  2,  3,  4,  5,  6,  7,  8, 9 10, 11, 12, 13, 14, [15, 16]. Also, I extracted features from 150,000-point chunks with 25,000 shifts, which resulted in a significant overlap between adjacent chunks, but it worked fine in my NN model.</p>\n\n<p><strong>Scaling</strong> \nI identified the best scaling factor for the train set (by treating it as a hyper-parameter), then adjusted it based on the info about the test data (upscaled it a bit). I used 2 coefficients: ~1.05 and ~1.13. The first one gave me a better public LB and it correlated well with my CV, so I used it to check my model. The way I came up with 1.05 is pretty random and makes little sense, but this is the multiplier that brings the mean of my CV predictions to the target mean ttf in the train set. The second coefficient is based on what I expected the mean test set ttf to be.</p></li>\n</ul>\n\n<p>Overall, it was a fun but stressful ride. I think predicting time fraction instead of ttf could be useful to other winning models here and it could be a better target to predict in real life, as it indicates the relative state of the system (kind of like a danger level).</p>\n\n<p>Finally, here is my CV predictions plotted:\n<img src=\"https://i.ibb.co/HXmdGqH/fig2.png\" alt=\"CV predcitions\"></p>\n\n<p>UPDATE: Just adjusted my best submission's mean from ~5.7 to 6.32 and my score improved to 2.24127 :O Oh well, this is the lottery we all played.</p>",
  "messages": [
    {
      "id": "543783",
      "postDate": "06/04/2019 20:19:22",
      "content": "<p>Hello kaggers!</p>\n\n<p>There have been many good solutions posted already, but I haven't yet seen some of the things I did, so I'd like to share some of my insights in this competition.</p>\n\n<ul>\n<li><p>First thing I realized is that no public kernel or my own features can distinguish between high and low quake intervals. Models usually settle for something that can be described as the mean quake interval multiplied by a factor between 0 and 1 with 1 being far from the next quake and 0 being right before the quake.</p></li>\n<li><p>Knowing this, I decided to eliminate the influence of the time between quakes on the target and instead predict that value between 0 and 1. I called it time fraction and it's <code>time_to_failure/time_between_quakes</code>. Then I just multiply the predicted values by the mean time between quakes (actually, a slightly lower value worked better, so I treated it as a hyper-parameter). This allowed my models to run longer before early stopping, converging to a better solution.</p></li>\n</ul>\n\n<p>*<em>Feature engineering *</em></p>\n\n<ul>\n<li><p>I used ~20 features in the end. Here is the picture of features' permutation importances:\n<img src=\"https://i.ibb.co/LYZGPF4/fig1.png\" alt=\"permutation importances\"></p></li>\n<li><p>Most of them involve some sort of signal filtering followed by some feature extraction. To filter the signal, I used Butterworth filters (<code>filtered_f1-f2_feature_name</code>), wavelet decompositions of various levels (<code>wavelet_name_decomposition_level_feature_name</code>), or removing large peaks. My most common features were histogram-based. Basically, it's the summed values at the peak or the tail of the histogram.  Some other features are peak-based or from public kernels.</p></li>\n<li><p>I did extensive feature selection, which improved my result significantly. I basically tried all possible wavelets in the pywt library with histogram-based features and did a stage-wise forward feature selection on them. To speed up feature selection, I used a simple SVR with a gaussian kernel instead of my final model.</p>\n\n<p><strong>Model</strong> \nI used a simple feed-forward neural net in pytorch, I found that it works much better with my features than LGBM. I added some bells and whistles to it, like adjusting learning rate on plateaus, batch normalization, adding 0 mean, 0.05 std gaussian noise to the training batches, training several models on each fold and picking the best one to avoid unlucky random weights initialization.</p>\n\n<p><strong>Cross-validation</strong> \nI loaded training data on a per-quake basis, so there was no overlap between quakes. I didn't test anything else, it just made sense to me to do it that way. I did 15-fold quake-based CV uniting the first and the last bits to the neighboring quake. So, my folds were [0,1],  2,  3,  4,  5,  6,  7,  8, 9 10, 11, 12, 13, 14, [15, 16]. Also, I extracted features from 150,000-point chunks with 25,000 shifts, which resulted in a significant overlap between adjacent chunks, but it worked fine in my NN model.</p>\n\n<p><strong>Scaling</strong> \nI identified the best scaling factor for the train set (by treating it as a hyper-parameter), then adjusted it based on the info about the test data (upscaled it a bit). I used 2 coefficients: ~1.05 and ~1.13. The first one gave me a better public LB and it correlated well with my CV, so I used it to check my model. The way I came up with 1.05 is pretty random and makes little sense, but this is the multiplier that brings the mean of my CV predictions to the target mean ttf in the train set. The second coefficient is based on what I expected the mean test set ttf to be.</p></li>\n</ul>\n\n<p>Overall, it was a fun but stressful ride. I think predicting time fraction instead of ttf could be useful to other winning models here and it could be a better target to predict in real life, as it indicates the relative state of the system (kind of like a danger level).</p>\n\n<p>Finally, here is my CV predictions plotted:\n<img src=\"https://i.ibb.co/HXmdGqH/fig2.png\" alt=\"CV predcitions\"></p>\n\n<p>UPDATE: Just adjusted my best submission's mean from ~5.7 to 6.32 and my score improved to 2.24127 :O Oh well, this is the lottery we all played.</p>",
      "rawMarkdown": "Hello kaggers!\n\nThere have been many good solutions posted already, but I haven't yet seen some of the things I did, so I'd like to share some of my insights in this competition.\n\n - First thing I realized is that no public kernel or my own features can distinguish between high and low quake intervals. Models usually settle for something that can be described as the mean quake interval multiplied by a factor between 0 and 1 with 1 being far from the next quake and 0 being right before the quake.\n\n - Knowing this, I decided to eliminate the influence of the time between quakes on the target and instead predict that value between 0 and 1. I called it time fraction and it's `time_to_failure/time_between_quakes`. Then I just multiply the predicted values by the mean time between quakes (actually, a slightly lower value worked better, so I treated it as a hyper-parameter). This allowed my models to run longer before early stopping, converging to a better solution.\n\n**Feature engineering **\n\n- I used ~20 features in the end. Here is the picture of features' permutation importances:\n![permutation importances](https://i.ibb.co/LYZGPF4/fig1.png)\n\n- Most of them involve some sort of signal filtering followed by some feature extraction. To filter the signal, I used Butterworth filters (`filtered_f1-f2_feature_name`), wavelet decompositions of various levels (`wavelet_name_decomposition_level_feature_name`), or removing large peaks. My most common features were histogram-based. Basically, it's the summed values at the peak or the tail of the histogram.  Some other features are peak-based or from public kernels.\n\n- I did extensive feature selection, which improved my result significantly. I basically tried all possible wavelets in the pywt library with histogram-based features and did a stage-wise forward feature selection on them. To speed up feature selection, I used a simple SVR with a gaussian kernel instead of my final model.\n\n **Model** \nI used a simple feed-forward neural net in pytorch, I found that it works much better with my features than LGBM. I added some bells and whistles to it, like adjusting learning rate on plateaus, batch normalization, adding 0 mean, 0.05 std gaussian noise to the training batches, training several models on each fold and picking the best one to avoid unlucky random weights initialization.\n\n **Cross-validation** \nI loaded training data on a per-quake basis, so there was no overlap between quakes. I didn't test anything else, it just made sense to me to do it that way. I did 15-fold quake-based CV uniting the first and the last bits to the neighboring quake. So, my folds were [0,1],  2,  3,  4,  5,  6,  7,  8, 9 10, 11, 12, 13, 14, [15, 16]. Also, I extracted features from 150,000-point chunks with 25,000 shifts, which resulted in a significant overlap between adjacent chunks, but it worked fine in my NN model.\n\n **Scaling** \nI identified the best scaling factor for the train set (by treating it as a hyper-parameter), then adjusted it based on the info about the test data (upscaled it a bit). I used 2 coefficients: ~1.05 and ~1.13. The first one gave me a better public LB and it correlated well with my CV, so I used it to check my model. The way I came up with 1.05 is pretty random and makes little sense, but this is the multiplier that brings the mean of my CV predictions to the target mean ttf in the train set. The second coefficient is based on what I expected the mean test set ttf to be.\n\nOverall, it was a fun but stressful ride. I think predicting time fraction instead of ttf could be useful to other winning models here and it could be a better target to predict in real life, as it indicates the relative state of the system (kind of like a danger level).\n\nFinally, here is my CV predictions plotted:\n![CV predcitions](https://i.ibb.co/HXmdGqH/fig2.png)\n\nUPDATE: Just adjusted my best submission's mean from ~5.7 to 6.32 and my score improved to 2.24127 :O Oh well, this is the lottery we all played.",
      "votes": null
    },
    {
      "id": "543823",
      "postDate": "06/04/2019 21:24:14",
      "content": "<p>Congrats, that's a great solution, normalizing the ttf is similar of what we did. Thanks for sharing.</p>",
      "rawMarkdown": "Congrats, that's a great solution, normalizing the ttf is similar of what we did. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "543950",
      "postDate": "06/05/2019 00:47:46",
      "content": "<p>Thanks! I just read the description of your solution and I'm surprised normalizing ttf didn't help in the end. But I also found through feature selection that the features that predict ttf well are not necessarily gonna be the best features to predict normalized ttf. I had to redo feature selection when I switched to normalized ttf.</p>",
      "rawMarkdown": "Thanks! I just read the description of your solution and I'm surprised normalizing ttf didn't help in the end. But I also found through feature selection that the features that predict ttf well are not necessarily gonna be the best features to predict normalized ttf. I had to redo feature selection when I switched to normalized ttf.",
      "votes": null
    },
    {
      "id": "543979",
      "postDate": "06/05/2019 02:14:47",
      "content": "<p>Congratulations Ivan. Thanks for posting your solution. </p>",
      "rawMarkdown": "Congratulations Ivan. Thanks for posting your solution.",
      "votes": null
    },
    {
      "id": "544389",
      "postDate": "06/05/2019 13:05:04",
      "content": "<p>Normalizing ttf is an elegant trick. Congratulations for your great result, and thank you for sharing this!</p>",
      "rawMarkdown": "Normalizing ttf is an elegant trick. Congratulations for your great result, and thank you for sharing this!",
      "votes": null
    },
    {
      "id": "591795",
      "postDate": "08/04/2019 09:08:32",
      "content": "<p>Excellent</p>",
      "rawMarkdown": "Excellent",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 543823,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/04/2019 21:24:14",
      "content": "<p>Congrats, that's a great solution, normalizing the ttf is similar of what we did. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 543950,
          "author_name": "batalov",
          "author_url": "",
          "post_date": "06/05/2019 00:47:46",
          "content": "<p>Thanks! I just read the description of your solution and I'm surprised normalizing ttf didn't help in the end. But I also found through feature selection that the features that predict ttf well are not necessarily gonna be the best features to predict normalized ttf. I had to redo feature selection when I switched to normalized ttf.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 543979,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "06/05/2019 02:14:47",
      "content": "<p>Congratulations Ivan. Thanks for posting your solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 544389,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "06/05/2019 13:05:04",
      "content": "<p>Normalizing ttf is an elegant trick. Congratulations for your great result, and thank you for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 591795,
      "author_name": "jashandhillon",
      "author_url": "",
      "post_date": "08/04/2019 09:08:32",
      "content": "<p>Excellent</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543783": "Hello kaggers!\n\nThere have been many good solutions posted already, but I haven't yet seen some of the things I did, so I'd like to share some of my insights in this competition.\n\n - First thing I realized is that no public kernel or my own features can distinguish between high and low quake intervals. Models usually settle for something that can be described as the mean quake interval multiplied by a factor between 0 and 1 with 1 being far from the next quake and 0 being right before the quake.\n\n - Knowing this, I decided to eliminate the influence of the time between quakes on the target and instead predict that value between 0 and 1. I called it time fraction and it's `time_to_failure/time_between_quakes`. Then I just multiply the predicted values by the mean time between quakes (actually, a slightly lower value worked better, so I treated it as a hyper-parameter). This allowed my models to run longer before early stopping, converging to a better solution.\n\n**Feature engineering **\n\n- I used ~20 features in the end. Here is the picture of features' permutation importances:\n![permutation importances](https://i.ibb.co/LYZGPF4/fig1.png)\n\n- Most of them involve some sort of signal filtering followed by some feature extraction. To filter the signal, I used Butterworth filters (`filtered_f1-f2_feature_name`), wavelet decompositions of various levels (`wavelet_name_decomposition_level_feature_name`), or removing large peaks. My most common features were histogram-based. Basically, it's the summed values at the peak or the tail of the histogram.  Some other features are peak-based or from public kernels.\n\n- I did extensive feature selection, which improved my result significantly. I basically tried all possible wavelets in the pywt library with histogram-based features and did a stage-wise forward feature selection on them. To speed up feature selection, I used a simple SVR with a gaussian kernel instead of my final model.\n\n **Model** \nI used a simple feed-forward neural net in pytorch, I found that it works much better with my features than LGBM. I added some bells and whistles to it, like adjusting learning rate on plateaus, batch normalization, adding 0 mean, 0.05 std gaussian noise to the training batches, training several models on each fold and picking the best one to avoid unlucky random weights initialization.\n\n **Cross-validation** \nI loaded training data on a per-quake basis, so there was no overlap between quakes. I didn't test anything else, it just made sense to me to do it that way. I did 15-fold quake-based CV uniting the first and the last bits to the neighboring quake. So, my folds were [0,1],  2,  3,  4,  5,  6,  7,  8, 9 10, 11, 12, 13, 14, [15, 16]. Also, I extracted features from 150,000-point chunks with 25,000 shifts, which resulted in a significant overlap between adjacent chunks, but it worked fine in my NN model.\n\n **Scaling** \nI identified the best scaling factor for the train set (by treating it as a hyper-parameter), then adjusted it based on the info about the test data (upscaled it a bit). I used 2 coefficients: ~1.05 and ~1.13. The first one gave me a better public LB and it correlated well with my CV, so I used it to check my model. The way I came up with 1.05 is pretty random and makes little sense, but this is the multiplier that brings the mean of my CV predictions to the target mean ttf in the train set. The second coefficient is based on what I expected the mean test set ttf to be.\n\nOverall, it was a fun but stressful ride. I think predicting time fraction instead of ttf could be useful to other winning models here and it could be a better target to predict in real life, as it indicates the relative state of the system (kind of like a danger level).\n\nFinally, here is my CV predictions plotted:\n![CV predcitions](https://i.ibb.co/HXmdGqH/fig2.png)\n\nUPDATE: Just adjusted my best submission's mean from ~5.7 to 6.32 and my score improved to 2.24127 :O Oh well, this is the lottery we all played.",
    "543823": "Congrats, that's a great solution, normalizing the ttf is similar of what we did. Thanks for sharing.",
    "543950": "Thanks! I just read the description of your solution and I'm surprised normalizing ttf didn't help in the end. But I also found through feature selection that the features that predict ttf well are not necessarily gonna be the best features to predict normalized ttf. I had to redo feature selection when I switched to normalized ttf.",
    "543979": "Congratulations Ivan. Thanks for posting your solution.",
    "544389": "Normalizing ttf is an elegant trick. Congratulations for your great result, and thank you for sharing this!",
    "591795": "Excellent"
  },
  "source": "meta"
}