{
  "id": 94556,
  "title": "15th Place Memo and CV Scheme",
  "url": "/competitions/LANL-Earthquake-Prediction/writeups/berni-15th-place-memo-and-cv-scheme",
  "author_name": "",
  "post_date": "2019-06-05T07:44:28.412182700Z",
  "votes": 18,
  "comment_count": 9,
  "views": 0,
  "content": "<h2>Thank You!</h2>\n\n<p>It was my first Kaggle challenge. (... I pursued to the end. I got into <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">Quora Insincere Questions Classification</a> but got distracted with other stuff in live.) I'm pretty happy to have survived the shake up and have finished so high up in the ranking. I'd like to thank the community to share insights and ideas so well (especially to @CPMPml and his <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125\">no magic</a> feature selection)! That's the spirit of Kaggle! It was a great learning experience, and I will try to give something back here.</p>\n\n<h2>What Matters</h2>\n\n<p>I made that huge jump in the LB and finished well. I'm still trying to figure out, to what degree it was luck and how much it was skill and a robust model. The shakeup might suggest a large influence of luck. Notching the model slightly, like adjusting the mean, can have a big impact on your standing in the final LB as @sushize showed <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-543708\">here</a>. So the final scoring was indeed quite fragile and one had to be lucky to win a gold medal. I don't think, it's pure luck despite the big shakeup, though. We still see the grandmasters in high spots and the winning solutions used models with few features and robust CV strategies.</p>\n\n<p>I think this challenge taught us that applying fancy ML techniques is not enough and it sometimes matters more to understand the basics. We have to:\n* get a good grasp of the problem and understand the data well\n* build a good and robust CV strategy\n* avoid overfitting by all means</p>\n\n<p>As the data was so little, these key ML ingredients played a major role in this competition and models ware relatively less important. Trust your thinking, think well, and do not apply stuff, you don't fully understand (though try to!).</p>\n\n<h2>The Heart: Cross Validation and Significance</h2>\n\n<p>After some initial struggle with the large data and spending quite some time trying to get NN's to work, time was getting short for me. That helped me to concentrate on the important. I committed myself to the models that worked best, which were GBM's (I used a blend of LightGBM, XGBoost, and CatBoost in the end). Most importantly, I realised that a robust CV strategy <em>you stick to</em> is crucial.</p>\n\n<p>We've seen that the quality of practically all models depended on the actual TTF. They do well for intermediate values and bad for very low or high values of TTF. So, depending on the distribution of TTF in the test / validation portion affects the score a lot. Hence, I tried to split the training data, such that this effect was minimised. At the same time, I got the feeling that I should take entire cycles in or out. Hence, I searched for splits of the 17 cycles into groups of three to four cycles that had similar TTF <em>and</em> size. One of those splits was\n| cycles | fraction | mean TTF |\n| --- | --- | --- |\n| 0, 3, 7, 8 | 21.0% | 5.86 |\n| 1, 10, 11 | 20.8% | 5.67 |\n| 2, 12, 15 | 19.9% | 5.68 |\n| 4, 5, 9, 16 | 19.7% | 5.56 |\n| 6, 13, 14 | 18.6% | 5.61 |</p>\n\n<p>Groups 0 and 16 are the very short \"cycles\" before and after the first and last earth quake. Note that with this split, I also don't have leakage between overlapping segments.</p>\n\n<p>These folds still didn't give me the stable CV scores I wanted. I created some more of these even splits and decided to actually do four of these 5-fold group splits and take the mean MAE of each of the four 5-fold splits, such that I had four scores. This resulted in a rather slow CV scheme, but it was stable and I could not only work with a CV score, but also an uncertainty: <code>np.std(cv_mae,ddof=1)</code>, where <code>cv_mae</code> are the four averaged scores of the 5-folg group splits. These uncertainties were typically of order <code>0.005</code>. Note that this scheme assumes that the (private) test set has the same TTF distribution as the training set.</p>\n\n<p>After establishing that, I was in a situation where I could not only test models and features with a robust CV scheme, but also had a measure for what a significant improve in the score is.</p>\n\n<h2>What We Predict On</h2>\n\n<p>The assumption that the distributions of features and TTF in the test set, I've made in setting up my CV scheme, is quite reasonable. Some said otherwise, but I contradict. It was claimed by @gpreda that <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\">the feature distribution was different for train and test set</a>. I could not find that for my features as I've subtracted the mean of the acoustic data on each segment. That makes physically sense and I already did it before I knew about the comparison.</p>\n\n<p>Also the assumption that the TTF distributions are similar should be reasonable. @mykper did some <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583\">nice work</a> on the public test set. We've seem a somewhat different distribution of the TTF and they also maxed out at some value below 10, whereas the training data had values up to 16. That's true. However, he also found out, that there are only two cycles in the public test set. These were quite typical, if you compare them with the training set. No reason to assume that we will predict on structurally different data than what we have seen in the training data.</p>\n\n<p>Also one <strong>note on the public LB</strong>:</p>\n\n<p>The public LB was calculated on just about 350 segments. This small number together with the high variance in the predictions alone should tell you that you cannot trust the public LB. Now compare that number with the size of the CV: we use several fold and will eventually use the entire training set for the score. That is 4200 segments. That's more than 10 times as much! When using overlapping segments - as I also did - the ratio gets even bigger, although the information gain will not grow linearly with the number of segments anymore. In anyway: the public LB score is calculated on a tiny set, especially when compared with the CV on the training data. I've never trusted the public LB in this competition.</p>\n\n<h2>The Features and Minimalism</h2>\n\n<p>I took @CPMPml approach of feature selection he explained in the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125\">no magic</a> discussion and augmented it with my uncertainty of my CV scores. As we have little data, and I would have ended up with almost no features, if I'd only taken features with high significance. I decided to take features that improved my model by at least half a standard deviation.</p>\n\n<p>Some side note from a physicist:\nIf you have two measurements <code>m_1</code> and <code>m_2</code> with standard deviations <code>s_1</code> and <code>s_2</code>, the difference of the measurements <code>m_1 - m_2</code> does <em>not</em> have an uncertainty / standard deviation of <code>s_1 + s_2</code>, but of <code>sqrt(s_1^2 + s_2^2)</code>.</p>\n\n<p>With that approach, I ended up with just 7 features!</p>\n\n<h2>Final Words</h2>\n\n<p>That's pretty much it. Some ensembeling of different GBM's (LightGBM, XGBoost, CatBoost) with (a little) optimised hyper-parameters and, as I've said in the beginning, probably a bit of luck.</p>\n\n<p>There are some things I've seen in other solutions, that I'd like to have used, too. There is the special treatment of outliers as them ABC did (see their approach <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731\">here</a>). Also nested CV would probably have done some good.</p>\n\n<p>Much of the detective work on the test set, was not so fruitful, I think. Despite we even <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664\">know where it came from</a>. The only think that was useful - and probably quite a bit, although I didn't use it - was the information about its mean TTF. It seems to have helped a few teams. So after all, my assumption that the test set is not as strictly fulfilled as I've assumed in my CV scheme.</p>",
  "messages": [
    {
      "id": "544148",
      "postDate": "06/05/2019 07:44:28",
      "content": "<h2>Thank You!</h2>\n\n<p>It was my first Kaggle challenge. (... I pursued to the end. I got into <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\">Quora Insincere Questions Classification</a> but got distracted with other stuff in live.) I'm pretty happy to have survived the shake up and have finished so high up in the ranking. I'd like to thank the community to share insights and ideas so well (especially to @CPMPml and his <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125\">no magic</a> feature selection)! That's the spirit of Kaggle! It was a great learning experience, and I will try to give something back here.</p>\n\n<h2>What Matters</h2>\n\n<p>I made that huge jump in the LB and finished well. I'm still trying to figure out, to what degree it was luck and how much it was skill and a robust model. The shakeup might suggest a large influence of luck. Notching the model slightly, like adjusting the mean, can have a big impact on your standing in the final LB as @sushize showed <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-543708\">here</a>. So the final scoring was indeed quite fragile and one had to be lucky to win a gold medal. I don't think, it's pure luck despite the big shakeup, though. We still see the grandmasters in high spots and the winning solutions used models with few features and robust CV strategies.</p>\n\n<p>I think this challenge taught us that applying fancy ML techniques is not enough and it sometimes matters more to understand the basics. We have to:\n* get a good grasp of the problem and understand the data well\n* build a good and robust CV strategy\n* avoid overfitting by all means</p>\n\n<p>As the data was so little, these key ML ingredients played a major role in this competition and models ware relatively less important. Trust your thinking, think well, and do not apply stuff, you don't fully understand (though try to!).</p>\n\n<h2>The Heart: Cross Validation and Significance</h2>\n\n<p>After some initial struggle with the large data and spending quite some time trying to get NN's to work, time was getting short for me. That helped me to concentrate on the important. I committed myself to the models that worked best, which were GBM's (I used a blend of LightGBM, XGBoost, and CatBoost in the end). Most importantly, I realised that a robust CV strategy <em>you stick to</em> is crucial.</p>\n\n<p>We've seen that the quality of practically all models depended on the actual TTF. They do well for intermediate values and bad for very low or high values of TTF. So, depending on the distribution of TTF in the test / validation portion affects the score a lot. Hence, I tried to split the training data, such that this effect was minimised. At the same time, I got the feeling that I should take entire cycles in or out. Hence, I searched for splits of the 17 cycles into groups of three to four cycles that had similar TTF <em>and</em> size. One of those splits was\n| cycles | fraction | mean TTF |\n| --- | --- | --- |\n| 0, 3, 7, 8 | 21.0% | 5.86 |\n| 1, 10, 11 | 20.8% | 5.67 |\n| 2, 12, 15 | 19.9% | 5.68 |\n| 4, 5, 9, 16 | 19.7% | 5.56 |\n| 6, 13, 14 | 18.6% | 5.61 |</p>\n\n<p>Groups 0 and 16 are the very short \"cycles\" before and after the first and last earth quake. Note that with this split, I also don't have leakage between overlapping segments.</p>\n\n<p>These folds still didn't give me the stable CV scores I wanted. I created some more of these even splits and decided to actually do four of these 5-fold group splits and take the mean MAE of each of the four 5-fold splits, such that I had four scores. This resulted in a rather slow CV scheme, but it was stable and I could not only work with a CV score, but also an uncertainty: <code>np.std(cv_mae,ddof=1)</code>, where <code>cv_mae</code> are the four averaged scores of the 5-folg group splits. These uncertainties were typically of order <code>0.005</code>. Note that this scheme assumes that the (private) test set has the same TTF distribution as the training set.</p>\n\n<p>After establishing that, I was in a situation where I could not only test models and features with a robust CV scheme, but also had a measure for what a significant improve in the score is.</p>\n\n<h2>What We Predict On</h2>\n\n<p>The assumption that the distributions of features and TTF in the test set, I've made in setting up my CV scheme, is quite reasonable. Some said otherwise, but I contradict. It was claimed by @gpreda that <a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\">the feature distribution was different for train and test set</a>. I could not find that for my features as I've subtracted the mean of the acoustic data on each segment. That makes physically sense and I already did it before I knew about the comparison.</p>\n\n<p>Also the assumption that the TTF distributions are similar should be reasonable. @mykper did some <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583\">nice work</a> on the public test set. We've seem a somewhat different distribution of the TTF and they also maxed out at some value below 10, whereas the training data had values up to 16. That's true. However, he also found out, that there are only two cycles in the public test set. These were quite typical, if you compare them with the training set. No reason to assume that we will predict on structurally different data than what we have seen in the training data.</p>\n\n<p>Also one <strong>note on the public LB</strong>:</p>\n\n<p>The public LB was calculated on just about 350 segments. This small number together with the high variance in the predictions alone should tell you that you cannot trust the public LB. Now compare that number with the size of the CV: we use several fold and will eventually use the entire training set for the score. That is 4200 segments. That's more than 10 times as much! When using overlapping segments - as I also did - the ratio gets even bigger, although the information gain will not grow linearly with the number of segments anymore. In anyway: the public LB score is calculated on a tiny set, especially when compared with the CV on the training data. I've never trusted the public LB in this competition.</p>\n\n<h2>The Features and Minimalism</h2>\n\n<p>I took @CPMPml approach of feature selection he explained in the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125\">no magic</a> discussion and augmented it with my uncertainty of my CV scores. As we have little data, and I would have ended up with almost no features, if I'd only taken features with high significance. I decided to take features that improved my model by at least half a standard deviation.</p>\n\n<p>Some side note from a physicist:\nIf you have two measurements <code>m_1</code> and <code>m_2</code> with standard deviations <code>s_1</code> and <code>s_2</code>, the difference of the measurements <code>m_1 - m_2</code> does <em>not</em> have an uncertainty / standard deviation of <code>s_1 + s_2</code>, but of <code>sqrt(s_1^2 + s_2^2)</code>.</p>\n\n<p>With that approach, I ended up with just 7 features!</p>\n\n<h2>Final Words</h2>\n\n<p>That's pretty much it. Some ensembeling of different GBM's (LightGBM, XGBoost, CatBoost) with (a little) optimised hyper-parameters and, as I've said in the beginning, probably a bit of luck.</p>\n\n<p>There are some things I've seen in other solutions, that I'd like to have used, too. There is the special treatment of outliers as them ABC did (see their approach <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731\">here</a>). Also nested CV would probably have done some good.</p>\n\n<p>Much of the detective work on the test set, was not so fruitful, I think. Despite we even <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664\">know where it came from</a>. The only think that was useful - and probably quite a bit, although I didn't use it - was the information about its mean TTF. It seems to have helped a few teams. So after all, my assumption that the test set is not as strictly fulfilled as I've assumed in my CV scheme.</p>",
      "rawMarkdown": "## Thank You!\n\nIt was my first Kaggle challenge. (... I pursued to the end. I got into [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification) but got distracted with other stuff in live.) I'm pretty happy to have survived the shake up and have finished so high up in the ranking. I'd like to thank the community to share insights and ideas so well (especially to @CPMPml and his [no magic](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125) feature selection)! That's the spirit of Kaggle! It was a great learning experience, and I will try to give something back here.\n\n## What Matters\n\nI made that huge jump in the LB and finished well. I'm still trying to figure out, to what degree it was luck and how much it was skill and a robust model. The shakeup might suggest a large influence of luck. Notching the model slightly, like adjusting the mean, can have a big impact on your standing in the final LB as @sushize showed [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-543708). So the final scoring was indeed quite fragile and one had to be lucky to win a gold medal. I don't think, it's pure luck despite the big shakeup, though. We still see the grandmasters in high spots and the winning solutions used models with few features and robust CV strategies.\n\nI think this challenge taught us that applying fancy ML techniques is not enough and it sometimes matters more to understand the basics. We have to:\n* get a good grasp of the problem and understand the data well\n* build a good and robust CV strategy\n* avoid overfitting by all means\n\nAs the data was so little, these key ML ingredients played a major role in this competition and models ware relatively less important. Trust your thinking, think well, and do not apply stuff, you don't fully understand (though try to!).\n\n## The Heart: Cross Validation and Significance\n\nAfter some initial struggle with the large data and spending quite some time trying to get NN's to work, time was getting short for me. That helped me to concentrate on the important. I committed myself to the models that worked best, which were GBM's (I used a blend of LightGBM, XGBoost, and CatBoost in the end). Most importantly, I realised that a robust CV strategy *you stick to* is crucial.\n\nWe've seen that the quality of practically all models depended on the actual TTF. They do well for intermediate values and bad for very low or high values of TTF. So, depending on the distribution of TTF in the test / validation portion affects the score a lot. Hence, I tried to split the training data, such that this effect was minimised. At the same time, I got the feeling that I should take entire cycles in or out. Hence, I searched for splits of the 17 cycles into groups of three to four cycles that had similar TTF *and* size. One of those splits was\n| cycles | fraction | mean TTF |\n| --- | --- | --- |\n| 0, 3, 7, 8 | 21.0% | 5.86 |\n| 1, 10, 11 | 20.8% | 5.67 |\n| 2, 12, 15 | 19.9% | 5.68 |\n| 4, 5, 9, 16 | 19.7% | 5.56 |\n| 6, 13, 14 | 18.6% | 5.61 |\n\nGroups 0 and 16 are the very short \"cycles\" before and after the first and last earth quake. Note that with this split, I also don't have leakage between overlapping segments.\n\nThese folds still didn't give me the stable CV scores I wanted. I created some more of these even splits and decided to actually do four of these 5-fold group splits and take the mean MAE of each of the four 5-fold splits, such that I had four scores. This resulted in a rather slow CV scheme, but it was stable and I could not only work with a CV score, but also an uncertainty: `np.std(cv_mae,ddof=1)`, where `cv_mae` are the four averaged scores of the 5-folg group splits. These uncertainties were typically of order `0.005`. Note that this scheme assumes that the (private) test set has the same TTF distribution as the training set.\n\nAfter establishing that, I was in a situation where I could not only test models and features with a robust CV scheme, but also had a measure for what a significant improve in the score is.\n\n## What We Predict On\n\nThe assumption that the distributions of features and TTF in the test set, I've made in setting up my CV scheme, is quite reasonable. Some said otherwise, but I contradict. It was claimed by @gpreda that [the feature distribution was different for train and test set](https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda). I could not find that for my features as I've subtracted the mean of the acoustic data on each segment. That makes physically sense and I already did it before I knew about the comparison.\n\nAlso the assumption that the TTF distributions are similar should be reasonable. @mykper did some [nice work](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583) on the public test set. We've seem a somewhat different distribution of the TTF and they also maxed out at some value below 10, whereas the training data had values up to 16. That's true. However, he also found out, that there are only two cycles in the public test set. These were quite typical, if you compare them with the training set. No reason to assume that we will predict on structurally different data than what we have seen in the training data.\n\nAlso one **note on the public LB**:\n\nThe public LB was calculated on just about 350 segments. This small number together with the high variance in the predictions alone should tell you that you cannot trust the public LB. Now compare that number with the size of the CV: we use several fold and will eventually use the entire training set for the score. That is 4200 segments. That's more than 10 times as much! When using overlapping segments - as I also did - the ratio gets even bigger, although the information gain will not grow linearly with the number of segments anymore. In anyway: the public LB score is calculated on a tiny set, especially when compared with the CV on the training data. I've never trusted the public LB in this competition.\n\n## The Features and Minimalism\n\nI took @CPMPml approach of feature selection he explained in the [no magic](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125) discussion and augmented it with my uncertainty of my CV scores. As we have little data, and I would have ended up with almost no features, if I'd only taken features with high significance. I decided to take features that improved my model by at least half a standard deviation.\n\nSome side note from a physicist:\nIf you have two measurements `m_1` and `m_2` with standard deviations `s_1` and `s_2`, the difference of the measurements `m_1 - m_2` does *not* have an uncertainty / standard deviation of `s_1 + s_2`, but of `sqrt(s_1^2 + s_2^2)`.\n\nWith that approach, I ended up with just 7 features!\n\n## Final Words\n\nThat's pretty much it. Some ensembeling of different GBM's (LightGBM, XGBoost, CatBoost) with (a little) optimised hyper-parameters and, as I've said in the beginning, probably a bit of luck.\n\nThere are some things I've seen in other solutions, that I'd like to have used, too. There is the special treatment of outliers as them ABC did (see their approach [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731)). Also nested CV would probably have done some good.\n\nMuch of the detective work on the test set, was not so fruitful, I think. Despite we even [know where it came from](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664). The only think that was useful - and probably quite a bit, although I didn't use it - was the information about its mean TTF. It seems to have helped a few teams. So after all, my assumption that the test set is not as strictly fulfilled as I've assumed in my CV scheme.",
      "votes": null
    },
    {
      "id": "544198",
      "postDate": "06/05/2019 08:57:55",
      "content": "<p>Thanks for sharing, happy you find my post useful.  Congrats on the result!</p>\n\n<blockquote>\n  <p>Some side note from a physicist:</p>\n</blockquote>\n\n<p>It should be known by every data scientist that variance of a sum is the sum of variance.</p>",
      "rawMarkdown": "Thanks for sharing, happy you find my post useful.  Congrats on the result!\n\n&gt; Some side note from a physicist:\n\nIt should be known by every data scientist that variance of a sum is the sum of variance.",
      "votes": null
    },
    {
      "id": "544228",
      "postDate": "06/05/2019 09:37:15",
      "content": "<p>It should. I've encountered quite a few that don't, though... (Or they forget, get distracted by fancy stuff, ...). It doesn't hurt to refresh the basics from time to time.</p>",
      "rawMarkdown": "It should. I've encountered quite a few that don't, though... (Or they forget, get distracted by fancy stuff, ...). It doesn't hurt to refresh the basics from time to time.",
      "votes": null
    },
    {
      "id": "544255",
      "postDate": "06/05/2019 10:12:48",
      "content": "<p>So you didn't use the leak, right?\nAnd as far as I understand you didn't try to make the train set similar to the test set? You made the folds of the train set homogenous (by several ways).</p>",
      "rawMarkdown": "So you didn't use the leak, right?\nAnd as far as I understand you didn't try to make the train set similar to the test set? You made the folds of the train set homogenous (by several ways).",
      "votes": null
    },
    {
      "id": "544267",
      "postDate": "06/05/2019 10:36:25",
      "content": "<p>Correct. I didn’t make the training set similar to the test set as there was no (strong enough) evidence we would actually habe a different test set. For me the risk was to high to overadjust or even adjust the wrong direction.</p>",
      "rawMarkdown": "Correct. I didn’t make the training set similar to the test set as there was no (strong enough) evidence we would actually habe a different test set. For me the risk was to high to overadjust or even adjust the wrong direction.",
      "votes": null
    },
    {
      "id": "544270",
      "postDate": "06/05/2019 10:41:13",
      "content": "<p>Congrats on the gold medal!</p>\n\n<p>Can you please disclose what are the 7 selected features, do they include the mean?</p>",
      "rawMarkdown": "Congrats on the gold medal!\n\nCan you please disclose what are the 7 selected features, do they include the mean?",
      "votes": null
    },
    {
      "id": "544271",
      "postDate": "06/05/2019 10:45:15",
      "content": "<p><a href=\"/bernir\">@bernir</a> \nVery cool! It seems that you're the only one in the gold zone who didn't use the leak or the test set.</p>",
      "rawMarkdown": "bernir \nVery cool! It seems that you're the only one in the gold zone who didn't use the leak or the test set.",
      "votes": null
    },
    {
      "id": "544294",
      "postDate": "06/05/2019 11:10:10",
      "content": "<p>I don't have the 7 features at hand at the moment, but no, I did not use the mean. I substracted the mean from each and every segment, so I would be zero all the time anyways. (Note, that I did not normalise the standard deviation, as it contains information!)</p>\n\n<p>The 7 features are no \"magic\" features and I'm pretty sure that one could find 5-10 other features that would perform equally good. Using only the features that significantly improve your model prevents overfitting. That's all whats going on here. I had engineered 900 features to choose from. Using all of them on just 4200 training examples (or a bit more with overlapping segments) would give you less than 5 data points per feature... That's badly constrained (although some features correlate strongly) and genearlises poorly.</p>",
      "rawMarkdown": "I don't have the 7 features at hand at the moment, but no, I did not use the mean. I substracted the mean from each and every segment, so I would be zero all the time anyways. (Note, that I did not normalise the standard deviation, as it contains information!)\n\nThe 7 features are no \"magic\" features and I'm pretty sure that one could find 5-10 other features that would perform equally good. Using only the features that significantly improve your model prevents overfitting. That's all whats going on here. I had engineered 900 features to choose from. Using all of them on just 4200 training examples (or a bit more with overlapping segments) would give you less than 5 data points per feature... That's badly constrained (although some features correlate strongly) and genearlises poorly.",
      "votes": null
    },
    {
      "id": "546001",
      "postDate": "06/06/2019 05:55:47",
      "content": "<p>My seven features are named like that, if it helps:\n<code>python\nselected_features = [\n    'roll100mean_num_peaks_5',\n    'spectral_centroid_perc90',\n    'FFT_1650_1800',\n    'noise_perc75',\n    'raw_num_peaks_8',\n    'spectral_bandwidth_perc90',\n    'perio_iqr25',\n]\n</code>\nSo not necessarily very straight forward feaatures, but nothing too fancy or even magic either.</p>",
      "rawMarkdown": "My seven features are named like that, if it helps:\n```python\nselected_features = [\n    'roll100mean_num_peaks_5',\n    'spectral_centroid_perc90',\n    'FFT_1650_1800',\n    'noise_perc75',\n    'raw_num_peaks_8',\n    'spectral_bandwidth_perc90',\n    'perio_iqr25',\n]\n```\nSo not necessarily very straight forward feaatures, but nothing too fancy or even magic either.",
      "votes": null
    },
    {
      "id": "553091",
      "postDate": "06/15/2019 05:24:35",
      "content": "<p>Thank you for sharing your selected features.</p>\n\n<p><a href=\"https://www.jstor.org/stable/40285694\">Roles for Spectral Centroid and Other Factors in Determining \"Blended\" Instrument Pairings in Orchestration</a> states: \"Blend worsened as a function of the overall centroid height of the combination or as the amount of difference between the centroids of the two instruments increased.\"</p>\n\n<p>MusicQuake and EarthQuake seem to be related!</p>",
      "rawMarkdown": "Thank you for sharing your selected features.\n\n[Roles for Spectral Centroid and Other Factors in Determining \"Blended\" Instrument Pairings in Orchestration](https://www.jstor.org/stable/40285694) states: \"Blend worsened as a function of the overall centroid height of the combination or as the amount of difference between the centroids of the two instruments increased.\"\n\nMusicQuake and EarthQuake seem to be related!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 544198,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2019 08:57:55",
      "content": "<p>Thanks for sharing, happy you find my post useful.  Congrats on the result!</p>\n\n<blockquote>\n  <p>Some side note from a physicist:</p>\n</blockquote>\n\n<p>It should be known by every data scientist that variance of a sum is the sum of variance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 544228,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "06/05/2019 09:37:15",
          "content": "<p>It should. I've encountered quite a few that don't, though... (Or they forget, get distracted by fancy stuff, ...). It doesn't hurt to refresh the basics from time to time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544255,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "06/05/2019 10:12:48",
      "content": "<p>So you didn't use the leak, right?\nAnd as far as I understand you didn't try to make the train set similar to the test set? You made the folds of the train set homogenous (by several ways).</p>",
      "votes": null,
      "replies": [
        {
          "id": 544267,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "06/05/2019 10:36:25",
          "content": "<p>Correct. I didn’t make the training set similar to the test set as there was no (strong enough) evidence we would actually habe a different test set. For me the risk was to high to overadjust or even adjust the wrong direction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544271,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "06/05/2019 10:45:15",
          "content": "<p><a href=\"/bernir\">@bernir</a> \nVery cool! It seems that you're the only one in the gold zone who didn't use the leak or the test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544270,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "06/05/2019 10:41:13",
      "content": "<p>Congrats on the gold medal!</p>\n\n<p>Can you please disclose what are the 7 selected features, do they include the mean?</p>",
      "votes": null,
      "replies": [
        {
          "id": 544294,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "06/05/2019 11:10:10",
          "content": "<p>I don't have the 7 features at hand at the moment, but no, I did not use the mean. I substracted the mean from each and every segment, so I would be zero all the time anyways. (Note, that I did not normalise the standard deviation, as it contains information!)</p>\n\n<p>The 7 features are no \"magic\" features and I'm pretty sure that one could find 5-10 other features that would perform equally good. Using only the features that significantly improve your model prevents overfitting. That's all whats going on here. I had engineered 900 features to choose from. Using all of them on just 4200 training examples (or a bit more with overlapping segments) would give you less than 5 data points per feature... That's badly constrained (although some features correlate strongly) and genearlises poorly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 546001,
          "author_name": "bernir",
          "author_url": "",
          "post_date": "06/06/2019 05:55:47",
          "content": "<p>My seven features are named like that, if it helps:\n<code>python\nselected_features = [\n    'roll100mean_num_peaks_5',\n    'spectral_centroid_perc90',\n    'FFT_1650_1800',\n    'noise_perc75',\n    'raw_num_peaks_8',\n    'spectral_bandwidth_perc90',\n    'perio_iqr25',\n]\n</code>\nSo not necessarily very straight forward feaatures, but nothing too fancy or even magic either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 553091,
          "author_name": "lalitapatel",
          "author_url": "",
          "post_date": "06/15/2019 05:24:35",
          "content": "<p>Thank you for sharing your selected features.</p>\n\n<p><a href=\"https://www.jstor.org/stable/40285694\">Roles for Spectral Centroid and Other Factors in Determining \"Blended\" Instrument Pairings in Orchestration</a> states: \"Blend worsened as a function of the overall centroid height of the combination or as the amount of difference between the centroids of the two instruments increased.\"</p>\n\n<p>MusicQuake and EarthQuake seem to be related!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "544148": "## Thank You!\n\nIt was my first Kaggle challenge. (... I pursued to the end. I got into [Quora Insincere Questions Classification](https://www.kaggle.com/c/quora-insincere-questions-classification) but got distracted with other stuff in live.) I'm pretty happy to have survived the shake up and have finished so high up in the ranking. I'd like to thank the community to share insights and ideas so well (especially to @CPMPml and his [no magic](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125) feature selection)! That's the spirit of Kaggle! It was a great learning experience, and I will try to give something back here.\n\n## What Matters\n\nI made that huge jump in the LB and finished well. I'm still trying to figure out, to what degree it was luck and how much it was skill and a robust model. The shakeup might suggest a large influence of luck. Notching the model slightly, like adjusting the mean, can have a big impact on your standing in the final LB as @sushize showed [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-543708). So the final scoring was indeed quite fragile and one had to be lucky to win a gold medal. I don't think, it's pure luck despite the big shakeup, though. We still see the grandmasters in high spots and the winning solutions used models with few features and robust CV strategies.\n\nI think this challenge taught us that applying fancy ML techniques is not enough and it sometimes matters more to understand the basics. We have to:\n* get a good grasp of the problem and understand the data well\n* build a good and robust CV strategy\n* avoid overfitting by all means\n\nAs the data was so little, these key ML ingredients played a major role in this competition and models ware relatively less important. Trust your thinking, think well, and do not apply stuff, you don't fully understand (though try to!).\n\n## The Heart: Cross Validation and Significance\n\nAfter some initial struggle with the large data and spending quite some time trying to get NN's to work, time was getting short for me. That helped me to concentrate on the important. I committed myself to the models that worked best, which were GBM's (I used a blend of LightGBM, XGBoost, and CatBoost in the end). Most importantly, I realised that a robust CV strategy *you stick to* is crucial.\n\nWe've seen that the quality of practically all models depended on the actual TTF. They do well for intermediate values and bad for very low or high values of TTF. So, depending on the distribution of TTF in the test / validation portion affects the score a lot. Hence, I tried to split the training data, such that this effect was minimised. At the same time, I got the feeling that I should take entire cycles in or out. Hence, I searched for splits of the 17 cycles into groups of three to four cycles that had similar TTF *and* size. One of those splits was\n| cycles | fraction | mean TTF |\n| --- | --- | --- |\n| 0, 3, 7, 8 | 21.0% | 5.86 |\n| 1, 10, 11 | 20.8% | 5.67 |\n| 2, 12, 15 | 19.9% | 5.68 |\n| 4, 5, 9, 16 | 19.7% | 5.56 |\n| 6, 13, 14 | 18.6% | 5.61 |\n\nGroups 0 and 16 are the very short \"cycles\" before and after the first and last earth quake. Note that with this split, I also don't have leakage between overlapping segments.\n\nThese folds still didn't give me the stable CV scores I wanted. I created some more of these even splits and decided to actually do four of these 5-fold group splits and take the mean MAE of each of the four 5-fold splits, such that I had four scores. This resulted in a rather slow CV scheme, but it was stable and I could not only work with a CV score, but also an uncertainty: `np.std(cv_mae,ddof=1)`, where `cv_mae` are the four averaged scores of the 5-folg group splits. These uncertainties were typically of order `0.005`. Note that this scheme assumes that the (private) test set has the same TTF distribution as the training set.\n\nAfter establishing that, I was in a situation where I could not only test models and features with a robust CV scheme, but also had a measure for what a significant improve in the score is.\n\n## What We Predict On\n\nThe assumption that the distributions of features and TTF in the test set, I've made in setting up my CV scheme, is quite reasonable. Some said otherwise, but I contradict. It was claimed by @gpreda that [the feature distribution was different for train and test set](https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda). I could not find that for my features as I've subtracted the mean of the acoustic data on each segment. That makes physically sense and I already did it before I knew about the comparison.\n\nAlso the assumption that the TTF distributions are similar should be reasonable. @mykper did some [nice work](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91583) on the public test set. We've seem a somewhat different distribution of the TTF and they also maxed out at some value below 10, whereas the training data had values up to 16. That's true. However, he also found out, that there are only two cycles in the public test set. These were quite typical, if you compare them with the training set. No reason to assume that we will predict on structurally different data than what we have seen in the training data.\n\nAlso one **note on the public LB**:\n\nThe public LB was calculated on just about 350 segments. This small number together with the high variance in the predictions alone should tell you that you cannot trust the public LB. Now compare that number with the size of the CV: we use several fold and will eventually use the entire training set for the score. That is 4200 segments. That's more than 10 times as much! When using overlapping segments - as I also did - the ratio gets even bigger, although the information gain will not grow linearly with the number of segments anymore. In anyway: the public LB score is calculated on a tiny set, especially when compared with the CV on the training data. I've never trusted the public LB in this competition.\n\n## The Features and Minimalism\n\nI took @CPMPml approach of feature selection he explained in the [no magic](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/92679#latest-543125) discussion and augmented it with my uncertainty of my CV scores. As we have little data, and I would have ended up with almost no features, if I'd only taken features with high significance. I decided to take features that improved my model by at least half a standard deviation.\n\nSome side note from a physicist:\nIf you have two measurements `m_1` and `m_2` with standard deviations `s_1` and `s_2`, the difference of the measurements `m_1 - m_2` does *not* have an uncertainty / standard deviation of `s_1 + s_2`, but of `sqrt(s_1^2 + s_2^2)`.\n\nWith that approach, I ended up with just 7 features!\n\n## Final Words\n\nThat's pretty much it. Some ensembeling of different GBM's (LightGBM, XGBoost, CatBoost) with (a little) optimised hyper-parameters and, as I've said in the beginning, probably a bit of luck.\n\nThere are some things I've seen in other solutions, that I'd like to have used, too. There is the special treatment of outliers as them ABC did (see their approach [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731)). Also nested CV would probably have done some good.\n\nMuch of the detective work on the test set, was not so fruitful, I think. Despite we even [know where it came from](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664). The only think that was useful - and probably quite a bit, although I didn't use it - was the information about its mean TTF. It seems to have helped a few teams. So after all, my assumption that the test set is not as strictly fulfilled as I've assumed in my CV scheme.",
    "544198": "Thanks for sharing, happy you find my post useful.  Congrats on the result!\n\n&gt; Some side note from a physicist:\n\nIt should be known by every data scientist that variance of a sum is the sum of variance.",
    "544228": "It should. I've encountered quite a few that don't, though... (Or they forget, get distracted by fancy stuff, ...). It doesn't hurt to refresh the basics from time to time.",
    "544255": "So you didn't use the leak, right?\nAnd as far as I understand you didn't try to make the train set similar to the test set? You made the folds of the train set homogenous (by several ways).",
    "544267": "Correct. I didn’t make the training set similar to the test set as there was no (strong enough) evidence we would actually habe a different test set. For me the risk was to high to overadjust or even adjust the wrong direction.",
    "544270": "Congrats on the gold medal!\n\nCan you please disclose what are the 7 selected features, do they include the mean?",
    "544271": "bernir \nVery cool! It seems that you're the only one in the gold zone who didn't use the leak or the test set.",
    "544294": "I don't have the 7 features at hand at the moment, but no, I did not use the mean. I substracted the mean from each and every segment, so I would be zero all the time anyways. (Note, that I did not normalise the standard deviation, as it contains information!)\n\nThe 7 features are no \"magic\" features and I'm pretty sure that one could find 5-10 other features that would perform equally good. Using only the features that significantly improve your model prevents overfitting. That's all whats going on here. I had engineered 900 features to choose from. Using all of them on just 4200 training examples (or a bit more with overlapping segments) would give you less than 5 data points per feature... That's badly constrained (although some features correlate strongly) and genearlises poorly.",
    "546001": "My seven features are named like that, if it helps:\n```python\nselected_features = [\n    'roll100mean_num_peaks_5',\n    'spectral_centroid_perc90',\n    'FFT_1650_1800',\n    'noise_perc75',\n    'raw_num_peaks_8',\n    'spectral_bandwidth_perc90',\n    'perio_iqr25',\n]\n```\nSo not necessarily very straight forward feaatures, but nothing too fancy or even magic either.",
    "553091": "Thank you for sharing your selected features.\n\n[Roles for Spectral Centroid and Other Factors in Determining \"Blended\" Instrument Pairings in Orchestration](https://www.jstor.org/stable/40285694) states: \"Blend worsened as a function of the overall centroid height of the combination or as the amount of difference between the centroids of the two instruments increased.\"\n\nMusicQuake and EarthQuake seem to be related!"
  },
  "source": "meta"
}