{
  "id": 91298,
  "title": "Posting My Masters Project",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91298",
  "author_name": "",
  "post_date": "2019-05-03T02:25:28.304062300Z",
  "votes": 95,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I have been doing the Kaggle LANL Earthquake Prediction Challenge as a final project for my Masters degree in Data Science.  My understanding of Kaggle rules is that by turning in my model code to my university professor and class, then I will be sharing the code, and Kaggle rules prohibit sharing for active competitions unless also shared on Kaggle.  I am therefore required, by rules and ethics, to post the kernels on Kaggle and give away what is currently a high public leader board solution.  \"Nobody\" posts kernels for high leader board solutions while a contest is in progress, but I feel I must because of the intersection of university needs and Kaggle rules.  This may trash my Kaggle leader board standing but it appears to be the right thing to do.  </p>\n\n<p>I never expected to be anywhere near so high on the public leader board for this project when I started the it.  Though, as has been discussed many times, the public leader board may be especially unreliable for this contest.  </p>\n\n<p>The kernels were written for a university course, I have simply copied them over here before posting for my class.  They were not really written for Kaggle, they were written for the university, so not all code will execute in the Kaggle environment, or run here at all.  I will work to fix plotting when my school work is completed.  These are my first ever posts to Kaggle's kernels, so please I hope that the community is supportive and kind.</p>\n\n<p>Special thanks to Preda and Lukayenko, plus the scripts that they cite, for the knowledge imparted which greatly helped my start on this journey. </p>\n\n<p>Vettejeep</p>",
  "messages": [
    {
      "id": "526416",
      "postDate": "05/03/2019 02:25:28",
      "content": "<p>I have been doing the Kaggle LANL Earthquake Prediction Challenge as a final project for my Masters degree in Data Science.  My understanding of Kaggle rules is that by turning in my model code to my university professor and class, then I will be sharing the code, and Kaggle rules prohibit sharing for active competitions unless also shared on Kaggle.  I am therefore required, by rules and ethics, to post the kernels on Kaggle and give away what is currently a high public leader board solution.  \"Nobody\" posts kernels for high leader board solutions while a contest is in progress, but I feel I must because of the intersection of university needs and Kaggle rules.  This may trash my Kaggle leader board standing but it appears to be the right thing to do.  </p>\n\n<p>I never expected to be anywhere near so high on the public leader board for this project when I started the it.  Though, as has been discussed many times, the public leader board may be especially unreliable for this contest.  </p>\n\n<p>The kernels were written for a university course, I have simply copied them over here before posting for my class.  They were not really written for Kaggle, they were written for the university, so not all code will execute in the Kaggle environment, or run here at all.  I will work to fix plotting when my school work is completed.  These are my first ever posts to Kaggle's kernels, so please I hope that the community is supportive and kind.</p>\n\n<p>Special thanks to Preda and Lukayenko, plus the scripts that they cite, for the knowledge imparted which greatly helped my start on this journey. </p>\n\n<p>Vettejeep</p>",
      "rawMarkdown": "I have been doing the Kaggle LANL Earthquake Prediction Challenge as a final project for my Masters degree in Data Science.  My understanding of Kaggle rules is that by turning in my model code to my university professor and class, then I will be sharing the code, and Kaggle rules prohibit sharing for active competitions unless also shared on Kaggle.  I am therefore required, by rules and ethics, to post the kernels on Kaggle and give away what is currently a high public leader board solution.  \"Nobody\" posts kernels for high leader board solutions while a contest is in progress, but I feel I must because of the intersection of university needs and Kaggle rules.  This may trash my Kaggle leader board standing but it appears to be the right thing to do.  \n\nI never expected to be anywhere near so high on the public leader board for this project when I started the it.  Though, as has been discussed many times, the public leader board may be especially unreliable for this contest.  \n\nThe kernels were written for a university course, I have simply copied them over here before posting for my class.  They were not really written for Kaggle, they were written for the university, so not all code will execute in the Kaggle environment, or run here at all.  I will work to fix plotting when my school work is completed.  These are my first ever posts to Kaggle's kernels, so please I hope that the community is supportive and kind.\n\nSpecial thanks to Preda and Lukayenko, plus the scripts that they cite, for the knowledge imparted which greatly helped my start on this journey. \n\nVettejeep",
      "votes": null
    },
    {
      "id": "526449",
      "postDate": "05/03/2019 04:45:13",
      "content": "<p>Amazing Congratulations!!!</p>",
      "rawMarkdown": "Amazing Congratulations!!!",
      "votes": null
    },
    {
      "id": "526465",
      "postDate": "05/03/2019 05:41:43",
      "content": "<p>Thank you for sharing this!\nYour EDA, feature creating and feature selection approaches are really great and interesting!</p>\n\n<p>I think that competition will benefit from this information and will push the boundaries of model quality even further.</p>",
      "rawMarkdown": "Thank you for sharing this!\nYour EDA, feature creating and feature selection approaches are really great and interesting!\n\nI think that competition will benefit from this information and will push the boundaries of model quality even further.",
      "votes": null
    },
    {
      "id": "526506",
      "postDate": "05/03/2019 07:24:05",
      "content": "<p>Great sharing!  I wonder why your EDA notebook isn’t getting as many votes as your model one.  It deserves better than this.  Also I like that you don’t provide a ready to fork/run/submit kernel.  Impact on LB will be smooth as a result.  Congrats on the good work so far.  I hope you’ll stay active on this competition and get to a gold medal.</p>",
      "rawMarkdown": "Great sharing!  I wonder why your EDA notebook isn’t getting as many votes as your model one.  It deserves better than this.  Also I like that you don’t provide a ready to fork/run/submit kernel.  Impact on LB will be smooth as a result.  Congrats on the good work so far.  I hope you’ll stay active on this competition and get to a gold medal.",
      "votes": null
    },
    {
      "id": "526530",
      "postDate": "05/03/2019 08:28:46",
      "content": "<p>Thanks, CPMP.  Probably the EDA will not do so well until I fix it so it runs on Kaggle and the charts show up.  As of now, it needs to be downloaded to run.  Also, it is late in the competition and models are probably more interesting than EDA at this point.   I am working on new ideas - no idea if they will pan out or not - but I am still trying.</p>",
      "rawMarkdown": "Thanks, CPMP.  Probably the EDA will not do so well until I fix it so it runs on Kaggle and the charts show up.  As of now, it needs to be downloaded to run.  Also, it is late in the competition and models are probably more interesting than EDA at this point.   I am working on new ideas - no idea if they will pan out or not - but I am still trying.",
      "votes": null
    },
    {
      "id": "526543",
      "postDate": "05/03/2019 08:51:22",
      "content": "<p>You're confirming a hunch I have: people vote for colorful pictures and not so much for the insights in a kernel. ;)</p>",
      "rawMarkdown": "You're confirming a hunch I have: people vote for colorful pictures and not so much for the insights in a kernel. ;)",
      "votes": null
    },
    {
      "id": "526721",
      "postDate": "05/03/2019 15:46:16",
      "content": "<p>Great sharing!!!! </p>",
      "rawMarkdown": "Great sharing!!!!",
      "votes": null
    },
    {
      "id": "526755",
      "postDate": "05/03/2019 17:14:26",
      "content": "<p>Thanks for sharing. Did I understand correctly that your models made the difference from Andrew's and Preda's kernels with augmentation (24,000 samples vs 4100)? \nI found it interesting that your models averaged out nicely. I tried averaging my lgb model (LB 1.399) with catboost model (LB 1.428) and got 1.401 (face palm), but hopefully they will do better on the private LB.</p>",
      "rawMarkdown": "Thanks for sharing. Did I understand correctly that your models made the difference from Andrew's and Preda's kernels with augmentation (24,000 samples vs 4100)? \nI found it interesting that your models averaged out nicely. I tried averaging my lgb model (LB 1.399) with catboost model (LB 1.428) and got 1.401 (face palm), but hopefully they will do better on the private LB.",
      "votes": null
    },
    {
      "id": "526764",
      "postDate": "05/03/2019 17:47:06",
      "content": "<p>Thanks for this.\nGood luck in your Masters degree in Data Science.  (I don't think you need to worry judging by the excellent quality of your notebooks)</p>",
      "rawMarkdown": "Thanks for this.\nGood luck in your Masters degree in Data Science.  (I don't think you need to worry judging by the excellent quality of your notebooks)",
      "votes": null
    },
    {
      "id": "526778",
      "postDate": "05/03/2019 18:25:33",
      "content": "<p>Since I did not have time to run all combos, maybe it is hard to tell what helped.  My main differences are:\n- 24k samples vs 4194\n- FFT Magnitude\n- 24k samples allows more features, so the low and band pass filters appear to add info to the model</p>\n\n<p>Even though the MSDS part is finished - I am still working to understand my own model.  Thanks.</p>",
      "rawMarkdown": "Since I did not have time to run all combos, maybe it is hard to tell what helped.  My main differences are:\n- 24k samples vs 4194\n- FFT Magnitude\n- 24k samples allows more features, so the low and band pass filters appear to add info to the model\n\nEven though the MSDS part is finished - I am still working to understand my own model.  Thanks.",
      "votes": null
    },
    {
      "id": "526804",
      "postDate": "05/03/2019 20:06:58",
      "content": "<p>Thanks for sharing both the EDA and model kernels. Will spend some part of the weekend going through them. Awesome work. 👍 </p>",
      "rawMarkdown": "Thanks for sharing both the EDA and model kernels. Will spend some part of the weekend going through them. Awesome work. 👍",
      "votes": null
    },
    {
      "id": "526838",
      "postDate": "05/03/2019 22:25:03",
      "content": "<p><a href=\"/vettejeep\">@vettejeep</a>, thanks for sharing. Filtering the data is a good idea. I upvoted your main kernel and aslo surprised it did not get much more votes than it has at this point. Good luck with your M.Sc degree.</p>",
      "rawMarkdown": "vettejeep, thanks for sharing. Filtering the data is a good idea. I upvoted your main kernel and aslo surprised it did not get much more votes than it has at this point. Good luck with your M.Sc degree.",
      "votes": null
    },
    {
      "id": "526940",
      "postDate": "05/04/2019 06:46:38",
      "content": "<p>I have fixed my EDA kernel so it runs on Kaggle.  Thanks to <a href=\"/corochann\">@corochann</a> who also posted the same fixes I found.</p>",
      "rawMarkdown": "I have fixed my EDA kernel so it runs on Kaggle.  Thanks to @corochann who also posted the same fixes I found.",
      "votes": null
    },
    {
      "id": "526991",
      "postDate": "05/04/2019 09:44:29",
      "content": "<p>Thks a lot for your sharing! It really helps!</p>",
      "rawMarkdown": "Thks a lot for your sharing! It really helps!",
      "votes": null
    },
    {
      "id": "527000",
      "postDate": "05/04/2019 10:55:23",
      "content": "<p><a href=\"/vettejeep\">@vettejeep</a> Very appreciate your excellent work and good luck in your Masters degree in Data Science :)! </p>",
      "rawMarkdown": "vettejeep Very appreciate your excellent work and good luck in your Masters degree in Data Science :)!",
      "votes": null
    },
    {
      "id": "527276",
      "postDate": "05/05/2019 01:57:12",
      "content": "<p>Thank you for sharing and congratulations again on your Masters project. I have a general question about your data augmentation strategy. I understand that you split the training data into 6 segments, and your feature dataframe was constructed from 4000 random start indices in each segment. You then employed a shuffled 8-fold CV with the total data. </p>\n\n<p>My question is how you achieved such success with this method given its potential for leakage. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always decreases. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score substantially decreased with the oversampled data compared to the regular data.</p>\n\n<p>Am I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score increases and the LB score gets worse. </p>",
      "rawMarkdown": "Thank you for sharing and congratulations again on your Masters project. I have a general question about your data augmentation strategy. I understand that you split the training data into 6 segments, and your feature dataframe was constructed from 4000 random start indices in each segment. You then employed a shuffled 8-fold CV with the total data. \n\nMy question is how you achieved such success with this method given its potential for leakage. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always decreases. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score substantially decreased with the oversampled data compared to the regular data.\n\nAm I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score increases and the LB score gets worse.",
      "votes": null
    },
    {
      "id": "527292",
      "postDate": "05/05/2019 03:00:08",
      "content": "<p>It is a <em>great</em> question.  I think it worked out for me because going to 24k samples allowed more features, specifically those created by generating bandwidth limited versions of the signals.  Even the filters leak a little, since brick wall filters are harder to create.  Because of leakage, CV score is better than LB score, due to random sampling in the CV.  Another option is to stratify the CV by slice, this will reduce leakage but has not so far helped with LB score.  LB score may be unreliable, so...</p>\n\n<p>One thing \"going big\" with sampling let me do was add features since it helps with the \"big p small n\" problem.</p>\n\n<p>CV by slice worsens problem of sampling the large failure time signals.  Some have complained that their models have trouble with long quake signals.  Models that stall out at 11 seconds or so.  I looked at one of my submission files and it got to almost 14.  Also, I can get over 14 with more tuning than is in my project files.  But LB score went down.  But, LB may be unreliable...</p>",
      "rawMarkdown": "It is a *great* question.  I think it worked out for me because going to 24k samples allowed more features, specifically those created by generating bandwidth limited versions of the signals.  Even the filters leak a little, since brick wall filters are harder to create.  Because of leakage, CV score is better than LB score, due to random sampling in the CV.  Another option is to stratify the CV by slice, this will reduce leakage but has not so far helped with LB score.  LB score may be unreliable, so...\n\nOne thing \"going big\" with sampling let me do was add features since it helps with the \"big p small n\" problem.\n\nCV by slice worsens problem of sampling the large failure time signals.  Some have complained that their models have trouble with long quake signals.  Models that stall out at 11 seconds or so.  I looked at one of my submission files and it got to almost 14.  Also, I can get over 14 with more tuning than is in my project files.  But LB score went down.  But, LB may be unreliable...",
      "votes": null
    },
    {
      "id": "527422",
      "postDate": "05/05/2019 13:05:09",
      "content": "<p>I can see that an increased number of samples would help determine features with a good p-value. A real concern with smaller datasets is that models will find spurious feature importance due to natural variance. <a href=\"/ogrellier\">@ogrellier</a> has <a href=\"https://www.kaggle.com/ogrellier/feature-selection-with-null-importances\">an excellent kernel here</a> for identifying these features. The basic principle is to shuffle the target value, so that the 'null' importance of features can be found and used as a baseline. This baseline can then be compared to the feature importances derived from models trained on the actual target. </p>\n\n<p>I must admit, when I tried oversampling, I didn't try any additional measures like feature reduction. I used the exact same features and parameters since I wanted to be 'scientific' when comparing the original and augmented datasets. Perhaps I have been taking the wrong approach.</p>\n\n<p>I would really appreciate some input from any resident grandmasters! I feel there is a serious gap in my understanding here. I simply do not understand why data augmentation is giving me such bad results. I was considering generating a new dataset that would triple the rows, and for the generated rows using 50% actual features data, and 50% random feature values generated from the existing distribution.</p>",
      "rawMarkdown": "I can see that an increased number of samples would help determine features with a good p-value. A real concern with smaller datasets is that models will find spurious feature importance due to natural variance. @ogrellier has [an excellent kernel here](https://www.kaggle.com/ogrellier/feature-selection-with-null-importances) for identifying these features. The basic principle is to shuffle the target value, so that the 'null' importance of features can be found and used as a baseline. This baseline can then be compared to the feature importances derived from models trained on the actual target. \n\nI must admit, when I tried oversampling, I didn't try any additional measures like feature reduction. I used the exact same features and parameters since I wanted to be 'scientific' when comparing the original and augmented datasets. Perhaps I have been taking the wrong approach.\n\nI would really appreciate some input from any resident grandmasters! I feel there is a serious gap in my understanding here. I simply do not understand why data augmentation is giving me such bad results. I was considering generating a new dataset that would triple the rows, and for the generated rows using 50% actual features data, and 50% random feature values generated from the existing distribution.",
      "votes": null
    },
    {
      "id": "527474",
      "postDate": "05/05/2019 16:07:41",
      "content": "<p>thanks sharing!!</p>",
      "rawMarkdown": "thanks sharing!!",
      "votes": null
    },
    {
      "id": "528037",
      "postDate": "05/06/2019 23:38:53",
      "content": "<p>Tks!! that's awesome</p>",
      "rawMarkdown": "Tks!! that's awesome",
      "votes": null
    },
    {
      "id": "528962",
      "postDate": "05/09/2019 00:24:01",
      "content": "<p>It's wonderful!!</p>",
      "rawMarkdown": "It's wonderful!!",
      "votes": null
    },
    {
      "id": "528966",
      "postDate": "05/09/2019 00:50:56",
      "content": "<p>good job</p>",
      "rawMarkdown": "good job",
      "votes": null
    },
    {
      "id": "534571",
      "postDate": "05/21/2019 13:30:23",
      "content": "<p>Kudos 💯 </p>",
      "rawMarkdown": "Kudos 💯",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 526449,
      "author_name": "akhileshrai",
      "author_url": "",
      "post_date": "05/03/2019 04:45:13",
      "content": "<p>Amazing Congratulations!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526465,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "05/03/2019 05:41:43",
      "content": "<p>Thank you for sharing this!\nYour EDA, feature creating and feature selection approaches are really great and interesting!</p>\n\n<p>I think that competition will benefit from this information and will push the boundaries of model quality even further.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526506,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/03/2019 07:24:05",
      "content": "<p>Great sharing!  I wonder why your EDA notebook isn’t getting as many votes as your model one.  It deserves better than this.  Also I like that you don’t provide a ready to fork/run/submit kernel.  Impact on LB will be smooth as a result.  Congrats on the good work so far.  I hope you’ll stay active on this competition and get to a gold medal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 526530,
          "author_name": "vettejeep",
          "author_url": "",
          "post_date": "05/03/2019 08:28:46",
          "content": "<p>Thanks, CPMP.  Probably the EDA will not do so well until I fix it so it runs on Kaggle and the charts show up.  As of now, it needs to be downloaded to run.  Also, it is late in the competition and models are probably more interesting than EDA at this point.   I am working on new ideas - no idea if they will pan out or not - but I am still trying.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 526543,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/03/2019 08:51:22",
          "content": "<p>You're confirming a hunch I have: people vote for colorful pictures and not so much for the insights in a kernel. ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 526721,
      "author_name": "himusoni",
      "author_url": "",
      "post_date": "05/03/2019 15:46:16",
      "content": "<p>Great sharing!!!! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526755,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "05/03/2019 17:14:26",
      "content": "<p>Thanks for sharing. Did I understand correctly that your models made the difference from Andrew's and Preda's kernels with augmentation (24,000 samples vs 4100)? \nI found it interesting that your models averaged out nicely. I tried averaging my lgb model (LB 1.399) with catboost model (LB 1.428) and got 1.401 (face palm), but hopefully they will do better on the private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 526778,
          "author_name": "vettejeep",
          "author_url": "",
          "post_date": "05/03/2019 18:25:33",
          "content": "<p>Since I did not have time to run all combos, maybe it is hard to tell what helped.  My main differences are:\n- 24k samples vs 4194\n- FFT Magnitude\n- 24k samples allows more features, so the low and band pass filters appear to add info to the model</p>\n\n<p>Even though the MSDS part is finished - I am still working to understand my own model.  Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 526764,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/03/2019 17:47:06",
      "content": "<p>Thanks for this.\nGood luck in your Masters degree in Data Science.  (I don't think you need to worry judging by the excellent quality of your notebooks)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526804,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "05/03/2019 20:06:58",
      "content": "<p>Thanks for sharing both the EDA and model kernels. Will spend some part of the weekend going through them. Awesome work. 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526838,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/03/2019 22:25:03",
      "content": "<p><a href=\"/vettejeep\">@vettejeep</a>, thanks for sharing. Filtering the data is a good idea. I upvoted your main kernel and aslo surprised it did not get much more votes than it has at this point. Good luck with your M.Sc degree.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526940,
      "author_name": "vettejeep",
      "author_url": "",
      "post_date": "05/04/2019 06:46:38",
      "content": "<p>I have fixed my EDA kernel so it runs on Kaggle.  Thanks to <a href=\"/corochann\">@corochann</a> who also posted the same fixes I found.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 526991,
      "author_name": "lyf19950404",
      "author_url": "",
      "post_date": "05/04/2019 09:44:29",
      "content": "<p>Thks a lot for your sharing! It really helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527000,
      "author_name": "hsinwenchang",
      "author_url": "",
      "post_date": "05/04/2019 10:55:23",
      "content": "<p><a href=\"/vettejeep\">@vettejeep</a> Very appreciate your excellent work and good luck in your Masters degree in Data Science :)! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527276,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "05/05/2019 01:57:12",
      "content": "<p>Thank you for sharing and congratulations again on your Masters project. I have a general question about your data augmentation strategy. I understand that you split the training data into 6 segments, and your feature dataframe was constructed from 4000 random start indices in each segment. You then employed a shuffled 8-fold CV with the total data. </p>\n\n<p>My question is how you achieved such success with this method given its potential for leakage. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always decreases. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score substantially decreased with the oversampled data compared to the regular data.</p>\n\n<p>Am I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score increases and the LB score gets worse. </p>",
      "votes": null,
      "replies": [
        {
          "id": 527292,
          "author_name": "vettejeep",
          "author_url": "",
          "post_date": "05/05/2019 03:00:08",
          "content": "<p>It is a <em>great</em> question.  I think it worked out for me because going to 24k samples allowed more features, specifically those created by generating bandwidth limited versions of the signals.  Even the filters leak a little, since brick wall filters are harder to create.  Because of leakage, CV score is better than LB score, due to random sampling in the CV.  Another option is to stratify the CV by slice, this will reduce leakage but has not so far helped with LB score.  LB score may be unreliable, so...</p>\n\n<p>One thing \"going big\" with sampling let me do was add features since it helps with the \"big p small n\" problem.</p>\n\n<p>CV by slice worsens problem of sampling the large failure time signals.  Some have complained that their models have trouble with long quake signals.  Models that stall out at 11 seconds or so.  I looked at one of my submission files and it got to almost 14.  Also, I can get over 14 with more tuning than is in my project files.  But LB score went down.  But, LB may be unreliable...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527422,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "05/05/2019 13:05:09",
          "content": "<p>I can see that an increased number of samples would help determine features with a good p-value. A real concern with smaller datasets is that models will find spurious feature importance due to natural variance. <a href=\"/ogrellier\">@ogrellier</a> has <a href=\"https://www.kaggle.com/ogrellier/feature-selection-with-null-importances\">an excellent kernel here</a> for identifying these features. The basic principle is to shuffle the target value, so that the 'null' importance of features can be found and used as a baseline. This baseline can then be compared to the feature importances derived from models trained on the actual target. </p>\n\n<p>I must admit, when I tried oversampling, I didn't try any additional measures like feature reduction. I used the exact same features and parameters since I wanted to be 'scientific' when comparing the original and augmented datasets. Perhaps I have been taking the wrong approach.</p>\n\n<p>I would really appreciate some input from any resident grandmasters! I feel there is a serious gap in my understanding here. I simply do not understand why data augmentation is giving me such bad results. I was considering generating a new dataset that would triple the rows, and for the generated rows using 50% actual features data, and 50% random feature values generated from the existing distribution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527474,
      "author_name": "meistermorxrc",
      "author_url": "",
      "post_date": "05/05/2019 16:07:41",
      "content": "<p>thanks sharing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528037,
      "author_name": "frmatias",
      "author_url": "",
      "post_date": "05/06/2019 23:38:53",
      "content": "<p>Tks!! that's awesome</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528962,
      "author_name": "bladerunnerrachel",
      "author_url": "",
      "post_date": "05/09/2019 00:24:01",
      "content": "<p>It's wonderful!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 528966,
      "author_name": "qlittle",
      "author_url": "",
      "post_date": "05/09/2019 00:50:56",
      "content": "<p>good job</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 534571,
      "author_name": "princeanddatascience",
      "author_url": "",
      "post_date": "05/21/2019 13:30:23",
      "content": "<p>Kudos 💯 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "526416": "I have been doing the Kaggle LANL Earthquake Prediction Challenge as a final project for my Masters degree in Data Science.  My understanding of Kaggle rules is that by turning in my model code to my university professor and class, then I will be sharing the code, and Kaggle rules prohibit sharing for active competitions unless also shared on Kaggle.  I am therefore required, by rules and ethics, to post the kernels on Kaggle and give away what is currently a high public leader board solution.  \"Nobody\" posts kernels for high leader board solutions while a contest is in progress, but I feel I must because of the intersection of university needs and Kaggle rules.  This may trash my Kaggle leader board standing but it appears to be the right thing to do.  \n\nI never expected to be anywhere near so high on the public leader board for this project when I started the it.  Though, as has been discussed many times, the public leader board may be especially unreliable for this contest.  \n\nThe kernels were written for a university course, I have simply copied them over here before posting for my class.  They were not really written for Kaggle, they were written for the university, so not all code will execute in the Kaggle environment, or run here at all.  I will work to fix plotting when my school work is completed.  These are my first ever posts to Kaggle's kernels, so please I hope that the community is supportive and kind.\n\nSpecial thanks to Preda and Lukayenko, plus the scripts that they cite, for the knowledge imparted which greatly helped my start on this journey. \n\nVettejeep",
    "526449": "Amazing Congratulations!!!",
    "526465": "Thank you for sharing this!\nYour EDA, feature creating and feature selection approaches are really great and interesting!\n\nI think that competition will benefit from this information and will push the boundaries of model quality even further.",
    "526506": "Great sharing!  I wonder why your EDA notebook isn’t getting as many votes as your model one.  It deserves better than this.  Also I like that you don’t provide a ready to fork/run/submit kernel.  Impact on LB will be smooth as a result.  Congrats on the good work so far.  I hope you’ll stay active on this competition and get to a gold medal.",
    "526530": "Thanks, CPMP.  Probably the EDA will not do so well until I fix it so it runs on Kaggle and the charts show up.  As of now, it needs to be downloaded to run.  Also, it is late in the competition and models are probably more interesting than EDA at this point.   I am working on new ideas - no idea if they will pan out or not - but I am still trying.",
    "526543": "You're confirming a hunch I have: people vote for colorful pictures and not so much for the insights in a kernel. ;)",
    "526721": "Great sharing!!!!",
    "526755": "Thanks for sharing. Did I understand correctly that your models made the difference from Andrew's and Preda's kernels with augmentation (24,000 samples vs 4100)? \nI found it interesting that your models averaged out nicely. I tried averaging my lgb model (LB 1.399) with catboost model (LB 1.428) and got 1.401 (face palm), but hopefully they will do better on the private LB.",
    "526764": "Thanks for this.\nGood luck in your Masters degree in Data Science.  (I don't think you need to worry judging by the excellent quality of your notebooks)",
    "526778": "Since I did not have time to run all combos, maybe it is hard to tell what helped.  My main differences are:\n- 24k samples vs 4194\n- FFT Magnitude\n- 24k samples allows more features, so the low and band pass filters appear to add info to the model\n\nEven though the MSDS part is finished - I am still working to understand my own model.  Thanks.",
    "526804": "Thanks for sharing both the EDA and model kernels. Will spend some part of the weekend going through them. Awesome work. 👍",
    "526838": "vettejeep, thanks for sharing. Filtering the data is a good idea. I upvoted your main kernel and aslo surprised it did not get much more votes than it has at this point. Good luck with your M.Sc degree.",
    "526940": "I have fixed my EDA kernel so it runs on Kaggle.  Thanks to @corochann who also posted the same fixes I found.",
    "526991": "Thks a lot for your sharing! It really helps!",
    "527000": "vettejeep Very appreciate your excellent work and good luck in your Masters degree in Data Science :)!",
    "527276": "Thank you for sharing and congratulations again on your Masters project. I have a general question about your data augmentation strategy. I understand that you split the training data into 6 segments, and your feature dataframe was constructed from 4000 random start indices in each segment. You then employed a shuffled 8-fold CV with the total data. \n\nMy question is how you achieved such success with this method given its potential for leakage. When I try naive 2x oversampling by running feature generation on 150,000 rows with a 75,000 row step, my LB score always decreases. When I attributed this to shuffled K-fold leakage, I tried a quake-wise split with no (or at least limited) leakage. Again, my score substantially decreased with the oversampled data compared to the regular data.\n\nAm I missing something fundamental to the data augmentation process? My approach was designed to limit the number of times any range of the data could be sampled more than once. But every time I use it, my CV score increases and the LB score gets worse.",
    "527292": "It is a *great* question.  I think it worked out for me because going to 24k samples allowed more features, specifically those created by generating bandwidth limited versions of the signals.  Even the filters leak a little, since brick wall filters are harder to create.  Because of leakage, CV score is better than LB score, due to random sampling in the CV.  Another option is to stratify the CV by slice, this will reduce leakage but has not so far helped with LB score.  LB score may be unreliable, so...\n\nOne thing \"going big\" with sampling let me do was add features since it helps with the \"big p small n\" problem.\n\nCV by slice worsens problem of sampling the large failure time signals.  Some have complained that their models have trouble with long quake signals.  Models that stall out at 11 seconds or so.  I looked at one of my submission files and it got to almost 14.  Also, I can get over 14 with more tuning than is in my project files.  But LB score went down.  But, LB may be unreliable...",
    "527422": "I can see that an increased number of samples would help determine features with a good p-value. A real concern with smaller datasets is that models will find spurious feature importance due to natural variance. @ogrellier has [an excellent kernel here](https://www.kaggle.com/ogrellier/feature-selection-with-null-importances) for identifying these features. The basic principle is to shuffle the target value, so that the 'null' importance of features can be found and used as a baseline. This baseline can then be compared to the feature importances derived from models trained on the actual target. \n\nI must admit, when I tried oversampling, I didn't try any additional measures like feature reduction. I used the exact same features and parameters since I wanted to be 'scientific' when comparing the original and augmented datasets. Perhaps I have been taking the wrong approach.\n\nI would really appreciate some input from any resident grandmasters! I feel there is a serious gap in my understanding here. I simply do not understand why data augmentation is giving me such bad results. I was considering generating a new dataset that would triple the rows, and for the generated rows using 50% actual features data, and 50% random feature values generated from the existing distribution.",
    "527474": "thanks sharing!!",
    "528037": "Tks!! that's awesome",
    "528962": "It's wonderful!!",
    "528966": "good job",
    "534571": "Kudos 💯"
  },
  "source": "meta"
}