{
  "id": 89366,
  "title": "Interesting insight from shuffling ?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89366",
  "author_name": "",
  "post_date": "2019-04-13T11:43:00.026605100Z",
  "votes": 54,
  "comment_count": 49,
  "views": 0,
  "content": "<p>I've been wondering what made the difference between local CV and LB and decided to conduct a few experiments...</p>\n\n<p>Here is something I find particularly interesting, using almost the same features as <a href=\"https://www.kaggle.com/inversion/basic-feature-benchmark\">Inversion</a></p>\n\n<p>Using 10-fold cross-validation with shuffling, i.e. <code>KFold(10, True, 1)</code>:\n<code>\nFold 1 MAE : 2.2725\nFold 2 MAE : 2.1927\nFold 3 MAE : 2.1419\nFold 4 MAE : 2.2639\nFold 5 MAE : 2.2061\nFold 6 MAE : 2.1068\nFold 7 MAE : 2.1524\nFold 8 MAE : 2.1555\nFold 9 MAE : 2.2781\nFold 10 MAE : 2.2490\nFull MAE : 2.2019\n</code></p>\n\n<p>Now without shuffling, i.e. <code>KFold(10, False, 1)</code>\n<code>\nFold 1 MAE : 2.5836\nFold 2 MAE : 1.8352\nFold 3 MAE : 2.1289\nFold 4 MAE : 2.9439\nFold 5 MAE : 3.1463\nFold 6 MAE : 2.0426\nFold 7 MAE : 1.5974\nFold 8 MAE : 1.5868\nFold 9 MAE : 3.3176\nFold 10 MAE : 2.1363\nFull MAE : 2.3317\n</code></p>\n\n<p>Note that I shuffle the training set inside the cross-validation loop.</p>\n\n<p>My bet is that there is sufficient correlation between training samples, which may lead to leakage between training and validation sets if you shuffle samples before cross-validation.</p>\n\n<p>I really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !</p>\n\n<p>What do you think ?</p>",
  "messages": [
    {
      "id": "515946",
      "postDate": "04/13/2019 11:43:00",
      "content": "<p>I've been wondering what made the difference between local CV and LB and decided to conduct a few experiments...</p>\n\n<p>Here is something I find particularly interesting, using almost the same features as <a href=\"https://www.kaggle.com/inversion/basic-feature-benchmark\">Inversion</a></p>\n\n<p>Using 10-fold cross-validation with shuffling, i.e. <code>KFold(10, True, 1)</code>:\n<code>\nFold 1 MAE : 2.2725\nFold 2 MAE : 2.1927\nFold 3 MAE : 2.1419\nFold 4 MAE : 2.2639\nFold 5 MAE : 2.2061\nFold 6 MAE : 2.1068\nFold 7 MAE : 2.1524\nFold 8 MAE : 2.1555\nFold 9 MAE : 2.2781\nFold 10 MAE : 2.2490\nFull MAE : 2.2019\n</code></p>\n\n<p>Now without shuffling, i.e. <code>KFold(10, False, 1)</code>\n<code>\nFold 1 MAE : 2.5836\nFold 2 MAE : 1.8352\nFold 3 MAE : 2.1289\nFold 4 MAE : 2.9439\nFold 5 MAE : 3.1463\nFold 6 MAE : 2.0426\nFold 7 MAE : 1.5974\nFold 8 MAE : 1.5868\nFold 9 MAE : 3.3176\nFold 10 MAE : 2.1363\nFull MAE : 2.3317\n</code></p>\n\n<p>Note that I shuffle the training set inside the cross-validation loop.</p>\n\n<p>My bet is that there is sufficient correlation between training samples, which may lead to leakage between training and validation sets if you shuffle samples before cross-validation.</p>\n\n<p>I really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !</p>\n\n<p>What do you think ?</p>",
      "rawMarkdown": "I've been wondering what made the difference between local CV and LB and decided to conduct a few experiments...\n\nHere is something I find particularly interesting, using almost the same features as [Inversion](https://www.kaggle.com/inversion/basic-feature-benchmark)\n\nUsing 10-fold cross-validation with shuffling, i.e. `KFold(10, True, 1)`:\n```\nFold 1 MAE : 2.2725\nFold 2 MAE : 2.1927\nFold 3 MAE : 2.1419\nFold 4 MAE : 2.2639\nFold 5 MAE : 2.2061\nFold 6 MAE : 2.1068\nFold 7 MAE : 2.1524\nFold 8 MAE : 2.1555\nFold 9 MAE : 2.2781\nFold 10 MAE : 2.2490\nFull MAE : 2.2019\n```\n\nNow without shuffling, i.e. `KFold(10, False, 1)`\n```\nFold 1 MAE : 2.5836\nFold 2 MAE : 1.8352\nFold 3 MAE : 2.1289\nFold 4 MAE : 2.9439\nFold 5 MAE : 3.1463\nFold 6 MAE : 2.0426\nFold 7 MAE : 1.5974\nFold 8 MAE : 1.5868\nFold 9 MAE : 3.3176\nFold 10 MAE : 2.1363\nFull MAE : 2.3317\n```\n\nNote that I shuffle the training set inside the cross-validation loop.\n\nMy bet is that there is sufficient correlation between training samples, which may lead to leakage between training and validation sets if you shuffle samples before cross-validation.\n\nI really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !\n\nWhat do you think ?",
      "votes": null
    },
    {
      "id": "515963",
      "postDate": "04/13/2019 12:33:59",
      "content": "<p>Interesting, Olivier! \nAm I right that you killed all time related information within folds and get better score? </p>",
      "rawMarkdown": "Interesting, Olivier! \nAm I right that you killed all time related information within folds and get better score?",
      "votes": null
    },
    {
      "id": "515976",
      "postDate": "04/13/2019 13:11:07",
      "content": "<p>!!!  Maybe we can find some leaks here. </p>",
      "rawMarkdown": "!!!  Maybe we can find some leaks here.",
      "votes": null
    },
    {
      "id": "515977",
      "postDate": "04/13/2019 13:15:49",
      "content": "<p>the most important thing in this competition is the fact that the test set segments are not continous as are in the train set. Hence, i believe that validating with suffling may lead us to better LB scores</p>",
      "rawMarkdown": "the most important thing in this competition is the fact that the test set segments are not continous as are in the train set. Hence, i believe that validating with suffling may lead us to better LB scores",
      "votes": null
    },
    {
      "id": "515998",
      "postDate": "04/13/2019 14:05:36",
      "content": "<p>Early stopping would be problematic without shuffling, just my thought</p>",
      "rawMarkdown": "Early stopping would be problematic without shuffling, just my thought",
      "votes": null
    },
    {
      "id": "516035",
      "postDate": "04/13/2019 14:50:13",
      "content": "<p><a href=\"/alexfir\">@alexfir</a>, yes if you shuffle (remove time) then you get a better CV and a better LB.</p>\n\n<p>However I would assume that taking 13% of the private LB as a validation is fairly risky  ;-)</p>\n\n<p>I don't know if it's about time or about the fact that if 2 values (std, max, min or anything else) are close then the samples must be close as well and therefore finding the *time_to_failure* is easy...</p>",
      "rawMarkdown": "alexfir, yes if you shuffle (remove time) then you get a better CV and a better LB.\n\nHowever I would assume that taking 13% of the private LB as a validation is fairly risky  ;-)\n\nI don't know if it's about time or about the fact that if 2 values (std, max, min or anything else) are close then the samples must be close as well and therefore finding the *time_to_failure* is easy...",
      "votes": null
    },
    {
      "id": "516040",
      "postDate": "04/13/2019 14:54:27",
      "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a>, you may be right. </p>\n\n<p>I just wanted to draw people's attention on the fact that the only way to reproduce public LB score is to not shuffle.</p>",
      "rawMarkdown": "dkaraflos, you may be right. \n\nI just wanted to draw people's attention on the fact that the only way to reproduce public LB score is to not shuffle.",
      "votes": null
    },
    {
      "id": "516092",
      "postDate": "04/13/2019 15:47:41",
      "content": "<p>the test set might not be continues, but they are part of a continues signal</p>",
      "rawMarkdown": "the test set might not be continues, but they are part of a continues signal",
      "votes": null
    },
    {
      "id": "516102",
      "postDate": "04/13/2019 16:02:10",
      "content": "<p>exactly. That is why the FE should treat each segment individually </p>",
      "rawMarkdown": "exactly. That is why the FE should treat each segment individually",
      "votes": null
    },
    {
      "id": "516222",
      "postDate": "04/13/2019 19:54:59",
      "content": "<p>I have the feeling the  leakage may occur when segments from the same quake are in both training and validation. This happens a lot with shuffling and less so without shuffling. \nI have used stratified shuffle sampling for a while and then turned to quake-wise split for this reason. </p>",
      "rawMarkdown": "I have the feeling the  leakage may occur when segments from the same quake are in both training and validation. This happens a lot with shuffling and less so without shuffling. \nI have used stratified shuffle sampling for a while and then turned to quake-wise split for this reason.",
      "votes": null
    },
    {
      "id": "516243",
      "postDate": "04/13/2019 20:45:29",
      "content": "<p>Thanks <a href=\"/stecasasso\">@stecasasso</a> for your comments. I totally agree with that ;-)</p>",
      "rawMarkdown": "Thanks @stecasasso for your comments. I totally agree with that ;-)",
      "votes": null
    },
    {
      "id": "516254",
      "postDate": "04/13/2019 21:22:52",
      "content": "<p>I am new to this competition, so i am making my baseline models and i started my EDA.\nSO, in this competition, each segment is a part of an earthquake, but there are many sperate earthquakes? then the validation setup must ensure that we must shuffle the segments and also ensure that there are no segments of the same earthquake between training data and validation data</p>",
      "rawMarkdown": "I am new to this competition, so i am making my baseline models and i started my EDA.\nSO, in this competition, each segment is a part of an earthquake, but there are many sperate earthquakes? then the validation setup must ensure that we must shuffle the segments and also ensure that there are no segments of the same earthquake between training data and validation data",
      "votes": null
    },
    {
      "id": "516327",
      "postDate": "04/14/2019 02:27:40",
      "content": "<p>In the training set there are 16 places where ttf reaches zero (actually it never reaches zero, but close enough) and resets, these 16 resets of the countdown are taken as the boundaries between the quakes. </p>\n\n<p>The starter kernel and most others since then, breaks the continuous 600 million row csv into non-overlapping segments of 150k rows starting from row zero, and computes whatever features per segment for these ~4100 segments. This seems like a natural way to treat the training data as the test data consists of ~2k individual segments of 150k rows of acoustic data and it is from these that we are to predict the time to failure, in seconds, from the acoustic signal. Doing it this way means 16 of the segments will have a ttf reset somewhere in their rows.</p>\n\n<p>We know that the test data supposedly came from the same experiment that generated the training data, we do not know if the testing segments also contain places where ttf resets but it seems reasonable to me to assume so, to assume there are somewhere around 8 resets in test, and that what this means is quite unclear. </p>\n\n<p>The validation problem here is really interesting, many have suggested breaking up the quakes and using the rows from one or more quakes as validation for the others. The LB seems to be almost entirely useless here and strongly biased towards low-ttf quakes which is easy to check, you can make a submission with some cv score, then make one with a worse cv but a lower mean predicted ttf, the one with lower mean will usually score better. I don't know the best way to do it yet.</p>",
      "rawMarkdown": "In the training set there are 16 places where ttf reaches zero (actually it never reaches zero, but close enough) and resets, these 16 resets of the countdown are taken as the boundaries between the quakes. \n\nThe starter kernel and most others since then, breaks the continuous 600 million row csv into non-overlapping segments of 150k rows starting from row zero, and computes whatever features per segment for these ~4100 segments. This seems like a natural way to treat the training data as the test data consists of ~2k individual segments of 150k rows of acoustic data and it is from these that we are to predict the time to failure, in seconds, from the acoustic signal. Doing it this way means 16 of the segments will have a ttf reset somewhere in their rows.\n\nWe know that the test data supposedly came from the same experiment that generated the training data, we do not know if the testing segments also contain places where ttf resets but it seems reasonable to me to assume so, to assume there are somewhere around 8 resets in test, and that what this means is quite unclear. \n\nThe validation problem here is really interesting, many have suggested breaking up the quakes and using the rows from one or more quakes as validation for the others. The LB seems to be almost entirely useless here and strongly biased towards low-ttf quakes which is easy to check, you can make a submission with some cv score, then make one with a worse cv but a lower mean predicted ttf, the one with lower mean will usually score better. I don't know the best way to do it yet.",
      "votes": null
    },
    {
      "id": "516329",
      "postDate": "04/14/2019 02:28:50",
      "content": "<p>I counted 16 quakes, no improvement yet...</p>",
      "rawMarkdown": "I counted 16 quakes, no improvement yet...",
      "votes": null
    },
    {
      "id": "516400",
      "postDate": "04/14/2019 05:38:19",
      "content": "<p>What is ttf?</p>",
      "rawMarkdown": "What is ttf?",
      "votes": null
    },
    {
      "id": "516402",
      "postDate": "04/14/2019 05:49:02",
      "content": "<p>I think that the huge fluctuation when using kfold without shuffle is because the model can't predict high TTFs. Each fold has 1~2 earthquakes for validation; if this quake has high TTF the score is bad (e.g. fold 5 is probably the earthquake with 16s). On the other hand, if the quake has around 10s the score will be quite good on that fold.</p>",
      "rawMarkdown": "I think that the huge fluctuation when using kfold without shuffle is because the model can't predict high TTFs. Each fold has 1~2 earthquakes for validation; if this quake has high TTF the score is bad (e.g. fold 5 is probably the earthquake with 16s). On the other hand, if the quake has around 10s the score will be quite good on that fold.",
      "votes": null
    },
    {
      "id": "516410",
      "postDate": "04/14/2019 06:05:41",
      "content": "<p><a href=\"/timmmmmms\">@timmmmmms</a>, I guess this stands for Time To Failure ;-)</p>",
      "rawMarkdown": "timmmmmms, I guess this stands for Time To Failure ;-)",
      "votes": null
    },
    {
      "id": "516486",
      "postDate": "04/14/2019 09:20:59",
      "content": "<p><a href=\"/interneuron\">@interneuron</a> you are right about the earthquake numbers, but what puzzles me is the fact that the competition host did not mention such a data property.\nThus, in this competition, the validation setup is a really crucial step. We must think of it very robustly</p>",
      "rawMarkdown": "interneuron you are right about the earthquake numbers, but what puzzles me is the fact that the competition host did not mention such a data property.\nThus, in this competition, the validation setup is a really crucial step. We must think of it very robustly",
      "votes": null
    },
    {
      "id": "516620",
      "postDate": "04/14/2019 15:08:42",
      "content": "<p>I've spent too long on this data... I can literally see the two earthquakes that gave him MAE &gt; 3s</p>",
      "rawMarkdown": "I've spent too long on this data... I can literally see the two earthquakes that gave him MAE &gt; 3s",
      "votes": null
    },
    {
      "id": "516997",
      "postDate": "04/15/2019 10:32:40",
      "content": "<p>Similar experiment.\nWithout shuffling.\n<code>\nFold  1 | Mean TTF : 6.7294 | MAE : 2.3739\nFold  2 | Mean TTF : 5.6401 | MAE : 1.9997\nFold  3 | Mean TTF : 5.4331 | MAE : 2.0602\nFold  4 | Mean TTF : 4.8600 | MAE : 1.9073\nFold  5 | Mean TTF : 7.2094 | MAE : 3.1402\nFold  6 | Mean TTF : 4.3884 | MAE : 1.4852\nFold  7 | Mean TTF : 6.3776 | MAE : 1.3512\nFold  8 | Mean TTF : 4.2407 | MAE : 1.0225\nFold  9 | Mean TTF : 7.2118 | MAE : 3.2313\nFold 10 | Mean TTF : 4.7368 | MAE : 1.4027\nFull    | Mean TTF : 5.6827 | MAE : 1.9975\n</code>\nAnd with shuffling.\n<code>\nFold  1 | Mean TTF : 5.6867 | MAE : 2.1739\nFold  2 | Mean TTF : 5.8734 | MAE : 2.1774\nFold  3 | Mean TTF : 5.6180 | MAE : 2.1584\nFold  4 | Mean TTF : 5.5127 | MAE : 2.0146\nFold  5 | Mean TTF : 5.7491 | MAE : 2.1442\nFold  6 | Mean TTF : 5.5263 | MAE : 1.9951\nFold  7 | Mean TTF : 5.4967 | MAE : 2.0595\nFold  8 | Mean TTF : 5.7904 | MAE : 2.0183\nFold  9 | Mean TTF : 6.0617 | MAE : 2.1592\nFold 10 | Mean TTF : 5.5121 | MAE : 2.0782\nFull    | Mean TTF : 5.6827 | MAE : 2.0979\n</code>\nMean TTF for Public is 4.017</p>",
      "rawMarkdown": "Similar experiment.\nWithout shuffling.\n```\nFold  1 | Mean TTF : 6.7294 | MAE : 2.3739\nFold  2 | Mean TTF : 5.6401 | MAE : 1.9997\nFold  3 | Mean TTF : 5.4331 | MAE : 2.0602\nFold  4 | Mean TTF : 4.8600 | MAE : 1.9073\nFold  5 | Mean TTF : 7.2094 | MAE : 3.1402\nFold  6 | Mean TTF : 4.3884 | MAE : 1.4852\nFold  7 | Mean TTF : 6.3776 | MAE : 1.3512\nFold  8 | Mean TTF : 4.2407 | MAE : 1.0225\nFold  9 | Mean TTF : 7.2118 | MAE : 3.2313\nFold 10 | Mean TTF : 4.7368 | MAE : 1.4027\nFull    | Mean TTF : 5.6827 | MAE : 1.9975\n```\nAnd with shuffling.\n```\nFold  1 | Mean TTF : 5.6867 | MAE : 2.1739\nFold  2 | Mean TTF : 5.8734 | MAE : 2.1774\nFold  3 | Mean TTF : 5.6180 | MAE : 2.1584\nFold  4 | Mean TTF : 5.5127 | MAE : 2.0146\nFold  5 | Mean TTF : 5.7491 | MAE : 2.1442\nFold  6 | Mean TTF : 5.5263 | MAE : 1.9951\nFold  7 | Mean TTF : 5.4967 | MAE : 2.0595\nFold  8 | Mean TTF : 5.7904 | MAE : 2.0183\nFold  9 | Mean TTF : 6.0617 | MAE : 2.1592\nFold 10 | Mean TTF : 5.5121 | MAE : 2.0782\nFull    | Mean TTF : 5.6827 | MAE : 2.0979\n```\nMean TTF for Public is 4.017",
      "votes": null
    },
    {
      "id": "517149",
      "postDate": "04/15/2019 16:00:29",
      "content": "<p>What about LB?</p>",
      "rawMarkdown": "What about LB?",
      "votes": null
    },
    {
      "id": "517167",
      "postDate": "04/15/2019 16:42:25",
      "content": "<p>If you shuffle, you basically validate on the training data (nearly same characteristic, distribution, highly correlated, ...). Give your model enough capacity and you'll be able to push your cv score to 0.</p>",
      "rawMarkdown": "If you shuffle, you basically validate on the training data (nearly same characteristic, distribution, highly correlated, ...). Give your model enough capacity and you'll be able to push your cv score to 0.",
      "votes": null
    },
    {
      "id": "517199",
      "postDate": "04/15/2019 17:41:30",
      "content": "<p>Thanks and good insight!</p>\n\n<p>What do you think it would happen when you change the K in the KFold? It seems that a big part of the variability of the model (when not shuffling) could come from the fact that the validation set is just <em>on the wrong place</em>.</p>",
      "rawMarkdown": "Thanks and good insight!\n\nWhat do you think it would happen when you change the K in the KFold? It seems that a big part of the variability of the model (when not shuffling) could come from the fact that the validation set is just _on the wrong place_.",
      "votes": null
    },
    {
      "id": "517291",
      "postDate": "04/15/2019 21:10:26",
      "content": "<p>As long as you don't see the earthquakes in your dreams i wouldn't worry too much :)</p>",
      "rawMarkdown": "As long as you don't see the earthquakes in your dreams i wouldn't worry too much :)",
      "votes": null
    },
    {
      "id": "517504",
      "postDate": "04/16/2019 05:42:20",
      "content": "<p><a href=\"/ricarddelgado\">@ricarddelgado</a>, higher K will give you higher variance in fold' scores but we have to remember public LB is only 13% of the test data.</p>\n\n<p>Shuffling dispatches samples from all earthquakes in all folds and you will get no idea how your models generalize on new earthquakes.</p>",
      "rawMarkdown": "ricarddelgado, higher K will give you higher variance in fold' scores but we have to remember public LB is only 13% of the test data.\n\nShuffling dispatches samples from all earthquakes in all folds and you will get no idea how your models generalize on new earthquakes.",
      "votes": null
    },
    {
      "id": "517580",
      "postDate": "04/16/2019 08:21:30",
      "content": "<p>That's a good point. Shuffling everything together won't give us a sense of how the model generalizes to new earthquakes. However, there are some earthquakes that start with very high values ttf. Would you put them in training or testing? </p>",
      "rawMarkdown": "That's a good point. Shuffling everything together won't give us a sense of how the model generalizes to new earthquakes. However, there are some earthquakes that start with very high values ttf. Would you put them in training or testing?",
      "votes": null
    },
    {
      "id": "518589",
      "postDate": "04/17/2019 13:47:43",
      "content": "<p>My thoughts exactly.</p>\n\n<p>How do you overcome this problem?</p>",
      "rawMarkdown": "My thoughts exactly.\n\nHow do you overcome this problem?",
      "votes": null
    },
    {
      "id": "519618",
      "postDate": "04/19/2019 09:45:59",
      "content": "<p>Can you show the code? I am confused what exactly you did - \"Note that I shuffle the training set inside the cross-validation loop.\" how is it inside? IMHO the leak is obvious if you split train/test from shuffled data, which I think will be the case here.</p>",
      "rawMarkdown": "Can you show the code? I am confused what exactly you did - \"Note that I shuffle the training set inside the cross-validation loop.\" how is it inside? IMHO the leak is obvious if you split train/test from shuffled data, which I think will be the case here.",
      "votes": null
    },
    {
      "id": "519850",
      "postDate": "04/19/2019 18:22:39",
      "content": "<p>I am shuffling. The best model I have scores 1.435 (shuffled 10 folds) and &gt;1.5 using leave one quake out CV weighted by quake length. So I wonder how others are doing it quake-wise.</p>",
      "rawMarkdown": "I am shuffling. The best model I have scores 1.435 (shuffled 10 folds) and &gt;1.5 using leave one quake out CV weighted by quake length. So I wonder how others are doing it quake-wise.",
      "votes": null
    },
    {
      "id": "519922",
      "postDate": "04/19/2019 20:24:00",
      "content": "<p>It seems to me that custom folds with data equally distributed in time is the way to go, but I'm still testing</p>",
      "rawMarkdown": "It seems to me that custom folds with data equally distributed in time is the way to go, but I'm still testing",
      "votes": null
    },
    {
      "id": "520857",
      "postDate": "04/21/2019 21:27:03",
      "content": "<p>My model doesn't do well on public LB for leave-oneQuake-out too. </p>",
      "rawMarkdown": "My model doesn't do well on public LB for leave-oneQuake-out too.",
      "votes": null
    },
    {
      "id": "520876",
      "postDate": "04/21/2019 23:03:18",
      "content": "<blockquote>\n  <p>Note that I shuffle the training set inside the cross-validation loop.</p>\n</blockquote>\n\n<p><a href=\"/ogrellier\">@ogrellier</a>, if you have shuffle set to True in your KFold, how are you only shuffling inside the cross validation loop? Do you mean you do shuffling twice?</p>",
      "rawMarkdown": "&gt;Note that I shuffle the training set inside the cross-validation loop.\n\n@ogrellier, if you have shuffle set to True in your KFold, how are you only shuffling inside the cross validation loop? Do you mean you do shuffling twice?",
      "votes": null
    },
    {
      "id": "520918",
      "postDate": "04/22/2019 01:39:57",
      "content": "<p>While I am still deciding whether to use validation with or without shuffle, I found the following link from SO:\n<a href=\"https://stackoverflow.com/questions/12237127/how-to-use-shuffle-in-kfold-in-scikit-learn\">https://stackoverflow.com/questions/12237127/how-to-use-shuffle-in-kfold-in-scikit-learn</a></p>\n\n<p>```\nIf shuffle is True, the whole data is first shuffled and then split into the K-Folds. For repeatable behavior, you can set the random_state, for example to an integer seed (random_state=0). If your parameters depend on the shuffling, this means your parameter selection is very unstable. Probably you have very little training data or you use to little folds (like 2 or 3).</p>\n\n<p>The \"shuffle\" is mainly useful if your data is somehow sorted by classes, because then each fold might contain only samples from one class (in particular for stochastic gradient decent classifiers sorted classes are dangerous). For other classifiers, it should make no differences. If shuffling is very unstable, your parameter selection is likely to be uninformative (aka garbage).\n```</p>\n\n<p>In our case, if \"shuffle\" is true, then there will be sample data coming from each earthquake group. Please correct me if I am wrong.</p>",
      "rawMarkdown": "While I am still deciding whether to use validation with or without shuffle, I found the following link from SO:\nhttps://stackoverflow.com/questions/12237127/how-to-use-shuffle-in-kfold-in-scikit-learn\n\n```\nIf shuffle is True, the whole data is first shuffled and then split into the K-Folds. For repeatable behavior, you can set the random_state, for example to an integer seed (random_state=0). If your parameters depend on the shuffling, this means your parameter selection is very unstable. Probably you have very little training data or you use to little folds (like 2 or 3).\n\nThe \"shuffle\" is mainly useful if your data is somehow sorted by classes, because then each fold might contain only samples from one class (in particular for stochastic gradient decent classifiers sorted classes are dangerous). For other classifiers, it should make no differences. If shuffling is very unstable, your parameter selection is likely to be uninformative (aka garbage).\n```\n\nIn our case, if \"shuffle\" is true, then there will be sample data coming from each earthquake group. Please correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "521346",
      "postDate": "04/22/2019 19:58:33",
      "content": "<p>I feel your frustration bro.</p>\n\n<p>pad pad</p>",
      "rawMarkdown": "I feel your frustration bro.\n\npad pad",
      "votes": null
    },
    {
      "id": "521353",
      "postDate": "04/22/2019 20:22:31",
      "content": "<p><a href=\"/sheriytm\">@sheriytm</a>, sorry for the misunderstanding. When I don't use shuffle in the Kfold split, i.e. <code>KFold(5, False)</code>, I shuffle the training data and keep the validation data as is. I don't think it's necessary to shuffle the training data since subsampling should take care of this.</p>",
      "rawMarkdown": "sheriytm, sorry for the misunderstanding. When I don't use shuffle in the Kfold split, i.e. `KFold(5, False)`, I shuffle the training data and keep the validation data as is. I don't think it's necessary to shuffle the training data since subsampling should take care of this.",
      "votes": null
    },
    {
      "id": "521366",
      "postDate": "04/22/2019 20:39:30",
      "content": "<p>CV by experiments is the better option because the test data is also from different experiments. Doing 10-Fold CV is similar to CV by experiments.  Your observations confirms that leakage is introduced by shuffling.  </p>",
      "rawMarkdown": "CV by experiments is the better option because the test data is also from different experiments. Doing 10-Fold CV is similar to CV by experiments.  Your observations confirms that leakage is introduced by shuffling.",
      "votes": null
    },
    {
      "id": "521371",
      "postDate": "04/22/2019 20:48:22",
      "content": "<p><a href=\"/danijelk\">@danijelk</a> what do u mean by CV by experiments?</p>",
      "rawMarkdown": "danijelk what do u mean by CV by experiments?",
      "votes": null
    },
    {
      "id": "521401",
      "postDate": "04/22/2019 21:26:28",
      "content": "<p>I think he means earthquake periods.</p>",
      "rawMarkdown": "I think he means earthquake periods.",
      "votes": null
    },
    {
      "id": "521404",
      "postDate": "04/22/2019 21:29:16",
      "content": "<p>I think that validation by earthquake group makes sense for preventing leakage.</p>",
      "rawMarkdown": "I think that validation by earthquake group makes sense for preventing leakage.",
      "votes": null
    },
    {
      "id": "521429",
      "postDate": "04/22/2019 22:37:17",
      "content": "<p>Yes by experiment I mean by earthquake period.</p>",
      "rawMarkdown": "Yes by experiment I mean by earthquake period.",
      "votes": null
    },
    {
      "id": "521606",
      "postDate": "04/23/2019 05:29:38",
      "content": "<p>Thanks <a href=\"/ogrellier\">@ogrellier</a>, I now understand that the above statement refers to the 2nd case which I also believe is the better CV scheme as you stated.</p>\n\n<p>&gt; I really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !</p>\n\n<p>How would that alert us to the LB shakeup? With the public LB being only based on 13% of the data and looking at all the intricacies of this problem I am ignoring the public LB alltogether.</p>",
      "rawMarkdown": "Thanks @ogrellier, I now understand that the above statement refers to the 2nd case which I also believe is the better CV scheme as you stated.\n\n&gt; I really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !\n\nHow would that alert us to the LB shakeup? With the public LB being only based on 13% of the data and looking at all the intricacies of this problem I am ignoring the public LB alltogether.",
      "votes": null
    },
    {
      "id": "521819",
      "postDate": "04/23/2019 13:28:10",
      "content": "<p>I love this discussion - thank you <a href=\"/ogrellier\">@ogrellier</a> ! </p>\n\n<p>I have not done so yet, but I will use the LOO (Leave One Out) technique now that I have read this discussion. In LOO, as suggested by <a href=\"/stecasasso\">@stecasasso</a> in this thread, we must isolate one of the 16 Earthquake Experiments provided in the Training Set, and use it only for validation. We then use the data from the remaining 15 experiments for training. This is repeated 16 times, to see if indeed the algorithm is working correctly. </p>\n\n<p>If it is, the final model can use all 16 for training, or can use an average of the 16 LOO results for evaluating the Test data.</p>\n\n<p>P.S. I have used LOO with success for Medical Device algorithms, especially when there are only a few volunteers compared to the total population. </p>\n\n<p>P.P.S. If medical devices are of interest to you, I made a Kernel comparing this Earthquake Warning system to a Medical Warning System, and related it to the recently released Food and Drug Administration Proposal, also discussed in the kernel.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/earthquake-pred-cnn-medical-analogy-v07\">https://www.kaggle.com/pnussbaum/earthquake-pred-cnn-medical-analogy-v07</a> </p>",
      "rawMarkdown": "I love this discussion - thank you @ogrellier ! \n\nI have not done so yet, but I will use the LOO (Leave One Out) technique now that I have read this discussion. In LOO, as suggested by @stecasasso in this thread, we must isolate one of the 16 Earthquake Experiments provided in the Training Set, and use it only for validation. We then use the data from the remaining 15 experiments for training. This is repeated 16 times, to see if indeed the algorithm is working correctly. \n\nIf it is, the final model can use all 16 for training, or can use an average of the 16 LOO results for evaluating the Test data.\n\nP.S. I have used LOO with success for Medical Device algorithms, especially when there are only a few volunteers compared to the total population. \n\nP.P.S. If medical devices are of interest to you, I made a Kernel comparing this Earthquake Warning system to a Medical Warning System, and related it to the recently released Food and Drug Administration Proposal, also discussed in the kernel.\n\nhttps://www.kaggle.com/pnussbaum/earthquake-pred-cnn-medical-analogy-v07",
      "votes": null
    },
    {
      "id": "525899",
      "postDate": "05/02/2019 00:52:21",
      "content": "<p>The reason is that some cycles are tougher, because in the middle of them there is some release of energy but not enough to generate a lab earthquake, so the error of the predicion is higher in those folds. If you shuffle you have some of those tougher cycles in all folds so you get a similar score.</p>",
      "rawMarkdown": "The reason is that some cycles are tougher, because in the middle of them there is some release of energy but not enough to generate a lab earthquake, so the error of the predicion is higher in those folds. If you shuffle you have some of those tougher cycles in all folds so you get a similar score.",
      "votes": null
    },
    {
      "id": "529519",
      "postDate": "05/10/2019 05:06:37",
      "content": "<p>The shuffling is the right way to predict?</p>",
      "rawMarkdown": "The shuffling is the right way to predict?",
      "votes": null
    },
    {
      "id": "529742",
      "postDate": "05/10/2019 16:13:26",
      "content": "<p>I explained how I understand this differences between CV when shuffling or not and LB in a new topic: \n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91726\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91726</a></p>\n\n<p>Please let me know if it make sense to you or if I am missing something. Thanks</p>",
      "rawMarkdown": "I explained how I understand this differences between CV when shuffling or not and LB in a new topic: \nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91726\n\nPlease let me know if it make sense to you or if I am missing something. Thanks",
      "votes": null
    },
    {
      "id": "530430",
      "postDate": "05/12/2019 19:43:40",
      "content": "<p>Have you tried LOO in this competition and is it helpful for you?</p>",
      "rawMarkdown": "Have you tried LOO in this competition and is it helpful for you?",
      "votes": null
    },
    {
      "id": "530865",
      "postDate": "05/13/2019 21:01:10",
      "content": "<p>I have, and I continue to use it as I refine my model. There is an issue of cross-experiment boundaries. You can either fill in those segments with dummy data, or skip those segments altogether. Here is a snippet of code to save you some time... Let me know if anything is confusing...</p>\n\n<pre><code># find start and endpoints of the sixteen (16) Earthquake Trials\n#\nstart = np.zeros(17, dtype=np.int32)\nend = np.zeros(17, dtype=np.int32)\nindex = 0\nfor i in tqdm(range (0,y_raw.shape[0] - rows, rows)) :\n    if i == 0:\n       start[index] = i;\n       end[index] = i;\n    if y_raw[i+rows] &amp;gt; y_raw[i] :\n        end[index] = i\n        index += 1\n       start[index] = i\n        end[index] = i + rows\n        # \"Clean Up\" the segment that traverses the experiment boundary\n        # This will be the beginning segment of a new experiment.\n        boundary = 0\n        for j in range(rows - 1) :\n            if y_raw[i + j + 1] &amp;gt; y_raw[i + j] :\n                boundary = j\n        # \"Clean up\" for now means zero out seismic data.\n        X_train[i:i+boundary] = 0\n# Now let's see how big each experiment is, in terms of the maximum number of 150000 sample rows \nrunning_count = 0\nfor i in range (16) :\n    count = np.int32((end[i] - start[i])/rows)\n    print (\"Experiment\", i, \"has\", count, \"segments of 150000 samples each\")\n    running_count += count\n    start[i] = np.int32((start[i])/rows)\n    end[i] = np.int32((end[i])/rows)\n\n# The total number of segments is\nsegments = np.int32(end[15])\n</code></pre>",
      "rawMarkdown": "I have, and I continue to use it as I refine my model. There is an issue of cross-experiment boundaries. You can either fill in those segments with dummy data, or skip those segments altogether. Here is a snippet of code to save you some time... Let me know if anything is confusing...\n\n    # find start and endpoints of the sixteen (16) Earthquake Trials\n    #\n    start = np.zeros(17, dtype=np.int32)\n    end = np.zeros(17, dtype=np.int32)\n    index = 0\n    for i in tqdm(range (0,y_raw.shape[0] - rows, rows)) :\n        if i == 0:\n           start[index] = i;\n           end[index] = i;\n        if y_raw[i+rows] &gt; y_raw[i] :\n            end[index] = i\n            index += 1\n           start[index] = i\n            end[index] = i + rows\n            # \"Clean Up\" the segment that traverses the experiment boundary\n            # This will be the beginning segment of a new experiment.\n            boundary = 0\n            for j in range(rows - 1) :\n                if y_raw[i + j + 1] &gt; y_raw[i + j] :\n                    boundary = j\n            # \"Clean up\" for now means zero out seismic data.\n            X_train[i:i+boundary] = 0\n    # Now let's see how big each experiment is, in terms of the maximum number of 150000 sample rows \n    running_count = 0\n    for i in range (16) :\n        count = np.int32((end[i] - start[i])/rows)\n        print (\"Experiment\", i, \"has\", count, \"segments of 150000 samples each\")\n        running_count += count\n        start[i] = np.int32((start[i])/rows)\n        end[i] = np.int32((end[i])/rows)\n\n    # The total number of segments is\n    segments = np.int32(end[15])",
      "votes": null
    },
    {
      "id": "530984",
      "postDate": "05/14/2019 04:54:56",
      "content": "<p>Paul, your code is hard to read.  Rather than using bulleted list, just add 4 spaces at teh start of each code line.  Code test will be formatted as code that way.</p>",
      "rawMarkdown": "Paul, your code is hard to read.  Rather than using bulleted list, just add 4 spaces at teh start of each code line.  Code test will be formatted as code that way.",
      "votes": null
    },
    {
      "id": "531147",
      "postDate": "05/14/2019 11:41:26",
      "content": "<p>Thank you! I couldn't find that formatting trick when I looked for it - appreciate your help <a href=\"/cpmpml\">@cpmpml</a> </p>",
      "rawMarkdown": "Thank you! I couldn't find that formatting trick when I looked for it - appreciate your help @cpmpml",
      "votes": null
    },
    {
      "id": "540084",
      "postDate": "05/31/2019 02:24:48",
      "content": "<p>Apologies for the late posting. </p>\n\n<p>I finished a Kernel for you to look over. </p>\n\n<p>This Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.</p>\n\n<p>This Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)</p>\n\n<p>Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01</a></p>\n\n<p>Enjoy, and good luck!</p>",
      "rawMarkdown": "Apologies for the late posting. \n\nI finished a Kernel for you to look over. \n\nThis Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.\n\nThis Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)\n\nGets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\n\nEnjoy, and good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 515963,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "04/13/2019 12:33:59",
      "content": "<p>Interesting, Olivier! \nAm I right that you killed all time related information within folds and get better score? </p>",
      "votes": null,
      "replies": [
        {
          "id": 516035,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/13/2019 14:50:13",
          "content": "<p><a href=\"/alexfir\">@alexfir</a>, yes if you shuffle (remove time) then you get a better CV and a better LB.</p>\n\n<p>However I would assume that taking 13% of the private LB as a validation is fairly risky  ;-)</p>\n\n<p>I don't know if it's about time or about the fact that if 2 values (std, max, min or anything else) are close then the samples must be close as well and therefore finding the *time_to_failure* is easy...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 515976,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "04/13/2019 13:11:07",
      "content": "<p>!!!  Maybe we can find some leaks here. </p>",
      "votes": null,
      "replies": [
        {
          "id": 516329,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/14/2019 02:28:50",
          "content": "<p>I counted 16 quakes, no improvement yet...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 515977,
      "author_name": "dkaraflos",
      "author_url": "",
      "post_date": "04/13/2019 13:15:49",
      "content": "<p>the most important thing in this competition is the fact that the test set segments are not continous as are in the train set. Hence, i believe that validating with suffling may lead us to better LB scores</p>",
      "votes": null,
      "replies": [
        {
          "id": 516040,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/13/2019 14:54:27",
          "content": "<p><a href=\"/dkaraflos\">@dkaraflos</a>, you may be right. </p>\n\n<p>I just wanted to draw people's attention on the fact that the only way to reproduce public LB score is to not shuffle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516092,
          "author_name": "greenwing1985",
          "author_url": "",
          "post_date": "04/13/2019 15:47:41",
          "content": "<p>the test set might not be continues, but they are part of a continues signal</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516102,
          "author_name": "dkaraflos",
          "author_url": "",
          "post_date": "04/13/2019 16:02:10",
          "content": "<p>exactly. That is why the FE should treat each segment individually </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 515998,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "04/13/2019 14:05:36",
      "content": "<p>Early stopping would be problematic without shuffling, just my thought</p>",
      "votes": null,
      "replies": [
        {
          "id": 518589,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "04/17/2019 13:47:43",
          "content": "<p>My thoughts exactly.</p>\n\n<p>How do you overcome this problem?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519850,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "04/19/2019 18:22:39",
          "content": "<p>I am shuffling. The best model I have scores 1.435 (shuffled 10 folds) and &gt;1.5 using leave one quake out CV weighted by quake length. So I wonder how others are doing it quake-wise.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 520857,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/21/2019 21:27:03",
          "content": "<p>My model doesn't do well on public LB for leave-oneQuake-out too. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 516222,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "04/13/2019 19:54:59",
      "content": "<p>I have the feeling the  leakage may occur when segments from the same quake are in both training and validation. This happens a lot with shuffling and less so without shuffling. \nI have used stratified shuffle sampling for a while and then turned to quake-wise split for this reason. </p>",
      "votes": null,
      "replies": [
        {
          "id": 516243,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/13/2019 20:45:29",
          "content": "<p>Thanks <a href=\"/stecasasso\">@stecasasso</a> for your comments. I totally agree with that ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516254,
          "author_name": "dkaraflos",
          "author_url": "",
          "post_date": "04/13/2019 21:22:52",
          "content": "<p>I am new to this competition, so i am making my baseline models and i started my EDA.\nSO, in this competition, each segment is a part of an earthquake, but there are many sperate earthquakes? then the validation setup must ensure that we must shuffle the segments and also ensure that there are no segments of the same earthquake between training data and validation data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516327,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/14/2019 02:27:40",
          "content": "<p>In the training set there are 16 places where ttf reaches zero (actually it never reaches zero, but close enough) and resets, these 16 resets of the countdown are taken as the boundaries between the quakes. </p>\n\n<p>The starter kernel and most others since then, breaks the continuous 600 million row csv into non-overlapping segments of 150k rows starting from row zero, and computes whatever features per segment for these ~4100 segments. This seems like a natural way to treat the training data as the test data consists of ~2k individual segments of 150k rows of acoustic data and it is from these that we are to predict the time to failure, in seconds, from the acoustic signal. Doing it this way means 16 of the segments will have a ttf reset somewhere in their rows.</p>\n\n<p>We know that the test data supposedly came from the same experiment that generated the training data, we do not know if the testing segments also contain places where ttf resets but it seems reasonable to me to assume so, to assume there are somewhere around 8 resets in test, and that what this means is quite unclear. </p>\n\n<p>The validation problem here is really interesting, many have suggested breaking up the quakes and using the rows from one or more quakes as validation for the others. The LB seems to be almost entirely useless here and strongly biased towards low-ttf quakes which is easy to check, you can make a submission with some cv score, then make one with a worse cv but a lower mean predicted ttf, the one with lower mean will usually score better. I don't know the best way to do it yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 516486,
          "author_name": "dkaraflos",
          "author_url": "",
          "post_date": "04/14/2019 09:20:59",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> you are right about the earthquake numbers, but what puzzles me is the fact that the competition host did not mention such a data property.\nThus, in this competition, the validation setup is a really crucial step. We must think of it very robustly</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 516400,
      "author_name": "timmmmmms",
      "author_url": "",
      "post_date": "04/14/2019 05:38:19",
      "content": "<p>What is ttf?</p>",
      "votes": null,
      "replies": [
        {
          "id": 516410,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/14/2019 06:05:41",
          "content": "<p><a href=\"/timmmmmms\">@timmmmmms</a>, I guess this stands for Time To Failure ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 516402,
      "author_name": "jsaguiar",
      "author_url": "",
      "post_date": "04/14/2019 05:49:02",
      "content": "<p>I think that the huge fluctuation when using kfold without shuffle is because the model can't predict high TTFs. Each fold has 1~2 earthquakes for validation; if this quake has high TTF the score is bad (e.g. fold 5 is probably the earthquake with 16s). On the other hand, if the quake has around 10s the score will be quite good on that fold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 516620,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/14/2019 15:08:42",
          "content": "<p>I've spent too long on this data... I can literally see the two earthquakes that gave him MAE &gt; 3s</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517291,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "04/15/2019 21:10:26",
          "content": "<p>As long as you don't see the earthquakes in your dreams i wouldn't worry too much :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 516997,
      "author_name": "mykper",
      "author_url": "",
      "post_date": "04/15/2019 10:32:40",
      "content": "<p>Similar experiment.\nWithout shuffling.\n<code>\nFold  1 | Mean TTF : 6.7294 | MAE : 2.3739\nFold  2 | Mean TTF : 5.6401 | MAE : 1.9997\nFold  3 | Mean TTF : 5.4331 | MAE : 2.0602\nFold  4 | Mean TTF : 4.8600 | MAE : 1.9073\nFold  5 | Mean TTF : 7.2094 | MAE : 3.1402\nFold  6 | Mean TTF : 4.3884 | MAE : 1.4852\nFold  7 | Mean TTF : 6.3776 | MAE : 1.3512\nFold  8 | Mean TTF : 4.2407 | MAE : 1.0225\nFold  9 | Mean TTF : 7.2118 | MAE : 3.2313\nFold 10 | Mean TTF : 4.7368 | MAE : 1.4027\nFull    | Mean TTF : 5.6827 | MAE : 1.9975\n</code>\nAnd with shuffling.\n<code>\nFold  1 | Mean TTF : 5.6867 | MAE : 2.1739\nFold  2 | Mean TTF : 5.8734 | MAE : 2.1774\nFold  3 | Mean TTF : 5.6180 | MAE : 2.1584\nFold  4 | Mean TTF : 5.5127 | MAE : 2.0146\nFold  5 | Mean TTF : 5.7491 | MAE : 2.1442\nFold  6 | Mean TTF : 5.5263 | MAE : 1.9951\nFold  7 | Mean TTF : 5.4967 | MAE : 2.0595\nFold  8 | Mean TTF : 5.7904 | MAE : 2.0183\nFold  9 | Mean TTF : 6.0617 | MAE : 2.1592\nFold 10 | Mean TTF : 5.5121 | MAE : 2.0782\nFull    | Mean TTF : 5.6827 | MAE : 2.0979\n</code>\nMean TTF for Public is 4.017</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517149,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/15/2019 16:00:29",
      "content": "<p>What about LB?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517167,
      "author_name": "svenhinderer",
      "author_url": "",
      "post_date": "04/15/2019 16:42:25",
      "content": "<p>If you shuffle, you basically validate on the training data (nearly same characteristic, distribution, highly correlated, ...). Give your model enough capacity and you'll be able to push your cv score to 0.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517199,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "04/15/2019 17:41:30",
      "content": "<p>Thanks and good insight!</p>\n\n<p>What do you think it would happen when you change the K in the KFold? It seems that a big part of the variability of the model (when not shuffling) could come from the fact that the validation set is just <em>on the wrong place</em>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 517504,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/16/2019 05:42:20",
          "content": "<p><a href=\"/ricarddelgado\">@ricarddelgado</a>, higher K will give you higher variance in fold' scores but we have to remember public LB is only 13% of the test data.</p>\n\n<p>Shuffling dispatches samples from all earthquakes in all folds and you will get no idea how your models generalize on new earthquakes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517580,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/16/2019 08:21:30",
          "content": "<p>That's a good point. Shuffling everything together won't give us a sense of how the model generalizes to new earthquakes. However, there are some earthquakes that start with very high values ttf. Would you put them in training or testing? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 519618,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "04/19/2019 09:45:59",
      "content": "<p>Can you show the code? I am confused what exactly you did - \"Note that I shuffle the training set inside the cross-validation loop.\" how is it inside? IMHO the leak is obvious if you split train/test from shuffled data, which I think will be the case here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 519922,
      "author_name": "daijin12",
      "author_url": "",
      "post_date": "04/19/2019 20:24:00",
      "content": "<p>It seems to me that custom folds with data equally distributed in time is the way to go, but I'm still testing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 520876,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "04/21/2019 23:03:18",
      "content": "<blockquote>\n  <p>Note that I shuffle the training set inside the cross-validation loop.</p>\n</blockquote>\n\n<p><a href=\"/ogrellier\">@ogrellier</a>, if you have shuffle set to True in your KFold, how are you only shuffling inside the cross validation loop? Do you mean you do shuffling twice?</p>",
      "votes": null,
      "replies": [
        {
          "id": 521353,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "04/22/2019 20:22:31",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a>, sorry for the misunderstanding. When I don't use shuffle in the Kfold split, i.e. <code>KFold(5, False)</code>, I shuffle the training data and keep the validation data as is. I don't think it's necessary to shuffle the training data since subsampling should take care of this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521606,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "04/23/2019 05:29:38",
          "content": "<p>Thanks <a href=\"/ogrellier\">@ogrellier</a>, I now understand that the above statement refers to the 2nd case which I also believe is the better CV scheme as you stated.</p>\n\n<p>&gt; I really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !</p>\n\n<p>How would that alert us to the LB shakeup? With the public LB being only based on 13% of the data and looking at all the intricacies of this problem I am ignoring the public LB alltogether.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 520918,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "04/22/2019 01:39:57",
      "content": "<p>While I am still deciding whether to use validation with or without shuffle, I found the following link from SO:\n<a href=\"https://stackoverflow.com/questions/12237127/how-to-use-shuffle-in-kfold-in-scikit-learn\">https://stackoverflow.com/questions/12237127/how-to-use-shuffle-in-kfold-in-scikit-learn</a></p>\n\n<p>```\nIf shuffle is True, the whole data is first shuffled and then split into the K-Folds. For repeatable behavior, you can set the random_state, for example to an integer seed (random_state=0). If your parameters depend on the shuffling, this means your parameter selection is very unstable. Probably you have very little training data or you use to little folds (like 2 or 3).</p>\n\n<p>The \"shuffle\" is mainly useful if your data is somehow sorted by classes, because then each fold might contain only samples from one class (in particular for stochastic gradient decent classifiers sorted classes are dangerous). For other classifiers, it should make no differences. If shuffling is very unstable, your parameter selection is likely to be uninformative (aka garbage).\n```</p>\n\n<p>In our case, if \"shuffle\" is true, then there will be sample data coming from each earthquake group. Please correct me if I am wrong.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 521346,
      "author_name": "",
      "author_url": "",
      "post_date": "04/22/2019 19:58:33",
      "content": "<p>I feel your frustration bro.</p>\n\n<p>pad pad</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 521366,
      "author_name": "danijelk",
      "author_url": "",
      "post_date": "04/22/2019 20:39:30",
      "content": "<p>CV by experiments is the better option because the test data is also from different experiments. Doing 10-Fold CV is similar to CV by experiments.  Your observations confirms that leakage is introduced by shuffling.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 521371,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/22/2019 20:48:22",
          "content": "<p><a href=\"/danijelk\">@danijelk</a> what do u mean by CV by experiments?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521401,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "04/22/2019 21:26:28",
          "content": "<p>I think he means earthquake periods.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521404,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/22/2019 21:29:16",
          "content": "<p>I think that validation by earthquake group makes sense for preventing leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 521429,
          "author_name": "danijelk",
          "author_url": "",
          "post_date": "04/22/2019 22:37:17",
          "content": "<p>Yes by experiment I mean by earthquake period.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 521819,
      "author_name": "pnussbaum",
      "author_url": "",
      "post_date": "04/23/2019 13:28:10",
      "content": "<p>I love this discussion - thank you <a href=\"/ogrellier\">@ogrellier</a> ! </p>\n\n<p>I have not done so yet, but I will use the LOO (Leave One Out) technique now that I have read this discussion. In LOO, as suggested by <a href=\"/stecasasso\">@stecasasso</a> in this thread, we must isolate one of the 16 Earthquake Experiments provided in the Training Set, and use it only for validation. We then use the data from the remaining 15 experiments for training. This is repeated 16 times, to see if indeed the algorithm is working correctly. </p>\n\n<p>If it is, the final model can use all 16 for training, or can use an average of the 16 LOO results for evaluating the Test data.</p>\n\n<p>P.S. I have used LOO with success for Medical Device algorithms, especially when there are only a few volunteers compared to the total population. </p>\n\n<p>P.P.S. If medical devices are of interest to you, I made a Kernel comparing this Earthquake Warning system to a Medical Warning System, and related it to the recently released Food and Drug Administration Proposal, also discussed in the kernel.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/earthquake-pred-cnn-medical-analogy-v07\">https://www.kaggle.com/pnussbaum/earthquake-pred-cnn-medical-analogy-v07</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 530430,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "05/12/2019 19:43:40",
          "content": "<p>Have you tried LOO in this competition and is it helpful for you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530865,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "05/13/2019 21:01:10",
          "content": "<p>I have, and I continue to use it as I refine my model. There is an issue of cross-experiment boundaries. You can either fill in those segments with dummy data, or skip those segments altogether. Here is a snippet of code to save you some time... Let me know if anything is confusing...</p>\n\n<pre><code># find start and endpoints of the sixteen (16) Earthquake Trials\n#\nstart = np.zeros(17, dtype=np.int32)\nend = np.zeros(17, dtype=np.int32)\nindex = 0\nfor i in tqdm(range (0,y_raw.shape[0] - rows, rows)) :\n    if i == 0:\n       start[index] = i;\n       end[index] = i;\n    if y_raw[i+rows] &amp;gt; y_raw[i] :\n        end[index] = i\n        index += 1\n       start[index] = i\n        end[index] = i + rows\n        # \"Clean Up\" the segment that traverses the experiment boundary\n        # This will be the beginning segment of a new experiment.\n        boundary = 0\n        for j in range(rows - 1) :\n            if y_raw[i + j + 1] &amp;gt; y_raw[i + j] :\n                boundary = j\n        # \"Clean up\" for now means zero out seismic data.\n        X_train[i:i+boundary] = 0\n# Now let's see how big each experiment is, in terms of the maximum number of 150000 sample rows \nrunning_count = 0\nfor i in range (16) :\n    count = np.int32((end[i] - start[i])/rows)\n    print (\"Experiment\", i, \"has\", count, \"segments of 150000 samples each\")\n    running_count += count\n    start[i] = np.int32((start[i])/rows)\n    end[i] = np.int32((end[i])/rows)\n\n# The total number of segments is\nsegments = np.int32(end[15])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 530984,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/14/2019 04:54:56",
          "content": "<p>Paul, your code is hard to read.  Rather than using bulleted list, just add 4 spaces at teh start of each code line.  Code test will be formatted as code that way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 531147,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "05/14/2019 11:41:26",
          "content": "<p>Thank you! I couldn't find that formatting trick when I looked for it - appreciate your help <a href=\"/cpmpml\">@cpmpml</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540084,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "05/31/2019 02:24:48",
          "content": "<p>Apologies for the late posting. </p>\n\n<p>I finished a Kernel for you to look over. </p>\n\n<p>This Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.</p>\n\n<p>This Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)</p>\n\n<p>Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01</a></p>\n\n<p>Enjoy, and good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 525899,
      "author_name": "carlospk",
      "author_url": "",
      "post_date": "05/02/2019 00:52:21",
      "content": "<p>The reason is that some cycles are tougher, because in the middle of them there is some release of energy but not enough to generate a lab earthquake, so the error of the predicion is higher in those folds. If you shuffle you have some of those tougher cycles in all folds so you get a similar score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529519,
      "author_name": "timmmmmms",
      "author_url": "",
      "post_date": "05/10/2019 05:06:37",
      "content": "<p>The shuffling is the right way to predict?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 529742,
      "author_name": "carlospk",
      "author_url": "",
      "post_date": "05/10/2019 16:13:26",
      "content": "<p>I explained how I understand this differences between CV when shuffling or not and LB in a new topic: \n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91726\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91726</a></p>\n\n<p>Please let me know if it make sense to you or if I am missing something. Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "515946": "I've been wondering what made the difference between local CV and LB and decided to conduct a few experiments...\n\nHere is something I find particularly interesting, using almost the same features as [Inversion](https://www.kaggle.com/inversion/basic-feature-benchmark)\n\nUsing 10-fold cross-validation with shuffling, i.e. `KFold(10, True, 1)`:\n```\nFold 1 MAE : 2.2725\nFold 2 MAE : 2.1927\nFold 3 MAE : 2.1419\nFold 4 MAE : 2.2639\nFold 5 MAE : 2.2061\nFold 6 MAE : 2.1068\nFold 7 MAE : 2.1524\nFold 8 MAE : 2.1555\nFold 9 MAE : 2.2781\nFold 10 MAE : 2.2490\nFull MAE : 2.2019\n```\n\nNow without shuffling, i.e. `KFold(10, False, 1)`\n```\nFold 1 MAE : 2.5836\nFold 2 MAE : 1.8352\nFold 3 MAE : 2.1289\nFold 4 MAE : 2.9439\nFold 5 MAE : 3.1463\nFold 6 MAE : 2.0426\nFold 7 MAE : 1.5974\nFold 8 MAE : 1.5868\nFold 9 MAE : 3.3176\nFold 10 MAE : 2.1363\nFull MAE : 2.3317\n```\n\nNote that I shuffle the training set inside the cross-validation loop.\n\nMy bet is that there is sufficient correlation between training samples, which may lead to leakage between training and validation sets if you shuffle samples before cross-validation.\n\nI really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !\n\nWhat do you think ?",
    "515963": "Interesting, Olivier! \nAm I right that you killed all time related information within folds and get better score?",
    "515976": "!!!  Maybe we can find some leaks here.",
    "515977": "the most important thing in this competition is the fact that the test set segments are not continous as are in the train set. Hence, i believe that validating with suffling may lead us to better LB scores",
    "515998": "Early stopping would be problematic without shuffling, just my thought",
    "516035": "alexfir, yes if you shuffle (remove time) then you get a better CV and a better LB.\n\nHowever I would assume that taking 13% of the private LB as a validation is fairly risky  ;-)\n\nI don't know if it's about time or about the fact that if 2 values (std, max, min or anything else) are close then the samples must be close as well and therefore finding the *time_to_failure* is easy...",
    "516040": "dkaraflos, you may be right. \n\nI just wanted to draw people's attention on the fact that the only way to reproduce public LB score is to not shuffle.",
    "516092": "the test set might not be continues, but they are part of a continues signal",
    "516102": "exactly. That is why the FE should treat each segment individually",
    "516222": "I have the feeling the  leakage may occur when segments from the same quake are in both training and validation. This happens a lot with shuffling and less so without shuffling. \nI have used stratified shuffle sampling for a while and then turned to quake-wise split for this reason.",
    "516243": "Thanks @stecasasso for your comments. I totally agree with that ;-)",
    "516254": "I am new to this competition, so i am making my baseline models and i started my EDA.\nSO, in this competition, each segment is a part of an earthquake, but there are many sperate earthquakes? then the validation setup must ensure that we must shuffle the segments and also ensure that there are no segments of the same earthquake between training data and validation data",
    "516327": "In the training set there are 16 places where ttf reaches zero (actually it never reaches zero, but close enough) and resets, these 16 resets of the countdown are taken as the boundaries between the quakes. \n\nThe starter kernel and most others since then, breaks the continuous 600 million row csv into non-overlapping segments of 150k rows starting from row zero, and computes whatever features per segment for these ~4100 segments. This seems like a natural way to treat the training data as the test data consists of ~2k individual segments of 150k rows of acoustic data and it is from these that we are to predict the time to failure, in seconds, from the acoustic signal. Doing it this way means 16 of the segments will have a ttf reset somewhere in their rows.\n\nWe know that the test data supposedly came from the same experiment that generated the training data, we do not know if the testing segments also contain places where ttf resets but it seems reasonable to me to assume so, to assume there are somewhere around 8 resets in test, and that what this means is quite unclear. \n\nThe validation problem here is really interesting, many have suggested breaking up the quakes and using the rows from one or more quakes as validation for the others. The LB seems to be almost entirely useless here and strongly biased towards low-ttf quakes which is easy to check, you can make a submission with some cv score, then make one with a worse cv but a lower mean predicted ttf, the one with lower mean will usually score better. I don't know the best way to do it yet.",
    "516329": "I counted 16 quakes, no improvement yet...",
    "516400": "What is ttf?",
    "516402": "I think that the huge fluctuation when using kfold without shuffle is because the model can't predict high TTFs. Each fold has 1~2 earthquakes for validation; if this quake has high TTF the score is bad (e.g. fold 5 is probably the earthquake with 16s). On the other hand, if the quake has around 10s the score will be quite good on that fold.",
    "516410": "timmmmmms, I guess this stands for Time To Failure ;-)",
    "516486": "interneuron you are right about the earthquake numbers, but what puzzles me is the fact that the competition host did not mention such a data property.\nThus, in this competition, the validation setup is a really crucial step. We must think of it very robustly",
    "516620": "I've spent too long on this data... I can literally see the two earthquakes that gave him MAE &gt; 3s",
    "516997": "Similar experiment.\nWithout shuffling.\n```\nFold  1 | Mean TTF : 6.7294 | MAE : 2.3739\nFold  2 | Mean TTF : 5.6401 | MAE : 1.9997\nFold  3 | Mean TTF : 5.4331 | MAE : 2.0602\nFold  4 | Mean TTF : 4.8600 | MAE : 1.9073\nFold  5 | Mean TTF : 7.2094 | MAE : 3.1402\nFold  6 | Mean TTF : 4.3884 | MAE : 1.4852\nFold  7 | Mean TTF : 6.3776 | MAE : 1.3512\nFold  8 | Mean TTF : 4.2407 | MAE : 1.0225\nFold  9 | Mean TTF : 7.2118 | MAE : 3.2313\nFold 10 | Mean TTF : 4.7368 | MAE : 1.4027\nFull    | Mean TTF : 5.6827 | MAE : 1.9975\n```\nAnd with shuffling.\n```\nFold  1 | Mean TTF : 5.6867 | MAE : 2.1739\nFold  2 | Mean TTF : 5.8734 | MAE : 2.1774\nFold  3 | Mean TTF : 5.6180 | MAE : 2.1584\nFold  4 | Mean TTF : 5.5127 | MAE : 2.0146\nFold  5 | Mean TTF : 5.7491 | MAE : 2.1442\nFold  6 | Mean TTF : 5.5263 | MAE : 1.9951\nFold  7 | Mean TTF : 5.4967 | MAE : 2.0595\nFold  8 | Mean TTF : 5.7904 | MAE : 2.0183\nFold  9 | Mean TTF : 6.0617 | MAE : 2.1592\nFold 10 | Mean TTF : 5.5121 | MAE : 2.0782\nFull    | Mean TTF : 5.6827 | MAE : 2.0979\n```\nMean TTF for Public is 4.017",
    "517149": "What about LB?",
    "517167": "If you shuffle, you basically validate on the training data (nearly same characteristic, distribution, highly correlated, ...). Give your model enough capacity and you'll be able to push your cv score to 0.",
    "517199": "Thanks and good insight!\n\nWhat do you think it would happen when you change the K in the KFold? It seems that a big part of the variability of the model (when not shuffling) could come from the fact that the validation set is just _on the wrong place_.",
    "517291": "As long as you don't see the earthquakes in your dreams i wouldn't worry too much :)",
    "517504": "ricarddelgado, higher K will give you higher variance in fold' scores but we have to remember public LB is only 13% of the test data.\n\nShuffling dispatches samples from all earthquakes in all folds and you will get no idea how your models generalize on new earthquakes.",
    "517580": "That's a good point. Shuffling everything together won't give us a sense of how the model generalizes to new earthquakes. However, there are some earthquakes that start with very high values ttf. Would you put them in training or testing?",
    "518589": "My thoughts exactly.\n\nHow do you overcome this problem?",
    "519618": "Can you show the code? I am confused what exactly you did - \"Note that I shuffle the training set inside the cross-validation loop.\" how is it inside? IMHO the leak is obvious if you split train/test from shuffled data, which I think will be the case here.",
    "519850": "I am shuffling. The best model I have scores 1.435 (shuffled 10 folds) and &gt;1.5 using leave one quake out CV weighted by quake length. So I wonder how others are doing it quake-wise.",
    "519922": "It seems to me that custom folds with data equally distributed in time is the way to go, but I'm still testing",
    "520857": "My model doesn't do well on public LB for leave-oneQuake-out too.",
    "520876": "&gt;Note that I shuffle the training set inside the cross-validation loop.\n\n@ogrellier, if you have shuffle set to True in your KFold, how are you only shuffling inside the cross validation loop? Do you mean you do shuffling twice?",
    "520918": "While I am still deciding whether to use validation with or without shuffle, I found the following link from SO:\nhttps://stackoverflow.com/questions/12237127/how-to-use-shuffle-in-kfold-in-scikit-learn\n\n```\nIf shuffle is True, the whole data is first shuffled and then split into the K-Folds. For repeatable behavior, you can set the random_state, for example to an integer seed (random_state=0). If your parameters depend on the shuffling, this means your parameter selection is very unstable. Probably you have very little training data or you use to little folds (like 2 or 3).\n\nThe \"shuffle\" is mainly useful if your data is somehow sorted by classes, because then each fold might contain only samples from one class (in particular for stochastic gradient decent classifiers sorted classes are dangerous). For other classifiers, it should make no differences. If shuffling is very unstable, your parameter selection is likely to be uninformative (aka garbage).\n```\n\nIn our case, if \"shuffle\" is true, then there will be sample data coming from each earthquake group. Please correct me if I am wrong.",
    "521346": "I feel your frustration bro.\n\npad pad",
    "521353": "sheriytm, sorry for the misunderstanding. When I don't use shuffle in the Kfold split, i.e. `KFold(5, False)`, I shuffle the training data and keep the validation data as is. I don't think it's necessary to shuffle the training data since subsampling should take care of this.",
    "521366": "CV by experiments is the better option because the test data is also from different experiments. Doing 10-Fold CV is similar to CV by experiments.  Your observations confirms that leakage is introduced by shuffling.",
    "521371": "danijelk what do u mean by CV by experiments?",
    "521401": "I think he means earthquake periods.",
    "521404": "I think that validation by earthquake group makes sense for preventing leakage.",
    "521429": "Yes by experiment I mean by earthquake period.",
    "521606": "Thanks @ogrellier, I now understand that the above statement refers to the 2nd case which I also believe is the better CV scheme as you stated.\n\n&gt; I really prefer the 2nd cross validation setup, although it may alert us on the coming LB earthquake !\n\nHow would that alert us to the LB shakeup? With the public LB being only based on 13% of the data and looking at all the intricacies of this problem I am ignoring the public LB alltogether.",
    "521819": "I love this discussion - thank you @ogrellier ! \n\nI have not done so yet, but I will use the LOO (Leave One Out) technique now that I have read this discussion. In LOO, as suggested by @stecasasso in this thread, we must isolate one of the 16 Earthquake Experiments provided in the Training Set, and use it only for validation. We then use the data from the remaining 15 experiments for training. This is repeated 16 times, to see if indeed the algorithm is working correctly. \n\nIf it is, the final model can use all 16 for training, or can use an average of the 16 LOO results for evaluating the Test data.\n\nP.S. I have used LOO with success for Medical Device algorithms, especially when there are only a few volunteers compared to the total population. \n\nP.P.S. If medical devices are of interest to you, I made a Kernel comparing this Earthquake Warning system to a Medical Warning System, and related it to the recently released Food and Drug Administration Proposal, also discussed in the kernel.\n\nhttps://www.kaggle.com/pnussbaum/earthquake-pred-cnn-medical-analogy-v07",
    "525899": "The reason is that some cycles are tougher, because in the middle of them there is some release of energy but not enough to generate a lab earthquake, so the error of the predicion is higher in those folds. If you shuffle you have some of those tougher cycles in all folds so you get a similar score.",
    "529519": "The shuffling is the right way to predict?",
    "529742": "I explained how I understand this differences between CV when shuffling or not and LB in a new topic: \nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91726\n\nPlease let me know if it make sense to you or if I am missing something. Thanks",
    "530430": "Have you tried LOO in this competition and is it helpful for you?",
    "530865": "I have, and I continue to use it as I refine my model. There is an issue of cross-experiment boundaries. You can either fill in those segments with dummy data, or skip those segments altogether. Here is a snippet of code to save you some time... Let me know if anything is confusing...\n\n    # find start and endpoints of the sixteen (16) Earthquake Trials\n    #\n    start = np.zeros(17, dtype=np.int32)\n    end = np.zeros(17, dtype=np.int32)\n    index = 0\n    for i in tqdm(range (0,y_raw.shape[0] - rows, rows)) :\n        if i == 0:\n           start[index] = i;\n           end[index] = i;\n        if y_raw[i+rows] &gt; y_raw[i] :\n            end[index] = i\n            index += 1\n           start[index] = i\n            end[index] = i + rows\n            # \"Clean Up\" the segment that traverses the experiment boundary\n            # This will be the beginning segment of a new experiment.\n            boundary = 0\n            for j in range(rows - 1) :\n                if y_raw[i + j + 1] &gt; y_raw[i + j] :\n                    boundary = j\n            # \"Clean up\" for now means zero out seismic data.\n            X_train[i:i+boundary] = 0\n    # Now let's see how big each experiment is, in terms of the maximum number of 150000 sample rows \n    running_count = 0\n    for i in range (16) :\n        count = np.int32((end[i] - start[i])/rows)\n        print (\"Experiment\", i, \"has\", count, \"segments of 150000 samples each\")\n        running_count += count\n        start[i] = np.int32((start[i])/rows)\n        end[i] = np.int32((end[i])/rows)\n\n    # The total number of segments is\n    segments = np.int32(end[15])",
    "530984": "Paul, your code is hard to read.  Rather than using bulleted list, just add 4 spaces at teh start of each code line.  Code test will be formatted as code that way.",
    "531147": "Thank you! I couldn't find that formatting trick when I looked for it - appreciate your help @cpmpml",
    "540084": "Apologies for the late posting. \n\nI finished a Kernel for you to look over. \n\nThis Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.\n\nThis Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)\n\nGets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\n\nEnjoy, and good luck!"
  },
  "source": "meta"
}