{
  "id": 90565,
  "title": "KFold strategies",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90565",
  "author_name": "",
  "post_date": "2019-04-24T19:11:25.452804600Z",
  "votes": 12,
  "comment_count": 16,
  "views": 0,
  "content": "<p>There are lots of discussions on CV strategies, and I just want to share my experience so far. I'm currently using two methods of KFold:\n<code>float_data = pd.read_csv('../input/train/LANL-Earthquake/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})</code></p>\n\n<p><code>ttf = float_data['time_to_failure']</code></p>\n\n<p>[1] Quake based\n<code>diff = ttf.diff()</code>\n<code>quake_position = np.where(diff&gt;0)[0]</code>\n<code>quake_interval = np.diff(quake_position)</code>\n<code>quake_interval = np.append(quake_position[0], quake_interval)</code>\n<code>quake_idx = [np.repeat(i, quake_interval[i]) for i in range(len(quake_interval))]</code>\n<code>quake_idx = np.concatenate(quake_idx)</code>\n<code>quake_idx= np.append(quake_idx, np.repeat(15, len(ttf)-len(quake_idx)))</code></p>\n\n<p>This way we can separate the training data by earthquake occurrence and we can split the whole data with equal proportion of each earthquake in each fold. I'm currently using this approach for my lgb (best fold is around 1.85, worst fold around 2.1). Even though my public LB score is only around 1.56, I still have faith in my CV.</p>\n\n<p>[2] ttf value based\n<code>sampling_idx = ttf.astype('int')</code>\n<code>sampling_idx[sampling_idx==16] = 15</code></p>\n\n<p>This way we make sure each fold contain equal proportion of ttf values. Based on the fact the majority of the error comes from ttfs with high values and ttf close to 0 (see the picture attached below). I'm using this approach for my RNN models. </p>\n\n<p>By the way, I have created my first ever public kernel <a href=\"https://www.kaggle.com/cjinny/cudnngru-with-stratifiedkfold\">here</a> using RNN with my tff value based StratifiedKFold. feel free to check it out, upvote and give comments!</p>\n\n<p>I've also tried my RNN model without using any shuffling or KFold (just split the data into 4 chunks and using 3 as train, 1 as valid, but the result looked aweful, so I think KFold strategy will be important if you want a reliable CV score).</p>",
  "messages": [
    {
      "id": "522653",
      "postDate": "04/24/2019 19:11:25",
      "content": "<p>There are lots of discussions on CV strategies, and I just want to share my experience so far. I'm currently using two methods of KFold:\n<code>float_data = pd.read_csv('../input/train/LANL-Earthquake/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})</code></p>\n\n<p><code>ttf = float_data['time_to_failure']</code></p>\n\n<p>[1] Quake based\n<code>diff = ttf.diff()</code>\n<code>quake_position = np.where(diff&gt;0)[0]</code>\n<code>quake_interval = np.diff(quake_position)</code>\n<code>quake_interval = np.append(quake_position[0], quake_interval)</code>\n<code>quake_idx = [np.repeat(i, quake_interval[i]) for i in range(len(quake_interval))]</code>\n<code>quake_idx = np.concatenate(quake_idx)</code>\n<code>quake_idx= np.append(quake_idx, np.repeat(15, len(ttf)-len(quake_idx)))</code></p>\n\n<p>This way we can separate the training data by earthquake occurrence and we can split the whole data with equal proportion of each earthquake in each fold. I'm currently using this approach for my lgb (best fold is around 1.85, worst fold around 2.1). Even though my public LB score is only around 1.56, I still have faith in my CV.</p>\n\n<p>[2] ttf value based\n<code>sampling_idx = ttf.astype('int')</code>\n<code>sampling_idx[sampling_idx==16] = 15</code></p>\n\n<p>This way we make sure each fold contain equal proportion of ttf values. Based on the fact the majority of the error comes from ttfs with high values and ttf close to 0 (see the picture attached below). I'm using this approach for my RNN models. </p>\n\n<p>By the way, I have created my first ever public kernel <a href=\"https://www.kaggle.com/cjinny/cudnngru-with-stratifiedkfold\">here</a> using RNN with my tff value based StratifiedKFold. feel free to check it out, upvote and give comments!</p>\n\n<p>I've also tried my RNN model without using any shuffling or KFold (just split the data into 4 chunks and using 3 as train, 1 as valid, but the result looked aweful, so I think KFold strategy will be important if you want a reliable CV score).</p>",
      "rawMarkdown": "There are lots of discussions on CV strategies, and I just want to share my experience so far. I'm currently using two methods of KFold:\n`float_data = pd.read_csv('../input/train/LANL-Earthquake/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})`\n\n`ttf = float_data['time_to_failure']`\n\n[1] Quake based\n`diff = ttf.diff()`\n`quake_position = np.where(diff&gt;0)[0]`\n`quake_interval = np.diff(quake_position)`\n`quake_interval = np.append(quake_position[0], quake_interval)`\n`quake_idx = [np.repeat(i, quake_interval[i]) for i in range(len(quake_interval))]`\n`quake_idx = np.concatenate(quake_idx)`\n`quake_idx= np.append(quake_idx, np.repeat(15, len(ttf)-len(quake_idx)))`\n\nThis way we can separate the training data by earthquake occurrence and we can split the whole data with equal proportion of each earthquake in each fold. I'm currently using this approach for my lgb (best fold is around 1.85, worst fold around 2.1). Even though my public LB score is only around 1.56, I still have faith in my CV.\n\n[2] ttf value based\n`sampling_idx = ttf.astype('int')`\n`sampling_idx[sampling_idx==16] = 15`\n\nThis way we make sure each fold contain equal proportion of ttf values. Based on the fact the majority of the error comes from ttfs with high values and ttf close to 0 (see the picture attached below). I'm using this approach for my RNN models. \n\nBy the way, I have created my first ever public kernel [here](https://www.kaggle.com/cjinny/cudnngru-with-stratifiedkfold) using RNN with my tff value based StratifiedKFold. feel free to check it out, upvote and give comments!\n\nI've also tried my RNN model without using any shuffling or KFold (just split the data into 4 chunks and using 3 as train, 1 as valid, but the result looked aweful, so I think KFold strategy will be important if you want a reliable CV score).",
      "votes": null
    },
    {
      "id": "522668",
      "postDate": "04/24/2019 19:39:22",
      "content": "<p>I am very surprised to see that you get 1.85 on a quake-wise CV! I am using the same approach and I am well above 2.</p>",
      "rawMarkdown": "I am very surprised to see that you get 1.85 on a quake-wise CV! I am using the same approach and I am well above 2.",
      "votes": null
    },
    {
      "id": "522679",
      "postDate": "04/24/2019 19:56:53",
      "content": "<p><code>Fold_0 MAE is 1.8816469681606713</code>\n<code>Fold_1 MAE is 2.092382525302977</code>\n<code>Fold_2 MAE is 2.023297700165216</code>\n<code>Fold_3 MAE is 1.8488641125271774</code>\n<code>Fold_4 MAE is 2.024186248515722</code>\n<code>Fold_5 MAE is 1.9653747969286623</code>\n<code>Fold_6 MAE is 2.0466579107974647</code>\n<code>Fold_7 MAE is 1.945615334998304</code>\n<code>Fold_8 MAE is 2.0013864415150326</code>\n<code>Final MAE is 1.9809075156544214</code>\n<code>Public LB is 1.596 !!! I have no idea why</code>\nThere is still variation between each fold, and if I change the random seed, I can get CV anywhere around 1.981 ~ 1.99, so I'm not sure how the private LB would look like...</p>",
      "rawMarkdown": "`Fold_0 MAE is 1.8816469681606713`\n`Fold_1 MAE is 2.092382525302977`\n`Fold_2 MAE is 2.023297700165216`\n`Fold_3 MAE is 1.8488641125271774`\n`Fold_4 MAE is 2.024186248515722`\n`Fold_5 MAE is 1.9653747969286623`\n`Fold_6 MAE is 2.0466579107974647`\n`Fold_7 MAE is 1.945615334998304`\n`Fold_8 MAE is 2.0013864415150326`\n`Final MAE is 1.9809075156544214`\n`Public LB is 1.596 !!! I have no idea why`\nThere is still variation between each fold, and if I change the random seed, I can get CV anywhere around 1.981 ~ 1.99, so I'm not sure how the private LB would look like...",
      "votes": null
    },
    {
      "id": "522717",
      "postDate": "04/24/2019 22:08:29",
      "content": "<p><a href=\"/cjinny\">@cjinny</a> , thanks for sharing. Is your result repeatable? i.e. do you get the same result with every run?</p>",
      "rawMarkdown": "cjinny , thanks for sharing. Is your result repeatable? i.e. do you get the same result with every run?",
      "votes": null
    },
    {
      "id": "522753",
      "postDate": "04/25/2019 00:29:56",
      "content": "<p>Yeah, I made sure the seeds stay the same in my lgb model, but as I changed the StratifiedKFold seed I saw some variations in my CV (1.981 ~ 1.995)</p>",
      "rawMarkdown": "Yeah, I made sure the seeds stay the same in my lgb model, but as I changed the StratifiedKFold seed I saw some variations in my CV (1.981 ~ 1.995)",
      "votes": null
    },
    {
      "id": "523180",
      "postDate": "04/25/2019 17:29:55",
      "content": "<p>You early stop on the validation data in your RNN kernel --&gt; you train until the 2? quakes of your validation data fit best. Correct me if i'm wrong, but this explains the low cv score.</p>",
      "rawMarkdown": "You early stop on the validation data in your RNN kernel --&gt; you train until the 2? quakes of your validation data fit best. Correct me if i'm wrong, but this explains the low cv score.",
      "votes": null
    },
    {
      "id": "523288",
      "postDate": "04/25/2019 23:34:36",
      "content": "<p>Do you mean these lines? I'm not sure what you meant by train until the 2.\n<code>ckpt = ModelCheckpoint(\"RNN_model_{}.hdf5\".format(n_fold), save_best_only=True, period=3)</code>\n<code>es = EarlyStopping(monitor='val_loss', patience=10)</code></p>",
      "rawMarkdown": "Do you mean these lines? I'm not sure what you meant by train until the 2.\n`ckpt = ModelCheckpoint(\"RNN_model_{}.hdf5\".format(n_fold), save_best_only=True, period=3)`\n`es = EarlyStopping(monitor='val_loss', patience=10)`",
      "votes": null
    },
    {
      "id": "523291",
      "postDate": "04/25/2019 23:54:10",
      "content": "<p>Yes i counted 8 folds in your other post (and just realized it's actually 9), so i guess you use 14 quakes for training and 2 for validation each split, right? </p>\n\n<p>If you use the ModelCheckpoint like in the example above, you train until your validation score doesn't improve for 10 epochs and save the model from the epoch that minimizes your validation error. </p>\n\n<p>However, your validation data only consists of data from 2 earthquakes. I think the model you get, when you take the one that fits those 2 earthquakes best, is probably not one that generalized well to new data, therefore low cv, but worse then expected lb. \nI saw similar training/validation error curves in some of my own DL approaches and i tried quite many of them since i wanted to get more into keras with this challenge. The validation score is really noisy in later epochs, which i think makes it even harder to evaluate the quality of the model you get from early stopping on one of these epochs. I always tried to smooth the validation curve, which sometimes worked, but not always.</p>",
      "rawMarkdown": "Yes i counted 8 folds in your other post (and just realized it's actually 9), so i guess you use 14 quakes for training and 2 for validation each split, right? \n\nIf you use the ModelCheckpoint like in the example above, you train until your validation score doesn't improve for 10 epochs and save the model from the epoch that minimizes your validation error. \n\nHowever, your validation data only consists of data from 2 earthquakes. I think the model you get, when you take the one that fits those 2 earthquakes best, is probably not one that generalized well to new data, therefore low cv, but worse then expected lb. \nI saw similar training/validation error curves in some of my own DL approaches and i tried quite many of them since i wanted to get more into keras with this challenge. The validation score is really noisy in later epochs, which i think makes it even harder to evaluate the quality of the model you get from early stopping on one of these epochs. I always tried to smooth the validation curve, which sometimes worked, but not always.",
      "votes": null
    },
    {
      "id": "523292",
      "postDate": "04/26/2019 00:01:22",
      "content": "<p><a href=\"/cjinny\">@cjinny</a> thanks and I am not sure I understand how you split the data for lgbm nFolds in your Quake based approach. Do you mind sharing a kernel or a snippet of how u do that?</p>",
      "rawMarkdown": "cjinny thanks and I am not sure I understand how you split the data for lgbm nFolds in your Quake based approach. Do you mind sharing a kernel or a snippet of how u do that?",
      "votes": null
    },
    {
      "id": "523351",
      "postDate": "04/26/2019 04:14:20",
      "content": "<p>Hi, my quake-based idea is like this:</p>\n\n<p>I put a label to every quake (0-15) as my quake-idx (so that data in first quake would get a label of 0) and then do StratifiedKFold to split quake-idx to 9 Fold with each fold consisting of equal proportion of quake_idx.</p>\n\n<p>For example: let's say in 1% data has quake-idx == 0,  7% of data has quake-idx == 1. Then in every fold both my training and validation data would consist of around 1% of data coming from quake-idx==0 and 7% of data coming from quake-idx==1.</p>\n\n<p>In short, my validation data consists of data from all earthquakes, mixed up in similar proportion as the entire train data.  In fact, I've tried before using data from 14 quakes as train and validate on 2 but the val_error varies greatly in each fold (something like 1.5 ~ 2.5).</p>",
      "rawMarkdown": "Hi, my quake-based idea is like this:\n\nI put a label to every quake (0-15) as my quake-idx (so that data in first quake would get a label of 0) and then do StratifiedKFold to split quake-idx to 9 Fold with each fold consisting of equal proportion of quake_idx.\n\nFor example: let's say in 1% data has quake-idx == 0,  7% of data has quake-idx == 1. Then in every fold both my training and validation data would consist of around 1% of data coming from quake-idx==0 and 7% of data coming from quake-idx==1.\n\nIn short, my validation data consists of data from all earthquakes, mixed up in similar proportion as the entire train data.  In fact, I've tried before using data from 14 quakes as train and validate on 2 but the val_error varies greatly in each fold (something like 1.5 ~ 2.5).",
      "votes": null
    },
    {
      "id": "523353",
      "postDate": "04/26/2019 04:16:44",
      "content": "<p><a href=\"/sheriytm\">@sheriytm</a> I'll do that this weekend. By the way, would you mind sharing your CV score? Your LB is a lot better than mine, LOL</p>",
      "rawMarkdown": "sheriytm I'll do that this weekend. By the way, would you mind sharing your CV score? Your LB is a lot better than mine, LOL",
      "votes": null
    },
    {
      "id": "523362",
      "postDate": "04/26/2019 05:02:27",
      "content": "<p><a href=\"/cjinny\">@cjinny</a>, my current LB score is as a result of weighted blending of 4 models so I do not have a CV score. However, my best model's scores are CV= 2.0167, LB=1.441</p>",
      "rawMarkdown": "cjinny, my current LB score is as a result of weighted blending of 4 models so I do not have a CV score. However, my best model's scores are CV= 2.0167, LB=1.441",
      "votes": null
    },
    {
      "id": "523367",
      "postDate": "04/26/2019 05:26:52",
      "content": "<p>I see, does blending/stacking help in your experience? I'll probably try some stacking during the late stage</p>",
      "rawMarkdown": "I see, does blending/stacking help in your experience? I'll probably try some stacking during the late stage",
      "votes": null
    },
    {
      "id": "523373",
      "postDate": "04/26/2019 05:48:59",
      "content": "<p>I don't normally do stacking at this point of a competition but this is something I tried a month ago hence I can't say much at this point for lack of enough experiments. I will go back to it in the last week of the competition.</p>\n\n<p>I look forward to seeing how you used your Quake based CV with LGBM over the weekend:-)</p>",
      "rawMarkdown": "I don't normally do stacking at this point of a competition but this is something I tried a month ago hence I can't say much at this point for lack of enough experiments. I will go back to it in the last week of the competition.\n\n I look forward to seeing how you used your Quake based CV with LGBM over the weekend:-)",
      "votes": null
    },
    {
      "id": "524619",
      "postDate": "04/29/2019 07:44:56",
      "content": "<p>Hi I just made the kernel using LGBM with quake-based and ttf based splits, it doesn't seem to help much though as compared to the normal KFold, but I think the feature engineering could be useful...</p>",
      "rawMarkdown": "Hi I just made the kernel using LGBM with quake-based and ttf based splits, it doesn't seem to help much though as compared to the normal KFold, but I think the feature engineering could be useful...",
      "votes": null
    },
    {
      "id": "524629",
      "postDate": "04/29/2019 08:14:10",
      "content": "<p>Thanks for sharing <a href=\"/cjinny\">@cjinny</a> and I just upvoted it. I tried your NN kernel with my features and it did not do well compared to the KFold either.</p>",
      "rawMarkdown": "Thanks for sharing @cjinny and I just upvoted it. I tried your NN kernel with my features and it did not do well compared to the KFold either.",
      "votes": null
    },
    {
      "id": "524825",
      "postDate": "04/29/2019 15:30:20",
      "content": "<p>Neither do I, I've tried more extensive features with RNN but CV is always above 2. Anyway, I think I'll stick to lgb/xgb for my final submissions :)</p>",
      "rawMarkdown": "Neither do I, I've tried more extensive features with RNN but CV is always above 2. Anyway, I think I'll stick to lgb/xgb for my final submissions :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 522668,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "04/24/2019 19:39:22",
      "content": "<p>I am very surprised to see that you get 1.85 on a quake-wise CV! I am using the same approach and I am well above 2.</p>",
      "votes": null,
      "replies": [
        {
          "id": 522679,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/24/2019 19:56:53",
          "content": "<p><code>Fold_0 MAE is 1.8816469681606713</code>\n<code>Fold_1 MAE is 2.092382525302977</code>\n<code>Fold_2 MAE is 2.023297700165216</code>\n<code>Fold_3 MAE is 1.8488641125271774</code>\n<code>Fold_4 MAE is 2.024186248515722</code>\n<code>Fold_5 MAE is 1.9653747969286623</code>\n<code>Fold_6 MAE is 2.0466579107974647</code>\n<code>Fold_7 MAE is 1.945615334998304</code>\n<code>Fold_8 MAE is 2.0013864415150326</code>\n<code>Final MAE is 1.9809075156544214</code>\n<code>Public LB is 1.596 !!! I have no idea why</code>\nThere is still variation between each fold, and if I change the random seed, I can get CV anywhere around 1.981 ~ 1.99, so I'm not sure how the private LB would look like...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522717,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "04/24/2019 22:08:29",
      "content": "<p><a href=\"/cjinny\">@cjinny</a> , thanks for sharing. Is your result repeatable? i.e. do you get the same result with every run?</p>",
      "votes": null,
      "replies": [
        {
          "id": 522753,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/25/2019 00:29:56",
          "content": "<p>Yeah, I made sure the seeds stay the same in my lgb model, but as I changed the StratifiedKFold seed I saw some variations in my CV (1.981 ~ 1.995)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523292,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "04/26/2019 00:01:22",
          "content": "<p><a href=\"/cjinny\">@cjinny</a> thanks and I am not sure I understand how you split the data for lgbm nFolds in your Quake based approach. Do you mind sharing a kernel or a snippet of how u do that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523353,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/26/2019 04:16:44",
          "content": "<p><a href=\"/sheriytm\">@sheriytm</a> I'll do that this weekend. By the way, would you mind sharing your CV score? Your LB is a lot better than mine, LOL</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523362,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "04/26/2019 05:02:27",
          "content": "<p><a href=\"/cjinny\">@cjinny</a>, my current LB score is as a result of weighted blending of 4 models so I do not have a CV score. However, my best model's scores are CV= 2.0167, LB=1.441</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523367,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/26/2019 05:26:52",
          "content": "<p>I see, does blending/stacking help in your experience? I'll probably try some stacking during the late stage</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523373,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "04/26/2019 05:48:59",
          "content": "<p>I don't normally do stacking at this point of a competition but this is something I tried a month ago hence I can't say much at this point for lack of enough experiments. I will go back to it in the last week of the competition.</p>\n\n<p>I look forward to seeing how you used your Quake based CV with LGBM over the weekend:-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524619,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/29/2019 07:44:56",
          "content": "<p>Hi I just made the kernel using LGBM with quake-based and ttf based splits, it doesn't seem to help much though as compared to the normal KFold, but I think the feature engineering could be useful...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524629,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "04/29/2019 08:14:10",
          "content": "<p>Thanks for sharing <a href=\"/cjinny\">@cjinny</a> and I just upvoted it. I tried your NN kernel with my features and it did not do well compared to the KFold either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 524825,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/29/2019 15:30:20",
          "content": "<p>Neither do I, I've tried more extensive features with RNN but CV is always above 2. Anyway, I think I'll stick to lgb/xgb for my final submissions :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 523180,
      "author_name": "svenhinderer",
      "author_url": "",
      "post_date": "04/25/2019 17:29:55",
      "content": "<p>You early stop on the validation data in your RNN kernel --&gt; you train until the 2? quakes of your validation data fit best. Correct me if i'm wrong, but this explains the low cv score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 523288,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/25/2019 23:34:36",
          "content": "<p>Do you mean these lines? I'm not sure what you meant by train until the 2.\n<code>ckpt = ModelCheckpoint(\"RNN_model_{}.hdf5\".format(n_fold), save_best_only=True, period=3)</code>\n<code>es = EarlyStopping(monitor='val_loss', patience=10)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523291,
          "author_name": "svenhinderer",
          "author_url": "",
          "post_date": "04/25/2019 23:54:10",
          "content": "<p>Yes i counted 8 folds in your other post (and just realized it's actually 9), so i guess you use 14 quakes for training and 2 for validation each split, right? </p>\n\n<p>If you use the ModelCheckpoint like in the example above, you train until your validation score doesn't improve for 10 epochs and save the model from the epoch that minimizes your validation error. </p>\n\n<p>However, your validation data only consists of data from 2 earthquakes. I think the model you get, when you take the one that fits those 2 earthquakes best, is probably not one that generalized well to new data, therefore low cv, but worse then expected lb. \nI saw similar training/validation error curves in some of my own DL approaches and i tried quite many of them since i wanted to get more into keras with this challenge. The validation score is really noisy in later epochs, which i think makes it even harder to evaluate the quality of the model you get from early stopping on one of these epochs. I always tried to smooth the validation curve, which sometimes worked, but not always.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 523351,
          "author_name": "cjinny",
          "author_url": "",
          "post_date": "04/26/2019 04:14:20",
          "content": "<p>Hi, my quake-based idea is like this:</p>\n\n<p>I put a label to every quake (0-15) as my quake-idx (so that data in first quake would get a label of 0) and then do StratifiedKFold to split quake-idx to 9 Fold with each fold consisting of equal proportion of quake_idx.</p>\n\n<p>For example: let's say in 1% data has quake-idx == 0,  7% of data has quake-idx == 1. Then in every fold both my training and validation data would consist of around 1% of data coming from quake-idx==0 and 7% of data coming from quake-idx==1.</p>\n\n<p>In short, my validation data consists of data from all earthquakes, mixed up in similar proportion as the entire train data.  In fact, I've tried before using data from 14 quakes as train and validate on 2 but the val_error varies greatly in each fold (something like 1.5 ~ 2.5).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "522653": "There are lots of discussions on CV strategies, and I just want to share my experience so far. I'm currently using two methods of KFold:\n`float_data = pd.read_csv('../input/train/LANL-Earthquake/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})`\n\n`ttf = float_data['time_to_failure']`\n\n[1] Quake based\n`diff = ttf.diff()`\n`quake_position = np.where(diff&gt;0)[0]`\n`quake_interval = np.diff(quake_position)`\n`quake_interval = np.append(quake_position[0], quake_interval)`\n`quake_idx = [np.repeat(i, quake_interval[i]) for i in range(len(quake_interval))]`\n`quake_idx = np.concatenate(quake_idx)`\n`quake_idx= np.append(quake_idx, np.repeat(15, len(ttf)-len(quake_idx)))`\n\nThis way we can separate the training data by earthquake occurrence and we can split the whole data with equal proportion of each earthquake in each fold. I'm currently using this approach for my lgb (best fold is around 1.85, worst fold around 2.1). Even though my public LB score is only around 1.56, I still have faith in my CV.\n\n[2] ttf value based\n`sampling_idx = ttf.astype('int')`\n`sampling_idx[sampling_idx==16] = 15`\n\nThis way we make sure each fold contain equal proportion of ttf values. Based on the fact the majority of the error comes from ttfs with high values and ttf close to 0 (see the picture attached below). I'm using this approach for my RNN models. \n\nBy the way, I have created my first ever public kernel [here](https://www.kaggle.com/cjinny/cudnngru-with-stratifiedkfold) using RNN with my tff value based StratifiedKFold. feel free to check it out, upvote and give comments!\n\nI've also tried my RNN model without using any shuffling or KFold (just split the data into 4 chunks and using 3 as train, 1 as valid, but the result looked aweful, so I think KFold strategy will be important if you want a reliable CV score).",
    "522668": "I am very surprised to see that you get 1.85 on a quake-wise CV! I am using the same approach and I am well above 2.",
    "522679": "`Fold_0 MAE is 1.8816469681606713`\n`Fold_1 MAE is 2.092382525302977`\n`Fold_2 MAE is 2.023297700165216`\n`Fold_3 MAE is 1.8488641125271774`\n`Fold_4 MAE is 2.024186248515722`\n`Fold_5 MAE is 1.9653747969286623`\n`Fold_6 MAE is 2.0466579107974647`\n`Fold_7 MAE is 1.945615334998304`\n`Fold_8 MAE is 2.0013864415150326`\n`Final MAE is 1.9809075156544214`\n`Public LB is 1.596 !!! I have no idea why`\nThere is still variation between each fold, and if I change the random seed, I can get CV anywhere around 1.981 ~ 1.99, so I'm not sure how the private LB would look like...",
    "522717": "cjinny , thanks for sharing. Is your result repeatable? i.e. do you get the same result with every run?",
    "522753": "Yeah, I made sure the seeds stay the same in my lgb model, but as I changed the StratifiedKFold seed I saw some variations in my CV (1.981 ~ 1.995)",
    "523180": "You early stop on the validation data in your RNN kernel --&gt; you train until the 2? quakes of your validation data fit best. Correct me if i'm wrong, but this explains the low cv score.",
    "523288": "Do you mean these lines? I'm not sure what you meant by train until the 2.\n`ckpt = ModelCheckpoint(\"RNN_model_{}.hdf5\".format(n_fold), save_best_only=True, period=3)`\n`es = EarlyStopping(monitor='val_loss', patience=10)`",
    "523291": "Yes i counted 8 folds in your other post (and just realized it's actually 9), so i guess you use 14 quakes for training and 2 for validation each split, right? \n\nIf you use the ModelCheckpoint like in the example above, you train until your validation score doesn't improve for 10 epochs and save the model from the epoch that minimizes your validation error. \n\nHowever, your validation data only consists of data from 2 earthquakes. I think the model you get, when you take the one that fits those 2 earthquakes best, is probably not one that generalized well to new data, therefore low cv, but worse then expected lb. \nI saw similar training/validation error curves in some of my own DL approaches and i tried quite many of them since i wanted to get more into keras with this challenge. The validation score is really noisy in later epochs, which i think makes it even harder to evaluate the quality of the model you get from early stopping on one of these epochs. I always tried to smooth the validation curve, which sometimes worked, but not always.",
    "523292": "cjinny thanks and I am not sure I understand how you split the data for lgbm nFolds in your Quake based approach. Do you mind sharing a kernel or a snippet of how u do that?",
    "523351": "Hi, my quake-based idea is like this:\n\nI put a label to every quake (0-15) as my quake-idx (so that data in first quake would get a label of 0) and then do StratifiedKFold to split quake-idx to 9 Fold with each fold consisting of equal proportion of quake_idx.\n\nFor example: let's say in 1% data has quake-idx == 0,  7% of data has quake-idx == 1. Then in every fold both my training and validation data would consist of around 1% of data coming from quake-idx==0 and 7% of data coming from quake-idx==1.\n\nIn short, my validation data consists of data from all earthquakes, mixed up in similar proportion as the entire train data.  In fact, I've tried before using data from 14 quakes as train and validate on 2 but the val_error varies greatly in each fold (something like 1.5 ~ 2.5).",
    "523353": "sheriytm I'll do that this weekend. By the way, would you mind sharing your CV score? Your LB is a lot better than mine, LOL",
    "523362": "cjinny, my current LB score is as a result of weighted blending of 4 models so I do not have a CV score. However, my best model's scores are CV= 2.0167, LB=1.441",
    "523367": "I see, does blending/stacking help in your experience? I'll probably try some stacking during the late stage",
    "523373": "I don't normally do stacking at this point of a competition but this is something I tried a month ago hence I can't say much at this point for lack of enough experiments. I will go back to it in the last week of the competition.\n\n I look forward to seeing how you used your Quake based CV with LGBM over the weekend:-)",
    "524619": "Hi I just made the kernel using LGBM with quake-based and ttf based splits, it doesn't seem to help much though as compared to the normal KFold, but I think the feature engineering could be useful...",
    "524629": "Thanks for sharing @cjinny and I just upvoted it. I tried your NN kernel with my features and it did not do well compared to the KFold either.",
    "524825": "Neither do I, I've tried more extensive features with RNN but CV is always above 2. Anyway, I think I'll stick to lgb/xgb for my final submissions :)"
  },
  "source": "meta"
}