{
  "id": 78809,
  "title": "CV Strategy",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/78809",
  "author_name": "",
  "post_date": "2019-01-28T06:35:06.639162Z",
  "votes": 34,
  "comment_count": 39,
  "views": 0,
  "content": "<p>In order to get a less biased CV, especially if we augment the data,</p>\n\n<p>Try to split by the earthquake ID ( 16 or 15 in total, however you define)</p>\n\n<p>For example, for a 5 fold CV,\nFold 1: use the first 13 earthquakes to train, and the last 3 to validate\n etc.</p>",
  "messages": [
    {
      "id": "462341",
      "postDate": "01/28/2019 06:35:06",
      "content": "<p>In order to get a less biased CV, especially if we augment the data,</p>\n\n<p>Try to split by the earthquake ID ( 16 or 15 in total, however you define)</p>\n\n<p>For example, for a 5 fold CV,\nFold 1: use the first 13 earthquakes to train, and the last 3 to validate\n etc.</p>",
      "rawMarkdown": "In order to get a less biased CV, especially if we augment the data,\n\nTry to split by the earthquake ID ( 16 or 15 in total, however you define)\n\nFor example, for a 5 fold CV,\nFold 1: use the first 13 earthquakes to train, and the last 3 to validate\n etc.",
      "votes": null
    },
    {
      "id": "462472",
      "postDate": "01/28/2019 10:47:09",
      "content": "<p>Thank you for the advice. However, depending on the validation fold size (number of earthquakes in the validation set) I get CV scores between 1.8-2.1. These best submissions are around 1.65. </p>\n\n<p>I'm not sure how many earthquakes should there be in validation sets.</p>",
      "rawMarkdown": "Thank you for the advice. However, depending on the validation fold size (number of earthquakes in the validation set) I get CV scores between 1.8-2.1. These best submissions are around 1.65. \n\nI'm not sure how many earthquakes should there be in validation sets.",
      "votes": null
    },
    {
      "id": "462714",
      "postDate": "01/28/2019 18:55:55",
      "content": "<p>But i have a question. I was seeing a test data segment and the acoustic data of one segment had a high acoustic signal at start that means probably earthquake occurred at start. This shows 1 segment can have data of 2 earthquakes. This CV technique would work in such case? </p>",
      "rawMarkdown": "But i have a question. I was seeing a test data segment and the acoustic data of one segment had a high acoustic signal at start that means probably earthquake occurred at start. This shows 1 segment can have data of 2 earthquakes. This CV technique would work in such case?",
      "votes": null
    },
    {
      "id": "462732",
      "postDate": "01/28/2019 19:42:53",
      "content": "<p>I am not sure we can say that there are 2 earthquakes. High peak in the acoustic data is not always followed by an earthquake (resetting the time) . Also high peak itself is not the where the actual earthquake is(at least not in the train file), it is hidden somewhere after the peak. </p>",
      "rawMarkdown": "I am not sure we can say that there are 2 earthquakes. High peak in the acoustic data is not always followed by an earthquake (resetting the time) . Also high peak itself is not the where the actual earthquake is(at least not in the train file), it is hidden somewhere after the peak.",
      "votes": null
    },
    {
      "id": "462736",
      "postDate": "01/28/2019 19:53:24",
      "content": "<p>Thank you Elliot and DavidSfor sharing your ideas!\nIf we split it in Elliot's way we will have proportionally eq /no eq, both in train and in validation.</p>\n\n<p>Overlapping of the batches will insert same sequence once at the beginning and once at the end  of every train/test batch. But I can't say what the effect on the model will be. Is it too much if we have several different CV techniques? </p>",
      "rawMarkdown": "Thank you Elliot and DavidSfor sharing your ideas!\nIf we split it in Elliot's way we will have proportionally eq /no eq, both in train and in validation.\n\nOverlapping of the batches will insert same sequence once at the beginning and once at the end  of every train/test batch. But I can't say what the effect on the model will be. Is it too much if we have several different CV techniques?",
      "votes": null
    },
    {
      "id": "462767",
      "postDate": "01/28/2019 21:21:33",
      "content": "<p>Thx for your input Elliot. I am also using this exact strategy, maybe with different groupings. One thing I am generally (not limited to this strategy) not happy about is that the fold containing the longest chunk will automatically get a bad CV score, at least if working with decision trees. They cannot extrapolate without tweaks. So maybe it might also be an option to make e.g. 50 chunks of length 12 mio each and then build the folds by randomly regrouping them into five folds to partially fight this problem. </p>\n\n<p>Happy for any idea.</p>",
      "rawMarkdown": "Thx for your input Elliot. I am also using this exact strategy, maybe with different groupings. One thing I am generally (not limited to this strategy) not happy about is that the fold containing the longest chunk will automatically get a bad CV score, at least if working with decision trees. They cannot extrapolate without tweaks. So maybe it might also be an option to make e.g. 50 chunks of length 12 mio each and then build the folds by randomly regrouping them into five folds to partially fight this problem. \n\nHappy for any idea.",
      "votes": null
    },
    {
      "id": "462809",
      "postDate": "01/28/2019 23:58:01",
      "content": "<p>Hi Micheal,</p>\n\n<p>If those edge values are really of concern, we can just have them always grouped in the training set.\nThe current implementation of the boosting trees don't have higher order basis implemented,\nso we can't combat the extrapolation problem head-on, but to contain them.</p>\n\n<p>In our case, however, the extrapolation is far from the top concern.\nThe bad CV score for the longest chunk, is mostly, due to the systematically under-estimated 'time to failure'.</p>\n\n<p>These are my thoughts to your response :)</p>",
      "rawMarkdown": "Hi Micheal,\n\nIf those edge values are really of concern, we can just have them always grouped in the training set.\nThe current implementation of the boosting trees don't have higher order basis implemented,\nso we can't combat the extrapolation problem head-on, but to contain them.\n\nIn our case, however, the extrapolation is far from the top concern.\nThe bad CV score for the longest chunk, is mostly, due to the systematically under-estimated 'time to failure'.\n\nThese are my thoughts to your response :)",
      "votes": null
    },
    {
      "id": "462813",
      "postDate": "01/29/2019 00:05:53",
      "content": "<p>Hi David,\nWhat CV can offer is a relative benchmark for you to compare among the models.\nAs for how accurate they can be served as an direct estimate on the test dataset?\nThat would depend on the similarity between the distributions're offered.\nIn our case, the mean value for the public LB is only 4.0xxx. This is completely different from our training set. So we will need to have some other methods to give the desired result.</p>",
      "rawMarkdown": "Hi David,\nWhat CV can offer is a relative benchmark for you to compare among the models.\nAs for how accurate they can be served as an direct estimate on the test dataset?\nThat would depend on the similarity between the distributions're offered.\nIn our case, the mean value for the public LB is only 4.0xxx. This is completely different from our training set. So we will need to have some other methods to give the desired result.",
      "votes": null
    },
    {
      "id": "462814",
      "postDate": "01/29/2019 00:09:15",
      "content": "<p>I would say this belongs to data cleaning issue :)\nIn my script, I don't allow segments that cross the borders.</p>",
      "rawMarkdown": "I would say this belongs to data cleaning issue :)\nIn my script, I don't allow segments that cross the borders.",
      "votes": null
    },
    {
      "id": "462819",
      "postDate": "01/29/2019 00:17:17",
      "content": "<p>Hi Ammad,\nFirstly, in the data preparation phase, I would recommend removing segments that cross borders. It is still fine if you keep them. Their number is tiny anyway.</p>\n\n<p>And second, I wouldn't decide every high peak to be near an earthquake. This is the most challenging part in this competition. If you can check out the 3rd, 8th and 15th earthquake in our training data, you will find high peaks everywhere even in the early stages.</p>",
      "rawMarkdown": "Hi Ammad,\nFirstly, in the data preparation phase, I would recommend removing segments that cross borders. It is still fine if you keep them. Their number is tiny anyway.\n\nAnd second, I wouldn't decide every high peak to be near an earthquake. This is the most challenging part in this competition. If you can check out the 3rd, 8th and 15th earthquake in our training data, you will find high peaks everywhere even in the early stages.",
      "votes": null
    },
    {
      "id": "462899",
      "postDate": "01/29/2019 04:36:50",
      "content": "<p>Hi All,</p>\n\n<p>I'm just curious as to how close people's CV scores are to the public leaderboard. I'm getting OK CV but I assume it will blow up when I submit?</p>",
      "rawMarkdown": "Hi All,\n\nI'm just curious as to how close people's CV scores are to the public leaderboard. I'm getting OK CV but I assume it will blow up when I submit?",
      "votes": null
    },
    {
      "id": "463184",
      "postDate": "01/29/2019 15:24:55",
      "content": "<p>That's interesting, I didn't notice the train mean value is ~5.6. Thx!</p>",
      "rawMarkdown": "That's interesting, I didn't notice the train mean value is ~5.6. Thx!",
      "votes": null
    },
    {
      "id": "463279",
      "postDate": "01/29/2019 18:25:44",
      "content": "<p>One thing to keep in mind is that leaderboard climbing is a beguiling as the sirens in Greek mythology - I can guarantee that a lot of people's best entries will be about 30 submissions below their final one!</p>",
      "rawMarkdown": "One thing to keep in mind is that leaderboard climbing is a beguiling as the sirens in Greek mythology - I can guarantee that a lot of people's best entries will be about 30 submissions below their final one!",
      "votes": null
    },
    {
      "id": "463360",
      "postDate": "01/29/2019 21:26:42",
      "content": "<p>Hi Elliot, thanks for this advice. I have not had much success with parsing the raw signal in ways other than demonstrated in the starter kernel yet, but I will try again. Your comment makes me wonder a few things and I am curious about the logic those with more knowledge of geoscience and physics than me (probably everyone) are using.</p>\n\n<p>First, it is interesting that the training data is one continuous signal which we can presumably chop up in any way we like. Assuming what I've read is true about the sampling frequency, the training signal consists of 157.275 seconds with 16 points where time to failure reaches zero, i.e., a labquake has occurred (and this timing was acquired by a different device than was used to obtain the acoustic data and which is not available to us).  It seems rational to begin by cutting the training signal to resemble the test set, as the starter did, yielding ~4200 segments of 150000 rows each comprising 37.5 milliseconds where the time to failure (target) is the last row of each segment. Now, we are evaluated on the predicted time to failure for each of the ~2k segments of 37ms of acoustic signal -- since the predictions are in seconds I am assuming there are likely no quakes in the test data, but I cannot say that for a certainty. My feeling is it does not matter if there are quakes in test or not, since we are not asked to identify quakes per se but how much time remains before a likely quake given current sensor data.</p>\n\n<p>My questions, so far, are these. Holding out one or two quake cycles from train to use as validation still requires breaking up into test-length segments? I agree that using overlapping segments it probably a bad idea and when I've tried it my score is way worse, I'm wondering if you think this an artifact of how we are being scored or if this is common knowledge for researchers working with seismic data? I'll stop here for now and thanks again for the insights. Best of luck to you!</p>",
      "rawMarkdown": "Hi Elliot, thanks for this advice. I have not had much success with parsing the raw signal in ways other than demonstrated in the starter kernel yet, but I will try again. Your comment makes me wonder a few things and I am curious about the logic those with more knowledge of geoscience and physics than me (probably everyone) are using.\n\nFirst, it is interesting that the training data is one continuous signal which we can presumably chop up in any way we like. Assuming what I've read is true about the sampling frequency, the training signal consists of 157.275 seconds with 16 points where time to failure reaches zero, i.e., a labquake has occurred (and this timing was acquired by a different device than was used to obtain the acoustic data and which is not available to us).  It seems rational to begin by cutting the training signal to resemble the test set, as the starter did, yielding ~4200 segments of 150000 rows each comprising 37.5 milliseconds where the time to failure (target) is the last row of each segment. Now, we are evaluated on the predicted time to failure for each of the ~2k segments of 37ms of acoustic signal -- since the predictions are in seconds I am assuming there are likely no quakes in the test data, but I cannot say that for a certainty. My feeling is it does not matter if there are quakes in test or not, since we are not asked to identify quakes per se but how much time remains before a likely quake given current sensor data.\n\nMy questions, so far, are these. Holding out one or two quake cycles from train to use as validation still requires breaking up into test-length segments? I agree that using overlapping segments it probably a bad idea and when I've tried it my score is way worse, I'm wondering if you think this an artifact of how we are being scored or if this is common knowledge for researchers working with seismic data? I'll stop here for now and thanks again for the insights. Best of luck to you!",
      "votes": null
    },
    {
      "id": "463432",
      "postDate": "01/30/2019 01:18:37",
      "content": "<p>I agree with your reasoning in the first two paragraphs though I did not re-check the exact numbers.</p>\n\n<p>The questions: \nI think the most straightforward way is to break into 150K segments but one can envision more complicated ways of handling the data - so I'm not sure about this.\nOn overlaps: It's not quite clear - are you talking about overlapping train and validation segments? If so, that is not a good thing to do as a leak is established that compromises the information provided by validation.</p>\n\n<p>If you're talking about overlapping segments within a train or test data set - I don't see anything wrong with that - it resembles image zoom augmentation, but in one dimension. There is also nothing seismically wrong with this.  In this case, was your LB score lowered or your CV or both?</p>",
      "rawMarkdown": "I agree with your reasoning in the first two paragraphs though I did not re-check the exact numbers.\n\nThe questions: \nI think the most straightforward way is to break into 150K segments but one can envision more complicated ways of handling the data - so I'm not sure about this.\nOn overlaps: It's not quite clear - are you talking about overlapping train and validation segments? If so, that is not a good thing to do as a leak is established that compromises the information provided by validation.\n\nIf you're talking about overlapping segments within a train or test data set - I don't see anything wrong with that - it resembles image zoom augmentation, but in one dimension. There is also nothing seismically wrong with this.  In this case, was your LB score lowered or your CV or both?",
      "votes": null
    },
    {
      "id": "463444",
      "postDate": "01/30/2019 01:54:39",
      "content": "<p>Here, I plot, sorted, the predicted times for a submission. The mean is low because long times are not present. I also get a stair step effect - anyone else get that?</p>",
      "rawMarkdown": "Here, I plot, sorted, the predicted times for a submission. The mean is low because long times are not present. I also get a stair step effect - anyone else get that?",
      "votes": null
    },
    {
      "id": "463459",
      "postDate": "01/30/2019 02:51:30",
      "content": "<p>Both were lower, by a fair amount. By overlap I meant taking the same time points from train to aggregate into more than one derived segment, though I admit I don’t quite understand why this would be the case (and so far I sort of doubt it is the case as many of the most useful features per segment are obtained from rolling windows). </p>\n\n<p>Mostly I’ve been following the starter kernel, segments of 150000 rows, no overlapping and starting from row 0 are used to make the ~4K segments. There are some public kernels that increase the number of segments to train on but they seem to perform worse than taking 150k, all else equal. </p>\n\n<p>As to validation, this part I find very difficult to make sense of. I’ve been following the starter here also for the most part though I’ve tried a few things. Once we aggregate the acoustic signal into segments, it’s no longer a time series problem as all we have are static features to predict a continuous target, and the question is what range and proportion of targets is appropriate. This is kind of why I’m keen on the rnn methods as they seem to preserve the temporal information though I really have no experience there so far. One thing I have not tried yet that I am planning to is taking the histogram from my best submission and using that to draw validation targets, though that comes with some obvious dangers. </p>",
      "rawMarkdown": "Both were lower, by a fair amount. By overlap I meant taking the same time points from train to aggregate into more than one derived segment, though I admit I don’t quite understand why this would be the case (and so far I sort of doubt it is the case as many of the most useful features per segment are obtained from rolling windows). \n\nMostly I’ve been following the starter kernel, segments of 150000 rows, no overlapping and starting from row 0 are used to make the ~4K segments. There are some public kernels that increase the number of segments to train on but they seem to perform worse than taking 150k, all else equal. \n\nAs to validation, this part I find very difficult to make sense of. I’ve been following the starter here also for the most part though I’ve tried a few things. Once we aggregate the acoustic signal into segments, it’s no longer a time series problem as all we have are static features to predict a continuous target, and the question is what range and proportion of targets is appropriate. This is kind of why I’m keen on the rnn methods as they seem to preserve the temporal information though I really have no experience there so far. One thing I have not tried yet that I am planning to is taking the histogram from my best submission and using that to draw validation targets, though that comes with some obvious dangers.",
      "votes": null
    },
    {
      "id": "463512",
      "postDate": "01/30/2019 05:38:40",
      "content": "<p>Your model(s) could be under-trained if you have 'early-stopping-round' turned on.\nCheck your code and see if the models are returned in primitive state.\nThis happens when your feature(s) don't generalize the data well. The validation sets sometimes score the best when the models are not really fitted yet. </p>",
      "rawMarkdown": "Your model(s) could be under-trained if you have 'early-stopping-round' turned on.\nCheck your code and see if the models are returned in primitive state.\nThis happens when your feature(s) don't generalize the data well. The validation sets sometimes score the best when the models are not really fitted yet.",
      "votes": null
    },
    {
      "id": "463516",
      "postDate": "01/30/2019 05:41:08",
      "content": "<p>Have you checked if your model(s) have converged yet? Some early-stopping-round might return you a badly fitted (not yet converged) models, which is not desired at all.</p>",
      "rawMarkdown": "Have you checked if your model(s) have converged yet? Some early-stopping-round might return you a badly fitted (not yet converged) models, which is not desired at all.",
      "votes": null
    },
    {
      "id": "463518",
      "postDate": "01/30/2019 05:45:03",
      "content": "<p>Interesting! I haven't thought about that at all. I am a novice in ML and I can't see clearly the benefit of excluding border cases from the training set. So I will split it your way and try to see the difference. <br>\nThank you, again!</p>",
      "rawMarkdown": "Interesting! I haven't thought about that at all. I am a novice in ML and I can't see clearly the benefit of excluding border cases from the training set. So I will split it your way and try to see the difference.  \nThank you, again!",
      "votes": null
    },
    {
      "id": "463537",
      "postDate": "01/30/2019 06:43:47",
      "content": "<p>I had such results when training only on quantiles of the raw signal. \nThe raw signal has steps (most values are integers between -10 and 10). \nThe 90% quantile, which is a good feature, can take only a few values, so a model trained on such features will have steps.</p>",
      "rawMarkdown": "I had such results when training only on quantiles of the raw signal. \nThe raw signal has steps (most values are integers between -10 and 10). \nThe 90% quantile, which is a good feature, can take only a few values, so a model trained on such features will have steps.",
      "votes": null
    },
    {
      "id": "463644",
      "postDate": "01/30/2019 11:02:45",
      "content": "<p>Early stopping, to me, is a must to prevent overfitting. A catboost model has training error around 1.8-1.9, CV around 2.0x, while with the same CV, my lgb model has training error to even 1.2 (even with early stopping). Underfitting is inevitable if the features are poor, whether early stopping is turned on or off. But if it’s turned off, training error can even reach 0. I can’t control the parameters to well-fit the features. </p>",
      "rawMarkdown": "Early stopping, to me, is a must to prevent overfitting. A catboost model has training error around 1.8-1.9, CV around 2.0x, while with the same CV, my lgb model has training error to even 1.2 (even with early stopping). Underfitting is inevitable if the features are poor, whether early stopping is turned on or off. But if it’s turned off, training error can even reach 0. I can’t control the parameters to well-fit the features.",
      "votes": null
    },
    {
      "id": "463730",
      "postDate": "01/30/2019 13:47:18",
      "content": "<p>Hi Kha Vo,</p>\n\n<p>Please refer to parameter \"num iteration\" in lgbm predict() function;\nand \"ntree end\" for the catboost counterpart.</p>",
      "rawMarkdown": "Hi Kha Vo,\n\nPlease refer to parameter \"num iteration\" in lgbm predict() function;\nand \"ntree end\" for the catboost counterpart.",
      "votes": null
    },
    {
      "id": "463784",
      "postDate": "01/30/2019 15:57:56",
      "content": "<p>Hi Elliot, if you use this method, how do you judge the fitting capability of your model (you use the public LB?)? If that’s the case, do you think you could extremely overfit on training set and public LB, but underfit on CV and private LB? I’m really scared of that. </p>",
      "rawMarkdown": "Hi Elliot, if you use this method, how do you judge the fitting capability of your model (you use the public LB?)? If that’s the case, do you think you could extremely overfit on training set and public LB, but underfit on CV and private LB? I’m really scared of that.",
      "votes": null
    },
    {
      "id": "463793",
      "postDate": "01/30/2019 16:09:51",
      "content": "<p>Yes, I tried not using early stopping with lgb but it badly overfit. I'll play around with the stopping criteria some more. Thanks!</p>",
      "rawMarkdown": "Yes, I tried not using early stopping with lgb but it badly overfit. I'll play around with the stopping criteria some more. Thanks!",
      "votes": null
    },
    {
      "id": "463878",
      "postDate": "01/30/2019 20:08:29",
      "content": "<p>Hi Kha Vo, calm down and relax. </p>\n\n<p>When one is doing CV, one should not use validation set for early stopping round.  The best practice is to further split the training set into a smaller training set and an early stopping set.</p>\n\n<p>Early stopping round is a direct involvement with the training process itself.  Therefore the data set used for early stopping is also a training set. </p>\n\n<p>Further more, early stopping is meant to prevent overfitting. It shouldn't hinder a model from being fitted at all in the first place. For this,  one can set a minimum/ hard threshold niterations before early stopping can set into place. </p>\n\n<p>There are many more ways to do this.  Im sure you will find the one suit your situation best.  Hope this clarifies. </p>",
      "rawMarkdown": "Hi Kha Vo, calm down and relax. \n\nWhen one is doing CV, one should not use validation set for early stopping round.  The best practice is to further split the training set into a smaller training set and an early stopping set.\n\nEarly stopping round is a direct involvement with the training process itself.  Therefore the data set used for early stopping is also a training set. \n\nFurther more, early stopping is meant to prevent overfitting. It shouldn't hinder a model from being fitted at all in the first place. For this,  one can set a minimum/ hard threshold niterations before early stopping can set into place. \n\nThere are many more ways to do this.  Im sure you will find the one suit your situation best.  Hope this clarifies.",
      "votes": null
    },
    {
      "id": "463888",
      "postDate": "01/30/2019 20:50:26",
      "content": "<p>hi interneuron, i think my reply here is related to your situation <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/78809#463784\">jump</a></p>",
      "rawMarkdown": "hi interneuron, i think my reply here is related to your situation [jump][1]\n\n\n  [1]: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/78809#463784 \"jump\"",
      "votes": null
    },
    {
      "id": "465968",
      "postDate": "02/04/2019 12:30:58",
      "content": "<p>Hi Elliot, I believe there's a problem with the advice you give about the CV strategy. If we assume that test chunks were generated randomly from an initially continuous dataset, then longer earthquake cycles would have more weight in the ultimate score. If we just average errors across earthquake cycles without weighting them by length, then we get a biased estimate.</p>",
      "rawMarkdown": "Hi Elliot, I believe there's a problem with the advice you give about the CV strategy. If we assume that test chunks were generated randomly from an initially continuous dataset, then longer earthquake cycles would have more weight in the ultimate score. If we just average errors across earthquake cycles without weighting them by length, then we get a biased estimate.",
      "votes": null
    },
    {
      "id": "466851",
      "postDate": "02/06/2019 04:14:21",
      "content": "<p>Hi KNMW,</p>\n\n<p>That is a separate issue.\nYou can tackle it by tweaking the sample weights or any other unbalanced data techniques.</p>\n\n<p>And considering the data is strongly correlated within each earthquake, it is therefore recommended not to randomize everything as a whole.</p>",
      "rawMarkdown": "Hi KNMW,\n\nThat is a separate issue.\nYou can tackle it by tweaking the sample weights or any other unbalanced data techniques.\n\nAnd considering the data is strongly correlated within each earthquake, it is therefore recommended not to randomize everything as a whole.",
      "votes": null
    },
    {
      "id": "466899",
      "postDate": "02/06/2019 07:05:29",
      "content": "<p>Between-corelation is important also, not just within-correlation.  There’s no guarantee that a high public LB score ensures a high private LB score. </p>",
      "rawMarkdown": "Between-corelation is important also, not just within-correlation.  There’s no guarantee that a high public LB score ensures a high private LB score.",
      "votes": null
    },
    {
      "id": "493301",
      "postDate": "03/18/2019 14:21:59",
      "content": "<p>This split gives very different distribution between train and CV in terms of ttf. In my case the LB seems to be quite unstable compared to CV with this split. Are you able to have a strong correlation between the CV with the split by earthquake and LB ?</p>",
      "rawMarkdown": "This split gives very different distribution between train and CV in terms of ttf. In my case the LB seems to be quite unstable compared to CV with this split. Are you able to have a strong correlation between the CV with the split by earthquake and LB ?",
      "votes": null
    },
    {
      "id": "501213",
      "postDate": "03/27/2019 03:27:03",
      "content": "<p>Sorry I get back late. Have been busy lately, haven't got a chance to look at the competition.</p>\n\n<p>As far as I can remember, I, myself, used leave-one-earth-quake out CV.\nSome earth quakes could have validation error as low as 0.7 ~ 0.8, while some others\ncould have CV scores as high as 3.7. And these two ends are no rare instances. \nConsequently, the standard error is very very high. It is so high\nto a point where all your public LB scores are legit within the probability range.</p>\n\n<p>My CV showed that short period earth quakes are usually predicted well, \nwhile the longer ones, e.g., TTF starts from 16sec, are usually predicted badly.\nIn the mean time,  we know that the public LB has a very low average TTF. \nSo if you really want to use CV as a test score predictor (instead of just a model selection criteria),\nthen maybe, just maybe, it could be more beneficial to just look at your CV scores for earthquakes having low starting TTFs.</p>\n\n<p>Hope that explains. </p>",
      "rawMarkdown": "Sorry I get back late. Have been busy lately, haven't got a chance to look at the competition.\n\nAs far as I can remember, I, myself, used leave-one-earth-quake out CV.\nSome earth quakes could have validation error as low as 0.7 ~ 0.8, while some others\ncould have CV scores as high as 3.7. And these two ends are no rare instances. \nConsequently, the standard error is very very high. It is so high\nto a point where all your public LB scores are legit within the probability range.\n\nMy CV showed that short period earth quakes are usually predicted well, \nwhile the longer ones, e.g., TTF starts from 16sec, are usually predicted badly.\nIn the mean time,  we know that the public LB has a very low average TTF. \nSo if you really want to use CV as a test score predictor (instead of just a model selection criteria),\nthen maybe, just maybe, it could be more beneficial to just look at your CV scores for earthquakes having low starting TTFs.\n\nHope that explains.",
      "votes": null
    },
    {
      "id": "503102",
      "postDate": "03/29/2019 13:16:18",
      "content": "<p>Thanks Elliot for the advice. I am curious, do you try to counter the unbalanced class problem (say for the 13 earth quakes you used to train, it has a histogram of TTF that has long tails) when you train through over/under sampling? When i tried over-weighing the under represented long TTF classes, it actually decreases my CV. Secondly, what would you say is the merit of picking a few earth quakes as CV, rather than just some randomly some sampled 150k continuous data points? (apart from fact that those used for CV in this case might have been used in training). Thank you so much for any help / advice!!</p>",
      "rawMarkdown": "Thanks Elliot for the advice. I am curious, do you try to counter the unbalanced class problem (say for the 13 earth quakes you used to train, it has a histogram of TTF that has long tails) when you train through over/under sampling? When i tried over-weighing the under represented long TTF classes, it actually decreases my CV. Secondly, what would you say is the merit of picking a few earth quakes as CV, rather than just some randomly some sampled 150k continuous data points? (apart from fact that those used for CV in this case might have been used in training). Thank you so much for any help / advice!!",
      "votes": null
    },
    {
      "id": "503352",
      "postDate": "03/29/2019 21:11:39",
      "content": "<p>Not Elliot, but I'll take a stab at it:</p>\n\n<p>There is an issue I (potentially) see with sampling to try and accommodate earthquakes with large TTFs and it's that, realistically, earthquakes with large TTFs are just noise. It's like using data collected today to try and predict an earthquake occurring in two years. </p>\n\n<p>So by over-weighing/sampling earthquakes with large TTFs (like say, over 15), you're just over-weighing noise. So it's not too surprising your model, according to CV, actually performs worse overall.</p>\n\n<p>Conversely, most models do very well on earthquakes with small TTFs. You can look through other discussions and the general gist of it is that the public LB has a much higher proportion of earthquakes with small TTFs. How much this translates to the private LB, no one knows, but it's why most everyone's CV is higher than their LB score. </p>\n\n<p>As for why to use StratifiedKFold instead of a pure KFold, I would say it's because our data is temporal (I mean, these are timed readings, so randomizing the data causes us to lose that signal) while also grouped (you can check other discussions/kernels, there are essentially 16 occurrences of where an earthquake occurred).  It's essentially a combination of 16 timed readings of an earthquake, if you will.</p>\n\n<p>By using StatifiedKFold instead of pure K-Fold, we are able to keep the temporal signals in both our training and test set, as well as maintain the grouped nature of how the data is being inputted. As Elliot mentions above, this is to give us the (theoretically) least biased estimate of our model performance.</p>",
      "rawMarkdown": "Not Elliot, but I'll take a stab at it:\n\nThere is an issue I (potentially) see with sampling to try and accommodate earthquakes with large TTFs and it's that, realistically, earthquakes with large TTFs are just noise. It's like using data collected today to try and predict an earthquake occurring in two years. \n\nSo by over-weighing/sampling earthquakes with large TTFs (like say, over 15), you're just over-weighing noise. So it's not too surprising your model, according to CV, actually performs worse overall.\n\nConversely, most models do very well on earthquakes with small TTFs. You can look through other discussions and the general gist of it is that the public LB has a much higher proportion of earthquakes with small TTFs. How much this translates to the private LB, no one knows, but it's why most everyone's CV is higher than their LB score. \n\n\nAs for why to use StratifiedKFold instead of a pure KFold, I would say it's because our data is temporal (I mean, these are timed readings, so randomizing the data causes us to lose that signal) while also grouped (you can check other discussions/kernels, there are essentially 16 occurrences of where an earthquake occurred).  It's essentially a combination of 16 timed readings of an earthquake, if you will.\n\nBy using StatifiedKFold instead of pure K-Fold, we are able to keep the temporal signals in both our training and test set, as well as maintain the grouped nature of how the data is being inputted. As Elliot mentions above, this is to give us the (theoretically) least biased estimate of our model performance.",
      "votes": null
    },
    {
      "id": "503451",
      "postDate": "03/30/2019 02:17:59",
      "content": "<p>That's awesome Conor, thanks! On the second point re K-Fold vs Stratified K Fold, if i may ask further, actually what i was wondering is: is it important that the validation data come from within a few particular earth quakes? If i use just any random 150k continuous data point (since they are continuous, it should retain any temporal info), would that work just as well for CV?</p>",
      "rawMarkdown": "That's awesome Conor, thanks! On the second point re K-Fold vs Stratified K Fold, if i may ask further, actually what i was wondering is: is it important that the validation data come from within a few particular earth quakes? If i use just any random 150k continuous data point (since they are continuous, it should retain any temporal info), would that work just as well for CV?",
      "votes": null
    },
    {
      "id": "503725",
      "postDate": "03/30/2019 13:01:07",
      "content": "<p>Thanks a lot Elliot ! I agree that this seems to be the best way to proceed. The only problem is the huge std of CV. Even with 5 folds the std of CV is still very important. Thanks for your time.</p>",
      "rawMarkdown": "Thanks a lot Elliot ! I agree that this seems to be the best way to proceed. The only problem is the huge std of CV. Even with 5 folds the std of CV is still very important. Thanks for your time.",
      "votes": null
    },
    {
      "id": "507532",
      "postDate": "04/04/2019 20:30:47",
      "content": "<p>I think it depends on how you go about with it.</p>\n\n<p>For example, let's say I take a random sample of 150k as my training set and another 150k as my test set. I could very well get an instance where my training set came AFTER my test set. In short, I'm using the future to predict the past. </p>\n\n<p>That's not very useful nor is it anything we would get in a regular application. So what it means is that our CV would be (potentially) over-optimistic about our model performance. </p>\n\n<p>If you still want to maintain some form of K-Fold Cross Validation rather than Stratified K-Fold, use something like TimeSeriesSplit. It allows you to maintain the temporal nature of the data and prevents you from using future data to predict the past. </p>\n\n<p>In short, it should provide a more accurate estimate of model performance. </p>",
      "rawMarkdown": "I think it depends on how you go about with it.\n\nFor example, let's say I take a random sample of 150k as my training set and another 150k as my test set. I could very well get an instance where my training set came AFTER my test set. In short, I'm using the future to predict the past. \n\nThat's not very useful nor is it anything we would get in a regular application. So what it means is that our CV would be (potentially) over-optimistic about our model performance. \n\nIf you still want to maintain some form of K-Fold Cross Validation rather than Stratified K-Fold, use something like TimeSeriesSplit. It allows you to maintain the temporal nature of the data and prevents you from using future data to predict the past. \n\nIn short, it should provide a more accurate estimate of model performance.",
      "votes": null
    },
    {
      "id": "507776",
      "postDate": "04/05/2019 07:26:36",
      "content": "<p><a href=\"/conormcnamara\">@conormcnamara</a> I agree that while working with time series it's better to keep validation after train. Although, in some case the bias could be very small, and I think this is one of this cases. It's not clear how a separate quake cycle should be influenced by the one that came before: changes in the experimental apparatus/settings? If so, these changes must be rather small, otherwise the whole quality of the experiment is somehow compromised and the validity of a machine learning model is weakened. </p>\n\n<p>One word on TimeSeriesSplit. It can be useful, but it's not a breakthrough compared to a simple holdout dataset, if you think about it. In order to have a competitive model, you need several earthquakes in your train set. Otherwise you will choose the parameters of your model by averaging the performances of models with large train set and small train set, which are going to be very different. This is a problem if you plan to refit on the whole dataset at the end, because at that point the splits with large train set are obviously more representative of what is a good choice of parameters. \nThe advantage of a flat k-fold is precisely that all folds are on equal footing and have the same right to contribute to the choose of hyper-parameters.</p>",
      "rawMarkdown": "conormcnamara I agree that while working with time series it's better to keep validation after train. Although, in some case the bias could be very small, and I think this is one of this cases. It's not clear how a separate quake cycle should be influenced by the one that came before: changes in the experimental apparatus/settings? If so, these changes must be rather small, otherwise the whole quality of the experiment is somehow compromised and the validity of a machine learning model is weakened. \n\nOne word on TimeSeriesSplit. It can be useful, but it's not a breakthrough compared to a simple holdout dataset, if you think about it. In order to have a competitive model, you need several earthquakes in your train set. Otherwise you will choose the parameters of your model by averaging the performances of models with large train set and small train set, which are going to be very different. This is a problem if you plan to refit on the whole dataset at the end, because at that point the splits with large train set are obviously more representative of what is a good choice of parameters. \nThe advantage of a flat k-fold is precisely that all folds are on equal footing and have the same right to contribute to the choose of hyper-parameters.",
      "votes": null
    },
    {
      "id": "508217",
      "postDate": "04/05/2019 19:56:09",
      "content": "<p>It's interesting you mentioned the separate quake cycle. If I recall, recent research has indicated that <a href=\"https://www.npr.org/2013/08/23/214619037/can-a-big-earthquake-trigger-another-one\">large earthquakes actually increase the likelihood of subsequent earthquakes, even along different fault lines</a>. Here's another<a href=\"https://www.sciencedaily.com/releases/2018/08/180802102352.htm\"> paper</a> that discusses similar things. Now, whether the experiments here even remotely model that, I don't know. But, assuming we want this to apply to real world applications, it is something we should take into consideration. </p>\n\n<p>As for Time Series Split, one way to mitigate this problem is to instead do a Rolling Time Series Split, where each validation chunk is the same. I use this one a lot when working with large Time Series datasets and don't have the resources to handle the final step of TimeSeriesSplit.</p>\n\n<p>For example, suppose that I split my dataset into 5 \"chunks\".\nThen for validation step 1: chunk 1 is used to predict chunk 2.\nFor validation step 2: ONLY chunk 2 is used to predict chunk 3....</p>\n\n<p>Doing this, you should roughly have 3 quakes worth of data in each chunk. Alternatively (due to the small dataset), you can probably drop this down to 3 chunks to get roughly 5 quakes worth of data in each chunk. </p>\n\n<p>You can specify this in TimeSeriesSplit using max_train_size.</p>\n\n<p>As an alternative alternative, you could use <a href=\"https://www.sciencedirect.com/science/article/pii/S0304407600000300\">hv-block </a>cross validation as described in the paper. I have personally never used it so can't point to the efficacy of it. </p>",
      "rawMarkdown": "It's interesting you mentioned the separate quake cycle. If I recall, recent research has indicated that [large earthquakes actually increase the likelihood of subsequent earthquakes, even along different fault lines](https://www.npr.org/2013/08/23/214619037/can-a-big-earthquake-trigger-another-one). Here's another[ paper](https://www.sciencedaily.com/releases/2018/08/180802102352.htm) that discusses similar things. Now, whether the experiments here even remotely model that, I don't know. But, assuming we want this to apply to real world applications, it is something we should take into consideration. \n\nAs for Time Series Split, one way to mitigate this problem is to instead do a Rolling Time Series Split, where each validation chunk is the same. I use this one a lot when working with large Time Series datasets and don't have the resources to handle the final step of TimeSeriesSplit.\n\nFor example, suppose that I split my dataset into 5 \"chunks\".\nThen for validation step 1: chunk 1 is used to predict chunk 2.\nFor validation step 2: ONLY chunk 2 is used to predict chunk 3....\n\nDoing this, you should roughly have 3 quakes worth of data in each chunk. Alternatively (due to the small dataset), you can probably drop this down to 3 chunks to get roughly 5 quakes worth of data in each chunk. \n\nYou can specify this in TimeSeriesSplit using max_train_size.\n\n\nAs an alternative alternative, you could use [hv-block ](https://www.sciencedirect.com/science/article/pii/S0304407600000300)cross validation as described in the paper. I have personally never used it so can't point to the efficacy of it.",
      "votes": null
    },
    {
      "id": "2066735",
      "postDate": "12/16/2022 02:51:09",
      "content": "<p>Thanks for Your Discussion</p>",
      "rawMarkdown": "Thanks for Your Discussion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2066735,
      "author_name": "mjbyeon",
      "author_url": "",
      "post_date": "12/16/2022 02:51:09",
      "content": "<p>Thanks for Your Discussion</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 462472,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "01/28/2019 10:47:09",
      "content": "<p>Thank you for the advice. However, depending on the validation fold size (number of earthquakes in the validation set) I get CV scores between 1.8-2.1. These best submissions are around 1.65. </p>\n\n<p>I'm not sure how many earthquakes should there be in validation sets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 462813,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/29/2019 00:05:53",
          "content": "<p>Hi David,\nWhat CV can offer is a relative benchmark for you to compare among the models.\nAs for how accurate they can be served as an direct estimate on the test dataset?\nThat would depend on the similarity between the distributions're offered.\nIn our case, the mean value for the public LB is only 4.0xxx. This is completely different from our training set. So we will need to have some other methods to give the desired result.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463184,
          "author_name": "davids1992",
          "author_url": "",
          "post_date": "01/29/2019 15:24:55",
          "content": "<p>That's interesting, I didn't notice the train mean value is ~5.6. Thx!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 462714,
      "author_name": "deepdreamx",
      "author_url": "",
      "post_date": "01/28/2019 18:55:55",
      "content": "<p>But i have a question. I was seeing a test data segment and the acoustic data of one segment had a high acoustic signal at start that means probably earthquake occurred at start. This shows 1 segment can have data of 2 earthquakes. This CV technique would work in such case? </p>",
      "votes": null,
      "replies": [
        {
          "id": 462732,
          "author_name": "petya5q",
          "author_url": "",
          "post_date": "01/28/2019 19:42:53",
          "content": "<p>I am not sure we can say that there are 2 earthquakes. High peak in the acoustic data is not always followed by an earthquake (resetting the time) . Also high peak itself is not the where the actual earthquake is(at least not in the train file), it is hidden somewhere after the peak. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462819,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/29/2019 00:17:17",
          "content": "<p>Hi Ammad,\nFirstly, in the data preparation phase, I would recommend removing segments that cross borders. It is still fine if you keep them. Their number is tiny anyway.</p>\n\n<p>And second, I wouldn't decide every high peak to be near an earthquake. This is the most challenging part in this competition. If you can check out the 3rd, 8th and 15th earthquake in our training data, you will find high peaks everywhere even in the early stages.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 462736,
      "author_name": "petya5q",
      "author_url": "",
      "post_date": "01/28/2019 19:53:24",
      "content": "<p>Thank you Elliot and DavidSfor sharing your ideas!\nIf we split it in Elliot's way we will have proportionally eq /no eq, both in train and in validation.</p>\n\n<p>Overlapping of the batches will insert same sequence once at the beginning and once at the end  of every train/test batch. But I can't say what the effect on the model will be. Is it too much if we have several different CV techniques? </p>",
      "votes": null,
      "replies": [
        {
          "id": 462814,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/29/2019 00:09:15",
          "content": "<p>I would say this belongs to data cleaning issue :)\nIn my script, I don't allow segments that cross the borders.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463518,
          "author_name": "petya5q",
          "author_url": "",
          "post_date": "01/30/2019 05:45:03",
          "content": "<p>Interesting! I haven't thought about that at all. I am a novice in ML and I can't see clearly the benefit of excluding border cases from the training set. So I will split it your way and try to see the difference. <br>\nThank you, again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 462767,
      "author_name": "mayer79",
      "author_url": "",
      "post_date": "01/28/2019 21:21:33",
      "content": "<p>Thx for your input Elliot. I am also using this exact strategy, maybe with different groupings. One thing I am generally (not limited to this strategy) not happy about is that the fold containing the longest chunk will automatically get a bad CV score, at least if working with decision trees. They cannot extrapolate without tweaks. So maybe it might also be an option to make e.g. 50 chunks of length 12 mio each and then build the folds by randomly regrouping them into five folds to partially fight this problem. </p>\n\n<p>Happy for any idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 462809,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/28/2019 23:58:01",
          "content": "<p>Hi Micheal,</p>\n\n<p>If those edge values are really of concern, we can just have them always grouped in the training set.\nThe current implementation of the boosting trees don't have higher order basis implemented,\nso we can't combat the extrapolation problem head-on, but to contain them.</p>\n\n<p>In our case, however, the extrapolation is far from the top concern.\nThe bad CV score for the longest chunk, is mostly, due to the systematically under-estimated 'time to failure'.</p>\n\n<p>These are my thoughts to your response :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 462899,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "01/29/2019 04:36:50",
      "content": "<p>Hi All,</p>\n\n<p>I'm just curious as to how close people's CV scores are to the public leaderboard. I'm getting OK CV but I assume it will blow up when I submit?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463279,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "01/29/2019 18:25:44",
      "content": "<p>One thing to keep in mind is that leaderboard climbing is a beguiling as the sirens in Greek mythology - I can guarantee that a lot of people's best entries will be about 30 submissions below their final one!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463360,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "01/29/2019 21:26:42",
      "content": "<p>Hi Elliot, thanks for this advice. I have not had much success with parsing the raw signal in ways other than demonstrated in the starter kernel yet, but I will try again. Your comment makes me wonder a few things and I am curious about the logic those with more knowledge of geoscience and physics than me (probably everyone) are using.</p>\n\n<p>First, it is interesting that the training data is one continuous signal which we can presumably chop up in any way we like. Assuming what I've read is true about the sampling frequency, the training signal consists of 157.275 seconds with 16 points where time to failure reaches zero, i.e., a labquake has occurred (and this timing was acquired by a different device than was used to obtain the acoustic data and which is not available to us).  It seems rational to begin by cutting the training signal to resemble the test set, as the starter did, yielding ~4200 segments of 150000 rows each comprising 37.5 milliseconds where the time to failure (target) is the last row of each segment. Now, we are evaluated on the predicted time to failure for each of the ~2k segments of 37ms of acoustic signal -- since the predictions are in seconds I am assuming there are likely no quakes in the test data, but I cannot say that for a certainty. My feeling is it does not matter if there are quakes in test or not, since we are not asked to identify quakes per se but how much time remains before a likely quake given current sensor data.</p>\n\n<p>My questions, so far, are these. Holding out one or two quake cycles from train to use as validation still requires breaking up into test-length segments? I agree that using overlapping segments it probably a bad idea and when I've tried it my score is way worse, I'm wondering if you think this an artifact of how we are being scored or if this is common knowledge for researchers working with seismic data? I'll stop here for now and thanks again for the insights. Best of luck to you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 463432,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "01/30/2019 01:18:37",
          "content": "<p>I agree with your reasoning in the first two paragraphs though I did not re-check the exact numbers.</p>\n\n<p>The questions: \nI think the most straightforward way is to break into 150K segments but one can envision more complicated ways of handling the data - so I'm not sure about this.\nOn overlaps: It's not quite clear - are you talking about overlapping train and validation segments? If so, that is not a good thing to do as a leak is established that compromises the information provided by validation.</p>\n\n<p>If you're talking about overlapping segments within a train or test data set - I don't see anything wrong with that - it resembles image zoom augmentation, but in one dimension. There is also nothing seismically wrong with this.  In this case, was your LB score lowered or your CV or both?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463459,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "01/30/2019 02:51:30",
          "content": "<p>Both were lower, by a fair amount. By overlap I meant taking the same time points from train to aggregate into more than one derived segment, though I admit I don’t quite understand why this would be the case (and so far I sort of doubt it is the case as many of the most useful features per segment are obtained from rolling windows). </p>\n\n<p>Mostly I’ve been following the starter kernel, segments of 150000 rows, no overlapping and starting from row 0 are used to make the ~4K segments. There are some public kernels that increase the number of segments to train on but they seem to perform worse than taking 150k, all else equal. </p>\n\n<p>As to validation, this part I find very difficult to make sense of. I’ve been following the starter here also for the most part though I’ve tried a few things. Once we aggregate the acoustic signal into segments, it’s no longer a time series problem as all we have are static features to predict a continuous target, and the question is what range and proportion of targets is appropriate. This is kind of why I’m keen on the rnn methods as they seem to preserve the temporal information though I really have no experience there so far. One thing I have not tried yet that I am planning to is taking the histogram from my best submission and using that to draw validation targets, though that comes with some obvious dangers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463516,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/30/2019 05:41:08",
          "content": "<p>Have you checked if your model(s) have converged yet? Some early-stopping-round might return you a badly fitted (not yet converged) models, which is not desired at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463793,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "01/30/2019 16:09:51",
          "content": "<p>Yes, I tried not using early stopping with lgb but it badly overfit. I'll play around with the stopping criteria some more. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463888,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/30/2019 20:50:26",
          "content": "<p>hi interneuron, i think my reply here is related to your situation <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/78809#463784\">jump</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 463444,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "01/30/2019 01:54:39",
      "content": "<p>Here, I plot, sorted, the predicted times for a submission. The mean is low because long times are not present. I also get a stair step effect - anyone else get that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 463512,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/30/2019 05:38:40",
          "content": "<p>Your model(s) could be under-trained if you have 'early-stopping-round' turned on.\nCheck your code and see if the models are returned in primitive state.\nThis happens when your feature(s) don't generalize the data well. The validation sets sometimes score the best when the models are not really fitted yet. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463537,
          "author_name": "zidmie",
          "author_url": "",
          "post_date": "01/30/2019 06:43:47",
          "content": "<p>I had such results when training only on quantiles of the raw signal. \nThe raw signal has steps (most values are integers between -10 and 10). \nThe 90% quantile, which is a good feature, can take only a few values, so a model trained on such features will have steps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463644,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "01/30/2019 11:02:45",
          "content": "<p>Early stopping, to me, is a must to prevent overfitting. A catboost model has training error around 1.8-1.9, CV around 2.0x, while with the same CV, my lgb model has training error to even 1.2 (even with early stopping). Underfitting is inevitable if the features are poor, whether early stopping is turned on or off. But if it’s turned off, training error can even reach 0. I can’t control the parameters to well-fit the features. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463730,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/30/2019 13:47:18",
          "content": "<p>Hi Kha Vo,</p>\n\n<p>Please refer to parameter \"num iteration\" in lgbm predict() function;\nand \"ntree end\" for the catboost counterpart.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463784,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "01/30/2019 15:57:56",
          "content": "<p>Hi Elliot, if you use this method, how do you judge the fitting capability of your model (you use the public LB?)? If that’s the case, do you think you could extremely overfit on training set and public LB, but underfit on CV and private LB? I’m really scared of that. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 463878,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "01/30/2019 20:08:29",
          "content": "<p>Hi Kha Vo, calm down and relax. </p>\n\n<p>When one is doing CV, one should not use validation set for early stopping round.  The best practice is to further split the training set into a smaller training set and an early stopping set.</p>\n\n<p>Early stopping round is a direct involvement with the training process itself.  Therefore the data set used for early stopping is also a training set. </p>\n\n<p>Further more, early stopping is meant to prevent overfitting. It shouldn't hinder a model from being fitted at all in the first place. For this,  one can set a minimum/ hard threshold niterations before early stopping can set into place. </p>\n\n<p>There are many more ways to do this.  Im sure you will find the one suit your situation best.  Hope this clarifies. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465968,
      "author_name": "kostyanomatterwhat",
      "author_url": "",
      "post_date": "02/04/2019 12:30:58",
      "content": "<p>Hi Elliot, I believe there's a problem with the advice you give about the CV strategy. If we assume that test chunks were generated randomly from an initially continuous dataset, then longer earthquake cycles would have more weight in the ultimate score. If we just average errors across earthquake cycles without weighting them by length, then we get a biased estimate.</p>",
      "votes": null,
      "replies": [
        {
          "id": 466851,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "02/06/2019 04:14:21",
          "content": "<p>Hi KNMW,</p>\n\n<p>That is a separate issue.\nYou can tackle it by tweaking the sample weights or any other unbalanced data techniques.</p>\n\n<p>And considering the data is strongly correlated within each earthquake, it is therefore recommended not to randomize everything as a whole.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466899,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "02/06/2019 07:05:29",
          "content": "<p>Between-corelation is important also, not just within-correlation.  There’s no guarantee that a high public LB score ensures a high private LB score. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 493301,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "03/18/2019 14:21:59",
      "content": "<p>This split gives very different distribution between train and CV in terms of ttf. In my case the LB seems to be quite unstable compared to CV with this split. Are you able to have a strong correlation between the CV with the split by earthquake and LB ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 501213,
          "author_name": "tclf90",
          "author_url": "",
          "post_date": "03/27/2019 03:27:03",
          "content": "<p>Sorry I get back late. Have been busy lately, haven't got a chance to look at the competition.</p>\n\n<p>As far as I can remember, I, myself, used leave-one-earth-quake out CV.\nSome earth quakes could have validation error as low as 0.7 ~ 0.8, while some others\ncould have CV scores as high as 3.7. And these two ends are no rare instances. \nConsequently, the standard error is very very high. It is so high\nto a point where all your public LB scores are legit within the probability range.</p>\n\n<p>My CV showed that short period earth quakes are usually predicted well, \nwhile the longer ones, e.g., TTF starts from 16sec, are usually predicted badly.\nIn the mean time,  we know that the public LB has a very low average TTF. \nSo if you really want to use CV as a test score predictor (instead of just a model selection criteria),\nthen maybe, just maybe, it could be more beneficial to just look at your CV scores for earthquakes having low starting TTFs.</p>\n\n<p>Hope that explains. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503725,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "03/30/2019 13:01:07",
          "content": "<p>Thanks a lot Elliot ! I agree that this seems to be the best way to proceed. The only problem is the huge std of CV. Even with 5 folds the std of CV is still very important. Thanks for your time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 503102,
      "author_name": "heisenger",
      "author_url": "",
      "post_date": "03/29/2019 13:16:18",
      "content": "<p>Thanks Elliot for the advice. I am curious, do you try to counter the unbalanced class problem (say for the 13 earth quakes you used to train, it has a histogram of TTF that has long tails) when you train through over/under sampling? When i tried over-weighing the under represented long TTF classes, it actually decreases my CV. Secondly, what would you say is the merit of picking a few earth quakes as CV, rather than just some randomly some sampled 150k continuous data points? (apart from fact that those used for CV in this case might have been used in training). Thank you so much for any help / advice!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 503352,
          "author_name": "conormcnamara",
          "author_url": "",
          "post_date": "03/29/2019 21:11:39",
          "content": "<p>Not Elliot, but I'll take a stab at it:</p>\n\n<p>There is an issue I (potentially) see with sampling to try and accommodate earthquakes with large TTFs and it's that, realistically, earthquakes with large TTFs are just noise. It's like using data collected today to try and predict an earthquake occurring in two years. </p>\n\n<p>So by over-weighing/sampling earthquakes with large TTFs (like say, over 15), you're just over-weighing noise. So it's not too surprising your model, according to CV, actually performs worse overall.</p>\n\n<p>Conversely, most models do very well on earthquakes with small TTFs. You can look through other discussions and the general gist of it is that the public LB has a much higher proportion of earthquakes with small TTFs. How much this translates to the private LB, no one knows, but it's why most everyone's CV is higher than their LB score. </p>\n\n<p>As for why to use StratifiedKFold instead of a pure KFold, I would say it's because our data is temporal (I mean, these are timed readings, so randomizing the data causes us to lose that signal) while also grouped (you can check other discussions/kernels, there are essentially 16 occurrences of where an earthquake occurred).  It's essentially a combination of 16 timed readings of an earthquake, if you will.</p>\n\n<p>By using StatifiedKFold instead of pure K-Fold, we are able to keep the temporal signals in both our training and test set, as well as maintain the grouped nature of how the data is being inputted. As Elliot mentions above, this is to give us the (theoretically) least biased estimate of our model performance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503451,
          "author_name": "heisenger",
          "author_url": "",
          "post_date": "03/30/2019 02:17:59",
          "content": "<p>That's awesome Conor, thanks! On the second point re K-Fold vs Stratified K Fold, if i may ask further, actually what i was wondering is: is it important that the validation data come from within a few particular earth quakes? If i use just any random 150k continuous data point (since they are continuous, it should retain any temporal info), would that work just as well for CV?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 507532,
          "author_name": "conormcnamara",
          "author_url": "",
          "post_date": "04/04/2019 20:30:47",
          "content": "<p>I think it depends on how you go about with it.</p>\n\n<p>For example, let's say I take a random sample of 150k as my training set and another 150k as my test set. I could very well get an instance where my training set came AFTER my test set. In short, I'm using the future to predict the past. </p>\n\n<p>That's not very useful nor is it anything we would get in a regular application. So what it means is that our CV would be (potentially) over-optimistic about our model performance. </p>\n\n<p>If you still want to maintain some form of K-Fold Cross Validation rather than Stratified K-Fold, use something like TimeSeriesSplit. It allows you to maintain the temporal nature of the data and prevents you from using future data to predict the past. </p>\n\n<p>In short, it should provide a more accurate estimate of model performance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 507776,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "04/05/2019 07:26:36",
          "content": "<p><a href=\"/conormcnamara\">@conormcnamara</a> I agree that while working with time series it's better to keep validation after train. Although, in some case the bias could be very small, and I think this is one of this cases. It's not clear how a separate quake cycle should be influenced by the one that came before: changes in the experimental apparatus/settings? If so, these changes must be rather small, otherwise the whole quality of the experiment is somehow compromised and the validity of a machine learning model is weakened. </p>\n\n<p>One word on TimeSeriesSplit. It can be useful, but it's not a breakthrough compared to a simple holdout dataset, if you think about it. In order to have a competitive model, you need several earthquakes in your train set. Otherwise you will choose the parameters of your model by averaging the performances of models with large train set and small train set, which are going to be very different. This is a problem if you plan to refit on the whole dataset at the end, because at that point the splits with large train set are obviously more representative of what is a good choice of parameters. \nThe advantage of a flat k-fold is precisely that all folds are on equal footing and have the same right to contribute to the choose of hyper-parameters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 508217,
          "author_name": "conormcnamara",
          "author_url": "",
          "post_date": "04/05/2019 19:56:09",
          "content": "<p>It's interesting you mentioned the separate quake cycle. If I recall, recent research has indicated that <a href=\"https://www.npr.org/2013/08/23/214619037/can-a-big-earthquake-trigger-another-one\">large earthquakes actually increase the likelihood of subsequent earthquakes, even along different fault lines</a>. Here's another<a href=\"https://www.sciencedaily.com/releases/2018/08/180802102352.htm\"> paper</a> that discusses similar things. Now, whether the experiments here even remotely model that, I don't know. But, assuming we want this to apply to real world applications, it is something we should take into consideration. </p>\n\n<p>As for Time Series Split, one way to mitigate this problem is to instead do a Rolling Time Series Split, where each validation chunk is the same. I use this one a lot when working with large Time Series datasets and don't have the resources to handle the final step of TimeSeriesSplit.</p>\n\n<p>For example, suppose that I split my dataset into 5 \"chunks\".\nThen for validation step 1: chunk 1 is used to predict chunk 2.\nFor validation step 2: ONLY chunk 2 is used to predict chunk 3....</p>\n\n<p>Doing this, you should roughly have 3 quakes worth of data in each chunk. Alternatively (due to the small dataset), you can probably drop this down to 3 chunks to get roughly 5 quakes worth of data in each chunk. </p>\n\n<p>You can specify this in TimeSeriesSplit using max_train_size.</p>\n\n<p>As an alternative alternative, you could use <a href=\"https://www.sciencedirect.com/science/article/pii/S0304407600000300\">hv-block </a>cross validation as described in the paper. I have personally never used it so can't point to the efficacy of it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "462341": "In order to get a less biased CV, especially if we augment the data,\n\nTry to split by the earthquake ID ( 16 or 15 in total, however you define)\n\nFor example, for a 5 fold CV,\nFold 1: use the first 13 earthquakes to train, and the last 3 to validate\n etc.",
    "462472": "Thank you for the advice. However, depending on the validation fold size (number of earthquakes in the validation set) I get CV scores between 1.8-2.1. These best submissions are around 1.65. \n\nI'm not sure how many earthquakes should there be in validation sets.",
    "462714": "But i have a question. I was seeing a test data segment and the acoustic data of one segment had a high acoustic signal at start that means probably earthquake occurred at start. This shows 1 segment can have data of 2 earthquakes. This CV technique would work in such case?",
    "462732": "I am not sure we can say that there are 2 earthquakes. High peak in the acoustic data is not always followed by an earthquake (resetting the time) . Also high peak itself is not the where the actual earthquake is(at least not in the train file), it is hidden somewhere after the peak.",
    "462736": "Thank you Elliot and DavidSfor sharing your ideas!\nIf we split it in Elliot's way we will have proportionally eq /no eq, both in train and in validation.\n\nOverlapping of the batches will insert same sequence once at the beginning and once at the end  of every train/test batch. But I can't say what the effect on the model will be. Is it too much if we have several different CV techniques?",
    "462767": "Thx for your input Elliot. I am also using this exact strategy, maybe with different groupings. One thing I am generally (not limited to this strategy) not happy about is that the fold containing the longest chunk will automatically get a bad CV score, at least if working with decision trees. They cannot extrapolate without tweaks. So maybe it might also be an option to make e.g. 50 chunks of length 12 mio each and then build the folds by randomly regrouping them into five folds to partially fight this problem. \n\nHappy for any idea.",
    "462809": "Hi Micheal,\n\nIf those edge values are really of concern, we can just have them always grouped in the training set.\nThe current implementation of the boosting trees don't have higher order basis implemented,\nso we can't combat the extrapolation problem head-on, but to contain them.\n\nIn our case, however, the extrapolation is far from the top concern.\nThe bad CV score for the longest chunk, is mostly, due to the systematically under-estimated 'time to failure'.\n\nThese are my thoughts to your response :)",
    "462813": "Hi David,\nWhat CV can offer is a relative benchmark for you to compare among the models.\nAs for how accurate they can be served as an direct estimate on the test dataset?\nThat would depend on the similarity between the distributions're offered.\nIn our case, the mean value for the public LB is only 4.0xxx. This is completely different from our training set. So we will need to have some other methods to give the desired result.",
    "462814": "I would say this belongs to data cleaning issue :)\nIn my script, I don't allow segments that cross the borders.",
    "462819": "Hi Ammad,\nFirstly, in the data preparation phase, I would recommend removing segments that cross borders. It is still fine if you keep them. Their number is tiny anyway.\n\nAnd second, I wouldn't decide every high peak to be near an earthquake. This is the most challenging part in this competition. If you can check out the 3rd, 8th and 15th earthquake in our training data, you will find high peaks everywhere even in the early stages.",
    "462899": "Hi All,\n\nI'm just curious as to how close people's CV scores are to the public leaderboard. I'm getting OK CV but I assume it will blow up when I submit?",
    "463184": "That's interesting, I didn't notice the train mean value is ~5.6. Thx!",
    "463279": "One thing to keep in mind is that leaderboard climbing is a beguiling as the sirens in Greek mythology - I can guarantee that a lot of people's best entries will be about 30 submissions below their final one!",
    "463360": "Hi Elliot, thanks for this advice. I have not had much success with parsing the raw signal in ways other than demonstrated in the starter kernel yet, but I will try again. Your comment makes me wonder a few things and I am curious about the logic those with more knowledge of geoscience and physics than me (probably everyone) are using.\n\nFirst, it is interesting that the training data is one continuous signal which we can presumably chop up in any way we like. Assuming what I've read is true about the sampling frequency, the training signal consists of 157.275 seconds with 16 points where time to failure reaches zero, i.e., a labquake has occurred (and this timing was acquired by a different device than was used to obtain the acoustic data and which is not available to us).  It seems rational to begin by cutting the training signal to resemble the test set, as the starter did, yielding ~4200 segments of 150000 rows each comprising 37.5 milliseconds where the time to failure (target) is the last row of each segment. Now, we are evaluated on the predicted time to failure for each of the ~2k segments of 37ms of acoustic signal -- since the predictions are in seconds I am assuming there are likely no quakes in the test data, but I cannot say that for a certainty. My feeling is it does not matter if there are quakes in test or not, since we are not asked to identify quakes per se but how much time remains before a likely quake given current sensor data.\n\nMy questions, so far, are these. Holding out one or two quake cycles from train to use as validation still requires breaking up into test-length segments? I agree that using overlapping segments it probably a bad idea and when I've tried it my score is way worse, I'm wondering if you think this an artifact of how we are being scored or if this is common knowledge for researchers working with seismic data? I'll stop here for now and thanks again for the insights. Best of luck to you!",
    "463432": "I agree with your reasoning in the first two paragraphs though I did not re-check the exact numbers.\n\nThe questions: \nI think the most straightforward way is to break into 150K segments but one can envision more complicated ways of handling the data - so I'm not sure about this.\nOn overlaps: It's not quite clear - are you talking about overlapping train and validation segments? If so, that is not a good thing to do as a leak is established that compromises the information provided by validation.\n\nIf you're talking about overlapping segments within a train or test data set - I don't see anything wrong with that - it resembles image zoom augmentation, but in one dimension. There is also nothing seismically wrong with this.  In this case, was your LB score lowered or your CV or both?",
    "463444": "Here, I plot, sorted, the predicted times for a submission. The mean is low because long times are not present. I also get a stair step effect - anyone else get that?",
    "463459": "Both were lower, by a fair amount. By overlap I meant taking the same time points from train to aggregate into more than one derived segment, though I admit I don’t quite understand why this would be the case (and so far I sort of doubt it is the case as many of the most useful features per segment are obtained from rolling windows). \n\nMostly I’ve been following the starter kernel, segments of 150000 rows, no overlapping and starting from row 0 are used to make the ~4K segments. There are some public kernels that increase the number of segments to train on but they seem to perform worse than taking 150k, all else equal. \n\nAs to validation, this part I find very difficult to make sense of. I’ve been following the starter here also for the most part though I’ve tried a few things. Once we aggregate the acoustic signal into segments, it’s no longer a time series problem as all we have are static features to predict a continuous target, and the question is what range and proportion of targets is appropriate. This is kind of why I’m keen on the rnn methods as they seem to preserve the temporal information though I really have no experience there so far. One thing I have not tried yet that I am planning to is taking the histogram from my best submission and using that to draw validation targets, though that comes with some obvious dangers.",
    "463512": "Your model(s) could be under-trained if you have 'early-stopping-round' turned on.\nCheck your code and see if the models are returned in primitive state.\nThis happens when your feature(s) don't generalize the data well. The validation sets sometimes score the best when the models are not really fitted yet.",
    "463516": "Have you checked if your model(s) have converged yet? Some early-stopping-round might return you a badly fitted (not yet converged) models, which is not desired at all.",
    "463518": "Interesting! I haven't thought about that at all. I am a novice in ML and I can't see clearly the benefit of excluding border cases from the training set. So I will split it your way and try to see the difference.  \nThank you, again!",
    "463537": "I had such results when training only on quantiles of the raw signal. \nThe raw signal has steps (most values are integers between -10 and 10). \nThe 90% quantile, which is a good feature, can take only a few values, so a model trained on such features will have steps.",
    "463644": "Early stopping, to me, is a must to prevent overfitting. A catboost model has training error around 1.8-1.9, CV around 2.0x, while with the same CV, my lgb model has training error to even 1.2 (even with early stopping). Underfitting is inevitable if the features are poor, whether early stopping is turned on or off. But if it’s turned off, training error can even reach 0. I can’t control the parameters to well-fit the features.",
    "463730": "Hi Kha Vo,\n\nPlease refer to parameter \"num iteration\" in lgbm predict() function;\nand \"ntree end\" for the catboost counterpart.",
    "463784": "Hi Elliot, if you use this method, how do you judge the fitting capability of your model (you use the public LB?)? If that’s the case, do you think you could extremely overfit on training set and public LB, but underfit on CV and private LB? I’m really scared of that.",
    "463793": "Yes, I tried not using early stopping with lgb but it badly overfit. I'll play around with the stopping criteria some more. Thanks!",
    "463878": "Hi Kha Vo, calm down and relax. \n\nWhen one is doing CV, one should not use validation set for early stopping round.  The best practice is to further split the training set into a smaller training set and an early stopping set.\n\nEarly stopping round is a direct involvement with the training process itself.  Therefore the data set used for early stopping is also a training set. \n\nFurther more, early stopping is meant to prevent overfitting. It shouldn't hinder a model from being fitted at all in the first place. For this,  one can set a minimum/ hard threshold niterations before early stopping can set into place. \n\nThere are many more ways to do this.  Im sure you will find the one suit your situation best.  Hope this clarifies.",
    "463888": "hi interneuron, i think my reply here is related to your situation [jump][1]\n\n\n  [1]: https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/78809#463784 \"jump\"",
    "465968": "Hi Elliot, I believe there's a problem with the advice you give about the CV strategy. If we assume that test chunks were generated randomly from an initially continuous dataset, then longer earthquake cycles would have more weight in the ultimate score. If we just average errors across earthquake cycles without weighting them by length, then we get a biased estimate.",
    "466851": "Hi KNMW,\n\nThat is a separate issue.\nYou can tackle it by tweaking the sample weights or any other unbalanced data techniques.\n\nAnd considering the data is strongly correlated within each earthquake, it is therefore recommended not to randomize everything as a whole.",
    "466899": "Between-corelation is important also, not just within-correlation.  There’s no guarantee that a high public LB score ensures a high private LB score.",
    "493301": "This split gives very different distribution between train and CV in terms of ttf. In my case the LB seems to be quite unstable compared to CV with this split. Are you able to have a strong correlation between the CV with the split by earthquake and LB ?",
    "501213": "Sorry I get back late. Have been busy lately, haven't got a chance to look at the competition.\n\nAs far as I can remember, I, myself, used leave-one-earth-quake out CV.\nSome earth quakes could have validation error as low as 0.7 ~ 0.8, while some others\ncould have CV scores as high as 3.7. And these two ends are no rare instances. \nConsequently, the standard error is very very high. It is so high\nto a point where all your public LB scores are legit within the probability range.\n\nMy CV showed that short period earth quakes are usually predicted well, \nwhile the longer ones, e.g., TTF starts from 16sec, are usually predicted badly.\nIn the mean time,  we know that the public LB has a very low average TTF. \nSo if you really want to use CV as a test score predictor (instead of just a model selection criteria),\nthen maybe, just maybe, it could be more beneficial to just look at your CV scores for earthquakes having low starting TTFs.\n\nHope that explains.",
    "503102": "Thanks Elliot for the advice. I am curious, do you try to counter the unbalanced class problem (say for the 13 earth quakes you used to train, it has a histogram of TTF that has long tails) when you train through over/under sampling? When i tried over-weighing the under represented long TTF classes, it actually decreases my CV. Secondly, what would you say is the merit of picking a few earth quakes as CV, rather than just some randomly some sampled 150k continuous data points? (apart from fact that those used for CV in this case might have been used in training). Thank you so much for any help / advice!!",
    "503352": "Not Elliot, but I'll take a stab at it:\n\nThere is an issue I (potentially) see with sampling to try and accommodate earthquakes with large TTFs and it's that, realistically, earthquakes with large TTFs are just noise. It's like using data collected today to try and predict an earthquake occurring in two years. \n\nSo by over-weighing/sampling earthquakes with large TTFs (like say, over 15), you're just over-weighing noise. So it's not too surprising your model, according to CV, actually performs worse overall.\n\nConversely, most models do very well on earthquakes with small TTFs. You can look through other discussions and the general gist of it is that the public LB has a much higher proportion of earthquakes with small TTFs. How much this translates to the private LB, no one knows, but it's why most everyone's CV is higher than their LB score. \n\n\nAs for why to use StratifiedKFold instead of a pure KFold, I would say it's because our data is temporal (I mean, these are timed readings, so randomizing the data causes us to lose that signal) while also grouped (you can check other discussions/kernels, there are essentially 16 occurrences of where an earthquake occurred).  It's essentially a combination of 16 timed readings of an earthquake, if you will.\n\nBy using StatifiedKFold instead of pure K-Fold, we are able to keep the temporal signals in both our training and test set, as well as maintain the grouped nature of how the data is being inputted. As Elliot mentions above, this is to give us the (theoretically) least biased estimate of our model performance.",
    "503451": "That's awesome Conor, thanks! On the second point re K-Fold vs Stratified K Fold, if i may ask further, actually what i was wondering is: is it important that the validation data come from within a few particular earth quakes? If i use just any random 150k continuous data point (since they are continuous, it should retain any temporal info), would that work just as well for CV?",
    "503725": "Thanks a lot Elliot ! I agree that this seems to be the best way to proceed. The only problem is the huge std of CV. Even with 5 folds the std of CV is still very important. Thanks for your time.",
    "507532": "I think it depends on how you go about with it.\n\nFor example, let's say I take a random sample of 150k as my training set and another 150k as my test set. I could very well get an instance where my training set came AFTER my test set. In short, I'm using the future to predict the past. \n\nThat's not very useful nor is it anything we would get in a regular application. So what it means is that our CV would be (potentially) over-optimistic about our model performance. \n\nIf you still want to maintain some form of K-Fold Cross Validation rather than Stratified K-Fold, use something like TimeSeriesSplit. It allows you to maintain the temporal nature of the data and prevents you from using future data to predict the past. \n\nIn short, it should provide a more accurate estimate of model performance.",
    "507776": "conormcnamara I agree that while working with time series it's better to keep validation after train. Although, in some case the bias could be very small, and I think this is one of this cases. It's not clear how a separate quake cycle should be influenced by the one that came before: changes in the experimental apparatus/settings? If so, these changes must be rather small, otherwise the whole quality of the experiment is somehow compromised and the validity of a machine learning model is weakened. \n\nOne word on TimeSeriesSplit. It can be useful, but it's not a breakthrough compared to a simple holdout dataset, if you think about it. In order to have a competitive model, you need several earthquakes in your train set. Otherwise you will choose the parameters of your model by averaging the performances of models with large train set and small train set, which are going to be very different. This is a problem if you plan to refit on the whole dataset at the end, because at that point the splits with large train set are obviously more representative of what is a good choice of parameters. \nThe advantage of a flat k-fold is precisely that all folds are on equal footing and have the same right to contribute to the choose of hyper-parameters.",
    "508217": "It's interesting you mentioned the separate quake cycle. If I recall, recent research has indicated that [large earthquakes actually increase the likelihood of subsequent earthquakes, even along different fault lines](https://www.npr.org/2013/08/23/214619037/can-a-big-earthquake-trigger-another-one). Here's another[ paper](https://www.sciencedaily.com/releases/2018/08/180802102352.htm) that discusses similar things. Now, whether the experiments here even remotely model that, I don't know. But, assuming we want this to apply to real world applications, it is something we should take into consideration. \n\nAs for Time Series Split, one way to mitigate this problem is to instead do a Rolling Time Series Split, where each validation chunk is the same. I use this one a lot when working with large Time Series datasets and don't have the resources to handle the final step of TimeSeriesSplit.\n\nFor example, suppose that I split my dataset into 5 \"chunks\".\nThen for validation step 1: chunk 1 is used to predict chunk 2.\nFor validation step 2: ONLY chunk 2 is used to predict chunk 3....\n\nDoing this, you should roughly have 3 quakes worth of data in each chunk. Alternatively (due to the small dataset), you can probably drop this down to 3 chunks to get roughly 5 quakes worth of data in each chunk. \n\nYou can specify this in TimeSeriesSplit using max_train_size.\n\n\nAs an alternative alternative, you could use [hv-block ](https://www.sciencedirect.com/science/article/pii/S0304407600000300)cross validation as described in the paper. I have personally never used it so can't point to the efficacy of it.",
    "2066735": "Thanks for Your Discussion"
  },
  "source": "meta"
}