{
  "id": 90281,
  "title": "Don't Overfit! III",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90281",
  "author_name": "CPMP",
  "post_date": "2019-04-22T11:49:10.700000",
  "votes": 57,
  "comment_count": 73,
  "views": 0,
  "content": "<p>This competition should be called \"Don't Overfit! III\"</p>\n\n<p>There will be a huge shakeup at the end IMHO.</p>\n\n<p>Edit: people rightfully ask how I came to this conclusion.  </p>\n\n<p>One reason is experiments I made to find a reliable cv setting.  </p>\n\n<p>Another reason, maybe more compelling if I don't explain the experiments I made, is that there are really only 16 quakes, i.e. 16 sample only.   You can argue there are 4.1k samples, but how many samples are there for each ttf value?  Answer is: 16.</p>\n\n<p>16 is a very low number. </p>",
  "messages": [
    {
      "id": 521129,
      "postDate": "2019-04-22T11:49:10.700Z",
      "content": "<p>This competition should be called \"Don't Overfit! III\"</p>\n\n<p>There will be a huge shakeup at the end IMHO.</p>\n\n<p>Edit: people rightfully ask how I came to this conclusion.  </p>\n\n<p>One reason is experiments I made to find a reliable cv setting.  </p>\n\n<p>Another reason, maybe more compelling if I don't explain the experiments I made, is that there are really only 16 quakes, i.e. 16 sample only.   You can argue there are 4.1k samples, but how many samples are there for each ttf value?  Answer is: 16.</p>\n\n<p>16 is a very low number. </p>",
      "rawMarkdown": "This competition should be called \"Don't Overfit! III\"\n\nThere will be a huge shakeup at the end IMHO.\n\nEdit: people rightfully ask how I came to this conclusion.  \n\nOne reason is experiments I made to find a reliable cv setting.  \n\nAnother reason, maybe more compelling if I don't explain the experiments I made, is that there are really only 16 quakes, i.e. 16 sample only.   You can argue there are 4.1k samples, but how many samples are there for each ttf value?  Answer is: 16.\n\n16 is a very low number. ",
      "votes": 57
    },
    {
      "id": 521276,
      "postDate": "2019-04-22T17:23:49.123Z",
      "content": "<p>New strategy:</p>\n\n<ol>\n<li>Assume there will be huge shake-up due to small, unrepresentative public test data</li>\n<li>Strive to obtain the worst possible LB score, ensuring a massive jump in your final ranking</li>\n</ol>\n\n<p>Genius!</p>",
      "rawMarkdown": "New strategy:\n\n1. Assume there will be huge shake-up due to small, unrepresentative public test data\n2. Strive to obtain the worst possible LB score, ensuring a massive jump in your final ranking\n\nGenius!",
      "votes": 13,
      "replies": [
        {
          "id": 522386,
          "postDate": "2019-04-24T11:11:38.397Z",
          "content": "<p>That's the spirit!</p>",
          "rawMarkdown": "That's the spirit!"
        },
        {
          "id": 524638,
          "postDate": "2019-04-29T08:37:32.473Z",
          "content": "<p>Haha nice!</p>",
          "rawMarkdown": "Haha nice!"
        },
        {
          "id": 529550,
          "postDate": "2019-05-10T07:00:52.163Z",
          "content": "<p>That is impossible!</p>",
          "rawMarkdown": "That is impossible!"
        },
        {
          "id": 537346,
          "postDate": "2019-05-26T19:26:03.513Z",
          "content": "<p>What Uncle means is that one should srtive for worst possible CV score))</p>",
          "rawMarkdown": "What Uncle means is that one should srtive for worst possible CV score))"
        }
      ]
    },
    {
      "id": 523937,
      "postDate": "2019-04-27T13:02:46.283Z",
      "content": "<p>I see people say: yes, we can overfit because public test is only 13% of test.  Well, risk of overfititng would be almost the same if public test was 87% of test.</p>\n\n<p>Issue is the small size of training data, not public vs private split.</p>",
      "rawMarkdown": "I see people say: yes, we can overfit because public test is only 13% of test.  Well, risk of overfititng would be almost the same if public test was 87% of test.\n\nIssue is the small size of training data, not public vs private split.",
      "votes": 9
    },
    {
      "id": 521422,
      "postDate": "2019-04-22T22:22:11.457Z",
      "content": "<p>Shake-up in shake prediction competition. It will be ironic if it happens :D</p>",
      "rawMarkdown": "Shake-up in shake prediction competition. It will be ironic if it happens :D",
      "votes": 7,
      "replies": [
        {
          "id": 529551,
          "postDate": "2019-05-10T07:02:14.227Z",
          "content": "<p>哈哈</p>",
          "rawMarkdown": "哈哈",
          "votes": -1
        }
      ]
    },
    {
      "id": 521662,
      "postDate": "2019-04-23T08:15:25.040Z",
      "content": "<p>I should at least find a way to overfit the LB and then think about these problems.</p>",
      "rawMarkdown": "I should at least find a way to overfit the LB and then think about these problems.",
      "votes": 5,
      "replies": [
        {
          "id": 521828,
          "postDate": "2019-04-23T13:53:47.513Z",
          "content": "<p>This is also what I am thinking - ok, let's not overfit to LB, but at least be able to do that...</p>",
          "rawMarkdown": "This is also what I am thinking - ok, let's not overfit to LB, but at least be able to do that...",
          "votes": 1
        },
        {
          "id": 521854,
          "postDate": "2019-04-23T14:36:40.450Z",
          "content": "<p>Well, overfitting is not correlated with public LB score.  It is correlated with (public LB score minus private LB score).  Maybe the current leader does not overfit at all while I am, although his public LB score is way lower than mine.</p>",
          "rawMarkdown": "Well, overfitting is not correlated with public LB score.  It is correlated with (public LB score minus private LB score).  Maybe the current leader does not overfit at all while I am, although his public LB score is way lower than mine.",
          "votes": 4
        },
        {
          "id": 522191,
          "postDate": "2019-04-24T02:04:00.407Z",
          "content": "<p>OK. I misunderstood.</p>",
          "rawMarkdown": "OK. I misunderstood.",
          "votes": 2
        }
      ]
    },
    {
      "id": 521565,
      "postDate": "2019-04-23T03:40:49.780Z",
      "content": "<p>Agree, it's a mind game... follow the right way to do cv and don't trust lb too much...</p>",
      "rawMarkdown": "Agree, it's a mind game... follow the right way to do cv and don't trust lb too much...",
      "votes": 5,
      "replies": [
        {
          "id": 521607,
          "postDate": "2019-04-23T05:32:12.093Z",
          "content": "<p>My sentiments exactly <a href=\"/khyeh0719\">@khyeh0719</a> .</p>",
          "rawMarkdown": "My sentiments exactly @khyeh0719 .",
          "votes": 2
        }
      ]
    },
    {
      "id": 534607,
      "postDate": "2019-05-21T14:33:50.123Z",
      "content": "<p>I agree with the shakeup because many reasons including CV settings, but specially because there are only 16 earthquakes in train. So... take care :)</p>",
      "rawMarkdown": "I agree with the shakeup because many reasons including CV settings, but specially because there are only 16 earthquakes in train. So... take care :)",
      "votes": 3,
      "replies": [
        {
          "id": 534614,
          "postDate": "2019-05-21T14:48:51.380Z",
          "content": "<p>I am very concerned that some of my models try to differentiate between these 16 earthquakes rather than trying to find time_to_failure in general.</p>",
          "rawMarkdown": "I am very concerned that some of my models try to differentiate between these 16 earthquakes rather than trying to find time_to_failure in general.",
          "votes": 1
        }
      ]
    },
    {
      "id": 522359,
      "postDate": "2019-04-24T09:56:46.693Z",
      "content": "<p>I think my LB score is lucky, true one is more around 1.43, slowly decreasing, but still not able to beat luck.</p>\n\n<p>Edit: luck is synonym of overfit ;)</p>",
      "rawMarkdown": "I think my LB score is lucky, true one is more around 1.43, slowly decreasing, but still not able to beat luck.\n\nEdit: luck is synonym of overfit ;)",
      "votes": 4,
      "replies": [
        {
          "id": 522654,
          "postDate": "2019-04-24T19:13:24.570Z",
          "content": "<p>Do you mean the private LB? How could you know that? I was expecting a true score around 1.9 at least...</p>",
          "rawMarkdown": "Do you mean the private LB? How could you know that? I was expecting a true score around 1.9 at least..."
        },
        {
          "id": 523066,
          "postDate": "2019-04-25T14:03:05.697Z",
          "content": "<p>I mean my current public LB score should be 1.43.  I have absolutely no clue on the private LB score.</p>",
          "rawMarkdown": "I mean my current public LB score should be 1.43.  I have absolutely no clue on the private LB score."
        }
      ]
    },
    {
      "id": 521832,
      "postDate": "2019-04-23T14:00:56.487Z",
      "content": "<p>Can u give us some advice on how we can find and choose a reliable cv setting ? and what are the patterns that can make us  say that this is a good cv setting or not ? Thank you in advance</p>",
      "rawMarkdown": "Can u give us some advice on how we can find and choose a reliable cv setting ? and what are the patterns that can make us  say that this is a good cv setting or not ? Thank you in advance",
      "votes": 4,
      "replies": [
        {
          "id": 521853,
          "postDate": "2019-04-23T14:35:00.043Z",
          "content": "<blockquote>\n  <p>Can u give us some advice on how we can find and choose a reliable cv setting ?</p>\n</blockquote>\n\n<p>I'm not sure I have found it.</p>\n\n<blockquote>\n  <p>what are the patterns that can make us say that this is a good cv setting or not ?</p>\n</blockquote>\n\n<p>You want it to be reliable, i.e. not exhibit too much variation when you make seemingly neutral changes like changing random seeds.  Also you want it to be correlated with LB, i.e. an improvement in CV score should translate into an improvement in LB score.  </p>\n\n<p>There are some good discussions in this forum on using shuffling or not, and on using quake based vs standard kfolds.</p>",
          "rawMarkdown": "&gt; Can u give us some advice on how we can find and choose a reliable cv setting ?\n\nI'm not sure I have found it.\n\n&gt; what are the patterns that can make us say that this is a good cv setting or not ?\n\nYou want it to be reliable, i.e. not exhibit too much variation when you make seemingly neutral changes like changing random seeds.  Also you want it to be correlated with LB, i.e. an improvement in CV score should translate into an improvement in LB score.  \n\nThere are some good discussions in this forum on using shuffling or not, and on using quake based vs standard kfolds.",
          "votes": 4
        }
      ]
    },
    {
      "id": 521236,
      "postDate": "2019-04-22T15:48:42.703Z",
      "content": "<p>The thing is although the 13% of public data can be similar to the 87% private data, it is unlikely the case in this competition. The old saying is to have one submission to do well in public LB, and then another submission to do well in CV and to avoid overfitting. I will try my best to follow this. </p>",
      "rawMarkdown": "The thing is although the 13% of public data can be similar to the 87% private data, it is unlikely the case in this competition. The old saying is to have one submission to do well in public LB, and then another submission to do well in CV and to avoid overfitting. I will try my best to follow this. ",
      "votes": 3
    },
    {
      "id": 529489,
      "postDate": "2019-05-10T03:14:14.237Z",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Why do you think there is only 16 \"real\" data points? I understood you are counting the top of each ttf (the quakes)? \nIf you translate this lab experiment to real world, the seconds on the y axis could be equivalent to years. This means if some one measure the acoustic signal right after the first quake in the data set it should give him the acoustic signal that will tell him the next quake will be after 12 seconds (or whatever in real life time equivalent). \nIf the person measure the signal after 2 seconds (whatever in real life time equivalent) of the first quake the signal should give him 10 seconds ttf for the next one so on and so forth. The point is all the ttf data points are as good as the one at the peak that you are counting. Am i missing anything here? I am not arguing or anything just trying to understand.\nThanks in advance for all the help and sharing. Your no magic no leak name and sharing of only using 14 features for your current score are really helping.</p>",
      "rawMarkdown": "@cpmpml Why do you think there is only 16 \"real\" data points? I understood you are counting the top of each ttf (the quakes)? \nIf you translate this lab experiment to real world, the seconds on the y axis could be equivalent to years. This means if some one measure the acoustic signal right after the first quake in the data set it should give him the acoustic signal that will tell him the next quake will be after 12 seconds (or whatever in real life time equivalent). \nIf the person measure the signal after 2 seconds (whatever in real life time equivalent) of the first quake the signal should give him 10 seconds ttf for the next one so on and so forth. The point is all the ttf data points are as good as the one at the peak that you are counting. Am i missing anything here? I am not arguing or anything just trying to understand.\nThanks in advance for all the help and sharing. Your no magic no leak name and sharing of only using 14 features for your current score are really helping.",
      "votes": 1,
      "replies": [
        {
          "id": 529510,
          "postDate": "2019-05-10T04:24:42.883Z",
          "content": "<blockquote>\n  <p>Why do you think there is only 16 \"real\" data points?</p>\n</blockquote>\n\n<p>I said it. Just look at how many rows have a given time to failure.</p>",
          "rawMarkdown": "&gt; Why do you think there is only 16 \"real\" data points?\n\nI said it. Just look at how many rows have a given time to failure."
        },
        {
          "id": 529635,
          "postDate": "2019-05-10T11:47:22.343Z",
          "content": "<p>Thank you again <a href=\"/cpmpml\">@cpmpml</a> There are way more than 16 ttf data I believe. I see raw data has 600 million ttf values, maybe i am looking at this wrong?.  <a href=\"/inversion\">@inversion</a> starter script aggregations has 4k again i might be looking at this wrong? I think what you are saying is if we know the 16 quake event times then all the other times in between can be known as well because it is time? My point is the acoustic signals should be different and they should tell the time decay from the first to second quake?</p>",
          "rawMarkdown": "Thank you again @cpmpml There are way more than 16 ttf data I believe. I see raw data has 600 million ttf values, maybe i am looking at this wrong?.  @inversion starter script aggregations has 4k again i might be looking at this wrong? I think what you are saying is if we know the 16 quake event times then all the other times in between can be known as well because it is time? My point is the acoustic signals should be different and they should tell the time decay from the first to second quake?"
        },
        {
          "id": 529663,
          "postDate": "2019-05-10T13:07:34.897Z",
          "content": "<blockquote>\n  <p>If you look at the raw data</p>\n</blockquote>\n\n<p>Thanks, I never thought of doing that!  My LB score will improve tremendously thanks to your incredibly wise advice. </p>\n\n<p>;)</p>",
          "rawMarkdown": "&gt; If you look at the raw data\n\nThanks, I never thought of doing that!  My LB score will improve tremendously thanks to your incredibly wise advice. \n\n;)",
          "votes": 1
        },
        {
          "id": 529671,
          "postDate": "2019-05-10T13:21:11.410Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> . Oh no, I did not mean that as a sarcasm or anything bad (I rephrased my question and English is not my first language, apologies). I was trying to understand the logic of considering the 16 quake as 16 useful data points and not considering the other ttf data. Sorry if i offended you i did not mean that. You are one of the people who shares and help idiots like me. Please ignore my question and thank you again for all the help.</p>",
          "rawMarkdown": "@cpmpml . Oh no, I did not mean that as a sarcasm or anything bad (I rephrased my question and English is not my first language, apologies). I was trying to understand the logic of considering the 16 quake as 16 useful data points and not considering the other ttf data. Sorry if i offended you i did not mean that. You are one of the people who shares and help idiots like me. Please ignore my question and thank you again for all the help."
        },
        {
          "id": 529726,
          "postDate": "2019-05-10T15:39:52.877Z",
          "content": "<p>I wrote you this:</p>\n\n<p>&gt; Just look at how many rows have a given time to failure.</p>\n\n<p>And your response was that I should look at data ;)</p>\n\n<p>Look at data, create a submission, and I'm sure we'll be on same page.</p>",
          "rawMarkdown": "I wrote you this:\n\n&gt; Just look at how many rows have a given time to failure.\n\nAnd your response was that I should look at data ;)\n\nLook at data, create a submission, and I'm sure we'll be on same page.\n"
        },
        {
          "id": 529791,
          "postDate": "2019-05-10T18:57:46.657Z",
          "content": "<p>Thank you again <a href=\"/cpmpml\">@cpmpml</a> \nI did some improvements on the starter script and made a submission. In my submission there are 4K data points. <br>\nI see your 16 data point claim because of the 16 quakes. My question was why are you not considering the time decay and the corresponding acoustic data that should show signals of the time decay? Let me phrase the question differently, let's assume there is a negative correlation between ttf and acoustics signal. Higher the signal the imminent a quake is. In that case the incremental increase in acoustic signal to ttf relationship and all the data showing this relationship is valid correct?</p>",
          "rawMarkdown": "Thank you again @cpmpml \nI did some improvements on the starter script and made a submission. In my submission there are 4K data points.  \nI see your 16 data point claim because of the 16 quakes. My question was why are you not considering the time decay and the corresponding acoustic data that should show signals of the time decay? Let me phrase the question differently, let's assume there is a negative correlation between ttf and acoustics signal. Higher the signal the imminent a quake is. In that case the incremental increase in acoustic signal to ttf relationship and all the data showing this relationship is valid correct?"
        },
        {
          "id": 529938,
          "postDate": "2019-05-11T07:54:57.697Z",
          "content": "<p>&gt; why are you not considering the time decay and the corresponding acoustic data that should show signals of the time decay? </p>\n\n<p>I won't answer questions on what I do or don't do in this competition.</p>\n\n<p>I am not sure if you want to convince me doing something you think I am not doing, or if you ask confirmation that you have a good idea.  If the latter, then my usual answer is: try it and see what works or not.</p>",
          "rawMarkdown": "&gt; why are you not considering the time decay and the corresponding acoustic data that should show signals of the time decay? \n\nI won't answer questions on what I do or don't do in this competition.\n\nI am not sure if you want to convince me doing something you think I am not doing, or if you ask confirmation that you have a good idea.  If the latter, then my usual answer is: try it and see what works or not.",
          "votes": 2
        },
        {
          "id": 530295,
          "postDate": "2019-05-12T12:31:37.223Z",
          "content": "<p>Fair enough, Thank you <a href=\"/cpmpml\">@cpmpml</a> </p>",
          "rawMarkdown": "Fair enough, Thank you @cpmpml "
        }
      ]
    },
    {
      "id": 525857,
      "postDate": "2019-05-01T20:50:25.643Z",
      "content": "<p>Each of the 16 quake's will have some variation in TTF -from the physics there is some inherent randomness. This defines the lower bound to the pred error (over many experiments) </p>\n\n<p>If one's model starts to infer which of the 16 experiments the data sample comes from to minimize the fit error,  this will not generalize to the unseen experiments in the test set.   </p>\n\n<p>See the graph attached to this other post below which possibly shows some variance by experiment('quake') and might indicate the possibility to overfit.</p>\n\n<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841</a></p>",
      "rawMarkdown": "Each of the 16 quake's will have some variation in TTF -from the physics there is some inherent randomness. This defines the lower bound to the pred error (over many experiments) \n\nIf one's model starts to infer which of the 16 experiments the data sample comes from to minimize the fit error,  this will not generalize to the unseen experiments in the test set.   \n\nSee the graph attached to this other post below which possibly shows some variance by experiment('quake') and might indicate the possibility to overfit.\n\n[https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841)\n\n",
      "votes": 1
    },
    {
      "id": 521990,
      "postDate": "2019-04-23T18:22:17.030Z",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": 1
    },
    {
      "id": 521488,
      "postDate": "2019-04-23T01:08:35.073Z",
      "content": "<p>I agree</p>",
      "rawMarkdown": "I agree",
      "votes": 1
    },
    {
      "id": 521344,
      "postDate": "2019-04-22T19:57:26.023Z",
      "content": "<p>What do you expect to change between the public and private data?</p>",
      "rawMarkdown": "What do you expect to change between the public and private data?",
      "votes": 1
    },
    {
      "id": 521144,
      "postDate": "2019-04-22T12:24:19.477Z",
      "content": "<p>Why does everybody think there will be a shakeup? Public LB is only 13% of the data, but when local CV goes down, public LB also goes down. Not seeing the discrepancy here as long as you are using 5-folds.</p>",
      "rawMarkdown": "Why does everybody think there will be a shakeup? Public LB is only 13% of the data, but when local CV goes down, public LB also goes down. Not seeing the discrepancy here as long as you are using 5-folds.",
      "votes": 1,
      "replies": [
        {
          "id": 521151,
          "postDate": "2019-04-22T12:52:03.330Z",
          "content": "<p>I'm not basing this on the public LB being small.  I'll say more after competition end.</p>",
          "rawMarkdown": "I'm not basing this on the public LB being small.  I'll say more after competition end.",
          "votes": 3
        }
      ]
    },
    {
      "id": 521141,
      "postDate": "2019-04-22T12:14:56.350Z",
      "content": "<p>I agree!. but it is not easy to trust your cv in this competition. Despite great effort I can not go below the magic limit of mae 2.0+ and LB 1.5+. I have developed a lot of features, but the cv improves only very slightly</p>",
      "rawMarkdown": "I agree!. but it is not easy to trust your cv in this competition. Despite great effort I can not go below the magic limit of mae 2.0+ and LB 1.5+. I have developed a lot of features, but the cv improves only very slightly",
      "votes": 1
    },
    {
      "id": 521338,
      "postDate": "2019-04-22T19:50:36.220Z",
      "content": "<p>Hi there,\n  May I ask how did you achieve this conclusion?</p>",
      "rawMarkdown": "Hi there,\n  May I ask how did you achieve this conclusion?",
      "votes": 2,
      "replies": [
        {
          "id": 521599,
          "postDate": "2019-04-23T05:11:21.197Z",
          "content": "<p>Based on experiments I made.</p>",
          "rawMarkdown": "Based on experiments I made.",
          "votes": 1
        },
        {
          "id": 528591,
          "postDate": "2019-05-08T06:19:06.317Z",
          "content": "<p>Hi <a href=\"/cpmpml\">@cpmpml</a>, I saw you made big progress over the weekend. I saw you only use 14 features to build your model, are these based on \"event\" instead of acoustic data?</p>",
          "rawMarkdown": "Hi @cpmpml, I saw you made big progress over the weekend. I saw you only use 14 features to build your model, are these based on \"event\" instead of acoustic data?"
        },
        {
          "id": 528686,
          "postDate": "2019-05-08T11:28:02.383Z",
          "content": "<blockquote>\n  <p>\"event\" instead of acoustic data</p>\n</blockquote>\n\n<p>Can you give an example of each feature type as I am not sure I understand the difference you make?  All we have in test is acoustic data, right?</p>",
          "rawMarkdown": "&gt; \"event\" instead of acoustic data\n\nCan you give an example of each feature type as I am not sure I understand the difference you make?  All we have in test is acoustic data, right?"
        },
        {
          "id": 529049,
          "postDate": "2019-05-09T04:35:18.237Z",
          "content": "<p>Let me explain. My \"event\" actually means the \"jumps in acoustic data\". I read the host's paper, and I'm thinking maybe split the training set w.r.t \"event\" can help reduce the feature numbers. Currently, I have around 900 features, which is highly probably overfitting.  Those features are generated based on acoustic data stats characteristics. I'm wondering how you generate features. Maybe not from acoustic data stats perspective?</p>\n\n<p>Those features are generated based on acoustic data stats characteristics. I'm wondering how you generate features. Maybe not from acoustic data stats perspective? Thanks!</p>",
          "rawMarkdown": "Let me explain. My \"event\" actually means the \"jumps in acoustic data\". I read the host's paper, and I'm thinking maybe split the training set w.r.t \"event\" can help reduce the feature numbers. Currently, I have around 900 features, which is highly probably overfitting.  Those features are generated based on acoustic data stats characteristics. I'm wondering how you generate features. Maybe not from acoustic data stats perspective?\n\nThose features are generated based on acoustic data stats characteristics. I'm wondering how you generate features. Maybe not from acoustic data stats perspective? Thanks!\n\n"
        }
      ]
    },
    {
      "id": 521137,
      "postDate": "2019-04-22T12:02:04.427Z",
      "content": "<p>Yep. Very huge.  </p>",
      "rawMarkdown": "Yep. Very huge.  ",
      "votes": 2
    },
    {
      "id": 535710,
      "postDate": "2019-05-23T11:06:57.633Z",
      "content": "<p>I hope we are not overfitting, but it is possible...</p>",
      "rawMarkdown": "I hope we are not overfitting, but it is possible...",
      "replies": [
        {
          "id": 535739,
          "postDate": "2019-05-23T11:35:41.270Z",
          "content": "<p>I hope you are not overfitting - I want to know what the feature is!</p>",
          "rawMarkdown": "I hope you are not overfitting - I want to know what the feature is!",
          "votes": 1
        },
        {
          "id": 535745,
          "postDate": "2019-05-23T11:41:29.300Z",
          "content": "<p>The feature? Didn't CPMP said there's no magic? </p>",
          "rawMarkdown": "The feature? Didn't CPMP said there's no magic? ",
          "votes": 1
        },
        {
          "id": 535747,
          "postDate": "2019-05-23T11:43:36.847Z",
          "content": "<p>No magic feature for me indeed, but a small set of good features.</p>",
          "rawMarkdown": "No magic feature for me indeed, but a small set of good features.",
          "votes": 1
        },
        {
          "id": 535760,
          "postDate": "2019-05-23T12:05:09.073Z",
          "content": "<p>An indicator of overfitting is mean of test predictions. Would you like to share it? </p>",
          "rawMarkdown": "An indicator of overfitting is mean of test predictions. Would you like to share it? "
        },
        {
          "id": 535770,
          "postDate": "2019-05-23T12:18:15.180Z",
          "content": "<p>How is it an indicator of overfit?</p>",
          "rawMarkdown": "How is it an indicator of overfit?",
          "votes": 1
        },
        {
          "id": 535778,
          "postDate": "2019-05-23T12:34:20.863Z",
          "content": "<p>If you share yours, I’ll share mine, and explain why. I think it is a good indicator. </p>",
          "rawMarkdown": "If you share yours, I’ll share mine, and explain why. I think it is a good indicator. "
        },
        {
          "id": 535786,
          "postDate": "2019-05-23T12:47:56.867Z",
          "content": "<p>Nice try ;)</p>\n\n<p>This would mean you have a fairly good idea of the mean of target for private test data.  Congrats.  Will be interesting to read how you found it.</p>\n\n<p>More generally, I won't share more info than I shared already in this forum.  You shared a bit too, but many others in the top of LB didn't share anything.  I may revisit my position if they share something interesting.</p>",
          "rawMarkdown": "Nice try ;)\n\nThis would mean you have a fairly good idea of the mean of target for private test data.  Congrats.  Will be interesting to read how you found it.\n\nMore generally, I won't share more info than I shared already in this forum.  You shared a bit too, but many others in the top of LB didn't share anything.  I may revisit my position if they share something interesting.",
          "votes": 2
        },
        {
          "id": 535788,
          "postDate": "2019-05-23T12:52:43Z",
          "content": "<p>what is your LB and CV gap ?</p>",
          "rawMarkdown": "what is your LB and CV gap ?"
        },
        {
          "id": 535791,
          "postDate": "2019-05-23T12:56:18.557Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I was about to grab some popcorn while reading the exchange. Now I'm disappointed</p>",
          "rawMarkdown": "@cpmpml I was about to grab some popcorn while reading the exchange. Now I'm disappointed",
          "votes": 1
        },
        {
          "id": 535793,
          "postDate": "2019-05-23T13:04:17.783Z",
          "content": "<p>Not sure why you're disappointed.  I wrote I would not share anything when I entered the competition, yet I ended up sharing things, like using a small set of features, and explaining what I look for in features.  I'm sure this helped some people already.  </p>\n\n<p>You shared as well, I am not targeting you at all here.</p>",
          "rawMarkdown": "Not sure why you're disappointed.  I wrote I would not share anything when I entered the competition, yet I ended up sharing things, like using a small set of features, and explaining what I look for in features.  I'm sure this helped some people already.  \n\nYou shared as well, I am not targeting you at all here.",
          "votes": 1
        },
        {
          "id": 535794,
          "postDate": "2019-05-23T13:09:45.523Z",
          "content": "<blockquote>\n  <p>what is your LB and CV gap ?</p>\n</blockquote>\n\n<p>This isn't relevant.  What is relevant is to have a CV score that correlates well with private LB score., for instance by having a constant gap.  Given we cannot assess this before competition end, a proxy is to look for CV setting such that CV score is well correlated with public LB score.  </p>\n\n<p>Our CV setting is such that CV score correlates reasonable well with LB score, but not always.</p>",
          "rawMarkdown": "&gt; what is your LB and CV gap ?\n\nThis isn't relevant.  What is relevant is to have a CV score that correlates well with private LB score., for instance by having a constant gap.  Given we cannot assess this before competition end, a proxy is to look for CV setting such that CV score is well correlated with public LB score.  \n\nOur CV setting is such that CV score correlates reasonable well with LB score, but not always.",
          "votes": 1
        },
        {
          "id": 535947,
          "postDate": "2019-05-23T17:14:11.197Z",
          "content": "<p>I'm using 17 features and my CV seems to be correlated with LB. I pay attention ONLY to CV. My test mean is, never touched 5.078. Any comment on that?</p>",
          "rawMarkdown": "I'm using 17 features and my CV seems to be correlated with LB. I pay attention ONLY to CV. My test mean is, never touched 5.078. Any comment on that?"
        }
      ]
    },
    {
      "id": 534587,
      "postDate": "2019-05-21T13:57:56.547Z",
      "content": "<p>I agree with the topic author @CPMP.</p>\n\n<p>The original work from LANL (eg\n<a href=\"https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL074677\">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL074677</a>\n) was based on a dataset much more easier because this dataset we are using on this competition <strike>there are some <em>minor shear stress drops</em></strike> we have only the time to failure and the failure was considered based on shear stress drop using a threashold making the longer experiments different from the shorter ones because in the case of the longer the shear stress may drop inside the threashold (this is kind of <em>minor shear stress drops</em>).</p>\n\n<p>I tried to rewrite the preceding paragraph but i think it still confuse, but my point is that: <strong>even though you have only 16 data points, the fact that the 16 experiment they come from are very different makes the problem harder</strong> </p>\n\n<p>Quoting <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview/additional-information\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview/additional-information</a></p>\n\n<p>&gt; Los Alamos' initial work showed that the prediction of laboratory earthquakes from continuous seismic data is possible in the case of quasi-periodic laboratory seismic cycles. In this competition, the team has provided a much more challenging dataset with considerably more aperiodic earthquake failures.</p>\n\n<p>I think you got the main challenge.</p>",
      "rawMarkdown": "I agree with the topic author @CPMP.\n\n The original work from LANL (eg\nhttps://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL074677\n) was based on a dataset much more easier because this dataset we are using on this competition <strike>there are some *minor shear stress drops*</strike> we have only the time to failure and the failure was considered based on shear stress drop using a threashold making the longer experiments different from the shorter ones because in the case of the longer the shear stress may drop inside the threashold (this is kind of *minor shear stress drops*).\n\nI tried to rewrite the preceding paragraph but i think it still confuse, but my point is that: **even though you have only 16 data points, the fact that the 16 experiment they come from are very different makes the problem harder** \n\nQuoting https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview/additional-information\n\n&gt; Los Alamos' initial work showed that the prediction of laboratory earthquakes from continuous seismic data is possible in the case of quasi-periodic laboratory seismic cycles. In this competition, the team has provided a much more challenging dataset with considerably more aperiodic earthquake failures.\n\nI think you got the main challenge."
    },
    {
      "id": 530009,
      "postDate": "2019-05-11T13:32:42.577Z",
      "content": "<p><a href=\"/achyuth1\">@achyuth1</a> see this </p>",
      "rawMarkdown": "@achyuth1 see this "
    },
    {
      "id": 529060,
      "postDate": "2019-05-09T05:23:21.013Z",
      "rawMarkdown": ""
    },
    {
      "id": 524586,
      "postDate": "2019-04-29T06:18:13.063Z",
      "content": "<p>I agree!</p>",
      "rawMarkdown": "I agree!"
    },
    {
      "id": 523637,
      "postDate": "2019-04-26T16:52:40.713Z",
      "content": "<p>Cool</p>",
      "rawMarkdown": "Cool"
    },
    {
      "id": 523587,
      "postDate": "2019-04-26T15:10:56.670Z",
      "content": "<p>thx</p>",
      "rawMarkdown": "thx"
    },
    {
      "id": 523259,
      "postDate": "2019-04-25T20:59:29.247Z",
      "content": "<p>thx</p>",
      "rawMarkdown": "thx"
    },
    {
      "id": 523079,
      "postDate": "2019-04-25T14:47:11.527Z",
      "content": "<p>Wow</p>",
      "rawMarkdown": "Wow"
    },
    {
      "id": 522761,
      "postDate": "2019-04-25T01:00:18.010Z",
      "content": "<p>Yes, seems much like a mind game. May be experience will help handling this problem well.. </p>",
      "rawMarkdown": "Yes, seems much like a mind game. May be experience will help handling this problem well.. "
    },
    {
      "id": 522688,
      "postDate": "2019-04-24T20:24:18.890Z",
      "content": "<p>May be this is a silly question but how did you arrived at number 16 ?. I am trying to understand this problem since quite a while still struggling to a have a well understanding also the size of data makes it hard to have a good summary. </p>",
      "rawMarkdown": "May be this is a silly question but how did you arrived at number 16 ?. I am trying to understand this problem since quite a while still struggling to a have a well understanding also the size of data makes it hard to have a good summary. ",
      "replies": [
        {
          "id": 522848,
          "postDate": "2019-04-25T06:07:31.387Z",
          "content": "<p>I said it.  Just look at how many rows have a given time to failure.</p>",
          "rawMarkdown": "I said it.  Just look at how many rows have a given time to failure.",
          "votes": 3
        },
        {
          "id": 523143,
          "postDate": "2019-04-25T16:33:17.040Z",
          "content": "<p>got it now, thanks :)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png\" alt=\"\"></p>",
          "rawMarkdown": "got it now, thanks :)\n\n![](https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png)",
          "votes": 4
        }
      ]
    },
    {
      "id": 521138,
      "postDate": "2019-04-22T12:06:58.273Z",
      "content": "<p>What would you do to avoid it?</p>",
      "rawMarkdown": "What would you do to avoid it?",
      "replies": [
        {
          "id": 521152,
          "postDate": "2019-04-22T12:52:30.950Z",
          "content": "<p>Ensembling many models I guess.  Haven't started yet.</p>",
          "rawMarkdown": "Ensembling many models I guess.  Haven't started yet.",
          "votes": 2
        }
      ]
    },
    {
      "id": 523054,
      "postDate": "2019-04-25T13:35:39.613Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 521197,
      "postDate": "2019-04-22T14:30:04.393Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 521276,
      "author_name": "RNA",
      "author_url": "",
      "post_date": "2019-04-22T17:23:49.123000",
      "content": "<p>New strategy:</p>\n\n<ol>\n<li>Assume there will be huge shake-up due to small, unrepresentative public test data</li>\n<li>Strive to obtain the worst possible LB score, ensuring a massive jump in your final ranking</li>\n</ol>\n\n<p>Genius!</p>",
      "votes": 13,
      "replies": [
        {
          "id": 522386,
          "author_name": "Sven Hinderer",
          "author_url": "",
          "post_date": "2019-04-24T11:11:38.397000",
          "content": "<p>That's the spirit!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 524638,
          "author_name": "jagdeep",
          "author_url": "",
          "post_date": "2019-04-29T08:37:32.473000",
          "content": "<p>Haha nice!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529550,
          "author_name": "张鉴鸾",
          "author_url": "",
          "post_date": "2019-05-10T07:00:52.163000",
          "content": "<p>That is impossible!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 537346,
          "author_name": "George Okromchedlishvili",
          "author_url": "",
          "post_date": "2019-05-26T19:26:03.513000",
          "content": "<p>What Uncle means is that one should srtive for worst possible CV score))</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 523937,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-27T13:02:46.283000",
      "content": "<p>I see people say: yes, we can overfit because public test is only 13% of test.  Well, risk of overfititng would be almost the same if public test was 87% of test.</p>\n\n<p>Issue is the small size of training data, not public vs private split.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 521422,
      "author_name": "Murat Korkmaz",
      "author_url": "",
      "post_date": "2019-04-22T22:22:11.457000",
      "content": "<p>Shake-up in shake prediction competition. It will be ironic if it happens :D</p>",
      "votes": 7,
      "replies": [
        {
          "id": 529551,
          "author_name": "张鉴鸾",
          "author_url": "",
          "post_date": "2019-05-10T07:02:14.227000",
          "content": "<p>哈哈</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 521662,
      "author_name": "Marcus Lin",
      "author_url": "",
      "post_date": "2019-04-23T08:15:25.040000",
      "content": "<p>I should at least find a way to overfit the LB and then think about these problems.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 521828,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-23T13:53:47.513000",
          "content": "<p>This is also what I am thinking - ok, let's not overfit to LB, but at least be able to do that...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 521854,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T14:36:40.450000",
          "content": "<p>Well, overfitting is not correlated with public LB score.  It is correlated with (public LB score minus private LB score).  Maybe the current leader does not overfit at all while I am, although his public LB score is way lower than mine.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 522191,
          "author_name": "Marcus Lin",
          "author_url": "",
          "post_date": "2019-04-24T02:04:00.407000",
          "content": "<p>OK. I misunderstood.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 521565,
      "author_name": "khyeh",
      "author_url": "",
      "post_date": "2019-04-23T03:40:49.780000",
      "content": "<p>Agree, it's a mind game... follow the right way to do cv and don't trust lb too much...</p>",
      "votes": 5,
      "replies": [
        {
          "id": 521607,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-04-23T05:32:12.093000",
          "content": "<p>My sentiments exactly <a href=\"/khyeh0719\">@khyeh0719</a> .</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 534607,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-05-21T14:33:50.123000",
      "content": "<p>I agree with the shakeup because many reasons including CV settings, but specially because there are only 16 earthquakes in train. So... take care :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 534614,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-21T14:48:51.380000",
          "content": "<p>I am very concerned that some of my models try to differentiate between these 16 earthquakes rather than trying to find time_to_failure in general.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 522359,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-24T09:56:46.693000",
      "content": "<p>I think my LB score is lucky, true one is more around 1.43, slowly decreasing, but still not able to beat luck.</p>\n\n<p>Edit: luck is synonym of overfit ;)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 522654,
          "author_name": "JinCui",
          "author_url": "",
          "post_date": "2019-04-24T19:13:24.570000",
          "content": "<p>Do you mean the private LB? How could you know that? I was expecting a true score around 1.9 at least...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 523066,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-25T14:03:05.697000",
          "content": "<p>I mean my current public LB score should be 1.43.  I have absolutely no clue on the private LB score.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 521832,
      "author_name": "propower",
      "author_url": "",
      "post_date": "2019-04-23T14:00:56.487000",
      "content": "<p>Can u give us some advice on how we can find and choose a reliable cv setting ? and what are the patterns that can make us  say that this is a good cv setting or not ? Thank you in advance</p>",
      "votes": 4,
      "replies": [
        {
          "id": 521853,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T14:35:00.043000",
          "content": "<blockquote>\n  <p>Can u give us some advice on how we can find and choose a reliable cv setting ?</p>\n</blockquote>\n\n<p>I'm not sure I have found it.</p>\n\n<blockquote>\n  <p>what are the patterns that can make us say that this is a good cv setting or not ?</p>\n</blockquote>\n\n<p>You want it to be reliable, i.e. not exhibit too much variation when you make seemingly neutral changes like changing random seeds.  Also you want it to be correlated with LB, i.e. an improvement in CV score should translate into an improvement in LB score.  </p>\n\n<p>There are some good discussions in this forum on using shuffling or not, and on using quake based vs standard kfolds.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 521236,
      "author_name": "FP",
      "author_url": "",
      "post_date": "2019-04-22T15:48:42.703000",
      "content": "<p>The thing is although the 13% of public data can be similar to the 87% private data, it is unlikely the case in this competition. The old saying is to have one submission to do well in public LB, and then another submission to do well in CV and to avoid overfitting. I will try my best to follow this. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 529489,
      "author_name": "Anon",
      "author_url": "",
      "post_date": "2019-05-10T03:14:14.237000",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Why do you think there is only 16 \"real\" data points? I understood you are counting the top of each ttf (the quakes)? \nIf you translate this lab experiment to real world, the seconds on the y axis could be equivalent to years. This means if some one measure the acoustic signal right after the first quake in the data set it should give him the acoustic signal that will tell him the next quake will be after 12 seconds (or whatever in real life time equivalent). \nIf the person measure the signal after 2 seconds (whatever in real life time equivalent) of the first quake the signal should give him 10 seconds ttf for the next one so on and so forth. The point is all the ttf data points are as good as the one at the peak that you are counting. Am i missing anything here? I am not arguing or anything just trying to understand.\nThanks in advance for all the help and sharing. Your no magic no leak name and sharing of only using 14 features for your current score are really helping.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 529510,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-10T04:24:42.883000",
          "content": "<blockquote>\n  <p>Why do you think there is only 16 \"real\" data points?</p>\n</blockquote>\n\n<p>I said it. Just look at how many rows have a given time to failure.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529635,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2019-05-10T11:47:22.343000",
          "content": "<p>Thank you again <a href=\"/cpmpml\">@cpmpml</a> There are way more than 16 ttf data I believe. I see raw data has 600 million ttf values, maybe i am looking at this wrong?.  <a href=\"/inversion\">@inversion</a> starter script aggregations has 4k again i might be looking at this wrong? I think what you are saying is if we know the 16 quake event times then all the other times in between can be known as well because it is time? My point is the acoustic signals should be different and they should tell the time decay from the first to second quake?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529663,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-10T13:07:34.897000",
          "content": "<blockquote>\n  <p>If you look at the raw data</p>\n</blockquote>\n\n<p>Thanks, I never thought of doing that!  My LB score will improve tremendously thanks to your incredibly wise advice. </p>\n\n<p>;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529671,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2019-05-10T13:21:11.410000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> . Oh no, I did not mean that as a sarcasm or anything bad (I rephrased my question and English is not my first language, apologies). I was trying to understand the logic of considering the 16 quake as 16 useful data points and not considering the other ttf data. Sorry if i offended you i did not mean that. You are one of the people who shares and help idiots like me. Please ignore my question and thank you again for all the help.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529726,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-10T15:39:52.877000",
          "content": "<p>I wrote you this:</p>\n\n<p>&gt; Just look at how many rows have a given time to failure.</p>\n\n<p>And your response was that I should look at data ;)</p>\n\n<p>Look at data, create a submission, and I'm sure we'll be on same page.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529791,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2019-05-10T18:57:46.657000",
          "content": "<p>Thank you again <a href=\"/cpmpml\">@cpmpml</a> \nI did some improvements on the starter script and made a submission. In my submission there are 4K data points. <br>\nI see your 16 data point claim because of the 16 quakes. My question was why are you not considering the time decay and the corresponding acoustic data that should show signals of the time decay? Let me phrase the question differently, let's assume there is a negative correlation between ttf and acoustics signal. Higher the signal the imminent a quake is. In that case the incremental increase in acoustic signal to ttf relationship and all the data showing this relationship is valid correct?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529938,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-11T07:54:57.697000",
          "content": "<p>&gt; why are you not considering the time decay and the corresponding acoustic data that should show signals of the time decay? </p>\n\n<p>I won't answer questions on what I do or don't do in this competition.</p>\n\n<p>I am not sure if you want to convince me doing something you think I am not doing, or if you ask confirmation that you have a good idea.  If the latter, then my usual answer is: try it and see what works or not.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 530295,
          "author_name": "Anon",
          "author_url": "",
          "post_date": "2019-05-12T12:31:37.223000",
          "content": "<p>Fair enough, Thank you <a href=\"/cpmpml\">@cpmpml</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 525857,
      "author_name": "Filip Mulier",
      "author_url": "",
      "post_date": "2019-05-01T20:50:25.643000",
      "content": "<p>Each of the 16 quake's will have some variation in TTF -from the physics there is some inherent randomness. This defines the lower bound to the pred error (over many experiments) </p>\n\n<p>If one's model starts to infer which of the 16 experiments the data sample comes from to minimize the fit error,  this will not generalize to the unseen experiments in the test set.   </p>\n\n<p>See the graph attached to this other post below which possibly shows some variance by experiment('quake') and might indicate the possibility to overfit.</p>\n\n<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 521990,
      "author_name": "Taj Alagawani",
      "author_url": "",
      "post_date": "2019-04-23T18:22:17.030000",
      "content": "<p>yes</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 521488,
      "author_name": "TY.Yang",
      "author_url": "",
      "post_date": "2019-04-23T01:08:35.073000",
      "content": "<p>I agree</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 521344,
      "author_name": "Zach Nagengast",
      "author_url": "",
      "post_date": "2019-04-22T19:57:26.023000",
      "content": "<p>What do you expect to change between the public and private data?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 521144,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2019-04-22T12:24:19.477000",
      "content": "<p>Why does everybody think there will be a shakeup? Public LB is only 13% of the data, but when local CV goes down, public LB also goes down. Not seeing the discrepancy here as long as you are using 5-folds.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 521151,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-22T12:52:03.330000",
          "content": "<p>I'm not basing this on the public LB being small.  I'll say more after competition end.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 521141,
      "author_name": "Bernd Allmendinger",
      "author_url": "",
      "post_date": "2019-04-22T12:14:56.350000",
      "content": "<p>I agree!. but it is not easy to trust your cv in this competition. Despite great effort I can not go below the magic limit of mae 2.0+ and LB 1.5+. I have developed a lot of features, but the cv improves only very slightly</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 521338,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-22T19:50:36.220000",
      "content": "<p>Hi there,\n  May I ask how did you achieve this conclusion?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 521599,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T05:11:21.197000",
          "content": "<p>Based on experiments I made.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 528591,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-08T06:19:06.317000",
          "content": "<p>Hi <a href=\"/cpmpml\">@cpmpml</a>, I saw you made big progress over the weekend. I saw you only use 14 features to build your model, are these based on \"event\" instead of acoustic data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 528686,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-08T11:28:02.383000",
          "content": "<blockquote>\n  <p>\"event\" instead of acoustic data</p>\n</blockquote>\n\n<p>Can you give an example of each feature type as I am not sure I understand the difference you make?  All we have in test is acoustic data, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529049,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-09T04:35:18.237000",
          "content": "<p>Let me explain. My \"event\" actually means the \"jumps in acoustic data\". I read the host's paper, and I'm thinking maybe split the training set w.r.t \"event\" can help reduce the feature numbers. Currently, I have around 900 features, which is highly probably overfitting.  Those features are generated based on acoustic data stats characteristics. I'm wondering how you generate features. Maybe not from acoustic data stats perspective?</p>\n\n<p>Those features are generated based on acoustic data stats characteristics. I'm wondering how you generate features. Maybe not from acoustic data stats perspective? Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 521137,
      "author_name": "DmitryS",
      "author_url": "",
      "post_date": "2019-04-22T12:02:04.427000",
      "content": "<p>Yep. Very huge.  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 535710,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-05-23T11:06:57.633000",
      "content": "<p>I hope we are not overfitting, but it is possible...</p>",
      "votes": 0,
      "replies": [
        {
          "id": 535739,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-05-23T11:35:41.270000",
          "content": "<p>I hope you are not overfitting - I want to know what the feature is!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535745,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-05-23T11:41:29.300000",
          "content": "<p>The feature? Didn't CPMP said there's no magic? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535747,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-23T11:43:36.847000",
          "content": "<p>No magic feature for me indeed, but a small set of good features.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535760,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-05-23T12:05:09.073000",
          "content": "<p>An indicator of overfitting is mean of test predictions. Would you like to share it? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535770,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-23T12:18:15.180000",
          "content": "<p>How is it an indicator of overfit?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535778,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2019-05-23T12:34:20.863000",
          "content": "<p>If you share yours, I’ll share mine, and explain why. I think it is a good indicator. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535786,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-23T12:47:56.867000",
          "content": "<p>Nice try ;)</p>\n\n<p>This would mean you have a fairly good idea of the mean of target for private test data.  Congrats.  Will be interesting to read how you found it.</p>\n\n<p>More generally, I won't share more info than I shared already in this forum.  You shared a bit too, but many others in the top of LB didn't share anything.  I may revisit my position if they share something interesting.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 535788,
          "author_name": "Manoj",
          "author_url": "",
          "post_date": "2019-05-23T12:52:43",
          "content": "<p>what is your LB and CV gap ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535791,
          "author_name": "Amjad",
          "author_url": "",
          "post_date": "2019-05-23T12:56:18.557000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> I was about to grab some popcorn while reading the exchange. Now I'm disappointed</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535793,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-23T13:04:17.783000",
          "content": "<p>Not sure why you're disappointed.  I wrote I would not share anything when I entered the competition, yet I ended up sharing things, like using a small set of features, and explaining what I look for in features.  I'm sure this helped some people already.  </p>\n\n<p>You shared as well, I am not targeting you at all here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535794,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-23T13:09:45.523000",
          "content": "<blockquote>\n  <p>what is your LB and CV gap ?</p>\n</blockquote>\n\n<p>This isn't relevant.  What is relevant is to have a CV score that correlates well with private LB score., for instance by having a constant gap.  Given we cannot assess this before competition end, a proxy is to look for CV setting such that CV score is well correlated with public LB score.  </p>\n\n<p>Our CV setting is such that CV score correlates reasonable well with LB score, but not always.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535947,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-05-23T17:14:11.197000",
          "content": "<p>I'm using 17 features and my CV seems to be correlated with LB. I pay attention ONLY to CV. My test mean is, never touched 5.078. Any comment on that?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 534587,
      "author_name": "Alex V B",
      "author_url": "",
      "post_date": "2019-05-21T13:57:56.547000",
      "content": "<p>I agree with the topic author @CPMP.</p>\n\n<p>The original work from LANL (eg\n<a href=\"https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL074677\">https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL074677</a>\n) was based on a dataset much more easier because this dataset we are using on this competition <strike>there are some <em>minor shear stress drops</em></strike> we have only the time to failure and the failure was considered based on shear stress drop using a threashold making the longer experiments different from the shorter ones because in the case of the longer the shear stress may drop inside the threashold (this is kind of <em>minor shear stress drops</em>).</p>\n\n<p>I tried to rewrite the preceding paragraph but i think it still confuse, but my point is that: <strong>even though you have only 16 data points, the fact that the 16 experiment they come from are very different makes the problem harder</strong> </p>\n\n<p>Quoting <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview/additional-information\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview/additional-information</a></p>\n\n<p>&gt; Los Alamos' initial work showed that the prediction of laboratory earthquakes from continuous seismic data is possible in the case of quasi-periodic laboratory seismic cycles. In this competition, the team has provided a much more challenging dataset with considerably more aperiodic earthquake failures.</p>\n\n<p>I think you got the main challenge.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 530009,
      "author_name": "Sayantan Das",
      "author_url": "",
      "post_date": "2019-05-11T13:32:42.577000",
      "content": "<p><a href=\"/achyuth1\">@achyuth1</a> see this </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 529060,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2019-05-09T05:23:21.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 524586,
      "author_name": "tornado76",
      "author_url": "",
      "post_date": "2019-04-29T06:18:13.063000",
      "content": "<p>I agree!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 523637,
      "author_name": "Abiola Sheikh",
      "author_url": "",
      "post_date": "2019-04-26T16:52:40.713000",
      "content": "<p>Cool</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 523587,
      "author_name": "Sharmiko",
      "author_url": "",
      "post_date": "2019-04-26T15:10:56.670000",
      "content": "<p>thx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 523259,
      "author_name": "Ihor Krutenko",
      "author_url": "",
      "post_date": "2019-04-25T20:59:29.247000",
      "content": "<p>thx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 523079,
      "author_name": "Debayan Ganguly",
      "author_url": "",
      "post_date": "2019-04-25T14:47:11.527000",
      "content": "<p>Wow</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 522761,
      "author_name": "Prabhu Rangaraju",
      "author_url": "",
      "post_date": "2019-04-25T01:00:18.010000",
      "content": "<p>Yes, seems much like a mind game. May be experience will help handling this problem well.. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 522688,
      "author_name": "Manoj",
      "author_url": "",
      "post_date": "2019-04-24T20:24:18.890000",
      "content": "<p>May be this is a silly question but how did you arrived at number 16 ?. I am trying to understand this problem since quite a while still struggling to a have a well understanding also the size of data makes it hard to have a good summary. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 522848,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-25T06:07:31.387000",
          "content": "<p>I said it.  Just look at how many rows have a given time to failure.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 523143,
          "author_name": "Manoj",
          "author_url": "",
          "post_date": "2019-04-25T16:33:17.040000",
          "content": "<p>got it now, thanks :)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/523143/13096/download.png\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 521138,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-22T12:06:58.273000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 521152,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-22T12:52:30.950000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 523054,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-25T13:35:39.613000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 521197,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-22T14:30:04.393000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "521129": "This competition should be called \"Don't Overfit! III\"\n\nThere will be a huge shakeup at the end IMHO.\n\nEdit: people rightfully ask how I came to this conclusion.  \n\nOne reason is experiments I made to find a reliable cv setting.  \n\nAnother reason, maybe more compelling if I don't explain the experiments I made, is that there are really only 16 quakes, i.e. 16 sample only.   You can argue there are 4.1k samples, but how many samples are there for each ttf value?  Answer is: 16.\n\n16 is a very low number. ",
    "521276": "New strategy:\n\n1. Assume there will be huge shake-up due to small, unrepresentative public test data\n2. Strive to obtain the worst possible LB score, ensuring a massive jump in your final ranking\n\nGenius!",
    "523937": "I see people say: yes, we can overfit because public test is only 13% of test.  Well, risk of overfititng would be almost the same if public test was 87% of test.\n\nIssue is the small size of training data, not public vs private split.",
    "521422": "Shake-up in shake prediction competition. It will be ironic if it happens :D",
    "521662": "I should at least find a way to overfit the LB and then think about these problems.",
    "521565": "Agree, it's a mind game... follow the right way to do cv and don't trust lb too much...",
    "534607": "I agree with the shakeup because many reasons including CV settings, but specially because there are only 16 earthquakes in train. So... take care :)",
    "522359": "I think my LB score is lucky, true one is more around 1.43, slowly decreasing, but still not able to beat luck.\n\nEdit: luck is synonym of overfit ;)",
    "521832": "Can u give us some advice on how we can find and choose a reliable cv setting ? and what are the patterns that can make us  say that this is a good cv setting or not ? Thank you in advance",
    "521236": "The thing is although the 13% of public data can be similar to the 87% private data, it is unlikely the case in this competition. The old saying is to have one submission to do well in public LB, and then another submission to do well in CV and to avoid overfitting. I will try my best to follow this. ",
    "529489": "@cpmpml Why do you think there is only 16 \"real\" data points? I understood you are counting the top of each ttf (the quakes)? \nIf you translate this lab experiment to real world, the seconds on the y axis could be equivalent to years. This means if some one measure the acoustic signal right after the first quake in the data set it should give him the acoustic signal that will tell him the next quake will be after 12 seconds (or whatever in real life time equivalent). \nIf the person measure the signal after 2 seconds (whatever in real life time equivalent) of the first quake the signal should give him 10 seconds ttf for the next one so on and so forth. The point is all the ttf data points are as good as the one at the peak that you are counting. Am i missing anything here? I am not arguing or anything just trying to understand.\nThanks in advance for all the help and sharing. Your no magic no leak name and sharing of only using 14 features for your current score are really helping.",
    "525857": "Each of the 16 quake's will have some variation in TTF -from the physics there is some inherent randomness. This defines the lower bound to the pred error (over many experiments) \n\nIf one's model starts to infer which of the 16 experiments the data sample comes from to minimize the fit error,  this will not generalize to the unseen experiments in the test set.   \n\nSee the graph attached to this other post below which possibly shows some variance by experiment('quake') and might indicate the possibility to overfit.\n\n[https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125#525841)\n\n",
    "521990": "yes",
    "521488": "I agree",
    "521344": "What do you expect to change between the public and private data?",
    "521144": "Why does everybody think there will be a shakeup? Public LB is only 13% of the data, but when local CV goes down, public LB also goes down. Not seeing the discrepancy here as long as you are using 5-folds.",
    "521141": "I agree!. but it is not easy to trust your cv in this competition. Despite great effort I can not go below the magic limit of mae 2.0+ and LB 1.5+. I have developed a lot of features, but the cv improves only very slightly",
    "521338": "Hi there,\n  May I ask how did you achieve this conclusion?",
    "521137": "Yep. Very huge.  ",
    "535710": "I hope we are not overfitting, but it is possible...",
    "534587": "I agree with the topic author @CPMP.\n\n The original work from LANL (eg\nhttps://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2017GL074677\n) was based on a dataset much more easier because this dataset we are using on this competition <strike>there are some *minor shear stress drops*</strike> we have only the time to failure and the failure was considered based on shear stress drop using a threashold making the longer experiments different from the shorter ones because in the case of the longer the shear stress may drop inside the threashold (this is kind of *minor shear stress drops*).\n\nI tried to rewrite the preceding paragraph but i think it still confuse, but my point is that: **even though you have only 16 data points, the fact that the 16 experiment they come from are very different makes the problem harder** \n\nQuoting https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview/additional-information\n\n&gt; Los Alamos' initial work showed that the prediction of laboratory earthquakes from continuous seismic data is possible in the case of quasi-periodic laboratory seismic cycles. In this competition, the team has provided a much more challenging dataset with considerably more aperiodic earthquake failures.\n\nI think you got the main challenge.",
    "530009": "@achyuth1 see this ",
    "529060": "",
    "524586": "I agree!",
    "523637": "Cool",
    "523587": "thx",
    "523259": "thx",
    "523079": "Wow",
    "522761": "Yes, seems much like a mind game. May be experience will help handling this problem well.. ",
    "522688": "May be this is a silly question but how did you arrived at number 16 ?. I am trying to understand this problem since quite a while still struggling to a have a well understanding also the size of data makes it hard to have a good summary. ",
    "521138": "What would you do to avoid it?",
    "523054": "",
    "521197": ""
  }
}