{
  "id": 90139,
  "title": "How Grandmasters are so good",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90139",
  "author_name": "DavidS",
  "post_date": "2019-04-20T20:45:54.509000",
  "votes": 17,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Both @tunguz and @cpmpml got top 40 just after joining the competition, in few submissions. </p>\n\n<p>I am trying to trust my CV and to wisely interpret all the facts containing the data inconsistency and distribution differences. So I can't get why are such Grandmasters overfitting LB so much - maybe that's what we should do too?</p>\n\n<p>You know I am not a pro but I work hard to be as good as possible and it is very frustrating for me not to be able to achieve high score such easly... some words of comfort?</p>",
  "messages": [
    {
      "id": 521115,
      "postDate": "2019-04-22T10:54:56.823Z",
      "content": "<p>Sorry for the pic :)\n<img src=\"https://i.postimg.cc/MKHd4cDS/image-c2682a69-8987-4449-96da-3cf825c18c4a20190422-162043.jpg\" alt=\"because of this\"></p>",
      "rawMarkdown": "Sorry for the pic :)\n![because of this](https://i.postimg.cc/MKHd4cDS/image-c2682a69-8987-4449-96da-3cf825c18c4a20190422-162043.jpg)",
      "votes": 20
    },
    {
      "id": 521131,
      "postDate": "2019-04-22T11:50:36.113Z",
      "content": "<p>One thing GMs are good at: they read carefully everything.  And they look at data, over and over again.</p>",
      "rawMarkdown": "One thing GMs are good at: they read carefully everything.  And they look at data, over and over again.",
      "votes": 15,
      "replies": [
        {
          "id": 521339,
          "postDate": "2019-04-22T19:52:18.693Z",
          "content": "<p>I agree. That's exactly what I'm lacking. Any suggestion how to improve it?</p>\n\n<p>Thank you very much.</p>",
          "rawMarkdown": "I agree. That's exactly what I'm lacking. Any suggestion how to improve it?\n\n  Thank you very much.",
          "votes": 1
        },
        {
          "id": 521815,
          "postDate": "2019-04-23T13:17:17.650Z",
          "content": "<blockquote>\n  <p>Any suggestion how to improve it?</p>\n</blockquote>\n\n<p>Read carefully what top teams share after each competition.  </p>",
          "rawMarkdown": "&gt; Any suggestion how to improve it?\n\nRead carefully what top teams share after each competition.  ",
          "votes": 3
        },
        {
          "id": 521830,
          "postDate": "2019-04-23T13:59:20.983Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> after the competition ends, I will kindly ask you to share the information on how could you get that high so fast, no matter if it's good or not, and also what you've meant writing this post :)</p>",
          "rawMarkdown": "@cpmpml after the competition ends, I will kindly ask you to share the information on how could you get that high so fast, no matter if it's good or not, and also what you've meant writing this post :)",
          "votes": 2
        },
        {
          "id": 521858,
          "postDate": "2019-04-23T14:42:08.380Z",
          "content": "<p>I usually share what I did after competition end.  Here are few examples with some Train/test validation (shameless self promotion here):</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84069\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84069</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75050\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75050</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44614\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44614</a></p>\n\n<p>But you should read all top teams sharing, not just mine, especially how they do CV and what kind of feature engineering they did. Based on questions asked I see readers way too focused on model parameters ( for lg/xgb) or architecture (for NNs) instead.</p>",
          "rawMarkdown": "I usually share what I did after competition end.  Here are few examples with some Train/test validation (shameless self promotion here):\n\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\n\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/84069\n\nhttps://www.kaggle.com/c/PLAsTiCC-2018/discussion/75050\n\nhttps://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44614\n\nBut you should read all top teams sharing, not just mine, especially how they do CV and what kind of feature engineering they did. Based on questions asked I see readers way too focused on model parameters ( for lg/xgb) or architecture (for NNs) instead.",
          "votes": 12
        },
        {
          "id": 521863,
          "postDate": "2019-04-23T14:51:19.737Z",
          "content": "<blockquote>\n  <p>what you've meant writing this post :)</p>\n</blockquote>\n\n<p>Well, for instance when a question answer is in the data page then it means that the one asking the question hasn't read that page carefully.  Also, a lot of great ideas have been shared here.  I'm not sure all have been read and understood and tested by participants.</p>",
          "rawMarkdown": "&gt; what you've meant writing this post :)\n\nWell, for instance when a question answer is in the data page then it means that the one asking the question hasn't read that page carefully.  Also, a lot of great ideas have been shared here.  I'm not sure all have been read and understood and tested by participants.",
          "votes": 3
        },
        {
          "id": 522037,
          "postDate": "2019-04-23T19:34:44.563Z",
          "content": "<p>Thanks for your sharing and explanation. I should learn to read successful kernels.</p>",
          "rawMarkdown": "Thanks for your sharing and explanation. I should learn to read successful kernels.",
          "votes": 1
        },
        {
          "id": 522068,
          "postDate": "2019-04-23T20:25:20.213Z",
          "content": "<p>I took your advice to heart and reread some of the academic papers I downloaded relevant to this competition. The new features I've added have decreased my KFold CV from ~2 to ~1.83. We'll see if it pays off on the leaderboard!</p>",
          "rawMarkdown": "I took your advice to heart and reread some of the academic papers I downloaded relevant to this competition. The new features I've added have decreased my KFold CV from ~2 to ~1.83. We'll see if it pays off on the leaderboard!",
          "votes": 4
        },
        {
          "id": 522082,
          "postDate": "2019-04-23T20:54:10.103Z",
          "content": "<p>That's a big jump!</p>",
          "rawMarkdown": "That's a big jump!",
          "votes": 1
        },
        {
          "id": 526021,
          "postDate": "2019-05-02T07:43:45.120Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Thank you for sharing and the excellent explanation  of these kernels, you help us learn a lot 👍 :)!</p>",
          "rawMarkdown": "@cpmpml Thank you for sharing and the excellent explanation  of these kernels, you help us learn a lot 👍 :)!",
          "votes": 1
        }
      ]
    },
    {
      "id": 520372,
      "postDate": "2019-04-20T20:45:54.510Z",
      "content": "<p>Both @tunguz and @cpmpml got top 40 just after joining the competition, in few submissions. </p>\n\n<p>I am trying to trust my CV and to wisely interpret all the facts containing the data inconsistency and distribution differences. So I can't get why are such Grandmasters overfitting LB so much - maybe that's what we should do too?</p>\n\n<p>You know I am not a pro but I work hard to be as good as possible and it is very frustrating for me not to be able to achieve high score such easly... some words of comfort?</p>",
      "rawMarkdown": "Both @tunguz and @cpmpml got top 40 just after joining the competition, in few submissions. \n\nI am trying to trust my CV and to wisely interpret all the facts containing the data inconsistency and distribution differences. So I can't get why are such Grandmasters overfitting LB so much - maybe that's what we should do too?\n\nYou know I am not a pro but I work hard to be as good as possible and it is very frustrating for me not to be able to achieve high score such easly... some words of comfort?",
      "votes": 16
    },
    {
      "id": 520397,
      "postDate": "2019-04-20T22:25:40.070Z",
      "content": "<p>it is quite common to join later in the competition as most of important details are already in kernels and discussions. I might join later on as well :)</p>",
      "rawMarkdown": "it is quite common to join later in the competition as most of important details are already in kernels and discussions. I might join later on as well :)",
      "votes": 11
    },
    {
      "id": 520885,
      "postDate": "2019-04-21T23:42:16.133Z",
      "content": "<p>Its a mix of experience, talent, hard work, resilience and luck. This happens not only in Kaggle but in all others areas.</p>",
      "rawMarkdown": "Its a mix of experience, talent, hard work, resilience and luck. This happens not only in Kaggle but in all others areas.",
      "votes": 8
    },
    {
      "id": 520765,
      "postDate": "2019-04-21T17:39:57.750Z",
      "content": "<p>You will see GMs coming and teaming with people already high on the LB.  </p>",
      "rawMarkdown": "You will see GMs coming and teaming with people already high on the LB.  ",
      "votes": 6
    },
    {
      "id": 520642,
      "postDate": "2019-04-21T13:22:57.763Z",
      "content": "<p>Seriously, trust your cv people who are leaderboard climbing are going to get bitten in the bottom - expect a huge shake-up on this one.</p>\n\n<p>One thing I notice about GM's is that their two submissions usually involve their best LB score and their best CV score.  Rarely will they choose two highly correlated submissions.</p>",
      "rawMarkdown": "Seriously, trust your cv people who are leaderboard climbing are going to get bitten in the bottom - expect a huge shake-up on this one.\n\nOne thing I notice about GM's is that their two submissions usually involve their best LB score and their best CV score.  Rarely will they choose two highly correlated submissions.\n",
      "votes": 6,
      "replies": [
        {
          "id": 521341,
          "postDate": "2019-04-22T19:54:49.580Z",
          "content": "<p>make sense</p>",
          "rawMarkdown": "make sense",
          "votes": 1
        }
      ]
    },
    {
      "id": 520677,
      "postDate": "2019-04-21T15:01:01.273Z",
      "content": "<p>Experience makes the difference. Also, GM will join late because they might have been focused on other competitions (like Santander).</p>",
      "rawMarkdown": "Experience makes the difference. Also, GM will join late because they might have been focused on other competitions (like Santander).",
      "votes": 3
    },
    {
      "id": 521808,
      "postDate": "2019-04-23T13:08:06.027Z",
      "content": "<p>Here are some words of comfort - You're doing great! Keep up the good work! Frustration is a necessary step towards progress. </p>",
      "rawMarkdown": "Here are some words of comfort - You're doing great! Keep up the good work! Frustration is a necessary step towards progress. ",
      "votes": 4
    },
    {
      "id": 520466,
      "postDate": "2019-04-21T03:43:05.210Z",
      "content": "<p>The number of submissions shown on leaderboard can be quite misleading. That number alone gives no information about the time or effort expended to produce it, which may have been substantial...or not. </p>",
      "rawMarkdown": "The number of submissions shown on leaderboard can be quite misleading. That number alone gives no information about the time or effort expended to produce it, which may have been substantial...or not. ",
      "votes": 4,
      "replies": [
        {
          "id": 525866,
          "postDate": "2019-05-01T21:46:38.957Z",
          "content": "<p>Having many submissions has a down side, because those iterations of parameter tuning based on the test score increase the risk of overfitting the test set.  </p>",
          "rawMarkdown": "Having many submissions has a down side, because those iterations of parameter tuning based on the test score increase the risk of overfitting the test set.  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 520377,
      "postDate": "2019-04-20T21:21:06.480Z",
      "content": "<p>I guess with experience comes intuition</p>",
      "rawMarkdown": "I guess with experience comes intuition",
      "votes": 4,
      "replies": [
        {
          "id": 520544,
          "postDate": "2019-04-21T08:31:18.670Z",
          "content": "<p>And in turn, intuition turns into viable plans to tackle almost any new problem. Solving (or at least trying) many ML problems enhances your creativity as well. </p>",
          "rawMarkdown": "And in turn, intuition turns into viable plans to tackle almost any new problem. Solving (or at least trying) many ML problems enhances your creativity as well. "
        }
      ]
    },
    {
      "id": 520425,
      "postDate": "2019-04-21T01:00:33.547Z",
      "content": "<p>Are you sure you are using a good CV ? It's really important in this competition to choose the right CV scheme. Which one are you currently using ?</p>",
      "rawMarkdown": "Are you sure you are using a good CV ? It's really important in this competition to choose the right CV scheme. Which one are you currently using ?",
      "votes": 3,
      "replies": [
        {
          "id": 520439,
          "postDate": "2019-04-21T01:26:42.797Z",
          "content": "<p>I want to know too. A good validation scheme can ensure the impact of overfitting is minimized. </p>",
          "rawMarkdown": "I want to know too. A good validation scheme can ensure the impact of overfitting is minimized. ",
          "votes": 2
        },
        {
          "id": 520775,
          "postDate": "2019-04-21T17:49:59.417Z",
          "content": "<p>I follow the validations proposed in the discussions. KFold, shuffling, not shuffling, leaving one earthquake out, time series splits, predicting forward, predicting backward, \"stratification\" using one sec bins - I investigate all these to choose the best. I focused on leaving a few earthquakes out, but now I have doubts because of the possibility to predict \"backward\" in time. </p>",
          "rawMarkdown": "I follow the validations proposed in the discussions. KFold, shuffling, not shuffling, leaving one earthquake out, time series splits, predicting forward, predicting backward, \"stratification\" using one sec bins - I investigate all these to choose the best. I focused on leaving a few earthquakes out, but now I have doubts because of the possibility to predict \"backward\" in time. ",
          "votes": 1
        },
        {
          "id": 520807,
          "postDate": "2019-04-21T19:29:18.917Z",
          "content": "<p>I don't think you should be worried about predicting backward in time. We can probably consider each earthquake to be independent from the previous one. Leaving a fews EQ is probably a safe option. Also LB is 13% of test data, so 341 segments, while your CV has 4000 segments.</p>\n\n<p>There is a good discussion about CV here:\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-519922\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-519922</a></p>",
          "rawMarkdown": "I don't think you should be worried about predicting backward in time. We can probably consider each earthquake to be independent from the previous one. Leaving a fews EQ is probably a safe option. Also LB is 13% of test data, so 341 segments, while your CV has 4000 segments.\n\nThere is a good discussion about CV here:\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-519922",
          "votes": 3
        }
      ]
    },
    {
      "id": 520542,
      "postDate": "2019-04-21T08:28:22.867Z",
      "content": "<p>Experience is key I guess. Also, getting a high ranking in Kaggle is correlated with ML (at least applied to competitions) mastery which is a good sign. Keep at it and you will see progress. All is good. :)</p>",
      "rawMarkdown": "Experience is key I guess. Also, getting a high ranking in Kaggle is correlated with ML (at least applied to competitions) mastery which is a good sign. Keep at it and you will see progress. All is good. :)",
      "votes": 1
    },
    {
      "id": 520517,
      "postDate": "2019-04-21T07:16:37.900Z",
      "content": "<p>I am also watching the LB for some time. Earlier it was mostly green, blue and purple. But, the gorgeous oranges and goldens have started shining now. </p>",
      "rawMarkdown": "I am also watching the LB for some time. Earlier it was mostly green, blue and purple. But, the gorgeous oranges and goldens have started shining now. ",
      "votes": 1
    },
    {
      "id": 520395,
      "postDate": "2019-04-20T22:21:48.957Z",
      "content": "<p>There is a clear bottleneck in this competition, that separates very high score from the rest so to speak. If you find a way to deal  with it and overcome it you'll naturally get to top 50-100 if not top 10.\nAt least that's what it looks like to me after having a look at kernels and discussions for a couple of days</p>",
      "rawMarkdown": "There is a clear bottleneck in this competition, that separates very high score from the rest so to speak. If you find a way to deal  with it and overcome it you'll naturally get to top 50-100 if not top 10.\nAt least that's what it looks like to me after having a look at kernels and discussions for a couple of days",
      "votes": 2
    },
    {
      "id": 520380,
      "postDate": "2019-04-20T21:36:27.167Z",
      "content": "<p>Why do you think we are overfitting the LB?</p>",
      "rawMarkdown": "Why do you think we are overfitting the LB?",
      "votes": 2,
      "replies": [
        {
          "id": 520384,
          "postDate": "2019-04-20T21:45:31.240Z",
          "content": "<p>Excuse me - it's a blunder. I meant \"fitting LB\". Overfitting is implicit - it's a common opinion that distributions of 13% LB and train data are different, so I assume a high score on LB equals overfitting, but it's just suspicion.</p>",
          "rawMarkdown": "Excuse me - it's a blunder. I meant \"fitting LB\". Overfitting is implicit - it's a common opinion that distributions of 13% LB and train data are different, so I assume a high score on LB equals overfitting, but it's just suspicion.",
          "votes": 1
        },
        {
          "id": 520396,
          "postDate": "2019-04-20T22:23:46.693Z",
          "content": "<p>Not necessarily. Boyan has recently scored well on VSB Powerline Competition and even trolled people who overfitted on public LB afterwards. This challenge is way less irregular and crazy than VSB in terms of Data, IMO</p>",
          "rawMarkdown": "Not necessarily. Boyan has recently scored well on VSB Powerline Competition and even trolled people who overfitted on public LB afterwards. This challenge is way less irregular and crazy than VSB in terms of Data, IMO",
          "votes": 1
        },
        {
          "id": 520524,
          "postDate": "2019-04-21T07:30:21.007Z",
          "content": "<p>I asked because we will know who overfits only when we will see the private LB.</p>\n\n<p>Regarding the time it took me to be where I am: one day of experiment after reading a bit for few days.  But I may be stuck there for a long time, this dataset is tricky.</p>",
          "rawMarkdown": "I asked because we will know who overfits only when we will see the private LB.\n\nRegarding the time it took me to be where I am: one day of experiment after reading a bit for few days.  But I may be stuck there for a long time, this dataset is tricky.",
          "votes": 2
        },
        {
          "id": 520525,
          "postDate": "2019-04-21T07:31:47.863Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 520819,
          "postDate": "2019-04-21T19:55:44.473Z",
          "content": "<blockquote>\n  <p>I may be stuck there for a long time</p>\n</blockquote>\n\n<p>Well we have believe In your work that you won't :)</p>",
          "rawMarkdown": "&gt;I may be stuck there for a long time\n\nWell we have believe In your work that you won't :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 540086,
      "postDate": "2019-05-31T02:26:33.310Z",
      "content": "<p>Apologies for the late posting.</p>\n\n<p>I finished a Kernel for you to look over.</p>\n\n<p>This Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.</p>\n\n<p>This Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)</p>\n\n<p>Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01</a></p>\n\n<p>Enjoy, and good luck!</p>",
      "rawMarkdown": "Apologies for the late posting.\n\nI finished a Kernel for you to look over.\n\nThis Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.\n\nThis Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)\n\nGets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\n\nEnjoy, and good luck!"
    },
    {
      "id": 520420,
      "postDate": "2019-04-21T00:24:57.713Z",
      "content": "<p><a href=\"/davids1992\">@davids1992</a>, please don't get frustrated.</p>",
      "rawMarkdown": "@davids1992, please don't get frustrated."
    }
  ],
  "comments": [
    {
      "id": 521115,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2019-04-22T10:54:56.823000",
      "content": "<p>Sorry for the pic :)\n<img src=\"https://i.postimg.cc/MKHd4cDS/image-c2682a69-8987-4449-96da-3cf825c18c4a20190422-162043.jpg\" alt=\"because of this\"></p>",
      "votes": 20,
      "replies": []
    },
    {
      "id": 521131,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-22T11:50:36.113000",
      "content": "<p>One thing GMs are good at: they read carefully everything.  And they look at data, over and over again.</p>",
      "votes": 15,
      "replies": [
        {
          "id": 521339,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-22T19:52:18.693000",
          "content": "<p>I agree. That's exactly what I'm lacking. Any suggestion how to improve it?</p>\n\n<p>Thank you very much.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 521815,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T13:17:17.650000",
          "content": "<blockquote>\n  <p>Any suggestion how to improve it?</p>\n</blockquote>\n\n<p>Read carefully what top teams share after each competition.  </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 521830,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-23T13:59:20.983000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> after the competition ends, I will kindly ask you to share the information on how could you get that high so fast, no matter if it's good or not, and also what you've meant writing this post :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 521858,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T14:42:08.380000",
          "content": "<p>I usually share what I did after competition end.  Here are few examples with some Train/test validation (shameless self promotion here):</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84069\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84069</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75050\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75050</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44614\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44614</a></p>\n\n<p>But you should read all top teams sharing, not just mine, especially how they do CV and what kind of feature engineering they did. Based on questions asked I see readers way too focused on model parameters ( for lg/xgb) or architecture (for NNs) instead.</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 521863,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-23T14:51:19.737000",
          "content": "<blockquote>\n  <p>what you've meant writing this post :)</p>\n</blockquote>\n\n<p>Well, for instance when a question answer is in the data page then it means that the one asking the question hasn't read that page carefully.  Also, a lot of great ideas have been shared here.  I'm not sure all have been read and understood and tested by participants.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 522037,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-23T19:34:44.563000",
          "content": "<p>Thanks for your sharing and explanation. I should learn to read successful kernels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 522068,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-04-23T20:25:20.213000",
          "content": "<p>I took your advice to heart and reread some of the academic papers I downloaded relevant to this competition. The new features I've added have decreased my KFold CV from ~2 to ~1.83. We'll see if it pays off on the leaderboard!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 522082,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-23T20:54:10.103000",
          "content": "<p>That's a big jump!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 526021,
          "author_name": "Beans",
          "author_url": "",
          "post_date": "2019-05-02T07:43:45.120000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Thank you for sharing and the excellent explanation  of these kernels, you help us learn a lot 👍 :)!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 520397,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2019-04-20T22:25:40.070000",
      "content": "<p>it is quite common to join later in the competition as most of important details are already in kernels and discussions. I might join later on as well :)</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 520885,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-04-21T23:42:16.133000",
      "content": "<p>Its a mix of experience, talent, hard work, resilience and luck. This happens not only in Kaggle but in all others areas.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 520765,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-21T17:39:57.750000",
      "content": "<p>You will see GMs coming and teaming with people already high on the LB.  </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 520642,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-04-21T13:22:57.763000",
      "content": "<p>Seriously, trust your cv people who are leaderboard climbing are going to get bitten in the bottom - expect a huge shake-up on this one.</p>\n\n<p>One thing I notice about GM's is that their two submissions usually involve their best LB score and their best CV score.  Rarely will they choose two highly correlated submissions.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 521341,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-22T19:54:49.580000",
          "content": "<p>make sense</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 520677,
      "author_name": "Gabriel Preda",
      "author_url": "",
      "post_date": "2019-04-21T15:01:01.273000",
      "content": "<p>Experience makes the difference. Also, GM will join late because they might have been focused on other competitions (like Santander).</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 521808,
      "author_name": "Paul Nussbaum, PhD",
      "author_url": "",
      "post_date": "2019-04-23T13:08:06.027000",
      "content": "<p>Here are some words of comfort - You're doing great! Keep up the good work! Frustration is a necessary step towards progress. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 520466,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-04-21T03:43:05.210000",
      "content": "<p>The number of submissions shown on leaderboard can be quite misleading. That number alone gives no information about the time or effort expended to produce it, which may have been substantial...or not. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 525866,
          "author_name": "Filip Mulier",
          "author_url": "",
          "post_date": "2019-05-01T21:46:38.957000",
          "content": "<p>Having many submissions has a down side, because those iterations of parameter tuning based on the test score increase the risk of overfitting the test set.  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 520377,
      "author_name": "Amjad",
      "author_url": "",
      "post_date": "2019-04-20T21:21:06.480000",
      "content": "<p>I guess with experience comes intuition</p>",
      "votes": 4,
      "replies": [
        {
          "id": 520544,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2019-04-21T08:31:18.670000",
          "content": "<p>And in turn, intuition turns into viable plans to tackle almost any new problem. Solving (or at least trying) many ML problems enhances your creativity as well. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 520425,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2019-04-21T01:00:33.547000",
      "content": "<p>Are you sure you are using a good CV ? It's really important in this competition to choose the right CV scheme. Which one are you currently using ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 520439,
          "author_name": "FP",
          "author_url": "",
          "post_date": "2019-04-21T01:26:42.797000",
          "content": "<p>I want to know too. A good validation scheme can ensure the impact of overfitting is minimized. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 520775,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-21T17:49:59.417000",
          "content": "<p>I follow the validations proposed in the discussions. KFold, shuffling, not shuffling, leaving one earthquake out, time series splits, predicting forward, predicting backward, \"stratification\" using one sec bins - I investigate all these to choose the best. I focused on leaving a few earthquakes out, but now I have doubts because of the possibility to predict \"backward\" in time. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 520807,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2019-04-21T19:29:18.917000",
          "content": "<p>I don't think you should be worried about predicting backward in time. We can probably consider each earthquake to be independent from the previous one. Leaving a fews EQ is probably a safe option. Also LB is 13% of test data, so 341 segments, while your CV has 4000 segments.</p>\n\n<p>There is a good discussion about CV here:\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-519922\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-519922</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 520542,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2019-04-21T08:28:22.867000",
      "content": "<p>Experience is key I guess. Also, getting a high ranking in Kaggle is correlated with ML (at least applied to competitions) mastery which is a good sign. Keep at it and you will see progress. All is good. :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 520517,
      "author_name": "arnab",
      "author_url": "",
      "post_date": "2019-04-21T07:16:37.900000",
      "content": "<p>I am also watching the LB for some time. Earlier it was mostly green, blue and purple. But, the gorgeous oranges and goldens have started shining now. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 520395,
      "author_name": "George Okromchedlishvili",
      "author_url": "",
      "post_date": "2019-04-20T22:21:48.957000",
      "content": "<p>There is a clear bottleneck in this competition, that separates very high score from the rest so to speak. If you find a way to deal  with it and overcome it you'll naturally get to top 50-100 if not top 10.\nAt least that's what it looks like to me after having a look at kernels and discussions for a couple of days</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 520380,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-04-20T21:36:27.167000",
      "content": "<p>Why do you think we are overfitting the LB?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 520384,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2019-04-20T21:45:31.240000",
          "content": "<p>Excuse me - it's a blunder. I meant \"fitting LB\". Overfitting is implicit - it's a common opinion that distributions of 13% LB and train data are different, so I assume a high score on LB equals overfitting, but it's just suspicion.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 520396,
          "author_name": "George Okromchedlishvili",
          "author_url": "",
          "post_date": "2019-04-20T22:23:46.693000",
          "content": "<p>Not necessarily. Boyan has recently scored well on VSB Powerline Competition and even trolled people who overfitted on public LB afterwards. This challenge is way less irregular and crazy than VSB in terms of Data, IMO</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 520524,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-04-21T07:30:21.007000",
          "content": "<p>I asked because we will know who overfits only when we will see the private LB.</p>\n\n<p>Regarding the time it took me to be where I am: one day of experiment after reading a bit for few days.  But I may be stuck there for a long time, this dataset is tricky.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 520525,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-21T07:31:47.863000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 520819,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2019-04-21T19:55:44.473000",
          "content": "<blockquote>\n  <p>I may be stuck there for a long time</p>\n</blockquote>\n\n<p>Well we have believe In your work that you won't :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 540086,
      "author_name": "Paul Nussbaum, PhD",
      "author_url": "",
      "post_date": "2019-05-31T02:26:33.310000",
      "content": "<p>Apologies for the late posting.</p>\n\n<p>I finished a Kernel for you to look over.</p>\n\n<p>This Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.</p>\n\n<p>This Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)</p>\n\n<p>Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01</a></p>\n\n<p>Enjoy, and good luck!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 520420,
      "author_name": "FP",
      "author_url": "",
      "post_date": "2019-04-21T00:24:57.713000",
      "content": "<p><a href=\"/davids1992\">@davids1992</a>, please don't get frustrated.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "521115": "Sorry for the pic :)\n![because of this](https://i.postimg.cc/MKHd4cDS/image-c2682a69-8987-4449-96da-3cf825c18c4a20190422-162043.jpg)",
    "521131": "One thing GMs are good at: they read carefully everything.  And they look at data, over and over again.",
    "520372": "Both @tunguz and @cpmpml got top 40 just after joining the competition, in few submissions. \n\nI am trying to trust my CV and to wisely interpret all the facts containing the data inconsistency and distribution differences. So I can't get why are such Grandmasters overfitting LB so much - maybe that's what we should do too?\n\nYou know I am not a pro but I work hard to be as good as possible and it is very frustrating for me not to be able to achieve high score such easly... some words of comfort?",
    "520397": "it is quite common to join later in the competition as most of important details are already in kernels and discussions. I might join later on as well :)",
    "520885": "Its a mix of experience, talent, hard work, resilience and luck. This happens not only in Kaggle but in all others areas.",
    "520765": "You will see GMs coming and teaming with people already high on the LB.  ",
    "520642": "Seriously, trust your cv people who are leaderboard climbing are going to get bitten in the bottom - expect a huge shake-up on this one.\n\nOne thing I notice about GM's is that their two submissions usually involve their best LB score and their best CV score.  Rarely will they choose two highly correlated submissions.\n",
    "520677": "Experience makes the difference. Also, GM will join late because they might have been focused on other competitions (like Santander).",
    "521808": "Here are some words of comfort - You're doing great! Keep up the good work! Frustration is a necessary step towards progress. ",
    "520466": "The number of submissions shown on leaderboard can be quite misleading. That number alone gives no information about the time or effort expended to produce it, which may have been substantial...or not. ",
    "520377": "I guess with experience comes intuition",
    "520425": "Are you sure you are using a good CV ? It's really important in this competition to choose the right CV scheme. Which one are you currently using ?",
    "520542": "Experience is key I guess. Also, getting a high ranking in Kaggle is correlated with ML (at least applied to competitions) mastery which is a good sign. Keep at it and you will see progress. All is good. :)",
    "520517": "I am also watching the LB for some time. Earlier it was mostly green, blue and purple. But, the gorgeous oranges and goldens have started shining now. ",
    "520395": "There is a clear bottleneck in this competition, that separates very high score from the rest so to speak. If you find a way to deal  with it and overcome it you'll naturally get to top 50-100 if not top 10.\nAt least that's what it looks like to me after having a look at kernels and discussions for a couple of days",
    "520380": "Why do you think we are overfitting the LB?",
    "540086": "Apologies for the late posting.\n\nI finished a Kernel for you to look over.\n\nThis Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.\n\nThis Kernel also demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)\n\nGets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\n\nEnjoy, and good luck!",
    "520420": "@davids1992, please don't get frustrated."
  }
}