{
  "id": 94411,
  "title": "Disappointed by Kaggle",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94411",
  "author_name": "nosound",
  "post_date": "2019-06-04T10:27:02.252000",
  "votes": 58,
  "comment_count": 79,
  "views": 0,
  "content": "<p>This competition was poorly organized. </p>\n\n<ol>\n<li>The major blunder is that the test part durations can be found online, as described in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">here</a>. It is exacerbated by the fact that the organizers neither reported it in the official information nor confirmed it later on. I don't think Kaggle wants its competitions to be about noticing or not that some article online contains a picture of the test set distribution. </li>\n<li>If the test set distribution was not intended to be known (which I assume is true), then I don't think such huge difference between the train and the test is justified and can be predicted, - bad competition data design. A simple multiplication by a factor of random model predictions behaves better than anything that can be learnt, as illustrated by this <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-542970\">solution</a>.</li>\n<li>The organizers stopped responding to legitimate participants questions in the official thread about 3-4 months before the competition end. It feels to me like \"well, we messed up here. Let's hide our head into the sand!\".</li>\n<li>Further displaying the organizers sloppiness are small facts that in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#latest-541180\">additional info</a> post 12 millisecond was edited to 12 microseconds at some point silently, leaving people perplexed, and the<a href=\"https://www.kaggle.com/inversion/basic-feature-benchmark\"> official benchmark kernel</a> still uses R2 metric apparently used during the competition preparation.</li>\n<li>Public test size of 348 is so small, that it is possible to discover public vs private segments in less than 90 submissions (will share about it later), giving further advantage to people who do it.</li>\n</ol>\n\n<p>I learnt a lot during the competition, huge thanks to all the forum discussion participants, and people sharing their work in kernels! Really priceless. And I hope that like me and you, Kaggle as an organization will learn from it as well.</p>\n\n<p>UPDATE: after reading the write-ups of the winning solutions and other forum discussions I backtrack and admit that point 2 is not a valid point. It was possible and optimal to score high without using the test set distribution, <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390\">reference to 1st place solution</a>. But only few did that correctly, or even adjusted to test distribution at all, and that is why lucky \"random\" solutions advanced so much. Kudos to few of us who did it correctly, it is inspiring.</p>",
  "messages": [
    {
      "id": 543100,
      "postDate": "2019-06-04T10:27:02.253Z",
      "content": "<p>This competition was poorly organized. </p>\n\n<ol>\n<li>The major blunder is that the test part durations can be found online, as described in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844\">here</a>. It is exacerbated by the fact that the organizers neither reported it in the official information nor confirmed it later on. I don't think Kaggle wants its competitions to be about noticing or not that some article online contains a picture of the test set distribution. </li>\n<li>If the test set distribution was not intended to be known (which I assume is true), then I don't think such huge difference between the train and the test is justified and can be predicted, - bad competition data design. A simple multiplication by a factor of random model predictions behaves better than anything that can be learnt, as illustrated by this <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-542970\">solution</a>.</li>\n<li>The organizers stopped responding to legitimate participants questions in the official thread about 3-4 months before the competition end. It feels to me like \"well, we messed up here. Let's hide our head into the sand!\".</li>\n<li>Further displaying the organizers sloppiness are small facts that in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#latest-541180\">additional info</a> post 12 millisecond was edited to 12 microseconds at some point silently, leaving people perplexed, and the<a href=\"https://www.kaggle.com/inversion/basic-feature-benchmark\"> official benchmark kernel</a> still uses R2 metric apparently used during the competition preparation.</li>\n<li>Public test size of 348 is so small, that it is possible to discover public vs private segments in less than 90 submissions (will share about it later), giving further advantage to people who do it.</li>\n</ol>\n\n<p>I learnt a lot during the competition, huge thanks to all the forum discussion participants, and people sharing their work in kernels! Really priceless. And I hope that like me and you, Kaggle as an organization will learn from it as well.</p>\n\n<p>UPDATE: after reading the write-ups of the winning solutions and other forum discussions I backtrack and admit that point 2 is not a valid point. It was possible and optimal to score high without using the test set distribution, <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390\">reference to 1st place solution</a>. But only few did that correctly, or even adjusted to test distribution at all, and that is why lucky \"random\" solutions advanced so much. Kudos to few of us who did it correctly, it is inspiring.</p>",
      "rawMarkdown": "This competition was poorly organized. \n\n1. The major blunder is that the test part durations can be found online, as described in [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844). It is exacerbated by the fact that the organizers neither reported it in the official information nor confirmed it later on. I don't think Kaggle wants its competitions to be about noticing or not that some article online contains a picture of the test set distribution. \n2. If the test set distribution was not intended to be known (which I assume is true), then I don't think such huge difference between the train and the test is justified and can be predicted, - bad competition data design. A simple multiplication by a factor of random model predictions behaves better than anything that can be learnt, as illustrated by this [solution](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-542970).\n3. The organizers stopped responding to legitimate participants questions in the official thread about 3-4 months before the competition end. It feels to me like \"well, we messed up here. Let's hide our head into the sand!\".\n4. Further displaying the organizers sloppiness are small facts that in [additional info](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#latest-541180) post 12 millisecond was edited to 12 microseconds at some point silently, leaving people perplexed, and the[ official benchmark kernel](https://www.kaggle.com/inversion/basic-feature-benchmark) still uses R2 metric apparently used during the competition preparation.\n5. Public test size of 348 is so small, that it is possible to discover public vs private segments in less than 90 submissions (will share about it later), giving further advantage to people who do it.\n\nI learnt a lot during the competition, huge thanks to all the forum discussion participants, and people sharing their work in kernels! Really priceless. And I hope that like me and you, Kaggle as an organization will learn from it as well.\n\nUPDATE: after reading the write-ups of the winning solutions and other forum discussions I backtrack and admit that point 2 is not a valid point. It was possible and optimal to score high without using the test set distribution, [reference to 1st place solution](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390). But only few did that correctly, or even adjusted to test distribution at all, and that is why lucky \"random\" solutions advanced so much. Kudos to few of us who did it correctly, it is inspiring.",
      "votes": 57
    },
    {
      "id": 543194,
      "postDate": "2019-06-04T11:21:03.623Z",
      "content": "<p>What I find really disappointing is the lack of follow up from organizers.  it was the same in recent research competitions I entered.  I see a difference with commercial sponsor who look way more interested in what we can do to help them.</p>",
      "rawMarkdown": "What I find really disappointing is the lack of follow up from organizers.  it was the same in recent research competitions I entered.  I see a difference with commercial sponsor who look way more interested in what we can do to help them.",
      "votes": 20,
      "replies": [
        {
          "id": 543214,
          "postDate": "2019-06-04T11:33:02.097Z",
          "content": "<p>I found it really odd how communication died with months to go we aren't talking about a couple of weeks for a holiday, conference etc.  The same has happened with product feedback - maybe Kaggle has a staff shortage.</p>\n\n<p>The good thing about this competition at least shows that metrics R2 and MAE may not be the best metrics.  If that helps better predictions then great.</p>",
          "rawMarkdown": "I found it really odd how communication died with months to go we aren't talking about a couple of weeks for a holiday, conference etc.  The same has happened with product feedback - maybe Kaggle has a staff shortage.\n\nThe good thing about this competition at least shows that metrics R2 and MAE may not be the best metrics.  If that helps better predictions then great.",
          "votes": 4
        },
        {
          "id": 543495,
          "postDate": "2019-06-04T15:07:29.123Z",
          "content": "<p>It surprises me though, usually scientists are much more interested in what is going on.</p>",
          "rawMarkdown": "It surprises me though, usually scientists are much more interested in what is going on.",
          "votes": 3
        },
        {
          "id": 543523,
          "postDate": "2019-06-04T15:18:42.357Z",
          "content": "<p>My guess is, they never wanted us to find their original paper on exp 4677. As soon as this was out, the results were no longer of value for them. \nI mean, they linked 3 papers but not the one containing the experiment 4677? That is no coincidence.</p>",
          "rawMarkdown": "My guess is, they never wanted us to find their original paper on exp 4677. As soon as this was out, the results were no longer of value for them. \nI mean, they linked 3 papers but not the one containing the experiment 4677? That is no coincidence.",
          "votes": 4
        },
        {
          "id": 543527,
          "postDate": "2019-06-04T15:21:42.760Z",
          "content": "<p>The paper with exp 4677 is not from same LANL authors as the ones that were shared.</p>",
          "rawMarkdown": "The paper with exp 4677 is not from same LANL authors as the ones that were shared.",
          "votes": 2
        },
        {
          "id": 543534,
          "postDate": "2019-06-04T15:23:57.460Z",
          "content": "<p>They don't need to bother.  The training data only has 15 earthquakes.  It's really hard to judge any of predicting methods are good or bad. </p>",
          "rawMarkdown": "They don't need to bother.  The training data only has 15 earthquakes.  It's really hard to judge any of predicting methods are good or bad. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 543626,
      "postDate": "2019-06-04T16:35:32.603Z",
      "content": "<p>I wonder why don’t they just randomly pick the shuffled segments from the whole length of data, 1 set for train, 1 for public, and 1 for private? That will make the competition authentic. The data preparation is poor, supporting host is even poorer. Sad for a 4500 contestant competition. </p>",
      "rawMarkdown": "I wonder why don’t they just randomly pick the shuffled segments from the whole length of data, 1 set for train, 1 for public, and 1 for private? That will make the competition authentic. The data preparation is poor, supporting host is even poorer. Sad for a 4500 contestant competition. ",
      "votes": 11,
      "replies": [
        {
          "id": 543631,
          "postDate": "2019-06-04T16:39:10.593Z",
          "content": "<p>Come on now they have only had 5 months ;P</p>",
          "rawMarkdown": "Come on now they have only had 5 months ;P",
          "votes": 4
        },
        {
          "id": 543763,
          "postDate": "2019-06-04T19:42:09.347Z",
          "content": "<p>Could not agree more. </p>",
          "rawMarkdown": "Could not agree more. ",
          "votes": 2
        },
        {
          "id": 546072,
          "postDate": "2019-06-06T08:21:28.120Z",
          "content": "<p>Or, they can scale up our predictions and give out second batch of medals ... </p>",
          "rawMarkdown": "Or, they can scale up our predictions and give out second batch of medals ... ",
          "votes": 1
        }
      ]
    },
    {
      "id": 543536,
      "postDate": "2019-06-04T15:24:33.140Z",
      "content": "<p>I can't agree more with 1. This is a big leakage and for sure changed competitions results. </p>",
      "rawMarkdown": "I can't agree more with 1. This is a big leakage and for sure changed competitions results. ",
      "votes": 9,
      "replies": [
        {
          "id": 543551,
          "postDate": "2019-06-04T15:34:54.570Z",
          "content": "<p>This is the first competition I know of that benefited from using rulers on print-outs for sure!</p>",
          "rawMarkdown": "This is the first competition I know of that benefited from using rulers on print-outs for sure!",
          "votes": 7
        },
        {
          "id": 543556,
          "postDate": "2019-06-04T15:39:19.587Z",
          "content": "<p>It is definetly a kind of leakage, but as you can see from our explanation we didn't use that information directly in our model. We optimized to oof/predictions similarity to account for the large train/test difference (for which the the data available at kaggle was already enough).</p>\n\n<p><a href=\"/scirpus\">@scirpus</a> we did use our rulers, but in the end we didn't use it for our final models</p>",
          "rawMarkdown": "It is definetly a kind of leakage, but as you can see from our explanation we didn't use that information directly in our model. We optimized to oof/predictions similarity to account for the large train/test difference (for which the the data available at kaggle was already enough).\n\n@scirpus we did use our rulers, but in the end we didn't use it for our final models",
          "votes": 3
        },
        {
          "id": 543562,
          "postDate": "2019-06-04T15:42:07.840Z",
          "content": "<p>I noticed your CV gives a mean of about 6.2 - so you did it the right way - the fact people can just scale it to 6.2 and do well from the paper is what people are sore about.  Congrats by the way.</p>",
          "rawMarkdown": "I noticed your CV gives a mean of about 6.2 - so you did it the right way - the fact people can just scale it to 6.2 and do well from the paper is what people are sore about.  Congrats by the way.",
          "votes": 2
        },
        {
          "id": 543567,
          "postDate": "2019-06-04T15:44:40.080Z",
          "content": "<p>I get you point, no worries. \nStill, up to now nobody published a better score with \"simple upscaling\"\nSo upscaling alone doesn't make a great model?</p>",
          "rawMarkdown": "I get you point, no worries. \nStill, up to now nobody published a better score with \"simple upscaling\"\nSo upscaling alone doesn't make a great model?",
          "votes": 1
        },
        {
          "id": 543573,
          "postDate": "2019-06-04T15:46:38.797Z",
          "content": "<p>You can get a gold medal just by picking a random model and multiplying the ttf's by ~ 1.1 to 1.2\nI may sound salty but I used this competition to finally get rid of all my Genetic Programming bugs from my code so it is celebration time for me! ;)  It has only taken 4 years!</p>",
          "rawMarkdown": "You can get a gold medal just by picking a random model and multiplying the ttf's by ~ 1.1 to 1.2\nI may sound salty but I used this competition to finally get rid of all my Genetic Programming bugs from my code so it is celebration time for me! ;)  It has only taken 4 years!",
          "votes": 1
        },
        {
          "id": 543575,
          "postDate": "2019-06-04T15:47:46.967Z",
          "content": "<p>That was very predictable and i expected many more people to do exactly that. And then it would again come down to the rest of the model and not \"just\" the mean.</p>",
          "rawMarkdown": "That was very predictable and i expected many more people to do exactly that. And then it would again come down to the rest of the model and not \"just\" the mean.\n"
        },
        {
          "id": 543621,
          "postDate": "2019-06-04T16:30:07.203Z",
          "content": "<p>Haha.. sometimes you don't need fancy stuff to get a good result. A simple ruler can help ;)\nI work in industry. I don't care much about the methods but about results. Many times I see people spending a lot of energy and time using very complex and fancy methods, and even when they present their results, they focus more on how they did than what they achieved. I am more practical. If I can get a good result using a linear regression I will do it. I could have certainly measured the length of test cycles with a software, but measuring it with a ruler took me 1 minute and it was a good starting point. It worked ;)</p>\n\n<p>I also used the length of the cycles in the paper as the first approach, but then I selected subsets of the training data to match better the distributions of the features between train and test.</p>\n\n<p>Even though I benefited here from the leakage (and despite it was disclosed for all competitors), I am also dissapointed that the 3 competitions I have participated on have been won by those who found a way to exploit a leakage. I really hope I can participate in a leakage-free competition in the future.</p>",
          "rawMarkdown": "Haha.. sometimes you don't need fancy stuff to get a good result. A simple ruler can help ;)\nI work in industry. I don't care much about the methods but about results. Many times I see people spending a lot of energy and time using very complex and fancy methods, and even when they present their results, they focus more on how they did than what they achieved. I am more practical. If I can get a good result using a linear regression I will do it. I could have certainly measured the length of test cycles with a software, but measuring it with a ruler took me 1 minute and it was a good starting point. It worked ;)\n\nI also used the length of the cycles in the paper as the first approach, but then I selected subsets of the training data to match better the distributions of the features between train and test.\n\nEven though I benefited here from the leakage (and despite it was disclosed for all competitors), I am also dissapointed that the 3 competitions I have participated on have been won by those who found a way to exploit a leakage. I really hope I can participate in a leakage-free competition in the future.",
          "votes": 5
        },
        {
          "id": 543686,
          "postDate": "2019-06-04T17:42:40.597Z",
          "content": "<blockquote>\n  <p>I may sound salty but I used this competition to finally get rid of all my Genetic Programming bugs from my code so it is celebration time for me! ;) It has only taken 4 years!</p>\n</blockquote>\n\n<p>Congrats! But I remember your program worked well in the past competitons, e.g. Plasticc: the famous Scirpus' formula! :)</p>",
          "rawMarkdown": "&gt; I may sound salty but I used this competition to finally get rid of all my Genetic Programming bugs from my code so it is celebration time for me! ;) It has only taken 4 years!\n\nCongrats! But I remember your program worked well in the past competitons, e.g. Plasticc: the famous Scirpus' formula! :)",
          "votes": 2
        },
        {
          "id": 543703,
          "postDate": "2019-06-04T17:56:46.417Z",
          "content": "<p>I used MS Paint and Excel to find the mean :)</p>",
          "rawMarkdown": "I used MS Paint and Excel to find the mean :)",
          "votes": 9
        },
        {
          "id": 543705,
          "postDate": "2019-06-04T17:58:45.670Z",
          "content": "<p>Image here</p>",
          "rawMarkdown": "Image here",
          "votes": 5
        },
        {
          "id": 543749,
          "postDate": "2019-06-04T19:02:32.293Z",
          "content": "<blockquote>\n  <p>Image here</p>\n</blockquote>\n\n<p>Cool, thanks! Now it's clear! :)</p>",
          "rawMarkdown": "&gt; Image here\n\nCool, thanks! Now it's clear! :)"
        },
        {
          "id": 543772,
          "postDate": "2019-06-04T20:02:47.367Z",
          "content": "<p>i actually used a \"data from image\" extraction tool ;)</p>",
          "rawMarkdown": "i actually used a \"data from image\" extraction tool ;)",
          "votes": 3
        },
        {
          "id": 543774,
          "postDate": "2019-06-04T20:07:01.620Z",
          "content": "<p>Ooh interesting - tell us more</p>",
          "rawMarkdown": "Ooh interesting - tell us more",
          "votes": 1
        },
        {
          "id": 543777,
          "postDate": "2019-06-04T20:10:32.443Z",
          "content": "<p><a href=\"/scirpus\">@scirpus</a>  <a href=\"https://automeris.io/WebPlotDigitizer/\">https://automeris.io/WebPlotDigitizer/</a> this is one online option for it</p>\n\n<p>Can get in handy with non-linear data where you are only provided a picture but no tabulated data.</p>",
          "rawMarkdown": "@scirpus  https://automeris.io/WebPlotDigitizer/ this is one online option for it\n\nCan get in handy with non-linear data where you are only provided a picture but no tabulated data.",
          "votes": 2
        },
        {
          "id": 543784,
          "postDate": "2019-06-04T20:21:07.857Z",
          "content": "<p>Thankyou</p>",
          "rawMarkdown": "Thankyou",
          "votes": 1
        },
        {
          "id": 543787,
          "postDate": "2019-06-04T20:23:20.080Z",
          "content": "<p>Thanks for sharing <a href=\"/ilu000\">@ilu000</a> </p>",
          "rawMarkdown": "Thanks for sharing @ilu000 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 543809,
      "postDate": "2019-06-04T21:08:12.470Z",
      "content": "<p>I don't think people should be too upset about the data leak from the paper. The time series raw data shown in the paper is a whole lot longer than just the train and test data. While it ended up that the test data was coming from the exact region labeled as test in the paper, that was far from a guarantee. The organizers could have easily taken the test data from either before the train data, or a little bit further down the road after the train data. That would lead to a different distribution, and probably very different leaderboard.</p>\n\n<p>I believe the teams that exploited the data shown in the paper took an educated risk, and congratulations to them that it paid off for this competition. But had the organizers picked differently, it would have been a losing strategy.</p>",
      "rawMarkdown": "I don't think people should be too upset about the data leak from the paper. The time series raw data shown in the paper is a whole lot longer than just the train and test data. While it ended up that the test data was coming from the exact region labeled as test in the paper, that was far from a guarantee. The organizers could have easily taken the test data from either before the train data, or a little bit further down the road after the train data. That would lead to a different distribution, and probably very different leaderboard.\n\nI believe the teams that exploited the data shown in the paper took an educated risk, and congratulations to them that it paid off for this competition. But had the organizers picked differently, it would have been a losing strategy.",
      "votes": 3
    },
    {
      "id": 543153,
      "postDate": "2019-06-04T11:00:31.577Z",
      "content": "<p>At least, here, everyone had the info about test and could work with it.\nIt doesnt apply to real world problems, but thats the issue with many kaggle challenges.</p>",
      "rawMarkdown": "At least, here, everyone had the info about test and could work with it.\nIt doesnt apply to real world problems, but thats the issue with many kaggle challenges.",
      "votes": 3,
      "replies": [
        {
          "id": 543171,
          "postDate": "2019-06-04T11:09:42.483Z",
          "content": "<p>In case the test distribution is known to everyone I agree it becomes an interesting problem how to transfer learn from train to test. What worries me is that it was not by design. Just imagine we didn't know the test distribution, - private LB in this case would have been completely random, instead of 80% random now.</p>",
          "rawMarkdown": "In case the test distribution is known to everyone I agree it becomes an interesting problem how to transfer learn from train to test. What worries me is that it was not by design. Just imagine we didn't know the test distribution, - private LB in this case would have been completely random, instead of 80% random now.",
          "votes": 1
        },
        {
          "id": 543173,
          "postDate": "2019-06-04T11:11:56.307Z",
          "content": "<p>I agree but our solution basically does not need to use information from the paper on test distributions but we rather used the test data itself and feature dists to better match the training data.</p>",
          "rawMarkdown": "I agree but our solution basically does not need to use information from the paper on test distributions but we rather used the test data itself and feature dists to better match the training data.",
          "votes": 6
        },
        {
          "id": 543187,
          "postDate": "2019-06-04T11:18:56.407Z",
          "content": "<p>I think the word \"prediction\" does not mean much if you need the test info!</p>",
          "rawMarkdown": "I think the word \"prediction\" does not mean much if you need the test info!",
          "votes": 4
        },
        {
          "id": 543189,
          "postDate": "2019-06-04T11:19:51.947Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> is right, and apparently this solution was even superior to \"simple upscaling\" as many have done. \nPeaking into test is always a problem that we see in kaggle challenges, and even cannot be prevented in Kernel only challenges as <a href=\"/cdeotte\">@cdeotte</a> has shown lately. </p>",
          "rawMarkdown": "@philippsinger is right, and apparently this solution was even superior to \"simple upscaling\" as many have done. \nPeaking into test is always a problem that we see in kaggle challenges, and even cannot be prevented in Kernel only challenges as @cdeotte has shown lately. ",
          "votes": 3
        },
        {
          "id": 543499,
          "postDate": "2019-06-04T15:08:44.403Z",
          "content": "<p>Looking at test set data distribution then go back to re-select training data is not good practice.  If you deploy a ML system, you don't really know what kind of data will come in.  One should build a good generic model instead of a model fitting one set of data.  They should have gotten test set from random fragments from a number of waves (simulations).  </p>",
          "rawMarkdown": "Looking at test set data distribution then go back to re-select training data is not good practice.  If you deploy a ML system, you don't really know what kind of data will come in.  One should build a good generic model instead of a model fitting one set of data.  They should have gotten test set from random fragments from a number of waves (simulations).  ",
          "votes": 1
        },
        {
          "id": 543512,
          "postDate": "2019-06-04T15:13:39.607Z",
          "content": "<p>Sadly, kaggle challenges don't reward best practice but only the best score. So, you got to use that information in your favor.</p>",
          "rawMarkdown": "Sadly, kaggle challenges don't reward best practice but only the best score. So, you got to use that information in your favor.",
          "votes": 2
        },
        {
          "id": 543647,
          "postDate": "2019-06-04T16:55:32.597Z",
          "content": "<p>Sometimes you have some test info to make your predictions. For instance, I have had to predict the geology in deeper levels given shallow information, and as a geologist I have some prior knowledge about what to expect deeper. I know the distribution of grades and rocks will change deeper.</p>",
          "rawMarkdown": "Sometimes you have some test info to make your predictions. For instance, I have had to predict the geology in deeper levels given shallow information, and as a geologist I have some prior knowledge about what to expect deeper. I know the distribution of grades and rocks will change deeper.",
          "votes": 1
        },
        {
          "id": 543871,
          "postDate": "2019-06-04T22:40:23.377Z",
          "content": "<blockquote>\n  <p><strong>Ilu wrote</strong></p>\n  \n  <blockquote>\n    <p>At least, here, everyone had the info about test and could work with it.</p>\n  </blockquote>\n</blockquote>\n\n<p>I only saw the paper mentioned in a separate post one day before the deadline. I am sure I am not the only one.</p>",
          "rawMarkdown": "&gt; **Ilu wrote**\n&gt; &gt; At least, here, everyone had the info about test and could work with it.\n\nI only saw the paper mentioned in a separate post one day before the deadline. I am sure I am not the only one.",
          "votes": 2
        },
        {
          "id": 544330,
          "postDate": "2019-06-05T11:55:58.333Z",
          "content": "<p><a href=\"/stocks\">@stocks</a> the paper is listed in the intro posted by the organizer 5 months ago:\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782</a></p>",
          "rawMarkdown": "@stocks the paper is listed in the intro posted by the organizer 5 months ago:\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782",
          "votes": 1
        },
        {
          "id": 544532,
          "postDate": "2019-06-05T16:11:36.663Z",
          "content": "<p>I am not sure we are talking about the same paper since I don't see p4677 mentioned in the intro.\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-543544\">p4677 leakage discussion</a></p>",
          "rawMarkdown": "I am not sure we are talking about the same paper since I don't see p4677 mentioned in the intro.\n[p4677 leakage discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-543544)"
        },
        {
          "id": 546285,
          "postDate": "2019-06-06T12:36:37.353Z",
          "content": "<p><a href=\"/stocks\">@stocks</a> <a href=\"/ilu000\">@ilu000</a> the image I am talking about, which let me figure out that the length of training experimets was the same as the ones in our training data, is the figure 1D of the second paper shared by them <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782\">here</a>\n(the link: <a href=\"https://doi.org/10.1002/2017GL076708\">https://doi.org/10.1002/2017GL076708</a>)</p>\n\n<p>This was shared as the introduction to the problem 5 months ago, and it was the only picture I saw and used about experiment p4677 ;)</p>",
          "rawMarkdown": "@stocks @ilu000 the image I am talking about, which let me figure out that the length of training experimets was the same as the ones in our training data, is the figure 1D of the second paper shared by them [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782)\n(the link: https://doi.org/10.1002/2017GL076708)\n\nThis was shared as the introduction to the problem 5 months ago, and it was the only picture I saw and used about experiment p4677 ;)",
          "votes": 1
        },
        {
          "id": 546308,
          "postDate": "2019-06-06T13:06:55.097Z",
          "content": "<p>I used a ruler on the same picture as <a href=\"/carlospk\">@carlospk</a> ;)  And I noticed it when this was discussed at length a month ago in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664\">p4677 leakage discussion</a>.  </p>",
          "rawMarkdown": "I used a ruler on the same picture as @carlospk ;)  And I noticed it when this was discussed at length a month ago in [p4677 leakage discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664).  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 543873,
      "postDate": "2019-06-04T22:42:27.353Z",
      "content": "<p>I should also confess that I benefit from the leakage. I knew that public test data is not a good representative for the private data. However, I did not know in which direction I should step in: toward longer EQ or shorter?!!  And I did not have enough time nor knowledge to match test and train data in a systematic way. The leakage helped me to step in the longer EQs so I decided to remove 2 of the shortest EQs from the training set.\nThanks to relatively good features that I extracted from the audio sound (I mostly investigated in this part, as a signal processing expert rather than an ML expert) I got a relatively good rank. As for my first competition, it is a good motivating reason to continue working on ML However, prefer not to encounter leakage nor the possibility of test-train matching in future competitions.</p>",
      "rawMarkdown": "I should also confess that I benefit from the leakage. I knew that public test data is not a good representative for the private data. However, I did not know in which direction I should step in: toward longer EQ or shorter?!!  And I did not have enough time nor knowledge to match test and train data in a systematic way. The leakage helped me to step in the longer EQs so I decided to remove 2 of the shortest EQs from the training set.\nThanks to relatively good features that I extracted from the audio sound (I mostly investigated in this part, as a signal processing expert rather than an ML expert) I got a relatively good rank. As for my first competition, it is a good motivating reason to continue working on ML However, prefer not to encounter leakage nor the possibility of test-train matching in future competitions.",
      "votes": 4
    },
    {
      "id": 543848,
      "postDate": "2019-06-04T22:13:26.413Z",
      "content": "<p>One thing is now more evident for me after reading about the best ranked methods: \nBest ranked methods are those methods that got some information from the test data and employed that information in the training procedure. This information could be obtained by the leakage, or by matching the train and test data (see  <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731\">7th</a> or <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543645\">1st</a> for instance).\nIn my opinion, private test results are not good at all. They show that none of our machine learning (ML) methods was really successful to distinguish a small fracture from a big fracture.\nI think the organizer(s) should be probably more disappointed that competitors. In my opinion they were waiting for a better ML method that could predict EQ pretty much better than that. They did not provide long duration EQ in the public data because they did not want that competitors go toward over-fitting. But it did not work!!! \n<strong>In brief: I would preferred that long and short time EQ could be distinguished directly by the ML methods, and not by finding similarities between test and train data.</strong></p>",
      "rawMarkdown": "One thing is now more evident for me after reading about the best ranked methods: \nBest ranked methods are those methods that got some information from the test data and employed that information in the training procedure. This information could be obtained by the leakage, or by matching the train and test data (see  [7th](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731) or [1st](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543645) for instance).\nIn my opinion, private test results are not good at all. They show that none of our machine learning (ML) methods was really successful to distinguish a small fracture from a big fracture.\nI think the organizer(s) should be probably more disappointed that competitors. In my opinion they were waiting for a better ML method that could predict EQ pretty much better than that. They did not provide long duration EQ in the public data because they did not want that competitors go toward over-fitting. But it did not work!!! \n**In brief: I would preferred that long and short time EQ could be distinguished directly by the ML methods, and not by finding similarities between test and train data.**",
      "votes": 4,
      "replies": [
        {
          "id": 543856,
          "postDate": "2019-06-04T22:25:32.437Z",
          "content": "<p>As I said before in another discussions, I think the way the organizers set up the problem didn't allow to use all the power of Machine Learning. The chunks are too small, and there are too few experiments, so powefull algorithms for this kind of unstructured problems like an LSTM or CNN could not work optimally.</p>",
          "rawMarkdown": "As I said before in another discussions, I think the way the organizers set up the problem didn't allow to use all the power of Machine Learning. The chunks are too small, and there are too few experiments, so powefull algorithms for this kind of unstructured problems like an LSTM or CNN could not work optimally.",
          "votes": 3
        },
        {
          "id": 552744,
          "postDate": "2019-06-14T12:58:12.897Z",
          "content": "<p>Just like a starving Tyson</p>",
          "rawMarkdown": "Just like a starving Tyson"
        }
      ]
    },
    {
      "id": 543114,
      "postDate": "2019-06-04T10:37:09.030Z",
      "content": "<p>Don't be too disappointed, in real world data can be way poorer.  You did well in this tricky competition.</p>",
      "rawMarkdown": "Don't be too disappointed, in real world data can be way poorer.  You did well in this tricky competition.",
      "votes": 4,
      "replies": [
        {
          "id": 543140,
          "postDate": "2019-06-04T10:50:19.640Z",
          "content": "<p>I am indeed very happy with my result! With only 10% of top 100 public LB staying top 100 in private LB, I am lucky to be between the 10% survivors. But it has nothing to do with the topic of the post.</p>",
          "rawMarkdown": "I am indeed very happy with my result! With only 10% of top 100 public LB staying top 100 in private LB, I am lucky to be between the 10% survivors. But it has nothing to do with the topic of the post.",
          "votes": 5
        },
        {
          "id": 543309,
          "postDate": "2019-06-04T13:11:15.197Z",
          "content": "<p>Well, true. Real world data can be, and often is, much worse. Still, that does not mean, we should not strive to improve! That's especially true for a scientific organisation (such as LANL) and a platform like Kaggle. <a href=\"/zaharch\">@zaharch</a> (alias nosound) has some valid points here.</p>",
          "rawMarkdown": "Well, true. Real world data can be, and often is, much worse. Still, that does not mean, we should not strive to improve! That's especially true for a scientific organisation (such as LANL) and a platform like Kaggle. @zaharch (alias nosound) has some valid points here.",
          "votes": 3
        }
      ]
    },
    {
      "id": 544538,
      "postDate": "2019-06-05T16:15:56.423Z",
      "content": "<p>I am wondering how difficult/expensive those earthquake simulation runs really are.  I suggest we do another  Earthquake prediction challenge with 10-50x more data (<strong>unpublished runs only</strong>).</p>",
      "rawMarkdown": "I am wondering how difficult/expensive those earthquake simulation runs really are.  I suggest we do another  Earthquake prediction challenge with 10-50x more data (**unpublished runs only**).",
      "votes": 1,
      "replies": [
        {
          "id": 544542,
          "postDate": "2019-06-05T16:22:56.810Z",
          "content": "<p>That would be great. The only problem of a chunk size so large is what <a href=\"/cpmpml\">@cpmpml</a> pointed out <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93966#latest-541020\">here</a>.\nConsidering that equivalence between time in the lab vs real world, I think the chunks should be at least something between 3 to 10 times the size they are.</p>",
          "rawMarkdown": "That would be great. The only problem of a chunk size so large is what @cpmpml pointed out [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93966#latest-541020).\nConsidering that equivalence between time in the lab vs real world, I think the chunks should be at least something between 3 to 10 times the size they are.\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 543489,
      "postDate": "2019-06-04T15:00:11.070Z",
      "content": "<p>It is shocking the whole training was just one long wave.  How typical was that wave ?  The number of earthquakes (15) in that wave is very small.    I think they should have provide at least 50 simulation waves (at shorter length or reduced resolution).  Then getting 150k fragments of test data from another 50 waves.  You can't really build a reliable model to predict time to failure with just 15 earthquakes.  It's shocking the competition was designed by professional scientists at the national lab.</p>",
      "rawMarkdown": "It is shocking the whole training was just one long wave.  How typical was that wave ?  The number of earthquakes (15) in that wave is very small.    I think they should have provide at least 50 simulation waves (at shorter length or reduced resolution).  Then getting 150k fragments of test data from another 50 waves.  You can't really build a reliable model to predict time to failure with just 15 earthquakes.  It's shocking the competition was designed by professional scientists at the national lab.",
      "votes": 1,
      "replies": [
        {
          "id": 543493,
          "postDate": "2019-06-04T15:03:02.920Z",
          "content": "<p>The experiment drifts too much. \nThat's why they only used a small portion. Don't blame them for that.\nAlso, there is a lot more data available from other labquake experiments. But to my knowledge nothing of that was helpful here. </p>",
          "rawMarkdown": "The experiment drifts too much. \nThat's why they only used a small portion. Don't blame them for that.\nAlso, there is a lot more data available from other labquake experiments. But to my knowledge nothing of that was helpful here. ",
          "votes": 1
        },
        {
          "id": 543524,
          "postDate": "2019-06-04T15:18:54.510Z",
          "content": "<p>Still, they could have provide multiple experiments from shorter non-drifted data.  I would hesitate to include outside data, as it's hard to know if these experiments were done under the same setting.  </p>",
          "rawMarkdown": "Still, they could have provide multiple experiments from shorter non-drifted data.  I would hesitate to include outside data, as it's hard to know if these experiments were done under the same setting.  "
        },
        {
          "id": 543529,
          "postDate": "2019-06-04T15:22:42.877Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77240537429\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77240537429</a></p>\n\n<p>Go ahead. There is your multiple experiment data. The setup is always different. I mean, hey, we are talking about layers of crumbeling and solidifying material. You cant reproduce that 1:1.</p>",
          "rawMarkdown": "https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77240537429\n\nGo ahead. There is your multiple experiment data. The setup is always different. I mean, hey, we are talking about layers of crumbeling and solidifying material. You cant reproduce that 1:1."
        }
      ]
    },
    {
      "id": 543159,
      "postDate": "2019-06-04T11:03:37.757Z",
      "content": "<p>For me the metric was far too sensitive to the mean considering the means of public and private were so different.  They would have been better deleting the real mean from the targets or using a percentile metric such as MPSE. Still I had never heard of MFCC prior to this competition and I think it is very cool!</p>",
      "rawMarkdown": "For me the metric was far too sensitive to the mean considering the means of public and private were so different.  They would have been better deleting the real mean from the targets or using a percentile metric such as MPSE. Still I had never heard of MFCC prior to this competition and I think it is very cool!\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 543516,
          "postDate": "2019-06-04T15:14:52.813Z",
          "content": "<p>Just wondering, for MPE how would you deal with actual values of TTF being 0? Division by zero would be a problem there.</p>",
          "rawMarkdown": "Just wondering, for MPE how would you deal with actual values of TTF being 0? Division by zero would be a problem there.",
          "votes": 1
        },
        {
          "id": 543519,
          "postDate": "2019-06-04T15:16:42.143Z",
          "content": "<p>TTF is never Zero in train and can easily be hardcoded to be e.g. min. 0.00001s</p>",
          "rawMarkdown": "TTF is never Zero in train and can easily be hardcoded to be e.g. min. 0.00001s",
          "votes": 1
        },
        {
          "id": 543636,
          "postDate": "2019-06-04T16:44:34.253Z",
          "content": "<p>Agree about MFCC, my most important features came from this in the end. I experimented with other librosa tools but nothing helped as much. Ratio between MFCC coeffs was particularly useful.</p>",
          "rawMarkdown": "Agree about MFCC, my most important features came from this in the end. I experimented with other librosa tools but nothing helped as much. Ratio between MFCC coeffs was particularly useful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 546484,
      "postDate": "2019-06-06T16:04:24.053Z",
      "content": "<p><a href=\"https://www.kaggle.com/c/instant-gratification\">Then you will appreciate this</a></p>",
      "rawMarkdown": "[Then you will appreciate this](https://www.kaggle.com/c/instant-gratification)",
      "votes": 2
    },
    {
      "id": 546368,
      "postDate": "2019-06-06T14:17:37.123Z",
      "content": "<p>Kaggle has responded - sort of!\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94638\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94638</a></p>\n\n<p>Edit - Kaggle will address these issues in a few days - cool</p>",
      "rawMarkdown": "Kaggle has responded - sort of!\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94638\n\nEdit - Kaggle will address these issues in a few days - cool\n\n",
      "votes": 2
    },
    {
      "id": 544199,
      "postDate": "2019-06-05T08:58:49.797Z",
      "content": "<p>That's all true, guys. The problem with this competition was clear at the very start of it when a significant mismatch between train and test distributions became obvious. This resulted in people fitting models to the exact given test set at the cost of their real-life generalization value.</p>",
      "rawMarkdown": "That's all true, guys. The problem with this competition was clear at the very start of it when a significant mismatch between train and test distributions became obvious. This resulted in people fitting models to the exact given test set at the cost of their real-life generalization value.",
      "votes": 2,
      "replies": [
        {
          "id": 544231,
          "postDate": "2019-06-05T09:42:45.113Z",
          "content": "<p>That's a misconception in my view. The feature distribution are not different once you substract the mean from each and every segment individually. Also the TTF is not that different as many seem to think. In the public test we had just two cycles! That leaves us with a high uncertainty in the measurement of the TTF. I've also made a short comment on that in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544228\">my memo of my final submission</a>. I think, the misconception that the test data would be so different lead many into overfitting the test data and this is why they fell so far in the privat LB.</p>",
          "rawMarkdown": "That's a misconception in my view. The feature distribution are not different once you substract the mean from each and every segment individually. Also the TTF is not that different as many seem to think. In the public test we had just two cycles! That leaves us with a high uncertainty in the measurement of the TTF. I've also made a short comment on that in [my memo of my final submission](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544228). I think, the misconception that the test data would be so different lead many into overfitting the test data and this is why they fell so far in the privat LB.",
          "votes": 2
        },
        {
          "id": 544236,
          "postDate": "2019-06-05T09:46:21.470Z",
          "content": "<blockquote>\n  <p>The feature distribution are not different once you substract the mean from each and every segment individually. </p>\n</blockquote>\n\n<p>Indeed,.  We also subtracted the mean of each segment, I forgot to say it in my writeup but my team mate did say it.</p>",
          "rawMarkdown": "&gt; The feature distribution are not different once you substract the mean from each and every segment individually. \n\nIndeed,.  We also subtracted the mean of each segment, I forgot to say it in my writeup but my team mate did say it.",
          "votes": 2
        },
        {
          "id": 544247,
          "postDate": "2019-06-05T09:56:22.510Z",
          "content": "<p>Even if you subtract the mean and even divide by the standard deviation each segment data, still many features (autocorrelation-based features, spectral analysis-based features) would be distributed quite differently. P-values of corresponding statistical tests (t-test or KS test) are extremely low. You may refer to several top solutions, including 1st and 6th place solutions: selecting train segments that match the test set distribution is a crucial part of them. </p>",
          "rawMarkdown": "Even if you subtract the mean and even divide by the standard deviation each segment data, still many features (autocorrelation-based features, spectral analysis-based features) would be distributed quite differently. P-values of corresponding statistical tests (t-test or KS test) are extremely low. You may refer to several top solutions, including 1st and 6th place solutions: selecting train segments that match the test set distribution is a crucial part of them. ",
          "votes": 1
        },
        {
          "id": 544251,
          "postDate": "2019-06-05T10:06:19.637Z",
          "content": "<p>I strongly disagree, see my other post.  I was possible to get into gold without any use of test data.  Berni did as well.</p>",
          "rawMarkdown": "I strongly disagree, see my other post.  I was possible to get into gold without any use of test data.  Berni did as well."
        },
        {
          "id": 544266,
          "postDate": "2019-06-05T10:32:39.073Z",
          "content": "<p>I congratulate you both on winning gold! But I cannot get what exactly you disagree with. In your 7th place solution description (in section \"Train / Test Difference\") you say \"But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the <strong>significant difference</strong> between train and test data.\" It's obvious that many top solutions used train data adaptation to particular test data. [One of my submissions that was based on a subset of EQs, that were selected based on train-test distributions matching (via KS and t-tests), scores in top 20 (however finally I decided to train the model on all EQs which led to me failing my first Kaggle competition :))).] </p>",
          "rawMarkdown": "I congratulate you both on winning gold! But I cannot get what exactly you disagree with. In your 7th place solution description (in section \"Train / Test Difference\") you say \"But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the **significant difference** between train and test data.\" It's obvious that many top solutions used train data adaptation to particular test data. [One of my submissions that was based on a subset of EQs, that were selected based on train-test distributions matching (via KS and t-tests), scores in top 20 (however finally I decided to train the model on all EQs which led to me failing my first Kaggle competition :))).] ",
          "votes": 2
        },
        {
          "id": 544276,
          "postDate": "2019-06-05T10:49:06.523Z",
          "content": "<p>I disagree on everything you wrote in the comment I responded to.  Once you subtract mean in each segment then my features have very similar distributions between train and test.  Second, we used a stack based on 4 base models, two of them (knn and gam) not depending at all from test data.  Third, using the gam model alone would led to a score of 2.3485 i.e. 13th rank. Therefore, using test estimate moved us from 13 to 7th rank.  And part of the improvement comes from stacking as well.  Not that crucial, is it?</p>",
          "rawMarkdown": "I disagree on everything you wrote in the comment I responded to.  Once you subtract mean in each segment then my features have very similar distributions between train and test.  Second, we used a stack based on 4 base models, two of them (knn and gam) not depending at all from test data.  Third, using the gam model alone would led to a score of 2.3485 i.e. 13th rank. Therefore, using test estimate moved us from 13 to 7th rank.  And part of the improvement comes from stacking as well.  Not that crucial, is it?",
          "votes": 1
        },
        {
          "id": 544307,
          "postDate": "2019-06-05T11:32:02.857Z",
          "content": "<p>Then I believe you didn't get my point. If the competition was designed in some other way, that would not motivate people to perform \"dirty\" ML research (LB probing, test set distribution fitting, analyzing pictures with a ruler etc.), than the organizers and the community would be better off. </p>\n\n<p>About feature distributions: when you say \"similar distributions\" what p-values of what statistical tests you imply? Anyway, you take your own experience and some particular set of your features and disagree with me saying that in several top solutions (I mentioned 1st and 6th place solutions, you may also refer to 2nd place solution [and many others I believe]) these train-test adaptation is crucial?.. That's strange at least. You cannot state that it is not crucial for other solutions based just on your own experiments. And the end of day, the authors of these solutions made a dicision to include train-test adaptation tricks into their solutions, they spent time on thinking about how to make predictions for this particular test set rather than about how to invent the best real-life model (arguably, genuine purpose of organizers). I'm not saying that it was impossible to get a decent rank without looking at the test set at all, I'm saying that many people did this \"looking\" at the potential cost of real-life generalization ability of the model. It may well turn out that on some future real-life test data many of the top models would not perform well, because they're in a way fitted to a particular test set.</p>",
          "rawMarkdown": "Then I believe you didn't get my point. If the competition was designed in some other way, that would not motivate people to perform \"dirty\" ML research (LB probing, test set distribution fitting, analyzing pictures with a ruler etc.), than the organizers and the community would be better off. \n\nAbout feature distributions: when you say \"similar distributions\" what p-values of what statistical tests you imply? Anyway, you take your own experience and some particular set of your features and disagree with me saying that in several top solutions (I mentioned 1st and 6th place solutions, you may also refer to 2nd place solution [and many others I believe]) these train-test adaptation is crucial?.. That's strange at least. You cannot state that it is not crucial for other solutions based just on your own experiments. And the end of day, the authors of these solutions made a dicision to include train-test adaptation tricks into their solutions, they spent time on thinking about how to make predictions for this particular test set rather than about how to invent the best real-life model (arguably, genuine purpose of organizers). I'm not saying that it was impossible to get a decent rank without looking at the test set at all, I'm saying that many people did this \"looking\" at the potential cost of real-life generalization ability of the model. It may well turn out that on some future real-life test data many of the top models would not perform well, because they're in a way fitted to a particular test set.\n\n",
          "votes": 2
        },
        {
          "id": 544342,
          "postDate": "2019-06-05T12:18:44.920Z",
          "content": "<blockquote>\n  <p>Then I believe you didn't get my point. </p>\n</blockquote>\n\n<p>There is a difference between that and disagreeing with you.  To answer the same way as you, I am not sure you get what I disagree with.  That's fine, let's agree to disagree.  Not worth a fight IMHO ;)</p>\n\n<p>Where I agree with you is that seeing paper pictures led people, including me, to use a ruler on zoomed in paper image.  That was fun actually, bringing us back to early days of physics ;)</p>",
          "rawMarkdown": "&gt; Then I believe you didn't get my point. \n\nThere is a difference between that and disagreeing with you.  To answer the same way as you, I am not sure you get what I disagree with.  That's fine, let's agree to disagree.  Not worth a fight IMHO ;)\n\nWhere I agree with you is that seeing paper pictures led people, including me, to use a ruler on zoomed in paper image.  That was fun actually, bringing us back to early days of physics ;)",
          "votes": 3
        },
        {
          "id": 544380,
          "postDate": "2019-06-05T12:59:21.437Z",
          "content": "<p>I agree with <a href=\"/kostyanomatterwhat\">@kostyanomatterwhat</a> that there are incentives when competing in Kaggle that don't go in line with the purpose of the organizers, so Kaggle should do a much better job to avoid any leakage.</p>\n\n<p>Finally, as <a href=\"/cpmpml\">@cpmpml</a> said, at least in this competition this was not an impediment for Kagglers to find really good features and models that can generalize well and are useful for the real purpose of the competition.</p>",
          "rawMarkdown": "I agree with @kostyanomatterwhat that there are incentives when competing in Kaggle that don't go in line with the purpose of the organizers, so Kaggle should do a much better job to avoid any leakage.\n\nFinally, as @cpmpml said, at least in this competition this was not an impediment for Kagglers to find really good features and models that can generalize well and are useful for the real purpose of the competition.",
          "votes": 4
        },
        {
          "id": 544411,
          "postDate": "2019-06-05T13:36:26.487Z",
          "content": "<p>Competing with solutions that are overfitted to private test set is pointless. (unless yours is overfitted as well)\n<a href=\"/cpmpml\">@cpmpml</a> \nNot many chose to \"risk\" and I think that is the reason that your models would have gold score. With more overfitted solutions score needed for gold would be lower.</p>",
          "rawMarkdown": "Competing with solutions that are overfitted to private test set is pointless. (unless yours is overfitted as well)\n@cpmpml \nNot many chose to \"risk\" and I think that is the reason that your models would have gold score. With more overfitted solutions score needed for gold would be lower.",
          "votes": 3
        },
        {
          "id": 544414,
          "postDate": "2019-06-05T13:41:50.707Z",
          "content": "<blockquote>\n  <p>Not many chose to \"risk\" </p>\n</blockquote>\n\n<p>That's an interesting hypothesis, but what data do you have to back it?</p>\n\n<p>Anyway, you are confirming what I say: not many solutions are above good models trained on training data without adding bias estimated from test.</p>",
          "rawMarkdown": "&gt; Not many chose to \"risk\" \n\nThat's an interesting hypothesis, but what data do you have to back it?\n\nAnyway, you are confirming what I say: not many solutions are above good models trained on training data without adding bias estimated from test."
        },
        {
          "id": 544424,
          "postDate": "2019-06-05T13:50:50.420Z",
          "content": "<p>I'm judging from forum posts and my model's (adjusted to private test mean) score that was last day effort  to avoid being in ~2-3K place.</p>",
          "rawMarkdown": "I'm judging from forum posts and my model's (adjusted to private test mean) score that was last day effort  to avoid being in ~2-3K place.",
          "votes": 1
        },
        {
          "id": 544633,
          "postDate": "2019-06-05T18:11:23.783Z",
          "content": "<p>I reported the systematic downvoting here, very interesting to see if anything happens.  I don't think any of the people who actually wrote something do it, because the downvotes happen way later ;)</p>",
          "rawMarkdown": "I reported the systematic downvoting here, very interesting to see if anything happens.  I don't think any of the people who actually wrote something do it, because the downvotes happen way later ;)",
          "votes": 1
        },
        {
          "id": 544679,
          "postDate": "2019-06-05T19:01:27.293Z",
          "content": "<p>It looks like you have a pathetic campaign going on from the keyboard warriors.  We have a report feature for bad comments so why don't we remove the down votes arrow and just have first click +1 second click back to zero on the up arrow as it does now.  AND remove voting from a users profile as people can just go to the discussion section and do block down-voting.</p>",
          "rawMarkdown": "It looks like you have a pathetic campaign going on from the keyboard warriors.  We have a report feature for bad comments so why don't we remove the down votes arrow and just have first click +1 second click back to zero on the up arrow as it does now.  AND remove voting from a users profile as people can just go to the discussion section and do block down-voting.",
          "votes": 2
        },
        {
          "id": 546216,
          "postDate": "2019-06-06T11:13:25.273Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 546384,
          "postDate": "2019-06-06T14:31:09.570Z",
          "content": "<p>Reporting made the down vote disappear. :)</p>",
          "rawMarkdown": "Reporting made the down vote disappear. :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 547375,
      "postDate": "2019-06-07T16:25:07.370Z",
      "content": "<p>Nice Kernel</p>",
      "rawMarkdown": "Nice Kernel",
      "votes": -1
    },
    {
      "id": 543834,
      "postDate": "2019-06-04T21:47:03.123Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 543194,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T11:21:03.623000",
      "content": "<p>What I find really disappointing is the lack of follow up from organizers.  it was the same in recent research competitions I entered.  I see a difference with commercial sponsor who look way more interested in what we can do to help them.</p>",
      "votes": 20,
      "replies": [
        {
          "id": 543214,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T11:33:02.097000",
          "content": "<p>I found it really odd how communication died with months to go we aren't talking about a couple of weeks for a holiday, conference etc.  The same has happened with product feedback - maybe Kaggle has a staff shortage.</p>\n\n<p>The good thing about this competition at least shows that metrics R2 and MAE may not be the best metrics.  If that helps better predictions then great.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 543495,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-06-04T15:07:29.123000",
          "content": "<p>It surprises me though, usually scientists are much more interested in what is going on.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 543523,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:18:42.357000",
          "content": "<p>My guess is, they never wanted us to find their original paper on exp 4677. As soon as this was out, the results were no longer of value for them. \nI mean, they linked 3 papers but not the one containing the experiment 4677? That is no coincidence.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 543527,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T15:21:42.760000",
          "content": "<p>The paper with exp 4677 is not from same LANL authors as the ones that were shared.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 543534,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "2019-06-04T15:23:57.460000",
          "content": "<p>They don't need to bother.  The training data only has 15 earthquakes.  It's really hard to judge any of predicting methods are good or bad. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543626,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-06-04T16:35:32.603000",
      "content": "<p>I wonder why don’t they just randomly pick the shuffled segments from the whole length of data, 1 set for train, 1 for public, and 1 for private? That will make the competition authentic. The data preparation is poor, supporting host is even poorer. Sad for a 4500 contestant competition. </p>",
      "votes": 11,
      "replies": [
        {
          "id": 543631,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T16:39:10.593000",
          "content": "<p>Come on now they have only had 5 months ;P</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 543763,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2019-06-04T19:42:09.347000",
          "content": "<p>Could not agree more. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 546072,
          "author_name": "vroom",
          "author_url": "",
          "post_date": "2019-06-06T08:21:28.120000",
          "content": "<p>Or, they can scale up our predictions and give out second batch of medals ... </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543536,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T15:24:33.140000",
      "content": "<p>I can't agree more with 1. This is a big leakage and for sure changed competitions results. </p>",
      "votes": 9,
      "replies": [
        {
          "id": 543551,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T15:34:54.570000",
          "content": "<p>This is the first competition I know of that benefited from using rulers on print-outs for sure!</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 543556,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:39:19.587000",
          "content": "<p>It is definetly a kind of leakage, but as you can see from our explanation we didn't use that information directly in our model. We optimized to oof/predictions similarity to account for the large train/test difference (for which the the data available at kaggle was already enough).</p>\n\n<p><a href=\"/scirpus\">@scirpus</a> we did use our rulers, but in the end we didn't use it for our final models</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 543562,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T15:42:07.840000",
          "content": "<p>I noticed your CV gives a mean of about 6.2 - so you did it the right way - the fact people can just scale it to 6.2 and do well from the paper is what people are sore about.  Congrats by the way.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 543567,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:44:40.080000",
          "content": "<p>I get you point, no worries. \nStill, up to now nobody published a better score with \"simple upscaling\"\nSo upscaling alone doesn't make a great model?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543573,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T15:46:38.797000",
          "content": "<p>You can get a gold medal just by picking a random model and multiplying the ttf's by ~ 1.1 to 1.2\nI may sound salty but I used this competition to finally get rid of all my Genetic Programming bugs from my code so it is celebration time for me! ;)  It has only taken 4 years!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543575,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:47:46.967000",
          "content": "<p>That was very predictable and i expected many more people to do exactly that. And then it would again come down to the rest of the model and not \"just\" the mean.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543621,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-04T16:30:07.203000",
          "content": "<p>Haha.. sometimes you don't need fancy stuff to get a good result. A simple ruler can help ;)\nI work in industry. I don't care much about the methods but about results. Many times I see people spending a lot of energy and time using very complex and fancy methods, and even when they present their results, they focus more on how they did than what they achieved. I am more practical. If I can get a good result using a linear regression I will do it. I could have certainly measured the length of test cycles with a software, but measuring it with a ruler took me 1 minute and it was a good starting point. It worked ;)</p>\n\n<p>I also used the length of the cycles in the paper as the first approach, but then I selected subsets of the training data to match better the distributions of the features between train and test.</p>\n\n<p>Even though I benefited here from the leakage (and despite it was disclosed for all competitors), I am also dissapointed that the 3 competitions I have participated on have been won by those who found a way to exploit a leakage. I really hope I can participate in a leakage-free competition in the future.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 543686,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2019-06-04T17:42:40.597000",
          "content": "<blockquote>\n  <p>I may sound salty but I used this competition to finally get rid of all my Genetic Programming bugs from my code so it is celebration time for me! ;) It has only taken 4 years!</p>\n</blockquote>\n\n<p>Congrats! But I remember your program worked well in the past competitons, e.g. Plasticc: the famous Scirpus' formula! :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 543703,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-04T17:56:46.417000",
          "content": "<p>I used MS Paint and Excel to find the mean :)</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 543705,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-04T17:58:45.670000",
          "content": "<p>Image here</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 543749,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2019-06-04T19:02:32.293000",
          "content": "<blockquote>\n  <p>Image here</p>\n</blockquote>\n\n<p>Cool, thanks! Now it's clear! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543772,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T20:02:47.367000",
          "content": "<p>i actually used a \"data from image\" extraction tool ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 543774,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T20:07:01.620000",
          "content": "<p>Ooh interesting - tell us more</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543777,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T20:10:32.443000",
          "content": "<p><a href=\"/scirpus\">@scirpus</a>  <a href=\"https://automeris.io/WebPlotDigitizer/\">https://automeris.io/WebPlotDigitizer/</a> this is one online option for it</p>\n\n<p>Can get in handy with non-linear data where you are only provided a picture but no tabulated data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 543784,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T20:21:07.857000",
          "content": "<p>Thankyou</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543787,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-04T20:23:20.080000",
          "content": "<p>Thanks for sharing <a href=\"/ilu000\">@ilu000</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543809,
      "author_name": "Ming Zhao",
      "author_url": "",
      "post_date": "2019-06-04T21:08:12.470000",
      "content": "<p>I don't think people should be too upset about the data leak from the paper. The time series raw data shown in the paper is a whole lot longer than just the train and test data. While it ended up that the test data was coming from the exact region labeled as test in the paper, that was far from a guarantee. The organizers could have easily taken the test data from either before the train data, or a little bit further down the road after the train data. That would lead to a different distribution, and probably very different leaderboard.</p>\n\n<p>I believe the teams that exploited the data shown in the paper took an educated risk, and congratulations to them that it paid off for this competition. But had the organizers picked differently, it would have been a losing strategy.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 543153,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2019-06-04T11:00:31.577000",
      "content": "<p>At least, here, everyone had the info about test and could work with it.\nIt doesnt apply to real world problems, but thats the issue with many kaggle challenges.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 543171,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-06-04T11:09:42.483000",
          "content": "<p>In case the test distribution is known to everyone I agree it becomes an interesting problem how to transfer learn from train to test. What worries me is that it was not by design. Just imagine we didn't know the test distribution, - private LB in this case would have been completely random, instead of 80% random now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543173,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-06-04T11:11:56.307000",
          "content": "<p>I agree but our solution basically does not need to use information from the paper on test distributions but we rather used the test data itself and feature dists to better match the training data.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 543187,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-04T11:18:56.407000",
          "content": "<p>I think the word \"prediction\" does not mean much if you need the test info!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 543189,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T11:19:51.947000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> is right, and apparently this solution was even superior to \"simple upscaling\" as many have done. \nPeaking into test is always a problem that we see in kaggle challenges, and even cannot be prevented in Kernel only challenges as <a href=\"/cdeotte\">@cdeotte</a> has shown lately. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 543499,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "2019-06-04T15:08:44.403000",
          "content": "<p>Looking at test set data distribution then go back to re-select training data is not good practice.  If you deploy a ML system, you don't really know what kind of data will come in.  One should build a good generic model instead of a model fitting one set of data.  They should have gotten test set from random fragments from a number of waves (simulations).  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543512,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:13:39.607000",
          "content": "<p>Sadly, kaggle challenges don't reward best practice but only the best score. So, you got to use that information in your favor.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 543647,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-04T16:55:32.597000",
          "content": "<p>Sometimes you have some test info to make your predictions. For instance, I have had to predict the geology in deeper levels given shallow information, and as a geologist I have some prior knowledge about what to expect deeper. I know the distribution of grades and rocks will change deeper.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543871,
          "author_name": "Mycroft Holmes",
          "author_url": "",
          "post_date": "2019-06-04T22:40:23.377000",
          "content": "<blockquote>\n  <p><strong>Ilu wrote</strong></p>\n  \n  <blockquote>\n    <p>At least, here, everyone had the info about test and could work with it.</p>\n  </blockquote>\n</blockquote>\n\n<p>I only saw the paper mentioned in a separate post one day before the deadline. I am sure I am not the only one.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 544330,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-05T11:55:58.333000",
          "content": "<p><a href=\"/stocks\">@stocks</a> the paper is listed in the intro posted by the organizer 5 months ago:\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 544532,
          "author_name": "Mycroft Holmes",
          "author_url": "",
          "post_date": "2019-06-05T16:11:36.663000",
          "content": "<p>I am not sure we are talking about the same paper since I don't see p4677 mentioned in the intro.\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-543544\">p4677 leakage discussion</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 546285,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-06T12:36:37.353000",
          "content": "<p><a href=\"/stocks\">@stocks</a> <a href=\"/ilu000\">@ilu000</a> the image I am talking about, which let me figure out that the length of training experimets was the same as the ones in our training data, is the figure 1D of the second paper shared by them <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77525#latest-524782\">here</a>\n(the link: <a href=\"https://doi.org/10.1002/2017GL076708\">https://doi.org/10.1002/2017GL076708</a>)</p>\n\n<p>This was shared as the introduction to the problem 5 months ago, and it was the only picture I saw and used about experiment p4677 ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 546308,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-06T13:06:55.097000",
          "content": "<p>I used a ruler on the same picture as <a href=\"/carlospk\">@carlospk</a> ;)  And I noticed it when this was discussed at length a month ago in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664\">p4677 leakage discussion</a>.  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543873,
      "author_name": "Behnam Molaee",
      "author_url": "",
      "post_date": "2019-06-04T22:42:27.353000",
      "content": "<p>I should also confess that I benefit from the leakage. I knew that public test data is not a good representative for the private data. However, I did not know in which direction I should step in: toward longer EQ or shorter?!!  And I did not have enough time nor knowledge to match test and train data in a systematic way. The leakage helped me to step in the longer EQs so I decided to remove 2 of the shortest EQs from the training set.\nThanks to relatively good features that I extracted from the audio sound (I mostly investigated in this part, as a signal processing expert rather than an ML expert) I got a relatively good rank. As for my first competition, it is a good motivating reason to continue working on ML However, prefer not to encounter leakage nor the possibility of test-train matching in future competitions.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 543848,
      "author_name": "Behnam Molaee",
      "author_url": "",
      "post_date": "2019-06-04T22:13:26.413000",
      "content": "<p>One thing is now more evident for me after reading about the best ranked methods: \nBest ranked methods are those methods that got some information from the test data and employed that information in the training procedure. This information could be obtained by the leakage, or by matching the train and test data (see  <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731\">7th</a> or <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543645\">1st</a> for instance).\nIn my opinion, private test results are not good at all. They show that none of our machine learning (ML) methods was really successful to distinguish a small fracture from a big fracture.\nI think the organizer(s) should be probably more disappointed that competitors. In my opinion they were waiting for a better ML method that could predict EQ pretty much better than that. They did not provide long duration EQ in the public data because they did not want that competitors go toward over-fitting. But it did not work!!! \n<strong>In brief: I would preferred that long and short time EQ could be distinguished directly by the ML methods, and not by finding similarities between test and train data.</strong></p>",
      "votes": 4,
      "replies": [
        {
          "id": 543856,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-04T22:25:32.437000",
          "content": "<p>As I said before in another discussions, I think the way the organizers set up the problem didn't allow to use all the power of Machine Learning. The chunks are too small, and there are too few experiments, so powefull algorithms for this kind of unstructured problems like an LSTM or CNN could not work optimally.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 552744,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-14T12:58:12.897000",
          "content": "<p>Just like a starving Tyson</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543114,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T10:37:09.030000",
      "content": "<p>Don't be too disappointed, in real world data can be way poorer.  You did well in this tricky competition.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 543140,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-06-04T10:50:19.640000",
          "content": "<p>I am indeed very happy with my result! With only 10% of top 100 public LB staying top 100 in private LB, I am lucky to be between the 10% survivors. But it has nothing to do with the topic of the post.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 543309,
          "author_name": "Berni",
          "author_url": "",
          "post_date": "2019-06-04T13:11:15.197000",
          "content": "<p>Well, true. Real world data can be, and often is, much worse. Still, that does not mean, we should not strive to improve! That's especially true for a scientific organisation (such as LANL) and a platform like Kaggle. <a href=\"/zaharch\">@zaharch</a> (alias nosound) has some valid points here.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 544538,
      "author_name": "Mycroft Holmes",
      "author_url": "",
      "post_date": "2019-06-05T16:15:56.423000",
      "content": "<p>I am wondering how difficult/expensive those earthquake simulation runs really are.  I suggest we do another  Earthquake prediction challenge with 10-50x more data (<strong>unpublished runs only</strong>).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 544542,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-05T16:22:56.810000",
          "content": "<p>That would be great. The only problem of a chunk size so large is what <a href=\"/cpmpml\">@cpmpml</a> pointed out <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/93966#latest-541020\">here</a>.\nConsidering that equivalence between time in the lab vs real world, I think the chunks should be at least something between 3 to 10 times the size they are.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 543489,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2019-06-04T15:00:11.070000",
      "content": "<p>It is shocking the whole training was just one long wave.  How typical was that wave ?  The number of earthquakes (15) in that wave is very small.    I think they should have provide at least 50 simulation waves (at shorter length or reduced resolution).  Then getting 150k fragments of test data from another 50 waves.  You can't really build a reliable model to predict time to failure with just 15 earthquakes.  It's shocking the competition was designed by professional scientists at the national lab.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 543493,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:03:02.920000",
          "content": "<p>The experiment drifts too much. \nThat's why they only used a small portion. Don't blame them for that.\nAlso, there is a lot more data available from other labquake experiments. But to my knowledge nothing of that was helpful here. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543524,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "2019-06-04T15:18:54.510000",
          "content": "<p>Still, they could have provide multiple experiments from shorter non-drifted data.  I would hesitate to include outside data, as it's hard to know if these experiments were done under the same setting.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543529,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:22:42.877000",
          "content": "<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77240537429\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77240537429</a></p>\n\n<p>Go ahead. There is your multiple experiment data. The setup is always different. I mean, hey, we are talking about layers of crumbeling and solidifying material. You cant reproduce that 1:1.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543159,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-06-04T11:03:37.757000",
      "content": "<p>For me the metric was far too sensitive to the mean considering the means of public and private were so different.  They would have been better deleting the real mean from the targets or using a percentile metric such as MPSE. Still I had never heard of MFCC prior to this competition and I think it is very cool!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 543516,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2019-06-04T15:14:52.813000",
          "content": "<p>Just wondering, for MPE how would you deal with actual values of TTF being 0? Division by zero would be a problem there.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543519,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T15:16:42.143000",
          "content": "<p>TTF is never Zero in train and can easily be hardcoded to be e.g. min. 0.00001s</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543636,
          "author_name": "RNA",
          "author_url": "",
          "post_date": "2019-06-04T16:44:34.253000",
          "content": "<p>Agree about MFCC, my most important features came from this in the end. I experimented with other librosa tools but nothing helped as much. Ratio between MFCC coeffs was particularly useful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 546484,
      "author_name": "Noah Weber",
      "author_url": "",
      "post_date": "2019-06-06T16:04:24.053000",
      "content": "<p><a href=\"https://www.kaggle.com/c/instant-gratification\">Then you will appreciate this</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 546368,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2019-06-06T14:17:37.123000",
      "content": "<p>Kaggle has responded - sort of!\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94638\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94638</a></p>\n\n<p>Edit - Kaggle will address these issues in a few days - cool</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 544199,
      "author_name": "KostyaNoMatterWhat",
      "author_url": "",
      "post_date": "2019-06-05T08:58:49.797000",
      "content": "<p>That's all true, guys. The problem with this competition was clear at the very start of it when a significant mismatch between train and test distributions became obvious. This resulted in people fitting models to the exact given test set at the cost of their real-life generalization value.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 544231,
          "author_name": "Berni",
          "author_url": "",
          "post_date": "2019-06-05T09:42:45.113000",
          "content": "<p>That's a misconception in my view. The feature distribution are not different once you substract the mean from each and every segment individually. Also the TTF is not that different as many seem to think. In the public test we had just two cycles! That leaves us with a high uncertainty in the measurement of the TTF. I've also made a short comment on that in <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94556#latest-544228\">my memo of my final submission</a>. I think, the misconception that the test data would be so different lead many into overfitting the test data and this is why they fell so far in the privat LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 544236,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-05T09:46:21.470000",
          "content": "<blockquote>\n  <p>The feature distribution are not different once you substract the mean from each and every segment individually. </p>\n</blockquote>\n\n<p>Indeed,.  We also subtracted the mean of each segment, I forgot to say it in my writeup but my team mate did say it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 544247,
          "author_name": "KostyaNoMatterWhat",
          "author_url": "",
          "post_date": "2019-06-05T09:56:22.510000",
          "content": "<p>Even if you subtract the mean and even divide by the standard deviation each segment data, still many features (autocorrelation-based features, spectral analysis-based features) would be distributed quite differently. P-values of corresponding statistical tests (t-test or KS test) are extremely low. You may refer to several top solutions, including 1st and 6th place solutions: selecting train segments that match the test set distribution is a crucial part of them. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 544251,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-05T10:06:19.637000",
          "content": "<p>I strongly disagree, see my other post.  I was possible to get into gold without any use of test data.  Berni did as well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 544266,
          "author_name": "KostyaNoMatterWhat",
          "author_url": "",
          "post_date": "2019-06-05T10:32:39.073000",
          "content": "<p>I congratulate you both on winning gold! But I cannot get what exactly you disagree with. In your 7th place solution description (in section \"Train / Test Difference\") you say \"But here, using pictures from academic papers could lead to a good estimate of test data, and it was the way to go given the <strong>significant difference</strong> between train and test data.\" It's obvious that many top solutions used train data adaptation to particular test data. [One of my submissions that was based on a subset of EQs, that were selected based on train-test distributions matching (via KS and t-tests), scores in top 20 (however finally I decided to train the model on all EQs which led to me failing my first Kaggle competition :))).] </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 544276,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-05T10:49:06.523000",
          "content": "<p>I disagree on everything you wrote in the comment I responded to.  Once you subtract mean in each segment then my features have very similar distributions between train and test.  Second, we used a stack based on 4 base models, two of them (knn and gam) not depending at all from test data.  Third, using the gam model alone would led to a score of 2.3485 i.e. 13th rank. Therefore, using test estimate moved us from 13 to 7th rank.  And part of the improvement comes from stacking as well.  Not that crucial, is it?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 544307,
          "author_name": "KostyaNoMatterWhat",
          "author_url": "",
          "post_date": "2019-06-05T11:32:02.857000",
          "content": "<p>Then I believe you didn't get my point. If the competition was designed in some other way, that would not motivate people to perform \"dirty\" ML research (LB probing, test set distribution fitting, analyzing pictures with a ruler etc.), than the organizers and the community would be better off. </p>\n\n<p>About feature distributions: when you say \"similar distributions\" what p-values of what statistical tests you imply? Anyway, you take your own experience and some particular set of your features and disagree with me saying that in several top solutions (I mentioned 1st and 6th place solutions, you may also refer to 2nd place solution [and many others I believe]) these train-test adaptation is crucial?.. That's strange at least. You cannot state that it is not crucial for other solutions based just on your own experiments. And the end of day, the authors of these solutions made a dicision to include train-test adaptation tricks into their solutions, they spent time on thinking about how to make predictions for this particular test set rather than about how to invent the best real-life model (arguably, genuine purpose of organizers). I'm not saying that it was impossible to get a decent rank without looking at the test set at all, I'm saying that many people did this \"looking\" at the potential cost of real-life generalization ability of the model. It may well turn out that on some future real-life test data many of the top models would not perform well, because they're in a way fitted to a particular test set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 544342,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-05T12:18:44.920000",
          "content": "<blockquote>\n  <p>Then I believe you didn't get my point. </p>\n</blockquote>\n\n<p>There is a difference between that and disagreeing with you.  To answer the same way as you, I am not sure you get what I disagree with.  That's fine, let's agree to disagree.  Not worth a fight IMHO ;)</p>\n\n<p>Where I agree with you is that seeing paper pictures led people, including me, to use a ruler on zoomed in paper image.  That was fun actually, bringing us back to early days of physics ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 544380,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-06-05T12:59:21.437000",
          "content": "<p>I agree with <a href=\"/kostyanomatterwhat\">@kostyanomatterwhat</a> that there are incentives when competing in Kaggle that don't go in line with the purpose of the organizers, so Kaggle should do a much better job to avoid any leakage.</p>\n\n<p>Finally, as <a href=\"/cpmpml\">@cpmpml</a> said, at least in this competition this was not an impediment for Kagglers to find really good features and models that can generalize well and are useful for the real purpose of the competition.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 544411,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "2019-06-05T13:36:26.487000",
          "content": "<p>Competing with solutions that are overfitted to private test set is pointless. (unless yours is overfitted as well)\n<a href=\"/cpmpml\">@cpmpml</a> \nNot many chose to \"risk\" and I think that is the reason that your models would have gold score. With more overfitted solutions score needed for gold would be lower.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 544414,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-05T13:41:50.707000",
          "content": "<blockquote>\n  <p>Not many chose to \"risk\" </p>\n</blockquote>\n\n<p>That's an interesting hypothesis, but what data do you have to back it?</p>\n\n<p>Anyway, you are confirming what I say: not many solutions are above good models trained on training data without adding bias estimated from test.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 544424,
          "author_name": "mykper",
          "author_url": "",
          "post_date": "2019-06-05T13:50:50.420000",
          "content": "<p>I'm judging from forum posts and my model's (adjusted to private test mean) score that was last day effort  to avoid being in ~2-3K place.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 544633,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-05T18:11:23.783000",
          "content": "<p>I reported the systematic downvoting here, very interesting to see if anything happens.  I don't think any of the people who actually wrote something do it, because the downvotes happen way later ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 544679,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2019-06-05T19:01:27.293000",
          "content": "<p>It looks like you have a pathetic campaign going on from the keyboard warriors.  We have a report feature for bad comments so why don't we remove the down votes arrow and just have first click +1 second click back to zero on the up arrow as it does now.  AND remove voting from a users profile as people can just go to the discussion section and do block down-voting.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 546216,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-06T11:13:25.273000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 546384,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-06T14:31:09.570000",
          "content": "<p>Reporting made the down vote disappear. :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 547375,
      "author_name": "Adelmo",
      "author_url": "",
      "post_date": "2019-06-07T16:25:07.370000",
      "content": "<p>Nice Kernel</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 543834,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T21:47:03.123000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543100": "This competition was poorly organized. \n\n1. The major blunder is that the test part durations can be found online, as described in [here](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90664#latest-535844). It is exacerbated by the fact that the organizers neither reported it in the official information nor confirmed it later on. I don't think Kaggle wants its competitions to be about noticing or not that some article online contains a picture of the test set distribution. \n2. If the test set distribution was not intended to be known (which I assume is true), then I don't think such huge difference between the train and the test is justified and can be predicted, - bad competition data design. A simple multiplication by a factor of random model predictions behaves better than anything that can be learnt, as illustrated by this [solution](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94324#latest-542970).\n3. The organizers stopped responding to legitimate participants questions in the official thread about 3-4 months before the competition end. It feels to me like \"well, we messed up here. Let's hide our head into the sand!\".\n4. Further displaying the organizers sloppiness are small facts that in [additional info](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/77526#latest-541180) post 12 millisecond was edited to 12 microseconds at some point silently, leaving people perplexed, and the[ official benchmark kernel](https://www.kaggle.com/inversion/basic-feature-benchmark) still uses R2 metric apparently used during the competition preparation.\n5. Public test size of 348 is so small, that it is possible to discover public vs private segments in less than 90 submissions (will share about it later), giving further advantage to people who do it.\n\nI learnt a lot during the competition, huge thanks to all the forum discussion participants, and people sharing their work in kernels! Really priceless. And I hope that like me and you, Kaggle as an organization will learn from it as well.\n\nUPDATE: after reading the write-ups of the winning solutions and other forum discussions I backtrack and admit that point 2 is not a valid point. It was possible and optimal to score high without using the test set distribution, [reference to 1st place solution](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390). But only few did that correctly, or even adjusted to test distribution at all, and that is why lucky \"random\" solutions advanced so much. Kudos to few of us who did it correctly, it is inspiring.",
    "543194": "What I find really disappointing is the lack of follow up from organizers.  it was the same in recent research competitions I entered.  I see a difference with commercial sponsor who look way more interested in what we can do to help them.",
    "543626": "I wonder why don’t they just randomly pick the shuffled segments from the whole length of data, 1 set for train, 1 for public, and 1 for private? That will make the competition authentic. The data preparation is poor, supporting host is even poorer. Sad for a 4500 contestant competition. ",
    "543536": "I can't agree more with 1. This is a big leakage and for sure changed competitions results. ",
    "543809": "I don't think people should be too upset about the data leak from the paper. The time series raw data shown in the paper is a whole lot longer than just the train and test data. While it ended up that the test data was coming from the exact region labeled as test in the paper, that was far from a guarantee. The organizers could have easily taken the test data from either before the train data, or a little bit further down the road after the train data. That would lead to a different distribution, and probably very different leaderboard.\n\nI believe the teams that exploited the data shown in the paper took an educated risk, and congratulations to them that it paid off for this competition. But had the organizers picked differently, it would have been a losing strategy.",
    "543153": "At least, here, everyone had the info about test and could work with it.\nIt doesnt apply to real world problems, but thats the issue with many kaggle challenges.",
    "543873": "I should also confess that I benefit from the leakage. I knew that public test data is not a good representative for the private data. However, I did not know in which direction I should step in: toward longer EQ or shorter?!!  And I did not have enough time nor knowledge to match test and train data in a systematic way. The leakage helped me to step in the longer EQs so I decided to remove 2 of the shortest EQs from the training set.\nThanks to relatively good features that I extracted from the audio sound (I mostly investigated in this part, as a signal processing expert rather than an ML expert) I got a relatively good rank. As for my first competition, it is a good motivating reason to continue working on ML However, prefer not to encounter leakage nor the possibility of test-train matching in future competitions.",
    "543848": "One thing is now more evident for me after reading about the best ranked methods: \nBest ranked methods are those methods that got some information from the test data and employed that information in the training procedure. This information could be obtained by the leakage, or by matching the train and test data (see  [7th](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94407#latest-543731) or [1st](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390#latest-543645) for instance).\nIn my opinion, private test results are not good at all. They show that none of our machine learning (ML) methods was really successful to distinguish a small fracture from a big fracture.\nI think the organizer(s) should be probably more disappointed that competitors. In my opinion they were waiting for a better ML method that could predict EQ pretty much better than that. They did not provide long duration EQ in the public data because they did not want that competitors go toward over-fitting. But it did not work!!! \n**In brief: I would preferred that long and short time EQ could be distinguished directly by the ML methods, and not by finding similarities between test and train data.**",
    "543114": "Don't be too disappointed, in real world data can be way poorer.  You did well in this tricky competition.",
    "544538": "I am wondering how difficult/expensive those earthquake simulation runs really are.  I suggest we do another  Earthquake prediction challenge with 10-50x more data (**unpublished runs only**).",
    "543489": "It is shocking the whole training was just one long wave.  How typical was that wave ?  The number of earthquakes (15) in that wave is very small.    I think they should have provide at least 50 simulation waves (at shorter length or reduced resolution).  Then getting 150k fragments of test data from another 50 waves.  You can't really build a reliable model to predict time to failure with just 15 earthquakes.  It's shocking the competition was designed by professional scientists at the national lab.",
    "543159": "For me the metric was far too sensitive to the mean considering the means of public and private were so different.  They would have been better deleting the real mean from the targets or using a percentile metric such as MPSE. Still I had never heard of MFCC prior to this competition and I think it is very cool!\n\n",
    "546484": "[Then you will appreciate this](https://www.kaggle.com/c/instant-gratification)",
    "546368": "Kaggle has responded - sort of!\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94638\n\nEdit - Kaggle will address these issues in a few days - cool\n\n",
    "544199": "That's all true, guys. The problem with this competition was clear at the very start of it when a significant mismatch between train and test distributions became obvious. This resulted in people fitting models to the exact given test set at the cost of their real-life generalization value.",
    "547375": "Nice Kernel",
    "543834": ""
  }
}