{
  "id": 94390,
  "title": "1st place solution",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94390",
  "author_name": "Psi",
  "post_date": "2019-06-04T08:50:58.328000",
  "votes": 213,
  "comment_count": 113,
  "views": 0,
  "content": "<p>Thanks a lot to the hosts of this competition and congratz to all participants and of course to my amazing teammates.</p>\n\n<p>What made this competition tricky was to find a proper CV setup that you believe in as the public LB gave bad feedback for private LB. This was my first competition where this was the case and it took me a while to completely ignore public LB, but it was necessary.</p>\n\n<p>I will now try to summarize some of the main points that helped us to win this competition. I am posting these elaborations in the we-form as we are a team and everyone contributed ideas and knowledge. Special thanks to <a href=\"/ilu000\">@ilu000</a> <a href=\"/dott1718\">@dott1718</a> <a href=\"/returnofsputnik\">@returnofsputnik</a> <a href=\"/dkaraflos\">@dkaraflos</a> <a href=\"/pukkinming\">@pukkinming</a> who worked hard the last few weeks on the comp.</p>\n\n<p><strong>Acoustic signal manipulation and features</strong></p>\n\n<p>As has been discussed in the forums and shown by adversarial validation, the signal had a certain time-trend that caused some issues specifically on mean and quantile based features. To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating <code>np.random.normal(0, 0.5, 150_000)</code>. Additionally, after noise addition, we subtracted the median of the segment. </p>\n\n<p>Our features are then calculated on this manipulated signal. We mostly focused on similar features as most participants in this competition, namely finding peaks and volatility of the signal. One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. We sometimes used a few more features (like for the NN, see below) but they are usually very similar. Those 4 are decently uncorrelated between themselves, and add good diversity. For each feature we always only considered it if it has a p-value &gt;0.05 on a KS statistic of train vs test.</p>\n\n<p><strong>Differences between train and test features</strong></p>\n\n<p>After doing abovementioned signal manipulation, we had more trust in our calculated features and could focus on better studying differences between train and test data feature distributions. We found that the test data should look different to training data in a few ways when comparing features by e.g., applying KS statistics between train and test. That’s when we decided to sample the train data to make it look more like we expect test data to look like (only from looking at feature distributions). We started by manually upsampling certain areas of train data, but gave up on that after a few tries and then we found a very nice way of aligning train and test data.</p>\n\n<p>So what we did is that we calculated a handful of features for train and test and tried to find a good subset of full earth-quakes in train, so that the overall feature distributions are similar to those of the full test data. We did this by sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test. A visualization for this looks like this (this is a limited visualization and not necessarily the one we chose to make our final selection of EQs):</p>\n\n<p><img src=\"https://i.imgur.com/9evXuhV.png\" alt=\"KS statistic train subsample vs. test\"></p>\n\n<p>The x-axis is the average target of the selected EQs in train and the y-axis is the KS statistic on a bunch of features comparing the distribution of that feature for the selected EQs vs the full test data. We can see that the best average KS-statistic is somewhere in the range of 6.2-6.5. You can also see nicely here that a problematic feature like the green one deviates clearly from the rest, this would be a feature we would not select in the end.</p>\n\n<p>After careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. The mean of this sample is 6.258 and the median is 6.031.</p>\n\n<p><strong>CV</strong></p>\n\n<p>Now that we had sampled train data that we though to be similar to test just purely based on statistical analysis, and now that we had features that should not have any time leaks, we decided on doing a simple shuffled 3-fold on that data. Higher fold results are similar. We now tried to improve this CV as well as possible. </p>\n\n<p><strong>Models</strong></p>\n\n<p>Our final submit is a hillclimber blend of three types of models: (i) LGB, (ii) SVR, (iii) NN. The overall CV score on this was ~1.83. The LGB is using a fair loss with relatively moderate other hyperparameters. The SVR is also quite simply set-up. The NN is a bit more complicated with a few layers on top of a bunch of features. The real interesting thing here is that we do multi-task learning by specifying additional losses next to the ttf loss that we weight higher than the others. We have one additional binary logloss with the target specifying if the ttf is &lt;0.5 and one further MAE loss on the target of time-since-failure. This helped to balance some of the predictions out a bit and specifically helped to better predict some of the areas at the end of earthquakes that make some weird spikes. The NN had the best single MAE, but blending improved. Actually, just blending LGB and NN would have produced the best private LB score (2.25909). Adding SVR did improve CV though.</p>\n\n<p>With all the steps described above, we also managed to make the distribution of test predictions very similar tho oof predictions. The following image shows for a single LGB the oof (blue) vs. test predictions (orange). The KS-test between those two does not reject the null hypothesis of them being equally distributed.</p>\n\n<p><img src=\"https://i.imgur.com/bspTxd4.png\" alt=\"LGB oof (blue) vs. test (orange) prediction dist\"></p>\n\n<p><strong>Ideas that have potential</strong></p>\n\n<p>We had quite some ideas that have potential but did not make it into our final submission. One area is to better use the time-since-failure prediction, which we used only as an additional loss in our NN. Modeling tsf works better than ttf. It can help to manually adjust a few predictions which have large discrepancies between tsf and ttf predictions, like the end of EQs. Also, they can be a reasonable proxy for predicting the approximate length of the EQ. So for example, we had one model that normalized the ttf targets to be in range 0-1 and then predicts this normalized target and scales it by ttf+tsf prediction. This was usually very close to our simpler models so we did not tune it extensively, I just feel that this has further potential. </p>\n\n<p>The following kernel runs a LGB model on most of what I explained above and also would score 1st place with 2.279 private MAE:</p>\n\n<p><a href=\"https://www.kaggle.com/ilu000/1-private-lb-kernel-lanl-lgbm/\">https://www.kaggle.com/ilu000/1-private-lb-kernel-lanl-lgbm/</a></p>\n\n<p>The following kernel runs a blend between LGB and NN scoring 2.25993 on private LB:</p>\n\n<p><a href=\"https://www.kaggle.com/dkaraflos/1-geomean-nn-and-6featlgbm-2-259-private-lb\">https://www.kaggle.com/dkaraflos/1-geomean-nn-and-6featlgbm-2-259-private-lb</a></p>",
  "messages": [
    {
      "id": 542982,
      "postDate": "2019-06-04T08:50:58.327Z",
      "content": "<p>Thanks a lot to the hosts of this competition and congratz to all participants and of course to my amazing teammates.</p>\n\n<p>What made this competition tricky was to find a proper CV setup that you believe in as the public LB gave bad feedback for private LB. This was my first competition where this was the case and it took me a while to completely ignore public LB, but it was necessary.</p>\n\n<p>I will now try to summarize some of the main points that helped us to win this competition. I am posting these elaborations in the we-form as we are a team and everyone contributed ideas and knowledge. Special thanks to <a href=\"/ilu000\">@ilu000</a> <a href=\"/dott1718\">@dott1718</a> <a href=\"/returnofsputnik\">@returnofsputnik</a> <a href=\"/dkaraflos\">@dkaraflos</a> <a href=\"/pukkinming\">@pukkinming</a> who worked hard the last few weeks on the comp.</p>\n\n<p><strong>Acoustic signal manipulation and features</strong></p>\n\n<p>As has been discussed in the forums and shown by adversarial validation, the signal had a certain time-trend that caused some issues specifically on mean and quantile based features. To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating <code>np.random.normal(0, 0.5, 150_000)</code>. Additionally, after noise addition, we subtracted the median of the segment. </p>\n\n<p>Our features are then calculated on this manipulated signal. We mostly focused on similar features as most participants in this competition, namely finding peaks and volatility of the signal. One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. We sometimes used a few more features (like for the NN, see below) but they are usually very similar. Those 4 are decently uncorrelated between themselves, and add good diversity. For each feature we always only considered it if it has a p-value &gt;0.05 on a KS statistic of train vs test.</p>\n\n<p><strong>Differences between train and test features</strong></p>\n\n<p>After doing abovementioned signal manipulation, we had more trust in our calculated features and could focus on better studying differences between train and test data feature distributions. We found that the test data should look different to training data in a few ways when comparing features by e.g., applying KS statistics between train and test. That’s when we decided to sample the train data to make it look more like we expect test data to look like (only from looking at feature distributions). We started by manually upsampling certain areas of train data, but gave up on that after a few tries and then we found a very nice way of aligning train and test data.</p>\n\n<p>So what we did is that we calculated a handful of features for train and test and tried to find a good subset of full earth-quakes in train, so that the overall feature distributions are similar to those of the full test data. We did this by sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test. A visualization for this looks like this (this is a limited visualization and not necessarily the one we chose to make our final selection of EQs):</p>\n\n<p><img src=\"https://i.imgur.com/9evXuhV.png\" alt=\"KS statistic train subsample vs. test\"></p>\n\n<p>The x-axis is the average target of the selected EQs in train and the y-axis is the KS statistic on a bunch of features comparing the distribution of that feature for the selected EQs vs the full test data. We can see that the best average KS-statistic is somewhere in the range of 6.2-6.5. You can also see nicely here that a problematic feature like the green one deviates clearly from the rest, this would be a feature we would not select in the end.</p>\n\n<p>After careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. The mean of this sample is 6.258 and the median is 6.031.</p>\n\n<p><strong>CV</strong></p>\n\n<p>Now that we had sampled train data that we though to be similar to test just purely based on statistical analysis, and now that we had features that should not have any time leaks, we decided on doing a simple shuffled 3-fold on that data. Higher fold results are similar. We now tried to improve this CV as well as possible. </p>\n\n<p><strong>Models</strong></p>\n\n<p>Our final submit is a hillclimber blend of three types of models: (i) LGB, (ii) SVR, (iii) NN. The overall CV score on this was ~1.83. The LGB is using a fair loss with relatively moderate other hyperparameters. The SVR is also quite simply set-up. The NN is a bit more complicated with a few layers on top of a bunch of features. The real interesting thing here is that we do multi-task learning by specifying additional losses next to the ttf loss that we weight higher than the others. We have one additional binary logloss with the target specifying if the ttf is &lt;0.5 and one further MAE loss on the target of time-since-failure. This helped to balance some of the predictions out a bit and specifically helped to better predict some of the areas at the end of earthquakes that make some weird spikes. The NN had the best single MAE, but blending improved. Actually, just blending LGB and NN would have produced the best private LB score (2.25909). Adding SVR did improve CV though.</p>\n\n<p>With all the steps described above, we also managed to make the distribution of test predictions very similar tho oof predictions. The following image shows for a single LGB the oof (blue) vs. test predictions (orange). The KS-test between those two does not reject the null hypothesis of them being equally distributed.</p>\n\n<p><img src=\"https://i.imgur.com/bspTxd4.png\" alt=\"LGB oof (blue) vs. test (orange) prediction dist\"></p>\n\n<p><strong>Ideas that have potential</strong></p>\n\n<p>We had quite some ideas that have potential but did not make it into our final submission. One area is to better use the time-since-failure prediction, which we used only as an additional loss in our NN. Modeling tsf works better than ttf. It can help to manually adjust a few predictions which have large discrepancies between tsf and ttf predictions, like the end of EQs. Also, they can be a reasonable proxy for predicting the approximate length of the EQ. So for example, we had one model that normalized the ttf targets to be in range 0-1 and then predicts this normalized target and scales it by ttf+tsf prediction. This was usually very close to our simpler models so we did not tune it extensively, I just feel that this has further potential. </p>\n\n<p>The following kernel runs a LGB model on most of what I explained above and also would score 1st place with 2.279 private MAE:</p>\n\n<p><a href=\"https://www.kaggle.com/ilu000/1-private-lb-kernel-lanl-lgbm/\">https://www.kaggle.com/ilu000/1-private-lb-kernel-lanl-lgbm/</a></p>\n\n<p>The following kernel runs a blend between LGB and NN scoring 2.25993 on private LB:</p>\n\n<p><a href=\"https://www.kaggle.com/dkaraflos/1-geomean-nn-and-6featlgbm-2-259-private-lb\">https://www.kaggle.com/dkaraflos/1-geomean-nn-and-6featlgbm-2-259-private-lb</a></p>",
      "rawMarkdown": "Thanks a lot to the hosts of this competition and congratz to all participants and of course to my amazing teammates.\n\nWhat made this competition tricky was to find a proper CV setup that you believe in as the public LB gave bad feedback for private LB. This was my first competition where this was the case and it took me a while to completely ignore public LB, but it was necessary.\n\nI will now try to summarize some of the main points that helped us to win this competition. I am posting these elaborations in the we-form as we are a team and everyone contributed ideas and knowledge. Special thanks to @ilu000 @dott1718 @returnofsputnik @dkaraflos @pukkinming who worked hard the last few weeks on the comp.\n\n**Acoustic signal manipulation and features**\n\nAs has been discussed in the forums and shown by adversarial validation, the signal had a certain time-trend that caused some issues specifically on mean and quantile based features. To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating ```np.random.normal(0, 0.5, 150_000)```. Additionally, after noise addition, we subtracted the median of the segment. \n\nOur features are then calculated on this manipulated signal. We mostly focused on similar features as most participants in this competition, namely finding peaks and volatility of the signal. One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. We sometimes used a few more features (like for the NN, see below) but they are usually very similar. Those 4 are decently uncorrelated between themselves, and add good diversity. For each feature we always only considered it if it has a p-value &gt;0.05 on a KS statistic of train vs test.\n\n**Differences between train and test features**\n\nAfter doing abovementioned signal manipulation, we had more trust in our calculated features and could focus on better studying differences between train and test data feature distributions. We found that the test data should look different to training data in a few ways when comparing features by e.g., applying KS statistics between train and test. That’s when we decided to sample the train data to make it look more like we expect test data to look like (only from looking at feature distributions). We started by manually upsampling certain areas of train data, but gave up on that after a few tries and then we found a very nice way of aligning train and test data.\n\nSo what we did is that we calculated a handful of features for train and test and tried to find a good subset of full earth-quakes in train, so that the overall feature distributions are similar to those of the full test data. We did this by sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test. A visualization for this looks like this (this is a limited visualization and not necessarily the one we chose to make our final selection of EQs):\n\n ![KS statistic train subsample vs. test](https://i.imgur.com/9evXuhV.png)\n\nThe x-axis is the average target of the selected EQs in train and the y-axis is the KS statistic on a bunch of features comparing the distribution of that feature for the selected EQs vs the full test data. We can see that the best average KS-statistic is somewhere in the range of 6.2-6.5. You can also see nicely here that a problematic feature like the green one deviates clearly from the rest, this would be a feature we would not select in the end.\n\nAfter careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. The mean of this sample is 6.258 and the median is 6.031.\n\n**CV**\n\nNow that we had sampled train data that we though to be similar to test just purely based on statistical analysis, and now that we had features that should not have any time leaks, we decided on doing a simple shuffled 3-fold on that data. Higher fold results are similar. We now tried to improve this CV as well as possible. \n\n**Models**\n\nOur final submit is a hillclimber blend of three types of models: (i) LGB, (ii) SVR, (iii) NN. The overall CV score on this was ~1.83. The LGB is using a fair loss with relatively moderate other hyperparameters. The SVR is also quite simply set-up. The NN is a bit more complicated with a few layers on top of a bunch of features. The real interesting thing here is that we do multi-task learning by specifying additional losses next to the ttf loss that we weight higher than the others. We have one additional binary logloss with the target specifying if the ttf is &lt;0.5 and one further MAE loss on the target of time-since-failure. This helped to balance some of the predictions out a bit and specifically helped to better predict some of the areas at the end of earthquakes that make some weird spikes. The NN had the best single MAE, but blending improved. Actually, just blending LGB and NN would have produced the best private LB score (2.25909). Adding SVR did improve CV though.\n\nWith all the steps described above, we also managed to make the distribution of test predictions very similar tho oof predictions. The following image shows for a single LGB the oof (blue) vs. test predictions (orange). The KS-test between those two does not reject the null hypothesis of them being equally distributed.\n\n![LGB oof (blue) vs. test (orange) prediction dist](https://i.imgur.com/bspTxd4.png)\n\n**Ideas that have potential**\n\nWe had quite some ideas that have potential but did not make it into our final submission. One area is to better use the time-since-failure prediction, which we used only as an additional loss in our NN. Modeling tsf works better than ttf. It can help to manually adjust a few predictions which have large discrepancies between tsf and ttf predictions, like the end of EQs. Also, they can be a reasonable proxy for predicting the approximate length of the EQ. So for example, we had one model that normalized the ttf targets to be in range 0-1 and then predicts this normalized target and scales it by ttf+tsf prediction. This was usually very close to our simpler models so we did not tune it extensively, I just feel that this has further potential. \n\nThe following kernel runs a LGB model on most of what I explained above and also would score 1st place with 2.279 private MAE:\n\nhttps://www.kaggle.com/ilu000/1-private-lb-kernel-lanl-lgbm/\n\nThe following kernel runs a blend between LGB and NN scoring 2.25993 on private LB:\n\nhttps://www.kaggle.com/dkaraflos/1-geomean-nn-and-6featlgbm-2-259-private-lb\n\n\n\n\n",
      "votes": 212
    },
    {
      "id": 543397,
      "postDate": "2019-06-04T14:08:51.280Z",
      "content": "<p>In a properly set scientific experiment the data would be <strong>separated</strong> into training and testing. In other words the testing data is locked in a box, the scientist goes to the box only after he/she has developed a hypothesis (a method/model etc). The testing data is then used only once to <strong>test</strong> this particular hypothesis. If the researcher is allowed to peek, analyze the testing data then the test loses its validity. Once the researcher starts tuning a hypothesis with respect to a data set then he/she needs another data set to test the reached conclusions.</p>\n\n<p>TLDR: The top ranking method has been tuned to the testing dataset via the KS test. This means that the method has become specific to this particular testing set. If the LANL researchers' goal is to apply the obtained findings to other experiments then it is very likely that they will be better of using a lower ranking method.</p>\n\n<p>The competition/ranking has lost its practical value because the organizers made the testing data public. They should have kept the testing data on the server side and allow only for submission of executable code (see Matlab Cody competitions for instance), preventing the testing data being used for tuning.</p>",
      "rawMarkdown": "In a properly set scientific experiment the data would be **separated** into training and testing. In other words the testing data is locked in a box, the scientist goes to the box only after he/she has developed a hypothesis (a method/model etc). The testing data is then used only once to **test** this particular hypothesis. If the researcher is allowed to peek, analyze the testing data then the test loses its validity. Once the researcher starts tuning a hypothesis with respect to a data set then he/she needs another data set to test the reached conclusions.\n\nTLDR: The top ranking method has been tuned to the testing dataset via the KS test. This means that the method has become specific to this particular testing set. If the LANL researchers' goal is to apply the obtained findings to other experiments then it is very likely that they will be better of using a lower ranking method.\n\nThe competition/ranking has lost its practical value because the organizers made the testing data public. They should have kept the testing data on the server side and allow only for submission of executable code (see Matlab Cody competitions for instance), preventing the testing data being used for tuning.",
      "votes": 11,
      "replies": [
        {
          "id": 543403,
          "postDate": "2019-06-04T14:13:32.633Z",
          "content": "<p>Absolutely right. \nBut please look at the scientific paper with exp. 4677. The authors merged two quakes to one there (train set). One could say this was done on purpose to increase the mean TTF of train and to make train and test more similar.... </p>\n\n<p>As long as there are humans doing the research, you will always find leaks. </p>",
          "rawMarkdown": "Absolutely right. \nBut please look at the scientific paper with exp. 4677. The authors merged two quakes to one there (train set). One could say this was done on purpose to increase the mean TTF of train and to make train and test more similar.... \n\nAs long as there are humans doing the research, you will always find leaks. ",
          "votes": 4
        },
        {
          "id": 543418,
          "postDate": "2019-06-04T14:19:40.953Z",
          "content": "<blockquote>\n  <p>The authors merged two quakes to one there (train set). One could say this was done on purpose to increase the mean TTF of train and to make train and test more similar…. </p>\n</blockquote>\n\n<p>Very interesting observation, I wish I had done this !!!</p>",
          "rawMarkdown": "&gt; The authors merged two quakes to one there (train set). One could say this was done on purpose to increase the mean TTF of train and to make train and test more similar…. \n\nVery interesting observation, I wish I had done this !!!"
        },
        {
          "id": 543423,
          "postDate": "2019-06-04T14:22:21.437Z",
          "content": "<p>I believe i have mentioned it about a month ago already</p>\n\n<p>Obviously, i wasn't eager to point everyone at it over and over again, as soon as we knew how important this aspect was.</p>",
          "rawMarkdown": "I believe i have mentioned it about a month ago already\n\nObviously, i wasn't eager to point everyone at it over and over again, as soon as we knew how important this aspect was.",
          "votes": 1
        },
        {
          "id": 543433,
          "postDate": "2019-06-04T14:26:07.233Z",
          "content": "<p>Thank you for your reply Ilu. It is very hard to convey the scientific aspect to new students who are fascinated with the AI/ML approaches ability to optimize any given performance function.</p>\n\n<p>I have to add that you did a very good job exploiting the data, congratulations. I hope that the organizers also learn from this competition in terms of trying to ask the right question and setting the right performance criteria.</p>",
          "rawMarkdown": "Thank you for your reply Ilu. It is very hard to convey the scientific aspect to new students who are fascinated with the AI/ML approaches ability to optimize any given performance function.\n\nI have to add that you did a very good job exploiting the data, congratulations. I hope that the organizers also learn from this competition in terms of trying to ask the right question and setting the right performance criteria.",
          "votes": 1
        },
        {
          "id": 543459,
          "postDate": "2019-06-04T14:35:16.040Z",
          "content": "<p>I would have felt better, if we really found a feature to detect those minor quakes. I also said about a month ago, if someone would truely be able to do that there is a nature paper incoming and they would win this challenge.\nAt that time my hypothesis was, that this is impossible.\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125525582\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125525582</a></p>\n\n<p>But sadly we didn't (and nobody else did) and thus, we used all knowledge that was available to come up with the solution discribed above. Sadly, that's the reality on kaggle and peaking at test is not prevented. </p>",
          "rawMarkdown": "I would have felt better, if we really found a feature to detect those minor quakes. I also said about a month ago, if someone would truely be able to do that there is a nature paper incoming and they would win this challenge.\nAt that time my hypothesis was, that this is impossible.\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125525582\n\nBut sadly we didn't (and nobody else did) and thus, we used all knowledge that was available to come up with the solution discribed above. Sadly, that's the reality on kaggle and peaking at test is not prevented. "
        },
        {
          "id": 543566,
          "postDate": "2019-06-04T15:43:26.240Z",
          "content": "<p>it is absolutely right for research, scientists shouldn't publish a paper based on such solution. this is the biggest lesson i learnt here. but this is a Kaggle competition. </p>",
          "rawMarkdown": "it is absolutely right for research, scientists shouldn't publish a paper based on such solution. this is the biggest lesson i learnt here. but this is a Kaggle competition. "
        },
        {
          "id": 544322,
          "postDate": "2019-06-05T11:46:13.877Z",
          "content": "<p>I agree with your opinion, but on Kaggle it's a common practice to use test data. You can look at other competitions. </p>",
          "rawMarkdown": "I agree with your opinion, but on Kaggle it's a common practice to use test data. You can look at other competitions. "
        },
        {
          "id": 565444,
          "postDate": "2019-07-01T00:41:33.977Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 543360,
      "postDate": "2019-06-04T13:50:24.457Z",
      "content": "<p>That's amazing solution. Big congrats to The Zoo team, well deserved.  </p>\n\n<p>Mimic the test statistics in train, using the paper p4677 or using testset statistics was the key to survive to the shakeup. </p>",
      "rawMarkdown": "That's amazing solution. Big congrats to The Zoo team, well deserved.  \n\nMimic the test statistics in train, using the paper p4677 or using testset statistics was the key to survive to the shakeup. ",
      "votes": 9
    },
    {
      "id": 564289,
      "postDate": "2019-06-29T08:40:21.790Z",
      "content": "<p>Thanks a lot for sharing this solution and a heartily congrates too. </p>",
      "rawMarkdown": "Thanks a lot for sharing this solution and a heartily congrates too. ",
      "votes": 7
    },
    {
      "id": 542992,
      "postDate": "2019-06-04T08:57:01.707Z",
      "content": "<p>Thank you for the wonderful write-up <a href=\"/philippsinger\">@philippsinger</a>\nIt was a pleasure to work with all of you and i will gladly team up with you again, <a href=\"/dott1718\">@dott1718</a> <a href=\"/returnofsputnik\">@returnofsputnik</a> <a href=\"/dkaraflos\">@dkaraflos</a> <a href=\"/pukkinming\">@pukkinming</a> </p>",
      "rawMarkdown": "Thank you for the wonderful write-up @philippsinger\nIt was a pleasure to work with all of you and i will gladly team up with you again, @dott1718 @returnofsputnik @dkaraflos @pukkinming ",
      "votes": 6
    },
    {
      "id": 543069,
      "postDate": "2019-06-04T10:06:59.923Z",
      "content": "<p>Thanks for sharing, and congratulation on the result.  Resampling was the way to go indeed, I wish we had thought of it!</p>",
      "rawMarkdown": "Thanks for sharing, and congratulation on the result.  Resampling was the way to go indeed, I wish we had thought of it!",
      "votes": 4
    },
    {
      "id": 543044,
      "postDate": "2019-06-04T09:46:28.083Z",
      "content": "<p>Congratulation! \nBecause of the differences in distribution between train and test, I gave up the competition early.\nI never image that we can make a sub-train set which have the same distribution with test.\nThanks for your wonderful sharing.</p>",
      "rawMarkdown": "Congratulation! \nBecause of the differences in distribution between train and test, I gave up the competition early.\nI never image that we can make a sub-train set which have the same distribution with test.\nThanks for your wonderful sharing.",
      "votes": 4
    },
    {
      "id": 626214,
      "postDate": "2019-09-14T03:02:59.247Z",
      "content": "<p>Hi,</p>\n\n<p>I know this gratitude comes late but I have to say thank you very much for sharing your awesome ideas! I am trying to reproduce your solution and I'm totally surprised by the fact that by selecting the earthquakes to align the training data and test data, my private score increases from 2.58 to 2.31, a top 10 result!</p>\n\n<p>I'm still digesting some of your ideas. But I want to ask a more general question: Is it always a good idea to align training set and test set? Thanks in advance!</p>",
      "rawMarkdown": "Hi,\n\nI know this gratitude comes late but I have to say thank you very much for sharing your awesome ideas! I am trying to reproduce your solution and I'm totally surprised by the fact that by selecting the earthquakes to align the training data and test data, my private score increases from 2.58 to 2.31, a top 10 result!\n\nI'm still digesting some of your ideas. But I want to ask a more general question: Is it always a good idea to align training set and test set? Thanks in advance!",
      "votes": 1,
      "replies": [
        {
          "id": 629852,
          "postDate": "2019-09-19T09:58:17.447Z",
          "content": "<p>That really depends on the type of problem and type of model. Sometimes models will still generalize better if you give them diverser training examples and aligning training and test sets can be really overfitty. So can't give a clear answer here.</p>",
          "rawMarkdown": "That really depends on the type of problem and type of model. Sometimes models will still generalize better if you give them diverser training examples and aligning training and test sets can be really overfitty. So can't give a clear answer here.",
          "votes": 3
        },
        {
          "id": 632778,
          "postDate": "2019-09-24T03:44:21.190Z",
          "content": "<p>Thank you very much!</p>",
          "rawMarkdown": "Thank you very much!"
        }
      ]
    },
    {
      "id": 555526,
      "postDate": "2019-06-19T02:48:34.737Z",
      "content": "<p>Congratulations, impressive!</p>",
      "rawMarkdown": "Congratulations, impressive!",
      "votes": 1
    },
    {
      "id": 553228,
      "postDate": "2019-06-15T10:36:39.790Z",
      "content": "<p>Congratulations!  Lots of amazing ideas !\nI have a question about:</p>\n\n<blockquote>\n  <p>To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating np.random.normal(0, 0.5, 150_000). Additionally, after noise addition, we subtracted the median of the segment.</p>\n</blockquote>\n\n<p>Why this preprocessing can help those time-trend features more useful.  Could you share the reasons behind this ?\nSorry if it is stupid question.</p>",
      "rawMarkdown": "Congratulations!  Lots of amazing ideas !\nI have a question about:\n&gt; To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating np.random.normal(0, 0.5, 150_000). Additionally, after noise addition, we subtracted the median of the segment.\n\nWhy this preprocessing can help those time-trend features more useful.  Could you share the reasons behind this ?\nSorry if it is stupid question.",
      "votes": 1,
      "replies": [
        {
          "id": 553415,
          "postDate": "2019-06-15T16:34:17.290Z",
          "content": "<p>This is not a stupid question. Thanks for asking <a href=\"/dodo74614\">@dodo74614</a> </p>\n\n<p>I'll try to explain the reasoning:\nWe have discrete values with steps of 1 for our acoustic data. As we know from EDA, the mean (and the median) of the segments slowly drifts within the experiment. Thus, almost all teams subtracted the mean of each segment in each segment. Unfortunatly, this subtraction makes median based features prone to error. Imagine a feature that evaluates a median value near zero. It will heavily depend on the corrected mean if this value is either -1(-mean), 0(-mean) or +1(-mean). (or any other small value)</p>\n\n<p>One candidate for such a feature is <code>absmedian</code>\n<code>z = z - np.mean(z); absmedian = np.median(abs(z))</code></p>\n\n<p>it's value is highly skewed to either 1.5 or 2.5. And it is escpecially different for train (higher mean) and test (lower mean).</p>\n\n<p>After adding the \"noise\" (std 0.5), which is actually lower or in the range of the expected accuracy of the sensor (1), <code>absmedian</code> behaves very normally (same distribution for train and test) and can be used for the models. The same thing applies to other median based features. \n<img src=\"https://i.imgur.com/SnAnYAB.png\" alt=\"noise or no noise\">\n<img src=\"https://i.imgur.com/nxzbmdR.png\" alt=\"after noise\"></p>\n\n<p>Additionally, adding the noise it allows us to remove the median instead of the mean, which is more robust with strong outliers like we have in some segments. </p>\n\n<p>I hope i could help.</p>",
          "rawMarkdown": "This is not a stupid question. Thanks for asking @dodo74614 \n\nI'll try to explain the reasoning:\nWe have discrete values with steps of 1 for our acoustic data. As we know from EDA, the mean (and the median) of the segments slowly drifts within the experiment. Thus, almost all teams subtracted the mean of each segment in each segment. Unfortunatly, this subtraction makes median based features prone to error. Imagine a feature that evaluates a median value near zero. It will heavily depend on the corrected mean if this value is either -1(-mean), 0(-mean) or +1(-mean). (or any other small value)\n\nOne candidate for such a feature is `absmedian`\n`z = z - np.mean(z); absmedian = np.median(abs(z))`\n\nit's value is highly skewed to either 1.5 or 2.5. And it is escpecially different for train (higher mean) and test (lower mean).\n\nAfter adding the \"noise\" (std 0.5), which is actually lower or in the range of the expected accuracy of the sensor (1), `absmedian` behaves very normally (same distribution for train and test) and can be used for the models. The same thing applies to other median based features. \n![noise or no noise](https://i.imgur.com/SnAnYAB.png)\n![after noise](https://i.imgur.com/nxzbmdR.png)\n\nAdditionally, adding the noise it allows us to remove the median instead of the mean, which is more robust with strong outliers like we have in some segments. \n\nI hope i could help.",
          "votes": 5
        },
        {
          "id": 553986,
          "postDate": "2019-06-16T17:42:44.927Z",
          "content": "<p>Wow. Very amazing ideas !!\nNow I am curious about how could you figure out this method?\nIt is heuristic or have some references, since I have never seen anyone else mentions this method.</p>",
          "rawMarkdown": "Wow. Very amazing ideas !!\nNow I am curious about how could you figure out this method?\nIt is heuristic or have some references, since I have never seen anyone else mentions this method.",
          "votes": 1
        },
        {
          "id": 555870,
          "postDate": "2019-06-19T14:47:03.453Z",
          "content": "<p>If you take mean of train and test, you can see the testing set has a higher mean than the training set. Therefore we elect to subtract the mean to make the test more similar to the train. However, even after subtracting the mean, the feature related to median (e.g. <code>np.median(z)</code>) still had dissimilarities between the train and test segment. To solve this, we tried subtracting solely the median (instead of the mean), but this still did not solve the issue. Therefore we injected small amount of noise, and then subtracted the median.</p>",
          "rawMarkdown": "If you take mean of train and test, you can see the testing set has a higher mean than the training set. Therefore we elect to subtract the mean to make the test more similar to the train. However, even after subtracting the mean, the feature related to median (e.g. `np.median(z)`) still had dissimilarities between the train and test segment. To solve this, we tried subtracting solely the median (instead of the mean), but this still did not solve the issue. Therefore we injected small amount of noise, and then subtracted the median.",
          "votes": 3
        },
        {
          "id": 566239,
          "postDate": "2019-07-02T00:25:37.583Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 567393,
          "postDate": "2019-07-03T12:46:30.157Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 567782,
          "postDate": "2019-07-04T02:21:22.093Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 567786,
          "postDate": "2019-07-04T02:30:15.053Z",
          "content": "<p>Hi <a href=\"/ilu000\">@ilu000</a> , could you please explain why 'Additionally, adding the noise it allows us to remove the median instead of the mean', this theory seems very amazing.</p>",
          "rawMarkdown": "Hi @ilu000 , could you please explain why 'Additionally, adding the noise it allows us to remove the median instead of the mean', this theory seems very amazing."
        },
        {
          "id": 568338,
          "postDate": "2019-07-04T18:28:20.890Z",
          "content": "<p>I think the issue with median is a bit different.  With initial raw data, median takes only 3 values 4, 4.5 or 5.  This is why most teams used mean instead.  Once noise is added then median value distribution is not as sparse.</p>",
          "rawMarkdown": "I think the issue with median is a bit different.  With initial raw data, median takes only 3 values 4, 4.5 or 5.  This is why most teams used mean instead.  Once noise is added then median value distribution is not as sparse.",
          "votes": 1
        },
        {
          "id": 568366,
          "postDate": "2019-07-04T19:40:07.700Z",
          "content": "<p>Spot on, <a href=\"/cpmpml\">@cpmpml</a> </p>",
          "rawMarkdown": "Spot on, @cpmpml ",
          "votes": 1
        },
        {
          "id": 568449,
          "postDate": "2019-07-05T00:33:11.440Z",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Yes, I also found this. But why adding noise can de-sparse median, do you know if is there a theory behind this?</p>",
          "rawMarkdown": "@cpmpml Yes, I also found this. But why adding noise can de-sparse median, do you know if is there a theory behind this?"
        }
      ]
    },
    {
      "id": 548101,
      "postDate": "2019-06-08T19:16:56.133Z",
      "content": "<p>Using the test set statistics to identify and give extra weight to similar events in the training set is very clever, congratulations. I'm a bit perplexed that given that there was data in the training set that were similar to the test set that the training with equal weight given to all the training data did not seem to result in an algorithm that could generalize well to both the public and private leaderboards. Do you have any thoughts about why this is so? Without weighting your training data, would your algorithm have achieved similar scores on both the public and private leaderboard?</p>\n\n<p>I wonder if your success mimicking the test set distribution points to a possible weakness in the training-testing split: imagine if there were 5 populations that the events were drawn from (A, B, C, D, E). A strong split would group items from (A, B, C, D) into the training set and leave items from E in the test set. But is it the case that the training set is items (A, B, C, D and a bit of E)? In the case of a strong split, there would have been little overlap with items in the training set and the test set, and there wouldn't have been an easy way weight your training data. </p>",
      "rawMarkdown": "Using the test set statistics to identify and give extra weight to similar events in the training set is very clever, congratulations. I'm a bit perplexed that given that there was data in the training set that were similar to the test set that the training with equal weight given to all the training data did not seem to result in an algorithm that could generalize well to both the public and private leaderboards. Do you have any thoughts about why this is so? Without weighting your training data, would your algorithm have achieved similar scores on both the public and private leaderboard?\n\nI wonder if your success mimicking the test set distribution points to a possible weakness in the training-testing split: imagine if there were 5 populations that the events were drawn from (A, B, C, D, E). A strong split would group items from (A, B, C, D) into the training set and leave items from E in the test set. But is it the case that the training set is items (A, B, C, D and a bit of E)? In the case of a strong split, there would have been little overlap with items in the training set and the test set, and there wouldn't have been an easy way weight your training data. ",
      "votes": 1,
      "replies": [
        {
          "id": 548109,
          "postDate": "2019-06-08T19:31:28.597Z",
          "content": "<p>CV and LB improvements are in line, but we were just not overfitting the public LB. </p>",
          "rawMarkdown": "CV and LB improvements are in line, but we were just not overfitting the public LB. "
        }
      ]
    },
    {
      "id": 544143,
      "postDate": "2019-06-05T07:38:32.283Z",
      "content": "<p>Sorry, I made a mistake in calculation.  I deleted my post.  My score should be 2.40.  Congrats for the amazing work!!!</p>",
      "rawMarkdown": "Sorry, I made a mistake in calculation.  I deleted my post.  My score should be 2.40.  Congrats for the amazing work!!!",
      "votes": 1
    },
    {
      "id": 544042,
      "postDate": "2019-06-05T04:04:23.433Z",
      "content": "<p>Thanks for sharing, very neat solution ! </p>\n\n<p>Could you please explain why adding noise to the data improves the mean and quantile based features ? </p>",
      "rawMarkdown": "Thanks for sharing, very neat solution ! \n\nCould you please explain why adding noise to the data improves the mean and quantile based features ? ",
      "votes": 1,
      "replies": [
        {
          "id": 553432,
          "postDate": "2019-06-15T17:08:58.553Z",
          "content": "<p>Thanks for the question, please see my answer above.</p>",
          "rawMarkdown": "Thanks for the question, please see my answer above."
        }
      ]
    },
    {
      "id": 543858,
      "postDate": "2019-06-04T22:26:45.677Z",
      "content": "<p>Wonderful solution! Thank you for sharing. By appreciating your great insight, I really want to do better next time. Good job!</p>",
      "rawMarkdown": "Wonderful solution! Thank you for sharing. By appreciating your great insight, I really want to do better next time. Good job!",
      "votes": 1
    },
    {
      "id": 543027,
      "postDate": "2019-06-04T09:26:02.720Z",
      "content": "<p>Congratulation for a very robust model,well deserved.</p>",
      "rawMarkdown": "Congratulation for a very robust model,well deserved.",
      "votes": 1
    },
    {
      "id": 543734,
      "postDate": "2019-06-04T18:33:47.827Z",
      "content": "<p>Amazing solution and well deserved winning! Great job!</p>",
      "rawMarkdown": "Amazing solution and well deserved winning! Great job!",
      "votes": 2
    },
    {
      "id": 553069,
      "postDate": "2019-06-15T04:43:47.263Z",
      "content": "<p>Out of curiosity ...</p>\n\n<p>Will it be ever possible to have a mathematical formula for predicting time-to-earthquake (based on acoustic or other signals)?</p>",
      "rawMarkdown": "Out of curiosity ...\n\nWill it be ever possible to have a mathematical formula for predicting time-to-earthquake (based on acoustic or other signals)?",
      "votes": -1
    },
    {
      "id": 3277687,
      "postDate": "2025-08-28T13:49:28.703Z",
      "content": "<p>Congratulations!<br>\nI gave up on the competition earlier because of the distribution shift between training and testing data.<br>\nI couldn’t imagine that a sub-train set could be built to align with the test distribution.<br>\nThank you for your excellent explanation.</p>",
      "rawMarkdown": "Congratulations!\nI gave up on the competition earlier because of the distribution shift between training and testing data.\nI couldn’t imagine that a sub-train set could be built to align with the test distribution.\nThank you for your excellent explanation."
    },
    {
      "id": 567770,
      "postDate": "2019-07-04T01:51:04.047Z",
      "content": "<p>Could you explain the theory of 'added noise to make median based features reliable.'? I am very curious about how this idea comes from.</p>",
      "rawMarkdown": "Could you explain the theory of 'added noise to make median based features reliable.'? I am very curious about how this idea comes from.",
      "replies": [
        {
          "id": 568158,
          "postDate": "2019-07-04T13:52:03.363Z",
          "content": "<p>This is elaborated in the comments below.</p>",
          "rawMarkdown": "This is elaborated in the comments below."
        }
      ]
    },
    {
      "id": 554034,
      "postDate": "2019-06-16T21:01:47.370Z",
      "content": "<p>Amazing solution!</p>",
      "rawMarkdown": "Amazing solution!"
    },
    {
      "id": 553057,
      "postDate": "2019-06-15T04:20:16.580Z",
      "content": "<p>Congratulations for winning this challenge of enormous significance.\nThank you for sharing your wisdom and your \"steps to the stardom\".\nThe way you have narrated will shine the subject of Statistics.</p>",
      "rawMarkdown": "Congratulations for winning this challenge of enormous significance.\nThank you for sharing your wisdom and your \"steps to the stardom\".\nThe way you have narrated will shine the subject of Statistics."
    },
    {
      "id": 549989,
      "postDate": "2019-06-11T07:42:01.877Z",
      "content": "<p>i canot fully understand your analysis with your ks test result(ks test plot ) .Can you share me your code  where create the picture1?</p>",
      "rawMarkdown": "i canot fully understand your analysis with your ks test result(ks test plot ) .Can you share me your code  where create the picture1?\n",
      "replies": [
        {
          "id": 550006,
          "postDate": "2019-06-11T08:02:36.470Z",
          "content": "<p>Each dot in that picture: (1) sample a certain number of earthquakes from train, (2) calculate KS statistic for each feature between sampled train and full test, (3) draw KS value on plot. Best sampling is based on lowest average KS statistic across features. </p>",
          "rawMarkdown": "Each dot in that picture: (1) sample a certain number of earthquakes from train, (2) calculate KS statistic for each feature between sampled train and full test, (3) draw KS value on plot. Best sampling is based on lowest average KS statistic across features. ",
          "votes": 3
        },
        {
          "id": 550082,
          "postDate": "2019-06-11T09:31:03.150Z",
          "content": "<p>ok,I understand.thank you for your reply.btw.Do you use scipy.stats.ks_2samp(train, test)?</p>",
          "rawMarkdown": "ok,I understand.thank you for your reply.btw.Do you use scipy.stats.ks_2samp(train, test)?"
        },
        {
          "id": 550104,
          "postDate": "2019-06-11T09:55:09.493Z",
          "content": "<p>Is  (ks statistic)/pvalue better than ks statistic?</p>",
          "rawMarkdown": "Is  (ks statistic)/pvalue better than ks statistic?"
        },
        {
          "id": 550455,
          "postDate": "2019-06-11T16:34:34.187Z",
          "content": "<p>Yeah, I use the scipy function. The p-value just gives you and indicator whether the null hypothesis of those two distributions being equal can be rejected. So in this case looking at the statistic or the p-value is kinda similar, we focused on the statistic here.</p>",
          "rawMarkdown": "Yeah, I use the scipy function. The p-value just gives you and indicator whether the null hypothesis of those two distributions being equal can be rejected. So in this case looking at the statistic or the p-value is kinda similar, we focused on the statistic here."
        },
        {
          "id": 550726,
          "postDate": "2019-06-12T01:18:00.253Z",
          "content": "<p>Thank you for your answer ! It gave me a lot of inspiration!</p>",
          "rawMarkdown": "Thank you for your answer ! It gave me a lot of inspiration!",
          "votes": 1
        }
      ]
    },
    {
      "id": 549410,
      "postDate": "2019-06-10T16:06:47.763Z",
      "content": "<p>Thanks for sharing this. Great analysis and a wonderful answer !!</p>",
      "rawMarkdown": "Thanks for sharing this. Great analysis and a wonderful answer !!"
    },
    {
      "id": 549117,
      "postDate": "2019-06-10T10:06:14.170Z",
      "content": "<p>i would like to get connected with dna sequence analysis</p>",
      "rawMarkdown": "i would like to get connected with dna sequence analysis"
    },
    {
      "id": 548897,
      "postDate": "2019-06-10T04:31:28.400Z",
      "content": "<p>Thank you for sharing! Congrats to The Zoo team!</p>",
      "rawMarkdown": "Thank you for sharing! Congrats to The Zoo team!"
    },
    {
      "id": 548333,
      "postDate": "2019-06-09T07:52:53.823Z",
      "content": "<p>Congratulations to the ZOO TEAM ! Outstanding achievement !</p>\n\n<p>Thank you Philip for the detailed explanations !</p>\n\n<p>Well deserved !</p>",
      "rawMarkdown": "Congratulations to the ZOO TEAM ! Outstanding achievement !\n\nThank you Philip for the detailed explanations !\n\nWell deserved !"
    },
    {
      "id": 547980,
      "postDate": "2019-06-08T15:37:06.717Z",
      "content": "<p>Congratulations !</p>",
      "rawMarkdown": "Congratulations !"
    },
    {
      "id": 547907,
      "postDate": "2019-06-08T13:52:27.663Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!"
    },
    {
      "id": 546719,
      "postDate": "2019-06-06T20:52:57.313Z",
      "content": "<p>Amazon work !!!</p>",
      "rawMarkdown": "Amazon work !!!"
    },
    {
      "id": 546132,
      "postDate": "2019-06-06T09:32:13.450Z",
      "content": "<p>So much great ideas and insights to study and use for future competitions. Brilliant solution!</p>",
      "rawMarkdown": "So much great ideas and insights to study and use for future competitions. Brilliant solution!"
    },
    {
      "id": 545936,
      "postDate": "2019-06-06T04:05:27.847Z",
      "content": "<p>This is beautiful. Congrats!</p>",
      "rawMarkdown": "This is beautiful. Congrats!"
    },
    {
      "id": 545869,
      "postDate": "2019-06-06T01:35:50.387Z",
      "content": "<p>Congrats! Thanks for sharing your solution!\nI learn a lot!\nespecially KS test part, very brilliant</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution!\nI learn a lot!\nespecially KS test part, very brilliant"
    },
    {
      "id": 544274,
      "postDate": "2019-06-05T10:47:19.990Z",
      "content": "<p>Thank you for sharing. What do you mean by  \"sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test\"? Did you compute your features for each of the test files, then randomly pick 10 earthquakes in the train data, randomly sampled them 10k times such that each sample is the same length as a test file (150000), computed your features on each sample, then computed the KS stat for features computed for this set of samples vs features computed for all test files? Forgive me if it's a stupid question.</p>",
      "rawMarkdown": "Thank you for sharing. What do you mean by  \"sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test\"? Did you compute your features for each of the test files, then randomly pick 10 earthquakes in the train data, randomly sampled them 10k times such that each sample is the same length as a test file (150000), computed your features on each sample, then computed the KS stat for features computed for this set of samples vs features computed for all test files? Forgive me if it's a stupid question.",
      "replies": [
        {
          "id": 544281,
          "postDate": "2019-06-05T10:56:12.030Z",
          "content": "<p>In training data we have 17 earthquakes. We sampled 10_000 times 10 earthquakes out of that (you could also just try all possible combinations) and compared how the feature distributions of those samples of 10 eqs map to the feature distributions in test. We used KS statistic as an estimate for the distribution similarity and finally picked those 10 eqs as training sample where the average KS statistic was the lowest. Does that make it more clear? </p>",
          "rawMarkdown": "In training data we have 17 earthquakes. We sampled 10_000 times 10 earthquakes out of that (you could also just try all possible combinations) and compared how the feature distributions of those samples of 10 eqs map to the feature distributions in test. We used KS statistic as an estimate for the distribution similarity and finally picked those 10 eqs as training sample where the average KS statistic was the lowest. Does that make it more clear? ",
          "votes": 5
        },
        {
          "id": 546100,
          "postDate": "2019-06-06T08:56:01.803Z",
          "content": "<p>Thanks for your reply (: Just to confirm, you randomly picked 10,000 combinations from the 19448 different combinations of 10 eqs you can get from the 17 eqs in train? (10,000 from 17 choose 10). Why 10 earthquakes? \nAnother question - you say you also only added features if the p value was larger than 0.05 for KS stat between test and train. Does this mean that you calculated a bunch of base features, then used those features to get the 10 earthquakes that you got, then got rid of the features which displayed a p value of 0.05 or less?</p>",
          "rawMarkdown": "Thanks for your reply (: Just to confirm, you randomly picked 10,000 combinations from the 19448 different combinations of 10 eqs you can get from the 17 eqs in train? (10,000 from 17 choose 10). Why 10 earthquakes? \nAnother question - you say you also only added features if the p value was larger than 0.05 for KS stat between test and train. Does this mean that you calculated a bunch of base features, then used those features to get the 10 earthquakes that you got, then got rid of the features which displayed a p value of 0.05 or less?"
        },
        {
          "id": 553411,
          "postDate": "2019-06-15T16:30:35.433Z",
          "content": "<p>To your first point: exactly that's how we did it. 10 EQs is close to how many EQs we expected in test, but kind of arbitrary. We found to have good agreement between train and test feature distributions choosing 10, but you can get to similar results choosing a different amount of EQs. To your second point: yes we only considered features with higher p-value. We also used a bunch of feature for the sampling procedure explained. </p>",
          "rawMarkdown": "To your first point: exactly that's how we did it. 10 EQs is close to how many EQs we expected in test, but kind of arbitrary. We found to have good agreement between train and test feature distributions choosing 10, but you can get to similar results choosing a different amount of EQs. To your second point: yes we only considered features with higher p-value. We also used a bunch of feature for the sampling procedure explained. "
        }
      ]
    },
    {
      "id": 544182,
      "postDate": "2019-06-05T08:42:21.913Z",
      "content": "<p>Amazing solution!\nCommit all 0 value,Private Score----mae 6.66993,Public Score----mae 4.01773.\nSo solution not important,I think.\nPrivate score depend on the same x，but very different y in public and private test set.\nDoes the test data is really produced by experiment?\nWhy it is very different?\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544182/13407/0.PNG\" alt=\"all 0 value,public and private mae\"></p>",
      "rawMarkdown": "Amazing solution!\nCommit all 0 value,Private Score----mae 6.66993,Public Score----mae 4.01773.\nSo solution not important,I think.\nPrivate score depend on the same x，but very different y in public and private test set.\nDoes the test data is really produced by experiment?\nWhy it is very different?\n![all 0 value,public and private mae](https://storage.googleapis.com/kaggle-forum-message-attachments/544182/13407/0.PNG)"
    },
    {
      "id": 543903,
      "postDate": "2019-06-04T23:32:44.137Z",
      "content": "<p>Congrats and thanks a lot for sharing! \nI think that I had that idea of sampling the train data looking at the test data. Basically, using the acoustic data feature, I tried to find in the train data the most similar 150k samples segment to each one of the test segments (using opencv and measuring similarity with R-Square). Then, in each of the segments I found, I created the features discussed in the public kernels. However it didn't work well and I am still wondering why...</p>",
      "rawMarkdown": "Congrats and thanks a lot for sharing! \nI think that I had that idea of sampling the train data looking at the test data. Basically, using the acoustic data feature, I tried to find in the train data the most similar 150k samples segment to each one of the test segments (using opencv and measuring similarity with R-Square). Then, in each of the segments I found, I created the features discussed in the public kernels. However it didn't work well and I am still wondering why...",
      "replies": [
        {
          "id": 545872,
          "postDate": "2019-06-06T01:45:45.553Z",
          "content": "<p>How did that work? Acoustic data has 150,000 raw datapoints. That's a lot of noise. Maybe taking absolute value could have helped, since I considered a positive acoustic data N to give the same information as a negative acoustic data N</p>",
          "rawMarkdown": "How did that work? Acoustic data has 150,000 raw datapoints. That's a lot of noise. Maybe taking absolute value could have helped, since I considered a positive acoustic data N to give the same information as a negative acoustic data N"
        }
      ]
    },
    {
      "id": 543843,
      "postDate": "2019-06-04T22:05:49.150Z",
      "content": "<p>&gt;&gt; KS statistic of train vs test</p>\n\n<p>very elegant! thank you for sharing</p>",
      "rawMarkdown": "&gt;&gt; KS statistic of train vs test\n\nvery elegant! thank you for sharing"
    },
    {
      "id": 543773,
      "postDate": "2019-06-04T20:05:29.693Z",
      "content": "<p>Thank you and congrat.\nWhat was the best corrcoef value that you could get between your features and ttf?\nMy best feature resulted a corrcoef of 0.679</p>",
      "rawMarkdown": "Thank you and congrat.\nWhat was the best corrcoef value that you could get between your features and ttf?\nMy best feature resulted a corrcoef of 0.679",
      "replies": [
        {
          "id": 543776,
          "postDate": "2019-06-04T20:09:20.637Z",
          "content": "<p>that depends on the training subset that we were using. But 0.67 pearson correlation seems about right for the \"number of peaks\" feature. The other features had less correlation but added second level information to the models. </p>",
          "rawMarkdown": "that depends on the training subset that we were using. But 0.67 pearson correlation seems about right for the \"number of peaks\" feature. The other features had less correlation but added second level information to the models. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 543601,
      "postDate": "2019-06-04T16:11:36.010Z",
      "content": "<p>Congrats! Thanks for sharing your solution. That's pretty amazing! </p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution. That's pretty amazing! "
    },
    {
      "id": 543559,
      "postDate": "2019-06-04T15:40:30.820Z",
      "content": "<p>tweak training dataset based on testing dataset is a big lesson I learnt. i suppose we should never ever touch testing set, but forgot that this is a Kaggle competition... </p>",
      "rawMarkdown": "tweak training dataset based on testing dataset is a big lesson I learnt. i suppose we should never ever touch testing set, but forgot that this is a Kaggle competition... ",
      "replies": [
        {
          "id": 553062,
          "postDate": "2019-06-15T04:28:12.527Z",
          "content": "<ul>\n<li>Not looking at the test data-set ::: Reading the textbook but not reviewing past exam papers</li>\n<li>Looking at the test data-set ::: Reading the textbook and also reviewing past exam papers</li>\n</ul>\n\n<p>Who will do better at an exam?</p>",
          "rawMarkdown": "- Not looking at the test data-set ::: Reading the textbook but not reviewing past exam papers\n- Looking at the test data-set ::: Reading the textbook and also reviewing past exam papers\n\nWho will do better at an exam?"
        }
      ]
    },
    {
      "id": 543530,
      "postDate": "2019-06-04T15:22:48.397Z",
      "content": "<p>Very well done! I tried something a little similar which I called AVMS or Adversarial Validation-Modulated Sampling. I assigned a weight to each row in train based on its AV score with a model that had an overal AUC of over 0.85 and used it to weight sampling of rows for training in each fold. This approach gave me the exact same LB score as when it wasn't used so I discarded it, but in hindsight it was a valuable discovery and I should have realised its importance. </p>",
      "rawMarkdown": "Very well done! I tried something a little similar which I called AVMS or Adversarial Validation-Modulated Sampling. I assigned a weight to each row in train based on its AV score with a model that had an overal AUC of over 0.85 and used it to weight sampling of rows for training in each fold. This approach gave me the exact same LB score as when it wasn't used so I discarded it, but in hindsight it was a valuable discovery and I should have realised its importance. ",
      "replies": [
        {
          "id": 543543,
          "postDate": "2019-06-04T15:28:51.460Z",
          "content": "<p>The problem still is that you need to fix the signal in some way to make certain features more consistent, like we did with the noise and median subtraction.</p>",
          "rawMarkdown": "The problem still is that you need to fix the signal in some way to make certain features more consistent, like we did with the noise and median subtraction.",
          "votes": 1
        },
        {
          "id": 543630,
          "postDate": "2019-06-04T16:39:08.890Z",
          "content": "<p>I did normalise the data, but I didn't denoise it. My rationale was that earthquakes are stochastic processes, and therefore measurements of stochastic noise might yield valuable features that would be lost with denoising. I extracted many useful features like various fractal dimensions and measures of entropy that helped my model, but ultimately it seems I took the wrong approach. The Petrosian Fractal Dimension in particular is a very strong feature. I should have denoised when extracting the other, less complex features though. Lesson learned!</p>",
          "rawMarkdown": "I did normalise the data, but I didn't denoise it. My rationale was that earthquakes are stochastic processes, and therefore measurements of stochastic noise might yield valuable features that would be lost with denoising. I extracted many useful features like various fractal dimensions and measures of entropy that helped my model, but ultimately it seems I took the wrong approach. The Petrosian Fractal Dimension in particular is a very strong feature. I should have denoised when extracting the other, less complex features though. Lesson learned!"
        },
        {
          "id": 543640,
          "postDate": "2019-06-04T16:49:23.337Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> is referring to the ADDED noise to make median based features reliable. Removing the mean is not enough for normalising. </p>\n\n<p>Denoising the data did give a boost for only few features (e.g. num_peaks).</p>",
          "rawMarkdown": "@philippsinger is referring to the ADDED noise to make median based features reliable. Removing the mean is not enough for normalising. \n\n\nDenoising the data did give a boost for only few features (e.g. num_peaks).",
          "votes": 1
        },
        {
          "id": 543645,
          "postDate": "2019-06-04T16:54:03.480Z",
          "content": "<p>Oh, the way he worded it made me think he was talking about noise subtraction. In either case, modifying the noise would have impacted my entropy-based features. I should have given it more consideration, but I can get tunnel vision sometimes and was probably focusing too much on the wrong areas! I hope we get another signal processing competition in the future, I learned a lot about it this time.</p>",
          "rawMarkdown": "Oh, the way he worded it made me think he was talking about noise subtraction. In either case, modifying the noise would have impacted my entropy-based features. I should have given it more consideration, but I can get tunnel vision sometimes and was probably focusing too much on the wrong areas! I hope we get another signal processing competition in the future, I learned a lot about it this time."
        }
      ]
    },
    {
      "id": 543498,
      "postDate": "2019-06-04T15:08:43.283Z",
      "content": "<p>Beautiful solution and very well explained. Congrats!!!</p>",
      "rawMarkdown": "Beautiful solution and very well explained. Congrats!!!"
    },
    {
      "id": 543497,
      "postDate": "2019-06-04T15:08:33.030Z",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution!"
    },
    {
      "id": 543422,
      "postDate": "2019-06-04T14:21:49.260Z",
      "content": "<p>Thanks for sharing and congratulations to all The Zoo!!\nQuick question: what is the score of the second submission, if I may ask?\nAlologies if I overlooked\nThx</p>",
      "rawMarkdown": "Thanks for sharing and congratulations to all The Zoo!!\nQuick question: what is the score of the second submission, if I may ask?\nAlologies if I overlooked\nThx",
      "replies": [
        {
          "id": 543427,
          "postDate": "2019-06-04T14:24:30.347Z",
          "content": "<p>2.29351 \nIt would have also gotten first place and is trained on another subset matching test quite good.</p>\n\n<p>Some other submission have scores in the 2.25 range, but those were not predictable at the time we submitted them. </p>",
          "rawMarkdown": "2.29351 \nIt would have also gotten first place and is trained on another subset matching test quite good.\n\nSome other submission have scores in the 2.25 range, but those were not predictable at the time we submitted them. ",
          "votes": 1
        },
        {
          "id": 543434,
          "postDate": "2019-06-04T14:26:27.643Z",
          "content": "<p>private 2.29351\npublic 1.61273</p>",
          "rawMarkdown": "private 2.29351\npublic 1.61273",
          "votes": 1
        },
        {
          "id": 543449,
          "postDate": "2019-06-04T14:31:55.757Z",
          "content": "<p>Amazing!</p>",
          "rawMarkdown": "Amazing!"
        },
        {
          "id": 543505,
          "postDate": "2019-06-04T15:10:01.387Z",
          "content": "<p>Yeah, for the second submit we basically added one more shorter EQ to the train data, as we got a bit of cold feet with the mean being so high to have a backup. Logically we should have taken an even higher mean. Actually the mean of our second submit prediction is only ~5.9, but as already said, would have also made first place.</p>",
          "rawMarkdown": "Yeah, for the second submit we basically added one more shorter EQ to the train data, as we got a bit of cold feet with the mean being so high to have a backup. Logically we should have taken an even higher mean. Actually the mean of our second submit prediction is only ~5.9, but as already said, would have also made first place."
        }
      ]
    },
    {
      "id": 543389,
      "postDate": "2019-06-04T14:03:28.843Z",
      "content": "<p>Did you try denoise autoencoder?</p>",
      "rawMarkdown": "Did you try denoise autoencoder?",
      "replies": [
        {
          "id": 543410,
          "postDate": "2019-06-04T14:14:55.533Z",
          "content": "<p>We thought about it at some point and we also tried to train models on raw data, but we didn't get any improvements out of it. </p>",
          "rawMarkdown": "We thought about it at some point and we also tried to train models on raw data, but we didn't get any improvements out of it. ",
          "votes": 1
        },
        {
          "id": 543436,
          "postDate": "2019-06-04T14:27:31.077Z",
          "content": "<p>Which type of denoise autoencoder did you try? Maybe wavenet denoiser <a href=\"https://arxiv.org/abs/1706.07162\">https://arxiv.org/abs/1706.07162</a>?</p>",
          "rawMarkdown": "Which type of denoise autoencoder did you try? Maybe wavenet denoiser https://arxiv.org/abs/1706.07162?\n"
        },
        {
          "id": 543461,
          "postDate": "2019-06-04T14:38:36.740Z",
          "content": "<p>I am sorry, i meant we did not try autoencoders as sadly no one of the team was an expert for that. Thats why i said, we were thinking about it (and maybe we would then need to learn about it). \nBut then all models (NNs mostly) trained on raw data failed, and we burried that idea. </p>",
          "rawMarkdown": "I am sorry, i meant we did not try autoencoders as sadly no one of the team was an expert for that. Thats why i said, we were thinking about it (and maybe we would then need to learn about it). \nBut then all models (NNs mostly) trained on raw data failed, and we burried that idea. "
        },
        {
          "id": 543473,
          "postDate": "2019-06-04T14:45:40.563Z",
          "content": "<p>And all my model trained on raw data failed too. I tried u-net models with weibull loss on raw data this model was burried too. </p>",
          "rawMarkdown": "And all my model trained on raw data failed too. I tried u-net models with weibull loss on raw data this model was burried too. "
        }
      ]
    },
    {
      "id": 543373,
      "postDate": "2019-06-04T13:56:52.467Z",
      "content": "<p>Congratulations !</p>",
      "rawMarkdown": "Congratulations !"
    },
    {
      "id": 543301,
      "postDate": "2019-06-04T13:01:47.377Z",
      "content": "<p>Thank you for sharing and congratulations. Getting 1st place without using any test ttf leak information from paper is great!\nI guess re-sampling method is the key, however it is not possible to check if this re-sampling affects better or not without looking LB. So you just trust your method works well until private LB opens, that is great.</p>",
      "rawMarkdown": "Thank you for sharing and congratulations. Getting 1st place without using any test ttf leak information from paper is great!\nI guess re-sampling method is the key, however it is not possible to check if this re-sampling affects better or not without looking LB. So you just trust your method works well until private LB opens, that is great."
    },
    {
      "id": 543195,
      "postDate": "2019-06-04T11:21:28.337Z",
      "content": "<p>Excellent <a href=\"/philippsinger\">@philippsinger</a>!\nQuick question(s):</p>\n\n<p>\"After careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. \"</p>\n\n<ol>\n<li><p>So you threw other EQ away or you added additional samples that were representative of earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10]? How big was you train data in the end anyways?</p></li>\n<li><p>Did you try using a model instead of KS statistic? For example using AV for individual features and only the ones that have low AUC score you would retain. (meaning they (should) have similiar distributions)</p></li>\n</ol>",
      "rawMarkdown": "Excellent @philippsinger!\nQuick question(s):\n\n\"After careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. \"\n\n1. So you threw other EQ away or you added additional samples that were representative of earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10]? How big was you train data in the end anyways?\n\n2. Did you try using a model instead of KS statistic? For example using AV for individual features and only the ones that have low AUC score you would retain. (meaning they (should) have similiar distributions)",
      "replies": [
        {
          "id": 543208,
          "postDate": "2019-06-04T11:29:21.780Z",
          "content": "<ol>\n<li>We only used the data of those quakes. There were other possible subsets, but this subset had the best KS statistic for oof and predictions. </li>\n<li>We based our feature selection on KS statistic and visual examination in train and test. Adding noise to the train and removing the median allowed us to use more reliable features. (e.g. median based features)</li>\n</ol>",
          "rawMarkdown": "1. We only used the data of those quakes. There were other possible subsets, but this subset had the best KS statistic for oof and predictions. \n2. We based our feature selection on KS statistic and visual examination in train and test. Adding noise to the train and removing the median allowed us to use more reliable features. (e.g. median based features)\n\n",
          "votes": 1
        },
        {
          "id": 543222,
          "postDate": "2019-06-04T11:46:34.610Z",
          "content": "<p>Thank you <a href=\"/ilu000\">@ilu000</a> !</p>\n\n<ol>\n<li><p>I understand, did you try or hypothesise more samples of these earthquakes, I would assume that if you agmented it according to this it would allow model to learn more than just throwing the others EQ away?</p></li>\n<li><p>I read the text. (Hopefully thoroughly) But my question was did you try something else to test the distributions instead of kolmogorov smirnoff test?  For example here is something I was playing around with: A model that gives a probability of a sample being in train or test (old idea):</p></li>\n</ol>\n\n<p>` </p>\n\n<pre><code>test['target'] = 0\n\ntrain['is_test'] = 0\ntest['is_test'] = 1\n\norig_train = train.copy()\n\n\ntrain = pd.concat(( orig_train, test ))\ntrain.reset_index( inplace = True, drop = True )\n\nx = train.drop( [ 'is_test', 'target' ], axis = 1 )\ny = train.is_test\n\n\nn_estimators = 100\nclf = RF( n_estimators = n_estimators, n_jobs = -1 )\n\npredictions = np.zeros( y.shape )\n\ncv = StratifiedKFold( n_splits = 5, shuffle = True, random_state =123 )\n\nfor train_i, test_i  in cv.split(train, y):\n\nprint (\"# fold {}, {}\".format( f + 1, ctime()))\n\nx_train = x.iloc[train_i]\nx_test = x.iloc[test_i]\ny_train = y.iloc[train_i]\ny_test = y.iloc[test_i]\n\nclf.fit( x_train, y_train ) \n\np = clf.predict_proba( x_test )[:,1]\n\nauc = AUC( y_test, p )\nprint(\"# AUC: {:.2%}\\n\".format( auc ))  \n\npredictions[ test_i ] = p\n\ntrain['p'] = predictions\n\ni = predictions.argsort()\ntrain_sorted = train.iloc[i]\n</code></pre>\n\n<p>`</p>\n\n<p>If yes, how did that go?</p>",
          "rawMarkdown": "Thank you @ilu000 !\n\n1. I understand, did you try or hypothesise more samples of these earthquakes, I would assume that if you agmented it according to this it would allow model to learn more than just throwing the others EQ away?\n\n2. I read the text. (Hopefully thoroughly) But my question was did you try something else to test the distributions instead of kolmogorov smirnoff test?  For example here is something I was playing around with: A model that gives a probability of a sample being in train or test (old idea):\n\n` \n\n\n    test['target'] = 0\n\n    train['is_test'] = 0\n    test['is_test'] = 1\n\n    orig_train = train.copy()\n\n\n    train = pd.concat(( orig_train, test ))\n    train.reset_index( inplace = True, drop = True )\n\n    x = train.drop( [ 'is_test', 'target' ], axis = 1 )\n    y = train.is_test\n\n\n    n_estimators = 100\n    clf = RF( n_estimators = n_estimators, n_jobs = -1 )\n\n    predictions = np.zeros( y.shape )\n\n    cv = StratifiedKFold( n_splits = 5, shuffle = True, random_state =123 )\n\n    for train_i, test_i  in cv.split(train, y):\n\n\tprint (\"# fold {}, {}\".format( f + 1, ctime()))\n\n\tx_train = x.iloc[train_i]\n\tx_test = x.iloc[test_i]\n\ty_train = y.iloc[train_i]\n\ty_test = y.iloc[test_i]\n\t\n\tclf.fit( x_train, y_train )\t\n\n\tp = clf.predict_proba( x_test )[:,1]\n\t\n\tauc = AUC( y_test, p )\n\tprint(\"# AUC: {:.2%}\\n\".format( auc ))\t\n\t\n\tpredictions[ test_i ] = p\n\n    train['p'] = predictions\n\t\n    i = predictions.argsort()\n    train_sorted = train.iloc[i]\n\n`\n\nIf yes, how did that go?"
        },
        {
          "id": 543228,
          "postDate": "2019-06-04T11:53:34.410Z",
          "content": "<p>No further augmenting did not help because you can only augment certain areas. We also only focused on KS statistic, I see no necessity to use something more complex for this data, Occams razor.</p>",
          "rawMarkdown": "No further augmenting did not help because you can only augment certain areas. We also only focused on KS statistic, I see no necessity to use something more complex for this data, Occams razor.",
          "votes": 1
        }
      ]
    },
    {
      "id": 543101,
      "postDate": "2019-06-04T10:27:22.147Z",
      "content": "<p>Congratulations and thanks for sharing. Very neat solution.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing. Very neat solution."
    },
    {
      "id": 543076,
      "postDate": "2019-06-04T10:12:37.903Z",
      "content": "<p>Congratulations! and thanks for sharing your great solution!!</p>",
      "rawMarkdown": "Congratulations! and thanks for sharing your great solution!!"
    },
    {
      "id": 566308,
      "postDate": "2019-07-02T03:21:48.270Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 544190,
      "postDate": "2019-06-05T08:50:42.040Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 566281,
      "postDate": "2019-07-02T02:01:43.217Z",
      "content": "<p>Thanks for sharing your wisdom.</p>",
      "rawMarkdown": "Thanks for sharing your wisdom.",
      "votes": 1
    },
    {
      "id": 543330,
      "postDate": "2019-06-04T13:27:50.940Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 542985,
      "postDate": "2019-06-04T08:52:24.783Z",
      "content": "<p>Congratulations! Thank you for sharing!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing!"
    },
    {
      "id": 2576195,
      "postDate": "2023-12-27T14:44:30.617Z",
      "content": "<p>Thanks a lot for sharing this solution</p>",
      "rawMarkdown": "Thanks a lot for sharing this solution"
    },
    {
      "id": 555604,
      "postDate": "2019-06-19T06:22:59.173Z",
      "content": "<p>Nice work, Thanks for sharing </p>",
      "rawMarkdown": "Nice work, Thanks for sharing "
    },
    {
      "id": 548389,
      "postDate": "2019-06-09T09:59:49.407Z",
      "content": "<p><strong>Congrats and thanks for sharing</strong></p>",
      "rawMarkdown": "**Congrats and thanks for sharing**"
    },
    {
      "id": 548371,
      "postDate": "2019-06-09T09:35:53.613Z",
      "content": "<p>Congratulations! Thank you for sharing!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing!"
    },
    {
      "id": 546528,
      "postDate": "2019-06-06T16:36:19.613Z",
      "content": "<p>Congratulations! Thanks for sharing :)</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing :)"
    },
    {
      "id": 546124,
      "postDate": "2019-06-06T09:21:06.137Z",
      "content": "<p>Thank you for sharing! Congrats !</p>",
      "rawMarkdown": "Thank you for sharing! Congrats !"
    },
    {
      "id": 544395,
      "postDate": "2019-06-05T13:12:31.287Z",
      "content": "<p>Congratulations! Thanks for sharing!</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing!"
    },
    {
      "id": 543249,
      "postDate": "2019-06-04T12:16:09.513Z",
      "content": "<p>Thank's</p>",
      "rawMarkdown": "Thank's"
    },
    {
      "id": 543059,
      "postDate": "2019-06-04T10:02:11.753Z",
      "content": "<p>Congratulations! Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing."
    },
    {
      "id": 543052,
      "postDate": "2019-06-04T09:53:31.410Z",
      "content": "<p>Congratulations! Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing."
    },
    {
      "id": 543048,
      "postDate": "2019-06-04T09:49:21.273Z",
      "content": "<p>Congrats! Thanks for sharing :) </p>",
      "rawMarkdown": "Congrats! Thanks for sharing :) "
    }
  ],
  "comments": [
    {
      "id": 543397,
      "author_name": "Satoshi",
      "author_url": "",
      "post_date": "2019-06-04T14:08:51.280000",
      "content": "<p>In a properly set scientific experiment the data would be <strong>separated</strong> into training and testing. In other words the testing data is locked in a box, the scientist goes to the box only after he/she has developed a hypothesis (a method/model etc). The testing data is then used only once to <strong>test</strong> this particular hypothesis. If the researcher is allowed to peek, analyze the testing data then the test loses its validity. Once the researcher starts tuning a hypothesis with respect to a data set then he/she needs another data set to test the reached conclusions.</p>\n\n<p>TLDR: The top ranking method has been tuned to the testing dataset via the KS test. This means that the method has become specific to this particular testing set. If the LANL researchers' goal is to apply the obtained findings to other experiments then it is very likely that they will be better of using a lower ranking method.</p>\n\n<p>The competition/ranking has lost its practical value because the organizers made the testing data public. They should have kept the testing data on the server side and allow only for submission of executable code (see Matlab Cody competitions for instance), preventing the testing data being used for tuning.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 543403,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T14:13:32.633000",
          "content": "<p>Absolutely right. \nBut please look at the scientific paper with exp. 4677. The authors merged two quakes to one there (train set). One could say this was done on purpose to increase the mean TTF of train and to make train and test more similar.... </p>\n\n<p>As long as there are humans doing the research, you will always find leaks. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 543418,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-04T14:19:40.953000",
          "content": "<blockquote>\n  <p>The authors merged two quakes to one there (train set). One could say this was done on purpose to increase the mean TTF of train and to make train and test more similar…. </p>\n</blockquote>\n\n<p>Very interesting observation, I wish I had done this !!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543423,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T14:22:21.437000",
          "content": "<p>I believe i have mentioned it about a month ago already</p>\n\n<p>Obviously, i wasn't eager to point everyone at it over and over again, as soon as we knew how important this aspect was.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543433,
          "author_name": "Satoshi",
          "author_url": "",
          "post_date": "2019-06-04T14:26:07.233000",
          "content": "<p>Thank you for your reply Ilu. It is very hard to convey the scientific aspect to new students who are fascinated with the AI/ML approaches ability to optimize any given performance function.</p>\n\n<p>I have to add that you did a very good job exploiting the data, congratulations. I hope that the organizers also learn from this competition in terms of trying to ask the right question and setting the right performance criteria.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543459,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-04T14:35:16.040000",
          "content": "<p>I would have felt better, if we really found a feature to detect those minor quakes. I also said about a month ago, if someone would truely be able to do that there is a nature paper incoming and they would win this challenge.\nAt that time my hypothesis was, that this is impossible.\n<a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125525582\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/91125525582</a></p>\n\n<p>But sadly we didn't (and nobody else did) and thus, we used all knowledge that was available to come up with the solution discribed above. Sadly, that's the reality on kaggle and peaking at test is not prevented. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543566,
          "author_name": "Z. Liu",
          "author_url": "",
          "post_date": "2019-06-04T15:43:26.240000",
          "content": "<p>it is absolutely right for research, scientists shouldn't publish a paper based on such solution. this is the biggest lesson i learnt here. but this is a Kaggle competition. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 544322,
          "author_name": "Derkanat",
          "author_url": "",
          "post_date": "2019-06-05T11:46:13.877000",
          "content": "<p>I agree with your opinion, but on Kaggle it's a common practice to use test data. You can look at other competitions. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 565444,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-01T00:41:33.977000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543360,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T13:50:24.457000",
      "content": "<p>That's amazing solution. Big congrats to The Zoo team, well deserved.  </p>\n\n<p>Mimic the test statistics in train, using the paper p4677 or using testset statistics was the key to survive to the shakeup. </p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 564289,
      "author_name": "Raju Kumar Mishra",
      "author_url": "",
      "post_date": "2019-06-29T08:40:21.790000",
      "content": "<p>Thanks a lot for sharing this solution and a heartily congrates too. </p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 542992,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2019-06-04T08:57:01.707000",
      "content": "<p>Thank you for the wonderful write-up <a href=\"/philippsinger\">@philippsinger</a>\nIt was a pleasure to work with all of you and i will gladly team up with you again, <a href=\"/dott1718\">@dott1718</a> <a href=\"/returnofsputnik\">@returnofsputnik</a> <a href=\"/dkaraflos\">@dkaraflos</a> <a href=\"/pukkinming\">@pukkinming</a> </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 543069,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T10:06:59.923000",
      "content": "<p>Thanks for sharing, and congratulation on the result.  Resampling was the way to go indeed, I wish we had thought of it!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 543044,
      "author_name": "Correlation",
      "author_url": "",
      "post_date": "2019-06-04T09:46:28.083000",
      "content": "<p>Congratulation! \nBecause of the differences in distribution between train and test, I gave up the competition early.\nI never image that we can make a sub-train set which have the same distribution with test.\nThanks for your wonderful sharing.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 626214,
      "author_name": "Dondon2019",
      "author_url": "",
      "post_date": "2019-09-14T03:02:59.247000",
      "content": "<p>Hi,</p>\n\n<p>I know this gratitude comes late but I have to say thank you very much for sharing your awesome ideas! I am trying to reproduce your solution and I'm totally surprised by the fact that by selecting the earthquakes to align the training data and test data, my private score increases from 2.58 to 2.31, a top 10 result!</p>\n\n<p>I'm still digesting some of your ideas. But I want to ask a more general question: Is it always a good idea to align training set and test set? Thanks in advance!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 629852,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-09-19T09:58:17.447000",
          "content": "<p>That really depends on the type of problem and type of model. Sometimes models will still generalize better if you give them diverser training examples and aligning training and test sets can be really overfitty. So can't give a clear answer here.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 632778,
          "author_name": "Dondon2019",
          "author_url": "",
          "post_date": "2019-09-24T03:44:21.190000",
          "content": "<p>Thank you very much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 555526,
      "author_name": "Vettejeep",
      "author_url": "",
      "post_date": "2019-06-19T02:48:34.737000",
      "content": "<p>Congratulations, impressive!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 553228,
      "author_name": "aaaa",
      "author_url": "",
      "post_date": "2019-06-15T10:36:39.790000",
      "content": "<p>Congratulations!  Lots of amazing ideas !\nI have a question about:</p>\n\n<blockquote>\n  <p>To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating np.random.normal(0, 0.5, 150_000). Additionally, after noise addition, we subtracted the median of the segment.</p>\n</blockquote>\n\n<p>Why this preprocessing can help those time-trend features more useful.  Could you share the reasons behind this ?\nSorry if it is stupid question.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 553415,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-15T16:34:17.290000",
          "content": "<p>This is not a stupid question. Thanks for asking <a href=\"/dodo74614\">@dodo74614</a> </p>\n\n<p>I'll try to explain the reasoning:\nWe have discrete values with steps of 1 for our acoustic data. As we know from EDA, the mean (and the median) of the segments slowly drifts within the experiment. Thus, almost all teams subtracted the mean of each segment in each segment. Unfortunatly, this subtraction makes median based features prone to error. Imagine a feature that evaluates a median value near zero. It will heavily depend on the corrected mean if this value is either -1(-mean), 0(-mean) or +1(-mean). (or any other small value)</p>\n\n<p>One candidate for such a feature is <code>absmedian</code>\n<code>z = z - np.mean(z); absmedian = np.median(abs(z))</code></p>\n\n<p>it's value is highly skewed to either 1.5 or 2.5. And it is escpecially different for train (higher mean) and test (lower mean).</p>\n\n<p>After adding the \"noise\" (std 0.5), which is actually lower or in the range of the expected accuracy of the sensor (1), <code>absmedian</code> behaves very normally (same distribution for train and test) and can be used for the models. The same thing applies to other median based features. \n<img src=\"https://i.imgur.com/SnAnYAB.png\" alt=\"noise or no noise\">\n<img src=\"https://i.imgur.com/nxzbmdR.png\" alt=\"after noise\"></p>\n\n<p>Additionally, adding the noise it allows us to remove the median instead of the mean, which is more robust with strong outliers like we have in some segments. </p>\n\n<p>I hope i could help.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 553986,
          "author_name": "aaaa",
          "author_url": "",
          "post_date": "2019-06-16T17:42:44.927000",
          "content": "<p>Wow. Very amazing ideas !!\nNow I am curious about how could you figure out this method?\nIt is heuristic or have some references, since I have never seen anyone else mentions this method.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 555870,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2019-06-19T14:47:03.453000",
          "content": "<p>If you take mean of train and test, you can see the testing set has a higher mean than the training set. Therefore we elect to subtract the mean to make the test more similar to the train. However, even after subtracting the mean, the feature related to median (e.g. <code>np.median(z)</code>) still had dissimilarities between the train and test segment. To solve this, we tried subtracting solely the median (instead of the mean), but this still did not solve the issue. Therefore we injected small amount of noise, and then subtracted the median.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 566239,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-02T00:25:37.583000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 567393,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-03T12:46:30.157000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 567782,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-04T02:21:22.093000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 567786,
          "author_name": "PJTucker",
          "author_url": "",
          "post_date": "2019-07-04T02:30:15.053000",
          "content": "<p>Hi <a href=\"/ilu000\">@ilu000</a> , could you please explain why 'Additionally, adding the noise it allows us to remove the median instead of the mean', this theory seems very amazing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 568338,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-07-04T18:28:20.890000",
          "content": "<p>I think the issue with median is a bit different.  With initial raw data, median takes only 3 values 4, 4.5 or 5.  This is why most teams used mean instead.  Once noise is added then median value distribution is not as sparse.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 568366,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-07-04T19:40:07.700000",
          "content": "<p>Spot on, <a href=\"/cpmpml\">@cpmpml</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 568449,
          "author_name": "PJTucker",
          "author_url": "",
          "post_date": "2019-07-05T00:33:11.440000",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Yes, I also found this. But why adding noise can de-sparse median, do you know if is there a theory behind this?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 548101,
      "author_name": "small yellow duck",
      "author_url": "",
      "post_date": "2019-06-08T19:16:56.133000",
      "content": "<p>Using the test set statistics to identify and give extra weight to similar events in the training set is very clever, congratulations. I'm a bit perplexed that given that there was data in the training set that were similar to the test set that the training with equal weight given to all the training data did not seem to result in an algorithm that could generalize well to both the public and private leaderboards. Do you have any thoughts about why this is so? Without weighting your training data, would your algorithm have achieved similar scores on both the public and private leaderboard?</p>\n\n<p>I wonder if your success mimicking the test set distribution points to a possible weakness in the training-testing split: imagine if there were 5 populations that the events were drawn from (A, B, C, D, E). A strong split would group items from (A, B, C, D) into the training set and leave items from E in the test set. But is it the case that the training set is items (A, B, C, D and a bit of E)? In the case of a strong split, there would have been little overlap with items in the training set and the test set, and there wouldn't have been an easy way weight your training data. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 548109,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-08T19:31:28.597000",
          "content": "<p>CV and LB improvements are in line, but we were just not overfitting the public LB. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 544143,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "2019-06-05T07:38:32.283000",
      "content": "<p>Sorry, I made a mistake in calculation.  I deleted my post.  My score should be 2.40.  Congrats for the amazing work!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 544042,
      "author_name": "Ouadada",
      "author_url": "",
      "post_date": "2019-06-05T04:04:23.433000",
      "content": "<p>Thanks for sharing, very neat solution ! </p>\n\n<p>Could you please explain why adding noise to the data improves the mean and quantile based features ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 553432,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2019-06-15T17:08:58.553000",
          "content": "<p>Thanks for the question, please see my answer above.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543858,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2019-06-04T22:26:45.677000",
      "content": "<p>Wonderful solution! Thank you for sharing. By appreciating your great insight, I really want to do better next time. Good job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543027,
      "author_name": "OverfitModel",
      "author_url": "",
      "post_date": "2019-06-04T09:26:02.720000",
      "content": "<p>Congratulation for a very robust model,well deserved.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543734,
      "author_name": "Shize Su",
      "author_url": "",
      "post_date": "2019-06-04T18:33:47.827000",
      "content": "<p>Amazing solution and well deserved winning! Great job!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 553069,
      "author_name": "Lalit",
      "author_url": "",
      "post_date": "2019-06-15T04:43:47.263000",
      "content": "<p>Out of curiosity ...</p>\n\n<p>Will it be ever possible to have a mathematical formula for predicting time-to-earthquake (based on acoustic or other signals)?</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3277687,
      "author_name": "Prashant Kumar",
      "author_url": "",
      "post_date": "2025-08-28T13:49:28.703000",
      "content": "<p>Congratulations!<br>\nI gave up on the competition earlier because of the distribution shift between training and testing data.<br>\nI couldn’t imagine that a sub-train set could be built to align with the test distribution.<br>\nThank you for your excellent explanation.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 567770,
      "author_name": "PJTucker",
      "author_url": "",
      "post_date": "2019-07-04T01:51:04.047000",
      "content": "<p>Could you explain the theory of 'added noise to make median based features reliable.'? I am very curious about how this idea comes from.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 568158,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-04T13:52:03.363000",
          "content": "<p>This is elaborated in the comments below.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 554034,
      "author_name": "Pierre-Adrien",
      "author_url": "",
      "post_date": "2019-06-16T21:01:47.370000",
      "content": "<p>Amazing solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 553057,
      "author_name": "Lalit",
      "author_url": "",
      "post_date": "2019-06-15T04:20:16.580000",
      "content": "<p>Congratulations for winning this challenge of enormous significance.\nThank you for sharing your wisdom and your \"steps to the stardom\".\nThe way you have narrated will shine the subject of Statistics.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549989,
      "author_name": "Xiang wang",
      "author_url": "",
      "post_date": "2019-06-11T07:42:01.877000",
      "content": "<p>i canot fully understand your analysis with your ks test result(ks test plot ) .Can you share me your code  where create the picture1?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 550006,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-06-11T08:02:36.470000",
          "content": "<p>Each dot in that picture: (1) sample a certain number of earthquakes from train, (2) calculate KS statistic for each feature between sampled train and full test, (3) draw KS value on plot. Best sampling is based on lowest average KS statistic across features. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 550082,
          "author_name": "Xiang wang",
          "author_url": "",
          "post_date": "2019-06-11T09:31:03.150000",
          "content": "<p>ok,I understand.thank you for your reply.btw.Do you use scipy.stats.ks_2samp(train, test)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 550104,
          "author_name": "Xiang wang",
          "author_url": "",
          "post_date": "2019-06-11T09:55:09.493000",
          "content": "<p>Is  (ks statistic)/pvalue better than ks statistic?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 550455,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-06-11T16:34:34.187000",
          "content": "<p>Yeah, I use the scipy function. The p-value just gives you and indicator whether the null hypothesis of those two distributions being equal can be rejected. So in this case looking at the statistic or the p-value is kinda similar, we focused on the statistic here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 550726,
          "author_name": "Xiang wang",
          "author_url": "",
          "post_date": "2019-06-12T01:18:00.253000",
          "content": "<p>Thank you for your answer ! It gave me a lot of inspiration!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 549410,
      "author_name": "Naman Bhatt",
      "author_url": "",
      "post_date": "2019-06-10T16:06:47.763000",
      "content": "<p>Thanks for sharing this. Great analysis and a wonderful answer !!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 549117,
      "author_name": "Mumthas T.K.",
      "author_url": "",
      "post_date": "2019-06-10T10:06:14.170000",
      "content": "<p>i would like to get connected with dna sequence analysis</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 548897,
      "author_name": "Joel Hanson",
      "author_url": "",
      "post_date": "2019-06-10T04:31:28.400000",
      "content": "<p>Thank you for sharing! Congrats to The Zoo team!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 548333,
      "author_name": "Firat Gonen",
      "author_url": "",
      "post_date": "2019-06-09T07:52:53.823000",
      "content": "<p>Congratulations to the ZOO TEAM ! Outstanding achievement !</p>\n\n<p>Thank you Philip for the detailed explanations !</p>\n\n<p>Well deserved !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 547980,
      "author_name": "yasu",
      "author_url": "",
      "post_date": "2019-06-08T15:37:06.717000",
      "content": "<p>Congratulations !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 547907,
      "author_name": "Vasco Carvalho",
      "author_url": "",
      "post_date": "2019-06-08T13:52:27.663000",
      "content": "<p>Congratulations!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 546719,
      "author_name": "ConradYen",
      "author_url": "",
      "post_date": "2019-06-06T20:52:57.313000",
      "content": "<p>Amazon work !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 546132,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2019-06-06T09:32:13.450000",
      "content": "<p>So much great ideas and insights to study and use for future competitions. Brilliant solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 545936,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-06T04:05:27.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 545869,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-06T01:35:50.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 544274,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-05T10:47:19.990000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 544281,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-05T10:56:12.030000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 546100,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-06T08:56:01.803000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 553411,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-15T16:30:35.433000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 544182,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-05T08:42:21.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543903,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T23:32:44.137000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 545872,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-06T01:45:45.553000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543843,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T22:05:49.150000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543773,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T20:05:29.693000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 543776,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T20:09:20.637000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543601,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T16:11:36.010000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543559,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T15:40:30.820000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 553062,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-15T04:28:12.527000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543530,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T15:22:48.397000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 543543,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T15:28:51.460000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543630,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T16:39:08.890000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543640,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T16:49:23.337000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543645,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T16:54:03.480000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543498,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T15:08:43.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543497,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T15:08:33.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543422,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T14:21:49.260000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 543427,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:24:30.347000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543434,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:26:27.643000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543449,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:31:55.757000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543505,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T15:10:01.387000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543389,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T14:03:28.843000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 543410,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:14:55.533000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543436,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:27:31.077000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543461,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:38:36.740000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543473,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T14:45:40.563000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543373,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T13:56:52.467000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543301,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T13:01:47.377000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543195,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T11:21:28.337000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 543208,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T11:29:21.780000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 543222,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T11:46:34.610000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 543228,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T11:53:34.410000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543101,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T10:27:22.147000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543076,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T10:12:37.903000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 566308,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-02T03:21:48.270000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 544190,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-05T08:50:42.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 566281,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-02T02:01:43.217000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543330,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T13:27:50.940000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542985,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T08:52:24.783000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2576195,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-27T14:44:30.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 555604,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-19T06:22:59.173000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 548389,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-09T09:59:49.407000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 548371,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-09T09:35:53.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 546528,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-06T16:36:19.613000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 546124,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-06T09:21:06.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 544395,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-05T13:12:31.287000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543249,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T12:16:09.513000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543059,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T10:02:11.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543052,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T09:53:31.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543048,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T09:49:21.273000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542982": "Thanks a lot to the hosts of this competition and congratz to all participants and of course to my amazing teammates.\n\nWhat made this competition tricky was to find a proper CV setup that you believe in as the public LB gave bad feedback for private LB. This was my first competition where this was the case and it took me a while to completely ignore public LB, but it was necessary.\n\nI will now try to summarize some of the main points that helped us to win this competition. I am posting these elaborations in the we-form as we are a team and everyone contributed ideas and knowledge. Special thanks to @ilu000 @dott1718 @returnofsputnik @dkaraflos @pukkinming who worked hard the last few weeks on the comp.\n\n**Acoustic signal manipulation and features**\n\nAs has been discussed in the forums and shown by adversarial validation, the signal had a certain time-trend that caused some issues specifically on mean and quantile based features. To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating ```np.random.normal(0, 0.5, 150_000)```. Additionally, after noise addition, we subtracted the median of the segment. \n\nOur features are then calculated on this manipulated signal. We mostly focused on similar features as most participants in this competition, namely finding peaks and volatility of the signal. One of our best final LGB model only used four features: (i) number of peaks of at least support 2 on the denoised signal, (ii) 20% percentile on std of rolling window of size 50, (iii) 4th and (iv) 18th Mel-frequency cepstral coefficients mean. We sometimes used a few more features (like for the NN, see below) but they are usually very similar. Those 4 are decently uncorrelated between themselves, and add good diversity. For each feature we always only considered it if it has a p-value &gt;0.05 on a KS statistic of train vs test.\n\n**Differences between train and test features**\n\nAfter doing abovementioned signal manipulation, we had more trust in our calculated features and could focus on better studying differences between train and test data feature distributions. We found that the test data should look different to training data in a few ways when comparing features by e.g., applying KS statistics between train and test. That’s when we decided to sample the train data to make it look more like we expect test data to look like (only from looking at feature distributions). We started by manually upsampling certain areas of train data, but gave up on that after a few tries and then we found a very nice way of aligning train and test data.\n\nSo what we did is that we calculated a handful of features for train and test and tried to find a good subset of full earth-quakes in train, so that the overall feature distributions are similar to those of the full test data. We did this by sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test. A visualization for this looks like this (this is a limited visualization and not necessarily the one we chose to make our final selection of EQs):\n\n ![KS statistic train subsample vs. test](https://i.imgur.com/9evXuhV.png)\n\nThe x-axis is the average target of the selected EQs in train and the y-axis is the KS statistic on a bunch of features comparing the distribution of that feature for the selected EQs vs the full test data. We can see that the best average KS-statistic is somewhere in the range of 6.2-6.5. You can also see nicely here that a problematic feature like the green one deviates clearly from the rest, this would be a feature we would not select in the end.\n\nAfter careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. The mean of this sample is 6.258 and the median is 6.031.\n\n**CV**\n\nNow that we had sampled train data that we though to be similar to test just purely based on statistical analysis, and now that we had features that should not have any time leaks, we decided on doing a simple shuffled 3-fold on that data. Higher fold results are similar. We now tried to improve this CV as well as possible. \n\n**Models**\n\nOur final submit is a hillclimber blend of three types of models: (i) LGB, (ii) SVR, (iii) NN. The overall CV score on this was ~1.83. The LGB is using a fair loss with relatively moderate other hyperparameters. The SVR is also quite simply set-up. The NN is a bit more complicated with a few layers on top of a bunch of features. The real interesting thing here is that we do multi-task learning by specifying additional losses next to the ttf loss that we weight higher than the others. We have one additional binary logloss with the target specifying if the ttf is &lt;0.5 and one further MAE loss on the target of time-since-failure. This helped to balance some of the predictions out a bit and specifically helped to better predict some of the areas at the end of earthquakes that make some weird spikes. The NN had the best single MAE, but blending improved. Actually, just blending LGB and NN would have produced the best private LB score (2.25909). Adding SVR did improve CV though.\n\nWith all the steps described above, we also managed to make the distribution of test predictions very similar tho oof predictions. The following image shows for a single LGB the oof (blue) vs. test predictions (orange). The KS-test between those two does not reject the null hypothesis of them being equally distributed.\n\n![LGB oof (blue) vs. test (orange) prediction dist](https://i.imgur.com/bspTxd4.png)\n\n**Ideas that have potential**\n\nWe had quite some ideas that have potential but did not make it into our final submission. One area is to better use the time-since-failure prediction, which we used only as an additional loss in our NN. Modeling tsf works better than ttf. It can help to manually adjust a few predictions which have large discrepancies between tsf and ttf predictions, like the end of EQs. Also, they can be a reasonable proxy for predicting the approximate length of the EQ. So for example, we had one model that normalized the ttf targets to be in range 0-1 and then predicts this normalized target and scales it by ttf+tsf prediction. This was usually very close to our simpler models so we did not tune it extensively, I just feel that this has further potential. \n\nThe following kernel runs a LGB model on most of what I explained above and also would score 1st place with 2.279 private MAE:\n\nhttps://www.kaggle.com/ilu000/1-private-lb-kernel-lanl-lgbm/\n\nThe following kernel runs a blend between LGB and NN scoring 2.25993 on private LB:\n\nhttps://www.kaggle.com/dkaraflos/1-geomean-nn-and-6featlgbm-2-259-private-lb\n\n\n\n\n",
    "543397": "In a properly set scientific experiment the data would be **separated** into training and testing. In other words the testing data is locked in a box, the scientist goes to the box only after he/she has developed a hypothesis (a method/model etc). The testing data is then used only once to **test** this particular hypothesis. If the researcher is allowed to peek, analyze the testing data then the test loses its validity. Once the researcher starts tuning a hypothesis with respect to a data set then he/she needs another data set to test the reached conclusions.\n\nTLDR: The top ranking method has been tuned to the testing dataset via the KS test. This means that the method has become specific to this particular testing set. If the LANL researchers' goal is to apply the obtained findings to other experiments then it is very likely that they will be better of using a lower ranking method.\n\nThe competition/ranking has lost its practical value because the organizers made the testing data public. They should have kept the testing data on the server side and allow only for submission of executable code (see Matlab Cody competitions for instance), preventing the testing data being used for tuning.",
    "543360": "That's amazing solution. Big congrats to The Zoo team, well deserved.  \n\nMimic the test statistics in train, using the paper p4677 or using testset statistics was the key to survive to the shakeup. ",
    "564289": "Thanks a lot for sharing this solution and a heartily congrates too. ",
    "542992": "Thank you for the wonderful write-up @philippsinger\nIt was a pleasure to work with all of you and i will gladly team up with you again, @dott1718 @returnofsputnik @dkaraflos @pukkinming ",
    "543069": "Thanks for sharing, and congratulation on the result.  Resampling was the way to go indeed, I wish we had thought of it!",
    "543044": "Congratulation! \nBecause of the differences in distribution between train and test, I gave up the competition early.\nI never image that we can make a sub-train set which have the same distribution with test.\nThanks for your wonderful sharing.",
    "626214": "Hi,\n\nI know this gratitude comes late but I have to say thank you very much for sharing your awesome ideas! I am trying to reproduce your solution and I'm totally surprised by the fact that by selecting the earthquakes to align the training data and test data, my private score increases from 2.58 to 2.31, a top 10 result!\n\nI'm still digesting some of your ideas. But I want to ask a more general question: Is it always a good idea to align training set and test set? Thanks in advance!",
    "555526": "Congratulations, impressive!",
    "553228": "Congratulations!  Lots of amazing ideas !\nI have a question about:\n&gt; To partly overcome this, we added a constant noise to each 150k segment (both in train and test) by calculating np.random.normal(0, 0.5, 150_000). Additionally, after noise addition, we subtracted the median of the segment.\n\nWhy this preprocessing can help those time-trend features more useful.  Could you share the reasons behind this ?\nSorry if it is stupid question.",
    "548101": "Using the test set statistics to identify and give extra weight to similar events in the training set is very clever, congratulations. I'm a bit perplexed that given that there was data in the training set that were similar to the test set that the training with equal weight given to all the training data did not seem to result in an algorithm that could generalize well to both the public and private leaderboards. Do you have any thoughts about why this is so? Without weighting your training data, would your algorithm have achieved similar scores on both the public and private leaderboard?\n\nI wonder if your success mimicking the test set distribution points to a possible weakness in the training-testing split: imagine if there were 5 populations that the events were drawn from (A, B, C, D, E). A strong split would group items from (A, B, C, D) into the training set and leave items from E in the test set. But is it the case that the training set is items (A, B, C, D and a bit of E)? In the case of a strong split, there would have been little overlap with items in the training set and the test set, and there wouldn't have been an easy way weight your training data. ",
    "544143": "Sorry, I made a mistake in calculation.  I deleted my post.  My score should be 2.40.  Congrats for the amazing work!!!",
    "544042": "Thanks for sharing, very neat solution ! \n\nCould you please explain why adding noise to the data improves the mean and quantile based features ? ",
    "543858": "Wonderful solution! Thank you for sharing. By appreciating your great insight, I really want to do better next time. Good job!",
    "543027": "Congratulation for a very robust model,well deserved.",
    "543734": "Amazing solution and well deserved winning! Great job!",
    "553069": "Out of curiosity ...\n\nWill it be ever possible to have a mathematical formula for predicting time-to-earthquake (based on acoustic or other signals)?",
    "3277687": "Congratulations!\nI gave up on the competition earlier because of the distribution shift between training and testing data.\nI couldn’t imagine that a sub-train set could be built to align with the test distribution.\nThank you for your excellent explanation.",
    "567770": "Could you explain the theory of 'added noise to make median based features reliable.'? I am very curious about how this idea comes from.",
    "554034": "Amazing solution!",
    "553057": "Congratulations for winning this challenge of enormous significance.\nThank you for sharing your wisdom and your \"steps to the stardom\".\nThe way you have narrated will shine the subject of Statistics.",
    "549989": "i canot fully understand your analysis with your ks test result(ks test plot ) .Can you share me your code  where create the picture1?\n",
    "549410": "Thanks for sharing this. Great analysis and a wonderful answer !!",
    "549117": "i would like to get connected with dna sequence analysis",
    "548897": "Thank you for sharing! Congrats to The Zoo team!",
    "548333": "Congratulations to the ZOO TEAM ! Outstanding achievement !\n\nThank you Philip for the detailed explanations !\n\nWell deserved !",
    "547980": "Congratulations !",
    "547907": "Congratulations!!",
    "546719": "Amazon work !!!",
    "546132": "So much great ideas and insights to study and use for future competitions. Brilliant solution!",
    "545936": "This is beautiful. Congrats!",
    "545869": "Congrats! Thanks for sharing your solution!\nI learn a lot!\nespecially KS test part, very brilliant",
    "544274": "Thank you for sharing. What do you mean by  \"sampling 10 full earthquakes multiple times (up to 10k times) on train, and comparing the average KS statistic of all selected features on the sampled earthquakes to the feature dists in full test\"? Did you compute your features for each of the test files, then randomly pick 10 earthquakes in the train data, randomly sampled them 10k times such that each sample is the same length as a test file (150000), computed your features on each sample, then computed the KS stat for features computed for this set of samples vs features computed for all test files? Forgive me if it's a stupid question.",
    "544182": "Amazing solution!\nCommit all 0 value,Private Score----mae 6.66993,Public Score----mae 4.01773.\nSo solution not important,I think.\nPrivate score depend on the same x，but very different y in public and private test set.\nDoes the test data is really produced by experiment?\nWhy it is very different?\n![all 0 value,public and private mae](https://storage.googleapis.com/kaggle-forum-message-attachments/544182/13407/0.PNG)",
    "543903": "Congrats and thanks a lot for sharing! \nI think that I had that idea of sampling the train data looking at the test data. Basically, using the acoustic data feature, I tried to find in the train data the most similar 150k samples segment to each one of the test segments (using opencv and measuring similarity with R-Square). Then, in each of the segments I found, I created the features discussed in the public kernels. However it didn't work well and I am still wondering why...",
    "543843": "&gt;&gt; KS statistic of train vs test\n\nvery elegant! thank you for sharing",
    "543773": "Thank you and congrat.\nWhat was the best corrcoef value that you could get between your features and ttf?\nMy best feature resulted a corrcoef of 0.679",
    "543601": "Congrats! Thanks for sharing your solution. That's pretty amazing! ",
    "543559": "tweak training dataset based on testing dataset is a big lesson I learnt. i suppose we should never ever touch testing set, but forgot that this is a Kaggle competition... ",
    "543530": "Very well done! I tried something a little similar which I called AVMS or Adversarial Validation-Modulated Sampling. I assigned a weight to each row in train based on its AV score with a model that had an overal AUC of over 0.85 and used it to weight sampling of rows for training in each fold. This approach gave me the exact same LB score as when it wasn't used so I discarded it, but in hindsight it was a valuable discovery and I should have realised its importance. ",
    "543498": "Beautiful solution and very well explained. Congrats!!!",
    "543497": "Congrats! Thanks for sharing your solution!",
    "543422": "Thanks for sharing and congratulations to all The Zoo!!\nQuick question: what is the score of the second submission, if I may ask?\nAlologies if I overlooked\nThx",
    "543389": "Did you try denoise autoencoder?",
    "543373": "Congratulations !",
    "543301": "Thank you for sharing and congratulations. Getting 1st place without using any test ttf leak information from paper is great!\nI guess re-sampling method is the key, however it is not possible to check if this re-sampling affects better or not without looking LB. So you just trust your method works well until private LB opens, that is great.",
    "543195": "Excellent @philippsinger!\nQuick question(s):\n\n\"After careful examination of these results, we decided in the end to subsample the train data to only consider earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10] numerating all 17 earthquake cycles we have in train. \"\n\n1. So you threw other EQ away or you added additional samples that were representative of earthquakes [2, 7, 0, 4, 11, 13, 9, 1, 14, 10]? How big was you train data in the end anyways?\n\n2. Did you try using a model instead of KS statistic? For example using AV for individual features and only the ones that have low AUC score you would retain. (meaning they (should) have similiar distributions)",
    "543101": "Congratulations and thanks for sharing. Very neat solution.",
    "543076": "Congratulations! and thanks for sharing your great solution!!",
    "566308": "",
    "544190": "",
    "566281": "Thanks for sharing your wisdom.",
    "543330": "Congratulations and thanks for sharing!",
    "542985": "Congratulations! Thank you for sharing!",
    "2576195": "Thanks a lot for sharing this solution",
    "555604": "Nice work, Thanks for sharing ",
    "548389": "**Congrats and thanks for sharing**",
    "548371": "Congratulations! Thank you for sharing!",
    "546528": "Congratulations! Thanks for sharing :)",
    "546124": "Thank you for sharing! Congrats !",
    "544395": "Congratulations! Thanks for sharing!",
    "543249": "Thank's",
    "543059": "Congratulations! Thanks for sharing.",
    "543052": "Congratulations! Thanks for sharing.",
    "543048": "Congrats! Thanks for sharing :) "
  }
}