{
  "id": 94363,
  "title": "11th place solution",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94363",
  "author_name": "Tom Van de Wiele",
  "post_date": "2019-06-04T04:45:15.451000",
  "votes": 48,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Many congratulations to the winners, The Zoo, you truly stood out on the private leaderboard! Also congratulations to all teams that survived the private earthquake or learned something during the competition. @Keita111 (+4024 to claim a gold with 2 submissions) is a great example that we should sometimes focus less on the public leaderboard even though we all know we should :-). </p>\n\n<p>Special thanks go to all forum contributors. People like @CPMPml, <a href=\"/mykper\">@mykper</a>, <a href=\"/scirpus\">@scirpus</a>, @Abhishek and several others are what make Kaggle great and I am deeply grateful for your contributions.</p>\n\n<p>Let’s start with what didn’t end up in the final submission. My efforts during this competition were mostly wasted on modeling the gaps (every 4095/4096 observations) in the data. The assumption was that if one could detect these gaps, it would be possible to order the test chunks, making the test prediction problem trivial. This model works perfectly on the validation data but it doesn't on the test data. I am certain that the test data is not contiguous as stated and is actually divided into 150,000 observation chunks with random gaps in between. It is a bit sad that the organizers ignored many good questions about the data distribution in the “Additional info” topic, requiring the competitors to make many assumptions. \nSide note: the gap analysis identified test chunks 1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c, d1eee8 as the start of an earthquake cycle, adding another piece of information to believe that the test data really comes from the test part of P4677.</p>\n\n<p>The second biggest chunk of my time was spent on unsupervised learning of features using the raw high frequency data with recurrent networks. The code was inspired by <a href=\"https://github.com/davidtellez/contrastive-predictive-coding\">David Tellez's</a> implementation of <a href=\"https://arxiv.org/pdf/1807.03748.pdf\">Contrastive Predictive Coding</a>. The hope was that this would lead to robust features for training and testing. I embedded each chunk of 1,500 observations to a shared embedding of size 32 and defined the following loss heads:\n- Predict if one embedding directly follows (with some small random gap) another embedding\n- Predict if a pair of embeddings comes from the same data chunk of 150,000 observations\n- Predict if an embedding comes from the train or test data\n- Predict if a pair of embeddings comes from the same earthquake\n- Predict Time To Failure (TTF)</p>\n\n<p>The first three losses are trained for both the train and test data of which the third one is trained using <a href=\"https://arxiv.org/pdf/1505.07818.pdf\">domain adversarial training</a>. The last two losses can only be trained for the train data. Even though I ended up not using any of these models, it became clear that it was very hard to distinguish if raw observations come from the same earthquake. That led me to redefine the learning objective in the final approach.</p>\n\n<p><strong>Actual submission</strong>\nAn alternative way to specify the learning objective is the following: predict the quantile TTF of an earthquake, meaning that the target is 0 at the beginning of an earthquake cycle and 1 for the last 150,000 observation chunk of each earthquake. Assuming that this model is trained to minimize the MAE of the quantile, what should each (1-prediction) be multiplied with to minimize the <strong>MAE</strong>? The median of [the earthquake lengths, repeated by the earthquake lengths]! For example, if the test data consisted of earthquake lengths 6, 3 and 2 this would result in median(6, 6, 6, 6, 6, 6, 3, 3, 3, 2, 2) = 6 * (1-quantile_pred) as the optimal predicted time to failure. The nice thing about this reformulation is that we can use the data from the P4677 experiment (I estimated the test median earthquake TTF to be 12) without having to break your head over what earthquake cycles to train on.</p>\n\n<p>The models themselves use basic FFT and quantile features for each 150,000 observation chunk and 100 1,500 observation subchunks. This resulted in 40*101 = 4040 features which were fed into LightGBM (feature fraction of 0.003) and neural networks. The first final submission averages the LightGBM and neural network predictions starting from the last 6 complete earthquake cycles. The second submission uses all data except for the first incomplete earthquake cycle (given that the cycle length is unknown we can’t determine its quantile). A nice property of this approach is that the predictions don’t change much when training on different folds so in hindsight it would have been better to submit with different estimated median private test earthquake times.</p>\n\n<p>I look forward to your feedback!</p>",
  "messages": [
    {
      "id": 542768,
      "postDate": "2019-06-04T04:45:15.453Z",
      "content": "<p>Many congratulations to the winners, The Zoo, you truly stood out on the private leaderboard! Also congratulations to all teams that survived the private earthquake or learned something during the competition. @Keita111 (+4024 to claim a gold with 2 submissions) is a great example that we should sometimes focus less on the public leaderboard even though we all know we should :-). </p>\n\n<p>Special thanks go to all forum contributors. People like @CPMPml, <a href=\"/mykper\">@mykper</a>, <a href=\"/scirpus\">@scirpus</a>, @Abhishek and several others are what make Kaggle great and I am deeply grateful for your contributions.</p>\n\n<p>Let’s start with what didn’t end up in the final submission. My efforts during this competition were mostly wasted on modeling the gaps (every 4095/4096 observations) in the data. The assumption was that if one could detect these gaps, it would be possible to order the test chunks, making the test prediction problem trivial. This model works perfectly on the validation data but it doesn't on the test data. I am certain that the test data is not contiguous as stated and is actually divided into 150,000 observation chunks with random gaps in between. It is a bit sad that the organizers ignored many good questions about the data distribution in the “Additional info” topic, requiring the competitors to make many assumptions. \nSide note: the gap analysis identified test chunks 1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c, d1eee8 as the start of an earthquake cycle, adding another piece of information to believe that the test data really comes from the test part of P4677.</p>\n\n<p>The second biggest chunk of my time was spent on unsupervised learning of features using the raw high frequency data with recurrent networks. The code was inspired by <a href=\"https://github.com/davidtellez/contrastive-predictive-coding\">David Tellez's</a> implementation of <a href=\"https://arxiv.org/pdf/1807.03748.pdf\">Contrastive Predictive Coding</a>. The hope was that this would lead to robust features for training and testing. I embedded each chunk of 1,500 observations to a shared embedding of size 32 and defined the following loss heads:\n- Predict if one embedding directly follows (with some small random gap) another embedding\n- Predict if a pair of embeddings comes from the same data chunk of 150,000 observations\n- Predict if an embedding comes from the train or test data\n- Predict if a pair of embeddings comes from the same earthquake\n- Predict Time To Failure (TTF)</p>\n\n<p>The first three losses are trained for both the train and test data of which the third one is trained using <a href=\"https://arxiv.org/pdf/1505.07818.pdf\">domain adversarial training</a>. The last two losses can only be trained for the train data. Even though I ended up not using any of these models, it became clear that it was very hard to distinguish if raw observations come from the same earthquake. That led me to redefine the learning objective in the final approach.</p>\n\n<p><strong>Actual submission</strong>\nAn alternative way to specify the learning objective is the following: predict the quantile TTF of an earthquake, meaning that the target is 0 at the beginning of an earthquake cycle and 1 for the last 150,000 observation chunk of each earthquake. Assuming that this model is trained to minimize the MAE of the quantile, what should each (1-prediction) be multiplied with to minimize the <strong>MAE</strong>? The median of [the earthquake lengths, repeated by the earthquake lengths]! For example, if the test data consisted of earthquake lengths 6, 3 and 2 this would result in median(6, 6, 6, 6, 6, 6, 3, 3, 3, 2, 2) = 6 * (1-quantile_pred) as the optimal predicted time to failure. The nice thing about this reformulation is that we can use the data from the P4677 experiment (I estimated the test median earthquake TTF to be 12) without having to break your head over what earthquake cycles to train on.</p>\n\n<p>The models themselves use basic FFT and quantile features for each 150,000 observation chunk and 100 1,500 observation subchunks. This resulted in 40*101 = 4040 features which were fed into LightGBM (feature fraction of 0.003) and neural networks. The first final submission averages the LightGBM and neural network predictions starting from the last 6 complete earthquake cycles. The second submission uses all data except for the first incomplete earthquake cycle (given that the cycle length is unknown we can’t determine its quantile). A nice property of this approach is that the predictions don’t change much when training on different folds so in hindsight it would have been better to submit with different estimated median private test earthquake times.</p>\n\n<p>I look forward to your feedback!</p>",
      "rawMarkdown": "Many congratulations to the winners, The Zoo, you truly stood out on the private leaderboard! Also congratulations to all teams that survived the private earthquake or learned something during the competition. @Keita111 (+4024 to claim a gold with 2 submissions) is a great example that we should sometimes focus less on the public leaderboard even though we all know we should :-). \n\nSpecial thanks go to all forum contributors. People like @CPMPml, @mykper, @scirpus, @Abhishek and several others are what make Kaggle great and I am deeply grateful for your contributions.\n\nLet’s start with what didn’t end up in the final submission. My efforts during this competition were mostly wasted on modeling the gaps (every 4095/4096 observations) in the data. The assumption was that if one could detect these gaps, it would be possible to order the test chunks, making the test prediction problem trivial. This model works perfectly on the validation data but it doesn't on the test data. I am certain that the test data is not contiguous as stated and is actually divided into 150,000 observation chunks with random gaps in between. It is a bit sad that the organizers ignored many good questions about the data distribution in the “Additional info” topic, requiring the competitors to make many assumptions. \nSide note: the gap analysis identified test chunks 1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c, d1eee8 as the start of an earthquake cycle, adding another piece of information to believe that the test data really comes from the test part of P4677.\n\nThe second biggest chunk of my time was spent on unsupervised learning of features using the raw high frequency data with recurrent networks. The code was inspired by [David Tellez's](https://github.com/davidtellez/contrastive-predictive-coding) implementation of [Contrastive Predictive Coding](https://arxiv.org/pdf/1807.03748.pdf). The hope was that this would lead to robust features for training and testing. I embedded each chunk of 1,500 observations to a shared embedding of size 32 and defined the following loss heads:\n- Predict if one embedding directly follows (with some small random gap) another embedding\n- Predict if a pair of embeddings comes from the same data chunk of 150,000 observations\n- Predict if an embedding comes from the train or test data\n- Predict if a pair of embeddings comes from the same earthquake\n- Predict Time To Failure (TTF)\n\nThe first three losses are trained for both the train and test data of which the third one is trained using [domain adversarial training](https://arxiv.org/pdf/1505.07818.pdf). The last two losses can only be trained for the train data. Even though I ended up not using any of these models, it became clear that it was very hard to distinguish if raw observations come from the same earthquake. That led me to redefine the learning objective in the final approach.\n\n\n**Actual submission**\nAn alternative way to specify the learning objective is the following: predict the quantile TTF of an earthquake, meaning that the target is 0 at the beginning of an earthquake cycle and 1 for the last 150,000 observation chunk of each earthquake. Assuming that this model is trained to minimize the MAE of the quantile, what should each (1-prediction) be multiplied with to minimize the **MAE**? The median of [the earthquake lengths, repeated by the earthquake lengths]! For example, if the test data consisted of earthquake lengths 6, 3 and 2 this would result in median(6, 6, 6, 6, 6, 6, 3, 3, 3, 2, 2) = 6 * (1-quantile_pred) as the optimal predicted time to failure. The nice thing about this reformulation is that we can use the data from the P4677 experiment (I estimated the test median earthquake TTF to be 12) without having to break your head over what earthquake cycles to train on.\n\nThe models themselves use basic FFT and quantile features for each 150,000 observation chunk and 100 1,500 observation subchunks. This resulted in 40*101 = 4040 features which were fed into LightGBM (feature fraction of 0.003) and neural networks. The first final submission averages the LightGBM and neural network predictions starting from the last 6 complete earthquake cycles. The second submission uses all data except for the first incomplete earthquake cycle (given that the cycle length is unknown we can’t determine its quantile). A nice property of this approach is that the predictions don’t change much when training on different folds so in hindsight it would have been better to submit with different estimated median private test earthquake times.\n\nI look forward to your feedback!",
      "votes": 47
    },
    {
      "id": 548737,
      "postDate": "2019-06-09T20:02:35.787Z",
      "content": "<p>Thanks for sharing! Great performance, especially considering the fact the private dataset was way different from the public one.</p>",
      "rawMarkdown": "Thanks for sharing! Great performance, especially considering the fact the private dataset was way different from the public one.",
      "votes": 1
    },
    {
      "id": 546196,
      "postDate": "2019-06-06T10:47:06.417Z",
      "content": "<p>Thank you for sharing! I think your idea is really fresh, and I'm pleasantly surprised that a model  'smudging' the predictions worked better than most models that tried to predict TTF directly. Just confused about some details - how did you assign quantile TTF value for intermediate points in the earthquake (not at the start, not at the last 150000 time points)? Also, why did you use the median of [earthquake length, repeated earthquake_length times], instead of simply the mean of earthquake lengths?</p>",
      "rawMarkdown": "Thank you for sharing! I think your idea is really fresh, and I'm pleasantly surprised that a model  'smudging' the predictions worked better than most models that tried to predict TTF directly. Just confused about some details - how did you assign quantile TTF value for intermediate points in the earthquake (not at the start, not at the last 150000 time points)? Also, why did you use the median of [earthquake length, repeated earthquake_length times], instead of simply the mean of earthquake lengths?",
      "votes": 1,
      "replies": [
        {
          "id": 546547,
          "postDate": "2019-06-06T17:15:39.717Z",
          "content": "<p>Thanks for the feedback. The quantile TTF values are simply computed by dividing by the earthquake cycle length.</p>\n\n<p>Your second question is quite interesting and I am surprised nobody else seems to have used this post-processing approach. Let's imagine that the quantile method is perfect and let's keep the example with three cycles of lengths 8, 4 and 3 in mind. What should the prediction then be if there is no way to tell what cycle a chunk of 150K observations belongs to? Each cycle would have an average absolute error of |cycle_length - 2*prediction_mean|/2 - The error would be cycle_length - 2*prediction_mean at the beginning of a cycle and decrease linearly to 0 by the end of a cycle. The average absolute error (MAE) would be Sum[P(cycle_i)*|cycle_length_i - 2*prediction_mean|/2]. This error is minimized when you set the prediction mean to median of [earthquake length, repeated earthquake_length times]/2. In the example that would lead to an MAE of 8/15*|8-8|/2 + 4/15*|4-8|/2 + 3/15*|3-8|/2 = 31/30 = 1.0333... If we set the prediction mean of the rescaled predictions to the true prediction mean the MAE is worse! The prediction mean of the 8, 4, 3 example is (8*8+4*4+3*3)/(2*15) = 89/30 - MAE of 8/15*|8-89/15|/2 + 4/15*|4-89/15|/2 + 3/15*|3-89/15|/2 = 1.10222... You can see why this is worse by realizing that if you change your mean prediction to a value less than 8, it is worse for more than half of the predictions (8/15) by the same value as it is better for 7/15 of the predictions. This optimal transformation still holds in expectation if the model predictions are unbiased estimates of the median quantile. Had there been one/two very long private test cycles, it would have made a bigger difference.</p>",
          "rawMarkdown": "Thanks for the feedback. The quantile TTF values are simply computed by dividing by the earthquake cycle length.\n\nYour second question is quite interesting and I am surprised nobody else seems to have used this post-processing approach. Let's imagine that the quantile method is perfect and let's keep the example with three cycles of lengths 8, 4 and 3 in mind. What should the prediction then be if there is no way to tell what cycle a chunk of 150K observations belongs to? Each cycle would have an average absolute error of |cycle\\_length - 2\\*prediction\\_mean|/2 - The error would be cycle\\_length - 2\\*prediction\\_mean at the beginning of a cycle and decrease linearly to 0 by the end of a cycle. The average absolute error (MAE) would be Sum[P(cycle\\_i)\\*|cycle\\_length\\_i - 2\\*prediction\\_mean|/2]. This error is minimized when you set the prediction mean to median of [earthquake length, repeated earthquake\\_length times]/2. In the example that would lead to an MAE of 8/15\\*|8-8|/2 + 4/15\\*|4-8|/2 + 3/15\\*|3-8|/2 = 31/30 = 1.0333... If we set the prediction mean of the rescaled predictions to the true prediction mean the MAE is worse! The prediction mean of the 8, 4, 3 example is (8\\*8+4\\*4+3\\*3)/(2*15) = 89/30 - MAE of 8/15\\*|8-89/15|/2 + 4/15\\*|4-89/15|/2 + 3/15\\*|3-89/15|/2 = 1.10222... You can see why this is worse by realizing that if you change your mean prediction to a value less than 8, it is worse for more than half of the predictions (8/15) by the same value as it is better for 7/15 of the predictions. This optimal transformation still holds in expectation if the model predictions are unbiased estimates of the median quantile. Had there been one/two very long private test cycles, it would have made a bigger difference.",
          "votes": 1
        },
        {
          "id": 547112,
          "postDate": "2019-06-07T09:50:58.287Z",
          "content": "<p>Thank you for your detailed reply!  So to generate training data for the model that would predict quantile TTFs, you divided each earthquake into quantiles, with each quantile containing 150k TTF data points, with the target being quantile index/number of quantiles?</p>\n\n<p>Just want to clarify what you mean by cycle length - would an earthquake which starts with a TTF of 12s have a cycle length of 12?</p>\n\n<p>Secondly, I don't quite understand why \"Each cycle would have an average absolute error of |cycle_length - 2*prediction_mean|/2 - The error would be cycle_length - 2*prediction_mean at the beginning of a cycle and decrease linearly to 0 by the end of a cycle\". </p>\n\n<p>Does prediction_mean mean the singular value you multiply with the quantile predictions to get the final TTF prediction (the median of [EQ_len, repeated EQ_len_times] in your case, if I understand correctly)? \nIf so, then shouldn't the error be, for a particular 150k section of an EQ, \n- TTF at end of section - prediction_mean, at the very start of the EQ;\n- 0, if TTF at end of section == prediction_mean*quantile prediction;\n- prediction_mean*quantile prediction, if the TTF at the end of the section is 0</p>\n\n<p>I think I need to make sure I understand these correctly before I try and work through the rest of your explanation.</p>\n\n<p>Thank you so much for your patience - I'm probably horrendously wrong, I need your help!</p>",
          "rawMarkdown": "Thank you for your detailed reply!  So to generate training data for the model that would predict quantile TTFs, you divided each earthquake into quantiles, with each quantile containing 150k TTF data points, with the target being quantile index/number of quantiles?\n\nJust want to clarify what you mean by cycle length - would an earthquake which starts with a TTF of 12s have a cycle length of 12?\n\nSecondly, I don't quite understand why \"Each cycle would have an average absolute error of |cycle_length - 2*prediction_mean|/2 - The error would be cycle_length - 2*prediction_mean at the beginning of a cycle and decrease linearly to 0 by the end of a cycle\". \n\nDoes prediction_mean mean the singular value you multiply with the quantile predictions to get the final TTF prediction (the median of [EQ_len, repeated EQ_len_times] in your case, if I understand correctly)? \nIf so, then shouldn't the error be, for a particular 150k section of an EQ, \n- TTF at end of section - prediction_mean, at the very start of the EQ;\n- 0, if TTF at end of section == prediction_mean*quantile prediction;\n- prediction_mean*quantile prediction, if the TTF at the end of the section is 0\n\nI think I need to make sure I understand these correctly before I try and work through the rest of your explanation.\n\nThank you so much for your patience - I'm probably horrendously wrong, I need your help!"
        },
        {
          "id": 547303,
          "postDate": "2019-06-07T14:36:38.113Z",
          "content": "<p>To be precise, the quantile targets were set to (number of observations at the end of each 150K until a TTF reset )/(number of observations in that cycle-149999). That way you have training targets that decrease linearly from 1 to 0. I also made sure not to introduce training chunks that overlapped the TTF reset.</p>\n\n<p>A cycle length is indeed the duration of the earthquake until a failure (reset of TTF)</p>\n\n<p>Does prediction_mean mean the scalar you multiply with the quantile predictions to get the final TTF prediction - Yes, to be precise, the prediction is (1-quantile_prediction) * C.</p>\n\n<p>The explanation assumes that the quantile model is perfect, so the error for quantile q is: (1-q)*|cycle_length - 2*prediction_mean|.</p>",
          "rawMarkdown": "To be precise, the quantile targets were set to (number of observations at the end of each 150K until a TTF reset )/(number of observations in that cycle-149999). That way you have training targets that decrease linearly from 1 to 0. I also made sure not to introduce training chunks that overlapped the TTF reset.\n\nA cycle length is indeed the duration of the earthquake until a failure (reset of TTF)\n\nDoes prediction\\_mean mean the scalar you multiply with the quantile predictions to get the final TTF prediction - Yes, to be precise, the prediction is (1-quantile\\_prediction) * C.\n\nThe explanation assumes that the quantile model is perfect, so the error for quantile q is: (1-q)\\*|cycle\\_length - 2*prediction\\_mean|.",
          "votes": 1
        },
        {
          "id": 547671,
          "postDate": "2019-06-08T05:17:05.613Z",
          "content": "<p>Got it.\nIf prediction_mean is C in your equation (1-q) * C, then the predicted TTF is (1-q)*C. The true TTF is (1-q)*cycle_length. Shouldn't the absolute error then be (1-q)*|cycle_length-C|? You use this for your calculations with the [8, 4, 3] example.\nI get the explanation as a whole now, though. I understand it as kind of minimizing the 'expected loss', given you don't know what cycle length the 150k sequence belongs to.  I like how you elegenatly multiplied the probability that a sequence belongs to a particular cycle length to the individual average absolute errors. I guess this has the assumption that the distribution of cycle lengths in test data not too far off from that of the training data, although that's essential for a problem to be machine learnable. \nThank you again for sharing!</p>",
          "rawMarkdown": "Got it.\nIf prediction_mean is C in your equation (1-q) * C, then the predicted TTF is (1-q)*C. The true TTF is (1-q)*cycle_length. Shouldn't the absolute error then be (1-q)*|cycle_length-C|? You use this for your calculations with the [8, 4, 3] example.\nI get the explanation as a whole now, though. I understand it as kind of minimizing the 'expected loss', given you don't know what cycle length the 150k sequence belongs to.  I like how you elegenatly multiplied the probability that a sequence belongs to a particular cycle length to the individual average absolute errors. I guess this has the assumption that the distribution of cycle lengths in test data not too far off from that of the training data, although that's essential for a problem to be machine learnable. \nThank you again for sharing!"
        },
        {
          "id": 547735,
          "postDate": "2019-06-08T07:37:32.777Z",
          "content": "<p>C should be twice the prediction_mean. I have adapted my [8, 4, 3] example since I made a \"divide by two\" mistake to compute the prediction mean which canceled out in the next step. Thank you for making me aware! The conclusion remains unchanged.</p>",
          "rawMarkdown": "C should be twice the prediction\\_mean. I have adapted my [8, 4, 3] example since I made a \"divide by two\" mistake to compute the prediction mean which canceled out in the next step. Thank you for making me aware! The conclusion remains unchanged.",
          "votes": 1
        }
      ]
    },
    {
      "id": 543503,
      "postDate": "2019-06-04T15:09:52.160Z",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "rawMarkdown": "Congrats! Thanks for sharing your solution!",
      "votes": 1
    },
    {
      "id": 543247,
      "postDate": "2019-06-04T12:13:05.897Z",
      "content": "<p>Congratulations and thank you for sharing solution!\nYour idea of changing the label to predict quantile of ttf is a really smart approach and I'm impressed.</p>\n\n<p>I also noticed that it is not possible to classify different quake, meaning it is not possible to predict whether this is long or short quake data. However I never come up with changing the problem formulation to predict quantile ttf value.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing solution!\nYour idea of changing the label to predict quantile of ttf is a really smart approach and I'm impressed.\n\nI also noticed that it is not possible to classify different quake, meaning it is not possible to predict whether this is long or short quake data. However I never come up with changing the problem formulation to predict quantile ttf value.",
      "votes": 1
    },
    {
      "id": 543143,
      "postDate": "2019-06-04T10:51:40.850Z",
      "content": "<p>Great approach. Congrats on your gold and thanks for sharing.</p>",
      "rawMarkdown": "Great approach. Congrats on your gold and thanks for sharing.",
      "votes": 1
    },
    {
      "id": 543086,
      "postDate": "2019-06-04T10:18:04.210Z",
      "content": "<p>Congrats on your gold, very interesting approach!</p>",
      "rawMarkdown": "Congrats on your gold, very interesting approach!",
      "votes": 1,
      "replies": [
        {
          "id": 543121,
          "postDate": "2019-06-04T10:41:45.317Z",
          "content": "<p>Thank you! It seems promising to combine this learning objective with superior feature engineering and selection demonstrated by your team and others.</p>",
          "rawMarkdown": "Thank you! It seems promising to combine this learning objective with superior feature engineering and selection demonstrated by your team and others."
        }
      ]
    },
    {
      "id": 542965,
      "postDate": "2019-06-04T08:39:19.110Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 542927,
      "postDate": "2019-06-04T07:56:38.440Z",
      "content": "<p>Thanks for sharing !  There is one question. \nWas there a particular reason for choosing that model? (Simple hypothesis, Better CV score, etc.)</p>\n\n<p>In my case, Changing the target into quantile target  makes cv score worse.</p>\n\n<p>Congrats :)</p>",
      "rawMarkdown": "Thanks for sharing !  There is one question. \nWas there a particular reason for choosing that model? (Simple hypothesis, Better CV score, etc.)\n\nIn my case, Changing the target into quantile target  makes cv score worse.\n\nCongrats :)",
      "votes": 1,
      "replies": [
        {
          "id": 542929,
          "postDate": "2019-06-04T07:57:48.873Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 542944,
          "postDate": "2019-06-04T08:09:11.277Z",
          "content": "<p>Using the quantile model was a lot better (CV of about 0.1) than modeling TTF directly in local validation, assuming that both models have access to the validation mean and median repeated cycle lengths. I compared the quantile approach to direct models with different normalization schemes to match the validation average/median statistics.</p>",
          "rawMarkdown": "Using the quantile model was a lot better (CV of about 0.1) than modeling TTF directly in local validation, assuming that both models have access to the validation mean and median repeated cycle lengths. I compared the quantile approach to direct models with different normalization schemes to match the validation average/median statistics.",
          "votes": 1
        },
        {
          "id": 542952,
          "postDate": "2019-06-04T08:18:32.787Z",
          "content": "<p>Great! Thanks.</p>",
          "rawMarkdown": "Great! Thanks."
        }
      ]
    },
    {
      "id": 542916,
      "postDate": "2019-06-04T07:32:05.100Z",
      "content": "<p>wow this is a really great approach</p>",
      "rawMarkdown": "wow this is a really great approach",
      "votes": 1
    },
    {
      "id": 542896,
      "postDate": "2019-06-04T07:09:08.527Z",
      "content": "<p>Congratulations and thank you for sharing! Can I ask how did you decide final submissions? I suppose there were lots of choices.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing! Can I ask how did you decide final submissions? I suppose there were lots of choices.",
      "votes": 1,
      "replies": [
        {
          "id": 542908,
          "postDate": "2019-06-04T07:17:24.533Z",
          "content": "<p>I used local validation to select the hyperparameters of the LightGBM and Neural net and averaged 10 random seeds/initializations for models trained on all or the last 6 complete training earthquake cycles. Local validation was split by time - models trained up to the kth earthquake were evaluated on later earthquakes. </p>",
          "rawMarkdown": "I used local validation to select the hyperparameters of the LightGBM and Neural net and averaged 10 random seeds/initializations for models trained on all or the last 6 complete training earthquake cycles. Local validation was split by time - models trained up to the kth earthquake were evaluated on later earthquakes. ",
          "votes": 1
        },
        {
          "id": 542923,
          "postDate": "2019-06-04T07:50:00.410Z",
          "content": "<p>I got it, thank you so much!</p>",
          "rawMarkdown": "I got it, thank you so much!"
        }
      ]
    },
    {
      "id": 542782,
      "postDate": "2019-06-04T04:56:27.927Z",
      "content": "<p>Congrats and thanks for sharing <a href=\"/tvdwiele\">@tvdwiele</a> !</p>",
      "rawMarkdown": "Congrats and thanks for sharing @tvdwiele !",
      "votes": 1
    },
    {
      "id": 543168,
      "postDate": "2019-06-04T11:07:10.993Z",
      "content": "<p>your approach is really cool, especially the unsupervised learning part. Is there any code for it? Also, for the quantile part, that was very smart. i think i remember in another competition, the 1st place instead of predicting pure regression, he predicted [binary 0 or 1, are you an outlier?] * [max value of regression]. I think the way you approached it is similar, and this technique seems powerful in regression. Congratulations on your gold medal</p>",
      "rawMarkdown": "your approach is really cool, especially the unsupervised learning part. Is there any code for it? Also, for the quantile part, that was very smart. i think i remember in another competition, the 1st place instead of predicting pure regression, he predicted [binary 0 or 1, are you an outlier?] * [max value of regression]. I think the way you approached it is similar, and this technique seems powerful in regression. Congratulations on your gold medal",
      "votes": 2,
      "replies": [
        {
          "id": 543215,
          "postDate": "2019-06-04T11:38:23.843Z",
          "content": "<p>Thanks for the feedback! I intend to put the code on GitHub but there is some cleaning up to do first. I ended up not using the CPC features because adding a head that predicts the MAE resulted in leakage, which I realized too close to the deadline. \nAs a starting point I would point to <a href=\"https://github.com/davidtellez/contrastive-predictive-coding\">this excellent resource for learning unsupervised features - CPC by David Tellez</a>.</p>",
          "rawMarkdown": "Thanks for the feedback! I intend to put the code on GitHub but there is some cleaning up to do first. I ended up not using the CPC features because adding a head that predicts the MAE resulted in leakage, which I realized too close to the deadline. \nAs a starting point I would point to [this excellent resource for learning unsupervised features - CPC by David Tellez](https://github.com/davidtellez/contrastive-predictive-coding).",
          "votes": 2
        },
        {
          "id": 547367,
          "postDate": "2019-06-07T16:18:01.303Z",
          "content": "<p>Hello again, all code has now been published on <a href=\"https://github.com/ttvand/Kaggle-LANL\">Github</a>. The unsupervised part can be found <a href=\"https://github.com/ttvand/Kaggle-LANL/blob/master/Logic/train_valid_test_cpc.py\">here</a>. Let me know if you have any questions.</p>",
          "rawMarkdown": "Hello again, all code has now been published on [Github](https://github.com/ttvand/Kaggle-LANL). The unsupervised part can be found [here](https://github.com/ttvand/Kaggle-LANL/blob/master/Logic/train_valid_test_cpc.py). Let me know if you have any questions.",
          "votes": 1
        },
        {
          "id": 547804,
          "postDate": "2019-06-08T09:54:00.603Z",
          "content": "<p>Wow, thank you for that! </p>\n\n<p>Just to supplement:\n<a href=\"https://github.com/pumpikano/tf-dann\">Domain-Adversarial Training of Neural Networks in Tensorflow</a></p>\n\n<p><a href=\"https://github.com/CuthbertCai/pytorch_DANN\">pytorch dann</a></p>\n\n<p><a href=\"https://github.com/ataakbari/DANN\">DANN keras</a></p>",
          "rawMarkdown": "Wow, thank you for that! \n\nJust to supplement:\n[Domain-Adversarial Training of Neural Networks in Tensorflow](https://github.com/pumpikano/tf-dann)\n\n[pytorch dann](https://github.com/CuthbertCai/pytorch_DANN)\n\n[DANN keras](https://github.com/ataakbari/DANN)",
          "votes": 1
        }
      ]
    },
    {
      "id": 543129,
      "postDate": "2019-06-04T10:44:42.340Z",
      "content": "<p>I'm a newbie in kaggle world. I learned a lot by participating in this competition. Thanks to all kagglers that have shared their greats kernels and warm congratulations to the winners.  </p>",
      "rawMarkdown": "I'm a newbie in kaggle world. I learned a lot by participating in this competition. Thanks to all kagglers that have shared their greats kernels and warm congratulations to the winners.  ",
      "votes": 2
    },
    {
      "id": 543040,
      "postDate": "2019-06-04T09:43:55.537Z",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats! Thanks for sharing :)",
      "votes": 1
    },
    {
      "id": 542848,
      "postDate": "2019-06-04T06:31:03.213Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 1
    },
    {
      "id": 542774,
      "postDate": "2019-06-04T04:48:44.397Z",
      "content": "<p>Thank you <a href=\"/tvdwiele\">@tvdwiele</a> !</p>\n\n<h1>👍 🥇</h1>",
      "rawMarkdown": "Thank you @tvdwiele !\n# 👍 🥇 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 548737,
      "author_name": "Alexis Laignelet",
      "author_url": "",
      "post_date": "2019-06-09T20:02:35.787000",
      "content": "<p>Thanks for sharing! Great performance, especially considering the fact the private dataset was way different from the public one.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 546196,
      "author_name": "June Kim",
      "author_url": "",
      "post_date": "2019-06-06T10:47:06.417000",
      "content": "<p>Thank you for sharing! I think your idea is really fresh, and I'm pleasantly surprised that a model  'smudging' the predictions worked better than most models that tried to predict TTF directly. Just confused about some details - how did you assign quantile TTF value for intermediate points in the earthquake (not at the start, not at the last 150000 time points)? Also, why did you use the median of [earthquake length, repeated earthquake_length times], instead of simply the mean of earthquake lengths?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 546547,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-06T17:15:39.717000",
          "content": "<p>Thanks for the feedback. The quantile TTF values are simply computed by dividing by the earthquake cycle length.</p>\n\n<p>Your second question is quite interesting and I am surprised nobody else seems to have used this post-processing approach. Let's imagine that the quantile method is perfect and let's keep the example with three cycles of lengths 8, 4 and 3 in mind. What should the prediction then be if there is no way to tell what cycle a chunk of 150K observations belongs to? Each cycle would have an average absolute error of |cycle_length - 2*prediction_mean|/2 - The error would be cycle_length - 2*prediction_mean at the beginning of a cycle and decrease linearly to 0 by the end of a cycle. The average absolute error (MAE) would be Sum[P(cycle_i)*|cycle_length_i - 2*prediction_mean|/2]. This error is minimized when you set the prediction mean to median of [earthquake length, repeated earthquake_length times]/2. In the example that would lead to an MAE of 8/15*|8-8|/2 + 4/15*|4-8|/2 + 3/15*|3-8|/2 = 31/30 = 1.0333... If we set the prediction mean of the rescaled predictions to the true prediction mean the MAE is worse! The prediction mean of the 8, 4, 3 example is (8*8+4*4+3*3)/(2*15) = 89/30 - MAE of 8/15*|8-89/15|/2 + 4/15*|4-89/15|/2 + 3/15*|3-89/15|/2 = 1.10222... You can see why this is worse by realizing that if you change your mean prediction to a value less than 8, it is worse for more than half of the predictions (8/15) by the same value as it is better for 7/15 of the predictions. This optimal transformation still holds in expectation if the model predictions are unbiased estimates of the median quantile. Had there been one/two very long private test cycles, it would have made a bigger difference.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 547112,
          "author_name": "June Kim",
          "author_url": "",
          "post_date": "2019-06-07T09:50:58.287000",
          "content": "<p>Thank you for your detailed reply!  So to generate training data for the model that would predict quantile TTFs, you divided each earthquake into quantiles, with each quantile containing 150k TTF data points, with the target being quantile index/number of quantiles?</p>\n\n<p>Just want to clarify what you mean by cycle length - would an earthquake which starts with a TTF of 12s have a cycle length of 12?</p>\n\n<p>Secondly, I don't quite understand why \"Each cycle would have an average absolute error of |cycle_length - 2*prediction_mean|/2 - The error would be cycle_length - 2*prediction_mean at the beginning of a cycle and decrease linearly to 0 by the end of a cycle\". </p>\n\n<p>Does prediction_mean mean the singular value you multiply with the quantile predictions to get the final TTF prediction (the median of [EQ_len, repeated EQ_len_times] in your case, if I understand correctly)? \nIf so, then shouldn't the error be, for a particular 150k section of an EQ, \n- TTF at end of section - prediction_mean, at the very start of the EQ;\n- 0, if TTF at end of section == prediction_mean*quantile prediction;\n- prediction_mean*quantile prediction, if the TTF at the end of the section is 0</p>\n\n<p>I think I need to make sure I understand these correctly before I try and work through the rest of your explanation.</p>\n\n<p>Thank you so much for your patience - I'm probably horrendously wrong, I need your help!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 547303,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-07T14:36:38.113000",
          "content": "<p>To be precise, the quantile targets were set to (number of observations at the end of each 150K until a TTF reset )/(number of observations in that cycle-149999). That way you have training targets that decrease linearly from 1 to 0. I also made sure not to introduce training chunks that overlapped the TTF reset.</p>\n\n<p>A cycle length is indeed the duration of the earthquake until a failure (reset of TTF)</p>\n\n<p>Does prediction_mean mean the scalar you multiply with the quantile predictions to get the final TTF prediction - Yes, to be precise, the prediction is (1-quantile_prediction) * C.</p>\n\n<p>The explanation assumes that the quantile model is perfect, so the error for quantile q is: (1-q)*|cycle_length - 2*prediction_mean|.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 547671,
          "author_name": "June Kim",
          "author_url": "",
          "post_date": "2019-06-08T05:17:05.613000",
          "content": "<p>Got it.\nIf prediction_mean is C in your equation (1-q) * C, then the predicted TTF is (1-q)*C. The true TTF is (1-q)*cycle_length. Shouldn't the absolute error then be (1-q)*|cycle_length-C|? You use this for your calculations with the [8, 4, 3] example.\nI get the explanation as a whole now, though. I understand it as kind of minimizing the 'expected loss', given you don't know what cycle length the 150k sequence belongs to.  I like how you elegenatly multiplied the probability that a sequence belongs to a particular cycle length to the individual average absolute errors. I guess this has the assumption that the distribution of cycle lengths in test data not too far off from that of the training data, although that's essential for a problem to be machine learnable. \nThank you again for sharing!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 547735,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-08T07:37:32.777000",
          "content": "<p>C should be twice the prediction_mean. I have adapted my [8, 4, 3] example since I made a \"divide by two\" mistake to compute the prediction mean which canceled out in the next step. Thank you for making me aware! The conclusion remains unchanged.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543503,
      "author_name": "dhaqui the kaggler",
      "author_url": "",
      "post_date": "2019-06-04T15:09:52.160000",
      "content": "<p>Congrats! Thanks for sharing your solution!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543247,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2019-06-04T12:13:05.897000",
      "content": "<p>Congratulations and thank you for sharing solution!\nYour idea of changing the label to predict quantile of ttf is a really smart approach and I'm impressed.</p>\n\n<p>I also noticed that it is not possible to classify different quake, meaning it is not possible to predict whether this is long or short quake data. However I never come up with changing the problem formulation to predict quantile ttf value.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543143,
      "author_name": "Karan Jakhar",
      "author_url": "",
      "post_date": "2019-06-04T10:51:40.850000",
      "content": "<p>Great approach. Congrats on your gold and thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543086,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T10:18:04.210000",
      "content": "<p>Congrats on your gold, very interesting approach!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 543121,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-04T10:41:45.317000",
          "content": "<p>Thank you! It seems promising to combine this learning objective with superior feature engineering and selection demonstrated by your team and others.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 542965,
      "author_name": "Stanislav Blinov",
      "author_url": "",
      "post_date": "2019-06-04T08:39:19.110000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542927,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2019-06-04T07:56:38.440000",
      "content": "<p>Thanks for sharing !  There is one question. \nWas there a particular reason for choosing that model? (Simple hypothesis, Better CV score, etc.)</p>\n\n<p>In my case, Changing the target into quantile target  makes cv score worse.</p>\n\n<p>Congrats :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 542929,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-04T07:57:48.873000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 542944,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-04T08:09:11.277000",
          "content": "<p>Using the quantile model was a lot better (CV of about 0.1) than modeling TTF directly in local validation, assuming that both models have access to the validation mean and median repeated cycle lengths. I compared the quantile approach to direct models with different normalization schemes to match the validation average/median statistics.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 542952,
          "author_name": "Wonho Song",
          "author_url": "",
          "post_date": "2019-06-04T08:18:32.787000",
          "content": "<p>Great! Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 542916,
      "author_name": "ArvindVepa",
      "author_url": "",
      "post_date": "2019-06-04T07:32:05.100000",
      "content": "<p>wow this is a really great approach</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542896,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2019-06-04T07:09:08.527000",
      "content": "<p>Congratulations and thank you for sharing! Can I ask how did you decide final submissions? I suppose there were lots of choices.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 542908,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-04T07:17:24.533000",
          "content": "<p>I used local validation to select the hyperparameters of the LightGBM and Neural net and averaged 10 random seeds/initializations for models trained on all or the last 6 complete training earthquake cycles. Local validation was split by time - models trained up to the kth earthquake were evaluated on later earthquakes. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 542923,
          "author_name": "u++",
          "author_url": "",
          "post_date": "2019-06-04T07:50:00.410000",
          "content": "<p>I got it, thank you so much!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 542782,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T04:56:27.927000",
      "content": "<p>Congrats and thanks for sharing <a href=\"/tvdwiele\">@tvdwiele</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543168,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2019-06-04T11:07:10.993000",
      "content": "<p>your approach is really cool, especially the unsupervised learning part. Is there any code for it? Also, for the quantile part, that was very smart. i think i remember in another competition, the 1st place instead of predicting pure regression, he predicted [binary 0 or 1, are you an outlier?] * [max value of regression]. I think the way you approached it is similar, and this technique seems powerful in regression. Congratulations on your gold medal</p>",
      "votes": 2,
      "replies": [
        {
          "id": 543215,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-04T11:38:23.843000",
          "content": "<p>Thanks for the feedback! I intend to put the code on GitHub but there is some cleaning up to do first. I ended up not using the CPC features because adding a head that predicts the MAE resulted in leakage, which I realized too close to the deadline. \nAs a starting point I would point to <a href=\"https://github.com/davidtellez/contrastive-predictive-coding\">this excellent resource for learning unsupervised features - CPC by David Tellez</a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 547367,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-07T16:18:01.303000",
          "content": "<p>Hello again, all code has now been published on <a href=\"https://github.com/ttvand/Kaggle-LANL\">Github</a>. The unsupervised part can be found <a href=\"https://github.com/ttvand/Kaggle-LANL/blob/master/Logic/train_valid_test_cpc.py\">here</a>. Let me know if you have any questions.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 547804,
          "author_name": "Noah Weber",
          "author_url": "",
          "post_date": "2019-06-08T09:54:00.603000",
          "content": "<p>Wow, thank you for that! </p>\n\n<p>Just to supplement:\n<a href=\"https://github.com/pumpikano/tf-dann\">Domain-Adversarial Training of Neural Networks in Tensorflow</a></p>\n\n<p><a href=\"https://github.com/CuthbertCai/pytorch_DANN\">pytorch dann</a></p>\n\n<p><a href=\"https://github.com/ataakbari/DANN\">DANN keras</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543129,
      "author_name": "Rachid Abida",
      "author_url": "",
      "post_date": "2019-06-04T10:44:42.340000",
      "content": "<p>I'm a newbie in kaggle world. I learned a lot by participating in this competition. Thanks to all kagglers that have shared their greats kernels and warm congratulations to the winners.  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 543040,
      "author_name": "Prashanth Thangavel",
      "author_url": "",
      "post_date": "2019-06-04T09:43:55.537000",
      "content": "<p>Congrats! Thanks for sharing :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542848,
      "author_name": "Timmmmmms",
      "author_url": "",
      "post_date": "2019-06-04T06:31:03.213000",
      "content": "<p>Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 542774,
      "author_name": "Nanashi",
      "author_url": "",
      "post_date": "2019-06-04T04:48:44.397000",
      "content": "<p>Thank you <a href=\"/tvdwiele\">@tvdwiele</a> !</p>\n\n<h1>👍 🥇</h1>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "542768": "Many congratulations to the winners, The Zoo, you truly stood out on the private leaderboard! Also congratulations to all teams that survived the private earthquake or learned something during the competition. @Keita111 (+4024 to claim a gold with 2 submissions) is a great example that we should sometimes focus less on the public leaderboard even though we all know we should :-). \n\nSpecial thanks go to all forum contributors. People like @CPMPml, @mykper, @scirpus, @Abhishek and several others are what make Kaggle great and I am deeply grateful for your contributions.\n\nLet’s start with what didn’t end up in the final submission. My efforts during this competition were mostly wasted on modeling the gaps (every 4095/4096 observations) in the data. The assumption was that if one could detect these gaps, it would be possible to order the test chunks, making the test prediction problem trivial. This model works perfectly on the validation data but it doesn't on the test data. I am certain that the test data is not contiguous as stated and is actually divided into 150,000 observation chunks with random gaps in between. It is a bit sad that the organizers ignored many good questions about the data distribution in the “Additional info” topic, requiring the competitors to make many assumptions. \nSide note: the gap analysis identified test chunks 1db8e8, 35a2d7, 35dd45, 395e0e, 62a403, 996c37, a35e7c, d1eee8 as the start of an earthquake cycle, adding another piece of information to believe that the test data really comes from the test part of P4677.\n\nThe second biggest chunk of my time was spent on unsupervised learning of features using the raw high frequency data with recurrent networks. The code was inspired by [David Tellez's](https://github.com/davidtellez/contrastive-predictive-coding) implementation of [Contrastive Predictive Coding](https://arxiv.org/pdf/1807.03748.pdf). The hope was that this would lead to robust features for training and testing. I embedded each chunk of 1,500 observations to a shared embedding of size 32 and defined the following loss heads:\n- Predict if one embedding directly follows (with some small random gap) another embedding\n- Predict if a pair of embeddings comes from the same data chunk of 150,000 observations\n- Predict if an embedding comes from the train or test data\n- Predict if a pair of embeddings comes from the same earthquake\n- Predict Time To Failure (TTF)\n\nThe first three losses are trained for both the train and test data of which the third one is trained using [domain adversarial training](https://arxiv.org/pdf/1505.07818.pdf). The last two losses can only be trained for the train data. Even though I ended up not using any of these models, it became clear that it was very hard to distinguish if raw observations come from the same earthquake. That led me to redefine the learning objective in the final approach.\n\n\n**Actual submission**\nAn alternative way to specify the learning objective is the following: predict the quantile TTF of an earthquake, meaning that the target is 0 at the beginning of an earthquake cycle and 1 for the last 150,000 observation chunk of each earthquake. Assuming that this model is trained to minimize the MAE of the quantile, what should each (1-prediction) be multiplied with to minimize the **MAE**? The median of [the earthquake lengths, repeated by the earthquake lengths]! For example, if the test data consisted of earthquake lengths 6, 3 and 2 this would result in median(6, 6, 6, 6, 6, 6, 3, 3, 3, 2, 2) = 6 * (1-quantile_pred) as the optimal predicted time to failure. The nice thing about this reformulation is that we can use the data from the P4677 experiment (I estimated the test median earthquake TTF to be 12) without having to break your head over what earthquake cycles to train on.\n\nThe models themselves use basic FFT and quantile features for each 150,000 observation chunk and 100 1,500 observation subchunks. This resulted in 40*101 = 4040 features which were fed into LightGBM (feature fraction of 0.003) and neural networks. The first final submission averages the LightGBM and neural network predictions starting from the last 6 complete earthquake cycles. The second submission uses all data except for the first incomplete earthquake cycle (given that the cycle length is unknown we can’t determine its quantile). A nice property of this approach is that the predictions don’t change much when training on different folds so in hindsight it would have been better to submit with different estimated median private test earthquake times.\n\nI look forward to your feedback!",
    "548737": "Thanks for sharing! Great performance, especially considering the fact the private dataset was way different from the public one.",
    "546196": "Thank you for sharing! I think your idea is really fresh, and I'm pleasantly surprised that a model  'smudging' the predictions worked better than most models that tried to predict TTF directly. Just confused about some details - how did you assign quantile TTF value for intermediate points in the earthquake (not at the start, not at the last 150000 time points)? Also, why did you use the median of [earthquake length, repeated earthquake_length times], instead of simply the mean of earthquake lengths?",
    "543503": "Congrats! Thanks for sharing your solution!",
    "543247": "Congratulations and thank you for sharing solution!\nYour idea of changing the label to predict quantile of ttf is a really smart approach and I'm impressed.\n\nI also noticed that it is not possible to classify different quake, meaning it is not possible to predict whether this is long or short quake data. However I never come up with changing the problem formulation to predict quantile ttf value.",
    "543143": "Great approach. Congrats on your gold and thanks for sharing.",
    "543086": "Congrats on your gold, very interesting approach!",
    "542965": "Congratulations!",
    "542927": "Thanks for sharing !  There is one question. \nWas there a particular reason for choosing that model? (Simple hypothesis, Better CV score, etc.)\n\nIn my case, Changing the target into quantile target  makes cv score worse.\n\nCongrats :)",
    "542916": "wow this is a really great approach",
    "542896": "Congratulations and thank you for sharing! Can I ask how did you decide final submissions? I suppose there were lots of choices.",
    "542782": "Congrats and thanks for sharing @tvdwiele !",
    "543168": "your approach is really cool, especially the unsupervised learning part. Is there any code for it? Also, for the quantile part, that was very smart. i think i remember in another competition, the 1st place instead of predicting pure regression, he predicted [binary 0 or 1, are you an outlier?] * [max value of regression]. I think the way you approached it is similar, and this technique seems powerful in regression. Congratulations on your gold medal",
    "543129": "I'm a newbie in kaggle world. I learned a lot by participating in this competition. Thanks to all kagglers that have shared their greats kernels and warm congratulations to the winners.  ",
    "543040": "Congrats! Thanks for sharing :)",
    "542848": "Thank you!",
    "542774": "Thank you @tvdwiele !\n# 👍 🥇 "
  }
}