{
  "id": 94559,
  "title": "CNNs would have worked...",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94559",
  "author_name": "",
  "post_date": "2019-06-05T08:12:12.012945900Z",
  "votes": 14,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi,\nanother 'lessons learned' story: I had played around a lot with CNNs working directly on the acoustic data, and discarded the idea when I couldn't get below 1.6 on the public LB. Turns out the private LB would have been 2.40 (top 25) <em>without any 'leak' information and without generating a single feature</em>. <a href=\"https://www.kaggle.com/friedchips/simple-cnn-would-have-been-top-25\">I've published the kernel here</a>, as it is different from all other approaches I've seen here.\nLessons:\n- Physics can't be beat. If anybody would have found a way to predict ttf (even ttf to the next major <em>or</em> minor quake!) correctly for individual quake cycles, that would have probably ended in a Nobel Prize. Not joking, as a physicist myself, I know that this would change statistical physics fundamentally...\n- Therefore, success is determined by averaging the cycles in the most intelligent way. This is why the leak information was so powerful :)\n- the power density used in my kernel is probably the most important physical quantity in this process. It actually corresponds to the energy released by the slips as can be seen in this plot. All the minor and major quakes in train are clearly visible:\n<img src=\"https://i.imgur.com/x7KLxmK.png\" alt=\"power\"></p>\n\n<p>I'd be interested to hear if anyone else has played around with the power density?</p>",
  "messages": [
    {
      "id": "544165",
      "postDate": "06/05/2019 08:12:12",
      "content": "<p>Hi,\nanother 'lessons learned' story: I had played around a lot with CNNs working directly on the acoustic data, and discarded the idea when I couldn't get below 1.6 on the public LB. Turns out the private LB would have been 2.40 (top 25) <em>without any 'leak' information and without generating a single feature</em>. <a href=\"https://www.kaggle.com/friedchips/simple-cnn-would-have-been-top-25\">I've published the kernel here</a>, as it is different from all other approaches I've seen here.\nLessons:\n- Physics can't be beat. If anybody would have found a way to predict ttf (even ttf to the next major <em>or</em> minor quake!) correctly for individual quake cycles, that would have probably ended in a Nobel Prize. Not joking, as a physicist myself, I know that this would change statistical physics fundamentally...\n- Therefore, success is determined by averaging the cycles in the most intelligent way. This is why the leak information was so powerful :)\n- the power density used in my kernel is probably the most important physical quantity in this process. It actually corresponds to the energy released by the slips as can be seen in this plot. All the minor and major quakes in train are clearly visible:\n<img src=\"https://i.imgur.com/x7KLxmK.png\" alt=\"power\"></p>\n\n<p>I'd be interested to hear if anyone else has played around with the power density?</p>",
      "rawMarkdown": "Hi,\nanother 'lessons learned' story: I had played around a lot with CNNs working directly on the acoustic data, and discarded the idea when I couldn't get below 1.6 on the public LB. Turns out the private LB would have been 2.40 (top 25) *without any 'leak' information and without generating a single feature*. [I've published the kernel here](https://www.kaggle.com/friedchips/simple-cnn-would-have-been-top-25), as it is different from all other approaches I've seen here.\nLessons:\n- Physics can't be beat. If anybody would have found a way to predict ttf (even ttf to the next major *or* minor quake!) correctly for individual quake cycles, that would have probably ended in a Nobel Prize. Not joking, as a physicist myself, I know that this would change statistical physics fundamentally...\n- Therefore, success is determined by averaging the cycles in the most intelligent way. This is why the leak information was so powerful :)\n- the power density used in my kernel is probably the most important physical quantity in this process. It actually corresponds to the energy released by the slips as can be seen in this plot. All the minor and major quakes in train are clearly visible:\n![power](https://i.imgur.com/x7KLxmK.png)\n\nI'd be interested to hear if anyone else has played around with the power density?",
      "votes": null
    },
    {
      "id": "544185",
      "postDate": "06/05/2019 08:45:32",
      "content": "<p>Nice approach.  I used CNN with straight TTF data or TTF plus std, got priLB 2.58 ,  pubLB 1.625.  Resnet 2.53, 1.747.  </p>",
      "rawMarkdown": "Nice approach.  I used CNN with straight TTF data or TTF plus std, got priLB 2.58 ,  pubLB 1.625.  Resnet 2.53, 1.747.",
      "votes": null
    },
    {
      "id": "544187",
      "postDate": "06/05/2019 08:49:50",
      "content": "<p>A lot of models would have worked. I have a KNN that scores 2.382 and would claim a gold medal and the 18th rank here, see attached picture.  It does not use any test info.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544187/13408/knn.png\" alt=\"knn\"></p>",
      "rawMarkdown": "A lot of models would have worked. I have a KNN that scores 2.382 and would claim a gold medal and the 18th rank here, see attached picture.  It does not use any test info.\n\n![knn](https://storage.googleapis.com/kaggle-forum-message-attachments/544187/13408/knn.png)",
      "votes": null
    },
    {
      "id": "544191",
      "postDate": "06/05/2019 08:52:40",
      "content": "<p>Thanks for your reply! You mean \"straight acoustic data\", right? Did you use a full 150,000 point window as input? My main motivation for using the power density was the ability to downsample a window from 150,000 to 300 points without losing accuracy.</p>",
      "rawMarkdown": "Thanks for your reply! You mean \"straight acoustic data\", right? Did you use a full 150,000 point window as input? My main motivation for using the power density was the ability to downsample a window from 150,000 to 300 points without losing accuracy.",
      "votes": null
    },
    {
      "id": "544193",
      "postDate": "06/05/2019 08:53:36",
      "content": "<p>How do you generate this plot? Actually the nature paper in the introduction thread clearly stated that acoustic power in the early stage of the cycle can predict TTF. I couldn't reproduce it. Your plot looks like what the paper said.</p>",
      "rawMarkdown": "How do you generate this plot? Actually the nature paper in the introduction thread clearly stated that acoustic power in the early stage of the cycle can predict TTF. I couldn't reproduce it. Your plot looks like what the paper said.",
      "votes": null
    },
    {
      "id": "544196",
      "postDate": "06/05/2019 08:57:35",
      "content": "<p>Yes, that's definitely true. The reason why I though this CNN might be of interest is because it doesn't use any aggregated features at all, only the (squared and downsampled) acoustic amplitude directly. There had been several discussion posts during the competition that said neural networks would perform poorly if used on the acoustic data directly.</p>",
      "rawMarkdown": "Yes, that's definitely true. The reason why I though this CNN might be of interest is because it doesn't use any aggregated features at all, only the (squared and downsampled) acoustic amplitude directly. There had been several discussion posts during the competition that said neural networks would perform poorly if used on the acoustic data directly.",
      "votes": null
    },
    {
      "id": "544207",
      "postDate": "06/05/2019 09:09:46",
      "content": "<p>If there's interest I will write a kernel that shows all of my acoustic power (AP) calculations. Short answer: this is the integral (np.cumsum) of the acoustic power in train minus the linear part to show the true power released by the slips.\nAs for the Nature paper, that is actually my main contention: you can tell from the AP approximately <em>where</em> you are in the cycle (think of it as % of the span between 2 quakes, the main remaining problem here are the minor quakes which reset the process a bit but don't count), but there is <em>no way</em> to tell whether it's going to be a 8 sec. or a 16 sec. quake cycle. Anything else would require a modification of physics as we know it.</p>",
      "rawMarkdown": "If there's interest I will write a kernel that shows all of my acoustic power (AP) calculations. Short answer: this is the integral (np.cumsum) of the acoustic power in train minus the linear part to show the true power released by the slips.\nAs for the Nature paper, that is actually my main contention: you can tell from the AP approximately *where* you are in the cycle (think of it as % of the span between 2 quakes, the main remaining problem here are the minor quakes which reset the process a bit but don't count), but there is *no way* to tell whether it's going to be a 8 sec. or a 16 sec. quake cycle. Anything else would require a modification of physics as we know it.",
      "votes": null
    },
    {
      "id": "544215",
      "postDate": "06/05/2019 09:18:00",
      "content": "<p>You're right, sorry if I appear to downplay what you have.  I will certainly look at what you shared!</p>",
      "rawMarkdown": "You're right, sorry if I appear to downplay what you have.  I will certainly look at what you shared!",
      "votes": null
    },
    {
      "id": "544229",
      "postDate": "06/05/2019 09:40:11",
      "content": "<p>No problem at all! The reason why I'm doing this (apart from valuable internet points :) ) is that, based on my physics knowledge, I believe that this model \"knows all there is to know\" about the data. I suspect that someone more knowledgeable in neural nets than me would be able to to push this quite a bit further, maybe even to a (late) #1 result. My specialty is model building, machine learning is something I'm just picking up right now :). Turning to the CHAMPS molecular NMR challenge now...</p>",
      "rawMarkdown": "No problem at all! The reason why I'm doing this (apart from valuable internet points :) ) is that, based on my physics knowledge, I believe that this model \"knows all there is to know\" about the data. I suspect that someone more knowledgeable in neural nets than me would be able to to push this quite a bit further, maybe even to a (late) #1 result. My specialty is model building, machine learning is something I'm just picking up right now :). Turning to the CHAMPS molecular NMR challenge now...",
      "votes": null
    },
    {
      "id": "544254",
      "postDate": "06/05/2019 10:12:43",
      "content": "<blockquote>\n  <p>Turning to the CHAMPS molecular NMR challenge now…</p>\n</blockquote>\n\n<p>I'll meet you there soon!</p>",
      "rawMarkdown": "&gt; Turning to the CHAMPS molecular NMR challenge now…\n\nI'll meet you there soon!",
      "votes": null
    },
    {
      "id": "544297",
      "postDate": "06/05/2019 11:11:18",
      "content": "<p>Isn't var close to power density?  4 of my 12 features are var based, and many teams used similar features.</p>\n\n<p>Also, you start from (x - x_mean)**2 hence you get std for free.  Saying you learn it is a bit of an exaggeration.  Am I missing something?</p>\n\n<p>Other than that it is a simple yet effective approach.</p>",
      "rawMarkdown": "Isn't var close to power density?  4 of my 12 features are var based, and many teams used similar features.\n\nAlso, you start from (x - x_mean)**2 hence you get std for free.  Saying you learn it is a bit of an exaggeration.  Am I missing something?\n\nOther than that it is a simple yet effective approach.",
      "votes": null
    },
    {
      "id": "544382",
      "postDate": "06/05/2019 13:01:03",
      "content": "<p>You are absolutely right. Power density is closely related to variance. The average power density over a sample window is identical to the signal variance in that window. But the neural network doesn't know that value. The power density is basically a kind of \"instantaneous\" variance. When I downsample the signal by a factor of 500 as done in my kernel, I am feeding the CNN a time series of 300 \"instantaneous variance\" measurements. Each one of them would be useless as a feature by itself, it is only their time sequence which contains any information about ttf and the quake cycle. This is why I said that the CNN learns about the variance/std.\nAs I'm still learning about neural networks, please correct me if this view is wrong. Thanks!</p>",
      "rawMarkdown": "You are absolutely right. Power density is closely related to variance. The average power density over a sample window is identical to the signal variance in that window. But the neural network doesn't know that value. The power density is basically a kind of \"instantaneous\" variance. When I downsample the signal by a factor of 500 as done in my kernel, I am feeding the CNN a time series of 300 \"instantaneous variance\" measurements. Each one of them would be useless as a feature by itself, it is only their time sequence which contains any information about ttf and the quake cycle. This is why I said that the CNN learns about the variance/std.\nAs I'm still learning about neural networks, please correct me if this view is wrong. Thanks!",
      "votes": null
    },
    {
      "id": "544403",
      "postDate": "06/05/2019 13:23:23",
      "content": "<blockquote>\n  <p>The average power density over a sample window is identical to the signal variance in that window.</p>\n</blockquote>\n\n<p>I agree, but NN work by alternating linear combination of input and activation.  Mean is the simplest linear combination, and a NN with a single unit can learn it ;)</p>\n\n<p>Don't get me wrong, it is a great idea to replace signal by power density.  A simple transfrom that makes CNN effective.</p>",
      "rawMarkdown": "&gt; The average power density over a sample window is identical to the signal variance in that window.\n\nI agree, but NN work by alternating linear combination of input and activation.  Mean is the simplest linear combination, and a NN with a single unit can learn it ;)\n\nDon't get me wrong, it is a great idea to replace signal by power density.  A simple transfrom that makes CNN effective.",
      "votes": null
    },
    {
      "id": "544447",
      "postDate": "06/05/2019 14:19:39",
      "content": "<p>Got it. Of course you're right. Thanks!</p>",
      "rawMarkdown": "Got it. Of course you're right. Thanks!",
      "votes": null
    },
    {
      "id": "544519",
      "postDate": "06/05/2019 15:46:06",
      "content": "<p>I down sample seq length to 15000 or 30000 with averaging.  I did 1st difference and 2nd difference after time shift, that didn't work for getting train_loss and val_loss down, although it could work for this competition.  I didn't submit anything.    I got CNN_LSTM and Inception models' Val_loss around 0.5-06.  The models also get similar mae on self_splitted  test sets comparing to val_mae.  But these models didn't do well at both LBs.   Since test data is rather different from training data.  With larger train data set,  these models could definitely learn the pattern very well.  A posted Wavenet_LSTM  seems to perform even better.  </p>",
      "rawMarkdown": "I down sample seq length to 15000 or 30000 with averaging.  I did 1st difference and 2nd difference after time shift, that didn't work for getting train_loss and val_loss down, although it could work for this competition.  I didn't submit anything.    I got CNN_LSTM and Inception models' Val_loss around 0.5-06.  The models also get similar mae on self_splitted  test sets comparing to val_mae.  But these models didn't do well at both LBs.   Since test data is rather different from training data.  With larger train data set,  these models could definitely learn the pattern very well.  A posted Wavenet_LSTM  seems to perform even better.",
      "votes": null
    },
    {
      "id": "544525",
      "postDate": "06/05/2019 15:55:17",
      "content": "<p>Did you down sample through averaging to 300 as comparison ?  </p>",
      "rawMarkdown": "Did you down sample through averaging to 300 as comparison ?",
      "votes": null
    },
    {
      "id": "544530",
      "postDate": "06/05/2019 16:05:32",
      "content": "<p>Markus.  NNs can be very different.  I definitely think some of them work very well.  It's really currently best methods for analyzing sequence data.  This competition only contains 15 earthquakes in the training data.  These models overfit 15 earthquakes, don't generalize well.  With larger training data set, these methods can definitely learn the patterns.   </p>",
      "rawMarkdown": "Markus.  NNs can be very different.  I definitely think some of them work very well.  It's really currently best methods for analyzing sequence data.  This competition only contains 15 earthquakes in the training data.  These models overfit 15 earthquakes, don't generalize well.  With larger training data set, these methods can definitely learn the patterns.",
      "votes": null
    },
    {
      "id": "544676",
      "postDate": "06/05/2019 19:00:08",
      "content": "<p>I tried several different downsampling factors and found to my own surprise that I could go as high as 500 without losing any of the information contained in the signal. The important thing was using the squared signal, because otherwise the information would have been lost in the downsampling process.</p>\n\n<p>This seems to fit well with the observation made by many other commenters that the frequency spectrum is not helpful at all. The relevant information seems to be in the shape of the peaks themselves.</p>",
      "rawMarkdown": "I tried several different downsampling factors and found to my own surprise that I could go as high as 500 without losing any of the information contained in the signal. The important thing was using the squared signal, because otherwise the information would have been lost in the downsampling process.\n\nThis seems to fit well with the observation made by many other commenters that the frequency spectrum is not helpful at all. The relevant information seems to be in the shape of the peaks themselves.",
      "votes": null
    },
    {
      "id": "544682",
      "postDate": "06/05/2019 19:04:12",
      "content": "<p>In my experience this dataset and its size actually works quite well. of course, you might argue that a neural network is not really needed here, and you would be right. Tree-based methods are more than enough to get top level results. The important thing that I did in my kernel was to stop early after about 16 epochs to make sure that the model doesn't overfit to the different quake cycle lengths and therefore loses its predictive power for the test set.</p>",
      "rawMarkdown": "In my experience this dataset and its size actually works quite well. of course, you might argue that a neural network is not really needed here, and you would be right. Tree-based methods are more than enough to get top level results. The important thing that I did in my kernel was to stop early after about 16 epochs to make sure that the model doesn't overfit to the different quake cycle lengths and therefore loses its predictive power for the test set.",
      "votes": null
    },
    {
      "id": "544687",
      "postDate": "06/05/2019 19:13:08",
      "content": "<p>Markus,</p>\n\n<p>This is  interesting - I just posted a thread asking if there were any cnn's that scored well (and embarked on a mission to prove it could be done). I just as quickly removed the thread when I saw this. I think many gave up on the cnn when they saw the poor LB scores - something to remember in the next competition!</p>\n\n<p>One thing I would push back on - the necessity of the use of power density.  In a cnn, the convolutional filters can easily handle an oscillating signal - they work on them all the time and there is no zeroing of the result because the filters themselves oscillate. And, when you directly downsample to a few hundred samples, surely you are losing information per Nyquist theorem (In my nets, I resample 150K to 37500 in prepossessing in the assumption that the data from 500K to 2 MHz has not much info). </p>\n\n<p>Progressive convolutions with stride, not max pool, can successfully reduce the number of samples from 37500 in an information-preserving way - in these layers, the cnn simply supplies a cost-effective-but-linear way to preserve the info. With enough reduction, perhaps a layer could be added to convert to power and non-linearity added to the remainder of the net?</p>\n\n<p>I still intend to try this, even though the data set is pretty poor.</p>",
      "rawMarkdown": "Markus,\n\nThis is  interesting - I just posted a thread asking if there were any cnn's that scored well (and embarked on a mission to prove it could be done). I just as quickly removed the thread when I saw this. I think many gave up on the cnn when they saw the poor LB scores - something to remember in the next competition!\n\nOne thing I would push back on - the necessity of the use of power density.  In a cnn, the convolutional filters can easily handle an oscillating signal - they work on them all the time and there is no zeroing of the result because the filters themselves oscillate. And, when you directly downsample to a few hundred samples, surely you are losing information per Nyquist theorem (In my nets, I resample 150K to 37500 in prepossessing in the assumption that the data from 500K to 2 MHz has not much info). \n\nProgressive convolutions with stride, not max pool, can successfully reduce the number of samples from 37500 in an information-preserving way - in these layers, the cnn simply supplies a cost-effective-but-linear way to preserve the info. With enough reduction, perhaps a layer could be added to convert to power and non-linearity added to the remainder of the net?\n\nI still intend to try this, even though the data set is pretty poor.",
      "votes": null
    },
    {
      "id": "546011",
      "postDate": "06/06/2019 06:26:30",
      "content": "<p>Hi Pete,\nyou are right, it isn't <em>necessary</em> to use the power density. I should have been more specific: it makes it a lot easier to downsample, which I definitely wanted to do.</p>\n\n<p>As you said, there is only noise above 500 kHz. So you can downsample the original acoustic data by 4x without losing any information. However, when you downsample further, you get into the frequency range where strong signals exist and therefore peaks will start to cancel each other and disappear. Using the power density eliminates this problem. I wanted to downsample further for two reasons:\n- I looked at the frequency spectrum for all samples and saw that the spectrum shape is practically uncorrelated with TTF. This fits well with the observations in the forum as well as in the organizer's publications that spectrum-based features are not very useful. So there was no motivation for me to use a 10,000+ pt input layer.\n- I realized that I could go down to 500x downsampling without losing accuracy (I started with 10x). That's only 300 pts in the input layer. My NN now was a lot faster. Also, I could now try LSTM and GRU as well, which I did. They actually work almost as well as CNN, but I never tried this any further.</p>\n\n<p>In summary, I used the power density because I think that there is no information in the oscillations of the signal. The only thing that matters are the peaks, their overall shape, height, frequency (which is ~ 1 kHz max)) / distance.</p>",
      "rawMarkdown": "Hi Pete,\nyou are right, it isn't *necessary* to use the power density. I should have been more specific: it makes it a lot easier to downsample, which I definitely wanted to do.\n\nAs you said, there is only noise above 500 kHz. So you can downsample the original acoustic data by 4x without losing any information. However, when you downsample further, you get into the frequency range where strong signals exist and therefore peaks will start to cancel each other and disappear. Using the power density eliminates this problem. I wanted to downsample further for two reasons:\n- I looked at the frequency spectrum for all samples and saw that the spectrum shape is practically uncorrelated with TTF. This fits well with the observations in the forum as well as in the organizer's publications that spectrum-based features are not very useful. So there was no motivation for me to use a 10,000+ pt input layer.\n- I realized that I could go down to 500x downsampling without losing accuracy (I started with 10x). That's only 300 pts in the input layer. My NN now was a lot faster. Also, I could now try LSTM and GRU as well, which I did. They actually work almost as well as CNN, but I never tried this any further.\n\n\n\n\n\n\nIn summary, I used the power density because I think that there is no information in the oscillations of the signal. The only thing that matters are the peaks, their overall shape, height, frequency (which is ~ 1 kHz max)) / distance.",
      "votes": null
    },
    {
      "id": "546940",
      "postDate": "06/07/2019 03:32:12",
      "content": "<p>I was in exactly the same boat. I had newly learned how to use deep learning and tensorflow, and was under the \"when all you have is a hammer, everything looks like a nail\" and was so convinced that neural nets would be the key to this competition. Much to my dismay, the public leaderboard scores were crap even though I had a fairly solid CV strategy. Turns out a 1.86 public leaderboard would have resulted in a 4th place finish with 2.29 private leaderboard. Big lessons learned here.</p>\n\n<p>Preprocessing step: Rolling 50 standard deviation, which is then downsampled by a factor of 10. I then take the log of these values to make it more approximately normally distributed, and then I subtract mean and standard deviation in order to center/scale. Each input is of size 14995 as a result.</p>\n\n<p>Network architecture: 1D CNN --&gt; 1D Max Pool --&gt; GRU --&gt; Dense --&gt; Output layer, although I'm not sure the GRU helped at all.</p>",
      "rawMarkdown": "I was in exactly the same boat. I had newly learned how to use deep learning and tensorflow, and was under the \"when all you have is a hammer, everything looks like a nail\" and was so convinced that neural nets would be the key to this competition. Much to my dismay, the public leaderboard scores were crap even though I had a fairly solid CV strategy. Turns out a 1.86 public leaderboard would have resulted in a 4th place finish with 2.29 private leaderboard. Big lessons learned here.\n\nPreprocessing step: Rolling 50 standard deviation, which is then downsampled by a factor of 10. I then take the log of these values to make it more approximately normally distributed, and then I subtract mean and standard deviation in order to center/scale. Each input is of size 14995 as a result.\n\nNetwork architecture: 1D CNN --&gt; 1D Max Pool --&gt; GRU --&gt; Dense --&gt; Output layer, although I'm not sure the GRU helped at all.",
      "votes": null
    },
    {
      "id": "547032",
      "postDate": "06/07/2019 07:13:18",
      "content": "<p>Very interesting. Your method and mine are really quite similar. as written in some comments above, the standard deviation or variance and the power density are practically the same.\nthe fact that you managed to get a significantly better score even than the 2.40 that I managed confirms my suspicion that using a neural network on the original signal (squared in one way or another and sufficiently down sampled) might in the end be the best approach for predicting new and really unknown test data.\nWould be interesting to know if your superior result is due to the different data preparation or due to your superior network topology and parameter choice...</p>",
      "rawMarkdown": "Very interesting. Your method and mine are really quite similar. as written in some comments above, the standard deviation or variance and the power density are practically the same.\nthe fact that you managed to get a significantly better score even than the 2.40 that I managed confirms my suspicion that using a neural network on the original signal (squared in one way or another and sufficiently down sampled) might in the end be the best approach for predicting new and really unknown test data.\nWould be interesting to know if your superior result is due to the different data preparation or due to your superior network topology and parameter choice...",
      "votes": null
    },
    {
      "id": "547233",
      "postDate": "06/07/2019 12:31:07",
      "content": "<p>I am planning on cleaning up the code a bit and posting it on Github. I'll be sure to let you know so you can take a look! My code is separated into several files, so not feeling like a kernel is the best route.</p>\n\n<p>I think the difference is likely due to a difference in my data generators. I fed the processed data in a gigantic 1D chunk to the generators, and the generators randomly decided where to start/stop in a stratified manner along the CV folds (Used a custom CV split that split along entire EQs, but was as balanced as possible based on certain basic features, like total EQ length). It also had a method that prevented me from training on any EQs that spanned between two consecutive EQs. I also threw out a few of the earthquakes (like the infamous 2 EQs for the price of 1 in the train set).</p>\n\n<p>I by no means optimized the architecture. I spent a ton of time optimizing the architecture at the beginning, only to have very little gains, and threw out hundreds of lines of code to end up with the relatively simple network that seemed to do just as well as the others.</p>",
      "rawMarkdown": "I am planning on cleaning up the code a bit and posting it on Github. I'll be sure to let you know so you can take a look! My code is separated into several files, so not feeling like a kernel is the best route.\n\nI think the difference is likely due to a difference in my data generators. I fed the processed data in a gigantic 1D chunk to the generators, and the generators randomly decided where to start/stop in a stratified manner along the CV folds (Used a custom CV split that split along entire EQs, but was as balanced as possible based on certain basic features, like total EQ length). It also had a method that prevented me from training on any EQs that spanned between two consecutive EQs. I also threw out a few of the earthquakes (like the infamous 2 EQs for the price of 1 in the train set).\n\nI by no means optimized the architecture. I spent a ton of time optimizing the architecture at the beginning, only to have very little gains, and threw out hundreds of lines of code to end up with the relatively simple network that seemed to do just as well as the others.",
      "votes": null
    },
    {
      "id": "548009",
      "postDate": "06/08/2019 16:15:58",
      "content": "<p>Here is the link as promised, let me know if you have any questions! Always down to chat deep learning ^^\n<a href=\"https://github.com/Atrus619/Top-Scoring-Kaggle-Model-LANL-EQ-Prediction-4th-Place-\">https://github.com/Atrus619/Top-Scoring-Kaggle-Model-LANL-EQ-Prediction-4th-Place-</a></p>",
      "rawMarkdown": "Here is the link as promised, let me know if you have any questions! Always down to chat deep learning ^^\nhttps://github.com/Atrus619/Top-Scoring-Kaggle-Model-LANL-EQ-Prediction-4th-Place-",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 544185,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "06/05/2019 08:45:32",
      "content": "<p>Nice approach.  I used CNN with straight TTF data or TTF plus std, got priLB 2.58 ,  pubLB 1.625.  Resnet 2.53, 1.747.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 544191,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 08:52:40",
          "content": "<p>Thanks for your reply! You mean \"straight acoustic data\", right? Did you use a full 150,000 point window as input? My main motivation for using the power density was the ability to downsample a window from 150,000 to 300 points without losing accuracy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544519,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "06/05/2019 15:46:06",
          "content": "<p>I down sample seq length to 15000 or 30000 with averaging.  I did 1st difference and 2nd difference after time shift, that didn't work for getting train_loss and val_loss down, although it could work for this competition.  I didn't submit anything.    I got CNN_LSTM and Inception models' Val_loss around 0.5-06.  The models also get similar mae on self_splitted  test sets comparing to val_mae.  But these models didn't do well at both LBs.   Since test data is rather different from training data.  With larger train data set,  these models could definitely learn the pattern very well.  A posted Wavenet_LSTM  seems to perform even better.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544525,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "06/05/2019 15:55:17",
          "content": "<p>Did you down sample through averaging to 300 as comparison ?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544676,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 19:00:08",
          "content": "<p>I tried several different downsampling factors and found to my own surprise that I could go as high as 500 without losing any of the information contained in the signal. The important thing was using the squared signal, because otherwise the information would have been lost in the downsampling process.</p>\n\n<p>This seems to fit well with the observation made by many other commenters that the frequency spectrum is not helpful at all. The relevant information seems to be in the shape of the peaks themselves.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544187,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2019 08:49:50",
      "content": "<p>A lot of models would have worked. I have a KNN that scores 2.382 and would claim a gold medal and the 18th rank here, see attached picture.  It does not use any test info.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/544187/13408/knn.png\" alt=\"knn\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 544196,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 08:57:35",
          "content": "<p>Yes, that's definitely true. The reason why I though this CNN might be of interest is because it doesn't use any aggregated features at all, only the (squared and downsampled) acoustic amplitude directly. There had been several discussion posts during the competition that said neural networks would perform poorly if used on the acoustic data directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544215,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 09:18:00",
          "content": "<p>You're right, sorry if I appear to downplay what you have.  I will certainly look at what you shared!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544229,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 09:40:11",
          "content": "<p>No problem at all! The reason why I'm doing this (apart from valuable internet points :) ) is that, based on my physics knowledge, I believe that this model \"knows all there is to know\" about the data. I suspect that someone more knowledgeable in neural nets than me would be able to to push this quite a bit further, maybe even to a (late) #1 result. My specialty is model building, machine learning is something I'm just picking up right now :). Turning to the CHAMPS molecular NMR challenge now...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544254,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 10:12:43",
          "content": "<blockquote>\n  <p>Turning to the CHAMPS molecular NMR challenge now…</p>\n</blockquote>\n\n<p>I'll meet you there soon!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544530,
          "author_name": "joejeo1",
          "author_url": "",
          "post_date": "06/05/2019 16:05:32",
          "content": "<p>Markus.  NNs can be very different.  I definitely think some of them work very well.  It's really currently best methods for analyzing sequence data.  This competition only contains 15 earthquakes in the training data.  These models overfit 15 earthquakes, don't generalize well.  With larger training data set, these methods can definitely learn the patterns.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544682,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 19:04:12",
          "content": "<p>In my experience this dataset and its size actually works quite well. of course, you might argue that a neural network is not really needed here, and you would be right. Tree-based methods are more than enough to get top level results. The important thing that I did in my kernel was to stop early after about 16 epochs to make sure that the model doesn't overfit to the different quake cycle lengths and therefore loses its predictive power for the test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544193,
      "author_name": "frankw",
      "author_url": "",
      "post_date": "06/05/2019 08:53:36",
      "content": "<p>How do you generate this plot? Actually the nature paper in the introduction thread clearly stated that acoustic power in the early stage of the cycle can predict TTF. I couldn't reproduce it. Your plot looks like what the paper said.</p>",
      "votes": null,
      "replies": [
        {
          "id": 544207,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 09:09:46",
          "content": "<p>If there's interest I will write a kernel that shows all of my acoustic power (AP) calculations. Short answer: this is the integral (np.cumsum) of the acoustic power in train minus the linear part to show the true power released by the slips.\nAs for the Nature paper, that is actually my main contention: you can tell from the AP approximately <em>where</em> you are in the cycle (think of it as % of the span between 2 quakes, the main remaining problem here are the minor quakes which reset the process a bit but don't count), but there is <em>no way</em> to tell whether it's going to be a 8 sec. or a 16 sec. quake cycle. Anything else would require a modification of physics as we know it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544297,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2019 11:11:18",
      "content": "<p>Isn't var close to power density?  4 of my 12 features are var based, and many teams used similar features.</p>\n\n<p>Also, you start from (x - x_mean)**2 hence you get std for free.  Saying you learn it is a bit of an exaggeration.  Am I missing something?</p>\n\n<p>Other than that it is a simple yet effective approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 544382,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 13:01:03",
          "content": "<p>You are absolutely right. Power density is closely related to variance. The average power density over a sample window is identical to the signal variance in that window. But the neural network doesn't know that value. The power density is basically a kind of \"instantaneous\" variance. When I downsample the signal by a factor of 500 as done in my kernel, I am feeding the CNN a time series of 300 \"instantaneous variance\" measurements. Each one of them would be useless as a feature by itself, it is only their time sequence which contains any information about ttf and the quake cycle. This is why I said that the CNN learns about the variance/std.\nAs I'm still learning about neural networks, please correct me if this view is wrong. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544403,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/05/2019 13:23:23",
          "content": "<blockquote>\n  <p>The average power density over a sample window is identical to the signal variance in that window.</p>\n</blockquote>\n\n<p>I agree, but NN work by alternating linear combination of input and activation.  Mean is the simplest linear combination, and a NN with a single unit can learn it ;)</p>\n\n<p>Don't get me wrong, it is a great idea to replace signal by power density.  A simple transfrom that makes CNN effective.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544447,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/05/2019 14:19:39",
          "content": "<p>Got it. Of course you're right. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 544687,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "06/05/2019 19:13:08",
      "content": "<p>Markus,</p>\n\n<p>This is  interesting - I just posted a thread asking if there were any cnn's that scored well (and embarked on a mission to prove it could be done). I just as quickly removed the thread when I saw this. I think many gave up on the cnn when they saw the poor LB scores - something to remember in the next competition!</p>\n\n<p>One thing I would push back on - the necessity of the use of power density.  In a cnn, the convolutional filters can easily handle an oscillating signal - they work on them all the time and there is no zeroing of the result because the filters themselves oscillate. And, when you directly downsample to a few hundred samples, surely you are losing information per Nyquist theorem (In my nets, I resample 150K to 37500 in prepossessing in the assumption that the data from 500K to 2 MHz has not much info). </p>\n\n<p>Progressive convolutions with stride, not max pool, can successfully reduce the number of samples from 37500 in an information-preserving way - in these layers, the cnn simply supplies a cost-effective-but-linear way to preserve the info. With enough reduction, perhaps a layer could be added to convert to power and non-linearity added to the remainder of the net?</p>\n\n<p>I still intend to try this, even though the data set is pretty poor.</p>",
      "votes": null,
      "replies": [
        {
          "id": 546011,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/06/2019 06:26:30",
          "content": "<p>Hi Pete,\nyou are right, it isn't <em>necessary</em> to use the power density. I should have been more specific: it makes it a lot easier to downsample, which I definitely wanted to do.</p>\n\n<p>As you said, there is only noise above 500 kHz. So you can downsample the original acoustic data by 4x without losing any information. However, when you downsample further, you get into the frequency range where strong signals exist and therefore peaks will start to cancel each other and disappear. Using the power density eliminates this problem. I wanted to downsample further for two reasons:\n- I looked at the frequency spectrum for all samples and saw that the spectrum shape is practically uncorrelated with TTF. This fits well with the observations in the forum as well as in the organizer's publications that spectrum-based features are not very useful. So there was no motivation for me to use a 10,000+ pt input layer.\n- I realized that I could go down to 500x downsampling without losing accuracy (I started with 10x). That's only 300 pts in the input layer. My NN now was a lot faster. Also, I could now try LSTM and GRU as well, which I did. They actually work almost as well as CNN, but I never tried this any further.</p>\n\n<p>In summary, I used the power density because I think that there is no information in the oscillations of the signal. The only thing that matters are the peaks, their overall shape, height, frequency (which is ~ 1 kHz max)) / distance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 546940,
      "author_name": "atrus619",
      "author_url": "",
      "post_date": "06/07/2019 03:32:12",
      "content": "<p>I was in exactly the same boat. I had newly learned how to use deep learning and tensorflow, and was under the \"when all you have is a hammer, everything looks like a nail\" and was so convinced that neural nets would be the key to this competition. Much to my dismay, the public leaderboard scores were crap even though I had a fairly solid CV strategy. Turns out a 1.86 public leaderboard would have resulted in a 4th place finish with 2.29 private leaderboard. Big lessons learned here.</p>\n\n<p>Preprocessing step: Rolling 50 standard deviation, which is then downsampled by a factor of 10. I then take the log of these values to make it more approximately normally distributed, and then I subtract mean and standard deviation in order to center/scale. Each input is of size 14995 as a result.</p>\n\n<p>Network architecture: 1D CNN --&gt; 1D Max Pool --&gt; GRU --&gt; Dense --&gt; Output layer, although I'm not sure the GRU helped at all.</p>",
      "votes": null,
      "replies": [
        {
          "id": 547032,
          "author_name": "friedchips",
          "author_url": "",
          "post_date": "06/07/2019 07:13:18",
          "content": "<p>Very interesting. Your method and mine are really quite similar. as written in some comments above, the standard deviation or variance and the power density are practically the same.\nthe fact that you managed to get a significantly better score even than the 2.40 that I managed confirms my suspicion that using a neural network on the original signal (squared in one way or another and sufficiently down sampled) might in the end be the best approach for predicting new and really unknown test data.\nWould be interesting to know if your superior result is due to the different data preparation or due to your superior network topology and parameter choice...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547233,
          "author_name": "atrus619",
          "author_url": "",
          "post_date": "06/07/2019 12:31:07",
          "content": "<p>I am planning on cleaning up the code a bit and posting it on Github. I'll be sure to let you know so you can take a look! My code is separated into several files, so not feeling like a kernel is the best route.</p>\n\n<p>I think the difference is likely due to a difference in my data generators. I fed the processed data in a gigantic 1D chunk to the generators, and the generators randomly decided where to start/stop in a stratified manner along the CV folds (Used a custom CV split that split along entire EQs, but was as balanced as possible based on certain basic features, like total EQ length). It also had a method that prevented me from training on any EQs that spanned between two consecutive EQs. I also threw out a few of the earthquakes (like the infamous 2 EQs for the price of 1 in the train set).</p>\n\n<p>I by no means optimized the architecture. I spent a ton of time optimizing the architecture at the beginning, only to have very little gains, and threw out hundreds of lines of code to end up with the relatively simple network that seemed to do just as well as the others.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 548009,
          "author_name": "atrus619",
          "author_url": "",
          "post_date": "06/08/2019 16:15:58",
          "content": "<p>Here is the link as promised, let me know if you have any questions! Always down to chat deep learning ^^\n<a href=\"https://github.com/Atrus619/Top-Scoring-Kaggle-Model-LANL-EQ-Prediction-4th-Place-\">https://github.com/Atrus619/Top-Scoring-Kaggle-Model-LANL-EQ-Prediction-4th-Place-</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "544165": "Hi,\nanother 'lessons learned' story: I had played around a lot with CNNs working directly on the acoustic data, and discarded the idea when I couldn't get below 1.6 on the public LB. Turns out the private LB would have been 2.40 (top 25) *without any 'leak' information and without generating a single feature*. [I've published the kernel here](https://www.kaggle.com/friedchips/simple-cnn-would-have-been-top-25), as it is different from all other approaches I've seen here.\nLessons:\n- Physics can't be beat. If anybody would have found a way to predict ttf (even ttf to the next major *or* minor quake!) correctly for individual quake cycles, that would have probably ended in a Nobel Prize. Not joking, as a physicist myself, I know that this would change statistical physics fundamentally...\n- Therefore, success is determined by averaging the cycles in the most intelligent way. This is why the leak information was so powerful :)\n- the power density used in my kernel is probably the most important physical quantity in this process. It actually corresponds to the energy released by the slips as can be seen in this plot. All the minor and major quakes in train are clearly visible:\n![power](https://i.imgur.com/x7KLxmK.png)\n\nI'd be interested to hear if anyone else has played around with the power density?",
    "544185": "Nice approach.  I used CNN with straight TTF data or TTF plus std, got priLB 2.58 ,  pubLB 1.625.  Resnet 2.53, 1.747.",
    "544187": "A lot of models would have worked. I have a KNN that scores 2.382 and would claim a gold medal and the 18th rank here, see attached picture.  It does not use any test info.\n\n![knn](https://storage.googleapis.com/kaggle-forum-message-attachments/544187/13408/knn.png)",
    "544191": "Thanks for your reply! You mean \"straight acoustic data\", right? Did you use a full 150,000 point window as input? My main motivation for using the power density was the ability to downsample a window from 150,000 to 300 points without losing accuracy.",
    "544193": "How do you generate this plot? Actually the nature paper in the introduction thread clearly stated that acoustic power in the early stage of the cycle can predict TTF. I couldn't reproduce it. Your plot looks like what the paper said.",
    "544196": "Yes, that's definitely true. The reason why I though this CNN might be of interest is because it doesn't use any aggregated features at all, only the (squared and downsampled) acoustic amplitude directly. There had been several discussion posts during the competition that said neural networks would perform poorly if used on the acoustic data directly.",
    "544207": "If there's interest I will write a kernel that shows all of my acoustic power (AP) calculations. Short answer: this is the integral (np.cumsum) of the acoustic power in train minus the linear part to show the true power released by the slips.\nAs for the Nature paper, that is actually my main contention: you can tell from the AP approximately *where* you are in the cycle (think of it as % of the span between 2 quakes, the main remaining problem here are the minor quakes which reset the process a bit but don't count), but there is *no way* to tell whether it's going to be a 8 sec. or a 16 sec. quake cycle. Anything else would require a modification of physics as we know it.",
    "544215": "You're right, sorry if I appear to downplay what you have.  I will certainly look at what you shared!",
    "544229": "No problem at all! The reason why I'm doing this (apart from valuable internet points :) ) is that, based on my physics knowledge, I believe that this model \"knows all there is to know\" about the data. I suspect that someone more knowledgeable in neural nets than me would be able to to push this quite a bit further, maybe even to a (late) #1 result. My specialty is model building, machine learning is something I'm just picking up right now :). Turning to the CHAMPS molecular NMR challenge now...",
    "544254": "&gt; Turning to the CHAMPS molecular NMR challenge now…\n\nI'll meet you there soon!",
    "544297": "Isn't var close to power density?  4 of my 12 features are var based, and many teams used similar features.\n\nAlso, you start from (x - x_mean)**2 hence you get std for free.  Saying you learn it is a bit of an exaggeration.  Am I missing something?\n\nOther than that it is a simple yet effective approach.",
    "544382": "You are absolutely right. Power density is closely related to variance. The average power density over a sample window is identical to the signal variance in that window. But the neural network doesn't know that value. The power density is basically a kind of \"instantaneous\" variance. When I downsample the signal by a factor of 500 as done in my kernel, I am feeding the CNN a time series of 300 \"instantaneous variance\" measurements. Each one of them would be useless as a feature by itself, it is only their time sequence which contains any information about ttf and the quake cycle. This is why I said that the CNN learns about the variance/std.\nAs I'm still learning about neural networks, please correct me if this view is wrong. Thanks!",
    "544403": "&gt; The average power density over a sample window is identical to the signal variance in that window.\n\nI agree, but NN work by alternating linear combination of input and activation.  Mean is the simplest linear combination, and a NN with a single unit can learn it ;)\n\nDon't get me wrong, it is a great idea to replace signal by power density.  A simple transfrom that makes CNN effective.",
    "544447": "Got it. Of course you're right. Thanks!",
    "544519": "I down sample seq length to 15000 or 30000 with averaging.  I did 1st difference and 2nd difference after time shift, that didn't work for getting train_loss and val_loss down, although it could work for this competition.  I didn't submit anything.    I got CNN_LSTM and Inception models' Val_loss around 0.5-06.  The models also get similar mae on self_splitted  test sets comparing to val_mae.  But these models didn't do well at both LBs.   Since test data is rather different from training data.  With larger train data set,  these models could definitely learn the pattern very well.  A posted Wavenet_LSTM  seems to perform even better.",
    "544525": "Did you down sample through averaging to 300 as comparison ?",
    "544530": "Markus.  NNs can be very different.  I definitely think some of them work very well.  It's really currently best methods for analyzing sequence data.  This competition only contains 15 earthquakes in the training data.  These models overfit 15 earthquakes, don't generalize well.  With larger training data set, these methods can definitely learn the patterns.",
    "544676": "I tried several different downsampling factors and found to my own surprise that I could go as high as 500 without losing any of the information contained in the signal. The important thing was using the squared signal, because otherwise the information would have been lost in the downsampling process.\n\nThis seems to fit well with the observation made by many other commenters that the frequency spectrum is not helpful at all. The relevant information seems to be in the shape of the peaks themselves.",
    "544682": "In my experience this dataset and its size actually works quite well. of course, you might argue that a neural network is not really needed here, and you would be right. Tree-based methods are more than enough to get top level results. The important thing that I did in my kernel was to stop early after about 16 epochs to make sure that the model doesn't overfit to the different quake cycle lengths and therefore loses its predictive power for the test set.",
    "544687": "Markus,\n\nThis is  interesting - I just posted a thread asking if there were any cnn's that scored well (and embarked on a mission to prove it could be done). I just as quickly removed the thread when I saw this. I think many gave up on the cnn when they saw the poor LB scores - something to remember in the next competition!\n\nOne thing I would push back on - the necessity of the use of power density.  In a cnn, the convolutional filters can easily handle an oscillating signal - they work on them all the time and there is no zeroing of the result because the filters themselves oscillate. And, when you directly downsample to a few hundred samples, surely you are losing information per Nyquist theorem (In my nets, I resample 150K to 37500 in prepossessing in the assumption that the data from 500K to 2 MHz has not much info). \n\nProgressive convolutions with stride, not max pool, can successfully reduce the number of samples from 37500 in an information-preserving way - in these layers, the cnn simply supplies a cost-effective-but-linear way to preserve the info. With enough reduction, perhaps a layer could be added to convert to power and non-linearity added to the remainder of the net?\n\nI still intend to try this, even though the data set is pretty poor.",
    "546011": "Hi Pete,\nyou are right, it isn't *necessary* to use the power density. I should have been more specific: it makes it a lot easier to downsample, which I definitely wanted to do.\n\nAs you said, there is only noise above 500 kHz. So you can downsample the original acoustic data by 4x without losing any information. However, when you downsample further, you get into the frequency range where strong signals exist and therefore peaks will start to cancel each other and disappear. Using the power density eliminates this problem. I wanted to downsample further for two reasons:\n- I looked at the frequency spectrum for all samples and saw that the spectrum shape is practically uncorrelated with TTF. This fits well with the observations in the forum as well as in the organizer's publications that spectrum-based features are not very useful. So there was no motivation for me to use a 10,000+ pt input layer.\n- I realized that I could go down to 500x downsampling without losing accuracy (I started with 10x). That's only 300 pts in the input layer. My NN now was a lot faster. Also, I could now try LSTM and GRU as well, which I did. They actually work almost as well as CNN, but I never tried this any further.\n\n\n\n\n\n\nIn summary, I used the power density because I think that there is no information in the oscillations of the signal. The only thing that matters are the peaks, their overall shape, height, frequency (which is ~ 1 kHz max)) / distance.",
    "546940": "I was in exactly the same boat. I had newly learned how to use deep learning and tensorflow, and was under the \"when all you have is a hammer, everything looks like a nail\" and was so convinced that neural nets would be the key to this competition. Much to my dismay, the public leaderboard scores were crap even though I had a fairly solid CV strategy. Turns out a 1.86 public leaderboard would have resulted in a 4th place finish with 2.29 private leaderboard. Big lessons learned here.\n\nPreprocessing step: Rolling 50 standard deviation, which is then downsampled by a factor of 10. I then take the log of these values to make it more approximately normally distributed, and then I subtract mean and standard deviation in order to center/scale. Each input is of size 14995 as a result.\n\nNetwork architecture: 1D CNN --&gt; 1D Max Pool --&gt; GRU --&gt; Dense --&gt; Output layer, although I'm not sure the GRU helped at all.",
    "547032": "Very interesting. Your method and mine are really quite similar. as written in some comments above, the standard deviation or variance and the power density are practically the same.\nthe fact that you managed to get a significantly better score even than the 2.40 that I managed confirms my suspicion that using a neural network on the original signal (squared in one way or another and sufficiently down sampled) might in the end be the best approach for predicting new and really unknown test data.\nWould be interesting to know if your superior result is due to the different data preparation or due to your superior network topology and parameter choice...",
    "547233": "I am planning on cleaning up the code a bit and posting it on Github. I'll be sure to let you know so you can take a look! My code is separated into several files, so not feeling like a kernel is the best route.\n\nI think the difference is likely due to a difference in my data generators. I fed the processed data in a gigantic 1D chunk to the generators, and the generators randomly decided where to start/stop in a stratified manner along the CV folds (Used a custom CV split that split along entire EQs, but was as balanced as possible based on certain basic features, like total EQ length). It also had a method that prevented me from training on any EQs that spanned between two consecutive EQs. I also threw out a few of the earthquakes (like the infamous 2 EQs for the price of 1 in the train set).\n\nI by no means optimized the architecture. I spent a ton of time optimizing the architecture at the beginning, only to have very little gains, and threw out hundreds of lines of code to end up with the relatively simple network that seemed to do just as well as the others.",
    "548009": "Here is the link as promised, let me know if you have any questions! Always down to chat deep learning ^^\nhttps://github.com/Atrus619/Top-Scoring-Kaggle-Model-LANL-EQ-Prediction-4th-Place-"
  },
  "source": "meta"
}