{
  "id": 89328,
  "title": "Vectorizing those eartquakes",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89328",
  "author_name": "",
  "post_date": "2019-04-13T02:36:57.388611600Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all - I'm a newbie and I would like to have your input please. in my mind, these acoustic signals are waves, and upon the completion of 150000 recordings of those signals, a quake will happen after X seconds. it sounds to me as if those waves (seen as window of 150K to match their size in the testing set) should be vectorized before being ingested into some ML magic. \nI have tried using some existing Kernels and simple adding a lot more features - things I didnt understand. Quantile reads at every percent incresaes, rolls, skew, and other functions I'm not at east with because I dont understand what's happening under the hood. Some submissions are better, others are worse. Overlapping the window reading is sometimes betteres, often worst.</p>\n\n<p>My LSTM architecture in Keras proved to be less significant than an XGBoost run ( I didnt even search the space for the right parameter). But I am wondering, would it make sense to try to convert this acoustic signal, and treat it as an acoustic wave. Does the machine work at a specific well know frequency? Maybe afterwards, some fourier analysis can be performed? I've read a paper where researchers were taking this approach to predict volcano erruptions.</p>\n\n<p>Feeding 150K lines into an RNN doesnt make sense, the vector is too scarce and not condensed enough.\nFeeding stats like mean and sd make no sense as well to me - what do they represent really? and how do we take into consideration the recurrence in learning, since we're lucky enough to be given a continuous file (and not 150K chunks). I believe that overlap of strides is necessary, and they should overlap at 4K (since it seems, this is when the time value drastically drops 0.0001 seconds). As a proof you can visualize time_to_failure between 4000 and 4200.</p>\n\n<p>I dont think I will get close to winning this, but it doesnt matter since I'm learning an endless amount of new stuff (and I'm trying to do it all in R :)) </p>\n\n<p>Thanks for sharing your thoughts and providing some course-correction if i've said something silly\nCheers, Eyas</p>",
  "messages": [
    {
      "id": "515717",
      "postDate": "04/13/2019 02:36:57",
      "content": "<p>Hi all - I'm a newbie and I would like to have your input please. in my mind, these acoustic signals are waves, and upon the completion of 150000 recordings of those signals, a quake will happen after X seconds. it sounds to me as if those waves (seen as window of 150K to match their size in the testing set) should be vectorized before being ingested into some ML magic. \nI have tried using some existing Kernels and simple adding a lot more features - things I didnt understand. Quantile reads at every percent incresaes, rolls, skew, and other functions I'm not at east with because I dont understand what's happening under the hood. Some submissions are better, others are worse. Overlapping the window reading is sometimes betteres, often worst.</p>\n\n<p>My LSTM architecture in Keras proved to be less significant than an XGBoost run ( I didnt even search the space for the right parameter). But I am wondering, would it make sense to try to convert this acoustic signal, and treat it as an acoustic wave. Does the machine work at a specific well know frequency? Maybe afterwards, some fourier analysis can be performed? I've read a paper where researchers were taking this approach to predict volcano erruptions.</p>\n\n<p>Feeding 150K lines into an RNN doesnt make sense, the vector is too scarce and not condensed enough.\nFeeding stats like mean and sd make no sense as well to me - what do they represent really? and how do we take into consideration the recurrence in learning, since we're lucky enough to be given a continuous file (and not 150K chunks). I believe that overlap of strides is necessary, and they should overlap at 4K (since it seems, this is when the time value drastically drops 0.0001 seconds). As a proof you can visualize time_to_failure between 4000 and 4200.</p>\n\n<p>I dont think I will get close to winning this, but it doesnt matter since I'm learning an endless amount of new stuff (and I'm trying to do it all in R :)) </p>\n\n<p>Thanks for sharing your thoughts and providing some course-correction if i've said something silly\nCheers, Eyas</p>",
      "rawMarkdown": "Hi all - I'm a newbie and I would like to have your input please. in my mind, these acoustic signals are waves, and upon the completion of 150000 recordings of those signals, a quake will happen after X seconds. it sounds to me as if those waves (seen as window of 150K to match their size in the testing set) should be vectorized before being ingested into some ML magic. \nI have tried using some existing Kernels and simple adding a lot more features - things I didnt understand. Quantile reads at every percent incresaes, rolls, skew, and other functions I'm not at east with because I dont understand what's happening under the hood. Some submissions are better, others are worse. Overlapping the window reading is sometimes betteres, often worst.\n\nMy LSTM architecture in Keras proved to be less significant than an XGBoost run ( I didnt even search the space for the right parameter). But I am wondering, would it make sense to try to convert this acoustic signal, and treat it as an acoustic wave. Does the machine work at a specific well know frequency? Maybe afterwards, some fourier analysis can be performed? I've read a paper where researchers were taking this approach to predict volcano erruptions.\n\nFeeding 150K lines into an RNN doesnt make sense, the vector is too scarce and not condensed enough.\nFeeding stats like mean and sd make no sense as well to me - what do they represent really? and how do we take into consideration the recurrence in learning, since we're lucky enough to be given a continuous file (and not 150K chunks). I believe that overlap of strides is necessary, and they should overlap at 4K (since it seems, this is when the time value drastically drops 0.0001 seconds). As a proof you can visualize time_to_failure between 4000 and 4200.\n\nI dont think I will get close to winning this, but it doesnt matter since I'm learning an endless amount of new stuff (and I'm trying to do it all in R :)) \n\nThanks for sharing your thoughts and providing some course-correction if i've said something silly\nCheers, Eyas",
      "votes": null
    },
    {
      "id": "516029",
      "postDate": "04/13/2019 14:39:05",
      "content": "<p>Hi Eyas, good luck with your exploration journey.</p>\n\n<p>Just yesterday I have <a href=\"https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89\">read something</a> which resembles me your line of thought. They used FFT in rolling window of constant size to create an image, which they then feed to regular CNN. I hope I will have time to try it.</p>",
      "rawMarkdown": "Hi Eyas, good luck with your exploration journey.\n\nJust yesterday I have [read something](https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89) which resembles me your line of thought. They used FFT in rolling window of constant size to create an image, which they then feed to regular CNN. I hope I will have time to try it.",
      "votes": null
    },
    {
      "id": "521961",
      "postDate": "04/23/2019 17:27:55",
      "content": "<p>That is exactly my approach. I did not have time to try it but i'm working on it. Frequential analysis can simplify the job for the CNN (reduce the number of layers, therefore help to easily extract meaningful features)</p>",
      "rawMarkdown": "That is exactly my approach. I did not have time to try it but i'm working on it. Frequential analysis can simplify the job for the CNN (reduce the number of layers, therefore help to easily extract meaningful features)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 516029,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "04/13/2019 14:39:05",
      "content": "<p>Hi Eyas, good luck with your exploration journey.</p>\n\n<p>Just yesterday I have <a href=\"https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89\">read something</a> which resembles me your line of thought. They used FFT in rolling window of constant size to create an image, which they then feed to regular CNN. I hope I will have time to try it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 521961,
          "author_name": "doctorsnake",
          "author_url": "",
          "post_date": "04/23/2019 17:27:55",
          "content": "<p>That is exactly my approach. I did not have time to try it but i'm working on it. Frequential analysis can simplify the job for the CNN (reduce the number of layers, therefore help to easily extract meaningful features)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "515717": "Hi all - I'm a newbie and I would like to have your input please. in my mind, these acoustic signals are waves, and upon the completion of 150000 recordings of those signals, a quake will happen after X seconds. it sounds to me as if those waves (seen as window of 150K to match their size in the testing set) should be vectorized before being ingested into some ML magic. \nI have tried using some existing Kernels and simple adding a lot more features - things I didnt understand. Quantile reads at every percent incresaes, rolls, skew, and other functions I'm not at east with because I dont understand what's happening under the hood. Some submissions are better, others are worse. Overlapping the window reading is sometimes betteres, often worst.\n\nMy LSTM architecture in Keras proved to be less significant than an XGBoost run ( I didnt even search the space for the right parameter). But I am wondering, would it make sense to try to convert this acoustic signal, and treat it as an acoustic wave. Does the machine work at a specific well know frequency? Maybe afterwards, some fourier analysis can be performed? I've read a paper where researchers were taking this approach to predict volcano erruptions.\n\nFeeding 150K lines into an RNN doesnt make sense, the vector is too scarce and not condensed enough.\nFeeding stats like mean and sd make no sense as well to me - what do they represent really? and how do we take into consideration the recurrence in learning, since we're lucky enough to be given a continuous file (and not 150K chunks). I believe that overlap of strides is necessary, and they should overlap at 4K (since it seems, this is when the time value drastically drops 0.0001 seconds). As a proof you can visualize time_to_failure between 4000 and 4200.\n\nI dont think I will get close to winning this, but it doesnt matter since I'm learning an endless amount of new stuff (and I'm trying to do it all in R :)) \n\nThanks for sharing your thoughts and providing some course-correction if i've said something silly\nCheers, Eyas",
    "516029": "Hi Eyas, good luck with your exploration journey.\n\nJust yesterday I have [read something](https://towardsdatascience.com/audio-classification-using-fastai-and-on-the-fly-frequency-transforms-4dbe1b540f89) which resembles me your line of thought. They used FFT in rolling window of constant size to create an image, which they then feed to regular CNN. I hope I will have time to try it.",
    "521961": "That is exactly my approach. I did not have time to try it but i'm working on it. Frequential analysis can simplify the job for the CNN (reduce the number of layers, therefore help to easily extract meaningful features)"
  },
  "source": "meta"
}