{
  "id": 91776,
  "title": "Earthquake melspectrogram images dataset",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91776",
  "author_name": "",
  "post_date": "2019-05-09T00:20:10.776067Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>A few weeks ago I wanted to to learn more about the fastai framework and was going through the tutorial videos. I created a bunch of melspectrograms of the LANL dataset using librosa, and trained on the images. Recently I've abandoned this approach since I've been working more on GBM models, but I figure others might have an interest in using them- so I've uploaded them as a public dataset.</p>\n\n<p>I also created a kernel with basic starter code with no tuning ~1.721 LB. I'm sure someone more skilled with NNs could do better.</p>\n\n<p><a href=\"https://www.kaggle.com/robikscube/lanl-earthquake-spectrogram-images\">https://www.kaggle.com/robikscube/lanl-earthquake-spectrogram-images</a>\n<a href=\"https://www.kaggle.com/robikscube/lanl-earthquake-melspectrogram-images-fastai-nn\">https://www.kaggle.com/robikscube/lanl-earthquake-melspectrogram-images-fastai-nn</a></p>",
  "messages": [
    {
      "id": "528960",
      "postDate": "05/09/2019 00:20:10",
      "content": "<p>A few weeks ago I wanted to to learn more about the fastai framework and was going through the tutorial videos. I created a bunch of melspectrograms of the LANL dataset using librosa, and trained on the images. Recently I've abandoned this approach since I've been working more on GBM models, but I figure others might have an interest in using them- so I've uploaded them as a public dataset.</p>\n\n<p>I also created a kernel with basic starter code with no tuning ~1.721 LB. I'm sure someone more skilled with NNs could do better.</p>\n\n<p><a href=\"https://www.kaggle.com/robikscube/lanl-earthquake-spectrogram-images\">https://www.kaggle.com/robikscube/lanl-earthquake-spectrogram-images</a>\n<a href=\"https://www.kaggle.com/robikscube/lanl-earthquake-melspectrogram-images-fastai-nn\">https://www.kaggle.com/robikscube/lanl-earthquake-melspectrogram-images-fastai-nn</a></p>",
      "rawMarkdown": "A few weeks ago I wanted to to learn more about the fastai framework and was going through the tutorial videos. I created a bunch of melspectrograms of the LANL dataset using librosa, and trained on the images. Recently I've abandoned this approach since I've been working more on GBM models, but I figure others might have an interest in using them- so I've uploaded them as a public dataset.\n\nI also created a kernel with basic starter code with no tuning ~1.721 LB. I'm sure someone more skilled with NNs could do better.\n\nhttps://www.kaggle.com/robikscube/lanl-earthquake-spectrogram-images\nhttps://www.kaggle.com/robikscube/lanl-earthquake-melspectrogram-images-fastai-nn",
      "votes": null
    },
    {
      "id": "529046",
      "postDate": "05/09/2019 04:33:35",
      "content": "<p>Great data! Thanks!</p>",
      "rawMarkdown": "Great data! Thanks!",
      "votes": null
    },
    {
      "id": "529083",
      "postDate": "05/09/2019 06:54:52",
      "content": "<p>Thanks for the nice work and great sharing! What are the meaning of the x and y axes of the figures? Are they time and frequency, respectively?</p>",
      "rawMarkdown": "Thanks for the nice work and great sharing! What are the meaning of the x and y axes of the figures? Are they time and frequency, respectively?",
      "votes": null
    },
    {
      "id": "529254",
      "postDate": "05/09/2019 14:17:30",
      "content": "<p>Great question! I used the librosa package <code>librosa.feature.melspectrogram</code> (link to docs below). My understanding is the x axis is time and the y axis are the frequency magnitudes mapped to a mel scale.</p>\n\n<p>As I was looking this up I'm realizing mel-scale representation are used commonly for speech, and probably isn't the best way to represent this type of audio- you may have given me an idea of something different I could go back and try! 😄 </p>\n\n<p><a href=\"https://librosa.github.io/librosa/generated/librosa.feature.melspectrogram.html\">https://librosa.github.io/librosa/generated/librosa.feature.melspectrogram.html</a> </p>",
      "rawMarkdown": "Great question! I used the librosa package `librosa.feature.melspectrogram` (link to docs below). My understanding is the x axis is time and the y axis are the frequency magnitudes mapped to a mel scale.\n\nAs I was looking this up I'm realizing mel-scale representation are used commonly for speech, and probably isn't the best way to represent this type of audio- you may have given me an idea of something different I could go back and try! 😄 \n\nhttps://librosa.github.io/librosa/generated/librosa.feature.melspectrogram.html",
      "votes": null
    },
    {
      "id": "529421",
      "postDate": "05/09/2019 21:13:55",
      "content": "<p>Do you lose the phase information?\nAlso: in mel, there are generally hardwired constants that need changing depending on bandwidth.</p>",
      "rawMarkdown": "Do you lose the phase information?\nAlso: in mel, there are generally hardwired constants that need changing depending on bandwidth.",
      "votes": null
    },
    {
      "id": "529437",
      "postDate": "05/09/2019 22:29:34",
      "content": "<p>Interesting. I'm new to learning about melspectrograms. Is there a different type of way to display this type of audio data you think would perform better?</p>\n\n<p>If you look at the notebook I provide the code I used to create the plots. I experimented with different sample rate and number of mels until I got something that struck a good balance between detail and image saturation.</p>\n\n<p><code>\nsr=10000\nn_mels=1000\nS = librosa.feature.melspectrogram(y, sr=sr, n_mels=n_mels)\nS = librosa.power_to_db(S, ref=np.max)\n</code></p>",
      "rawMarkdown": "Interesting. I'm new to learning about melspectrograms. Is there a different type of way to display this type of audio data you think would perform better?\n\nIf you look at the notebook I provide the code I used to create the plots. I experimented with different sample rate and number of mels until I got something that struck a good balance between detail and image saturation.\n\n```\nsr=10000\nn_mels=1000\nS = librosa.feature.melspectrogram(y, sr=sr, n_mels=n_mels)\nS = librosa.power_to_db(S, ref=np.max)\n```",
      "votes": null
    },
    {
      "id": "529457",
      "postDate": "05/10/2019 00:03:50",
      "content": "<p>I'm also new to this. I was playing with it for a while but as I recall I got discouraged by the phase issue. That is a lot of information.</p>",
      "rawMarkdown": "I'm also new to this. I was playing with it for a while but as I recall I got discouraged by the phase issue. That is a lot of information.",
      "votes": null
    },
    {
      "id": "529835",
      "postDate": "05/10/2019 22:48:04",
      "content": "<p>Thanks for sharing!</p>\n\n<p>I also tried to work with librosa's melspectrograms. However, it seems that they apply some manipulations on the images which may lose significant information (e.g. order of magnitude of the signal). I guess it can all be tuned, but since I'm afraid of other unknown blackbox features, I tend to prefer home-maid spectrograms for this competition :)</p>\n\n<p>I wrote some of my observations regarding this issue <a href=\"https://www.kaggle.com/idog90/lanl-competition-why-do-spectrograms-fail\">here</a>.</p>",
      "rawMarkdown": "Thanks for sharing!\n\nI also tried to work with librosa's melspectrograms. However, it seems that they apply some manipulations on the images which may lose significant information (e.g. order of magnitude of the signal). I guess it can all be tuned, but since I'm afraid of other unknown blackbox features, I tend to prefer home-maid spectrograms for this competition :)\n\nI wrote some of my observations regarding this issue [here](https://www.kaggle.com/idog90/lanl-competition-why-do-spectrograms-fail).",
      "votes": null
    },
    {
      "id": "570840",
      "postDate": "07/08/2019 20:28:11",
      "content": "<p>Dear Lenik Terenin!\nI greatly apologize for the off-top, but it is very difficult to contact you, although I have repeatedly tried to do it. I have a huge request for you, as the owner of the site <a href=\"http://www.samosud.org\">http://www.samosud.org</a>. Your resource contains copies of the court decisions: <a href=\"http://www.samosud.org/case_969293378\">http://www.samosud.org/case_969293378</a> and <a href=\"http://www.samosud.org/case_849613734\">http://www.samosud.org/case_849613734</a>, in which I appear as the complainant in the labor dispute. The problem is that while somebody enters my name into the Yandex search box, the links to these pages are in the first 2 positions of the search results. I did not commit any crime and did not inflict physical or material damage to anyone. Nevertheless, the information you control is very troubling me: I work as the leading researcher, teach students and practice phlebology. Currently, this kind of information can play a fatal role in Russia. I request your participation in the removal of the above pages from your site. Help me, please! Thank you and once again I apologize for invading a non-core site. </p>\n\n<p>Sincerely yours, </p>\n\n<p>Yuri Gustelev</p>",
      "rawMarkdown": "Dear Lenik Terenin!\nI greatly apologize for the off-top, but it is very difficult to contact you, although I have repeatedly tried to do it. I have a huge request for you, as the owner of the site http://www.samosud.org. Your resource contains copies of the court decisions: http://www.samosud.org/case_969293378 and http://www.samosud.org/case_849613734, in which I appear as the complainant in the labor dispute. The problem is that while somebody enters my name into the Yandex search box, the links to these pages are in the first 2 positions of the search results. I did not commit any crime and did not inflict physical or material damage to anyone. Nevertheless, the information you control is very troubling me: I work as the leading researcher, teach students and practice phlebology. Currently, this kind of information can play a fatal role in Russia. I request your participation in the removal of the above pages from your site. Help me, please! Thank you and once again I apologize for invading a non-core site. \n\nSincerely yours, \n\nYuri Gustelev",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 529046,
      "author_name": "terenin",
      "author_url": "",
      "post_date": "05/09/2019 04:33:35",
      "content": "<p>Great data! Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 570840,
          "author_name": "cavasurg",
          "author_url": "",
          "post_date": "07/08/2019 20:28:11",
          "content": "<p>Dear Lenik Terenin!\nI greatly apologize for the off-top, but it is very difficult to contact you, although I have repeatedly tried to do it. I have a huge request for you, as the owner of the site <a href=\"http://www.samosud.org\">http://www.samosud.org</a>. Your resource contains copies of the court decisions: <a href=\"http://www.samosud.org/case_969293378\">http://www.samosud.org/case_969293378</a> and <a href=\"http://www.samosud.org/case_849613734\">http://www.samosud.org/case_849613734</a>, in which I appear as the complainant in the labor dispute. The problem is that while somebody enters my name into the Yandex search box, the links to these pages are in the first 2 positions of the search results. I did not commit any crime and did not inflict physical or material damage to anyone. Nevertheless, the information you control is very troubling me: I work as the leading researcher, teach students and practice phlebology. Currently, this kind of information can play a fatal role in Russia. I request your participation in the removal of the above pages from your site. Help me, please! Thank you and once again I apologize for invading a non-core site. </p>\n\n<p>Sincerely yours, </p>\n\n<p>Yuri Gustelev</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529083,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "05/09/2019 06:54:52",
      "content": "<p>Thanks for the nice work and great sharing! What are the meaning of the x and y axes of the figures? Are they time and frequency, respectively?</p>",
      "votes": null,
      "replies": [
        {
          "id": 529254,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "05/09/2019 14:17:30",
          "content": "<p>Great question! I used the librosa package <code>librosa.feature.melspectrogram</code> (link to docs below). My understanding is the x axis is time and the y axis are the frequency magnitudes mapped to a mel scale.</p>\n\n<p>As I was looking this up I'm realizing mel-scale representation are used commonly for speech, and probably isn't the best way to represent this type of audio- you may have given me an idea of something different I could go back and try! 😄 </p>\n\n<p><a href=\"https://librosa.github.io/librosa/generated/librosa.feature.melspectrogram.html\">https://librosa.github.io/librosa/generated/librosa.feature.melspectrogram.html</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529421,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "05/09/2019 21:13:55",
      "content": "<p>Do you lose the phase information?\nAlso: in mel, there are generally hardwired constants that need changing depending on bandwidth.</p>",
      "votes": null,
      "replies": [
        {
          "id": 529437,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "05/09/2019 22:29:34",
          "content": "<p>Interesting. I'm new to learning about melspectrograms. Is there a different type of way to display this type of audio data you think would perform better?</p>\n\n<p>If you look at the notebook I provide the code I used to create the plots. I experimented with different sample rate and number of mels until I got something that struck a good balance between detail and image saturation.</p>\n\n<p><code>\nsr=10000\nn_mels=1000\nS = librosa.feature.melspectrogram(y, sr=sr, n_mels=n_mels)\nS = librosa.power_to_db(S, ref=np.max)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 529457,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "05/10/2019 00:03:50",
          "content": "<p>I'm also new to this. I was playing with it for a while but as I recall I got discouraged by the phase issue. That is a lot of information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 529835,
      "author_name": "idog90",
      "author_url": "",
      "post_date": "05/10/2019 22:48:04",
      "content": "<p>Thanks for sharing!</p>\n\n<p>I also tried to work with librosa's melspectrograms. However, it seems that they apply some manipulations on the images which may lose significant information (e.g. order of magnitude of the signal). I guess it can all be tuned, but since I'm afraid of other unknown blackbox features, I tend to prefer home-maid spectrograms for this competition :)</p>\n\n<p>I wrote some of my observations regarding this issue <a href=\"https://www.kaggle.com/idog90/lanl-competition-why-do-spectrograms-fail\">here</a>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "528960": "A few weeks ago I wanted to to learn more about the fastai framework and was going through the tutorial videos. I created a bunch of melspectrograms of the LANL dataset using librosa, and trained on the images. Recently I've abandoned this approach since I've been working more on GBM models, but I figure others might have an interest in using them- so I've uploaded them as a public dataset.\n\nI also created a kernel with basic starter code with no tuning ~1.721 LB. I'm sure someone more skilled with NNs could do better.\n\nhttps://www.kaggle.com/robikscube/lanl-earthquake-spectrogram-images\nhttps://www.kaggle.com/robikscube/lanl-earthquake-melspectrogram-images-fastai-nn",
    "529046": "Great data! Thanks!",
    "529083": "Thanks for the nice work and great sharing! What are the meaning of the x and y axes of the figures? Are they time and frequency, respectively?",
    "529254": "Great question! I used the librosa package `librosa.feature.melspectrogram` (link to docs below). My understanding is the x axis is time and the y axis are the frequency magnitudes mapped to a mel scale.\n\nAs I was looking this up I'm realizing mel-scale representation are used commonly for speech, and probably isn't the best way to represent this type of audio- you may have given me an idea of something different I could go back and try! 😄 \n\nhttps://librosa.github.io/librosa/generated/librosa.feature.melspectrogram.html",
    "529421": "Do you lose the phase information?\nAlso: in mel, there are generally hardwired constants that need changing depending on bandwidth.",
    "529437": "Interesting. I'm new to learning about melspectrograms. Is there a different type of way to display this type of audio data you think would perform better?\n\nIf you look at the notebook I provide the code I used to create the plots. I experimented with different sample rate and number of mels until I got something that struck a good balance between detail and image saturation.\n\n```\nsr=10000\nn_mels=1000\nS = librosa.feature.melspectrogram(y, sr=sr, n_mels=n_mels)\nS = librosa.power_to_db(S, ref=np.max)\n```",
    "529457": "I'm also new to this. I was playing with it for a while but as I recall I got discouraged by the phase issue. That is a lot of information.",
    "529835": "Thanks for sharing!\n\nI also tried to work with librosa's melspectrograms. However, it seems that they apply some manipulations on the images which may lose significant information (e.g. order of magnitude of the signal). I guess it can all be tuned, but since I'm afraid of other unknown blackbox features, I tend to prefer home-maid spectrograms for this competition :)\n\nI wrote some of my observations regarding this issue [here](https://www.kaggle.com/idog90/lanl-competition-why-do-spectrograms-fail).",
    "570840": "Dear Lenik Terenin!\nI greatly apologize for the off-top, but it is very difficult to contact you, although I have repeatedly tried to do it. I have a huge request for you, as the owner of the site http://www.samosud.org. Your resource contains copies of the court decisions: http://www.samosud.org/case_969293378 and http://www.samosud.org/case_849613734, in which I appear as the complainant in the labor dispute. The problem is that while somebody enters my name into the Yandex search box, the links to these pages are in the first 2 positions of the search results. I did not commit any crime and did not inflict physical or material damage to anyone. Nevertheless, the information you control is very troubling me: I work as the leading researcher, teach students and practice phlebology. Currently, this kind of information can play a fatal role in Russia. I request your participation in the removal of the above pages from your site. Help me, please! Thank you and once again I apologize for invading a non-core site. \n\nSincerely yours, \n\nYuri Gustelev"
  },
  "source": "meta"
}