{
  "id": 312458,
  "title": "Resampled 16k Dataset",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/312458",
  "author_name": "",
  "post_date": "2022-03-12T05:36:12.626435400Z",
  "votes": 12,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I like working with wav files more so I resampled the ogg files to wav files using ffmpeg.</p>\n<pre><code>Mono channel\nSample rate: 16000\nExtension: wav\n</code></pre>\n<p>If anyone thinks this will be useful, here is the <a href=\"https://www.kaggle.com/harveenchadha/pogmusicclassification\" target=\"_blank\">link</a> to the dataset.</p>\n<p>ffmpeg was used to resample the audio with the following command:</p>\n<pre><code>ffmpeg  -i input_filename -ar 16000 -ac 1 -bits_per_raw_sample 16 -vn output_filename\n</code></pre>",
  "messages": [
    {
      "id": "1719780",
      "postDate": "03/12/2022 05:36:12",
      "content": "<p>I like working with wav files more so I resampled the ogg files to wav files using ffmpeg.</p>\n<pre><code>Mono channel\nSample rate: 16000\nExtension: wav\n</code></pre>\n<p>If anyone thinks this will be useful, here is the <a href=\"https://www.kaggle.com/harveenchadha/pogmusicclassification\" target=\"_blank\">link</a> to the dataset.</p>\n<p>ffmpeg was used to resample the audio with the following command:</p>\n<pre><code>ffmpeg  -i input_filename -ar 16000 -ac 1 -bits_per_raw_sample 16 -vn output_filename\n</code></pre>",
      "rawMarkdown": "I like working with wav files more so I resampled the ogg files to wav files using ffmpeg.\n\n```\nMono channel\nSample rate: 16000\nExtension: wav\n```\n\nIf anyone thinks this will be useful, here is the [link](https://www.kaggle.com/harveenchadha/pogmusicclassification) to the dataset.\n\nffmpeg was used to resample the audio with the following command:\n\n```\nffmpeg  -i input_filename -ar 16000 -ac 1 -bits_per_raw_sample 16 -vn output_filename\n```",
      "votes": null
    },
    {
      "id": "1720618",
      "postDate": "03/13/2022 01:17:59",
      "content": "<p>Hi. Thanks for you work</p>",
      "rawMarkdown": "Hi. Thanks for you work",
      "votes": null
    },
    {
      "id": "1721338",
      "postDate": "03/13/2022 15:43:47",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> . I'm curious what the benefit is of having the data in wav format.</p>",
      "rawMarkdown": "Nice work @harveenchadha . I'm curious what the benefit is of having the data in wav format.",
      "votes": null
    },
    {
      "id": "1721373",
      "postDate": "03/13/2022 16:13:25",
      "content": "<p>When I started this I was thinking from the perspective of having audio baselines using CNN's, RNN's and Transformers.This is an excellent venture to try out different things that work for Music (and not just audio). </p>\n<p>All the modern day Transformer models like wav2vec2, hubert (for embeddings) unfortunately work at 16000 sample rate and love wav format to work with (<a href=\"https://arxiv.org/abs/2006.11477)\" target=\"_blank\">https://arxiv.org/abs/2006.11477)</a>, hence this was the first step for that pipeline. </p>\n<p>Why this experiment is interesting for me : I have seen wav2vec2 work brilliantly with noisy audio but struggle with music like beats, drums, rhythm, noise so was thinking of self supervised approach of pretraining a transformer using contrastive loss and then finetune it for music classification task to see if wav2vec2 can be extended to music or not. [ if will have compute to do the same :) ]</p>",
      "rawMarkdown": "When I started this I was thinking from the perspective of having audio baselines using CNN's, RNN's and Transformers.This is an excellent venture to try out different things that work for Music (and not just audio). \n\nAll the modern day Transformer models like wav2vec2, hubert (for embeddings) unfortunately work at 16000 sample rate and love wav format to work with (https://arxiv.org/abs/2006.11477), hence this was the first step for that pipeline. \n\nWhy this experiment is interesting for me : I have seen wav2vec2 work brilliantly with noisy audio but struggle with music like beats, drums, rhythm, noise so was thinking of self supervised approach of pretraining a transformer using contrastive loss and then finetune it for music classification task to see if wav2vec2 can be extended to music or not. [ if will have compute to do the same :) ]",
      "votes": null
    },
    {
      "id": "1721399",
      "postDate": "03/13/2022 16:51:01",
      "content": "<p>So cool! I'm learning so much already. That totally makes sense that you would convert to the standard sample rate.</p>",
      "rawMarkdown": "So cool! I'm learning so much already. That totally makes sense that you would convert to the standard sample rate.",
      "votes": null
    },
    {
      "id": "1724560",
      "postDate": "03/16/2022 10:26:45",
      "content": "<p>Thanks a lot !!! It's super helpful</p>",
      "rawMarkdown": "Thanks a lot !!! It's super helpful",
      "votes": null
    },
    {
      "id": "1724780",
      "postDate": "03/16/2022 13:57:01",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a>, this is super helpful . Does mono means you've considered input's 1 channel only?</p>",
      "rawMarkdown": "Thanks a lot @harveenchadha, this is super helpful . Does mono means you've considered input's 1 channel only?",
      "votes": null
    },
    {
      "id": "1724810",
      "postDate": "03/16/2022 14:27:23",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> I have updated the command in description which was used to resample the audio.</p>\n<p>Also, I checked how is ffmpeg doing the resampling internally, it takes signal from both the channels [and maybe average it out]. Although I am not sure about the part written in brackets.</p>\n<p>Here is a <a href=\"https://trac.ffmpeg.org/wiki/AudioChannelManipulation#stereomonostream\" target=\"_blank\">link</a> which explains the process of conversion from stereo to mono.</p>",
      "rawMarkdown": "pheadrus I have updated the command in description which was used to resample the audio.\n\nAlso, I checked how is ffmpeg doing the resampling internally, it takes signal from both the channels [and maybe average it out]. Although I am not sure about the part written in brackets.\n\nHere is a [link](https://trac.ffmpeg.org/wiki/AudioChannelManipulation#stereomonostream) which explains the process of conversion from stereo to mono.",
      "votes": null
    },
    {
      "id": "1725419",
      "postDate": "03/17/2022 04:44:10",
      "content": "<p>Got it, I did the same thing but used torchaudio for resampling to 16k. Earlier I was resampling within the pipeline, but then saved everything as np array. I kept 2 channels though. </p>",
      "rawMarkdown": "Got it, I did the same thing but used torchaudio for resampling to 16k. Earlier I was resampling within the pipeline, but then saved everything as np array. I kept 2 channels though.",
      "votes": null
    },
    {
      "id": "1725530",
      "postDate": "03/17/2022 07:18:56",
      "content": "<p>You can take average from both the channels to convert into mono audio. Librosa gives you output like that.</p>",
      "rawMarkdown": "You can take average from both the channels to convert into mono audio. Librosa gives you output like that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1720618,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "03/13/2022 01:17:59",
      "content": "<p>Hi. Thanks for you work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1721338,
      "author_name": "robikscube",
      "author_url": "",
      "post_date": "03/13/2022 15:43:47",
      "content": "<p>Nice work <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> . I'm curious what the benefit is of having the data in wav format.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1721373,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "03/13/2022 16:13:25",
          "content": "<p>When I started this I was thinking from the perspective of having audio baselines using CNN's, RNN's and Transformers.This is an excellent venture to try out different things that work for Music (and not just audio). </p>\n<p>All the modern day Transformer models like wav2vec2, hubert (for embeddings) unfortunately work at 16000 sample rate and love wav format to work with (<a href=\"https://arxiv.org/abs/2006.11477)\" target=\"_blank\">https://arxiv.org/abs/2006.11477)</a>, hence this was the first step for that pipeline. </p>\n<p>Why this experiment is interesting for me : I have seen wav2vec2 work brilliantly with noisy audio but struggle with music like beats, drums, rhythm, noise so was thinking of self supervised approach of pretraining a transformer using contrastive loss and then finetune it for music classification task to see if wav2vec2 can be extended to music or not. [ if will have compute to do the same :) ]</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721399,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/13/2022 16:51:01",
          "content": "<p>So cool! I'm learning so much already. That totally makes sense that you would convert to the standard sample rate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1724560,
      "author_name": "dienhoa",
      "author_url": "",
      "post_date": "03/16/2022 10:26:45",
      "content": "<p>Thanks a lot !!! It's super helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1724780,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "03/16/2022 13:57:01",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a>, this is super helpful . Does mono means you've considered input's 1 channel only?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1724810,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "03/16/2022 14:27:23",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> I have updated the command in description which was used to resample the audio.</p>\n<p>Also, I checked how is ffmpeg doing the resampling internally, it takes signal from both the channels [and maybe average it out]. Although I am not sure about the part written in brackets.</p>\n<p>Here is a <a href=\"https://trac.ffmpeg.org/wiki/AudioChannelManipulation#stereomonostream\" target=\"_blank\">link</a> which explains the process of conversion from stereo to mono.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725419,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/17/2022 04:44:10",
          "content": "<p>Got it, I did the same thing but used torchaudio for resampling to 16k. Earlier I was resampling within the pipeline, but then saved everything as np array. I kept 2 channels though. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725530,
          "author_name": "prateeksohlot",
          "author_url": "",
          "post_date": "03/17/2022 07:18:56",
          "content": "<p>You can take average from both the channels to convert into mono audio. Librosa gives you output like that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1719780": "I like working with wav files more so I resampled the ogg files to wav files using ffmpeg.\n\n```\nMono channel\nSample rate: 16000\nExtension: wav\n```\n\nIf anyone thinks this will be useful, here is the [link](https://www.kaggle.com/harveenchadha/pogmusicclassification) to the dataset.\n\nffmpeg was used to resample the audio with the following command:\n\n```\nffmpeg  -i input_filename -ar 16000 -ac 1 -bits_per_raw_sample 16 -vn output_filename\n```",
    "1720618": "Hi. Thanks for you work",
    "1721338": "Nice work @harveenchadha . I'm curious what the benefit is of having the data in wav format.",
    "1721373": "When I started this I was thinking from the perspective of having audio baselines using CNN's, RNN's and Transformers.This is an excellent venture to try out different things that work for Music (and not just audio). \n\nAll the modern day Transformer models like wav2vec2, hubert (for embeddings) unfortunately work at 16000 sample rate and love wav format to work with (https://arxiv.org/abs/2006.11477), hence this was the first step for that pipeline. \n\nWhy this experiment is interesting for me : I have seen wav2vec2 work brilliantly with noisy audio but struggle with music like beats, drums, rhythm, noise so was thinking of self supervised approach of pretraining a transformer using contrastive loss and then finetune it for music classification task to see if wav2vec2 can be extended to music or not. [ if will have compute to do the same :) ]",
    "1721399": "So cool! I'm learning so much already. That totally makes sense that you would convert to the standard sample rate.",
    "1724560": "Thanks a lot !!! It's super helpful",
    "1724780": "Thanks a lot @harveenchadha, this is super helpful . Does mono means you've considered input's 1 channel only?",
    "1724810": "pheadrus I have updated the command in description which was used to resample the audio.\n\nAlso, I checked how is ffmpeg doing the resampling internally, it takes signal from both the channels [and maybe average it out]. Although I am not sure about the part written in brackets.\n\nHere is a [link](https://trac.ffmpeg.org/wiki/AudioChannelManipulation#stereomonostream) which explains the process of conversion from stereo to mono.",
    "1725419": "Got it, I did the same thing but used torchaudio for resampling to 16k. Earlier I was resampling within the pipeline, but then saved everything as np array. I kept 2 channels though.",
    "1725530": "You can take average from both the channels to convert into mono audio. Librosa gives you output like that."
  },
  "source": "meta"
}