{
  "id": 308579,
  "title": "Why do some sound files have (X, 2) dim?",
  "url": "/competitions/birdclef-2022/discussion/308579",
  "author_name": "",
  "post_date": "2022-02-19T09:54:03.860589500Z",
  "votes": 21,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I found that some sound files had (X, 2) dim.<br>\nI'm not familiar with it, but is the difference due to the author's recording equipment?<br>\nAlso, is averaging a good way to handle this kind of audio?</p>\n<p><code>data = np.mean(data, 1)</code></p>\n<p><img src=\"https://user-images.githubusercontent.com/58644957/154795828-9bda47c1-bf27-4388-b114-4d45a0d6429a.png\" alt=\"image\"></p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1697007",
      "postDate": "02/19/2022 09:54:03",
      "content": "<p>I found that some sound files had (X, 2) dim.<br>\nI'm not familiar with it, but is the difference due to the author's recording equipment?<br>\nAlso, is averaging a good way to handle this kind of audio?</p>\n<p><code>data = np.mean(data, 1)</code></p>\n<p><img src=\"https://user-images.githubusercontent.com/58644957/154795828-9bda47c1-bf27-4388-b114-4d45a0d6429a.png\" alt=\"image\"></p>\n<p>Thanks.</p>",
      "rawMarkdown": "I found that some sound files had (X, 2) dim.\nI'm not familiar with it, but is the difference due to the author's recording equipment?\nAlso, is averaging a good way to handle this kind of audio?\n\n`data = np.mean(data, 1)`\n\n![image](https://user-images.githubusercontent.com/58644957/154795828-9bda47c1-bf27-4388-b114-4d45a0d6429a.png)\n\nThanks.",
      "votes": null
    },
    {
      "id": "1697035",
      "postDate": "02/19/2022 10:18:20",
      "content": "<p>That means binaural. You can use <code>librosa.to_mono</code> to uniformly convert to mono format.</p>",
      "rawMarkdown": "That means binaural. You can use `librosa.to_mono` to uniformly convert to mono format.",
      "votes": null
    },
    {
      "id": "1697068",
      "postDate": "02/19/2022 10:45:07",
      "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> Thanks for sharing the function, I wasn't aware of it! </p>\n<p>Don't you think we are discarding useful info by converting binaural to mono format? </p>\n<p>I picture it's like converting an RGB image to B/W, throwing away info that models could potentially use. </p>",
      "rawMarkdown": "zzy990106 Thanks for sharing the function, I wasn't aware of it! \n\nDon't you think we are discarding useful info by converting binaural to mono format? \n\nI picture it's like converting an RGB image to B/W, throwing away info that models could potentially use.",
      "votes": null
    },
    {
      "id": "1697071",
      "postDate": "02/19/2022 10:48:55",
      "content": "<p>Yes, there is information loss, but the impact on classification is far negligible.</p>",
      "rawMarkdown": "Yes, there is information loss, but the impact on classification is far negligible.",
      "votes": null
    },
    {
      "id": "1697388",
      "postDate": "02/19/2022 15:38:15",
      "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a>, is there a difference between binaural and stereo sound in this case ?</p>\n<p>Can't we just take the mean of the 2 channels ?</p>\n<p><code>y, sr = sf.read(path, always_2d=True)</code></p>\n<p><code>y = np.mean(y, 1)</code></p>",
      "rawMarkdown": "zzy990106, is there a difference between binaural and stereo sound in this case ?\n\nCan't we just take the mean of the 2 channels ?\n\n`y, sr = sf.read(path, always_2d=True)`\n\n`y = np.mean(y, 1)`",
      "votes": null
    },
    {
      "id": "1697390",
      "postDate": "02/19/2022 15:40:13",
      "content": "<p>Alex, I believe both are the same</p>\n<blockquote>\n  <p>Can't we just take the mean of the 2 channels ?</p>\n</blockquote>\n<p>I think this might be the best approach! I'll try this 👌. Thanks!</p>",
      "rawMarkdown": "Alex, I believe both are the same\n\n> Can't we just take the mean of the 2 channels ?\n\nI think this might be the best approach! I'll try this 👌. Thanks!",
      "votes": null
    },
    {
      "id": "1697899",
      "postDate": "02/20/2022 00:28:04",
      "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a>, Thanks for the answer! </p>",
      "rawMarkdown": "zzy990106, Thanks for the answer!",
      "votes": null
    },
    {
      "id": "1699305",
      "postDate": "02/21/2022 05:14:02",
      "content": "<p>I have experimented with 3 approaches to this problem.</p>\n<ol>\n<li>use only the first channel.</li>\n<li>calculate the mean.</li>\n<li>randomly sample at training which channel to use.</li>\n</ol>\n<p>I thought that approach 3 would work, but the results of LB were as follows</p>\n<table>\n<thead>\n<tr>\n<th>approach</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.70</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.69</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.64</td>\n</tr>\n</tbody>\n</table>\n<p>For reference, as you know, all test audio files seem to have only one channel. I don't know if this was recorded monaural, or if they do some way to make it only one channel…</p>",
      "rawMarkdown": "I have experimented with 3 approaches to this problem.\n\n1. use only the first channel.\n2. calculate the mean.\n3. randomly sample at training which channel to use.\n\nI thought that approach 3 would work, but the results of LB were as follows\n\n| approach | LB |\n| --- | --- |\n| 1 | 0.70 |\n| 2 | 0.69 |\n| 3 | 0.64 |\n\nFor reference, as you know, all test audio files seem to have only one channel. I don't know if this was recorded monaural, or if they do some way to make it only one channel...",
      "votes": null
    },
    {
      "id": "1699522",
      "postDate": "02/21/2022 08:41:28",
      "content": "<p>Could you share how you only use the first channel ?</p>",
      "rawMarkdown": "Could you share how you only use the first channel ?",
      "votes": null
    },
    {
      "id": "1699582",
      "postDate": "02/21/2022 09:41:13",
      "content": "<p>This is an example of assigning the first channel to y.</p>\n<pre><code>y, sr = sf.read(wav_path, always_2d=True)\ny = y[:, 0]  \n</code></pre>",
      "rawMarkdown": "This is an example of assigning the first channel to y.\n\n```\ny, sr = sf.read(wav_path, always_2d=True)\ny = y[:, 0]  \n```",
      "votes": null
    },
    {
      "id": "1701393",
      "postDate": "02/22/2022 18:10:03",
      "content": "<p>I found there are some date that the volume of one microphone is abnormal (for example, XC344134.ogg : Channel 1 broken ,XC521984.ogg : Channel 0 broken).<br>\nIf you normalize the data after transform to spectrogram, it might not be a problem, but be careful when dealing with volumes.</p>",
      "rawMarkdown": "I found there are some date that the volume of one microphone is abnormal (for example, XC344134.ogg : Channel 1 broken ,XC521984.ogg : Channel 0 broken).\nIf you normalize the data after transform to spectrogram, it might not be a problem, but be careful when dealing with volumes.",
      "votes": null
    },
    {
      "id": "1702177",
      "postDate": "02/23/2022 11:35:18",
      "content": "<p><a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@nomorevotch</a> I had same idea as you (randomly select one of the audio channel when training) and my CV is quite lower than when I use channel average.  Kind of confirm your experience.</p>",
      "rawMarkdown": "nomorevotch I had same idea as you (randomly select one of the audio channel when training) and my CV is quite lower than when I use channel average.  Kind of confirm your experience.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1697035,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "02/19/2022 10:18:20",
      "content": "<p>That means binaural. You can use <code>librosa.to_mono</code> to uniformly convert to mono format.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1697068,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/19/2022 10:45:07",
          "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a> Thanks for sharing the function, I wasn't aware of it! </p>\n<p>Don't you think we are discarding useful info by converting binaural to mono format? </p>\n<p>I picture it's like converting an RGB image to B/W, throwing away info that models could potentially use. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697071,
          "author_name": "zzy990106",
          "author_url": "",
          "post_date": "02/19/2022 10:48:55",
          "content": "<p>Yes, there is information loss, but the impact on classification is far negligible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697388,
          "author_name": "alexandrecc",
          "author_url": "",
          "post_date": "02/19/2022 15:38:15",
          "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a>, is there a difference between binaural and stereo sound in this case ?</p>\n<p>Can't we just take the mean of the 2 channels ?</p>\n<p><code>y, sr = sf.read(path, always_2d=True)</code></p>\n<p><code>y = np.mean(y, 1)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697390,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/19/2022 15:40:13",
          "content": "<p>Alex, I believe both are the same</p>\n<blockquote>\n  <p>Can't we just take the mean of the 2 channels ?</p>\n</blockquote>\n<p>I think this might be the best approach! I'll try this 👌. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697899,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "02/20/2022 00:28:04",
          "content": "<p><a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106</a>, Thanks for the answer! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699305,
          "author_name": "nomorevotch",
          "author_url": "",
          "post_date": "02/21/2022 05:14:02",
          "content": "<p>I have experimented with 3 approaches to this problem.</p>\n<ol>\n<li>use only the first channel.</li>\n<li>calculate the mean.</li>\n<li>randomly sample at training which channel to use.</li>\n</ol>\n<p>I thought that approach 3 would work, but the results of LB were as follows</p>\n<table>\n<thead>\n<tr>\n<th>approach</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.70</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.69</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.64</td>\n</tr>\n</tbody>\n</table>\n<p>For reference, as you know, all test audio files seem to have only one channel. I don't know if this was recorded monaural, or if they do some way to make it only one channel…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699522,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "02/21/2022 08:41:28",
          "content": "<p>Could you share how you only use the first channel ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1699582,
          "author_name": "nomorevotch",
          "author_url": "",
          "post_date": "02/21/2022 09:41:13",
          "content": "<p>This is an example of assigning the first channel to y.</p>\n<pre><code>y, sr = sf.read(wav_path, always_2d=True)\ny = y[:, 0]  \n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1702177,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/23/2022 11:35:18",
          "content": "<p><a href=\"https://www.kaggle.com/nomorevotch\" target=\"_blank\">@nomorevotch</a> I had same idea as you (randomly select one of the audio channel when training) and my CV is quite lower than when I use channel average.  Kind of confirm your experience.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1701393,
      "author_name": "asaliquid1011",
      "author_url": "",
      "post_date": "02/22/2022 18:10:03",
      "content": "<p>I found there are some date that the volume of one microphone is abnormal (for example, XC344134.ogg : Channel 1 broken ,XC521984.ogg : Channel 0 broken).<br>\nIf you normalize the data after transform to spectrogram, it might not be a problem, but be careful when dealing with volumes.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1697007": "I found that some sound files had (X, 2) dim.\nI'm not familiar with it, but is the difference due to the author's recording equipment?\nAlso, is averaging a good way to handle this kind of audio?\n\n`data = np.mean(data, 1)`\n\n![image](https://user-images.githubusercontent.com/58644957/154795828-9bda47c1-bf27-4388-b114-4d45a0d6429a.png)\n\nThanks.",
    "1697035": "That means binaural. You can use `librosa.to_mono` to uniformly convert to mono format.",
    "1697068": "zzy990106 Thanks for sharing the function, I wasn't aware of it! \n\nDon't you think we are discarding useful info by converting binaural to mono format? \n\nI picture it's like converting an RGB image to B/W, throwing away info that models could potentially use.",
    "1697071": "Yes, there is information loss, but the impact on classification is far negligible.",
    "1697388": "zzy990106, is there a difference between binaural and stereo sound in this case ?\n\nCan't we just take the mean of the 2 channels ?\n\n`y, sr = sf.read(path, always_2d=True)`\n\n`y = np.mean(y, 1)`",
    "1697390": "Alex, I believe both are the same\n\n> Can't we just take the mean of the 2 channels ?\n\nI think this might be the best approach! I'll try this 👌. Thanks!",
    "1697899": "zzy990106, Thanks for the answer!",
    "1699305": "I have experimented with 3 approaches to this problem.\n\n1. use only the first channel.\n2. calculate the mean.\n3. randomly sample at training which channel to use.\n\nI thought that approach 3 would work, but the results of LB were as follows\n\n| approach | LB |\n| --- | --- |\n| 1 | 0.70 |\n| 2 | 0.69 |\n| 3 | 0.64 |\n\nFor reference, as you know, all test audio files seem to have only one channel. I don't know if this was recorded monaural, or if they do some way to make it only one channel...",
    "1699522": "Could you share how you only use the first channel ?",
    "1699582": "This is an example of assigning the first channel to y.\n\n```\ny, sr = sf.read(wav_path, always_2d=True)\ny = y[:, 0]  \n```",
    "1701393": "I found there are some date that the volume of one microphone is abnormal (for example, XC344134.ogg : Channel 1 broken ,XC521984.ogg : Channel 0 broken).\nIf you normalize the data after transform to spectrogram, it might not be a problem, but be careful when dealing with volumes.",
    "1702177": "nomorevotch I had same idea as you (randomly select one of the audio channel when training) and my CV is quite lower than when I use channel average.  Kind of confirm your experience."
  },
  "source": "meta"
}