{
  "id": 316719,
  "title": "This is my first audio classification competiton.",
  "url": "/competitions/birdclef-2022/discussion/316719",
  "author_name": "",
  "post_date": "2022-04-03T11:56:40.639381Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is the first audio classification that I started to make a model for myself.<br>\nSince I don't know anything about audio classification, I regard this competition as image classification. <br>\nThe difference between this competition and the image classification competition is that the frequency direction is important.<br>\nNow I am very surprised by the current leaderboard score. It seems that my modeling is working.</p>\n<p>Good Luck!</p>",
  "messages": [
    {
      "id": "1743892",
      "postDate": "04/03/2022 11:56:40",
      "content": "<p>This is the first audio classification that I started to make a model for myself.<br>\nSince I don't know anything about audio classification, I regard this competition as image classification. <br>\nThe difference between this competition and the image classification competition is that the frequency direction is important.<br>\nNow I am very surprised by the current leaderboard score. It seems that my modeling is working.</p>\n<p>Good Luck!</p>",
      "rawMarkdown": "This is the first audio classification that I started to make a model for myself.\nSince I don't know anything about audio classification, I regard this competition as image classification. \nThe difference between this competition and the image classification competition is that the frequency direction is important.\nNow I am very surprised by the current leaderboard score. It seems that my modeling is working.\n\nGood Luck!",
      "votes": null
    },
    {
      "id": "1747105",
      "postDate": "04/06/2022 11:09:32",
      "content": "<p>Nice. Audio is pretty much an image classification contest. The part about converting to spectrograms is usually straightforward and similar amongst all contestants. From thereon, it is more about handling images. Of course there are some well documented tricks and hacks from winners of prev comp which will help. Trust you have gone thru' those. All the best for the private score!!</p>",
      "rawMarkdown": "Nice. Audio is pretty much an image classification contest. The part about converting to spectrograms is usually straightforward and similar amongst all contestants. From thereon, it is more about handling images. Of course there are some well documented tricks and hacks from winners of prev comp which will help. Trust you have gone thru' those. All the best for the private score!!",
      "votes": null
    },
    {
      "id": "1747568",
      "postDate": "04/06/2022 18:55:02",
      "content": "<p>I wonder why - do you have some insights to share? My previous works w/ audio have always used the raw audio data. Why would you want to convert it into an image? You can't magically get more information from it - you can only lose information … Is it the (F)FT(s) that somehow transform the signal into something that's easier to reason about? But then, I think I saw a lot of Mel spectrograms … why would you chose these? Isn't the \"Mel\" part about some characteristics of our human ear? Why should the machine care about that?</p>",
      "rawMarkdown": "I wonder why - do you have some insights to share? My previous works w/ audio have always used the raw audio data. Why would you want to convert it into an image? You can't magically get more information from it - you can only lose information ... Is it the (F)FT(s) that somehow transform the signal into something that's easier to reason about? But then, I think I saw a lot of Mel spectrograms ... why would you chose these? Isn't the \"Mel\" part about some characteristics of our human ear? Why should the machine care about that?",
      "votes": null
    },
    {
      "id": "1747902",
      "postDate": "04/07/2022 07:00:54",
      "content": "<p>Multiple reasons (note I am not expert though)</p>\n<ul>\n<li>Loss of information is good in a way so we can focus on much smaller data. Spectrogram is a century old concept to compress relevant audio data without losing much info</li>\n<li>SOTA Prertained models available for images - simpler than training a model from scratch using raw data</li>\n<li>Easy to augment images rather than augmenting raw data</li>\n<li>Mel spects perform (very slightly) better than regular spects. This is a good question and I dont have an answer why. Possibly this is a result of the way humans manually annotate the test dataset. I am not sure</li>\n<li>Yes, if we use image models, frequency axis should ideally be variant and not invariant. Yet it works, possibly due to all the other remaining advantages. one of the competitors CPMP in last comp came up with a novel way to handle frequency invariance which you should read about</li>\n</ul>\n<p>But coming back to the realm of speculation - If you have enough compute power, it would be worthwhile throwing the RAW DATA (not image) at a SOTA transformer model and train for a long long time. I suspect it might throw good results or at least some very interesting observations. As birdclef is a standard recurring event, a lot of GMs will enter the comp in the last month. It would be interesting to hear their comments on this particular aspect.  </p>",
      "rawMarkdown": "Multiple reasons (note I am not expert though)\n- Loss of information is good in a way so we can focus on much smaller data. Spectrogram is a century old concept to compress relevant audio data without losing much info\n- SOTA Prertained models available for images - simpler than training a model from scratch using raw data\n- Easy to augment images rather than augmenting raw data\n- Mel spects perform (very slightly) better than regular spects. This is a good question and I dont have an answer why. Possibly this is a result of the way humans manually annotate the test dataset. I am not sure\n- Yes, if we use image models, frequency axis should ideally be variant and not invariant. Yet it works, possibly due to all the other remaining advantages. one of the competitors CPMP in last comp came up with a novel way to handle frequency invariance which you should read about\n\nBut coming back to the realm of speculation - If you have enough compute power, it would be worthwhile throwing the RAW DATA (not image) at a SOTA transformer model and train for a long long time. I suspect it might throw good results or at least some very interesting observations. As birdclef is a standard recurring event, a lot of GMs will enter the comp in the last month. It would be interesting to hear their comments on this particular aspect.",
      "votes": null
    },
    {
      "id": "1757242",
      "postDate": "04/16/2022 12:26:05",
      "content": "<p>Thanks a lot for your reply!</p>",
      "rawMarkdown": "Thanks a lot for your reply!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1747105,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "04/06/2022 11:09:32",
      "content": "<p>Nice. Audio is pretty much an image classification contest. The part about converting to spectrograms is usually straightforward and similar amongst all contestants. From thereon, it is more about handling images. Of course there are some well documented tricks and hacks from winners of prev comp which will help. Trust you have gone thru' those. All the best for the private score!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1747568,
          "author_name": "m02ph3u5",
          "author_url": "",
          "post_date": "04/06/2022 18:55:02",
          "content": "<p>I wonder why - do you have some insights to share? My previous works w/ audio have always used the raw audio data. Why would you want to convert it into an image? You can't magically get more information from it - you can only lose information … Is it the (F)FT(s) that somehow transform the signal into something that's easier to reason about? But then, I think I saw a lot of Mel spectrograms … why would you chose these? Isn't the \"Mel\" part about some characteristics of our human ear? Why should the machine care about that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1747902,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "04/07/2022 07:00:54",
          "content": "<p>Multiple reasons (note I am not expert though)</p>\n<ul>\n<li>Loss of information is good in a way so we can focus on much smaller data. Spectrogram is a century old concept to compress relevant audio data without losing much info</li>\n<li>SOTA Prertained models available for images - simpler than training a model from scratch using raw data</li>\n<li>Easy to augment images rather than augmenting raw data</li>\n<li>Mel spects perform (very slightly) better than regular spects. This is a good question and I dont have an answer why. Possibly this is a result of the way humans manually annotate the test dataset. I am not sure</li>\n<li>Yes, if we use image models, frequency axis should ideally be variant and not invariant. Yet it works, possibly due to all the other remaining advantages. one of the competitors CPMP in last comp came up with a novel way to handle frequency invariance which you should read about</li>\n</ul>\n<p>But coming back to the realm of speculation - If you have enough compute power, it would be worthwhile throwing the RAW DATA (not image) at a SOTA transformer model and train for a long long time. I suspect it might throw good results or at least some very interesting observations. As birdclef is a standard recurring event, a lot of GMs will enter the comp in the last month. It would be interesting to hear their comments on this particular aspect.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1757242,
          "author_name": "m02ph3u5",
          "author_url": "",
          "post_date": "04/16/2022 12:26:05",
          "content": "<p>Thanks a lot for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1743892": "This is the first audio classification that I started to make a model for myself.\nSince I don't know anything about audio classification, I regard this competition as image classification. \nThe difference between this competition and the image classification competition is that the frequency direction is important.\nNow I am very surprised by the current leaderboard score. It seems that my modeling is working.\n\nGood Luck!",
    "1747105": "Nice. Audio is pretty much an image classification contest. The part about converting to spectrograms is usually straightforward and similar amongst all contestants. From thereon, it is more about handling images. Of course there are some well documented tricks and hacks from winners of prev comp which will help. Trust you have gone thru' those. All the best for the private score!!",
    "1747568": "I wonder why - do you have some insights to share? My previous works w/ audio have always used the raw audio data. Why would you want to convert it into an image? You can't magically get more information from it - you can only lose information ... Is it the (F)FT(s) that somehow transform the signal into something that's easier to reason about? But then, I think I saw a lot of Mel spectrograms ... why would you chose these? Isn't the \"Mel\" part about some characteristics of our human ear? Why should the machine care about that?",
    "1747902": "Multiple reasons (note I am not expert though)\n- Loss of information is good in a way so we can focus on much smaller data. Spectrogram is a century old concept to compress relevant audio data without losing much info\n- SOTA Prertained models available for images - simpler than training a model from scratch using raw data\n- Easy to augment images rather than augmenting raw data\n- Mel spects perform (very slightly) better than regular spects. This is a good question and I dont have an answer why. Possibly this is a result of the way humans manually annotate the test dataset. I am not sure\n- Yes, if we use image models, frequency axis should ideally be variant and not invariant. Yet it works, possibly due to all the other remaining advantages. one of the competitors CPMP in last comp came up with a novel way to handle frequency invariance which you should read about\n\nBut coming back to the realm of speculation - If you have enough compute power, it would be worthwhile throwing the RAW DATA (not image) at a SOTA transformer model and train for a long long time. I suspect it might throw good results or at least some very interesting observations. As birdclef is a standard recurring event, a lot of GMs will enter the comp in the last month. It would be interesting to hear their comments on this particular aspect.",
    "1757242": "Thanks a lot for your reply!"
  },
  "source": "meta"
}