{
  "id": 163657,
  "title": "CNN, RNN..? What's your approach?",
  "url": "/competitions/birdsong-recognition/discussion/163657",
  "author_name": "",
  "post_date": "2020-07-02T23:27:46.431618500Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I've entered this competition framing the problem as a sequence problem, and willing to try a RNN approach. But then I saw several kernels tackling the problem with CNNs trained on MEL Spectogram arrays (which is all new to me). Which makes more sense to you and why?</p>\n\n<p>Happy coding everyone</p>",
  "messages": [
    {
      "id": "913054",
      "postDate": "07/02/2020 23:27:46",
      "content": "<p>I've entered this competition framing the problem as a sequence problem, and willing to try a RNN approach. But then I saw several kernels tackling the problem with CNNs trained on MEL Spectogram arrays (which is all new to me). Which makes more sense to you and why?</p>\n\n<p>Happy coding everyone</p>",
      "rawMarkdown": "I've entered this competition framing the problem as a sequence problem, and willing to try a RNN approach. But then I saw several kernels tackling the problem with CNNs trained on MEL Spectogram arrays (which is all new to me). Which makes more sense to you and why?\n\nHappy coding everyone",
      "votes": null
    },
    {
      "id": "913116",
      "postDate": "07/03/2020 02:01:32",
      "content": "<p>Honestly, just thinking about the problem, RNN makes more sense as sound is more of a sequence data, but looks like using CNNs on the spectrograms, like you said, tends to perform better overall.</p>",
      "rawMarkdown": "Honestly, just thinking about the problem, RNN makes more sense as sound is more of a sequence data, but looks like using CNNs on the spectrograms, like you said, tends to perform better overall.",
      "votes": null
    },
    {
      "id": "913742",
      "postDate": "07/03/2020 11:37:05",
      "content": "<p>Most of the gold medalists in previous years used CNN. And CNN also performs better than RNN in my model. </p>",
      "rawMarkdown": "Most of the gold medalists in previous years used CNN. And CNN also performs better than RNN in my model.",
      "votes": null
    },
    {
      "id": "913832",
      "postDate": "07/03/2020 13:08:57",
      "content": "<p>Combination of both ConvLSTM2D</p>",
      "rawMarkdown": "Combination of both ConvLSTM2D",
      "votes": null
    },
    {
      "id": "914415",
      "postDate": "07/03/2020 21:05:20",
      "content": "<p>bidirectional cuda lstm</p>",
      "rawMarkdown": "bidirectional cuda lstm",
      "votes": null
    },
    {
      "id": "914444",
      "postDate": "07/03/2020 22:04:22",
      "content": "<p>I've been approaching it as an \"identification\" problem but could not come up with anything interesting. Now looking at it from NLP lenses.</p>",
      "rawMarkdown": "I've been approaching it as an \"identification\" problem but could not come up with anything interesting. Now looking at it from NLP lenses.",
      "votes": null
    },
    {
      "id": "982000",
      "postDate": "08/23/2020 00:04:13",
      "content": "<p>maybe using data as images works better because there are a lot of pre-trained models for image, it helps a lot to perform better. But i will not be surprised if anyone could develop a transformer architecture that performs very well too</p>",
      "rawMarkdown": "maybe using data as images works better because there are a lot of pre-trained models for image, it helps a lot to perform better. But i will not be surprised if anyone could develop a transformer architecture that performs very well too",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 982000,
      "author_name": "hidekin2011",
      "author_url": "",
      "post_date": "08/23/2020 00:04:13",
      "content": "<p>maybe using data as images works better because there are a lot of pre-trained models for image, it helps a lot to perform better. But i will not be surprised if anyone could develop a transformer architecture that performs very well too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 913116,
      "author_name": "jonykarki",
      "author_url": "",
      "post_date": "07/03/2020 02:01:32",
      "content": "<p>Honestly, just thinking about the problem, RNN makes more sense as sound is more of a sequence data, but looks like using CNNs on the spectrograms, like you said, tends to perform better overall.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 913742,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "07/03/2020 11:37:05",
      "content": "<p>Most of the gold medalists in previous years used CNN. And CNN also performs better than RNN in my model. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 913832,
      "author_name": "mlneo07",
      "author_url": "",
      "post_date": "07/03/2020 13:08:57",
      "content": "<p>Combination of both ConvLSTM2D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 914415,
      "author_name": "servietsky",
      "author_url": "",
      "post_date": "07/03/2020 21:05:20",
      "content": "<p>bidirectional cuda lstm</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 914444,
      "author_name": "ramarlina",
      "author_url": "",
      "post_date": "07/03/2020 22:04:22",
      "content": "<p>I've been approaching it as an \"identification\" problem but could not come up with anything interesting. Now looking at it from NLP lenses.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "913054": "I've entered this competition framing the problem as a sequence problem, and willing to try a RNN approach. But then I saw several kernels tackling the problem with CNNs trained on MEL Spectogram arrays (which is all new to me). Which makes more sense to you and why?\n\nHappy coding everyone",
    "913116": "Honestly, just thinking about the problem, RNN makes more sense as sound is more of a sequence data, but looks like using CNNs on the spectrograms, like you said, tends to perform better overall.",
    "913742": "Most of the gold medalists in previous years used CNN. And CNN also performs better than RNN in my model.",
    "913832": "Combination of both ConvLSTM2D",
    "914415": "bidirectional cuda lstm",
    "914444": "I've been approaching it as an \"identification\" problem but could not come up with anything interesting. Now looking at it from NLP lenses.",
    "982000": "maybe using data as images works better because there are a lot of pre-trained models for image, it helps a lot to perform better. But i will not be surprised if anyone could develop a transformer architecture that performs very well too"
  },
  "source": "meta"
}