{
  "id": 230181,
  "title": "What type of Deep learning models do you use about BirdCLEF ?",
  "url": "/competitions/birdclef-2021/discussion/230181",
  "author_name": "",
  "post_date": "2021-04-02T12:28:40.472435600Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I have studied the image classification, but I little know audio detection Deep learning models.<br>\nIt is often said that RNNs are used in speech recognition, but It seems that RNN models little used in this competitions. </p>\n<p>Why is RNN models not used so much? </p>\n<p>Please teach me this reason if you can.</p>",
  "messages": [
    {
      "id": "1260815",
      "postDate": "04/02/2021 12:28:40",
      "content": "<p>I have studied the image classification, but I little know audio detection Deep learning models.<br>\nIt is often said that RNNs are used in speech recognition, but It seems that RNN models little used in this competitions. </p>\n<p>Why is RNN models not used so much? </p>\n<p>Please teach me this reason if you can.</p>",
      "rawMarkdown": "I have studied the image classification, but I little know audio detection Deep learning models.\nIt is often said that RNNs are used in speech recognition, but It seems that RNN models little used in this competitions. \n\nWhy is RNN models not used so much? \n\nPlease teach me this reason if you can.",
      "votes": null
    },
    {
      "id": "1262505",
      "postDate": "04/04/2021 11:45:12",
      "content": "<p>You can look at what was used in last two birdsong competitions:</p>\n<p><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection</a></p>\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition</a></p>",
      "rawMarkdown": "You can look at what was used in last two birdsong competitions:\n\nhttps://www.kaggle.com/c/rfcx-species-audio-detection\n\nhttps://www.kaggle.com/c/birdsong-recognition",
      "votes": null
    },
    {
      "id": "1262925",
      "postDate": "04/04/2021 22:09:33",
      "content": "<p>Thank you for sharing the important competitions.<br>\nI will study hard.</p>",
      "rawMarkdown": "Thank you for sharing the important competitions.\nI will study hard.",
      "votes": null
    },
    {
      "id": "1262959",
      "postDate": "04/04/2021 23:43:38",
      "content": "<p>Mainly because the audio spectograms can be interpreted as images, ConvNets are a popular choice because of the ability to do Transfer Learning on state-of-the-art models. I personally think that CNNs (both 1d and 2d) are the best when dealing with this type of data.</p>",
      "rawMarkdown": "Mainly because the audio spectograms can be interpreted as images, ConvNets are a popular choice because of the ability to do Transfer Learning on state-of-the-art models. I personally think that CNNs (both 1d and 2d) are the best when dealing with this type of data.",
      "votes": null
    },
    {
      "id": "1264720",
      "postDate": "04/06/2021 11:47:04",
      "content": "<p>Thank you for appreciate replying. <br>\nThat is, we should focus on translating audio-data to image-data and then input the  Transfer Learning model , not considering the model selection,RNN or CNN.</p>\n<p>I think I could understand it roughly.</p>",
      "rawMarkdown": "Thank you for appreciate replying. \nThat is, we should focus on translating audio-data to image-data and then input the  Transfer Learning model , not considering the model selection,RNN or CNN.\n\nI think I could understand it roughly.",
      "votes": null
    },
    {
      "id": "1269509",
      "postDate": "04/10/2021 15:46:40",
      "content": "<p>My best guess is that you need strong label to use RNN, knowing the exact start time and end time for the exact bird (the same way in speech recognition task, you would require the transcription to train)</p>",
      "rawMarkdown": "My best guess is that you need strong label to use RNN, knowing the exact start time and end time for the exact bird (the same way in speech recognition task, you would require the transcription to train)",
      "votes": null
    },
    {
      "id": "1293536",
      "postDate": "05/05/2021 01:51:33",
      "content": "<p>Thank you for your reply. Sorry for replying late. </p>\n<p>I see. That is, the data label in this competition is weakly compared to the usual RNN audio models, so RNN model does not effective in this competitions.</p>",
      "rawMarkdown": "Thank you for your reply. Sorry for replying late. \n\nI see. That is, the data label in this competition is weakly compared to the usual RNN audio models, so RNN model does not effective in this competitions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1262505,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/04/2021 11:45:12",
      "content": "<p>You can look at what was used in last two birdsong competitions:</p>\n<p><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection</a></p>\n<p><a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1262925,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "04/04/2021 22:09:33",
          "content": "<p>Thank you for sharing the important competitions.<br>\nI will study hard.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1262959,
      "author_name": "sebastianponce",
      "author_url": "",
      "post_date": "04/04/2021 23:43:38",
      "content": "<p>Mainly because the audio spectograms can be interpreted as images, ConvNets are a popular choice because of the ability to do Transfer Learning on state-of-the-art models. I personally think that CNNs (both 1d and 2d) are the best when dealing with this type of data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264720,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "04/06/2021 11:47:04",
          "content": "<p>Thank you for appreciate replying. <br>\nThat is, we should focus on translating audio-data to image-data and then input the  Transfer Learning model , not considering the model selection,RNN or CNN.</p>\n<p>I think I could understand it roughly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1269509,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "04/10/2021 15:46:40",
      "content": "<p>My best guess is that you need strong label to use RNN, knowing the exact start time and end time for the exact bird (the same way in speech recognition task, you would require the transcription to train)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1293536,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "05/05/2021 01:51:33",
          "content": "<p>Thank you for your reply. Sorry for replying late. </p>\n<p>I see. That is, the data label in this competition is weakly compared to the usual RNN audio models, so RNN model does not effective in this competitions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1260815": "I have studied the image classification, but I little know audio detection Deep learning models.\nIt is often said that RNNs are used in speech recognition, but It seems that RNN models little used in this competitions. \n\nWhy is RNN models not used so much? \n\nPlease teach me this reason if you can.",
    "1262505": "You can look at what was used in last two birdsong competitions:\n\nhttps://www.kaggle.com/c/rfcx-species-audio-detection\n\nhttps://www.kaggle.com/c/birdsong-recognition",
    "1262925": "Thank you for sharing the important competitions.\nI will study hard.",
    "1262959": "Mainly because the audio spectograms can be interpreted as images, ConvNets are a popular choice because of the ability to do Transfer Learning on state-of-the-art models. I personally think that CNNs (both 1d and 2d) are the best when dealing with this type of data.",
    "1264720": "Thank you for appreciate replying. \nThat is, we should focus on translating audio-data to image-data and then input the  Transfer Learning model , not considering the model selection,RNN or CNN.\n\nI think I could understand it roughly.",
    "1269509": "My best guess is that you need strong label to use RNN, knowing the exact start time and end time for the exact bird (the same way in speech recognition task, you would require the transcription to train)",
    "1293536": "Thank you for your reply. Sorry for replying late. \n\nI see. That is, the data label in this competition is weakly compared to the usual RNN audio models, so RNN model does not effective in this competitions."
  },
  "source": "meta"
}