{
  "id": 169415,
  "title": "Professionals guide",
  "url": "/competitions/birdsong-recognition/discussion/169415",
  "author_name": "",
  "post_date": "2020-07-23T19:06:55.280463200Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I can say this challenge is one of the different competitions by Kaggle. What I have found so far shows it is not easy job to detect each bird by their voices. </p>\n\n<p>I am new to this field and would like to discover new skills but I have no idea from where to start. <strong>What algorithms, methods, or models do you recommend to study</strong>? I would really appreciate if you draw a roadmap. \nConsider I am good at general models and the latest algorithms such as XGBoost, CatBoost, LightGBM except deep learning and Neural Networks. </p>\n\n<p>Thank you</p>",
  "messages": [
    {
      "id": "942482",
      "postDate": "07/23/2020 19:06:55",
      "content": "<p>I can say this challenge is one of the different competitions by Kaggle. What I have found so far shows it is not easy job to detect each bird by their voices. </p>\n\n<p>I am new to this field and would like to discover new skills but I have no idea from where to start. <strong>What algorithms, methods, or models do you recommend to study</strong>? I would really appreciate if you draw a roadmap. \nConsider I am good at general models and the latest algorithms such as XGBoost, CatBoost, LightGBM except deep learning and Neural Networks. </p>\n\n<p>Thank you</p>",
      "rawMarkdown": "I can say this challenge is one of the different competitions by Kaggle. What I have found so far shows it is not easy job to detect each bird by their voices. \n\nI am new to this field and would like to discover new skills but I have no idea from where to start. **What algorithms, methods, or models do you recommend to study**? I would really appreciate if you draw a roadmap. \nConsider I am good at general models and the latest algorithms such as XGBoost, CatBoost, LightGBM except deep learning and Neural Networks. \n\nThank you",
      "votes": null
    },
    {
      "id": "943210",
      "postDate": "07/24/2020 07:51:00",
      "content": "<p>If you are looking at tabular methods you can have a look here <a href=\"http://ataspinar.com/2018/04/04/machine-learning-with-signal-processing-techniques/\">http://ataspinar.com/2018/04/04/machine-learning-with-signal-processing-techniques/</a>\nHowever it is very unlikely that any classical tabular approach will work here</p>",
      "rawMarkdown": "If you are looking at tabular methods you can have a look here http://ataspinar.com/2018/04/04/machine-learning-with-signal-processing-techniques/\nHowever it is very unlikely that any classical tabular approach will work here",
      "votes": null
    },
    {
      "id": "943363",
      "postDate": "07/24/2020 10:10:57",
      "content": "<p>Thank you for response.\nNot necessarily looking for tabular methods. I can start studying other subjects as well. What is the standard method or model for this case?</p>",
      "rawMarkdown": "Thank you for response.\nNot necessarily looking for tabular methods. I can start studying other subjects as well. What is the standard method or model for this case?",
      "votes": null
    },
    {
      "id": "943754",
      "postDate": "07/24/2020 15:01:16",
      "content": "<p>What I have understood so far is that one of the ways for classifying audio is by using fourier transfer based features. Essentially an audio is a wave. And each wave is composedof multiple waves which can be arrived at using fourier transform. So essentially now you have multiple waves, each having diff frequency an amplitude. These can be represented visually by  a spectrogram, where in color and intensity are used to represent frequency and amplitude. Mel transform is applied to frequency which essentially scales the waves so that they are differentiated by a factor more recognizable by human ear. Once we have audio represented by an image we can use usual image classification architectures, I have so far see resent50 beng used more often. \nAnother interesting solution I saw was an LSTM based approach where in we took 10 samples per seconds as agains the the usual sampling rate of 2.2 K and then LSTM was applied. This solution achieved similar accuracy levels and I found it quite interesting. Hope this gives you some intuition.</p>",
      "rawMarkdown": "What I have understood so far is that one of the ways for classifying audio is by using fourier transfer based features. Essentially an audio is a wave. And each wave is composedof multiple waves which can be arrived at using fourier transform. So essentially now you have multiple waves, each having diff frequency an amplitude. These can be represented visually by  a spectrogram, where in color and intensity are used to represent frequency and amplitude. Mel transform is applied to frequency which essentially scales the waves so that they are differentiated by a factor more recognizable by human ear. Once we have audio represented by an image we can use usual image classification architectures, I have so far see resent50 beng used more often. \nAnother interesting solution I saw was an LSTM based approach where in we took 10 samples per seconds as agains the the usual sampling rate of 2.2 K and then LSTM was applied. This solution achieved similar accuracy levels and I found it quite interesting. Hope this gives you some intuition.",
      "votes": null
    },
    {
      "id": "944449",
      "postDate": "07/25/2020 06:02:21",
      "content": "<p>Well yes, I read your explanation and got a lot information in this regard and I appreciate it. What I've got is to visualize voices and then step into image classification. </p>",
      "rawMarkdown": "Well yes, I read your explanation and got a lot information in this regard and I appreciate it. What I've got is to visualize voices and then step into image classification.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 943210,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "07/24/2020 07:51:00",
      "content": "<p>If you are looking at tabular methods you can have a look here <a href=\"http://ataspinar.com/2018/04/04/machine-learning-with-signal-processing-techniques/\">http://ataspinar.com/2018/04/04/machine-learning-with-signal-processing-techniques/</a>\nHowever it is very unlikely that any classical tabular approach will work here</p>",
      "votes": null,
      "replies": [
        {
          "id": 943363,
          "author_name": "abdolazimrezaei",
          "author_url": "",
          "post_date": "07/24/2020 10:10:57",
          "content": "<p>Thank you for response.\nNot necessarily looking for tabular methods. I can start studying other subjects as well. What is the standard method or model for this case?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 943754,
      "author_name": "shikha130vv",
      "author_url": "",
      "post_date": "07/24/2020 15:01:16",
      "content": "<p>What I have understood so far is that one of the ways for classifying audio is by using fourier transfer based features. Essentially an audio is a wave. And each wave is composedof multiple waves which can be arrived at using fourier transform. So essentially now you have multiple waves, each having diff frequency an amplitude. These can be represented visually by  a spectrogram, where in color and intensity are used to represent frequency and amplitude. Mel transform is applied to frequency which essentially scales the waves so that they are differentiated by a factor more recognizable by human ear. Once we have audio represented by an image we can use usual image classification architectures, I have so far see resent50 beng used more often. \nAnother interesting solution I saw was an LSTM based approach where in we took 10 samples per seconds as agains the the usual sampling rate of 2.2 K and then LSTM was applied. This solution achieved similar accuracy levels and I found it quite interesting. Hope this gives you some intuition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 944449,
          "author_name": "abdolazimrezaei",
          "author_url": "",
          "post_date": "07/25/2020 06:02:21",
          "content": "<p>Well yes, I read your explanation and got a lot information in this regard and I appreciate it. What I've got is to visualize voices and then step into image classification. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "942482": "I can say this challenge is one of the different competitions by Kaggle. What I have found so far shows it is not easy job to detect each bird by their voices. \n\nI am new to this field and would like to discover new skills but I have no idea from where to start. **What algorithms, methods, or models do you recommend to study**? I would really appreciate if you draw a roadmap. \nConsider I am good at general models and the latest algorithms such as XGBoost, CatBoost, LightGBM except deep learning and Neural Networks. \n\nThank you",
    "943210": "If you are looking at tabular methods you can have a look here http://ataspinar.com/2018/04/04/machine-learning-with-signal-processing-techniques/\nHowever it is very unlikely that any classical tabular approach will work here",
    "943363": "Thank you for response.\nNot necessarily looking for tabular methods. I can start studying other subjects as well. What is the standard method or model for this case?",
    "943754": "What I have understood so far is that one of the ways for classifying audio is by using fourier transfer based features. Essentially an audio is a wave. And each wave is composedof multiple waves which can be arrived at using fourier transform. So essentially now you have multiple waves, each having diff frequency an amplitude. These can be represented visually by  a spectrogram, where in color and intensity are used to represent frequency and amplitude. Mel transform is applied to frequency which essentially scales the waves so that they are differentiated by a factor more recognizable by human ear. Once we have audio represented by an image we can use usual image classification architectures, I have so far see resent50 beng used more often. \nAnother interesting solution I saw was an LSTM based approach where in we took 10 samples per seconds as agains the the usual sampling rate of 2.2 K and then LSTM was applied. This solution achieved similar accuracy levels and I found it quite interesting. Hope this gives you some intuition.",
    "944449": "Well yes, I read your explanation and got a lot information in this regard and I appreciate it. What I've got is to visualize voices and then step into image classification."
  },
  "source": "meta"
}