{
  "id": 176651,
  "title": "Song Bird NLP",
  "url": "/competitions/birdsong-recognition/discussion/176651",
  "author_name": "",
  "post_date": "2020-08-22T18:50:11.411951600Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p><a href=\"https://www.nature.com/articles/ncomms10986\" target=\"_blank\">https://www.nature.com/articles/ncomms10986</a></p>\n<p>This paper essentially purports that certain songbirds have a rudimentary language. Something that, perhaps, can be modeled with tf-idf or more contemporary neural nlp techniques, such as skipgrams, transformer models, etc. I don't have much exposure to this myself, but I do have some doves that live in the tree by my house. I've noticed that whenever people get close to them, they always let out the same, sharp, characteristic \"chirp\", which is immediately followed by any other doves in the area frantically taking off to fly away.</p>\n<p>I am uncertain if these approaches would be beneficial to the task at hand here with this competition, and perhaps even lesser so since we are lacking fine-grained ground truth annotations. But from a data science perspective, how cool would it be to have a bird species language model? 🙂 It would allow us to study what birds are saying much more intimately, and do things like bird species to bird species translation.</p>",
  "messages": [
    {
      "id": "981859",
      "postDate": "08/22/2020 18:50:11",
      "content": "<p><a href=\"https://www.nature.com/articles/ncomms10986\" target=\"_blank\">https://www.nature.com/articles/ncomms10986</a></p>\n<p>This paper essentially purports that certain songbirds have a rudimentary language. Something that, perhaps, can be modeled with tf-idf or more contemporary neural nlp techniques, such as skipgrams, transformer models, etc. I don't have much exposure to this myself, but I do have some doves that live in the tree by my house. I've noticed that whenever people get close to them, they always let out the same, sharp, characteristic \"chirp\", which is immediately followed by any other doves in the area frantically taking off to fly away.</p>\n<p>I am uncertain if these approaches would be beneficial to the task at hand here with this competition, and perhaps even lesser so since we are lacking fine-grained ground truth annotations. But from a data science perspective, how cool would it be to have a bird species language model? 🙂 It would allow us to study what birds are saying much more intimately, and do things like bird species to bird species translation.</p>",
      "rawMarkdown": "https://www.nature.com/articles/ncomms10986\n\nThis paper essentially purports that certain songbirds have a rudimentary language. Something that, perhaps, can be modeled with tf-idf or more contemporary neural nlp techniques, such as skipgrams, transformer models, etc. I don't have much exposure to this myself, but I do have some doves that live in the tree by my house. I've noticed that whenever people get close to them, they always let out the same, sharp, characteristic \"chirp\", which is immediately followed by any other doves in the area frantically taking off to fly away.\n\nI am uncertain if these approaches would be beneficial to the task at hand here with this competition, and perhaps even lesser so since we are lacking fine-grained ground truth annotations. But from a data science perspective, how cool would it be to have a bird species language model? 🙂 It would allow us to study what birds are saying much more intimately, and do things like bird species to bird species translation.",
      "votes": null
    },
    {
      "id": "982295",
      "postDate": "08/23/2020 08:27:41",
      "content": "<blockquote>\n  <p>can be modeled with tf-idf </p>\n</blockquote>\n<p>still though, without actual words, how would you use td-idf?</p>",
      "rawMarkdown": "> can be modeled with tf-idf \n\nstill though, without actual words, how would you use td-idf?",
      "votes": null
    },
    {
      "id": "982483",
      "postDate": "08/23/2020 12:05:24",
      "content": "<p>Great question. If all the chirps a particular species make are label encoded, the result will be a categorical variable that term and document frequencies can be computed over. I believe this can be accomplished even in an unsupervised manner with clustering on the chirps. In fact, such an approach would also help identify anomalous chirps as well. The drawbacks to look out for would then be sounds which are frequent enough in the audio that aren't chirps appearing as clusters (so the chirps need to be manually inspected), and of course the source audio containing multiple species chirps but not appropriately being labeled as such.</p>",
      "rawMarkdown": "Great question. If all the chirps a particular species make are label encoded, the result will be a categorical variable that term and document frequencies can be computed over. I believe this can be accomplished even in an unsupervised manner with clustering on the chirps. In fact, such an approach would also help identify anomalous chirps as well. The drawbacks to look out for would then be sounds which are frequent enough in the audio that aren't chirps appearing as clusters (so the chirps need to be manually inspected), and of course the source audio containing multiple species chirps but not appropriately being labeled as such.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 982295,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "08/23/2020 08:27:41",
      "content": "<blockquote>\n  <p>can be modeled with tf-idf </p>\n</blockquote>\n<p>still though, without actual words, how would you use td-idf?</p>",
      "votes": null,
      "replies": [
        {
          "id": 982483,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/23/2020 12:05:24",
          "content": "<p>Great question. If all the chirps a particular species make are label encoded, the result will be a categorical variable that term and document frequencies can be computed over. I believe this can be accomplished even in an unsupervised manner with clustering on the chirps. In fact, such an approach would also help identify anomalous chirps as well. The drawbacks to look out for would then be sounds which are frequent enough in the audio that aren't chirps appearing as clusters (so the chirps need to be manually inspected), and of course the source audio containing multiple species chirps but not appropriately being labeled as such.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "981859": "https://www.nature.com/articles/ncomms10986\n\nThis paper essentially purports that certain songbirds have a rudimentary language. Something that, perhaps, can be modeled with tf-idf or more contemporary neural nlp techniques, such as skipgrams, transformer models, etc. I don't have much exposure to this myself, but I do have some doves that live in the tree by my house. I've noticed that whenever people get close to them, they always let out the same, sharp, characteristic \"chirp\", which is immediately followed by any other doves in the area frantically taking off to fly away.\n\nI am uncertain if these approaches would be beneficial to the task at hand here with this competition, and perhaps even lesser so since we are lacking fine-grained ground truth annotations. But from a data science perspective, how cool would it be to have a bird species language model? 🙂 It would allow us to study what birds are saying much more intimately, and do things like bird species to bird species translation.",
    "982295": "> can be modeled with tf-idf \n\nstill though, without actual words, how would you use td-idf?",
    "982483": "Great question. If all the chirps a particular species make are label encoded, the result will be a categorical variable that term and document frequencies can be computed over. I believe this can be accomplished even in an unsupervised manner with clustering on the chirps. In fact, such an approach would also help identify anomalous chirps as well. The drawbacks to look out for would then be sounds which are frequent enough in the audio that aren't chirps appearing as clusters (so the chirps need to be manually inspected), and of course the source audio containing multiple species chirps but not appropriately being labeled as such."
  },
  "source": "meta"
}