{
  "id": 201046,
  "title": "But what is songtype_id? What does it represent?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/201046",
  "author_name": "",
  "post_date": "2020-12-02T23:12:01.542292500Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>While doing preliminary analysis on the dataset, I found out this. First, have a look.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F5850a01b0c2831b6778f1f4aa75c552d%2F__results___25_0.png?generation=1607367549139328&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2Fa0fb12e5b5281c05be2e898cd874fe92%2F__results___25_1.png?generation=1607367574032376&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F2d35144052f63ecd592572378c554b88%2F__results___28_0.png?generation=1607367514274035&amp;alt=media\" alt=\"\"></p>\n<p>All those class imbalance for species_id 23 and stuff are fine, but what is songtype_id? It clearly has 2 values, and type 1 dominates. In fact, no species do have type-4 songs only except  16, 17, and 23. That too for 16, there is only song type 4 is there and for 23, it's evenly distributed. Any idea what can be the take away from this?</p>",
  "messages": [
    {
      "id": "1100217",
      "postDate": "12/02/2020 23:12:01",
      "content": "<p>While doing preliminary analysis on the dataset, I found out this. First, have a look.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F5850a01b0c2831b6778f1f4aa75c552d%2F__results___25_0.png?generation=1607367549139328&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2Fa0fb12e5b5281c05be2e898cd874fe92%2F__results___25_1.png?generation=1607367574032376&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F2d35144052f63ecd592572378c554b88%2F__results___28_0.png?generation=1607367514274035&amp;alt=media\" alt=\"\"></p>\n<p>All those class imbalance for species_id 23 and stuff are fine, but what is songtype_id? It clearly has 2 values, and type 1 dominates. In fact, no species do have type-4 songs only except  16, 17, and 23. That too for 16, there is only song type 4 is there and for 23, it's evenly distributed. Any idea what can be the take away from this?</p>",
      "rawMarkdown": "While doing preliminary analysis on the dataset, I found out this. First, have a look.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F5850a01b0c2831b6778f1f4aa75c552d%2F__results___25_0.png?generation=1607367549139328&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2Fa0fb12e5b5281c05be2e898cd874fe92%2F__results___25_1.png?generation=1607367574032376&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F2d35144052f63ecd592572378c554b88%2F__results___28_0.png?generation=1607367514274035&alt=media)\n\nAll those class imbalance for species_id 23 and stuff are fine, but what is songtype_id? It clearly has 2 values, and type 1 dominates. In fact, no species do have type-4 songs only except  16, 17, and 23. That too for 16, there is only song type 4 is there and for 23, it's evenly distributed. Any idea what can be the take away from this?",
      "votes": null
    },
    {
      "id": "1101010",
      "postDate": "12/03/2020 14:49:51",
      "content": "<p>We could make a couple extra classes and use for train and infer, then reduce for submission<br>\nby taking max for 17 and 23. <br>\nspecies_id == 17 &amp; songtype_id == 4<br>\nspecies_id == 23 &amp; songtype_id == 4</p>\n<p>I don't know yet how this will affect the scores.</p>",
      "rawMarkdown": "We could make a couple extra classes and use for train and infer, then reduce for submission\nby taking max for 17 and 23. \nspecies_id == 17 & songtype_id == 4\nspecies_id == 23 & songtype_id == 4\n\nI don't know yet how this will affect the scores.",
      "votes": null
    },
    {
      "id": "1102309",
      "postDate": "12/04/2020 19:21:09",
      "content": "<p>A bird may sing multiple songs. So it's the same species, but a different song for that species. As Thomas suggested, you could make a model that predicts Species+Song, then use logic to reduce it to just Species.</p>\n<p>Song ID is just an identifier, so the number 1 vs 4 doesn't mean anything to us other than representing a different song for each species.</p>",
      "rawMarkdown": "A bird may sing multiple songs. So it's the same species, but a different song for that species. As Thomas suggested, you could make a model that predicts Species+Song, then use logic to reduce it to just Species.\n\nSong ID is just an identifier, so the number 1 vs 4 doesn't mean anything to us other than representing a different song for each species.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1101010,
      "author_name": "tomkkk",
      "author_url": "",
      "post_date": "12/03/2020 14:49:51",
      "content": "<p>We could make a couple extra classes and use for train and infer, then reduce for submission<br>\nby taking max for 17 and 23. <br>\nspecies_id == 17 &amp; songtype_id == 4<br>\nspecies_id == 23 &amp; songtype_id == 4</p>\n<p>I don't know yet how this will affect the scores.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1102309,
      "author_name": "maltonji",
      "author_url": "",
      "post_date": "12/04/2020 19:21:09",
      "content": "<p>A bird may sing multiple songs. So it's the same species, but a different song for that species. As Thomas suggested, you could make a model that predicts Species+Song, then use logic to reduce it to just Species.</p>\n<p>Song ID is just an identifier, so the number 1 vs 4 doesn't mean anything to us other than representing a different song for each species.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1100217": "While doing preliminary analysis on the dataset, I found out this. First, have a look.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F5850a01b0c2831b6778f1f4aa75c552d%2F__results___25_0.png?generation=1607367549139328&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2Fa0fb12e5b5281c05be2e898cd874fe92%2F__results___25_1.png?generation=1607367574032376&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4227513%2F2d35144052f63ecd592572378c554b88%2F__results___28_0.png?generation=1607367514274035&alt=media)\n\nAll those class imbalance for species_id 23 and stuff are fine, but what is songtype_id? It clearly has 2 values, and type 1 dominates. In fact, no species do have type-4 songs only except  16, 17, and 23. That too for 16, there is only song type 4 is there and for 23, it's evenly distributed. Any idea what can be the take away from this?",
    "1101010": "We could make a couple extra classes and use for train and infer, then reduce for submission\nby taking max for 17 and 23. \nspecies_id == 17 & songtype_id == 4\nspecies_id == 23 & songtype_id == 4\n\nI don't know yet how this will affect the scores.",
    "1102309": "A bird may sing multiple songs. So it's the same species, but a different song for that species. As Thomas suggested, you could make a model that predicts Species+Song, then use logic to reduce it to just Species.\n\nSong ID is just an identifier, so the number 1 vs 4 doesn't mean anything to us other than representing a different song for each species."
  },
  "source": "meta"
}