{
  "id": 217955,
  "title": " There are between 4 and 5 songs per recording on average",
  "url": "/competitions/rfcx-species-audio-detection/discussion/217955",
  "author_name": "CPMP",
  "post_date": "2021-02-08T22:05:43.109000",
  "votes": 19,
  "comment_count": 10,
  "views": 0,
  "content": "<p>See my notebook <a href=\"https://www.kaggle.com/cpmpml/there-are-between-4-and-5-ongs-per-recordings?scriptVersionId=53850583\" target=\"_blank\">https://www.kaggle.com/cpmpml/there-are-between-4-and-5-ongs-per-recordings?scriptVersionId=53850583</a>  for how I get to this conclusion.  TL;DR it is derived from the 0.185 score of sample submission.</p>\n<p>How can we use this information? I don't know for now. It would certainly be helpful if the metric was about rounded predictions, like in the Cornell birds song competition. Here the ordering of predictions is what maters, and knowing how many should be close to 1 is not very helpful. Yet, why not sharing this?</p>",
  "messages": [
    {
      "id": 1192061,
      "postDate": "2021-02-08T22:05:43.110Z",
      "content": "<p>See my notebook <a href=\"https://www.kaggle.com/cpmpml/there-are-between-4-and-5-ongs-per-recordings?scriptVersionId=53850583\" target=\"_blank\">https://www.kaggle.com/cpmpml/there-are-between-4-and-5-ongs-per-recordings?scriptVersionId=53850583</a>  for how I get to this conclusion.  TL;DR it is derived from the 0.185 score of sample submission.</p>\n<p>How can we use this information? I don't know for now. It would certainly be helpful if the metric was about rounded predictions, like in the Cornell birds song competition. Here the ordering of predictions is what maters, and knowing how many should be close to 1 is not very helpful. Yet, why not sharing this?</p>",
      "rawMarkdown": "See my notebook https://www.kaggle.com/cpmpml/there-are-between-4-and-5-ongs-per-recordings?scriptVersionId=53850583  for how I get to this conclusion.  TL;DR it is derived from the 0.185 score of sample submission.\n\nHow can we use this information? I don't know for now. It would certainly be helpful if the metric was about rounded predictions, like in the Cornell birds song competition. Here the ordering of predictions is what maters, and knowing how many should be close to 1 is not very helpful. Yet, why not sharing this?",
      "votes": 19
    },
    {
      "id": 1192892,
      "postDate": "2021-02-09T10:46:14.910Z",
      "content": "<p>i suppose you are assuming that all 24 classes are balanced?</p>\n<p>since the kaggle metric is label-weighted, I believe there is some class imbalance.<br>\nSo the number \"4 to 5\" is a very rough estimate.</p>",
      "rawMarkdown": "i suppose you are assuming that all 24 classes are balanced?\n\nsince the kaggle metric is label-weighted, I believe there is some class imbalance.\nSo the number \"4 to 5\" is a very rough estimate.",
      "votes": 1,
      "replies": [
        {
          "id": 1192998,
          "postDate": "2021-02-09T11:57:10.273Z",
          "content": "<p>No, I don't assume that classes are balanced.  I assume that the number of actual species per recording is almost constant.</p>\n<p>Yes, this is an estimate.  I don't know how rough it is, and I welcome your explanation of why it would be very rough.</p>\n<p>Edit.  Your comment made me think more.  Rows are weighted by the number of actual species in the row, i.e. the number of targets equal to 1.  And the score for the rows increases linearly with the number of targets equal to 1.  Therefore, rows with a high score are weighted more than rows with a low score.  With an uneven distribution, the actual average is lower than my first estimate.</p>",
          "rawMarkdown": "No, I don't assume that classes are balanced.  I assume that the number of actual species per recording is almost constant.\n\nYes, this is an estimate.  I don't know how rough it is, and I welcome your explanation of why it would be very rough.\n\nEdit.  Your comment made me think more.  Rows are weighted by the number of actual species in the row, i.e. the number of targets equal to 1.  And the score for the rows increases linearly with the number of targets equal to 1.  Therefore, rows with a high score are weighted more than rows with a low score.  With an uneven distribution, the actual average is lower than my first estimate."
        }
      ]
    },
    {
      "id": 1192893,
      "postDate": "2021-02-09T10:46:14.910Z",
      "content": "<p>‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ </p>",
      "rawMarkdown": "‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ "
    },
    {
      "id": 1192613,
      "postDate": "2021-02-09T08:03:32.363Z",
      "content": "<p>Hi CPMP,</p>\n<p>Are you saying that the order of the 4 to 5 songs in a prediction matter? Should it not be that as long as the 4 to 5 songs are higher ranking than the remaining species, the numbers themselves don't matter?</p>",
      "rawMarkdown": "Hi CPMP,\n\nAre you saying that the order of the 4 to 5 songs in a prediction matter? Should it not be that as long as the 4 to 5 songs are higher ranking than the remaining species, the numbers themselves don't matter?",
      "replies": [
        {
          "id": 1192707,
          "postDate": "2021-02-09T09:35:45.410Z",
          "content": "<p>I'm not saying anything like that.  I don't know what to do with this information.</p>",
          "rawMarkdown": "I'm not saying anything like that.  I don't know what to do with this information."
        },
        {
          "id": 1192923,
          "postDate": "2021-02-09T11:03:53.873Z",
          "content": "<p>Me neither, but I think it is useful to have a baseline expectation.</p>",
          "rawMarkdown": "Me neither, but I think it is useful to have a baseline expectation."
        },
        {
          "id": 1193000,
          "postDate": "2021-02-09T11:58:59.703Z",
          "content": "<p>What matters for the metric is that the predictions of actual songs is higher than prediction for other songs, for each row.  Knowing how many songs there are is not directly relevant.  But maybe it can help.  I don't know how for now.</p>",
          "rawMarkdown": "What matters for the metric is that the predictions of actual songs is higher than prediction for other songs, for each row.  Knowing how many songs there are is not directly relevant.  But maybe it can help.  I don't know how for now."
        }
      ]
    },
    {
      "id": 1192185,
      "postDate": "2021-02-09T03:06:46.230Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1192206,
          "postDate": "2021-02-09T03:32:50.030Z",
          "content": "<p>Where did I make a mistake then?</p>",
          "rawMarkdown": "Where did I make a mistake then?"
        },
        {
          "id": 1192236,
          "postDate": "2021-02-09T03:55:58.807Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1192892,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-09T10:46:14.910000",
      "content": "<p>i suppose you are assuming that all 24 classes are balanced?</p>\n<p>since the kaggle metric is label-weighted, I believe there is some class imbalance.<br>\nSo the number \"4 to 5\" is a very rough estimate.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1192998,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-02-09T11:57:10.273000",
          "content": "<p>No, I don't assume that classes are balanced.  I assume that the number of actual species per recording is almost constant.</p>\n<p>Yes, this is an estimate.  I don't know how rough it is, and I welcome your explanation of why it would be very rough.</p>\n<p>Edit.  Your comment made me think more.  Rows are weighted by the number of actual species in the row, i.e. the number of targets equal to 1.  And the score for the rows increases linearly with the number of targets equal to 1.  Therefore, rows with a high score are weighted more than rows with a low score.  With an uneven distribution, the actual average is lower than my first estimate.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1192893,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-09T10:46:14.910000",
      "content": "<p>‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1192613,
      "author_name": "Victor",
      "author_url": "",
      "post_date": "2021-02-09T08:03:32.363000",
      "content": "<p>Hi CPMP,</p>\n<p>Are you saying that the order of the 4 to 5 songs in a prediction matter? Should it not be that as long as the 4 to 5 songs are higher ranking than the remaining species, the numbers themselves don't matter?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1192707,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-02-09T09:35:45.410000",
          "content": "<p>I'm not saying anything like that.  I don't know what to do with this information.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1192923,
          "author_name": "Victor",
          "author_url": "",
          "post_date": "2021-02-09T11:03:53.873000",
          "content": "<p>Me neither, but I think it is useful to have a baseline expectation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1193000,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-02-09T11:58:59.703000",
          "content": "<p>What matters for the metric is that the predictions of actual songs is higher than prediction for other songs, for each row.  Knowing how many songs there are is not directly relevant.  But maybe it can help.  I don't know how for now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1192185,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-09T03:06:46.230000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1192206,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-02-09T03:32:50.030000",
          "content": "<p>Where did I make a mistake then?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1192236,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-09T03:55:58.807000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1192061": "See my notebook https://www.kaggle.com/cpmpml/there-are-between-4-and-5-ongs-per-recordings?scriptVersionId=53850583  for how I get to this conclusion.  TL;DR it is derived from the 0.185 score of sample submission.\n\nHow can we use this information? I don't know for now. It would certainly be helpful if the metric was about rounded predictions, like in the Cornell birds song competition. Here the ordering of predictions is what maters, and knowing how many should be close to 1 is not very helpful. Yet, why not sharing this?",
    "1192892": "i suppose you are assuming that all 24 classes are balanced?\n\nsince the kaggle metric is label-weighted, I believe there is some class imbalance.\nSo the number \"4 to 5\" is a very rough estimate.",
    "1192893": "‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ‎ ",
    "1192613": "Hi CPMP,\n\nAre you saying that the order of the 4 to 5 songs in a prediction matter? Should it not be that as long as the 4 to 5 songs are higher ranking than the remaining species, the numbers themselves don't matter?",
    "1192185": ""
  }
}