{
  "id": 500269,
  "title": "More thoughts on scoring... (are soundscapes scored together or separate?)",
  "url": "/competitions/birdclef-2024/discussion/500269",
  "author_name": "",
  "post_date": "2024-05-05T01:42:13.429568900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So - the competition describes: <br>\n\"a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"<br>\n<a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">https://www.kaggle.com/code/metric/birdclef-roc-auc</a><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring</a></p>\n<p>That code would seem to suggest a single DF goes in for scoring - and then kicks out an LB score.</p>\n<p>On that assumption - all the soundscapes would be scored together.</p>\n<p>In this case - for a class to be \"skipped\" - it must not have any true positive labels in any of the soundscapes.  Predicting a Purple Pigeon in a soundscape with no Purple Pigeons will result in your score being penalized (assuming another soundscape has a Purple Pigeon).  Scores for a given species would be relative to all scores for that species across soundscapes.</p>\n<p>Another possibility is for each soundscape to be scored separately - and then have the scores for the soundscapes averaged together.</p>\n<p>In that case - any species not present in a given soundscape is exempt from scoring for that soundscape.  Further - predictions for the same species across different soundscapes could be scaled or offset differently from each other without impact on the LB score.</p>\n<p>Does anyone have any thoughts on which it is (or something different)? Or better yet - is there something documented for the competition?</p>",
  "messages": [
    {
      "id": "2793825",
      "postDate": "05/05/2024 01:42:13",
      "content": "<p>So - the competition describes: <br>\n\"a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"<br>\n<a href=\"https://www.kaggle.com/code/metric/birdclef-roc-auc\" target=\"_blank\">https://www.kaggle.com/code/metric/birdclef-roc-auc</a><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring</a></p>\n<p>That code would seem to suggest a single DF goes in for scoring - and then kicks out an LB score.</p>\n<p>On that assumption - all the soundscapes would be scored together.</p>\n<p>In this case - for a class to be \"skipped\" - it must not have any true positive labels in any of the soundscapes.  Predicting a Purple Pigeon in a soundscape with no Purple Pigeons will result in your score being penalized (assuming another soundscape has a Purple Pigeon).  Scores for a given species would be relative to all scores for that species across soundscapes.</p>\n<p>Another possibility is for each soundscape to be scored separately - and then have the scores for the soundscapes averaged together.</p>\n<p>In that case - any species not present in a given soundscape is exempt from scoring for that soundscape.  Further - predictions for the same species across different soundscapes could be scaled or offset differently from each other without impact on the LB score.</p>\n<p>Does anyone have any thoughts on which it is (or something different)? Or better yet - is there something documented for the competition?</p>",
      "rawMarkdown": "So - the competition describes: \n\"a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"\nhttps://www.kaggle.com/code/metric/birdclef-roc-auc\nhttps://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\n\nThat code would seem to suggest a single DF goes in for scoring - and then kicks out an LB score.\n\nOn that assumption - all the soundscapes would be scored together.\n\nIn this case - for a class to be \"skipped\" - it must not have any true positive labels in any of the soundscapes.  Predicting a Purple Pigeon in a soundscape with no Purple Pigeons will result in your score being penalized (assuming another soundscape has a Purple Pigeon).  Scores for a given species would be relative to all scores for that species across soundscapes.\n\nAnother possibility is for each soundscape to be scored separately - and then have the scores for the soundscapes averaged together.\n\nIn that case - any species not present in a given soundscape is exempt from scoring for that soundscape.  Further - predictions for the same species across different soundscapes could be scaled or offset differently from each other without impact on the LB score.\n\nDoes anyone have any thoughts on which it is (or something different)? Or better yet - is there something documented for the competition?",
      "votes": null
    },
    {
      "id": "2793902",
      "postDate": "05/05/2024 03:58:47",
      "content": "<p>As far as I understand, complete results (<code>solution</code> and <code>submission</code>) are submitted to calculate the metric. That is, all the soundscapes.</p>",
      "rawMarkdown": "As far as I understand, complete results (`solution` and `submission`) are submitted to calculate the metric. That is, all the soundscapes.",
      "votes": null
    },
    {
      "id": "2795881",
      "postDate": "05/06/2024 03:21:51",
      "content": "<p>OK - so I did some tests - and I got some curious results…</p>\n<p>Version 6: 0.61 LB \"Baseline\"<br>\nAn existing model I've used / is used for all other versions in this test.<br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment</a></p>\n<p>Version 7: 0.50 LB<br>\nI select a random number between 0 and 1 for each soundscape - and added it to all the predictions for that soundscape. <br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175709383\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175709383</a></p>\n<p>Version 8: 0.61 LB (\"Best Score\"?!)<br>\nIn this notebook I choose a random number between 0.5 and 1 for each soundscape - and then -multiply- each prediction in that soundscape by it.  So - each soundscape will randomly have its score scaled to between 50% and 100% of what the model predicted. <br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175888223\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175888223</a></p>\n<p>Version 9: 0.60 LB<br>\nHere we scale each soundscapes prediction by a multiplier between 0.1 and 1.  Each soundscape has all it's scored scaled between 10% and 100% of what the model predicted.<br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175899536\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175899536</a></p>\n<p>So - I'm not 100% sure what this means - but maybe there are some clues about the nature of the test data / scoring.  I think my code is correct…</p>\n<p>0.50 LB score in Version 7 is what I'd expect for combined scoring.</p>\n<p>But.. I find that scaling predictions per-soundscape over a factor of 10x would have relatively little impact on score surprising.  That could be explained if separated-by-soundscape scoring was used.  Alternately - maybe it indicates something about the data (or is maybe a side-effect of doing these tests on a relatively-low-scoring model).</p>\n<p>Anyways - tossing it out there for discussion….</p>",
      "rawMarkdown": "OK - so I did some tests - and I got some curious results...\n\nVersion 6: 0.61 LB \"Baseline\"\nAn existing model I've used / is used for all other versions in this test.\nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment\n\nVersion 7: 0.50 LB\nI select a random number between 0 and 1 for each soundscape - and added it to all the predictions for that soundscape. \nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175709383\n\nVersion 8: 0.61 LB (\"Best Score\"?!)\nIn this notebook I choose a random number between 0.5 and 1 for each soundscape - and then -multiply- each prediction in that soundscape by it.  So - each soundscape will randomly have its score scaled to between 50% and 100% of what the model predicted. \nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175888223\n\nVersion 9: 0.60 LB\nHere we scale each soundscapes prediction by a multiplier between 0.1 and 1.  Each soundscape has all it's scored scaled between 10% and 100% of what the model predicted.\nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175899536\n\nSo - I'm not 100% sure what this means - but maybe there are some clues about the nature of the test data / scoring.  I think my code is correct...\n\n0.50 LB score in Version 7 is what I'd expect for combined scoring.\n\nBut.. I find that scaling predictions per-soundscape over a factor of 10x would have relatively little impact on score surprising.  That could be explained if separated-by-soundscape scoring was used.  Alternately - maybe it indicates something about the data (or is maybe a side-effect of doing these tests on a relatively-low-scoring model).\n\nAnyways - tossing it out there for discussion....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2793902,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "05/05/2024 03:58:47",
      "content": "<p>As far as I understand, complete results (<code>solution</code> and <code>submission</code>) are submitted to calculate the metric. That is, all the soundscapes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2795881,
      "author_name": "richolson",
      "author_url": "",
      "post_date": "05/06/2024 03:21:51",
      "content": "<p>OK - so I did some tests - and I got some curious results…</p>\n<p>Version 6: 0.61 LB \"Baseline\"<br>\nAn existing model I've used / is used for all other versions in this test.<br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment</a></p>\n<p>Version 7: 0.50 LB<br>\nI select a random number between 0 and 1 for each soundscape - and added it to all the predictions for that soundscape. <br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175709383\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175709383</a></p>\n<p>Version 8: 0.61 LB (\"Best Score\"?!)<br>\nIn this notebook I choose a random number between 0.5 and 1 for each soundscape - and then -multiply- each prediction in that soundscape by it.  So - each soundscape will randomly have its score scaled to between 50% and 100% of what the model predicted. <br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175888223\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175888223</a></p>\n<p>Version 9: 0.60 LB<br>\nHere we scale each soundscapes prediction by a multiplier between 0.1 and 1.  Each soundscape has all it's scored scaled between 10% and 100% of what the model predicted.<br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175899536\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175899536</a></p>\n<p>So - I'm not 100% sure what this means - but maybe there are some clues about the nature of the test data / scoring.  I think my code is correct…</p>\n<p>0.50 LB score in Version 7 is what I'd expect for combined scoring.</p>\n<p>But.. I find that scaling predictions per-soundscape over a factor of 10x would have relatively little impact on score surprising.  That could be explained if separated-by-soundscape scoring was used.  Alternately - maybe it indicates something about the data (or is maybe a side-effect of doing these tests on a relatively-low-scoring model).</p>\n<p>Anyways - tossing it out there for discussion….</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2793825": "So - the competition describes: \n\"a version of macro-averaged ROC-AUC that skips classes which have no true positive labels.\"\nhttps://www.kaggle.com/code/metric/birdclef-roc-auc\nhttps://www.kaggle.com/code/richolson/birdclef-2024-exploring-scoring\n\nThat code would seem to suggest a single DF goes in for scoring - and then kicks out an LB score.\n\nOn that assumption - all the soundscapes would be scored together.\n\nIn this case - for a class to be \"skipped\" - it must not have any true positive labels in any of the soundscapes.  Predicting a Purple Pigeon in a soundscape with no Purple Pigeons will result in your score being penalized (assuming another soundscape has a Purple Pigeon).  Scores for a given species would be relative to all scores for that species across soundscapes.\n\nAnother possibility is for each soundscape to be scored separately - and then have the scores for the soundscapes averaged together.\n\nIn that case - any species not present in a given soundscape is exempt from scoring for that soundscape.  Further - predictions for the same species across different soundscapes could be scaled or offset differently from each other without impact on the LB score.\n\nDoes anyone have any thoughts on which it is (or something different)? Or better yet - is there something documented for the competition?",
    "2793902": "As far as I understand, complete results (`solution` and `submission`) are submitted to calculate the metric. That is, all the soundscapes.",
    "2795881": "OK - so I did some tests - and I got some curious results...\n\nVersion 6: 0.61 LB \"Baseline\"\nAn existing model I've used / is used for all other versions in this test.\nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment\n\nVersion 7: 0.50 LB\nI select a random number between 0 and 1 for each soundscape - and added it to all the predictions for that soundscape. \nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175709383\n\nVersion 8: 0.61 LB (\"Best Score\"?!)\nIn this notebook I choose a random number between 0.5 and 1 for each soundscape - and then -multiply- each prediction in that soundscape by it.  So - each soundscape will randomly have its score scaled to between 50% and 100% of what the model predicted. \nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175888223\n\nVersion 9: 0.60 LB\nHere we scale each soundscapes prediction by a multiplier between 0.1 and 1.  Each soundscape has all it's scored scaled between 10% and 100% of what the model predicted.\nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-scoring-experiment?scriptVersionId=175899536\n\nSo - I'm not 100% sure what this means - but maybe there are some clues about the nature of the test data / scoring.  I think my code is correct...\n\n0.50 LB score in Version 7 is what I'd expect for combined scoring.\n\nBut.. I find that scaling predictions per-soundscape over a factor of 10x would have relatively little impact on score surprising.  That could be explained if separated-by-soundscape scoring was used.  Alternately - maybe it indicates something about the data (or is maybe a side-effect of doing these tests on a relatively-low-scoring model).\n\nAnyways - tossing it out there for discussion...."
  },
  "source": "meta"
}