{
  "id": 198235,
  "title": "In the submission, has the output score to be normalized? i.e. sum to 1?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/198235",
  "author_name": "",
  "post_date": "2020-11-20T10:49:17.843364500Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In the evaluation section of this competition it is reported: </p>\n<blockquote>\n  <p>For each recording_id in the test set, you must predict the probability of each species label being found in the audio sample. The file should contain a header (each species number with an s prefix) and have the following format:</p>\n</blockquote>\n<p>then in the <code>sample_submission.csv</code> each random component is set to 0.5.</p>\n<p>My question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?</p>\n<p>Thank you</p>",
  "messages": [
    {
      "id": "1084740",
      "postDate": "11/20/2020 10:49:17",
      "content": "<p>In the evaluation section of this competition it is reported: </p>\n<blockquote>\n  <p>For each recording_id in the test set, you must predict the probability of each species label being found in the audio sample. The file should contain a header (each species number with an s prefix) and have the following format:</p>\n</blockquote>\n<p>then in the <code>sample_submission.csv</code> each random component is set to 0.5.</p>\n<p>My question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?</p>\n<p>Thank you</p>",
      "rawMarkdown": "In the evaluation section of this competition it is reported: \n\n> For each recording_id in the test set, you must predict the probability of each species label being found in the audio sample. The file should contain a header (each species number with an s prefix) and have the following format:\n\nthen in the `sample_submission.csv` each random component is set to 0.5.\n\nMy question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?\n\nThank you",
      "votes": null
    },
    {
      "id": "1084830",
      "postDate": "11/20/2020 12:36:34",
      "content": "<p>Several labels can occur at the same time, then it makes sense that the total probability doesn't have to sum to 1.</p>\n<p>The lrap metric relies on the rank, then (as far as I know) normalization doesn't even matter.</p>",
      "rawMarkdown": "Several labels can occur at the same time, then it makes sense that the total probability doesn't have to sum to 1.\n\nThe lrap metric relies on the rank, then (as far as I know) normalization doesn't even matter.",
      "votes": null
    },
    {
      "id": "1084834",
      "postDate": "11/20/2020 12:38:59",
      "content": "<p><code>My question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?</code></p>\n<p>You can predict the classes with 0 or 1 but therefore your classifier has to be very good, it is sufficient to keep the predicted score within [0,1], you could also just implement a threshold value when its greater then this value you can pass 1 else 0, but for me it did decrease the auc-score in previous competitions when using a threshold. Thats why I stuck to values between [0,1] due to its error restistance.</p>\n<p>For more information about auc-score check: <a href=\"https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc\" target=\"_blank\">https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc</a></p>",
      "rawMarkdown": "`My question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?`\n\nYou can predict the classes with 0 or 1 but therefore your classifier has to be very good, it is sufficient to keep the predicted score within [0,1], you could also just implement a threshold value when its greater then this value you can pass 1 else 0, but for me it did decrease the auc-score in previous competitions when using a threshold. Thats why I stuck to values between [0,1] due to its error restistance.\n\nFor more information about auc-score check: https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc",
      "votes": null
    },
    {
      "id": "1084836",
      "postDate": "11/20/2020 12:42:07",
      "content": "<p>hey Ali,</p>\n<p>I think you misunderstood - I said [0, 1] not {0, 1}<br>\n:-)</p>",
      "rawMarkdown": "hey Ali,\n\nI think you misunderstood - I said [0, 1] not {0, 1}\n:-)",
      "votes": null
    },
    {
      "id": "1084837",
      "postDate": "11/20/2020 12:42:28",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "1084852",
      "postDate": "11/20/2020 13:03:42",
      "content": "<p>What do you mean by that exactly? </p>\n<p>[0,1] means to me an intervall between 0 and 1 including 0 and 1</p>",
      "rawMarkdown": "What do you mean by that exactly? \n\n[0,1] means to me an intervall between 0 and 1 including 0 and 1",
      "votes": null
    },
    {
      "id": "1085033",
      "postDate": "11/20/2020 16:03:50",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> is correct. If this were multi-class, it would have to sum to 1. But since it's multi-label, that constraint isn't there.</p>",
      "rawMarkdown": "theoviel is correct. If this were multi-class, it would have to sum to 1. But since it's multi-label, that constraint isn't there.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1084830,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/20/2020 12:36:34",
      "content": "<p>Several labels can occur at the same time, then it makes sense that the total probability doesn't have to sum to 1.</p>\n<p>The lrap metric relies on the rank, then (as far as I know) normalization doesn't even matter.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1084837,
          "author_name": "guglielmocamporese",
          "author_url": "",
          "post_date": "11/20/2020 12:42:28",
          "content": "<p>thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1085033,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "11/20/2020 16:03:50",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> is correct. If this were multi-class, it would have to sum to 1. But since it's multi-label, that constraint isn't there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1084834,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "11/20/2020 12:38:59",
      "content": "<p><code>My question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?</code></p>\n<p>You can predict the classes with 0 or 1 but therefore your classifier has to be very good, it is sufficient to keep the predicted score within [0,1], you could also just implement a threshold value when its greater then this value you can pass 1 else 0, but for me it did decrease the auc-score in previous competitions when using a threshold. Thats why I stuck to values between [0,1] due to its error restistance.</p>\n<p>For more information about auc-score check: <a href=\"https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc\" target=\"_blank\">https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1084836,
          "author_name": "guglielmocamporese",
          "author_url": "",
          "post_date": "11/20/2020 12:42:07",
          "content": "<p>hey Ali,</p>\n<p>I think you misunderstood - I said [0, 1] not {0, 1}<br>\n:-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084852,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/20/2020 13:03:42",
          "content": "<p>What do you mean by that exactly? </p>\n<p>[0,1] means to me an intervall between 0 and 1 including 0 and 1</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1084740": "In the evaluation section of this competition it is reported: \n\n> For each recording_id in the test set, you must predict the probability of each species label being found in the audio sample. The file should contain a header (each species number with an s prefix) and have the following format:\n\nthen in the `sample_submission.csv` each random component is set to 0.5.\n\nMy question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?\n\nThank you",
    "1084830": "Several labels can occur at the same time, then it makes sense that the total probability doesn't have to sum to 1.\n\nThe lrap metric relies on the rank, then (as far as I know) normalization doesn't even matter.",
    "1084834": "`My question is: has the output probably to be normalized (sum to 1) or it is sufficient to just have it in [0, 1]? Does this affect the leaderboard score?`\n\nYou can predict the classes with 0 or 1 but therefore your classifier has to be very good, it is sufficient to keep the predicted score within [0,1], you could also just implement a threshold value when its greater then this value you can pass 1 else 0, but for me it did decrease the auc-score in previous competitions when using a threshold. Thats why I stuck to values between [0,1] due to its error restistance.\n\nFor more information about auc-score check: https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc",
    "1084836": "hey Ali,\n\nI think you misunderstood - I said [0, 1] not {0, 1}\n:-)",
    "1084837": "thank you!",
    "1084852": "What do you mean by that exactly? \n\n[0,1] means to me an intervall between 0 and 1 including 0 and 1",
    "1085033": "theoviel is correct. If this were multi-class, it would have to sum to 1. But since it's multi-label, that constraint isn't there."
  },
  "source": "meta"
}