{
  "id": 503196,
  "title": "Hidden test data has label for every 5 second data?",
  "url": "/competitions/birdclef-2024/discussion/503196",
  "author_name": "",
  "post_date": "2024-05-16T12:01:45.945757Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I can see from the data that this is a multi-label problem. However, the submission file expects probabilities for all species every 5 seconds of the data. In that case, I assume the hidden test data has labels every 5 seconds.</p>\n<p>Please correct me if my understanding is wrong.</p>",
  "messages": [
    {
      "id": "2816575",
      "postDate": "05/16/2024 12:01:45",
      "content": "<p>I can see from the data that this is a multi-label problem. However, the submission file expects probabilities for all species every 5 seconds of the data. In that case, I assume the hidden test data has labels every 5 seconds.</p>\n<p>Please correct me if my understanding is wrong.</p>",
      "rawMarkdown": "I can see from the data that this is a multi-label problem. However, the submission file expects probabilities for all species every 5 seconds of the data. In that case, I assume the hidden test data has labels every 5 seconds.\n\nPlease correct me if my understanding is wrong.",
      "votes": null
    },
    {
      "id": "2816603",
      "postDate": "05/16/2024 12:27:26",
      "content": "<p>It's possible for there to be a \"no call\" chunk that gets all zeros</p>\n<p>Edit:<br>\nThis is at least my understanding. If this is wrong please advise.</p>",
      "rawMarkdown": "It's possible for there to be a \"no call\" chunk that gets all zeros\n\nEdit:\nThis is at least my understanding. If this is wrong please advise.",
      "votes": null
    },
    {
      "id": "2816805",
      "postDate": "05/16/2024 14:40:49",
      "content": "<p>Yes. Wel, we are asked to provide predictions of every 5 seconds slice. I assume they are all used.</p>",
      "rawMarkdown": "Yes. Wel, we are asked to provide predictions of every 5 seconds slice. I assume they are all used.",
      "votes": null
    },
    {
      "id": "2817679",
      "postDate": "05/17/2024 04:25:45",
      "content": "<p>Yes, all 5 second lines exist in the markup. It's just that some lines may contain all zeros. And in some there may be several 1. (judging by past similar competitions, when different birds sing in this interval). Therefore, it <strong>seems</strong> to me that it is better to predict the probability of each class (in this case, the sum of the probabilities in the row will not be equal to 1)</p>",
      "rawMarkdown": "Yes, all 5 second lines exist in the markup. It's just that some lines may contain all zeros. And in some there may be several 1. (judging by past similar competitions, when different birds sing in this interval). Therefore, it **seems** to me that it is better to predict the probability of each class (in this case, the sum of the probabilities in the row will not be equal to 1)",
      "votes": null
    },
    {
      "id": "2818307",
      "postDate": "05/17/2024 12:33:04",
      "content": "<p>It is weird to see downvotes for the only relevant answer to the post so far.</p>\n<p>We don't have explicit evidence that every row in the submission is used for scoring. We only have the indirect evidence that we have to supply predictions for each row. I assume, like the OP, that all our predictions are used, which would mean that there is a label for each row.</p>\n<p>The OP clearly states that the problem is a multilabel one, hence this is not the topic to clarify.</p>",
      "rawMarkdown": "It is weird to see downvotes for the only relevant answer to the post so far.\n\nWe don't have explicit evidence that every row in the submission is used for scoring. We only have the indirect evidence that we have to supply predictions for each row. I assume, like the OP, that all our predictions are used, which would mean that there is a label for each row.\n\nThe OP clearly states that the problem is a multilabel one, hence this is not the topic to clarify.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2816603,
      "author_name": "willrice",
      "author_url": "",
      "post_date": "05/16/2024 12:27:26",
      "content": "<p>It's possible for there to be a \"no call\" chunk that gets all zeros</p>\n<p>Edit:<br>\nThis is at least my understanding. If this is wrong please advise.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2816805,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/16/2024 14:40:49",
      "content": "<p>Yes. Wel, we are asked to provide predictions of every 5 seconds slice. I assume they are all used.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2818307,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/17/2024 12:33:04",
          "content": "<p>It is weird to see downvotes for the only relevant answer to the post so far.</p>\n<p>We don't have explicit evidence that every row in the submission is used for scoring. We only have the indirect evidence that we have to supply predictions for each row. I assume, like the OP, that all our predictions are used, which would mean that there is a label for each row.</p>\n<p>The OP clearly states that the problem is a multilabel one, hence this is not the topic to clarify.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2817679,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "05/17/2024 04:25:45",
      "content": "<p>Yes, all 5 second lines exist in the markup. It's just that some lines may contain all zeros. And in some there may be several 1. (judging by past similar competitions, when different birds sing in this interval). Therefore, it <strong>seems</strong> to me that it is better to predict the probability of each class (in this case, the sum of the probabilities in the row will not be equal to 1)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2816575": "I can see from the data that this is a multi-label problem. However, the submission file expects probabilities for all species every 5 seconds of the data. In that case, I assume the hidden test data has labels every 5 seconds.\n\nPlease correct me if my understanding is wrong.",
    "2816603": "It's possible for there to be a \"no call\" chunk that gets all zeros\n\nEdit:\nThis is at least my understanding. If this is wrong please advise.",
    "2816805": "Yes. Wel, we are asked to provide predictions of every 5 seconds slice. I assume they are all used.",
    "2817679": "Yes, all 5 second lines exist in the markup. It's just that some lines may contain all zeros. And in some there may be several 1. (judging by past similar competitions, when different birds sing in this interval). Therefore, it **seems** to me that it is better to predict the probability of each class (in this case, the sum of the probabilities in the row will not be equal to 1)",
    "2818307": "It is weird to see downvotes for the only relevant answer to the post so far.\n\nWe don't have explicit evidence that every row in the submission is used for scoring. We only have the indirect evidence that we have to supply predictions for each row. I assume, like the OP, that all our predictions are used, which would mean that there is a label for each row.\n\nThe OP clearly states that the problem is a multilabel one, hence this is not the topic to clarify."
  },
  "source": "meta"
}