{
  "id": 308009,
  "title": "Submission Clarification - Time Chunks and Predictions",
  "url": "/competitions/birdclef-2022/discussion/308009",
  "author_name": "",
  "post_date": "2022-02-16T16:59:50.954915Z",
  "votes": 13,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Awesome competition! Quick clarification question/concern though for <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/amandanavine\" target=\"_blank\">@amandanavine</a> or anyone else with good intuition. Figured I would ask this here in discussions as I figure it may come up in other comments.</p>\n<p>1st:</p>\n<ul>\n<li>Are the test sounds annotated in 5 second chunks or when bird calls are present?<ul>\n<li>In other words are we trying to predict instances of bird calls whenever they occur or whether bird calls exist in 5 second segments of audio files?</li></ul></li>\n</ul>\n<p>I'm trying to determine the best approach for dealing with birdcalls that are present between/across Time Chunks. Visually, this is the scenario I am trying to plan for:</p>\n<p>2nd: </p>\n<p><a href=\"https://postimg.cc/FfQnZHt4\" target=\"_blank\"><img src=\"https://i.postimg.cc/GpBwb4qH/67d1b906af4f88b2b1cc1687daecdbdf.jpg\" alt=\"67d1b906af4f88b2b1cc1687daecdbdf.jpg\"></a></p>\n<p>Would an accurate submission assume the target class is in both time chunks, neither, or potentially one and not the other? </p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>soundscape_453028782_greatfrigate_5</td>\n<td>True</td>\n</tr>\n<tr>\n<td>soundscape_453028782_greatfrigate_10</td>\n<td>True</td>\n</tr>\n</tbody>\n</table>\n<p>or </p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>soundscape_453028782_greatfrigate_5</td>\n<td>False</td>\n</tr>\n<tr>\n<td>soundscape_453028782_greatfrigate_10</td>\n<td>False</td>\n</tr>\n</tbody>\n</table>\n<p>or </p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>soundscape_453028782_greatfrigate_5</td>\n<td>False</td>\n</tr>\n<tr>\n<td>soundscape_453028782_greatfrigate_10</td>\n<td>True</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "1693427",
      "postDate": "02/16/2022 16:59:50",
      "content": "<p>Awesome competition! Quick clarification question/concern though for <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/amandanavine\" target=\"_blank\">@amandanavine</a> or anyone else with good intuition. Figured I would ask this here in discussions as I figure it may come up in other comments.</p>\n<p>1st:</p>\n<ul>\n<li>Are the test sounds annotated in 5 second chunks or when bird calls are present?<ul>\n<li>In other words are we trying to predict instances of bird calls whenever they occur or whether bird calls exist in 5 second segments of audio files?</li></ul></li>\n</ul>\n<p>I'm trying to determine the best approach for dealing with birdcalls that are present between/across Time Chunks. Visually, this is the scenario I am trying to plan for:</p>\n<p>2nd: </p>\n<p><a href=\"https://postimg.cc/FfQnZHt4\" target=\"_blank\"><img src=\"https://i.postimg.cc/GpBwb4qH/67d1b906af4f88b2b1cc1687daecdbdf.jpg\" alt=\"67d1b906af4f88b2b1cc1687daecdbdf.jpg\"></a></p>\n<p>Would an accurate submission assume the target class is in both time chunks, neither, or potentially one and not the other? </p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>soundscape_453028782_greatfrigate_5</td>\n<td>True</td>\n</tr>\n<tr>\n<td>soundscape_453028782_greatfrigate_10</td>\n<td>True</td>\n</tr>\n</tbody>\n</table>\n<p>or </p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>soundscape_453028782_greatfrigate_5</td>\n<td>False</td>\n</tr>\n<tr>\n<td>soundscape_453028782_greatfrigate_10</td>\n<td>False</td>\n</tr>\n</tbody>\n</table>\n<p>or </p>\n<table>\n<thead>\n<tr>\n<th>row_id</th>\n<th>target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>soundscape_453028782_greatfrigate_5</td>\n<td>False</td>\n</tr>\n<tr>\n<td>soundscape_453028782_greatfrigate_10</td>\n<td>True</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Awesome competition! Quick clarification question/concern though for @stefankahl @amandanavine or anyone else with good intuition. Figured I would ask this here in discussions as I figure it may come up in other comments.\n\n1st:\n- Are the test sounds annotated in 5 second chunks or when bird calls are present?\n - In other words are we trying to predict instances of bird calls whenever they occur or whether bird calls exist in 5 second segments of audio files?\n\nI'm trying to determine the best approach for dealing with birdcalls that are present between/across Time Chunks. Visually, this is the scenario I am trying to plan for:\n\n2nd: \n\n[![67d1b906af4f88b2b1cc1687daecdbdf.jpg](https://i.postimg.cc/GpBwb4qH/67d1b906af4f88b2b1cc1687daecdbdf.jpg)](https://postimg.cc/FfQnZHt4)\n\nWould an accurate submission assume the target class is in both time chunks, neither, or potentially one and not the other? \n\n| row_id | target |\n| --- | --- |\n| soundscape_453028782_greatfrigate_5 | True |\n| soundscape_453028782_greatfrigate_10 | True |\n\nor \n\n| row_id | target |\n| --- | --- |\n| soundscape_453028782_greatfrigate_5 | False |\n| soundscape_453028782_greatfrigate_10 | False |\n\nor \n\n| row_id | target |\n| --- | --- |\n| soundscape_453028782_greatfrigate_5 | False |\n| soundscape_453028782_greatfrigate_10 | True |",
      "votes": null
    },
    {
      "id": "1694212",
      "postDate": "02/17/2022 08:55:42",
      "content": "<p>Good question! Let me try to clarify a bit: </p>\n<ul>\n<li>For all 5-second segments of audio in the test data, we know if one of the scored birds occurs or not. Labels are for 5-second chunks, not call occurrences.</li>\n<li>If a bird call is present across two segments, both segments were labeled as \"True\" for this bird.</li>\n<li>However, overlap between segment and bird call needs to be at least 1 second. If this is not the case, e.g. when the call is audible in segment A for 3.5 seconds and only 0.5 seconds in segment B, the label would be \"True\" for segment A and \"False\" for segment B (but be aware, this is very rarely the case and should statistically be neglectable). </li>\n<li>Yet, if the entire call is within the segment, the length of the call does not matter, it can be only a fraction of a second long, the label would still be \"True\".</li>\n</ul>\n<p>Hope that helps, let me know in a comment, if you need more details. In your provided example, answer 1 would be the correct choice.</p>",
      "rawMarkdown": "Good question! Let me try to clarify a bit: \n\n- For all 5-second segments of audio in the test data, we know if one of the scored birds occurs or not. Labels are for 5-second chunks, not call occurrences.\n- If a bird call is present across two segments, both segments were labeled as \"True\" for this bird.\n- However, overlap between segment and bird call needs to be at least 1 second. If this is not the case, e.g. when the call is audible in segment A for 3.5 seconds and only 0.5 seconds in segment B, the label would be \"True\" for segment A and \"False\" for segment B (but be aware, this is very rarely the case and should statistically be neglectable). \n- Yet, if the entire call is within the segment, the length of the call does not matter, it can be only a fraction of a second long, the label would still be \"True\".\n\nHope that helps, let me know in a comment, if you need more details. In your provided example, answer 1 would be the correct choice.",
      "votes": null
    },
    {
      "id": "1694527",
      "postDate": "02/17/2022 13:50:10",
      "content": "<p>Thank you for your quick response <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> - perfectly clear now!</p>",
      "rawMarkdown": "Thank you for your quick response @stefankahl - perfectly clear now!",
      "votes": null
    },
    {
      "id": "1698496",
      "postDate": "02/20/2022 12:28:29",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> May I have one quick additional question? The description states:</p>\n<blockquote>\n  <p>For each row_id in the test set, you must predict a probability for the target variable.</p>\n</blockquote>\n<p>but the examples show binary classification in string format so would the following submission be also valid?</p>\n<pre><code>row_id,target\nsoundscape_1000170626_akiapo_5,0.25\nsoundscape_1000170626_akiapo_10,0.99\nsoundscape_1000170626_akiapo_15,0.78\netc.\n</code></pre>",
      "rawMarkdown": "stefankahl May I have one quick additional question? The description states:\n\n> For each row_id in the test set, you must predict a probability for the target variable.\n\nbut the examples show binary classification in string format so would the following submission be also valid?\n\n```\nrow_id,target\nsoundscape_1000170626_akiapo_5,0.25\nsoundscape_1000170626_akiapo_10,0.99\nsoundscape_1000170626_akiapo_15,0.78\netc.\n```",
      "votes": null
    },
    {
      "id": "1698753",
      "postDate": "02/20/2022 16:20:56",
      "content": "<p>I would assume that this format isn't valid, because it is unclear which cut-off threshold to use. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> should know more.</p>",
      "rawMarkdown": "I would assume that this format isn't valid, because it is unclear which cut-off threshold to use. @sohier should know more.",
      "votes": null
    },
    {
      "id": "1707907",
      "postDate": "02/28/2022 23:46:38",
      "content": "<p>Can there be more than one bird in a segment?</p>",
      "rawMarkdown": "Can there be more than one bird in a segment?",
      "votes": null
    },
    {
      "id": "1708175",
      "postDate": "03/01/2022 07:17:14",
      "content": "<p>said so I am a bit puzzled about the evaluation: <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/310335\" target=\"_blank\">Expected test data distribution</a></p>",
      "rawMarkdown": "said so I am a bit puzzled about the evaluation: [Expected test data distribution](https://www.kaggle.com/c/birdclef-2022/discussion/310335)",
      "votes": null
    },
    {
      "id": "1708176",
      "postDate": "03/01/2022 07:18:37",
      "content": "<p>with respect to the evaluation of fake data, I would say yes for sure…<br>\nsee <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/310335\" target=\"_blank\">Expected test data distribution</a></p>",
      "rawMarkdown": "with respect to the evaluation of fake data, I would say yes for sure...\nsee [Expected test data distribution](https://www.kaggle.com/c/birdclef-2022/discussion/310335)",
      "votes": null
    },
    {
      "id": "1708486",
      "postDate": "03/01/2022 13:40:04",
      "content": "<p>Yes, segments can have more than one label. Not sure how often that is actually the case, but despite only 21 target species, it's a multi-label classification problem.</p>",
      "rawMarkdown": "Yes, segments can have more than one label. Not sure how often that is actually the case, but despite only 21 target species, it's a multi-label classification problem.",
      "votes": null
    },
    {
      "id": "1709840",
      "postDate": "03/02/2022 14:25:08",
      "content": "<blockquote>\n  <p>but the examples show binary classification in string format so would the following submission be also valid?</p>\n</blockquote>\n<p>to comply with submission format you need to select a threshold (eg, 0.5) and turn your predictions to binary outcomes, i.e.  </p>\n<pre><code>row_id,target\nsoundscape_1000170626_akiapo_5, False  ### 0.25&gt;0.5\nsoundscape_1000170626_akiapo_10, True  ### 0.99&gt;0.5 \nsoundscape_1000170626_akiapo_15, True. ### 0.78&gt;0.5\netc.\n</code></pre>",
      "rawMarkdown": "> but the examples show binary classification in string format so would the following submission be also valid?\n\nto comply with submission format you need to select a threshold (eg, 0.5) and turn your predictions to binary outcomes, i.e.  \n\n```\nrow_id,target\nsoundscape_1000170626_akiapo_5, False  ### 0.25>0.5\nsoundscape_1000170626_akiapo_10, True  ### 0.99>0.5 \nsoundscape_1000170626_akiapo_15, True. ### 0.78>0.5\netc.\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1694212,
      "author_name": "stefankahl",
      "author_url": "",
      "post_date": "02/17/2022 08:55:42",
      "content": "<p>Good question! Let me try to clarify a bit: </p>\n<ul>\n<li>For all 5-second segments of audio in the test data, we know if one of the scored birds occurs or not. Labels are for 5-second chunks, not call occurrences.</li>\n<li>If a bird call is present across two segments, both segments were labeled as \"True\" for this bird.</li>\n<li>However, overlap between segment and bird call needs to be at least 1 second. If this is not the case, e.g. when the call is audible in segment A for 3.5 seconds and only 0.5 seconds in segment B, the label would be \"True\" for segment A and \"False\" for segment B (but be aware, this is very rarely the case and should statistically be neglectable). </li>\n<li>Yet, if the entire call is within the segment, the length of the call does not matter, it can be only a fraction of a second long, the label would still be \"True\".</li>\n</ul>\n<p>Hope that helps, let me know in a comment, if you need more details. In your provided example, answer 1 would be the correct choice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1694527,
          "author_name": "alexteboul",
          "author_url": "",
          "post_date": "02/17/2022 13:50:10",
          "content": "<p>Thank you for your quick response <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> - perfectly clear now!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1707907,
          "author_name": "ambrusattila",
          "author_url": "",
          "post_date": "02/28/2022 23:46:38",
          "content": "<p>Can there be more than one bird in a segment?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1708176,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/01/2022 07:18:37",
          "content": "<p>with respect to the evaluation of fake data, I would say yes for sure…<br>\nsee <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/310335\" target=\"_blank\">Expected test data distribution</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1708486,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "03/01/2022 13:40:04",
          "content": "<p>Yes, segments can have more than one label. Not sure how often that is actually the case, but despite only 21 target species, it's a multi-label classification problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1698496,
      "author_name": "jirkaborovec",
      "author_url": "",
      "post_date": "02/20/2022 12:28:29",
      "content": "<p><a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> May I have one quick additional question? The description states:</p>\n<blockquote>\n  <p>For each row_id in the test set, you must predict a probability for the target variable.</p>\n</blockquote>\n<p>but the examples show binary classification in string format so would the following submission be also valid?</p>\n<pre><code>row_id,target\nsoundscape_1000170626_akiapo_5,0.25\nsoundscape_1000170626_akiapo_10,0.99\nsoundscape_1000170626_akiapo_15,0.78\netc.\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1698753,
          "author_name": "stefankahl",
          "author_url": "",
          "post_date": "02/20/2022 16:20:56",
          "content": "<p>I would assume that this format isn't valid, because it is unclear which cut-off threshold to use. <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> should know more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1708175,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/01/2022 07:17:14",
          "content": "<p>said so I am a bit puzzled about the evaluation: <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/310335\" target=\"_blank\">Expected test data distribution</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1709840,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "03/02/2022 14:25:08",
          "content": "<blockquote>\n  <p>but the examples show binary classification in string format so would the following submission be also valid?</p>\n</blockquote>\n<p>to comply with submission format you need to select a threshold (eg, 0.5) and turn your predictions to binary outcomes, i.e.  </p>\n<pre><code>row_id,target\nsoundscape_1000170626_akiapo_5, False  ### 0.25&gt;0.5\nsoundscape_1000170626_akiapo_10, True  ### 0.99&gt;0.5 \nsoundscape_1000170626_akiapo_15, True. ### 0.78&gt;0.5\netc.\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1693427": "Awesome competition! Quick clarification question/concern though for @stefankahl @amandanavine or anyone else with good intuition. Figured I would ask this here in discussions as I figure it may come up in other comments.\n\n1st:\n- Are the test sounds annotated in 5 second chunks or when bird calls are present?\n - In other words are we trying to predict instances of bird calls whenever they occur or whether bird calls exist in 5 second segments of audio files?\n\nI'm trying to determine the best approach for dealing with birdcalls that are present between/across Time Chunks. Visually, this is the scenario I am trying to plan for:\n\n2nd: \n\n[![67d1b906af4f88b2b1cc1687daecdbdf.jpg](https://i.postimg.cc/GpBwb4qH/67d1b906af4f88b2b1cc1687daecdbdf.jpg)](https://postimg.cc/FfQnZHt4)\n\nWould an accurate submission assume the target class is in both time chunks, neither, or potentially one and not the other? \n\n| row_id | target |\n| --- | --- |\n| soundscape_453028782_greatfrigate_5 | True |\n| soundscape_453028782_greatfrigate_10 | True |\n\nor \n\n| row_id | target |\n| --- | --- |\n| soundscape_453028782_greatfrigate_5 | False |\n| soundscape_453028782_greatfrigate_10 | False |\n\nor \n\n| row_id | target |\n| --- | --- |\n| soundscape_453028782_greatfrigate_5 | False |\n| soundscape_453028782_greatfrigate_10 | True |",
    "1694212": "Good question! Let me try to clarify a bit: \n\n- For all 5-second segments of audio in the test data, we know if one of the scored birds occurs or not. Labels are for 5-second chunks, not call occurrences.\n- If a bird call is present across two segments, both segments were labeled as \"True\" for this bird.\n- However, overlap between segment and bird call needs to be at least 1 second. If this is not the case, e.g. when the call is audible in segment A for 3.5 seconds and only 0.5 seconds in segment B, the label would be \"True\" for segment A and \"False\" for segment B (but be aware, this is very rarely the case and should statistically be neglectable). \n- Yet, if the entire call is within the segment, the length of the call does not matter, it can be only a fraction of a second long, the label would still be \"True\".\n\nHope that helps, let me know in a comment, if you need more details. In your provided example, answer 1 would be the correct choice.",
    "1694527": "Thank you for your quick response @stefankahl - perfectly clear now!",
    "1698496": "stefankahl May I have one quick additional question? The description states:\n\n> For each row_id in the test set, you must predict a probability for the target variable.\n\nbut the examples show binary classification in string format so would the following submission be also valid?\n\n```\nrow_id,target\nsoundscape_1000170626_akiapo_5,0.25\nsoundscape_1000170626_akiapo_10,0.99\nsoundscape_1000170626_akiapo_15,0.78\netc.\n```",
    "1698753": "I would assume that this format isn't valid, because it is unclear which cut-off threshold to use. @sohier should know more.",
    "1707907": "Can there be more than one bird in a segment?",
    "1708175": "said so I am a bit puzzled about the evaluation: [Expected test data distribution](https://www.kaggle.com/c/birdclef-2022/discussion/310335)",
    "1708176": "with respect to the evaluation of fake data, I would say yes for sure...\nsee [Expected test data distribution](https://www.kaggle.com/c/birdclef-2022/discussion/310335)",
    "1708486": "Yes, segments can have more than one label. Not sure how often that is actually the case, but despite only 21 target species, it's a multi-label classification problem.",
    "1709840": "> but the examples show binary classification in string format so would the following submission be also valid?\n\nto comply with submission format you need to select a threshold (eg, 0.5) and turn your predictions to binary outcomes, i.e.  \n\n```\nrow_id,target\nsoundscape_1000170626_akiapo_5, False  ### 0.25>0.5\nsoundscape_1000170626_akiapo_10, True  ### 0.99>0.5 \nsoundscape_1000170626_akiapo_15, True. ### 0.78>0.5\netc.\n```"
  },
  "source": "meta"
}