{
  "id": 202456,
  "title": "BBoxes on Time and Frequency - Are [t_min, t_max, f_min, f_max] Always Correct?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/202456",
  "author_name": "",
  "post_date": "2020-12-10T03:12:02.215972500Z",
  "votes": 37,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I can share some patterns that I see when I visualize the bounding box [<strong>t_min</strong>, <strong>t_max</strong>, <strong>f_min</strong>, <strong>f_max</strong>] given by the competition. Here I reported the log-mel spectrum of some interesting samples (<strong>f_min</strong> and <strong>f_max</strong> are transformed from the frequency to the <em>mel</em> domain).</p>\n<table>\n<thead>\n<tr>\n<th><strong>Crops That Seem OK</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 0201197ec <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Ff9d7e937152ad7cce6025ae89f423c95%2FScreenshot%202020-12-10%20at%203.56.09%20AM.png?generation=1607569344764362&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>:  009b760e6 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F4519ccf76b4348da311f426b715c501b%2FScreenshot%202020-12-10%20at%203.53.17%20AM.png?generation=1607569172265426&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th><strong>Crop Size Along Time Seems Too Small</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 0099c367b <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F9c648d8b99652b9c9b531133b17a2960%2FScreenshot%202020-12-10%20at%203.53.13%20AM.png?generation=1607569440726620&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>: 01b41f92b <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F09952b81563f8b6775a504a3a4a871f6%2FScreenshot%202020-12-10%20at%203.53.43%20AM.png?generation=1607569507395742&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th><strong>Crop On Frequency Seems Shifted</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 0313e82cf <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F670d7651f7a51fc8190f5d34df97eed0%2FScreenshot%202020-12-10%20at%203.54.11%20AM.png?generation=1607569593086409&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>: 03b96f209 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fb8df2426290b188216aca77633932b2e%2FScreenshot%202020-12-10%20at%203.54.16%20AM.png?generation=1607569640591224&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th><strong>Only Part Of The Pattern Seems Labeled</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 011f25080 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F58888c52ae8d4a4283e7bccace2f3541%2FScreenshot%202020-12-10%20at%203.53.26%20AM.png?generation=1607569709906488&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>: 0295e3234 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fccac7cb9788d7c6f53dd5eb8559b4f57%2FScreenshot%202020-12-10%20at%203.53.56%20AM.png?generation=1607569766306295&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<p>Have you found any other pattern like these for the bboxes [<strong>t_min</strong>, <strong>t_max</strong>, <strong>f_min</strong>, <strong>f_max</strong>]?</p>",
  "messages": [
    {
      "id": "1107878",
      "postDate": "12/10/2020 03:12:02",
      "content": "<p>Hi everyone,</p>\n<p>I can share some patterns that I see when I visualize the bounding box [<strong>t_min</strong>, <strong>t_max</strong>, <strong>f_min</strong>, <strong>f_max</strong>] given by the competition. Here I reported the log-mel spectrum of some interesting samples (<strong>f_min</strong> and <strong>f_max</strong> are transformed from the frequency to the <em>mel</em> domain).</p>\n<table>\n<thead>\n<tr>\n<th><strong>Crops That Seem OK</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 0201197ec <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Ff9d7e937152ad7cce6025ae89f423c95%2FScreenshot%202020-12-10%20at%203.56.09%20AM.png?generation=1607569344764362&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>:  009b760e6 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F4519ccf76b4348da311f426b715c501b%2FScreenshot%202020-12-10%20at%203.53.17%20AM.png?generation=1607569172265426&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th><strong>Crop Size Along Time Seems Too Small</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 0099c367b <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F9c648d8b99652b9c9b531133b17a2960%2FScreenshot%202020-12-10%20at%203.53.13%20AM.png?generation=1607569440726620&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>: 01b41f92b <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F09952b81563f8b6775a504a3a4a871f6%2FScreenshot%202020-12-10%20at%203.53.43%20AM.png?generation=1607569507395742&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th><strong>Crop On Frequency Seems Shifted</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 0313e82cf <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F670d7651f7a51fc8190f5d34df97eed0%2FScreenshot%202020-12-10%20at%203.54.11%20AM.png?generation=1607569593086409&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>: 03b96f209 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fb8df2426290b188216aca77633932b2e%2FScreenshot%202020-12-10%20at%203.54.16%20AM.png?generation=1607569640591224&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th><strong>Only Part Of The Pattern Seems Labeled</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>rec_id</strong>: 011f25080 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F58888c52ae8d4a4283e7bccace2f3541%2FScreenshot%202020-12-10%20at%203.53.26%20AM.png?generation=1607569709906488&amp;alt=media\" alt=\"\"> <strong>rec_id</strong>: 0295e3234 <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fccac7cb9788d7c6f53dd5eb8559b4f57%2FScreenshot%202020-12-10%20at%203.53.56%20AM.png?generation=1607569766306295&amp;alt=media\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<p>Have you found any other pattern like these for the bboxes [<strong>t_min</strong>, <strong>t_max</strong>, <strong>f_min</strong>, <strong>f_max</strong>]?</p>",
      "rawMarkdown": "Hi everyone,\n\nI can share some patterns that I see when I visualize the bounding box [**t_min**, **t_max**, **f_min**, **f_max**] given by the competition. Here I reported the log-mel spectrum of some interesting samples (**f_min** and **f_max** are transformed from the frequency to the *mel* domain).\n\n| **Crops That Seem OK** |\n| --- |\n| **rec_id**: 0201197ec ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Ff9d7e937152ad7cce6025ae89f423c95%2FScreenshot%202020-12-10%20at%203.56.09%20AM.png?generation=1607569344764362&alt=media) **rec_id**:  009b760e6 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F4519ccf76b4348da311f426b715c501b%2FScreenshot%202020-12-10%20at%203.53.17%20AM.png?generation=1607569172265426&alt=media) |\n\n| **Crop Size Along Time Seems Too Small** |\n| --- |\n| **rec_id**: 0099c367b ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F9c648d8b99652b9c9b531133b17a2960%2FScreenshot%202020-12-10%20at%203.53.13%20AM.png?generation=1607569440726620&alt=media) **rec_id**: 01b41f92b ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F09952b81563f8b6775a504a3a4a871f6%2FScreenshot%202020-12-10%20at%203.53.43%20AM.png?generation=1607569507395742&alt=media)|\n\n| **Crop On Frequency Seems Shifted** |\n| --- |\n| **rec_id**: 0313e82cf ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F670d7651f7a51fc8190f5d34df97eed0%2FScreenshot%202020-12-10%20at%203.54.11%20AM.png?generation=1607569593086409&alt=media) **rec_id**: 03b96f209 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fb8df2426290b188216aca77633932b2e%2FScreenshot%202020-12-10%20at%203.54.16%20AM.png?generation=1607569640591224&alt=media)|\n\n| **Only Part Of The Pattern Seems Labeled** |\n| --- |\n| **rec_id**: 011f25080 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F58888c52ae8d4a4283e7bccace2f3541%2FScreenshot%202020-12-10%20at%203.53.26%20AM.png?generation=1607569709906488&alt=media) **rec_id**: 0295e3234 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fccac7cb9788d7c6f53dd5eb8559b4f57%2FScreenshot%202020-12-10%20at%203.53.56%20AM.png?generation=1607569766306295&alt=media) |\n\nHave you found any other pattern like these for the bboxes [**t_min**, **t_max**, **f_min**, **f_max**]?",
      "votes": null
    },
    {
      "id": "1108063",
      "postDate": "12/10/2020 08:22:09",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmocamporese\" target=\"_blank\">@guglielmocamporese</a> that is why they named the training file as the name <code>train_tp.csv</code> rather than <code>train.csv</code> :D. We do not know if the unlabeled segment contains species or not. </p>",
      "rawMarkdown": "guglielmocamporese that is why they named the training file as the name `train_tp.csv` rather than `train.csv` :D. We do not know if the unlabeled segment contains species or not.",
      "votes": null
    },
    {
      "id": "1108174",
      "postDate": "12/10/2020 11:13:49",
      "content": "<p>Upvote if you stance out the recordingIds… 😁</p>",
      "rawMarkdown": "Upvote if you stance out the recordingIds... 😁",
      "votes": null
    },
    {
      "id": "1108231",
      "postDate": "12/10/2020 12:40:53",
      "content": "<p>done! :)  </p>",
      "rawMarkdown": "done! :)",
      "votes": null
    },
    {
      "id": "1108328",
      "postDate": "12/10/2020 14:48:59",
      "content": "<p>Done, too! :)</p>\n<p>My first sound competition, so I'm trying to learn and with recordingId it makes my life easier in interpreting the data.</p>",
      "rawMarkdown": "Done, too! :)\n\nMy first sound competition, so I'm trying to learn and with recordingId it makes my life easier in interpreting the data.",
      "votes": null
    },
    {
      "id": "1108742",
      "postDate": "12/11/2020 00:57:32",
      "content": "<p>thanks for sharing. not sure if it is always correct though…. i saw there are some tp segments with very short duration like 0.3s, i wonder 1) how did a person label these with such accuracy/precision 2) the short duration shows up as a very strip in mel-spectrogram, how could the model deal with that? there are <strong>plenty</strong> place for the model to overfit/memorize on</p>\n<p>if you dont mind sharing, knowing some patterns of these boxes (t_min max, f_min max), how could it potentially help us? </p>",
      "rawMarkdown": "thanks for sharing. not sure if it is always correct though.... i saw there are some tp segments with very short duration like 0.3s, i wonder 1) how did a person label these with such accuracy/precision 2) the short duration shows up as a very strip in mel-spectrogram, how could the model deal with that? there are **plenty** place for the model to overfit/memorize on\n\nif you dont mind sharing, knowing some patterns of these boxes (t_min max, f_min max), how could it potentially help us?",
      "votes": null
    },
    {
      "id": "1109741",
      "postDate": "12/12/2020 02:58:01",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmocamporese\" target=\"_blank\">@guglielmocamporese</a> as I generally understand the 'manual' id process the scientists that know the <em>call</em> make the <em>call</em> as to the given species, both as to species and length of representative vocalization. In the case of train.tp, this might be the most representative. </p>\n<p>Extending into train.fp might either be who gets confused with whom (as species) or which species are generally in the same habitats and overlap/step on each other's frequencies, or perhaps mimic. The data, as presented, doesn't readily lend itself to permutation analysis of birds of a habitat habitat together.  </p>\n<p>But, one imagines that another expert has determined the segment is a false positive determined to be positive by what: Another algorithm, or by whom: Another expert. In either case, algo or expert, what will our algo make of this inter-rater disagreement, and do we even want to go there?  I don't know, but these are things I wonder about in the presence of warmed data.</p>\n<p><em>mel</em> is a transform that is applied to the data in the pre-transformed rec_id: 0295e3234 for example. Is there a way to extract the post-transformed corners without reference to the pre-transformed, then diff them? And apply the diff to the label, perhaps? </p>\n<p>Well, sorry for a bunch of speculative statements </p>\n<p>Chris</p>",
      "rawMarkdown": "guglielmocamporese as I generally understand the 'manual' id process the scientists that know the *call* make the *call* as to the given species, both as to species and length of representative vocalization. In the case of train.tp, this might be the most representative. \n\nExtending into train.fp might either be who gets confused with whom (as species) or which species are generally in the same habitats and overlap/step on each other's frequencies, or perhaps mimic. The data, as presented, doesn't readily lend itself to permutation analysis of birds of a habitat habitat together.  \n\nBut, one imagines that another expert has determined the segment is a false positive determined to be positive by what: Another algorithm, or by whom: Another expert. In either case, algo or expert, what will our algo make of this inter-rater disagreement, and do we even want to go there?  I don't know, but these are things I wonder about in the presence of warmed data.\n\n*mel* is a transform that is applied to the data in the pre-transformed rec_id: 0295e3234 for example. Is there a way to extract the post-transformed corners without reference to the pre-transformed, then diff them? And apply the diff to the label, perhaps? \n\nWell, sorry for a bunch of speculative statements \n\nChris",
      "votes": null
    },
    {
      "id": "1117950",
      "postDate": "12/18/2020 16:00:39",
      "content": "<p>Yes. I've seen many such cases. Only one of the repeated patterns is labeled. Apparently, the localization information is not complete, and I think that's why time localization is not part of the test predictions. This situation will raise an issue when we cut the clips and creating train labels according to the localization information. Not sure how to improve this yet.</p>",
      "rawMarkdown": "Yes. I've seen many such cases. Only one of the repeated patterns is labeled. Apparently, the localization information is not complete, and I think that's why time localization is not part of the test predictions. This situation will raise an issue when we cut the clips and creating train labels according to the localization information. Not sure how to improve this yet.",
      "votes": null
    },
    {
      "id": "1122761",
      "postDate": "12/22/2020 17:28:36",
      "content": "<p>I've noticed that for train_tp the species is present in the recording somewhere and or in multiple parts of the recording but not necessarily the time segment that is marked as having the species.</p>",
      "rawMarkdown": "I've noticed that for train_tp the species is present in the recording somewhere and or in multiple parts of the recording but not necessarily the time segment that is marked as having the species.",
      "votes": null
    },
    {
      "id": "1122862",
      "postDate": "12/22/2020 18:46:24",
      "content": "<p>Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings</p>",
      "rawMarkdown": "Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings",
      "votes": null
    },
    {
      "id": "1122892",
      "postDate": "12/22/2020 19:07:22",
      "content": "<blockquote>\n  <p>Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings</p>\n</blockquote>\n<p>I'm using the log-mel-spectrogram (with log10) so I'm considering the dB scale, and I'm normalizing the samples globally, with respect to the entire training set.</p>\n<p>Best,</p>\n<p>Guglielmo</p>",
      "rawMarkdown": "> Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings\n\nI'm using the log-mel-spectrogram (with log10) so I'm considering the dB scale, and I'm normalizing the samples globally, with respect to the entire training set.\n\nBest,\n\nGuglielmo",
      "votes": null
    },
    {
      "id": "1125800",
      "postDate": "12/25/2020 04:47:54",
      "content": "<p>I've noticed some weird recordings as well (orange rectangle is TP, red rectangles is FP):</p>\n<hr>\n<p><strong>rec_id:</strong> 7448edc55<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2Ff4a31d9522a3ac0c7c59c8b4ebb5fa87%2Fmel_1.png?generation=1608871279214045&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>rec_Id:</strong> de80fb815<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2F15c5682f0f703b2ef0093cd10189292b%2Fmel_2.png?generation=1608871601733051&amp;alt=media\" alt=\"\"></p>\n<p>I'll keep this comment updated with weird songs that I will find in the future.</p>",
      "rawMarkdown": "I've noticed some weird recordings as well (orange rectangle is TP, red rectangles is FP):\n\n----------------------------------------------------------------------\n\n**rec_id:** 7448edc55\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2Ff4a31d9522a3ac0c7c59c8b4ebb5fa87%2Fmel_1.png?generation=1608871279214045&alt=media)\n\n----------------------------------------------------------------------\n\n**rec_Id:** de80fb815\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2F15c5682f0f703b2ef0093cd10189292b%2Fmel_2.png?generation=1608871601733051&alt=media)\n\nI'll keep this comment updated with weird songs that I will find in the future.",
      "votes": null
    },
    {
      "id": "1147922",
      "postDate": "01/10/2021 20:01:21",
      "content": "<p><a href=\"https://www.kaggle.com/hav4ik\" target=\"_blank\">@hav4ik</a> thanks for sharing, does all the vertical line in the first image you shared make sense to you? I listened to its recording and it almost sounds like some audio glitches</p>",
      "rawMarkdown": "hav4ik thanks for sharing, does all the vertical line in the first image you shared make sense to you? I listened to its recording and it almost sounds like some audio glitches",
      "votes": null
    },
    {
      "id": "1182155",
      "postDate": "02/02/2021 10:33:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/guglielmocamporese\" target=\"_blank\">@guglielmocamporese</a>, could you please share the code that draws these great pictures?</p>",
      "rawMarkdown": "Hi @guglielmocamporese, could you please share the code that draws these great pictures?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1108063,
      "author_name": "dathudeptrai",
      "author_url": "",
      "post_date": "12/10/2020 08:22:09",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmocamporese\" target=\"_blank\">@guglielmocamporese</a> that is why they named the training file as the name <code>train_tp.csv</code> rather than <code>train.csv</code> :D. We do not know if the unlabeled segment contains species or not. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1108174,
      "author_name": "glimmung",
      "author_url": "",
      "post_date": "12/10/2020 11:13:49",
      "content": "<p>Upvote if you stance out the recordingIds… 😁</p>",
      "votes": null,
      "replies": [
        {
          "id": 1108231,
          "author_name": "guglielmocamporese",
          "author_url": "",
          "post_date": "12/10/2020 12:40:53",
          "content": "<p>done! :)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108328,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "12/10/2020 14:48:59",
          "content": "<p>Done, too! :)</p>\n<p>My first sound competition, so I'm trying to learn and with recordingId it makes my life easier in interpreting the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1108742,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "12/11/2020 00:57:32",
      "content": "<p>thanks for sharing. not sure if it is always correct though…. i saw there are some tp segments with very short duration like 0.3s, i wonder 1) how did a person label these with such accuracy/precision 2) the short duration shows up as a very strip in mel-spectrogram, how could the model deal with that? there are <strong>plenty</strong> place for the model to overfit/memorize on</p>\n<p>if you dont mind sharing, knowing some patterns of these boxes (t_min max, f_min max), how could it potentially help us? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1109741,
      "author_name": "cenglish",
      "author_url": "",
      "post_date": "12/12/2020 02:58:01",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmocamporese\" target=\"_blank\">@guglielmocamporese</a> as I generally understand the 'manual' id process the scientists that know the <em>call</em> make the <em>call</em> as to the given species, both as to species and length of representative vocalization. In the case of train.tp, this might be the most representative. </p>\n<p>Extending into train.fp might either be who gets confused with whom (as species) or which species are generally in the same habitats and overlap/step on each other's frequencies, or perhaps mimic. The data, as presented, doesn't readily lend itself to permutation analysis of birds of a habitat habitat together.  </p>\n<p>But, one imagines that another expert has determined the segment is a false positive determined to be positive by what: Another algorithm, or by whom: Another expert. In either case, algo or expert, what will our algo make of this inter-rater disagreement, and do we even want to go there?  I don't know, but these are things I wonder about in the presence of warmed data.</p>\n<p><em>mel</em> is a transform that is applied to the data in the pre-transformed rec_id: 0295e3234 for example. Is there a way to extract the post-transformed corners without reference to the pre-transformed, then diff them? And apply the diff to the label, perhaps? </p>\n<p>Well, sorry for a bunch of speculative statements </p>\n<p>Chris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1117950,
      "author_name": "blaxkdolphin",
      "author_url": "",
      "post_date": "12/18/2020 16:00:39",
      "content": "<p>Yes. I've seen many such cases. Only one of the repeated patterns is labeled. Apparently, the localization information is not complete, and I think that's why time localization is not part of the test predictions. This situation will raise an issue when we cut the clips and creating train labels according to the localization information. Not sure how to improve this yet.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122761,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/22/2020 17:28:36",
      "content": "<p>I've noticed that for train_tp the species is present in the recording somewhere and or in multiple parts of the recording but not necessarily the time segment that is marked as having the species.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122862,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "12/22/2020 18:46:24",
      "content": "<p>Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122892,
          "author_name": "guglielmocamporese",
          "author_url": "",
          "post_date": "12/22/2020 19:07:22",
          "content": "<blockquote>\n  <p>Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings</p>\n</blockquote>\n<p>I'm using the log-mel-spectrogram (with log10) so I'm considering the dB scale, and I'm normalizing the samples globally, with respect to the entire training set.</p>\n<p>Best,</p>\n<p>Guglielmo</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1125800,
      "author_name": "hav4ik",
      "author_url": "",
      "post_date": "12/25/2020 04:47:54",
      "content": "<p>I've noticed some weird recordings as well (orange rectangle is TP, red rectangles is FP):</p>\n<hr>\n<p><strong>rec_id:</strong> 7448edc55<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2Ff4a31d9522a3ac0c7c59c8b4ebb5fa87%2Fmel_1.png?generation=1608871279214045&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><strong>rec_Id:</strong> de80fb815<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2F15c5682f0f703b2ef0093cd10189292b%2Fmel_2.png?generation=1608871601733051&amp;alt=media\" alt=\"\"></p>\n<p>I'll keep this comment updated with weird songs that I will find in the future.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1147922,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "01/10/2021 20:01:21",
          "content": "<p><a href=\"https://www.kaggle.com/hav4ik\" target=\"_blank\">@hav4ik</a> thanks for sharing, does all the vertical line in the first image you shared make sense to you? I listened to its recording and it almost sounds like some audio glitches</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1182155,
      "author_name": "karlyukang",
      "author_url": "",
      "post_date": "02/02/2021 10:33:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/guglielmocamporese\" target=\"_blank\">@guglielmocamporese</a>, could you please share the code that draws these great pictures?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1107878": "Hi everyone,\n\nI can share some patterns that I see when I visualize the bounding box [**t_min**, **t_max**, **f_min**, **f_max**] given by the competition. Here I reported the log-mel spectrum of some interesting samples (**f_min** and **f_max** are transformed from the frequency to the *mel* domain).\n\n| **Crops That Seem OK** |\n| --- |\n| **rec_id**: 0201197ec ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Ff9d7e937152ad7cce6025ae89f423c95%2FScreenshot%202020-12-10%20at%203.56.09%20AM.png?generation=1607569344764362&alt=media) **rec_id**:  009b760e6 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F4519ccf76b4348da311f426b715c501b%2FScreenshot%202020-12-10%20at%203.53.17%20AM.png?generation=1607569172265426&alt=media) |\n\n| **Crop Size Along Time Seems Too Small** |\n| --- |\n| **rec_id**: 0099c367b ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F9c648d8b99652b9c9b531133b17a2960%2FScreenshot%202020-12-10%20at%203.53.13%20AM.png?generation=1607569440726620&alt=media) **rec_id**: 01b41f92b ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F09952b81563f8b6775a504a3a4a871f6%2FScreenshot%202020-12-10%20at%203.53.43%20AM.png?generation=1607569507395742&alt=media)|\n\n| **Crop On Frequency Seems Shifted** |\n| --- |\n| **rec_id**: 0313e82cf ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F670d7651f7a51fc8190f5d34df97eed0%2FScreenshot%202020-12-10%20at%203.54.11%20AM.png?generation=1607569593086409&alt=media) **rec_id**: 03b96f209 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fb8df2426290b188216aca77633932b2e%2FScreenshot%202020-12-10%20at%203.54.16%20AM.png?generation=1607569640591224&alt=media)|\n\n| **Only Part Of The Pattern Seems Labeled** |\n| --- |\n| **rec_id**: 011f25080 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2F58888c52ae8d4a4283e7bccace2f3541%2FScreenshot%202020-12-10%20at%203.53.26%20AM.png?generation=1607569709906488&alt=media) **rec_id**: 0295e3234 ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F758632%2Fccac7cb9788d7c6f53dd5eb8559b4f57%2FScreenshot%202020-12-10%20at%203.53.56%20AM.png?generation=1607569766306295&alt=media) |\n\nHave you found any other pattern like these for the bboxes [**t_min**, **t_max**, **f_min**, **f_max**]?",
    "1108063": "guglielmocamporese that is why they named the training file as the name `train_tp.csv` rather than `train.csv` :D. We do not know if the unlabeled segment contains species or not.",
    "1108174": "Upvote if you stance out the recordingIds... 😁",
    "1108231": "done! :)",
    "1108328": "Done, too! :)\n\nMy first sound competition, so I'm trying to learn and with recordingId it makes my life easier in interpreting the data.",
    "1108742": "thanks for sharing. not sure if it is always correct though.... i saw there are some tp segments with very short duration like 0.3s, i wonder 1) how did a person label these with such accuracy/precision 2) the short duration shows up as a very strip in mel-spectrogram, how could the model deal with that? there are **plenty** place for the model to overfit/memorize on\n\nif you dont mind sharing, knowing some patterns of these boxes (t_min max, f_min max), how could it potentially help us?",
    "1109741": "guglielmocamporese as I generally understand the 'manual' id process the scientists that know the *call* make the *call* as to the given species, both as to species and length of representative vocalization. In the case of train.tp, this might be the most representative. \n\nExtending into train.fp might either be who gets confused with whom (as species) or which species are generally in the same habitats and overlap/step on each other's frequencies, or perhaps mimic. The data, as presented, doesn't readily lend itself to permutation analysis of birds of a habitat habitat together.  \n\nBut, one imagines that another expert has determined the segment is a false positive determined to be positive by what: Another algorithm, or by whom: Another expert. In either case, algo or expert, what will our algo make of this inter-rater disagreement, and do we even want to go there?  I don't know, but these are things I wonder about in the presence of warmed data.\n\n*mel* is a transform that is applied to the data in the pre-transformed rec_id: 0295e3234 for example. Is there a way to extract the post-transformed corners without reference to the pre-transformed, then diff them? And apply the diff to the label, perhaps? \n\nWell, sorry for a bunch of speculative statements \n\nChris",
    "1117950": "Yes. I've seen many such cases. Only one of the repeated patterns is labeled. Apparently, the localization information is not complete, and I think that's why time localization is not part of the test predictions. This situation will raise an issue when we cut the clips and creating train labels according to the localization information. Not sure how to improve this yet.",
    "1122761": "I've noticed that for train_tp the species is present in the recording somewhere and or in multiple parts of the recording but not necessarily the time segment that is marked as having the species.",
    "1122862": "Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings",
    "1122892": "> Yep - didn't take me long to find them on my first EDA. One other thing to account for is DB scale (if you're using a relative reference point). It depends on the loudest noise in the recording - so sometimes a call might be much darker on some recordings\n\nI'm using the log-mel-spectrogram (with log10) so I'm considering the dB scale, and I'm normalizing the samples globally, with respect to the entire training set.\n\nBest,\n\nGuglielmo",
    "1125800": "I've noticed some weird recordings as well (orange rectangle is TP, red rectangles is FP):\n\n----------------------------------------------------------------------\n\n**rec_id:** 7448edc55\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2Ff4a31d9522a3ac0c7c59c8b4ebb5fa87%2Fmel_1.png?generation=1608871279214045&alt=media)\n\n----------------------------------------------------------------------\n\n**rec_Id:** de80fb815\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1009388%2F15c5682f0f703b2ef0093cd10189292b%2Fmel_2.png?generation=1608871601733051&alt=media)\n\nI'll keep this comment updated with weird songs that I will find in the future.",
    "1147922": "hav4ik thanks for sharing, does all the vertical line in the first image you shared make sense to you? I listened to its recording and it almost sounds like some audio glitches",
    "1182155": "Hi @guglielmocamporese, could you please share the code that draws these great pictures?"
  },
  "source": "meta"
}