{
  "id": 197782,
  "title": "Welcome to the Rainforest Connection challenge!",
  "url": "/competitions/rfcx-species-audio-detection/discussion/197782",
  "author_name": "Zephyr Gold",
  "post_date": "2020-11-18T02:30:53.822000",
  "votes": 33,
  "comment_count": 33,
  "views": 0,
  "content": "<p>We are so excited to see everyone's interest in this challenge! It's exciting and inspiring to see how much people truly care about our work. So first off, thank you to all for participating! The results of this challenge will allow biologists and ecologists to find extremely rare endangered species in large volumes of audio recordings collected from the field. This is crucial to the work they are doing in the conservation space; allowing them to do their research faster through automation helps all of us support important conservation interventions out there in the world. Let us know if you have any questions, and thanks again from the Rainforest Connection team!</p>",
  "messages": [
    {
      "id": 1082554,
      "postDate": "2020-11-18T02:30:53.823Z",
      "content": "<p>We are so excited to see everyone's interest in this challenge! It's exciting and inspiring to see how much people truly care about our work. So first off, thank you to all for participating! The results of this challenge will allow biologists and ecologists to find extremely rare endangered species in large volumes of audio recordings collected from the field. This is crucial to the work they are doing in the conservation space; allowing them to do their research faster through automation helps all of us support important conservation interventions out there in the world. Let us know if you have any questions, and thanks again from the Rainforest Connection team!</p>",
      "rawMarkdown": "We are so excited to see everyone's interest in this challenge! It's exciting and inspiring to see how much people truly care about our work. So first off, thank you to all for participating! The results of this challenge will allow biologists and ecologists to find extremely rare endangered species in large volumes of audio recordings collected from the field. This is crucial to the work they are doing in the conservation space; allowing them to do their research faster through automation helps all of us support important conservation interventions out there in the world. Let us know if you have any questions, and thanks again from the Rainforest Connection team!",
      "votes": 32
    },
    {
      "id": 1101468,
      "postDate": "2020-12-03T23:36:33.887Z",
      "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a> <br>\nI have one question regrading to the annotation. Was the test dataset annotated in the same way as that of train dataset? In that case, the labels for test dataset can be noisy because train annotation turns out to be incomplete.</p>",
      "rawMarkdown": "@jacklebien @zephyrgold \nI have one question regrading to the annotation. Was the test dataset annotated in the same way as that of train dataset? In that case, the labels for test dataset can be noisy because train annotation turns out to be incomplete.",
      "votes": 10,
      "replies": [
        {
          "id": 1154437,
          "postDate": "2021-01-15T16:27:25.963Z",
          "content": "<p>No, further effort was put into the test set annotation to ensure complete labels.</p>",
          "rawMarkdown": "No, further effort was put into the test set annotation to ensure complete labels.",
          "votes": 5,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1185635,
      "postDate": "2021-02-04T08:47:43.623Z",
      "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> , <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <br>\nWould  you please clarify on this issue <br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215808\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215808</a></p>",
      "rawMarkdown": "@juliaelliott , @jacklebien @inversion \nWould  you please clarify on this issue \nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215808",
      "votes": 1
    },
    {
      "id": 1185383,
      "postDate": "2021-02-04T06:23:22.193Z",
      "content": "<p>Hi we are having trouble running on TPUs something seems to have broken this week. And prir running code has stopped doing so <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216408#1185368\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216408#1185368</a></p>",
      "rawMarkdown": "Hi we are having trouble running on TPUs something seems to have broken this week. And prir running code has stopped doing so https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216408#1185368",
      "votes": 1,
      "replies": [
        {
          "id": 1185917,
          "postDate": "2021-02-04T13:38:31.010Z",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> , <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a></p>",
          "rawMarkdown": "@juliaelliott , @jacklebien @inversion"
        }
      ]
    },
    {
      "id": 1182999,
      "postDate": "2021-02-02T17:14:33.960Z",
      "content": "<p>some creative commons license have Non-commercial right stated <a href=\"https://en.wikipedia.org/wiki/Creative_Commons_license\" target=\"_blank\">https://en.wikipedia.org/wiki/Creative_Commons_license</a> , and according to the link Non-commercial right means <code>Licensees may copy, distribute, display, and perform the work and make derivative works and remixes based on it only for non-commercial purposes.</code> . I wonder if any component of the solution that has the common creative common license that states it has this Non-commercial right, is permitted under the rules? <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> </p>",
      "rawMarkdown": "some creative commons license have Non-commercial right stated https://en.wikipedia.org/wiki/Creative_Commons_license , and according to the link Non-commercial right means `Licensees may copy, distribute, display, and perform the work and make derivative works and remixes based on it only for non-commercial purposes.` . I wonder if any component of the solution that has the common creative common license that states it has this Non-commercial right, is permitted under the rules? @juliaelliott @jacklebien ",
      "votes": 1,
      "replies": [
        {
          "id": 1185180,
          "postDate": "2021-02-04T02:47:35.643Z",
          "content": "<p><a href=\"https://www.kaggle.com/dicksonchin93\" target=\"_blank\">@dicksonchin93</a> I have confirmed with the host that they permit Creative Commons licensed datasets, including those specifying non-commercial use. As long as (if you are a winner) you license your solution with a non-exclusive license per the competition rules.</p>",
          "rawMarkdown": "@dicksonchin93 I have confirmed with the host that they permit Creative Commons licensed datasets, including those specifying non-commercial use. As long as (if you are a winner) you license your solution with a non-exclusive license per the competition rules.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1097414,
      "postDate": "2020-12-01T03:44:36.770Z",
      "content": "<p>I just have a question and doubt on the correct predictions. The question is, how many possible species there might contain in one single test clip? I mean, if there can be no no uppter limit(up to 24), then how would those ground truths be determined, even by experts in biologist? Say, if there is 8 ground truth, I will assume that this is an extremely hard process for the biologist to figure out the ground truth exactly as it is. So that is my question.</p>",
      "rawMarkdown": "I just have a question and doubt on the correct predictions. The question is, how many possible species there might contain in one single test clip? I mean, if there can be no no uppter limit(up to 24), then how would those ground truths be determined, even by experts in biologist? Say, if there is 8 ground truth, I will assume that this is an extremely hard process for the biologist to figure out the ground truth exactly as it is. So that is my question.",
      "votes": 1,
      "replies": [
        {
          "id": 1106191,
          "postDate": "2020-12-08T15:25:20.503Z",
          "content": "<p>It is possible for there to be up to 24, I don't think I should disclose anything else. We have done our best using several methods of review to ensure accuracy of the test labels.</p>",
          "rawMarkdown": "It is possible for there to be up to 24, I don't think I should disclose anything else. We have done our best using several methods of review to ensure accuracy of the test labels.",
          "isDeleted": true
        },
        {
          "id": 1141034,
          "postDate": "2021-01-06T12:52:24.513Z",
          "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> sorry if thuis was asked and answered already, I am just entering.</p>\n<blockquote>\n  <p>Not all the target sounds present in the training data are annotated in train_tp.csv</p>\n</blockquote>\n<p>I sit the same in test or is test fully labelled?</p>",
          "rawMarkdown": "@jacklebien sorry if thuis was asked and answered already, I am just entering.\n\n> Not all the target sounds present in the training data are annotated in train_tp.csv\n\nI sit the same in test or is test fully labelled?"
        },
        {
          "id": 1141047,
          "postDate": "2021-01-06T12:57:35.653Z",
          "content": "<p>Nevermind, you responded in the discussion of TP vs FP topic.  Interesting semi supervised problem!</p>",
          "rawMarkdown": "Nevermind, you responded in the discussion of TP vs FP topic.  Interesting semi supervised problem!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1177181,
      "postDate": "2021-01-30T05:52:24.573Z",
      "content": "<p><a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a> </p>\n<p>Can you comment if public testset's label distribution is similar to that of private set ? </p>",
      "rawMarkdown": "@zephyrgold \n\nCan you comment if public testset's label distribution is similar to that of private set ? "
    },
    {
      "id": 1154321,
      "postDate": "2021-01-15T15:15:12.760Z",
      "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a> <br>\nJoined to the competition new.<br>\nI have a concrete question.<br>\nI was checking audio files based on train_tp.csv information.<br>\nspecies_id=6, checked two files on the train_tp: eb0cd97da.flac, that has species 6 in time: approx. 22.1 -24.1 sec.(train_tp.idx at 1135)<br>\nsame species (6),  has its train_tp line for e2ce1658e.flac and at time: 34.7-36.8 sec…(train_tp.idx at 1092)<br>\n<strong>NOW BOTH SONGTYPES ARE 1, but when you listen to those sections, although they might be from the same bird, the sounds (songs) are totally different,</strong> in a level that one does not have to be a Subject Matter Expert.<br>\n e2ce1658e bird is a calm one in the background, whereas eb0cd97da is an angry dominant bird…<br>\n<strong>Question: should't at least the songtypes be labelled different???\nWhat is going on here???</strong></p>",
      "rawMarkdown": "@jacklebien @zephyrgold \nJoined to the competition new.\nI have a concrete question.\nI was checking audio files based on train_tp.csv information.\nspecies_id=6, checked two files on the train_tp: eb0cd97da.flac, that has species 6 in time: approx. 22.1 -24.1 sec.(train_tp.idx at 1135)\nsame species (6),  has its train_tp line for e2ce1658e.flac and at time: 34.7-36.8 sec...(train_tp.idx at 1092)\n**NOW BOTH SONGTYPES ARE 1, but when you listen to those sections, although they might be from the same bird, the sounds (songs) are totally different,** in a level that one does not have to be a Subject Matter Expert.\n e2ce1658e bird is a calm one in the background, whereas eb0cd97da is an angry dominant bird...\n**Question: should't at least the songtypes be labelled different???\nWhat is going on here???**",
      "replies": [
        {
          "id": 1154447,
          "postDate": "2021-01-15T16:33:07.133Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/abdulkadirguner\" target=\"_blank\">@abdulkadirguner</a>, I asked a biologist to review these two cases you mentioned, and he confirmed it is the same species. It is sometimes questionable when to assign new call types, but in his opinion, these calls consist of similar notes, but different pacing and repetition. Some species like this one have greater variation in their calls.</p>",
          "rawMarkdown": "Hello @abdulkadirguner, I asked a biologist to review these two cases you mentioned, and he confirmed it is the same species. It is sometimes questionable when to assign new call types, but in his opinion, these calls consist of similar notes, but different pacing and repetition. Some species like this one have greater variation in their calls.",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 1154490,
          "postDate": "2021-01-15T16:53:59.663Z",
          "content": "<p>Thanks for the quick turn around. </p>",
          "rawMarkdown": "Thanks for the quick turn around. "
        }
      ]
    },
    {
      "id": 1144690,
      "postDate": "2021-01-08T16:00:25.457Z",
      "content": "<p>I am having a problem with getting the sound data (.flac file) into usable data.  I have inputted the .flac file in with python soundfile and have used python numpy.fft.fft for the Fourier Transform.  The problem I am having is it take a large chunk of data to get the frequency where I want it and I want to know the frequency response for a smaller part of the data.  Anyone please, help me with this part of the completion, because it is not part of the AI.</p>",
      "rawMarkdown": "I am having a problem with getting the sound data (.flac file) into usable data.  I have inputted the .flac file in with python soundfile and have used python numpy.fft.fft for the Fourier Transform.  The problem I am having is it take a large chunk of data to get the frequency where I want it and I want to know the frequency response for a smaller part of the data.  Anyone please, help me with this part of the completion, because it is not part of the AI.",
      "replies": [
        {
          "id": 1170353,
          "postDate": "2021-01-26T06:58:30.423Z",
          "content": "<p><a href=\"https://www.kaggle.com/edwardgfleming\" target=\"_blank\">@edwardgfleming</a> You could use tf or librosa fft. My kernels cover examples for both. <br>\n<a href=\"https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering\" target=\"_blank\">https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering</a><br>\nLet me know if I can help further </p>",
          "rawMarkdown": "@edwardgfleming You could use tf or librosa fft. My kernels cover examples for both. \nhttps://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering\nLet me know if I can help further "
        }
      ]
    },
    {
      "id": 1106137,
      "postDate": "2020-12-08T15:06:50.800Z",
      "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a>  may I get your reply? thanks. it looks like it is 7 days ago…</p>",
      "rawMarkdown": "@jacklebien  may I get your reply? thanks. it looks like it is 7 days ago..."
    },
    {
      "id": 1085745,
      "postDate": "2020-11-21T07:02:18.300Z",
      "content": "<p>I am new/novice to kaggle. I have one query. kaggle kernal is more than enough for reading these huge audio files or is it recommended to use external GPU from google collab or so? Please help me with this. I am checking if there is an option for a kernal but due to my lack of knowhow, i am not able to find one. </p>",
      "rawMarkdown": "I am new/novice to kaggle. I have one query. kaggle kernal is more than enough for reading these huge audio files or is it recommended to use external GPU from google collab or so? Please help me with this. I am checking if there is an option for a kernal but due to my lack of knowhow, i am not able to find one. ",
      "replies": [
        {
          "id": 1136624,
          "postDate": "2021-01-03T09:19:26.747Z",
          "content": "<p>You can read given audio files via kaggle kernel itself</p>",
          "rawMarkdown": "You can read given audio files via kaggle kernel itself"
        },
        {
          "id": 1170349,
          "postDate": "2021-01-26T06:57:31.050Z",
          "content": "<p>Pl see if this kernel helps:<br>\n<a href=\"https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering\" target=\"_blank\">https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering</a></p>",
          "rawMarkdown": "Pl see if this kernel helps:\nhttps://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering"
        }
      ]
    },
    {
      "id": 1083350,
      "postDate": "2020-11-18T22:12:46.057Z",
      "content": "<p>QQ, is it possible for there to be more than true positive species in an audio sample? </p>",
      "rawMarkdown": "QQ, is it possible for there to be more than true positive species in an audio sample? ",
      "replies": [
        {
          "id": 1083507,
          "postDate": "2020-11-19T04:04:35.797Z",
          "content": "<p>Yes there can be other species aside from the labeled species calling simultaneously. Though when also considering frequency bounds <code>f_min</code> and <code>f_max</code>, it is less likely that the labeled signal overlaps with a different species' call.</p>",
          "rawMarkdown": "Yes there can be other species aside from the labeled species calling simultaneously. Though when also considering frequency bounds `f_min` and `f_max`, it is less likely that the labeled signal overlaps with a different species' call.",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 1083515,
          "postDate": "2020-11-19T04:18:23.170Z",
          "content": "<p>Okay great, thanks! I assume that for each sample there won't be more than one true positive, is that correct? Ran a sanity check but just wanted to double check 😃</p>",
          "rawMarkdown": "Okay great, thanks! I assume that for each sample there won't be more than one true positive, is that correct? Ran a sanity check but just wanted to double check 😃"
        },
        {
          "id": 1083871,
          "postDate": "2020-11-19T13:31:29.460Z",
          "content": "<p>For a single row of <code>train_tp.csv</code>, it is unlikely that another signal (perhaps another target species) overlaps in time <em>and</em> frequency with the labeled signal, but it is possible.</p>",
          "rawMarkdown": "For a single row of `train_tp.csv`, it is unlikely that another signal (perhaps another target species) overlaps in time *and* frequency with the labeled signal, but it is possible.",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1082862,
      "postDate": "2020-11-18T10:07:09.737Z",
      "content": "<p>I am new to kaggle <br>\nI want to clear my doubt<br>\n1- traintp and trainfp is what <br>\n2- total rows in train combine fp and tp is about 9000 and I have only 4247 audio file in train folder and sme problem with test folder too. </p>\n<p>please resolve my these doubts<br>\nthanks</p>",
      "rawMarkdown": "I am new to kaggle \nI want to clear my doubt\n1- traintp and trainfp is what \n2- total rows in train combine fp and tp is about 9000 and I have only 4247 audio file in train folder and sme problem with test folder too. \n\nplease resolve my these doubts\nthanks",
      "replies": [
        {
          "id": 1082975,
          "postDate": "2020-11-18T13:05:41.757Z",
          "content": "<p>Hello, I have posted a new topic \"train_tp vs. train_fp\" that will hopefully help you out. About the file counts, each row in the CSV files corresponds to a single training sample (1 animal call), and there can be multiple calls per audio file.</p>",
          "rawMarkdown": "Hello, I have posted a new topic \"train_tp vs. train_fp\" that will hopefully help you out. About the file counts, each row in the CSV files corresponds to a single training sample (1 animal call), and there can be multiple calls per audio file.",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1082847,
      "postDate": "2020-11-18T09:46:09.430Z",
      "content": "<p>Hiiii, I want to know how to start the process..</p>",
      "rawMarkdown": "Hiiii, I want to know how to start the process..",
      "replies": [
        {
          "id": 1084106,
          "postDate": "2020-11-19T18:25:45.997Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1209673,
      "postDate": "2021-02-19T02:18:53.663Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1093677,
      "postDate": "2020-11-27T23:05:10.987Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a></p>\n<p>Could you give us some more details on the annotations process to obtain <strong>train_tp.csv</strong>.<br>\nIn particular:</p>\n<ul>\n<li>Were TP annotations performed by human experts, semi-automated, or fully automated ?</li>\n<li>How complete are the annotations; do you think they capture a high proportion of the targeted sounds ?</li>\n</ul>",
      "rawMarkdown": "Hi @jacklebien @zephyrgold\n\nCould you give us some more details on the annotations process to obtain **train_tp.csv**.\nIn particular:\n- Were TP annotations performed by human experts, semi-automated, or fully automated ?\n- How complete are the annotations; do you think they capture a high proportion of the targeted sounds ?",
      "votes": 6,
      "isDeleted": true,
      "replies": [
        {
          "id": 1106124,
          "postDate": "2020-12-08T14:55:50.280Z",
          "content": "<p>Hello, please see the discussion \"train_tp vs train_fp\" for more details about the data collection. Not all the target sounds present in the training data are annotated in train_tp.csv. The portion annotated will vary for the species based on how rarely they call.</p>",
          "rawMarkdown": "Hello, please see the discussion \"train_tp vs train_fp\" for more details about the data collection. Not all the target sounds present in the training data are annotated in train_tp.csv. The portion annotated will vary for the species based on how rarely they call.",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 1106444,
          "postDate": "2020-12-08T21:00:22.720Z",
          "content": "<p>Thanks Jack, now it is clear. <br>\nSo the challenge is to train high performing models under incomplete labels.<br>\nThis is weak supervision which is a relevant use-case because labeling large audio collections fully manual is often unfeasible. </p>",
          "rawMarkdown": "Thanks Jack, now it is clear. \nSo the challenge is to train high performing models under incomplete labels.\nThis is weak supervision which is a relevant use-case because labeling large audio collections fully manual is often unfeasible. ",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1101468,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2020-12-03T23:36:33.887000",
      "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a> <br>\nI have one question regrading to the annotation. Was the test dataset annotated in the same way as that of train dataset? In that case, the labels for test dataset can be noisy because train annotation turns out to be incomplete.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 1154437,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-15T16:27:25.963000",
          "content": "<p>No, further effort was put into the test set annotation to ensure complete labels.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1185635,
      "author_name": "RAHUL SINGH INDA",
      "author_url": "",
      "post_date": "2021-02-04T08:47:43.623000",
      "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> , <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <br>\nWould  you please clarify on this issue <br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215808\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215808</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1185383,
      "author_name": "Felipe Bivort Haiek",
      "author_url": "",
      "post_date": "2021-02-04T06:23:22.193000",
      "content": "<p>Hi we are having trouble running on TPUs something seems to have broken this week. And prir running code has stopped doing so <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216408#1185368\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216408#1185368</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1185917,
          "author_name": "Felipe Bivort Haiek",
          "author_url": "",
          "post_date": "2021-02-04T13:38:31.010000",
          "content": "<p><a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> , <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1182999,
      "author_name": "Ee Kin Chin",
      "author_url": "",
      "post_date": "2021-02-02T17:14:33.960000",
      "content": "<p>some creative commons license have Non-commercial right stated <a href=\"https://en.wikipedia.org/wiki/Creative_Commons_license\" target=\"_blank\">https://en.wikipedia.org/wiki/Creative_Commons_license</a> , and according to the link Non-commercial right means <code>Licensees may copy, distribute, display, and perform the work and make derivative works and remixes based on it only for non-commercial purposes.</code> . I wonder if any component of the solution that has the common creative common license that states it has this Non-commercial right, is permitted under the rules? <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1185180,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2021-02-04T02:47:35.643000",
          "content": "<p><a href=\"https://www.kaggle.com/dicksonchin93\" target=\"_blank\">@dicksonchin93</a> I have confirmed with the host that they permit Creative Commons licensed datasets, including those specifying non-commercial use. As long as (if you are a winner) you license your solution with a non-exclusive license per the competition rules.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1097414,
      "author_name": "Wubin Bai",
      "author_url": "",
      "post_date": "2020-12-01T03:44:36.770000",
      "content": "<p>I just have a question and doubt on the correct predictions. The question is, how many possible species there might contain in one single test clip? I mean, if there can be no no uppter limit(up to 24), then how would those ground truths be determined, even by experts in biologist? Say, if there is 8 ground truth, I will assume that this is an extremely hard process for the biologist to figure out the ground truth exactly as it is. So that is my question.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1106191,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-08T15:25:20.503000",
          "content": "<p>It is possible for there to be up to 24, I don't think I should disclose anything else. We have done our best using several methods of review to ensure accuracy of the test labels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141034,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-01-06T12:52:24.513000",
          "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> sorry if thuis was asked and answered already, I am just entering.</p>\n<blockquote>\n  <p>Not all the target sounds present in the training data are annotated in train_tp.csv</p>\n</blockquote>\n<p>I sit the same in test or is test fully labelled?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1141047,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-01-06T12:57:35.653000",
          "content": "<p>Nevermind, you responded in the discussion of TP vs FP topic.  Interesting semi supervised problem!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1177181,
      "author_name": "NakedKoala",
      "author_url": "",
      "post_date": "2021-01-30T05:52:24.573000",
      "content": "<p><a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a> </p>\n<p>Can you comment if public testset's label distribution is similar to that of private set ? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1154321,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2021-01-15T15:15:12.760000",
      "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a> <br>\nJoined to the competition new.<br>\nI have a concrete question.<br>\nI was checking audio files based on train_tp.csv information.<br>\nspecies_id=6, checked two files on the train_tp: eb0cd97da.flac, that has species 6 in time: approx. 22.1 -24.1 sec.(train_tp.idx at 1135)<br>\nsame species (6),  has its train_tp line for e2ce1658e.flac and at time: 34.7-36.8 sec…(train_tp.idx at 1092)<br>\n<strong>NOW BOTH SONGTYPES ARE 1, but when you listen to those sections, although they might be from the same bird, the sounds (songs) are totally different,</strong> in a level that one does not have to be a Subject Matter Expert.<br>\n e2ce1658e bird is a calm one in the background, whereas eb0cd97da is an angry dominant bird…<br>\n<strong>Question: should't at least the songtypes be labelled different???\nWhat is going on here???</strong></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1154447,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-15T16:33:07.133000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/abdulkadirguner\" target=\"_blank\">@abdulkadirguner</a>, I asked a biologist to review these two cases you mentioned, and he confirmed it is the same species. It is sometimes questionable when to assign new call types, but in his opinion, these calls consist of similar notes, but different pacing and repetition. Some species like this one have greater variation in their calls.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1154490,
          "author_name": "GUNER",
          "author_url": "",
          "post_date": "2021-01-15T16:53:59.663000",
          "content": "<p>Thanks for the quick turn around. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1144690,
      "author_name": "EDWARD G. FLEMING",
      "author_url": "",
      "post_date": "2021-01-08T16:00:25.457000",
      "content": "<p>I am having a problem with getting the sound data (.flac file) into usable data.  I have inputted the .flac file in with python soundfile and have used python numpy.fft.fft for the Fourier Transform.  The problem I am having is it take a large chunk of data to get the frequency where I want it and I want to know the frequency response for a smaller part of the data.  Anyone please, help me with this part of the completion, because it is not part of the AI.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1170353,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-26T06:58:30.423000",
          "content": "<p><a href=\"https://www.kaggle.com/edwardgfleming\" target=\"_blank\">@edwardgfleming</a> You could use tf or librosa fft. My kernels cover examples for both. <br>\n<a href=\"https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering\" target=\"_blank\">https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering</a><br>\nLet me know if I can help further </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1106137,
      "author_name": "Wubin Bai",
      "author_url": "",
      "post_date": "2020-12-08T15:06:50.800000",
      "content": "<p><a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a>  may I get your reply? thanks. it looks like it is 7 days ago…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1085745,
      "author_name": "ChangMaulee",
      "author_url": "",
      "post_date": "2020-11-21T07:02:18.300000",
      "content": "<p>I am new/novice to kaggle. I have one query. kaggle kernal is more than enough for reading these huge audio files or is it recommended to use external GPU from google collab or so? Please help me with this. I am checking if there is an option for a kernal but due to my lack of knowhow, i am not able to find one. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1136624,
          "author_name": "Kushal Motwani",
          "author_url": "",
          "post_date": "2021-01-03T09:19:26.747000",
          "content": "<p>You can read given audio files via kaggle kernel itself</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1170349,
          "author_name": "Allohvk",
          "author_url": "",
          "post_date": "2021-01-26T06:57:31.050000",
          "content": "<p>Pl see if this kernel helps:<br>\n<a href=\"https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering\" target=\"_blank\">https://www.kaggle.com/allohvk/the-nuts-and-bolts-of-sound-feature-engineering</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1083350,
      "author_name": "Zach Nussbaum",
      "author_url": "",
      "post_date": "2020-11-18T22:12:46.057000",
      "content": "<p>QQ, is it possible for there to be more than true positive species in an audio sample? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1083507,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-19T04:04:35.797000",
          "content": "<p>Yes there can be other species aside from the labeled species calling simultaneously. Though when also considering frequency bounds <code>f_min</code> and <code>f_max</code>, it is less likely that the labeled signal overlaps with a different species' call.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1083515,
          "author_name": "Zach Nussbaum",
          "author_url": "",
          "post_date": "2020-11-19T04:18:23.170000",
          "content": "<p>Okay great, thanks! I assume that for each sample there won't be more than one true positive, is that correct? Ran a sanity check but just wanted to double check 😃</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1083871,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-19T13:31:29.460000",
          "content": "<p>For a single row of <code>train_tp.csv</code>, it is unlikely that another signal (perhaps another target species) overlaps in time <em>and</em> frequency with the labeled signal, but it is possible.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1082862,
      "author_name": "Shivam Khandewal",
      "author_url": "",
      "post_date": "2020-11-18T10:07:09.737000",
      "content": "<p>I am new to kaggle <br>\nI want to clear my doubt<br>\n1- traintp and trainfp is what <br>\n2- total rows in train combine fp and tp is about 9000 and I have only 4247 audio file in train folder and sme problem with test folder too. </p>\n<p>please resolve my these doubts<br>\nthanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1082975,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-18T13:05:41.757000",
          "content": "<p>Hello, I have posted a new topic \"train_tp vs. train_fp\" that will hopefully help you out. About the file counts, each row in the CSV files corresponds to a single training sample (1 animal call), and there can be multiple calls per audio file.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1082847,
      "author_name": "AIMEN NAAZ SHAHEEN",
      "author_url": "",
      "post_date": "2020-11-18T09:46:09.430000",
      "content": "<p>Hiiii, I want to know how to start the process..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1084106,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-11-19T18:25:45.997000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1209673,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-19T02:18:53.663000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1093677,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-27T23:05:10.987000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> <a href=\"https://www.kaggle.com/zephyrgold\" target=\"_blank\">@zephyrgold</a></p>\n<p>Could you give us some more details on the annotations process to obtain <strong>train_tp.csv</strong>.<br>\nIn particular:</p>\n<ul>\n<li>Were TP annotations performed by human experts, semi-automated, or fully automated ?</li>\n<li>How complete are the annotations; do you think they capture a high proportion of the targeted sounds ?</li>\n</ul>",
      "votes": 6,
      "replies": [
        {
          "id": 1106124,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-08T14:55:50.280000",
          "content": "<p>Hello, please see the discussion \"train_tp vs train_fp\" for more details about the data collection. Not all the target sounds present in the training data are annotated in train_tp.csv. The portion annotated will vary for the species based on how rarely they call.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1106444,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-08T21:00:22.720000",
          "content": "<p>Thanks Jack, now it is clear. <br>\nSo the challenge is to train high performing models under incomplete labels.<br>\nThis is weak supervision which is a relevant use-case because labeling large audio collections fully manual is often unfeasible. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1082554": "We are so excited to see everyone's interest in this challenge! It's exciting and inspiring to see how much people truly care about our work. So first off, thank you to all for participating! The results of this challenge will allow biologists and ecologists to find extremely rare endangered species in large volumes of audio recordings collected from the field. This is crucial to the work they are doing in the conservation space; allowing them to do their research faster through automation helps all of us support important conservation interventions out there in the world. Let us know if you have any questions, and thanks again from the Rainforest Connection team!",
    "1101468": "@jacklebien @zephyrgold \nI have one question regrading to the annotation. Was the test dataset annotated in the same way as that of train dataset? In that case, the labels for test dataset can be noisy because train annotation turns out to be incomplete.",
    "1185635": "@juliaelliott , @jacklebien @inversion \nWould  you please clarify on this issue \nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215808",
    "1185383": "Hi we are having trouble running on TPUs something seems to have broken this week. And prir running code has stopped doing so https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/216408#1185368",
    "1182999": "some creative commons license have Non-commercial right stated https://en.wikipedia.org/wiki/Creative_Commons_license , and according to the link Non-commercial right means `Licensees may copy, distribute, display, and perform the work and make derivative works and remixes based on it only for non-commercial purposes.` . I wonder if any component of the solution that has the common creative common license that states it has this Non-commercial right, is permitted under the rules? @juliaelliott @jacklebien ",
    "1097414": "I just have a question and doubt on the correct predictions. The question is, how many possible species there might contain in one single test clip? I mean, if there can be no no uppter limit(up to 24), then how would those ground truths be determined, even by experts in biologist? Say, if there is 8 ground truth, I will assume that this is an extremely hard process for the biologist to figure out the ground truth exactly as it is. So that is my question.",
    "1177181": "@zephyrgold \n\nCan you comment if public testset's label distribution is similar to that of private set ? ",
    "1154321": "@jacklebien @zephyrgold \nJoined to the competition new.\nI have a concrete question.\nI was checking audio files based on train_tp.csv information.\nspecies_id=6, checked two files on the train_tp: eb0cd97da.flac, that has species 6 in time: approx. 22.1 -24.1 sec.(train_tp.idx at 1135)\nsame species (6),  has its train_tp line for e2ce1658e.flac and at time: 34.7-36.8 sec...(train_tp.idx at 1092)\n**NOW BOTH SONGTYPES ARE 1, but when you listen to those sections, although they might be from the same bird, the sounds (songs) are totally different,** in a level that one does not have to be a Subject Matter Expert.\n e2ce1658e bird is a calm one in the background, whereas eb0cd97da is an angry dominant bird...\n**Question: should't at least the songtypes be labelled different???\nWhat is going on here???**",
    "1144690": "I am having a problem with getting the sound data (.flac file) into usable data.  I have inputted the .flac file in with python soundfile and have used python numpy.fft.fft for the Fourier Transform.  The problem I am having is it take a large chunk of data to get the frequency where I want it and I want to know the frequency response for a smaller part of the data.  Anyone please, help me with this part of the completion, because it is not part of the AI.",
    "1106137": "@jacklebien  may I get your reply? thanks. it looks like it is 7 days ago...",
    "1085745": "I am new/novice to kaggle. I have one query. kaggle kernal is more than enough for reading these huge audio files or is it recommended to use external GPU from google collab or so? Please help me with this. I am checking if there is an option for a kernal but due to my lack of knowhow, i am not able to find one. ",
    "1083350": "QQ, is it possible for there to be more than true positive species in an audio sample? ",
    "1082862": "I am new to kaggle \nI want to clear my doubt\n1- traintp and trainfp is what \n2- total rows in train combine fp and tp is about 9000 and I have only 4247 audio file in train folder and sme problem with test folder too. \n\nplease resolve my these doubts\nthanks",
    "1082847": "Hiiii, I want to know how to start the process..",
    "1209673": "",
    "1093677": "Hi @jacklebien @zephyrgold\n\nCould you give us some more details on the annotations process to obtain **train_tp.csv**.\nIn particular:\n- Were TP annotations performed by human experts, semi-automated, or fully automated ?\n- How complete are the annotations; do you think they capture a high proportion of the targeted sounds ?"
  }
}