{
  "id": 206980,
  "title": "Is this what is going on with TP, FP and the test set?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/206980",
  "author_name": "",
  "post_date": "2020-12-27T14:08:03.652380700Z",
  "votes": 24,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Here's my best attempt of summarizing what I've understood (from various discussions such as <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206012\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866\" target=\"_blank\">here</a>), about the true positive, false positives and the test set, first as a diagram and then as a bulleted list.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F649192e00d29f7ee8aa1abb432e90095%2FScreenshot%20from%202020-12-27%2014-28-08.png?generation=1609075778207654&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Training data was run through a computer filter that is just a <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866\" target=\"_blank\">\"rudimentary detection algorithm\"</a>, so I guess we don't quite know how much we'd trust it (i.e. esp. things it does not flag might very well still be positives)</li>\n<li>If the filter did not highlight a clip, we do not know what is going on in it (although we might guess that the filter is more likely to be right than wrong)</li>\n<li>If the filter did not highlight a clip, it underwent human review for the one single species that the algorithm \"detected\". If this is truly that species, we get the record in true positives, otherwise in false positives (but the human reviewer would not flag other species - I'm sort of ignoring the possibility that the humans get it wrong).</li>\n<li>So records in true positives are real \"1\"s, those in false positives are \"0\"s (but ones that are somehow close to \"1\"s as far as a \"rudimentary detection algorithm\" is concerned), but if a clip does not have a record in true positives or false positives, we just don't know for sure what is going on (if the rudimentary detection algorithm is great, then they might really mostly be \"0\"s)?</li>\n<li>In contrast, the test set clips were completely labelled by humans.</li>\n</ul>\n<p>Did I get that right? The biggest question mark I really had was around whether no record in true positives and false negatives for a species in a clip really just means \"not flagged automatically\" = we don't know for sure (but not so probable depending on how good the algorithm used was)?</p>\n<p>Does that suggest that we should use a loss function that \"insists\" on the true positives and false negatives, but is a bit more \"lenient\" about the clips we are not sure about (but might guess are more likely to be negatives than positives)?</p>\n<p>One really big challenge is then to have audio data that corresponds to things that the correlational \"rudimentary detection algorithm\" would not flag, but which is definitely negatives. Is that right?</p>",
  "messages": [
    {
      "id": "1128518",
      "postDate": "12/27/2020 14:08:03",
      "content": "<p>Here's my best attempt of summarizing what I've understood (from various discussions such as <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206012\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866\" target=\"_blank\">here</a>), about the true positive, false positives and the test set, first as a diagram and then as a bulleted list.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F649192e00d29f7ee8aa1abb432e90095%2FScreenshot%20from%202020-12-27%2014-28-08.png?generation=1609075778207654&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Training data was run through a computer filter that is just a <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866\" target=\"_blank\">\"rudimentary detection algorithm\"</a>, so I guess we don't quite know how much we'd trust it (i.e. esp. things it does not flag might very well still be positives)</li>\n<li>If the filter did not highlight a clip, we do not know what is going on in it (although we might guess that the filter is more likely to be right than wrong)</li>\n<li>If the filter did not highlight a clip, it underwent human review for the one single species that the algorithm \"detected\". If this is truly that species, we get the record in true positives, otherwise in false positives (but the human reviewer would not flag other species - I'm sort of ignoring the possibility that the humans get it wrong).</li>\n<li>So records in true positives are real \"1\"s, those in false positives are \"0\"s (but ones that are somehow close to \"1\"s as far as a \"rudimentary detection algorithm\" is concerned), but if a clip does not have a record in true positives or false positives, we just don't know for sure what is going on (if the rudimentary detection algorithm is great, then they might really mostly be \"0\"s)?</li>\n<li>In contrast, the test set clips were completely labelled by humans.</li>\n</ul>\n<p>Did I get that right? The biggest question mark I really had was around whether no record in true positives and false negatives for a species in a clip really just means \"not flagged automatically\" = we don't know for sure (but not so probable depending on how good the algorithm used was)?</p>\n<p>Does that suggest that we should use a loss function that \"insists\" on the true positives and false negatives, but is a bit more \"lenient\" about the clips we are not sure about (but might guess are more likely to be negatives than positives)?</p>\n<p>One really big challenge is then to have audio data that corresponds to things that the correlational \"rudimentary detection algorithm\" would not flag, but which is definitely negatives. Is that right?</p>",
      "rawMarkdown": "Here's my best attempt of summarizing what I've understood (from various discussions such as [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782), [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206012) and [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866)), about the true positive, false positives and the test set, first as a diagram and then as a bulleted list.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F649192e00d29f7ee8aa1abb432e90095%2FScreenshot%20from%202020-12-27%2014-28-08.png?generation=1609075778207654&alt=media)\n* Training data was run through a computer filter that is just a [\"rudimentary detection algorithm\"](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866), so I guess we don't quite know how much we'd trust it (i.e. esp. things it does not flag might very well still be positives)\n* If the filter did not highlight a clip, we do not know what is going on in it (although we might guess that the filter is more likely to be right than wrong)\n* If the filter did not highlight a clip, it underwent human review for the one single species that the algorithm \"detected\". If this is truly that species, we get the record in true positives, otherwise in false positives (but the human reviewer would not flag other species - I'm sort of ignoring the possibility that the humans get it wrong).\n* So records in true positives are real \"1\"s, those in false positives are \"0\"s (but ones that are somehow close to \"1\"s as far as a \"rudimentary detection algorithm\" is concerned), but if a clip does not have a record in true positives or false positives, we just don't know for sure what is going on (if the rudimentary detection algorithm is great, then they might really mostly be \"0\"s)?\n* In contrast, the test set clips were completely labelled by humans.\n\nDid I get that right? The biggest question mark I really had was around whether no record in true positives and false negatives for a species in a clip really just means \"not flagged automatically\" = we don't know for sure (but not so probable depending on how good the algorithm used was)?\n\nDoes that suggest that we should use a loss function that \"insists\" on the true positives and false negatives, but is a bit more \"lenient\" about the clips we are not sure about (but might guess are more likely to be negatives than positives)?\n\nOne really big challenge is then to have audio data that corresponds to things that the correlational \"rudimentary detection algorithm\" would not flag, but which is definitely negatives. Is that right?",
      "votes": null
    },
    {
      "id": "1128649",
      "postDate": "12/27/2020 16:23:10",
      "content": "<p>great work<br>\n👍</p>",
      "rawMarkdown": "great work\n👍",
      "votes": null
    },
    {
      "id": "1131093",
      "postDate": "12/29/2020 14:40:37",
      "content": "<p>Thanks for this summary. I noticed you upvoted my question to the organisers <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1123751\" target=\"_blank\">here</a>. There I assumed that the test data was labelled in the same way. So my question to you is: which link says that the test data is 100% human labelled?</p>\n<p>If the test data is indeed human labelled that I do agree with your idea of context sensitive loss. Nice idea.</p>",
      "rawMarkdown": "Thanks for this summary. I noticed you upvoted my question to the organisers [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1123751). There I assumed that the test data was labelled in the same way. So my question to you is: which link says that the test data is 100% human labelled?\n\nIf the test data is indeed human labelled that I do agree with your idea of context sensitive loss. Nice idea.",
      "votes": null
    },
    {
      "id": "1131152",
      "postDate": "12/29/2020 15:10:45",
      "content": "<p>I interpreted <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1109273\" target=\"_blank\">this response</a> in that way, where <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> said:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>, no, unlike the training data, we tried to ensure presence/absence of any target species is labeled for every test file.</p>\n</blockquote>\n<p>This seems to suggest a different process was followed and at a minimum this would mean that any potential hit for a target species would have been given to a human (while it was sort of implied to me that not even all flagged suspected cases were even given for human review - I was not 100% sure on this). I guess it does not necessarily mean that every audio segment that was not even flagged for attention was given to a humans and it does not necessarily mean that a human review was done for species other than the flagged target species.</p>\n<p>I posted the topic, because I am not 100% sure on several aspects (incl. what you asked about), and obviously it makes a huge difference for how one should approach the problem.</p>\n<p>So, perhaps this should look more like this (with a question mark also on whether the test data was really all human reviewed, since that was not directly stated):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F50c4f648b64017056dd8fe7583ae8595%2FScreenshot%20from%202020-12-29%2016-16-26.png?generation=1609255058462146&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I interpreted [this response](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1109273) in that way, where @jacklebien said:\n> @selimsef, no, unlike the training data, we tried to ensure presence/absence of any target species is labeled for every test file.\n\nThis seems to suggest a different process was followed and at a minimum this would mean that any potential hit for a target species would have been given to a human (while it was sort of implied to me that not even all flagged suspected cases were even given for human review - I was not 100% sure on this). I guess it does not necessarily mean that every audio segment that was not even flagged for attention was given to a humans and it does not necessarily mean that a human review was done for species other than the flagged target species.\n\nI posted the topic, because I am not 100% sure on several aspects (incl. what you asked about), and obviously it makes a huge difference for how one should approach the problem.\n\nSo, perhaps this should look more like this (with a question mark also on whether the test data was really all human reviewed, since that was not directly stated):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F50c4f648b64017056dd8fe7583ae8595%2FScreenshot%20from%202020-12-29%2016-16-26.png?generation=1609255058462146&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1128649,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "12/27/2020 16:23:10",
      "content": "<p>great work<br>\n👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1131093,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "12/29/2020 14:40:37",
      "content": "<p>Thanks for this summary. I noticed you upvoted my question to the organisers <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1123751\" target=\"_blank\">here</a>. There I assumed that the test data was labelled in the same way. So my question to you is: which link says that the test data is 100% human labelled?</p>\n<p>If the test data is indeed human labelled that I do agree with your idea of context sensitive loss. Nice idea.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1131152,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/29/2020 15:10:45",
          "content": "<p>I interpreted <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1109273\" target=\"_blank\">this response</a> in that way, where <a href=\"https://www.kaggle.com/jacklebien\" target=\"_blank\">@jacklebien</a> said:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a>, no, unlike the training data, we tried to ensure presence/absence of any target species is labeled for every test file.</p>\n</blockquote>\n<p>This seems to suggest a different process was followed and at a minimum this would mean that any potential hit for a target species would have been given to a human (while it was sort of implied to me that not even all flagged suspected cases were even given for human review - I was not 100% sure on this). I guess it does not necessarily mean that every audio segment that was not even flagged for attention was given to a humans and it does not necessarily mean that a human review was done for species other than the flagged target species.</p>\n<p>I posted the topic, because I am not 100% sure on several aspects (incl. what you asked about), and obviously it makes a huge difference for how one should approach the problem.</p>\n<p>So, perhaps this should look more like this (with a question mark also on whether the test data was really all human reviewed, since that was not directly stated):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F50c4f648b64017056dd8fe7583ae8595%2FScreenshot%20from%202020-12-29%2016-16-26.png?generation=1609255058462146&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1128518": "Here's my best attempt of summarizing what I've understood (from various discussions such as [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782), [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206012) and [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866)), about the true positive, false positives and the test set, first as a diagram and then as a bulleted list.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F649192e00d29f7ee8aa1abb432e90095%2FScreenshot%20from%202020-12-27%2014-28-08.png?generation=1609075778207654&alt=media)\n* Training data was run through a computer filter that is just a [\"rudimentary detection algorithm\"](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866), so I guess we don't quite know how much we'd trust it (i.e. esp. things it does not flag might very well still be positives)\n* If the filter did not highlight a clip, we do not know what is going on in it (although we might guess that the filter is more likely to be right than wrong)\n* If the filter did not highlight a clip, it underwent human review for the one single species that the algorithm \"detected\". If this is truly that species, we get the record in true positives, otherwise in false positives (but the human reviewer would not flag other species - I'm sort of ignoring the possibility that the humans get it wrong).\n* So records in true positives are real \"1\"s, those in false positives are \"0\"s (but ones that are somehow close to \"1\"s as far as a \"rudimentary detection algorithm\" is concerned), but if a clip does not have a record in true positives or false positives, we just don't know for sure what is going on (if the rudimentary detection algorithm is great, then they might really mostly be \"0\"s)?\n* In contrast, the test set clips were completely labelled by humans.\n\nDid I get that right? The biggest question mark I really had was around whether no record in true positives and false negatives for a species in a clip really just means \"not flagged automatically\" = we don't know for sure (but not so probable depending on how good the algorithm used was)?\n\nDoes that suggest that we should use a loss function that \"insists\" on the true positives and false negatives, but is a bit more \"lenient\" about the clips we are not sure about (but might guess are more likely to be negatives than positives)?\n\nOne really big challenge is then to have audio data that corresponds to things that the correlational \"rudimentary detection algorithm\" would not flag, but which is definitely negatives. Is that right?",
    "1128649": "great work\n👍",
    "1131093": "Thanks for this summary. I noticed you upvoted my question to the organisers [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1123751). There I assumed that the test data was labelled in the same way. So my question to you is: which link says that the test data is 100% human labelled?\n\nIf the test data is indeed human labelled that I do agree with your idea of context sensitive loss. Nice idea.",
    "1131152": "I interpreted [this response](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197866#1109273) in that way, where @jacklebien said:\n> @selimsef, no, unlike the training data, we tried to ensure presence/absence of any target species is labeled for every test file.\n\nThis seems to suggest a different process was followed and at a minimum this would mean that any potential hit for a target species would have been given to a human (while it was sort of implied to me that not even all flagged suspected cases were even given for human review - I was not 100% sure on this). I guess it does not necessarily mean that every audio segment that was not even flagged for attention was given to a humans and it does not necessarily mean that a human review was done for species other than the flagged target species.\n\nI posted the topic, because I am not 100% sure on several aspects (incl. what you asked about), and obviously it makes a huge difference for how one should approach the problem.\n\nSo, perhaps this should look more like this (with a question mark also on whether the test data was really all human reviewed, since that was not directly stated):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F485314%2F50c4f648b64017056dd8fe7583ae8595%2FScreenshot%20from%202020-12-29%2016-16-26.png?generation=1609255058462146&alt=media)"
  },
  "source": "meta"
}