{
  "id": 197892,
  "title": "My understanding  of 'train_tp.csv' and 'train_fp.csv' . Am I correct ? Please Clarify .",
  "url": "/competitions/rfcx-species-audio-detection/discussion/197892",
  "author_name": "",
  "post_date": "2020-11-18T15:52:32.114633Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>So in this competition we have two csv files ** 'train_tp.csv' **and *<em>'train_fp.csv'</em>* , Apart from the fact that this two files have time localization . How they can be used for training I guess is like this .<br>\nSince this is a Multi Label Classification Competition in which an Audio can have multiple species sound present on it so while preparing the Data for training we need to have Labels in this format  . Where 'si' corresponds to i th species . So here we can use train_tp and train_fp files . Lets assume a recording has only one id numbered as 14 present in train_tp and remaining all present in train_fp so the label of this data point or recording will have 1 only at the 14th position and all zeroes . So this can be useful for training labels .<br>\nThis is what I think , Am I correct if Not please do correct me .</p>",
  "messages": [
    {
      "id": "1083108",
      "postDate": "11/18/2020 15:52:32",
      "content": "<p>So in this competition we have two csv files ** 'train_tp.csv' **and *<em>'train_fp.csv'</em>* , Apart from the fact that this two files have time localization . How they can be used for training I guess is like this .<br>\nSince this is a Multi Label Classification Competition in which an Audio can have multiple species sound present on it so while preparing the Data for training we need to have Labels in this format  . Where 'si' corresponds to i th species . So here we can use train_tp and train_fp files . Lets assume a recording has only one id numbered as 14 present in train_tp and remaining all present in train_fp so the label of this data point or recording will have 1 only at the 14th position and all zeroes . So this can be useful for training labels .<br>\nThis is what I think , Am I correct if Not please do correct me .</p>",
      "rawMarkdown": "So in this competition we have two csv files ** 'train_tp.csv' **and **'train_fp.csv'** , Apart from the fact that this two files have time localization . How they can be used for training I guess is like this .\nSince this is a Multi Label Classification Competition in which an Audio can have multiple species sound present on it so while preparing the Data for training we need to have Labels in this format <s0 , s1 , s2 , s3 , s4 , ...> . Where 'si' corresponds to i th species . So here we can use train_tp and train_fp files . Lets assume a recording has only one id numbered as 14 present in train_tp and remaining all present in train_fp so the label of this data point or recording will have 1 only at the 14th position and all zeroes . So this can be useful for training labels .\nThis is what I think , Am I correct if Not please do correct me .",
      "votes": null
    },
    {
      "id": "1131188",
      "postDate": "12/29/2020 15:28:30",
      "content": "<p>The tp and fp labels stem from the data generation process. Preliminary labelling was done using a rudimentary correlation-based algorithm, the labels were then verified by expert human annotators. The correctly labelled ones are the true positives and the false positives are, well, false positives.  <br>\nFrom what I can understand, the false positives tells us about the mistakes their labelling algorithm makes and can be used by associating high costs with those errors so that our models learn to not make the same mistakes. In my humble opinion, this will only help if a model trained on the tp data has similar false positives.  </p>",
      "rawMarkdown": "The tp and fp labels stem from the data generation process. Preliminary labelling was done using a rudimentary correlation-based algorithm, the labels were then verified by expert human annotators. The correctly labelled ones are the true positives and the false positives are, well, false positives.  \nFrom what I can understand, the false positives tells us about the mistakes their labelling algorithm makes and can be used by associating high costs with those errors so that our models learn to not make the same mistakes. In my humble opinion, this will only help if a model trained on the tp data has similar false positives.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1131188,
      "author_name": "rajsuryan",
      "author_url": "",
      "post_date": "12/29/2020 15:28:30",
      "content": "<p>The tp and fp labels stem from the data generation process. Preliminary labelling was done using a rudimentary correlation-based algorithm, the labels were then verified by expert human annotators. The correctly labelled ones are the true positives and the false positives are, well, false positives.  <br>\nFrom what I can understand, the false positives tells us about the mistakes their labelling algorithm makes and can be used by associating high costs with those errors so that our models learn to not make the same mistakes. In my humble opinion, this will only help if a model trained on the tp data has similar false positives.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1083108": "So in this competition we have two csv files ** 'train_tp.csv' **and **'train_fp.csv'** , Apart from the fact that this two files have time localization . How they can be used for training I guess is like this .\nSince this is a Multi Label Classification Competition in which an Audio can have multiple species sound present on it so while preparing the Data for training we need to have Labels in this format <s0 , s1 , s2 , s3 , s4 , ...> . Where 'si' corresponds to i th species . So here we can use train_tp and train_fp files . Lets assume a recording has only one id numbered as 14 present in train_tp and remaining all present in train_fp so the label of this data point or recording will have 1 only at the 14th position and all zeroes . So this can be useful for training labels .\nThis is what I think , Am I correct if Not please do correct me .",
    "1131188": "The tp and fp labels stem from the data generation process. Preliminary labelling was done using a rudimentary correlation-based algorithm, the labels were then verified by expert human annotators. The correctly labelled ones are the true positives and the false positives are, well, false positives.  \nFrom what I can understand, the false positives tells us about the mistakes their labelling algorithm makes and can be used by associating high costs with those errors so that our models learn to not make the same mistakes. In my humble opinion, this will only help if a model trained on the tp data has similar false positives."
  },
  "source": "meta"
}