{
  "id": 214508,
  "title": "Object Detection as augmentation",
  "url": "/competitions/rfcx-species-audio-detection/discussion/214508",
  "author_name": "ryches",
  "post_date": "2021-01-26T22:15:59.149000",
  "votes": 15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>From what I have seen so far in the discussion of this competition there seems to be two different avenues people are taking:</p>\n<ul>\n<li>Framewise training<ul>\n<li>The temporal information is used from the labels in order to make crops near the occurrence of the labeled noise</li>\n<li>model is trained on differing crops of this region as a form of augmentation to make sure the model is focused on the single noise. Or the model is trained on consistent segments 10s chunks overlapping by 5s, but for this data we do not have full labels for all segments so this can be problematic</li>\n<li>at inference time, chunks of overlapping audio are passed to the model in order to make predictions for segments and then those predictions are aggregated in some way. </li></ul></li>\n</ul>\n<p>The pros of this model is that you can focus models on the signal given specifically from the cropping information from the labels and can make sure your model is learning the signal and not the surrounding noise</p>\n<p>The con is that at inference time there is a bit of a gap between what you are doing during training and the way things are being aggregated on validation/test predictions. If the model only sees the areas surrounding the labeled noise then making it make predictions in other segments of the audio is a bit undocumented how it will behave. The hope is that false positives will not be triggered during these times, but it isn't a guarantee. </p>\n<ul>\n<li>Clipwise training<ul>\n<li>This training is somewhat throwing away the temporal/frequency bounding boxes we have been given and just operating as if we have weak supervision signal</li>\n<li>The pro of this is that you know your model is doing exactly the same thing between train and validation/test so performance you see on train is likely to more highly correlate at inference time</li>\n<li>The con of this is with very few clips this makes for very few training samples, strong augmentation or some other method will need to be introduced to prevent the model from learning the several hundred clips we have. </li>\n<li>Another con is that the input images are either too large to reasonably train on or too coarse of a resolution to show reasonable detail</li></ul></li>\n</ul>\n<p>My belief is that full clipwise training likely will not be fruitful, but framewise also has the downfall of missing many labels in this data. Depending on how training is organized for it it is either being shown frames that are mislabeled, model is told no important sound was there when it was simply unlabeled, or the model is only focusing on the area that was labeled and is missing out on opportunities to train on additional surrounding information. </p>\n<p>To combat both of these scenarios I think it might be reasonable to train an object detection model using varying crops around the labeled fmin, fmax, tmin, tmax as the bounding box/class labels and then applying this object detection model to try to apply labels to regions that are unlabeled. This model might be useful on its own to make predictions but it is more likely to just be valuable at augmenting the samples the framewise model looks at. </p>\n<p>For example, as noted in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/202456\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/202456</a> there are many potentially repeated occurrences of similar audio snippets from birds or frogs that went unlabeled that could likely be found by an object detection model. Like in the below picture, that squiggle is repeated a couple more times with alternate surrounding noise patterns. Might be valuable for generating realistic \"augmentations\", by simply extracting more missed labels from the real audio<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd603c20b494dc6260825a343b0258e70%2Fcrop.png?generation=1611699198552461&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1171509,
      "postDate": "2021-01-26T22:15:59.150Z",
      "content": "<p>From what I have seen so far in the discussion of this competition there seems to be two different avenues people are taking:</p>\n<ul>\n<li>Framewise training<ul>\n<li>The temporal information is used from the labels in order to make crops near the occurrence of the labeled noise</li>\n<li>model is trained on differing crops of this region as a form of augmentation to make sure the model is focused on the single noise. Or the model is trained on consistent segments 10s chunks overlapping by 5s, but for this data we do not have full labels for all segments so this can be problematic</li>\n<li>at inference time, chunks of overlapping audio are passed to the model in order to make predictions for segments and then those predictions are aggregated in some way. </li></ul></li>\n</ul>\n<p>The pros of this model is that you can focus models on the signal given specifically from the cropping information from the labels and can make sure your model is learning the signal and not the surrounding noise</p>\n<p>The con is that at inference time there is a bit of a gap between what you are doing during training and the way things are being aggregated on validation/test predictions. If the model only sees the areas surrounding the labeled noise then making it make predictions in other segments of the audio is a bit undocumented how it will behave. The hope is that false positives will not be triggered during these times, but it isn't a guarantee. </p>\n<ul>\n<li>Clipwise training<ul>\n<li>This training is somewhat throwing away the temporal/frequency bounding boxes we have been given and just operating as if we have weak supervision signal</li>\n<li>The pro of this is that you know your model is doing exactly the same thing between train and validation/test so performance you see on train is likely to more highly correlate at inference time</li>\n<li>The con of this is with very few clips this makes for very few training samples, strong augmentation or some other method will need to be introduced to prevent the model from learning the several hundred clips we have. </li>\n<li>Another con is that the input images are either too large to reasonably train on or too coarse of a resolution to show reasonable detail</li></ul></li>\n</ul>\n<p>My belief is that full clipwise training likely will not be fruitful, but framewise also has the downfall of missing many labels in this data. Depending on how training is organized for it it is either being shown frames that are mislabeled, model is told no important sound was there when it was simply unlabeled, or the model is only focusing on the area that was labeled and is missing out on opportunities to train on additional surrounding information. </p>\n<p>To combat both of these scenarios I think it might be reasonable to train an object detection model using varying crops around the labeled fmin, fmax, tmin, tmax as the bounding box/class labels and then applying this object detection model to try to apply labels to regions that are unlabeled. This model might be useful on its own to make predictions but it is more likely to just be valuable at augmenting the samples the framewise model looks at. </p>\n<p>For example, as noted in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/202456\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/202456</a> there are many potentially repeated occurrences of similar audio snippets from birds or frogs that went unlabeled that could likely be found by an object detection model. Like in the below picture, that squiggle is repeated a couple more times with alternate surrounding noise patterns. Might be valuable for generating realistic \"augmentations\", by simply extracting more missed labels from the real audio<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd603c20b494dc6260825a343b0258e70%2Fcrop.png?generation=1611699198552461&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "From what I have seen so far in the discussion of this competition there seems to be two different avenues people are taking:\n\n- Framewise training\n  - The temporal information is used from the labels in order to make crops near the occurrence of the labeled noise\n  - model is trained on differing crops of this region as a form of augmentation to make sure the model is focused on the single noise. Or the model is trained on consistent segments 10s chunks overlapping by 5s, but for this data we do not have full labels for all segments so this can be problematic\n  - at inference time, chunks of overlapping audio are passed to the model in order to make predictions for segments and then those predictions are aggregated in some way. \n\nThe pros of this model is that you can focus models on the signal given specifically from the cropping information from the labels and can make sure your model is learning the signal and not the surrounding noise\n\nThe con is that at inference time there is a bit of a gap between what you are doing during training and the way things are being aggregated on validation/test predictions. If the model only sees the areas surrounding the labeled noise then making it make predictions in other segments of the audio is a bit undocumented how it will behave. The hope is that false positives will not be triggered during these times, but it isn't a guarantee. \n\n- Clipwise training\n  - This training is somewhat throwing away the temporal/frequency bounding boxes we have been given and just operating as if we have weak supervision signal\n  - The pro of this is that you know your model is doing exactly the same thing between train and validation/test so performance you see on train is likely to more highly correlate at inference time\n  - The con of this is with very few clips this makes for very few training samples, strong augmentation or some other method will need to be introduced to prevent the model from learning the several hundred clips we have. \n  - Another con is that the input images are either too large to reasonably train on or too coarse of a resolution to show reasonable detail\n\nMy belief is that full clipwise training likely will not be fruitful, but framewise also has the downfall of missing many labels in this data. Depending on how training is organized for it it is either being shown frames that are mislabeled, model is told no important sound was there when it was simply unlabeled, or the model is only focusing on the area that was labeled and is missing out on opportunities to train on additional surrounding information. \n\nTo combat both of these scenarios I think it might be reasonable to train an object detection model using varying crops around the labeled fmin, fmax, tmin, tmax as the bounding box/class labels and then applying this object detection model to try to apply labels to regions that are unlabeled. This model might be useful on its own to make predictions but it is more likely to just be valuable at augmenting the samples the framewise model looks at. \n\n\nFor example, as noted in https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/202456 there are many potentially repeated occurrences of similar audio snippets from birds or frogs that went unlabeled that could likely be found by an object detection model. Like in the below picture, that squiggle is repeated a couple more times with alternate surrounding noise patterns. Might be valuable for generating realistic \"augmentations\", by simply extracting more missed labels from the real audio\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd603c20b494dc6260825a343b0258e70%2Fcrop.png?generation=1611699198552461&alt=media)",
      "votes": 15
    },
    {
      "id": 1171608,
      "postDate": "2021-01-27T01:43:34.190Z",
      "content": "<p>One thing I think people are missing is that the <strong>t_min</strong> and <strong>t_max</strong> are also very noisy. If you calculate <strong>length = t_max - t_min</strong> from both tp and fp files and group by classes, you can see there are only 1-3 different lengths per class. This is due to how the ground truth labels are generated. I believe it was done by some query-by-example search and the query lengths could tell us how many queries were used for each class.</p>\n<p>If the assumptions above were true, <strong>t_min</strong> and <strong>t_max</strong> would be very noisy. Plus there are so many missing labels due to the labelling procedure, I don't think object detection would work here without hand relabeling</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1605499%2F11b634688c482ee70acf678fa348c1fc%2FScreenshot%202021-01-26%20174305.png?generation=1611711811293659&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "One thing I think people are missing is that the **t_min** and **t_max** are also very noisy. If you calculate **length = t_max - t_min** from both tp and fp files and group by classes, you can see there are only 1-3 different lengths per class. This is due to how the ground truth labels are generated. I believe it was done by some query-by-example search and the query lengths could tell us how many queries were used for each class.\n\nIf the assumptions above were true, **t_min** and **t_max** would be very noisy. Plus there are so many missing labels due to the labelling procedure, I don't think object detection would work here without hand relabeling\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1605499%2F11b634688c482ee70acf678fa348c1fc%2FScreenshot%202021-01-26%20174305.png?generation=1611711811293659&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 1171627,
          "postDate": "2021-01-27T02:01:39.523Z",
          "content": "<p>Good points. I was not aware that there were only a few different lengths. I believe the organizers mentioned the process was some automatic system that was then reviewed by experts. Might have been some system similar to many are using that splits audio into frames and classifies and then had experts verify, hence the tp and fp files. </p>",
          "rawMarkdown": "Good points. I was not aware that there were only a few different lengths. I believe the organizers mentioned the process was some automatic system that was then reviewed by experts. Might have been some system similar to many are using that splits audio into frames and classifies and then had experts verify, hence the tp and fp files. ",
          "votes": 1
        },
        {
          "id": 1172256,
          "postDate": "2021-01-27T10:20:01.553Z",
          "content": "<p><a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954120300637</a><br>\nIf it follows the above procedure, it's template matching search on relevant frequencies based on correlation (section 2.2.2)</p>",
          "rawMarkdown": "https://www.sciencedirect.com/science/article/pii/S1574954120300637\nIf it follows the above procedure, it's template matching search on relevant frequencies based on correlation (section 2.2.2)",
          "votes": 4
        },
        {
          "id": 1173614,
          "postDate": "2021-01-28T02:22:47.840Z",
          "content": "<p>It seems this also gives away names of 24 species</p>",
          "rawMarkdown": "It seems this also gives away names of 24 species",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1183368,
      "postDate": "2021-02-02T23:43:18.490Z",
      "content": "<p>Tried Faster RCNN and got LB 0.750. Got similar problem like my classification models: species 3 and 7 got really low recall and precision.</p>",
      "rawMarkdown": "Tried Faster RCNN and got LB 0.750. Got similar problem like my classification models: species 3 and 7 got really low recall and precision.",
      "votes": 1
    },
    {
      "id": 1171547,
      "postDate": "2021-01-26T23:40:46.703Z",
      "content": "<p>I can't wait to see your LB score given all the ideas you kindly share.</p>",
      "rawMarkdown": "I can't wait to see your LB score given all the ideas you kindly share.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1171608,
      "author_name": "Darkate",
      "author_url": "",
      "post_date": "2021-01-27T01:43:34.190000",
      "content": "<p>One thing I think people are missing is that the <strong>t_min</strong> and <strong>t_max</strong> are also very noisy. If you calculate <strong>length = t_max - t_min</strong> from both tp and fp files and group by classes, you can see there are only 1-3 different lengths per class. This is due to how the ground truth labels are generated. I believe it was done by some query-by-example search and the query lengths could tell us how many queries were used for each class.</p>\n<p>If the assumptions above were true, <strong>t_min</strong> and <strong>t_max</strong> would be very noisy. Plus there are so many missing labels due to the labelling procedure, I don't think object detection would work here without hand relabeling</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1605499%2F11b634688c482ee70acf678fa348c1fc%2FScreenshot%202021-01-26%20174305.png?generation=1611711811293659&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1171627,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-01-27T02:01:39.523000",
          "content": "<p>Good points. I was not aware that there were only a few different lengths. I believe the organizers mentioned the process was some automatic system that was then reviewed by experts. Might have been some system similar to many are using that splits audio into frames and classifies and then had experts verify, hence the tp and fp files. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1172256,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2021-01-27T10:20:01.553000",
          "content": "<p><a href=\"https://www.sciencedirect.com/science/article/pii/S1574954120300637\" target=\"_blank\">https://www.sciencedirect.com/science/article/pii/S1574954120300637</a><br>\nIf it follows the above procedure, it's template matching search on relevant frequencies based on correlation (section 2.2.2)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1173614,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-28T02:22:47.840000",
          "content": "<p>It seems this also gives away names of 24 species</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1183368,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "2021-02-02T23:43:18.490000",
      "content": "<p>Tried Faster RCNN and got LB 0.750. Got similar problem like my classification models: species 3 and 7 got really low recall and precision.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1171547,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-01-26T23:40:46.703000",
      "content": "<p>I can't wait to see your LB score given all the ideas you kindly share.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1171509": "From what I have seen so far in the discussion of this competition there seems to be two different avenues people are taking:\n\n- Framewise training\n  - The temporal information is used from the labels in order to make crops near the occurrence of the labeled noise\n  - model is trained on differing crops of this region as a form of augmentation to make sure the model is focused on the single noise. Or the model is trained on consistent segments 10s chunks overlapping by 5s, but for this data we do not have full labels for all segments so this can be problematic\n  - at inference time, chunks of overlapping audio are passed to the model in order to make predictions for segments and then those predictions are aggregated in some way. \n\nThe pros of this model is that you can focus models on the signal given specifically from the cropping information from the labels and can make sure your model is learning the signal and not the surrounding noise\n\nThe con is that at inference time there is a bit of a gap between what you are doing during training and the way things are being aggregated on validation/test predictions. If the model only sees the areas surrounding the labeled noise then making it make predictions in other segments of the audio is a bit undocumented how it will behave. The hope is that false positives will not be triggered during these times, but it isn't a guarantee. \n\n- Clipwise training\n  - This training is somewhat throwing away the temporal/frequency bounding boxes we have been given and just operating as if we have weak supervision signal\n  - The pro of this is that you know your model is doing exactly the same thing between train and validation/test so performance you see on train is likely to more highly correlate at inference time\n  - The con of this is with very few clips this makes for very few training samples, strong augmentation or some other method will need to be introduced to prevent the model from learning the several hundred clips we have. \n  - Another con is that the input images are either too large to reasonably train on or too coarse of a resolution to show reasonable detail\n\nMy belief is that full clipwise training likely will not be fruitful, but framewise also has the downfall of missing many labels in this data. Depending on how training is organized for it it is either being shown frames that are mislabeled, model is told no important sound was there when it was simply unlabeled, or the model is only focusing on the area that was labeled and is missing out on opportunities to train on additional surrounding information. \n\nTo combat both of these scenarios I think it might be reasonable to train an object detection model using varying crops around the labeled fmin, fmax, tmin, tmax as the bounding box/class labels and then applying this object detection model to try to apply labels to regions that are unlabeled. This model might be useful on its own to make predictions but it is more likely to just be valuable at augmenting the samples the framewise model looks at. \n\n\nFor example, as noted in https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/202456 there are many potentially repeated occurrences of similar audio snippets from birds or frogs that went unlabeled that could likely be found by an object detection model. Like in the below picture, that squiggle is repeated a couple more times with alternate surrounding noise patterns. Might be valuable for generating realistic \"augmentations\", by simply extracting more missed labels from the real audio\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd603c20b494dc6260825a343b0258e70%2Fcrop.png?generation=1611699198552461&alt=media)",
    "1171608": "One thing I think people are missing is that the **t_min** and **t_max** are also very noisy. If you calculate **length = t_max - t_min** from both tp and fp files and group by classes, you can see there are only 1-3 different lengths per class. This is due to how the ground truth labels are generated. I believe it was done by some query-by-example search and the query lengths could tell us how many queries were used for each class.\n\nIf the assumptions above were true, **t_min** and **t_max** would be very noisy. Plus there are so many missing labels due to the labelling procedure, I don't think object detection would work here without hand relabeling\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1605499%2F11b634688c482ee70acf678fa348c1fc%2FScreenshot%202021-01-26%20174305.png?generation=1611711811293659&alt=media)",
    "1183368": "Tried Faster RCNN and got LB 0.750. Got similar problem like my classification models: species 3 and 7 got really low recall and precision.",
    "1171547": "I can't wait to see your LB score given all the ideas you kindly share."
  }
}