{
  "id": 300405,
  "title": "[placeholder] how to get lb 0.560 with single model",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/300405",
  "author_name": "hengck23",
  "post_date": "2022-01-12T14:01:03.464000",
  "votes": 152,
  "comment_count": 55,
  "views": 0,
  "content": "<p>After experimenting for a week, i am ready to share my intermediate solution.<br>\nFollow this thread for the continuous update!</p>\n<p><br>\nrefer to google drive:<br>\n<a href=\"https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing</a><br>\n2022-02-13.tar.gz</p>\n<hr>\n<p>summary of approach</p>\n<p>With designing augmentation for object detection, there is two main consideration:</p>\n<ol>\n<li>target object appearance and pose</li>\n<li>object layout in the scene</li>\n</ol>\n<p>for 1. it is easy to test.  the median size of the target COTS object is about 32 (ranges from 20 to 128). Hence we can actually use 32x32, 64x64, 128x128 image classifier. Recall is important in the f2 metric. you have to have recall of about 60% to 80%.  you can keep precision to about 80% to 90% and FPrate to 0.20 per image.</p>\n<p>you can use feature of image classifier and make a tsne map to see the distribution of the target objects. which is the rare subclass? (actually feature from object detector also will work)</p>\n<p>Now for 2. if you have a very good image classifier (i.e. good representative train data/augmentation),  the object detection model should work well too (assuming you have no problem to post processing like NMS). if you do not do well here, then it  could be:</p>\n<ul>\n<li>wrong anchor box size</li>\n<li>the object size in the validation set is different from the train</li>\n<li>sampling problem ( object vs background imbalance ratio)</li>\n</ul>\n<p>It is all about scaling and rotation ( convolution is invariant to translation, but not scale and rotation). This is why it fails when you move from classifier to object detection.</p>\n<p>when you do augmentation,  so bbox gets too small and ends up partially exceeding the image boundary. you should filter those if they hurt results.</p>\n<p>also note that the number of bbox per image is low. it is likely that after affine augmentation (rotate, enlarge, etc), you can end up with any empty image without bbox.</p>\n<p>(do note that the whole idea about augmentation is to create good train samples and also some good noise so that the model will not be overconfident on the prediction. noise balancing is difficult)</p>\n<p>one easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector</p>\n<p>for splitting, it is sufficient to use train=video0,1, valid=video2 to achieve lb 0.560. But for experimenting with augmentation, you have to try all 3 folds using video0,1,2 as validation.</p>\n<hr>\n<p>I am grateful to HP for providing a Z by HP workstation for this competition. Equipped with 2x A600 Nvidia cards, I can complete my experiments in a short time and share the results and findings with the kaggle community. </p>",
  "messages": [
    {
      "id": 1647337,
      "postDate": "2022-01-12T14:01:03.463Z",
      "content": "<p>After experimenting for a week, i am ready to share my intermediate solution.<br>\nFollow this thread for the continuous update!</p>\n<p><br>\nrefer to google drive:<br>\n<a href=\"https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing</a><br>\n2022-02-13.tar.gz</p>\n<hr>\n<p>summary of approach</p>\n<p>With designing augmentation for object detection, there is two main consideration:</p>\n<ol>\n<li>target object appearance and pose</li>\n<li>object layout in the scene</li>\n</ol>\n<p>for 1. it is easy to test.  the median size of the target COTS object is about 32 (ranges from 20 to 128). Hence we can actually use 32x32, 64x64, 128x128 image classifier. Recall is important in the f2 metric. you have to have recall of about 60% to 80%.  you can keep precision to about 80% to 90% and FPrate to 0.20 per image.</p>\n<p>you can use feature of image classifier and make a tsne map to see the distribution of the target objects. which is the rare subclass? (actually feature from object detector also will work)</p>\n<p>Now for 2. if you have a very good image classifier (i.e. good representative train data/augmentation),  the object detection model should work well too (assuming you have no problem to post processing like NMS). if you do not do well here, then it  could be:</p>\n<ul>\n<li>wrong anchor box size</li>\n<li>the object size in the validation set is different from the train</li>\n<li>sampling problem ( object vs background imbalance ratio)</li>\n</ul>\n<p>It is all about scaling and rotation ( convolution is invariant to translation, but not scale and rotation). This is why it fails when you move from classifier to object detection.</p>\n<p>when you do augmentation,  so bbox gets too small and ends up partially exceeding the image boundary. you should filter those if they hurt results.</p>\n<p>also note that the number of bbox per image is low. it is likely that after affine augmentation (rotate, enlarge, etc), you can end up with any empty image without bbox.</p>\n<p>(do note that the whole idea about augmentation is to create good train samples and also some good noise so that the model will not be overconfident on the prediction. noise balancing is difficult)</p>\n<p>one easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector</p>\n<p>for splitting, it is sufficient to use train=video0,1, valid=video2 to achieve lb 0.560. But for experimenting with augmentation, you have to try all 3 folds using video0,1,2 as validation.</p>\n<hr>\n<p>I am grateful to HP for providing a Z by HP workstation for this competition. Equipped with 2x A600 Nvidia cards, I can complete my experiments in a short time and share the results and findings with the kaggle community. </p>",
      "rawMarkdown": "After experimenting for a week, i am ready to share my intermediate solution.\nFollow this thread for the continuous update!\n\n\n~~Code (dirty code for illustration purpose) will be updated in the next few days.~~\nrefer to google drive:\nhttps://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\n2022-02-13.tar.gz\n\n----\nsummary of approach\n\nWith designing augmentation for object detection, there is two main consideration:\n1. target object appearance and pose\n2. object layout in the scene\n\nfor 1. it is easy to test.  the median size of the target COTS object is about 32 (ranges from 20 to 128). Hence we can actually use 32x32, 64x64, 128x128 image classifier. Recall is important in the f2 metric. you have to have recall of about 60% to 80%.  you can keep precision to about 80% to 90% and FPrate to 0.20 per image.\n\nyou can use feature of image classifier and make a tsne map to see the distribution of the target objects. which is the rare subclass? (actually feature from object detector also will work)\n\n\nNow for 2. if you have a very good image classifier (i.e. good representative train data/augmentation),  the object detection model should work well too (assuming you have no problem to post processing like NMS). if you do not do well here, then it  could be:\n- wrong anchor box size\n- the object size in the validation set is different from the train\n- sampling problem ( object vs background imbalance ratio)\n\nIt is all about scaling and rotation ( convolution is invariant to translation, but not scale and rotation). This is why it fails when you move from classifier to object detection.\n\nwhen you do augmentation,  so bbox gets too small and ends up partially exceeding the image boundary. you should filter those if they hurt results.\n\nalso note that the number of bbox per image is low. it is likely that after affine augmentation (rotate, enlarge, etc), you can end up with any empty image without bbox.\n\n(do note that the whole idea about augmentation is to create good train samples and also some good noise so that the model will not be overconfident on the prediction. noise balancing is difficult)\n\none easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector\n\n\n\nfor splitting, it is sufficient to use train=video0,1, valid=video2 to achieve lb 0.560. But for experimenting with augmentation, you have to try all 3 folds using video0,1,2 as validation.\n\n----\n\nI am grateful to HP for providing a Z by HP workstation for this competition. Equipped with 2x A600 Nvidia cards, I can complete my experiments in a short time and share the results and findings with the kaggle community. ",
      "votes": 151
    },
    {
      "id": 1650379,
      "postDate": "2022-01-15T02:29:16.157Z",
      "content": "<p>final solution would be someting like this:<br>\n<img src=\"https://i.ibb.co/FH9Lrzy/Selection-999-723.png\" alt=\"https://i.ibb.co/FH9Lrzy/Selection-999-723.png\"></p>",
      "rawMarkdown": "final solution would be someting like this:\n![https://i.ibb.co/FH9Lrzy/Selection-999-723.png](https://i.ibb.co/FH9Lrzy/Selection-999-723.png)",
      "votes": 13,
      "replies": [
        {
          "id": 1650413,
          "postDate": "2022-01-15T03:19:17.833Z",
          "content": "<p>Yeah, it seems upscaling the image, despite being compute expensive, is resulting in a large leaderboard jump right now. I think your approach of cropping regions of interest and then upscaling is a good solution for this performance trade off. Also as a side note, I've only had dissapointment with two-stage detectors compared to YOLO computation cost wise</p>",
          "rawMarkdown": "Yeah, it seems upscaling the image, despite being compute expensive, is resulting in a large leaderboard jump right now. I think your approach of cropping regions of interest and then upscaling is a good solution for this performance trade off. Also as a side note, I've only had dissapointment with two-stage detectors compared to YOLO computation cost wise"
        },
        {
          "id": 1650845,
          "postDate": "2022-01-15T11:20:49.967Z",
          "content": "<p>it takes time. I will try it and add some new ideas.</p>",
          "rawMarkdown": "it takes time. I will try it and add some new ideas.",
          "votes": -1
        },
        {
          "id": 1650871,
          "postDate": "2022-01-15T11:34:18.250Z",
          "content": "<p>After the end of the competition I will still need 1 more month to test all the hypotheses you raised. As always your posts are pure gold. Congratulations</p>",
          "rawMarkdown": "After the end of the competition I will still need 1 more month to test all the hypotheses you raised. As always your posts are pure gold. Congratulations",
          "votes": 1
        }
      ]
    },
    {
      "id": 1648383,
      "postDate": "2022-01-13T11:48:30.277Z",
      "content": "<p>here is a list of good papers you can read. Out use case is similar to object detection I'm monocular video for unmanned aerial vehicle (UAV), particularly underwater drone.</p>\n<ol>\n<li>TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. <br>\n<a href=\"https://arxiv.org/abs/2108.11539\" target=\"_blank\">https://arxiv.org/abs/2108.11539</a><br>\n5th in On VisDrone Challenge 2021<br>\n(1st to 4th place: DBNet, SOLOer, Swin-T, DPNetV4)</li>\n</ol>\n<p>ECAP-YOLO: Efficient Channel Attention Pyramid YOLO for Small Object Detection in Aerial Imag<br>\nanother modification</p>\n<p>remotesensing-13-04851-v2.pdf</p>\n<ol>\n<li>Bag of Freebies for Training Object Detection Neural Networks<br>\n<a href=\"https://arxiv.org/abs/1902.04103\" target=\"_blank\">https://arxiv.org/abs/1902.04103</a></li>\n</ol>\n<hr>\n<p>PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices<br>\n<a href=\"https://arxiv.org/abs/2111.00902\" target=\"_blank\">https://arxiv.org/abs/2111.00902</a></p>\n<p>A ConvNet for the 2020s<br>\n<a href=\"https://arxiv.org/abs/2201.03545\" target=\"_blank\">https://arxiv.org/abs/2201.03545</a></p>\n<p>Augmenting Convolutional networks with attention-based aggregation<br>\n<a href=\"https://arxiv.org/pdf/2112.13692v1\" target=\"_blank\">https://arxiv.org/pdf/2112.13692v1</a><br>\n\", as our approach is not pyramidal, we only use the final output of our network in Mask R-CNN\"</p>\n<hr>\n<p>slam+object detection as an alternative to detection+tracking<br>\n<a href=\"https://www.youtube.com/watch?v=m6sStUk3UVk\" target=\"_blank\">https://www.youtube.com/watch?v=m6sStUk3UVk</a></p>\n<hr>\n<p>MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks<br>\n<a href=\"https://arxiv.org/pdf/2009.08453\" target=\"_blank\">https://arxiv.org/pdf/2009.08453</a></p>\n<hr>\n<p>some zooming based network:<br>\nDynamic Zoom-in Network for Fast Object Detection in Large Images<br>\nDense and Small Object Detection in UAV Vision based on Cascade Network</p>\n<p>maybe we can add coordinate feature (and frame number feature) to make the model region aware.<br>\n<img src=\"https://images.deepai.org/converted-papers/1711.05187/x1.png\" alt=\"https://images.deepai.org/converted-papers/1711.05187/x1.png\"></p>\n<hr>\n<p>better moasic. we have single class with different subsclass<br>\nImage-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection<br>\n<a href=\"https://arxiv.org/pdf/2104.05702\" target=\"_blank\">https://arxiv.org/pdf/2104.05702</a></p>\n<hr>\n<p>better resolution</p>\n<p>Fixing the train-test resolution discrepancy<br>\n<img src=\"https://github.com/facebookresearch/FixRes/raw/main/image/image2.png\" alt=\"https://github.com/facebookresearch/FixRes/raw/main/image/image2.png\"></p>\n<p>MIX &amp; MATCH: TRAINING CONVNETS WITH MIXED IMAGE SIZES FOR IMPROVED ACCURACY, SPEED AND<br>\nSCALE RESILIENCY<br>\n\"For instance, we receive a 76.43% top-1 accuracy using ResNet50 with an image size of 160,\"</p>\n<p>Learning to Resize Images for Computer Vision Tasks<br>\n<a href=\"https://arxiv.org/pdf/2003.08237.pdf\" target=\"_blank\">https://arxiv.org/pdf/2003.08237.pdf</a></p>\n<p>temporal (and spatial) frame interpolation to create more training data?</p>\n<p>Perceptual Generative Adversarial Networks for Small Object Detection</p>\n<hr>\n<p>one can treat extreme augmentation as unlabelled data<br>\nBeyond Self-Supervision: A Simple Yet Effective Network Distillation<br>\nAlternative to Improve Backbones</p>\n<p>\"Ours(ResNet50-D+ResFix) 5.2M 84.0 (top1)\"</p>",
      "rawMarkdown": "here is a list of good papers you can read. Out use case is similar to object detection I'm monocular video for unmanned aerial vehicle (UAV), particularly underwater drone.\n\n\n1. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. \nhttps://arxiv.org/abs/2108.11539\n5th in On VisDrone Challenge 2021\n(1st to 4th place: DBNet, SOLOer, Swin-T, DPNetV4)\n\nECAP-YOLO: Efficient Channel Attention Pyramid YOLO for Small Object Detection in Aerial Imag\nanother modification\n\nremotesensing-13-04851-v2.pdf\n\n2. Bag of Freebies for Training Object Detection Neural Networks\n https://arxiv.org/abs/1902.04103\n\n---\n\nPP-PicoDet: A Better Real-Time Object Detector on Mobile Devices\nhttps://arxiv.org/abs/2111.00902\n\nA ConvNet for the 2020s\nhttps://arxiv.org/abs/2201.03545\n\nAugmenting Convolutional networks with attention-based aggregation\nhttps://arxiv.org/pdf/2112.13692v1\n\", as our approach is not pyramidal, we only use the final output of our network in Mask R-CNN\"\n\n---\nslam+object detection as an alternative to detection+tracking\nhttps://www.youtube.com/watch?v=m6sStUk3UVk\n\n---\nMEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks\nhttps://arxiv.org/pdf/2009.08453\n\n---\n\nsome zooming based network:\nDynamic Zoom-in Network for Fast Object Detection in Large Images\nDense and Small Object Detection in UAV Vision based on Cascade Network\n\nmaybe we can add coordinate feature (and frame number feature) to make the model region aware.\n![https://images.deepai.org/converted-papers/1711.05187/x1.png](https://images.deepai.org/converted-papers/1711.05187/x1.png)\n\n---\n\nbetter moasic. we have single class with different subsclass\nImage-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection\nhttps://arxiv.org/pdf/2104.05702\n\n---\n\nbetter resolution\n\nFixing the train-test resolution discrepancy\n![https://github.com/facebookresearch/FixRes/raw/main/image/image2.png](https://github.com/facebookresearch/FixRes/raw/main/image/image2.png)\n\nMIX & MATCH: TRAINING CONVNETS WITH MIXED IMAGE SIZES FOR IMPROVED ACCURACY, SPEED AND\nSCALE RESILIENCY\n\"For instance, we receive a 76.43% top-1 accuracy using ResNet50 with an image size of 160,\"\n\nLearning to Resize Images for Computer Vision Tasks\nhttps://arxiv.org/pdf/2003.08237.pdf\n\ntemporal (and spatial) frame interpolation to create more training data?\n\nPerceptual Generative Adversarial Networks for Small Object Detection\n\n---\n\none can treat extreme augmentation as unlabelled data\nBeyond Self-Supervision: A Simple Yet Effective Network Distillation\nAlternative to Improve Backbones\n\n\"Ours(ResNet50-D+ResFix) 5.2M 84.0 (top1)\"",
      "votes": 12
    },
    {
      "id": 1648132,
      "postDate": "2022-01-13T07:25:13.623Z",
      "content": "<p>experiment chart:<br>\n<img src=\"https://i.ibb.co/F7t2fjk/Selection-999-712.png\" alt=\"https://i.ibb.co/F7t2fjk/Selection-999-712.png\"></p>",
      "rawMarkdown": "experiment chart:\n![https://i.ibb.co/F7t2fjk/Selection-999-712.png](https://i.ibb.co/F7t2fjk/Selection-999-712.png)",
      "votes": 9,
      "replies": [
        {
          "id": 1648147,
          "postDate": "2022-01-13T07:49:42.907Z",
          "content": "<p>Would you retrain a final model using all videos as training data for final submission. I assume this model would have better generalization?</p>",
          "rawMarkdown": "Would you retrain a final model using all videos as training data for final submission. I assume this model would have better generalization?"
        },
        {
          "id": 1648245,
          "postDate": "2022-01-13T09:24:01.917Z",
          "content": "<ul>\n<li><p>instead of spending time to determine a better split, my strategy is to create my train samples and probe the LB to see what kind of data to generate</p></li>\n<li><p>my next step is to use multi-frame for detection (not really a tracking problem as we don't have to associate track ID)</p></li>\n</ul>",
          "rawMarkdown": "- instead of spending time to determine a better split, my strategy is to create my train samples and probe the LB to see what kind of data to generate\n\n- my next step is to use multi-frame for detection (not really a tracking problem as we don't have to associate track ID)",
          "votes": 3
        },
        {
          "id": 1648304,
          "postDate": "2022-01-13T10:23:28.090Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , thanks a lot for this useful analysis.</p>\n<p>May I ask how you deal with unlabeled images, i.e. do you include (part of) them in the training and are they included for the calculation of your CV scores above?</p>\n<p>I'm currently including a small amount of unlabeled images in my training (~3% of the labeled training data) and I include all the unlabeled images for CV scoring. I was wondering how you and maybe others deal with them.</p>",
          "rawMarkdown": "@hengck23 , thanks a lot for this useful analysis.\n\nMay I ask how you deal with unlabeled images, i.e. do you include (part of) them in the training and are they included for the calculation of your CV scores above?\n\nI'm currently including a small amount of unlabeled images in my training (~3% of the labeled training data) and I include all the unlabeled images for CV scoring. I was wondering how you and maybe others deal with them.",
          "votes": 1
        },
        {
          "id": 1648328,
          "postDate": "2022-01-13T10:55:17.047Z",
          "content": "<p><a href=\"https://www.kaggle.com/omallo\" target=\"_blank\">@omallo</a> your approach is ok and smililar to mine.<br>\nFor  N images in each batch or (for very N images in M batch iterations), just make sure :</p>\n<ul>\n<li>there are x% of empty image</li>\n<li>there are (1-x)% of non-empty image. And there are on average som eN numbers of truth bboxes, of various sizes and appearances</li>\n</ul>",
          "rawMarkdown": "@omallo your approach is ok and smililar to mine.\nFor  N images in each batch or (for very N images in M batch iterations), just make sure :\n- there are x% of empty image\n- there are (1-x)% of non-empty image. And there are on average som eN numbers of truth bboxes, of various sizes and appearances",
          "votes": 4
        },
        {
          "id": 1649188,
          "postDate": "2022-01-14T03:54:19.123Z",
          "content": "<p>Seem higher resolution is better</p>",
          "rawMarkdown": "Seem higher resolution is better"
        },
        {
          "id": 1650532,
          "postDate": "2022-01-15T05:20:22.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> do you mean to have better cv with <strong>aspect ratio</strong> or <strong>area</strong> play an important role</p>",
          "rawMarkdown": "@hengck23 do you mean to have better cv with **aspect ratio** or **area** play an important role",
          "votes": -1
        }
      ]
    },
    {
      "id": 1647410,
      "postDate": "2022-01-12T15:27:45.230Z",
      "content": "<p>Wow, leaderboard inflation is coming!</p>",
      "rawMarkdown": "Wow, leaderboard inflation is coming!",
      "votes": 7
    },
    {
      "id": 1648264,
      "postDate": "2022-01-13T09:45:59.667Z",
      "content": "<p>dirty code is up:<br>\n<a href=\"https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing</a><br>\n2022-02-13.tar.gz</p>\n<p>note:</p>\n<ul>\n<li>i no longer work on this code as i setup an updated version at my side</li>\n<li>feel free to use this code in any way, e.g. modify it and put it in your notebook, etc</li>\n<li>code is dirty, may have missing unimportant functions, etc (I copied the files from a larger project). But it should be easy to figure out the missing parts.</li>\n</ul>",
      "rawMarkdown": "dirty code is up:\nhttps://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\n2022-02-13.tar.gz\n\nnote:\n- i no longer work on this code as i setup an updated version at my side\n- feel free to use this code in any way, e.g. modify it and put it in your notebook, etc\n- code is dirty, may have missing unimportant functions, etc (I copied the files from a larger project). But it should be easy to figure out the missing parts.",
      "votes": 5,
      "replies": [
        {
          "id": 1650536,
          "postDate": "2022-01-15T05:23:06.497Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> will you add your code link in topic. So, everyone learn from it. Your knowledge sharing helped to learn new things faster. Thanks for sharing your experiments save time and energy.</p>",
          "rawMarkdown": "@hengck23 will you add your code link in topic. So, everyone learn from it. Your knowledge sharing helped to learn new things faster. Thanks for sharing your experiments save time and energy."
        },
        {
          "id": 1650544,
          "postDate": "2022-01-15T05:32:27.173Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> one question regarding data, in ur data set is cleaned version, did u resolved the missing COTS annotations in the sequence frame based or any other approach ? </p>",
          "rawMarkdown": "@hengck23 one question regarding data, in ur data set is cleaned version, did u resolved the missing COTS annotations in the sequence frame based or any other approach ? "
        },
        {
          "id": 1652470,
          "postDate": "2022-01-16T16:51:15.127Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> it's missing my_bbox.py  <br>\nWhat does this function do : xywh2cxcywh ? </p>",
          "rawMarkdown": "@hengck23 it's missing my_bbox.py  \nWhat does this function do : xywh2cxcywh ? "
        },
        {
          "id": 1652498,
          "postDate": "2022-01-16T17:26:11.817Z",
          "content": "<p>convert bbox format</p>\n<pre><code>def xywh2cxcywh(bbox):\n    bbox[:, 0] = bbox[:, 0] + bbox[:, 2] * 0.5\n    bbox[:, 1] = bbox[:, 1] + bbox[:, 3] * 0.5\n    return bbox\n</code></pre>\n<p>`</p>",
          "rawMarkdown": "convert bbox format\n\n```\n\ndef xywh2cxcywh(bbox):\n    bbox[:, 0] = bbox[:, 0] + bbox[:, 2] * 0.5\n    bbox[:, 1] = bbox[:, 1] + bbox[:, 3] * 0.5\n    return bbox\n\n```\n\n`",
          "votes": 2
        },
        {
          "id": 1653966,
          "postDate": "2022-01-18T02:44:32.153Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you, running my first training iteration at the moment</p>",
          "rawMarkdown": "@hengck23 Thank you, running my first training iteration at the moment"
        }
      ]
    },
    {
      "id": 1649008,
      "postDate": "2022-01-13T22:06:26.880Z",
      "content": "<p>Note: For those using rotations for augmentation, I believe you should be careful to not rotate too much since rotations make bounding boxes bigger once they are axis aligned after rotation. Having more loose bounding boxes probably hurts the model's performance (at least that's what I was experiencing).</p>",
      "rawMarkdown": "Note: For those using rotations for augmentation, I believe you should be careful to not rotate too much since rotations make bounding boxes bigger once they are axis aligned after rotation. Having more loose bounding boxes probably hurts the model's performance (at least that's what I was experiencing).",
      "votes": 4,
      "replies": [
        {
          "id": 1649070,
          "postDate": "2022-01-14T00:08:28.227Z",
          "content": "<p>create an inscribing ecllipse inside the bounding box. this is an approximate polygon annotation of the COTS object.</p>\n<p>rotate the polygon. the final bbox is computed from the rotated polygon</p>",
          "rawMarkdown": "create an inscribing ecllipse inside the bounding box. this is an approximate polygon annotation of the COTS object.\n\nrotate the polygon. the final bbox is computed from the rotated polygon",
          "votes": 3
        }
      ]
    },
    {
      "id": 1656348,
      "postDate": "2022-01-19T09:36:19.453Z",
      "content": "<p>you can test this on your validation set before probomg on your LB</p>\n<p>probing tricks?</p>\n<pre><code>f2 = 5*TP/(5*TP + 4*FN + 1*FP)\n\nLet T = total no. of true objects\n\nf2 = TP/(TP + FN - 0.2*FN + 0.2*FP)\n   = TP/(T - 0.2*FN + 0.2*FP)\n\nlet's choose a very high threshold so that we can almost sure FP=0,\nwe make a submission.\n\nf2_0 = TP/(T - 0.2*FN )         ---- eqn(1)\n\nwith the same threshold, we randomly discard 50% of the prediction,\nwe make another submission. We expect currect TP to be half of\nprevious and FN to double\n\nf1_1 = 0.5*TP/(T - 0.2*2*FN )   ---- eqn(2)\n\nyou know\nT = TP+FN  ---- eqn(3)\n\nwith eqn(1),(2),(3), we should be about to solve for T (and also TP, FN)\n\nby using different discarding different portion of predictions, we can\nget an estimate of T\n\n---\n\nlet's again choose a very high threshold so that we can almost sure FP=0.\nWe add a FP to each image e.g predict bbox = 0,0,1,1.\nThen if we submit\n\nf2_2 = TP/(T - 0.2*FN + FP )    ---- eqn(4)\nwhere FP = I =  no. of test images\n\nusing above results we can solve for I\n</code></pre>",
      "rawMarkdown": "you can test this on your validation set before probomg on your LB\n\nprobing tricks?\n\n```\n\nf2 = 5*TP/(5*TP + 4*FN + 1*FP)\n\nLet T = total no. of true objects\n\nf2 = TP/(TP + FN - 0.2*FN + 0.2*FP)\n   = TP/(T - 0.2*FN + 0.2*FP)\n\nlet's choose a very high threshold so that we can almost sure FP=0,\nwe make a submission.\n\nf2_0 = TP/(T - 0.2*FN )         ---- eqn(1)\n\nwith the same threshold, we randomly discard 50% of the prediction,\nwe make another submission. We expect currect TP to be half of\nprevious and FN to double\n\nf1_1 = 0.5*TP/(T - 0.2*2*FN )   ---- eqn(2)\n\nyou know\nT = TP+FN  ---- eqn(3)\n\nwith eqn(1),(2),(3), we should be about to solve for T (and also TP, FN)\n\nby using different discarding different portion of predictions, we can\nget an estimate of T\n\n---\n\nlet's again choose a very high threshold so that we can almost sure FP=0.\nWe add a FP to each image e.g predict bbox = 0,0,1,1.\nThen if we submit\n\nf2_2 = TP/(T - 0.2*FN + FP )    ---- eqn(4)\nwhere FP = I =  no. of test images\n\nusing above results we can solve for I\n\n```",
      "votes": 2,
      "replies": [
        {
          "id": 1656366,
          "postDate": "2022-01-19T09:51:34.323Z",
          "content": "<p>Now FPrate (FP per image) is a very relieable value you can estimate from your validation set (since you have sufficient empty image)<br>\nrecall, precision is quite relieable (Often test recall is about 10 to 20% lower than validation, test precision is quite relieable)</p>\n<p>we know the number of test image (given in the data desription page)<br>\n(if you really want to know how many there are, use a timer in your submission code.<br>\nfor each image do nothing and pause for xxx sec, then observed the time to make your submission)</p>\n<p>waht we don't know is the number of test objects.<br>\nwe make this table:</p>\n<pre><code>for\nrecall    = ...\nprecision = ...\nFPrate    = ...\n\nAt total image of I, FP = FPrate*I\n\ntotal no. of test objects |        TP | f2\n-------------------------------------------------------------------\n50                        |recall*50  | ...\n60                        |recall*60  | ...\n70                        |recall*70  | ...\n...                       |...        | ...\n...                       |...        | ...\n200                       |...        | ...\n</code></pre>\n<p>use this table to match the public LB score. this gives an estimate of total no. of test objects</p>",
          "rawMarkdown": "Now FPrate (FP per image) is a very relieable value you can estimate from your validation set (since you have sufficient empty image)\nrecall, precision is quite relieable (Often test recall is about 10 to 20% lower than validation, test precision is quite relieable)\n\nwe know the number of test image (given in the data desription page)\n(if you really want to know how many there are, use a timer in your submission code.\nfor each image do nothing and pause for xxx sec, then observed the time to make your submission)\n\nwaht we don't know is the number of test objects.\nwe make this table:\n\n```\n\nfor\nrecall    = ...\nprecision = ...\nFPrate    = ...\n\nAt total image of I, FP = FPrate*I\n\ntotal no. of test objects |        TP | f2\n-------------------------------------------------------------------\n50                        |recall*50  | ...\n60                        |recall*60  | ...\n70                        |recall*70  | ...\n...                       |...        | ...\n...                       |...        | ...\n200                       |...        | ...\n\n\n```\n\nuse this table to match the public LB score. this gives an estimate of total no. of test objects\n\n\n",
          "votes": 2
        },
        {
          "id": 1656370,
          "postDate": "2022-01-19T09:57:04.840Z",
          "content": "<p>what we really want to know from probing is:</p>\n<pre><code>no. of true objects with bbox size &lt;32 = ...\nno. of true objects with 32&lt; bbox size &lt;64 = ...\n....\n</code></pre>",
          "rawMarkdown": "what we really want to know from probing is:\n\n```\nno. of true objects with bbox size <32 = ...\nno. of true objects with 32< bbox size <64 = ...\n....\n\n```",
          "votes": 1
        },
        {
          "id": 1656392,
          "postDate": "2022-01-19T10:10:55.867Z",
          "content": "<p>one trick to improve lb is to improve iou of box prediction for small objects</p>",
          "rawMarkdown": "one trick to improve lb is to improve iou of box prediction for small objects",
          "votes": 2
        },
        {
          "id": 1657412,
          "postDate": "2022-01-20T07:00:11.287Z",
          "content": "<blockquote>\n  <p>by using different discarding different portion of predictions, we can<br>\n  get an estimate of T</p>\n</blockquote>\n<p>It’s clever idea! But I point out one thing.<br>\nThe equation 1-3 makes 3rd order simultaneous linear equations of shape A@X=0 (@ means matrix multiplication), so we have infinitely many solutions. We can know the TP/FP, TP/FN ratio from this method, but we can’t know each variables.</p>\n<p>All and all, it’s useful to know recall and precision of the public LB. </p>",
          "rawMarkdown": "> by using different discarding different portion of predictions, we can\nget an estimate of T\n\nIt’s clever idea! But I point out one thing.\nThe equation 1-3 makes 3rd order simultaneous linear equations of shape A@X=0 (@ means matrix multiplication), so we have infinitely many solutions. We can know the TP/FP, TP/FN ratio from this method, but we can’t know each variables.\n\nAll and all, it’s useful to know recall and precision of the public LB. "
        },
        {
          "id": 1657435,
          "postDate": "2022-01-20T07:12:42.677Z",
          "content": "<p>public test is 25% of data</p>\n<p>you can randomly drop 5% of the data.</p>\n<p>then you have different subset 20%. you can measure the variation due to data change.</p>",
          "rawMarkdown": "public test is 25% of data\n\nyou can randomly drop 5% of the data.\n\nthen you have different subset 20%. you can measure the variation due to data change.\n\n ",
          "votes": 1
        },
        {
          "id": 1657515,
          "postDate": "2022-01-20T08:48:40.500Z",
          "content": "<blockquote>\n  <p>you can randomly drop 5% of the data.</p>\n</blockquote>\n<p>Sorry, I don't understand. Is this mean dropping 5% of prediction?</p>\n<p>By the way, my comment is about the first comment of the thread (probing trick).<br>\nI haven't understand the second and later comments, so I'm sorry if you are confused.</p>",
          "rawMarkdown": "> you can randomly drop 5% of the data.\n\nSorry, I don't understand. Is this mean dropping 5% of prediction?\n\nBy the way, my comment is about the first comment of the thread (probing trick).\nI haven't understand the second and later comments, so I'm sorry if you are confused."
        },
        {
          "id": 1657531,
          "postDate": "2022-01-20T09:02:51.813Z",
          "content": "<p>If we changed dropping ratio (e.g. 50% -&gt; 5%), we still have equations form A'@X = 0. (Here, the valuables X are T, TP, FN, FP, and A' is their coefficient matrix).<br>\nThe right hand side is zero, and multiplying inverse matrix cannot solve this equation.<br>\nSo we can't know the value T, TP, FN, FP from this equation.</p>\n<p>But we can have the estimation of their ratio: e.g. TP/FN, TP/FP by dividing the original equations with one of the variables (e.g. FN, FP etc.).<br>\nAm I wrong?</p>",
          "rawMarkdown": "If we changed dropping ratio (e.g. 50% -> 5%), we still have equations form A'@X = 0. (Here, the valuables X are T, TP, FN, FP, and A' is their coefficient matrix).\nThe right hand side is zero, and multiplying inverse matrix cannot solve this equation.\nSo we can't know the value T, TP, FN, FP from this equation.\n\nBut we can have the estimation of their ratio: e.g. TP/FN, TP/FP by dividing the original equations with one of the variables (e.g. FN, FP etc.).\nAm I wrong?"
        }
      ]
    },
    {
      "id": 1647794,
      "postDate": "2022-01-12T23:24:36.170Z",
      "content": "<blockquote>\n  <p>Code (dirty code for illustration purpose) will be updated in the next few days.</p>\n</blockquote>\n<p>Thank you for posting this very informative discussion. This will definitely help the Kaggle community.</p>\n<p>However, there is one question I have.<br>\nA public notebook which anyone can post and get a good score will just end up in a big mess on the leaderboard.<br>\nIf you are going to share results, I would prefer it be conceptual, so that people can see what they are doing and how they are doing it.</p>",
      "rawMarkdown": "> Code (dirty code for illustration purpose) will be updated in the next few days.\n\nThank you for posting this very informative discussion. This will definitely help the Kaggle community.\n\nHowever, there is one question I have.\nA public notebook which anyone can post and get a good score will just end up in a big mess on the leaderboard.\nIf you are going to share results, I would prefer it be conceptual, so that people can see what they are doing and how they are doing it.",
      "votes": 2,
      "replies": [
        {
          "id": 1647856,
          "postDate": "2022-01-13T01:17:57.530Z",
          "content": "<p>code is public, and we have time.</p>",
          "rawMarkdown": "code is public, and we have time.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1648806,
      "postDate": "2022-01-13T17:45:50.927Z",
      "content": "<p>understanding video transect in reef survey:<br>\n<img src=\"https://biophysics.sbg.ac.at/aqaba/scan/video.jpg\" alt=\"https://biophysics.sbg.ac.at/aqaba/scan/video.jpg\"></p>\n<p>I wonder if anyone is exploring feature matching to do image mosaic/stitching?</p>",
      "rawMarkdown": "understanding video transect in reef survey:\n![https://biophysics.sbg.ac.at/aqaba/scan/video.jpg](https://biophysics.sbg.ac.at/aqaba/scan/video.jpg)\n\nI wonder if anyone is exploring feature matching to do image mosaic/stitching?",
      "votes": 1
    },
    {
      "id": 1671369,
      "postDate": "2022-02-01T12:02:58.537Z",
      "content": "<p>still stuck below LB 0.600? maybe this would help<br>\n<img src=\"https://i.ibb.co/w7jZPY0/Selection-021.png\" alt=\"https://i.ibb.co/w7jZPY0/Selection-021.png\"></p>",
      "rawMarkdown": "still stuck below LB 0.600? maybe this would help\n![https://i.ibb.co/w7jZPY0/Selection-021.png](https://i.ibb.co/w7jZPY0/Selection-021.png)\n",
      "votes": 2,
      "replies": [
        {
          "id": 1672535,
          "postDate": "2022-02-02T06:42:47.370Z",
          "content": "<p>Does this meaning higher resolution traing and testing is heavily overfitting to lb l?</p>\n<p>Fold 1 has over 2000 postive , maybe much better than lb for evaluate</p>",
          "rawMarkdown": "Does this meaning higher resolution traing and testing is heavily overfitting to lb l?\n\nFold 1 has over 2000 postive , maybe much better than lb for evaluate"
        }
      ]
    },
    {
      "id": 1648408,
      "postDate": "2022-01-13T12:10:23.503Z",
      "content": "<p>Thank you for the discussion and useful tips!<br>\nAfter considering your suggestion and the previous discussion, my LB has made a big improvement, wow😍</p>",
      "rawMarkdown": "Thank you for the discussion and useful tips!\nAfter considering your suggestion and the previous discussion, my LB has made a big improvement, wow😍",
      "votes": 2
    },
    {
      "id": 1647869,
      "postDate": "2022-01-13T01:34:16.453Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you always make teach new way to approach the problem. Happy to see you in this competition. It will be more easy and fun to join the competitions where you participate.</p>",
      "rawMarkdown": "@hengck23 you always make teach new way to approach the problem. Happy to see you in this competition. It will be more easy and fun to join the competitions where you participate.",
      "votes": 2
    },
    {
      "id": 1662541,
      "postDate": "2022-01-24T11:23:04.090Z",
      "content": "<p>Thank your for your discussion.It is beneficial for fresh kaggler.</p>",
      "rawMarkdown": "Thank your for your discussion.It is beneficial for fresh kaggler."
    },
    {
      "id": 1659085,
      "postDate": "2022-01-21T14:42:31.943Z",
      "content": "<p>great jobs, tkx</p>",
      "rawMarkdown": "great jobs, tkx"
    },
    {
      "id": 1649135,
      "postDate": "2022-01-14T01:53:41.547Z",
      "content": "<p>Thanks for your disscussion.Happy to see you again in this competition, I can learn a lot from you every time.</p>",
      "rawMarkdown": "Thanks for your disscussion.Happy to see you again in this competition, I can learn a lot from you every time."
    },
    {
      "id": 1648297,
      "postDate": "2022-01-13T10:14:56.523Z",
      "content": "<p>one quick  question:<br>\nduring training use enlarged test_size, or just do so when inference?</p>\n<p>slightly enlarge, I got LB improved a little bit, but when test_size enlarge too much, LB becomes worse.</p>\n<p>what kind of models suggesting 1.5x or 2.0X?    I am using yolox_L and yolox_X.</p>",
      "rawMarkdown": "one quick  question:\nduring training use enlarged test_size, or just do so when inference?\n\nslightly enlarge, I got LB improved a little bit, but when test_size enlarge too much, LB becomes worse.\n\nwhat kind of models suggesting 1.5x or 2.0X?    I am using yolox_L and yolox_X.",
      "replies": [
        {
          "id": 1648301,
          "postDate": "2022-01-13T10:18:15.247Z",
          "content": "<p>the better results with the large image just showed that </p>\n<ul>\n<li>there are likely to be smaller cots in the test</li>\n<li>large image may contain more information</li>\n</ul>\n<p>you would have to try training with larger images or modify your network model to keep high-resolution details. It would be by trial and error to see the results.</p>\n<p>(note: detecting smaller objects will definitely lead to more FP)</p>",
          "rawMarkdown": "the better results with the large image just showed that \n- there are likely to be smaller cots in the test\n- large image may contain more information\n\nyou would have to try training with larger images or modify your network model to keep high-resolution details. It would be by trial and error to see the results.\n\n(note: detecting smaller objects will definitely lead to more FP)",
          "votes": 2
        },
        {
          "id": 1648380,
          "postDate": "2022-01-13T11:43:06.673Z",
          "content": "<p>thx for reply.  It is difficult for try and error with limited GPU resource.</p>",
          "rawMarkdown": "thx for reply.  It is difficult for try and error with limited GPU resource."
        }
      ]
    },
    {
      "id": 1647983,
      "postDate": "2022-01-13T03:21:01.683Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> great insight! Could kindly teach how to ensemble multiple fold model into one? I never succeed in this after lots of experiments.</p>",
      "rawMarkdown": "Hi @hengck23 great insight! Could kindly teach how to ensemble multiple fold model into one? I never succeed in this after lots of experiments.",
      "replies": [
        {
          "id": 1648239,
          "postDate": "2022-01-13T09:20:27.483Z",
          "content": "<p>Assuming there is no code bug, it depends on the TP rate and FP rate of your individual models.</p>",
          "rawMarkdown": "Assuming there is no code bug, it depends on the TP rate and FP rate of your individual models.",
          "votes": 2
        },
        {
          "id": 1648258,
          "postDate": "2022-01-13T09:37:48.567Z",
          "content": "<p>Thanks for the tip.</p>",
          "rawMarkdown": "Thanks for the tip."
        }
      ]
    },
    {
      "id": 1647408,
      "postDate": "2022-01-12T15:25:14.627Z",
      "content": "<p>What detection model are you using?<br>\nand about the classification how are you doing that?</p>",
      "rawMarkdown": "What detection model are you using?\nand about the classification how are you doing that?"
    },
    {
      "id": 1647406,
      "postDate": "2022-01-12T15:23:51.163Z",
      "content": "<p>Thanks for sharing,<br>\nI have 90% Precision and 65% recall but problem with post processing, Is there any tip?  <br>\nBy the way, did you use all data  (labels, no labels) ?<br>\nThanks</p>",
      "rawMarkdown": "Thanks for sharing,\nI have 90% Precision and 65% recall but problem with post processing, Is there any tip?  \nBy the way, did you use all data  (labels, no labels) ?\nThanks",
      "replies": [
        {
          "id": 1648580,
          "postDate": "2022-01-13T14:22:49.190Z",
          "content": "<p>i think you mean no annotations cause there's no \"unlabelled\" data provided. <br>\nAlso this P, R gives about 0.83 on F2 score that's a great CV! I think if your evaluation methods are right this kind of CV should give decent LB results even without a solid post processing script</p>",
          "rawMarkdown": "i think you mean no annotations cause there's no \"unlabelled\" data provided. \nAlso this P, R gives about 0.83 on F2 score that's a great CV! I think if your evaluation methods are right this kind of CV should give decent LB results even without a solid post processing script",
          "votes": 1
        }
      ]
    },
    {
      "id": 1647396,
      "postDate": "2022-01-12T15:06:52.260Z",
      "content": "<p>Thank you for sharing the finding <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I have a question regarding this statement: <br>\n<code>one easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector</code></p>\n<p>In general, do we need to test with smaller image size as well for both image classifier and object detector or do image classifiers and object detectors generally perform pretty well with smaller sized input?</p>",
      "rawMarkdown": "Thank you for sharing the finding @hengck23 ! I have a question regarding this statement: \n`one easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector`\n\nIn general, do we need to test with smaller image size as well for both image classifier and object detector or do image classifiers and object detectors generally perform pretty well with smaller sized input?"
    },
    {
      "id": 1647384,
      "postDate": "2022-01-12T14:53:01.430Z",
      "content": "<p>Thank you for sharing and hope to be helpful！</p>",
      "rawMarkdown": "Thank you for sharing and hope to be helpful！"
    },
    {
      "id": 1647361,
      "postDate": "2022-01-12T14:26:52.530Z",
      "content": "<p>Great work. I've been loving all your posts and experiments! Keep it up!</p>",
      "rawMarkdown": "Great work. I've been loving all your posts and experiments! Keep it up!"
    },
    {
      "id": 1647898,
      "postDate": "2022-01-13T01:51:43.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1647375,
      "postDate": "2022-01-12T14:34:40.893Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1657172,
      "postDate": "2022-01-20T01:41:20.770Z",
      "content": "<p>thank you for your sharing</p>",
      "rawMarkdown": "thank you for your sharing"
    },
    {
      "id": 1650673,
      "postDate": "2022-01-15T08:08:50.710Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 1650379,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-15T02:29:16.157000",
      "content": "<p>final solution would be someting like this:<br>\n<img src=\"https://i.ibb.co/FH9Lrzy/Selection-999-723.png\" alt=\"https://i.ibb.co/FH9Lrzy/Selection-999-723.png\"></p>",
      "votes": 13,
      "replies": [
        {
          "id": 1650413,
          "author_name": "Max van Dijck",
          "author_url": "",
          "post_date": "2022-01-15T03:19:17.833000",
          "content": "<p>Yeah, it seems upscaling the image, despite being compute expensive, is resulting in a large leaderboard jump right now. I think your approach of cropping regions of interest and then upscaling is a good solution for this performance trade off. Also as a side note, I've only had dissapointment with two-stage detectors compared to YOLO computation cost wise</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1650845,
          "author_name": "shigengtian",
          "author_url": "",
          "post_date": "2022-01-15T11:20:49.967000",
          "content": "<p>it takes time. I will try it and add some new ideas.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1650871,
          "author_name": "Robson",
          "author_url": "",
          "post_date": "2022-01-15T11:34:18.250000",
          "content": "<p>After the end of the competition I will still need 1 more month to test all the hypotheses you raised. As always your posts are pure gold. Congratulations</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1648383,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-13T11:48:30.277000",
      "content": "<p>here is a list of good papers you can read. Out use case is similar to object detection I'm monocular video for unmanned aerial vehicle (UAV), particularly underwater drone.</p>\n<ol>\n<li>TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. <br>\n<a href=\"https://arxiv.org/abs/2108.11539\" target=\"_blank\">https://arxiv.org/abs/2108.11539</a><br>\n5th in On VisDrone Challenge 2021<br>\n(1st to 4th place: DBNet, SOLOer, Swin-T, DPNetV4)</li>\n</ol>\n<p>ECAP-YOLO: Efficient Channel Attention Pyramid YOLO for Small Object Detection in Aerial Imag<br>\nanother modification</p>\n<p>remotesensing-13-04851-v2.pdf</p>\n<ol>\n<li>Bag of Freebies for Training Object Detection Neural Networks<br>\n<a href=\"https://arxiv.org/abs/1902.04103\" target=\"_blank\">https://arxiv.org/abs/1902.04103</a></li>\n</ol>\n<hr>\n<p>PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices<br>\n<a href=\"https://arxiv.org/abs/2111.00902\" target=\"_blank\">https://arxiv.org/abs/2111.00902</a></p>\n<p>A ConvNet for the 2020s<br>\n<a href=\"https://arxiv.org/abs/2201.03545\" target=\"_blank\">https://arxiv.org/abs/2201.03545</a></p>\n<p>Augmenting Convolutional networks with attention-based aggregation<br>\n<a href=\"https://arxiv.org/pdf/2112.13692v1\" target=\"_blank\">https://arxiv.org/pdf/2112.13692v1</a><br>\n\", as our approach is not pyramidal, we only use the final output of our network in Mask R-CNN\"</p>\n<hr>\n<p>slam+object detection as an alternative to detection+tracking<br>\n<a href=\"https://www.youtube.com/watch?v=m6sStUk3UVk\" target=\"_blank\">https://www.youtube.com/watch?v=m6sStUk3UVk</a></p>\n<hr>\n<p>MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks<br>\n<a href=\"https://arxiv.org/pdf/2009.08453\" target=\"_blank\">https://arxiv.org/pdf/2009.08453</a></p>\n<hr>\n<p>some zooming based network:<br>\nDynamic Zoom-in Network for Fast Object Detection in Large Images<br>\nDense and Small Object Detection in UAV Vision based on Cascade Network</p>\n<p>maybe we can add coordinate feature (and frame number feature) to make the model region aware.<br>\n<img src=\"https://images.deepai.org/converted-papers/1711.05187/x1.png\" alt=\"https://images.deepai.org/converted-papers/1711.05187/x1.png\"></p>\n<hr>\n<p>better moasic. we have single class with different subsclass<br>\nImage-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection<br>\n<a href=\"https://arxiv.org/pdf/2104.05702\" target=\"_blank\">https://arxiv.org/pdf/2104.05702</a></p>\n<hr>\n<p>better resolution</p>\n<p>Fixing the train-test resolution discrepancy<br>\n<img src=\"https://github.com/facebookresearch/FixRes/raw/main/image/image2.png\" alt=\"https://github.com/facebookresearch/FixRes/raw/main/image/image2.png\"></p>\n<p>MIX &amp; MATCH: TRAINING CONVNETS WITH MIXED IMAGE SIZES FOR IMPROVED ACCURACY, SPEED AND<br>\nSCALE RESILIENCY<br>\n\"For instance, we receive a 76.43% top-1 accuracy using ResNet50 with an image size of 160,\"</p>\n<p>Learning to Resize Images for Computer Vision Tasks<br>\n<a href=\"https://arxiv.org/pdf/2003.08237.pdf\" target=\"_blank\">https://arxiv.org/pdf/2003.08237.pdf</a></p>\n<p>temporal (and spatial) frame interpolation to create more training data?</p>\n<p>Perceptual Generative Adversarial Networks for Small Object Detection</p>\n<hr>\n<p>one can treat extreme augmentation as unlabelled data<br>\nBeyond Self-Supervision: A Simple Yet Effective Network Distillation<br>\nAlternative to Improve Backbones</p>\n<p>\"Ours(ResNet50-D+ResFix) 5.2M 84.0 (top1)\"</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 1648132,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-13T07:25:13.623000",
      "content": "<p>experiment chart:<br>\n<img src=\"https://i.ibb.co/F7t2fjk/Selection-999-712.png\" alt=\"https://i.ibb.co/F7t2fjk/Selection-999-712.png\"></p>",
      "votes": 9,
      "replies": [
        {
          "id": 1648147,
          "author_name": "Max van Dijck",
          "author_url": "",
          "post_date": "2022-01-13T07:49:42.907000",
          "content": "<p>Would you retrain a final model using all videos as training data for final submission. I assume this model would have better generalization?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1648245,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-13T09:24:01.917000",
          "content": "<ul>\n<li><p>instead of spending time to determine a better split, my strategy is to create my train samples and probe the LB to see what kind of data to generate</p></li>\n<li><p>my next step is to use multi-frame for detection (not really a tracking problem as we don't have to associate track ID)</p></li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1648304,
          "author_name": "omallo",
          "author_url": "",
          "post_date": "2022-01-13T10:23:28.090000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , thanks a lot for this useful analysis.</p>\n<p>May I ask how you deal with unlabeled images, i.e. do you include (part of) them in the training and are they included for the calculation of your CV scores above?</p>\n<p>I'm currently including a small amount of unlabeled images in my training (~3% of the labeled training data) and I include all the unlabeled images for CV scoring. I was wondering how you and maybe others deal with them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1648328,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-13T10:55:17.047000",
          "content": "<p><a href=\"https://www.kaggle.com/omallo\" target=\"_blank\">@omallo</a> your approach is ok and smililar to mine.<br>\nFor  N images in each batch or (for very N images in M batch iterations), just make sure :</p>\n<ul>\n<li>there are x% of empty image</li>\n<li>there are (1-x)% of non-empty image. And there are on average som eN numbers of truth bboxes, of various sizes and appearances</li>\n</ul>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1649188,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2022-01-14T03:54:19.123000",
          "content": "<p>Seem higher resolution is better</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1650532,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2022-01-15T05:20:22.137000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> do you mean to have better cv with <strong>aspect ratio</strong> or <strong>area</strong> play an important role</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1647410,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2022-01-12T15:27:45.230000",
      "content": "<p>Wow, leaderboard inflation is coming!</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1648264,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-13T09:45:59.667000",
      "content": "<p>dirty code is up:<br>\n<a href=\"https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\" target=\"_blank\">https://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing</a><br>\n2022-02-13.tar.gz</p>\n<p>note:</p>\n<ul>\n<li>i no longer work on this code as i setup an updated version at my side</li>\n<li>feel free to use this code in any way, e.g. modify it and put it in your notebook, etc</li>\n<li>code is dirty, may have missing unimportant functions, etc (I copied the files from a larger project). But it should be easy to figure out the missing parts.</li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 1650536,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2022-01-15T05:23:06.497000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> will you add your code link in topic. So, everyone learn from it. Your knowledge sharing helped to learn new things faster. Thanks for sharing your experiments save time and energy.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1650544,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2022-01-15T05:32:27.173000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> one question regarding data, in ur data set is cleaned version, did u resolved the missing COTS annotations in the sequence frame based or any other approach ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1652470,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2022-01-16T16:51:15.127000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> it's missing my_bbox.py  <br>\nWhat does this function do : xywh2cxcywh ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1652498,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-16T17:26:11.817000",
          "content": "<p>convert bbox format</p>\n<pre><code>def xywh2cxcywh(bbox):\n    bbox[:, 0] = bbox[:, 0] + bbox[:, 2] * 0.5\n    bbox[:, 1] = bbox[:, 1] + bbox[:, 3] * 0.5\n    return bbox\n</code></pre>\n<p>`</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1653966,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2022-01-18T02:44:32.153000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you, running my first training iteration at the moment</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1649008,
      "author_name": "omallo",
      "author_url": "",
      "post_date": "2022-01-13T22:06:26.880000",
      "content": "<p>Note: For those using rotations for augmentation, I believe you should be careful to not rotate too much since rotations make bounding boxes bigger once they are axis aligned after rotation. Having more loose bounding boxes probably hurts the model's performance (at least that's what I was experiencing).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1649070,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-14T00:08:28.227000",
          "content": "<p>create an inscribing ecllipse inside the bounding box. this is an approximate polygon annotation of the COTS object.</p>\n<p>rotate the polygon. the final bbox is computed from the rotated polygon</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1656348,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-19T09:36:19.453000",
      "content": "<p>you can test this on your validation set before probomg on your LB</p>\n<p>probing tricks?</p>\n<pre><code>f2 = 5*TP/(5*TP + 4*FN + 1*FP)\n\nLet T = total no. of true objects\n\nf2 = TP/(TP + FN - 0.2*FN + 0.2*FP)\n   = TP/(T - 0.2*FN + 0.2*FP)\n\nlet's choose a very high threshold so that we can almost sure FP=0,\nwe make a submission.\n\nf2_0 = TP/(T - 0.2*FN )         ---- eqn(1)\n\nwith the same threshold, we randomly discard 50% of the prediction,\nwe make another submission. We expect currect TP to be half of\nprevious and FN to double\n\nf1_1 = 0.5*TP/(T - 0.2*2*FN )   ---- eqn(2)\n\nyou know\nT = TP+FN  ---- eqn(3)\n\nwith eqn(1),(2),(3), we should be about to solve for T (and also TP, FN)\n\nby using different discarding different portion of predictions, we can\nget an estimate of T\n\n---\n\nlet's again choose a very high threshold so that we can almost sure FP=0.\nWe add a FP to each image e.g predict bbox = 0,0,1,1.\nThen if we submit\n\nf2_2 = TP/(T - 0.2*FN + FP )    ---- eqn(4)\nwhere FP = I =  no. of test images\n\nusing above results we can solve for I\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1656366,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-19T09:51:34.323000",
          "content": "<p>Now FPrate (FP per image) is a very relieable value you can estimate from your validation set (since you have sufficient empty image)<br>\nrecall, precision is quite relieable (Often test recall is about 10 to 20% lower than validation, test precision is quite relieable)</p>\n<p>we know the number of test image (given in the data desription page)<br>\n(if you really want to know how many there are, use a timer in your submission code.<br>\nfor each image do nothing and pause for xxx sec, then observed the time to make your submission)</p>\n<p>waht we don't know is the number of test objects.<br>\nwe make this table:</p>\n<pre><code>for\nrecall    = ...\nprecision = ...\nFPrate    = ...\n\nAt total image of I, FP = FPrate*I\n\ntotal no. of test objects |        TP | f2\n-------------------------------------------------------------------\n50                        |recall*50  | ...\n60                        |recall*60  | ...\n70                        |recall*70  | ...\n...                       |...        | ...\n...                       |...        | ...\n200                       |...        | ...\n</code></pre>\n<p>use this table to match the public LB score. this gives an estimate of total no. of test objects</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1656370,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-19T09:57:04.840000",
          "content": "<p>what we really want to know from probing is:</p>\n<pre><code>no. of true objects with bbox size &lt;32 = ...\nno. of true objects with 32&lt; bbox size &lt;64 = ...\n....\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1656392,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-19T10:10:55.867000",
          "content": "<p>one trick to improve lb is to improve iou of box prediction for small objects</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1657412,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-20T07:00:11.287000",
          "content": "<blockquote>\n  <p>by using different discarding different portion of predictions, we can<br>\n  get an estimate of T</p>\n</blockquote>\n<p>It’s clever idea! But I point out one thing.<br>\nThe equation 1-3 makes 3rd order simultaneous linear equations of shape A@X=0 (@ means matrix multiplication), so we have infinitely many solutions. We can know the TP/FP, TP/FN ratio from this method, but we can’t know each variables.</p>\n<p>All and all, it’s useful to know recall and precision of the public LB. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1657435,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-20T07:12:42.677000",
          "content": "<p>public test is 25% of data</p>\n<p>you can randomly drop 5% of the data.</p>\n<p>then you have different subset 20%. you can measure the variation due to data change.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1657515,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-20T08:48:40.500000",
          "content": "<blockquote>\n  <p>you can randomly drop 5% of the data.</p>\n</blockquote>\n<p>Sorry, I don't understand. Is this mean dropping 5% of prediction?</p>\n<p>By the way, my comment is about the first comment of the thread (probing trick).<br>\nI haven't understand the second and later comments, so I'm sorry if you are confused.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1657531,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-20T09:02:51.813000",
          "content": "<p>If we changed dropping ratio (e.g. 50% -&gt; 5%), we still have equations form A'@X = 0. (Here, the valuables X are T, TP, FN, FP, and A' is their coefficient matrix).<br>\nThe right hand side is zero, and multiplying inverse matrix cannot solve this equation.<br>\nSo we can't know the value T, TP, FN, FP from this equation.</p>\n<p>But we can have the estimation of their ratio: e.g. TP/FN, TP/FP by dividing the original equations with one of the variables (e.g. FN, FP etc.).<br>\nAm I wrong?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1647794,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-12T23:24:36.170000",
      "content": "<blockquote>\n  <p>Code (dirty code for illustration purpose) will be updated in the next few days.</p>\n</blockquote>\n<p>Thank you for posting this very informative discussion. This will definitely help the Kaggle community.</p>\n<p>However, there is one question I have.<br>\nA public notebook which anyone can post and get a good score will just end up in a big mess on the leaderboard.<br>\nIf you are going to share results, I would prefer it be conceptual, so that people can see what they are doing and how they are doing it.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1647856,
          "author_name": "Tian",
          "author_url": "",
          "post_date": "2022-01-13T01:17:57.530000",
          "content": "<p>code is public, and we have time.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1648806,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-13T17:45:50.927000",
      "content": "<p>understanding video transect in reef survey:<br>\n<img src=\"https://biophysics.sbg.ac.at/aqaba/scan/video.jpg\" alt=\"https://biophysics.sbg.ac.at/aqaba/scan/video.jpg\"></p>\n<p>I wonder if anyone is exploring feature matching to do image mosaic/stitching?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1671369,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-02-01T12:02:58.537000",
      "content": "<p>still stuck below LB 0.600? maybe this would help<br>\n<img src=\"https://i.ibb.co/w7jZPY0/Selection-021.png\" alt=\"https://i.ibb.co/w7jZPY0/Selection-021.png\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1672535,
          "author_name": "Drzhuzhe",
          "author_url": "",
          "post_date": "2022-02-02T06:42:47.370000",
          "content": "<p>Does this meaning higher resolution traing and testing is heavily overfitting to lb l?</p>\n<p>Fold 1 has over 2000 postive , maybe much better than lb for evaluate</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1648408,
      "author_name": "ShengzheLiu",
      "author_url": "",
      "post_date": "2022-01-13T12:10:23.503000",
      "content": "<p>Thank you for the discussion and useful tips!<br>\nAfter considering your suggestion and the previous discussion, my LB has made a big improvement, wow😍</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1647869,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2022-01-13T01:34:16.453000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you always make teach new way to approach the problem. Happy to see you in this competition. It will be more easy and fun to join the competitions where you participate.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1662541,
      "author_name": "Good Moon",
      "author_url": "",
      "post_date": "2022-01-24T11:23:04.090000",
      "content": "<p>Thank your for your discussion.It is beneficial for fresh kaggler.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1659085,
      "author_name": "Fred Shi",
      "author_url": "",
      "post_date": "2022-01-21T14:42:31.943000",
      "content": "<p>great jobs, tkx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1649135,
      "author_name": "Roc",
      "author_url": "",
      "post_date": "2022-01-14T01:53:41.547000",
      "content": "<p>Thanks for your disscussion.Happy to see you again in this competition, I can learn a lot from you every time.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1648297,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-01-13T10:14:56.523000",
      "content": "<p>one quick  question:<br>\nduring training use enlarged test_size, or just do so when inference?</p>\n<p>slightly enlarge, I got LB improved a little bit, but when test_size enlarge too much, LB becomes worse.</p>\n<p>what kind of models suggesting 1.5x or 2.0X?    I am using yolox_L and yolox_X.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1648301,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-13T10:18:15.247000",
          "content": "<p>the better results with the large image just showed that </p>\n<ul>\n<li>there are likely to be smaller cots in the test</li>\n<li>large image may contain more information</li>\n</ul>\n<p>you would have to try training with larger images or modify your network model to keep high-resolution details. It would be by trial and error to see the results.</p>\n<p>(note: detecting smaller objects will definitely lead to more FP)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1648380,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-01-13T11:43:06.673000",
          "content": "<p>thx for reply.  It is difficult for try and error with limited GPU resource.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1647983,
      "author_name": "DeepInvolution",
      "author_url": "",
      "post_date": "2022-01-13T03:21:01.683000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> great insight! Could kindly teach how to ensemble multiple fold model into one? I never succeed in this after lots of experiments.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1648239,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-13T09:20:27.483000",
          "content": "<p>Assuming there is no code bug, it depends on the TP rate and FP rate of your individual models.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1648258,
          "author_name": "DeepInvolution",
          "author_url": "",
          "post_date": "2022-01-13T09:37:48.567000",
          "content": "<p>Thanks for the tip.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1647408,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-01-12T15:25:14.627000",
      "content": "<p>What detection model are you using?<br>\nand about the classification how are you doing that?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1647406,
      "author_name": "Ichimaru Gin",
      "author_url": "",
      "post_date": "2022-01-12T15:23:51.163000",
      "content": "<p>Thanks for sharing,<br>\nI have 90% Precision and 65% recall but problem with post processing, Is there any tip?  <br>\nBy the way, did you use all data  (labels, no labels) ?<br>\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1648580,
          "author_name": "Hannah B",
          "author_url": "",
          "post_date": "2022-01-13T14:22:49.190000",
          "content": "<p>i think you mean no annotations cause there's no \"unlabelled\" data provided. <br>\nAlso this P, R gives about 0.83 on F2 score that's a great CV! I think if your evaluation methods are right this kind of CV should give decent LB results even without a solid post processing script</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1647396,
      "author_name": "hide on bread",
      "author_url": "",
      "post_date": "2022-01-12T15:06:52.260000",
      "content": "<p>Thank you for sharing the finding <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I have a question regarding this statement: <br>\n<code>one easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector</code></p>\n<p>In general, do we need to test with smaller image size as well for both image classifier and object detector or do image classifiers and object detectors generally perform pretty well with smaller sized input?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1647384,
      "author_name": "AIhouhaoxiong",
      "author_url": "",
      "post_date": "2022-01-12T14:53:01.430000",
      "content": "<p>Thank you for sharing and hope to be helpful！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1647361,
      "author_name": "outwrest",
      "author_url": "",
      "post_date": "2022-01-12T14:26:52.530000",
      "content": "<p>Great work. I've been loving all your posts and experiments! Keep it up!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1647898,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-13T01:51:43.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1647375,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-12T14:34:40.893000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1657172,
      "author_name": "Gong",
      "author_url": "",
      "post_date": "2022-01-20T01:41:20.770000",
      "content": "<p>thank you for your sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1650673,
      "author_name": "Akhil Sachar",
      "author_url": "",
      "post_date": "2022-01-15T08:08:50.710000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1647337": "After experimenting for a week, i am ready to share my intermediate solution.\nFollow this thread for the continuous update!\n\n\n~~Code (dirty code for illustration purpose) will be updated in the next few days.~~\nrefer to google drive:\nhttps://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\n2022-02-13.tar.gz\n\n----\nsummary of approach\n\nWith designing augmentation for object detection, there is two main consideration:\n1. target object appearance and pose\n2. object layout in the scene\n\nfor 1. it is easy to test.  the median size of the target COTS object is about 32 (ranges from 20 to 128). Hence we can actually use 32x32, 64x64, 128x128 image classifier. Recall is important in the f2 metric. you have to have recall of about 60% to 80%.  you can keep precision to about 80% to 90% and FPrate to 0.20 per image.\n\nyou can use feature of image classifier and make a tsne map to see the distribution of the target objects. which is the rare subclass? (actually feature from object detector also will work)\n\n\nNow for 2. if you have a very good image classifier (i.e. good representative train data/augmentation),  the object detection model should work well too (assuming you have no problem to post processing like NMS). if you do not do well here, then it  could be:\n- wrong anchor box size\n- the object size in the validation set is different from the train\n- sampling problem ( object vs background imbalance ratio)\n\nIt is all about scaling and rotation ( convolution is invariant to translation, but not scale and rotation). This is why it fails when you move from classifier to object detection.\n\nwhen you do augmentation,  so bbox gets too small and ends up partially exceeding the image boundary. you should filter those if they hurt results.\n\nalso note that the number of bbox per image is low. it is likely that after affine augmentation (rotate, enlarge, etc), you can end up with any empty image without bbox.\n\n(do note that the whole idea about augmentation is to create good train samples and also some good noise so that the model will not be overconfident on the prediction. noise balancing is difficult)\n\none easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector\n\n\n\nfor splitting, it is sufficient to use train=video0,1, valid=video2 to achieve lb 0.560. But for experimenting with augmentation, you have to try all 3 folds using video0,1,2 as validation.\n\n----\n\nI am grateful to HP for providing a Z by HP workstation for this competition. Equipped with 2x A600 Nvidia cards, I can complete my experiments in a short time and share the results and findings with the kaggle community. ",
    "1650379": "final solution would be someting like this:\n![https://i.ibb.co/FH9Lrzy/Selection-999-723.png](https://i.ibb.co/FH9Lrzy/Selection-999-723.png)",
    "1648383": "here is a list of good papers you can read. Out use case is similar to object detection I'm monocular video for unmanned aerial vehicle (UAV), particularly underwater drone.\n\n\n1. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. \nhttps://arxiv.org/abs/2108.11539\n5th in On VisDrone Challenge 2021\n(1st to 4th place: DBNet, SOLOer, Swin-T, DPNetV4)\n\nECAP-YOLO: Efficient Channel Attention Pyramid YOLO for Small Object Detection in Aerial Imag\nanother modification\n\nremotesensing-13-04851-v2.pdf\n\n2. Bag of Freebies for Training Object Detection Neural Networks\n https://arxiv.org/abs/1902.04103\n\n---\n\nPP-PicoDet: A Better Real-Time Object Detector on Mobile Devices\nhttps://arxiv.org/abs/2111.00902\n\nA ConvNet for the 2020s\nhttps://arxiv.org/abs/2201.03545\n\nAugmenting Convolutional networks with attention-based aggregation\nhttps://arxiv.org/pdf/2112.13692v1\n\", as our approach is not pyramidal, we only use the final output of our network in Mask R-CNN\"\n\n---\nslam+object detection as an alternative to detection+tracking\nhttps://www.youtube.com/watch?v=m6sStUk3UVk\n\n---\nMEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks\nhttps://arxiv.org/pdf/2009.08453\n\n---\n\nsome zooming based network:\nDynamic Zoom-in Network for Fast Object Detection in Large Images\nDense and Small Object Detection in UAV Vision based on Cascade Network\n\nmaybe we can add coordinate feature (and frame number feature) to make the model region aware.\n![https://images.deepai.org/converted-papers/1711.05187/x1.png](https://images.deepai.org/converted-papers/1711.05187/x1.png)\n\n---\n\nbetter moasic. we have single class with different subsclass\nImage-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection\nhttps://arxiv.org/pdf/2104.05702\n\n---\n\nbetter resolution\n\nFixing the train-test resolution discrepancy\n![https://github.com/facebookresearch/FixRes/raw/main/image/image2.png](https://github.com/facebookresearch/FixRes/raw/main/image/image2.png)\n\nMIX & MATCH: TRAINING CONVNETS WITH MIXED IMAGE SIZES FOR IMPROVED ACCURACY, SPEED AND\nSCALE RESILIENCY\n\"For instance, we receive a 76.43% top-1 accuracy using ResNet50 with an image size of 160,\"\n\nLearning to Resize Images for Computer Vision Tasks\nhttps://arxiv.org/pdf/2003.08237.pdf\n\ntemporal (and spatial) frame interpolation to create more training data?\n\nPerceptual Generative Adversarial Networks for Small Object Detection\n\n---\n\none can treat extreme augmentation as unlabelled data\nBeyond Self-Supervision: A Simple Yet Effective Network Distillation\nAlternative to Improve Backbones\n\n\"Ours(ResNet50-D+ResFix) 5.2M 84.0 (top1)\"",
    "1648132": "experiment chart:\n![https://i.ibb.co/F7t2fjk/Selection-999-712.png](https://i.ibb.co/F7t2fjk/Selection-999-712.png)",
    "1647410": "Wow, leaderboard inflation is coming!",
    "1648264": "dirty code is up:\nhttps://drive.google.com/drive/folders/1GX4maMF4ZwQ77mx8NNnVpZVNOCPvBCPk?usp=sharing\n2022-02-13.tar.gz\n\nnote:\n- i no longer work on this code as i setup an updated version at my side\n- feel free to use this code in any way, e.g. modify it and put it in your notebook, etc\n- code is dirty, may have missing unimportant functions, etc (I copied the files from a larger project). But it should be easy to figure out the missing parts.",
    "1649008": "Note: For those using rotations for augmentation, I believe you should be careful to not rotate too much since rotations make bounding boxes bigger once they are axis aligned after rotation. Having more loose bounding boxes probably hurts the model's performance (at least that's what I was experiencing).",
    "1656348": "you can test this on your validation set before probomg on your LB\n\nprobing tricks?\n\n```\n\nf2 = 5*TP/(5*TP + 4*FN + 1*FP)\n\nLet T = total no. of true objects\n\nf2 = TP/(TP + FN - 0.2*FN + 0.2*FP)\n   = TP/(T - 0.2*FN + 0.2*FP)\n\nlet's choose a very high threshold so that we can almost sure FP=0,\nwe make a submission.\n\nf2_0 = TP/(T - 0.2*FN )         ---- eqn(1)\n\nwith the same threshold, we randomly discard 50% of the prediction,\nwe make another submission. We expect currect TP to be half of\nprevious and FN to double\n\nf1_1 = 0.5*TP/(T - 0.2*2*FN )   ---- eqn(2)\n\nyou know\nT = TP+FN  ---- eqn(3)\n\nwith eqn(1),(2),(3), we should be about to solve for T (and also TP, FN)\n\nby using different discarding different portion of predictions, we can\nget an estimate of T\n\n---\n\nlet's again choose a very high threshold so that we can almost sure FP=0.\nWe add a FP to each image e.g predict bbox = 0,0,1,1.\nThen if we submit\n\nf2_2 = TP/(T - 0.2*FN + FP )    ---- eqn(4)\nwhere FP = I =  no. of test images\n\nusing above results we can solve for I\n\n```",
    "1647794": "> Code (dirty code for illustration purpose) will be updated in the next few days.\n\nThank you for posting this very informative discussion. This will definitely help the Kaggle community.\n\nHowever, there is one question I have.\nA public notebook which anyone can post and get a good score will just end up in a big mess on the leaderboard.\nIf you are going to share results, I would prefer it be conceptual, so that people can see what they are doing and how they are doing it.",
    "1648806": "understanding video transect in reef survey:\n![https://biophysics.sbg.ac.at/aqaba/scan/video.jpg](https://biophysics.sbg.ac.at/aqaba/scan/video.jpg)\n\nI wonder if anyone is exploring feature matching to do image mosaic/stitching?",
    "1671369": "still stuck below LB 0.600? maybe this would help\n![https://i.ibb.co/w7jZPY0/Selection-021.png](https://i.ibb.co/w7jZPY0/Selection-021.png)\n",
    "1648408": "Thank you for the discussion and useful tips!\nAfter considering your suggestion and the previous discussion, my LB has made a big improvement, wow😍",
    "1647869": "@hengck23 you always make teach new way to approach the problem. Happy to see you in this competition. It will be more easy and fun to join the competitions where you participate.",
    "1662541": "Thank your for your discussion.It is beneficial for fresh kaggler.",
    "1659085": "great jobs, tkx",
    "1649135": "Thanks for your disscussion.Happy to see you again in this competition, I can learn a lot from you every time.",
    "1648297": "one quick  question:\nduring training use enlarged test_size, or just do so when inference?\n\nslightly enlarge, I got LB improved a little bit, but when test_size enlarge too much, LB becomes worse.\n\nwhat kind of models suggesting 1.5x or 2.0X?    I am using yolox_L and yolox_X.",
    "1647983": "Hi @hengck23 great insight! Could kindly teach how to ensemble multiple fold model into one? I never succeed in this after lots of experiments.",
    "1647408": "What detection model are you using?\nand about the classification how are you doing that?",
    "1647406": "Thanks for sharing,\nI have 90% Precision and 65% recall but problem with post processing, Is there any tip?  \nBy the way, did you use all data  (labels, no labels) ?\nThanks",
    "1647396": "Thank you for sharing the finding @hengck23 ! I have a question regarding this statement: \n`one easier way to test size problem is to train in size say 1280 and test with an enlarged image(1.25x, 1.50x, 2.00x). it works for both image classifier and object detector`\n\nIn general, do we need to test with smaller image size as well for both image classifier and object detector or do image classifiers and object detectors generally perform pretty well with smaller sized input?",
    "1647384": "Thank you for sharing and hope to be helpful！",
    "1647361": "Great work. I've been loving all your posts and experiments! Keep it up!",
    "1647898": "",
    "1647375": "",
    "1657172": "thank you for your sharing",
    "1650673": "Thanks for sharing."
  }
}