{
  "id": 298427,
  "title": "Why higher mAP leads to a significantly lower public score?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/298427",
  "author_name": "Early Bird",
  "post_date": "2022-01-02T20:43:23.816000",
  "votes": 26,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>This is doing my head in… not sure if anyone else experienced this or someone may be able to give me some advice. </p>\n<p>I used this model checkpoint in the public dataset to get score of 0.539 just like many people did. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/dragonzhang/f0-25-yolox-pth</a></p>\n<p>Then, I moved on to train my own yolox model with tuned hyperparameters and more epochs. I was able to get better Precision, Recall and mAP at the end of the training loop. However, when I submitted to the leaderboard, I got much lower results than the previous which used the checkpoint in the public dataset mentioned above. </p>\n<p>Leaderboard score<br>\nPublic checkpoint - 0.539<br>\nMy checkpoint - 0.384</p>\n<p><strong>Public checkpoint</strong><br>\nAverage forward time: 46.69 ms, Average NMS time: 3.08 ms, Average inference time: 49.77 ms<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.262<br>\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.547<br>\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.205<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.120<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.300<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.153<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.318<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.318<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.160<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.361<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000</p>\n<p>2021-12-18 09:57:42.160 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config<br>\n2021-12-18 09:57:43.408 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is <strong>47.58</strong></p>\n<p><strong>My checkpoint</strong><br>\nAverage forward time: 46.02 ms, Average NMS time: 2.13 ms, Average inference time: 48.15 ms<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.586<br>\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.969<br>\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.655<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.479<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.607<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.668<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.281<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.612<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.651<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.563<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.672<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.686</p>\n<p>2022-01-02 06:45:52.360 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config<br>\n2022-01-02 06:45:54.322 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is <strong>58.64</strong></p>\n<p>As you can see from the above, the best AP from my checkpoint is 58.64 comparing to 47.58 from the public checkpoint. </p>\n<p>I understand we are using the average F2 score which is different to AP (Average Precision) as F2 puts more weight on recall. Nevertheless, I should be getting a better result if both of my precision and recall are better right?</p>\n<p>![<a href=\"https://blog.paperspace.com/content/images/2020/09/Fig04-1.jpg\" target=\"_blank\">https://blog.paperspace.com/content/images/2020/09/Fig04-1.jpg</a>]</p>\n<p>The inference notebook I used is very similar to this one:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539?kernelSessionId=83072655</a></p>\n<p>And I change the checkpoint by updating the variable 'CHECKPOINT_FILE'. </p>\n<p>The only thing I can think of at the moment… is maybe I need to tune the thresholds as my model is different now? it shouldn't really make too big of difference tho… <br>\nconfthre = 0.35<br>\nnmsthre = 0.4</p>\n<p>Any comments will be much appreciated!!</p>",
  "messages": [
    {
      "id": 1636401,
      "postDate": "2022-01-02T20:43:23.817Z",
      "content": "<p>Hi everyone, </p>\n<p>This is doing my head in… not sure if anyone else experienced this or someone may be able to give me some advice. </p>\n<p>I used this model checkpoint in the public dataset to get score of 0.539 just like many people did. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/dragonzhang/f0-25-yolox-pth</a></p>\n<p>Then, I moved on to train my own yolox model with tuned hyperparameters and more epochs. I was able to get better Precision, Recall and mAP at the end of the training loop. However, when I submitted to the leaderboard, I got much lower results than the previous which used the checkpoint in the public dataset mentioned above. </p>\n<p>Leaderboard score<br>\nPublic checkpoint - 0.539<br>\nMy checkpoint - 0.384</p>\n<p><strong>Public checkpoint</strong><br>\nAverage forward time: 46.69 ms, Average NMS time: 3.08 ms, Average inference time: 49.77 ms<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.262<br>\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.547<br>\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.205<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.120<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.300<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.153<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.318<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.318<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.160<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.361<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000</p>\n<p>2021-12-18 09:57:42.160 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config<br>\n2021-12-18 09:57:43.408 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is <strong>47.58</strong></p>\n<p><strong>My checkpoint</strong><br>\nAverage forward time: 46.02 ms, Average NMS time: 2.13 ms, Average inference time: 48.15 ms<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.586<br>\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.969<br>\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.655<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.479<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.607<br>\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.668<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.281<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.612<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.651<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.563<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.672<br>\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.686</p>\n<p>2022-01-02 06:45:52.360 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config<br>\n2022-01-02 06:45:54.322 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is <strong>58.64</strong></p>\n<p>As you can see from the above, the best AP from my checkpoint is 58.64 comparing to 47.58 from the public checkpoint. </p>\n<p>I understand we are using the average F2 score which is different to AP (Average Precision) as F2 puts more weight on recall. Nevertheless, I should be getting a better result if both of my precision and recall are better right?</p>\n<p>![<a href=\"https://blog.paperspace.com/content/images/2020/09/Fig04-1.jpg\" target=\"_blank\">https://blog.paperspace.com/content/images/2020/09/Fig04-1.jpg</a>]</p>\n<p>The inference notebook I used is very similar to this one:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539?kernelSessionId=83072655</a></p>\n<p>And I change the checkpoint by updating the variable 'CHECKPOINT_FILE'. </p>\n<p>The only thing I can think of at the moment… is maybe I need to tune the thresholds as my model is different now? it shouldn't really make too big of difference tho… <br>\nconfthre = 0.35<br>\nnmsthre = 0.4</p>\n<p>Any comments will be much appreciated!!</p>",
      "rawMarkdown": "Hi everyone, \n\nThis is doing my head in... not sure if anyone else experienced this or someone may be able to give me some advice. \n\nI used this model checkpoint in the public dataset to get score of 0.539 just like many people did. [https://www.kaggle.com/dragonzhang/f0-25-yolox-pth](url)\n\nThen, I moved on to train my own yolox model with tuned hyperparameters and more epochs. I was able to get better Precision, Recall and mAP at the end of the training loop. However, when I submitted to the leaderboard, I got much lower results than the previous which used the checkpoint in the public dataset mentioned above. \n\nLeaderboard score\nPublic checkpoint - 0.539\nMy checkpoint - 0.384\n\n**Public checkpoint**\nAverage forward time: 46.69 ms, Average NMS time: 3.08 ms, Average inference time: 49.77 ms\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.262\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.547\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.205\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.120\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.300\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.153\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.318\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.318\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.160\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.361\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000\n\n2021-12-18 09:57:42.160 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config\n2021-12-18 09:57:43.408 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is **47.58**\n\n**My checkpoint**\nAverage forward time: 46.02 ms, Average NMS time: 2.13 ms, Average inference time: 48.15 ms\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.586\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.969\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.655\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.479\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.607\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.668\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.281\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.612\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.651\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.563\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.672\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.686\n\n2022-01-02 06:45:52.360 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config\n2022-01-02 06:45:54.322 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is **58.64**\n\nAs you can see from the above, the best AP from my checkpoint is 58.64 comparing to 47.58 from the public checkpoint. \n\nI understand we are using the average F2 score which is different to AP (Average Precision) as F2 puts more weight on recall. Nevertheless, I should be getting a better result if both of my precision and recall are better right?\n\n![https://blog.paperspace.com/content/images/2020/09/Fig04-1.jpg]\n\nThe inference notebook I used is very similar to this one:\n[https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539?kernelSessionId=83072655](url)\n\nAnd I change the checkpoint by updating the variable 'CHECKPOINT_FILE'. \n\nThe only thing I can think of at the moment... is maybe I need to tune the thresholds as my model is different now? it shouldn't really make too big of difference tho... \nconfthre = 0.35\nnmsthre = 0.4\n\nAny comments will be much appreciated!!",
      "votes": 26
    },
    {
      "id": 1638482,
      "postDate": "2022-01-04T18:56:28.563Z",
      "content": "<p>i think there is going to be video_3, video_4 in the test (while  video_0, video_1, video_2 are the given train). according to the COTS dataset paper, there are 5 locations for data collection. each location has about 6 to 8k images</p>\n<p>you may want to try these, e.g</p>\n<ol>\n<li>if I use train = subset of  video_n only, test = of  video_n only, what is the metric performance?</li>\n<li>if I use train = subset of  video_n only, test = of  video_m only, what is the metric performance?</li>\n</ol>\n<p>(we are trying to measure domain difference between the different video id) pay attention if cross domain will increase number of FP and decrease the confidence score of target object.</p>\n<ol>\n<li>what augmentation/tricks can I use so that my model is robust against the change of video id.</li>\n</ol>\n<p>if you are successful, you should be able to e.g. train on video_0 + video_1, and get good results when testing on video_2</p>",
      "rawMarkdown": "i think there is going to be video\\_3, video\\_4 in the test (while  video\\_0, video\\_1, video\\_2 are the given train). according to the COTS dataset paper, there are 5 locations for data collection. each location has about 6 to 8k images\n\nyou may want to try these, e.g\n1. if I use train = subset of  video\\_n only, test = of  video\\_n only, what is the metric performance?\n2. if I use train = subset of  video\\_n only, test = of  video\\_m only, what is the metric performance?\n\n(we are trying to measure domain difference between the different video id) pay attention if cross domain will increase number of FP and decrease the confidence score of target object.\n\n3. what augmentation/tricks can I use so that my model is robust against the change of video id.\n\nif you are successful, you should be able to e.g. train on video\\_0 + video\\_1, and get good results when testing on video\\_2",
      "votes": 14,
      "replies": [
        {
          "id": 1639147,
          "postDate": "2022-01-05T12:30:34.140Z",
          "content": "<p>here are the experimental results</p>\n<p><img src=\"https://i.ibb.co/j5Vkb9h/Selection-999-591.png\" alt=\"\"></p>",
          "rawMarkdown": "here are the experimental results\n\n![](https://i.ibb.co/j5Vkb9h/Selection-999-591.png)",
          "votes": 18
        },
        {
          "id": 1639867,
          "postDate": "2022-01-06T02:11:02.437Z",
          "content": "<p>effect of better augmentation !!!!<br>\n<img src=\"https://i.ibb.co/NFry0MS/Selection-999-592.png\" alt=\"\"></p>",
          "rawMarkdown": "effect of better augmentation !!!!\n![](https://i.ibb.co/NFry0MS/Selection-999-592.png)",
          "votes": 9
        },
        {
          "id": 1639904,
          "postDate": "2022-01-06T03:20:27.413Z",
          "content": "<p>Nice ablation experiments,  what is the difference between your weak augmentation and the strong augmentation?  I have tried some strong augmentation methods such as copy paste but only obtain a little improvement.</p>",
          "rawMarkdown": "Nice ablation experiments,  what is the difference between your weak augmentation and the strong augmentation?  I have tried some strong augmentation methods such as copy paste but only obtain a little improvement."
        },
        {
          "id": 1639938,
          "postDate": "2022-01-06T04:31:09.947Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! Agree with you about videos 3 and 4 being in the test. To clarify may I ask if one uses the video-id as the train set, wouldn't that also decrease the number of training samples -  hence the need for data augmentation? For me, that was an aversion to splitting like that, so I split using sequences. I am currently experimenting with ways to expand/augment the training dataset. </p>",
          "rawMarkdown": "Thank you @hengck23! Agree with you about videos 3 and 4 being in the test. To clarify may I ask if one uses the video-id as the train set, wouldn't that also decrease the number of training samples -  hence the need for data augmentation? For me, that was an aversion to splitting like that, so I split using sequences. I am currently experimenting with ways to expand/augment the training dataset. "
        },
        {
          "id": 1639954,
          "postDate": "2022-01-06T04:55:21.243Z",
          "content": "<p>\"video-id as the train set, wouldn't that also decrease the number of training samples?\"</p>\n<p>using video id as validation split is for me to design augmentation methods. you can see that the initial CV results range from 0.37 to 0.51. My target is to design augmentation so that this difference is minimized, preferably all above 0.55 regardless of video id split (which I think is possible) and the difference between LB and CV should be smaller than 0.04</p>\n<p>after that, maybe I will change to split by seq or other cv splitting strategy.</p>\n<p>from my experiments, it seems that the public test data is close and smliar to some subset of video0,1,2.<br>\nthat is why some fold performs better than others.</p>\n<p>But this results can be unreliable because  I am not sure about the hidden private test set (i.e. if the hidden  private test set is as similar, compared to the public, to the train or not)</p>",
          "rawMarkdown": "\"video-id as the train set, wouldn't that also decrease the number of training samples?\"\n\nusing video id as validation split is for me to design augmentation methods. you can see that the initial CV results range from 0.37 to 0.51. My target is to design augmentation so that this difference is minimized, preferably all above 0.55 regardless of video id split (which I think is possible) and the difference between LB and CV should be smaller than 0.04\n\nafter that, maybe I will change to split by seq or other cv splitting strategy.\n\nfrom my experiments, it seems that the public test data is close and smliar to some subset of video0,1,2.\nthat is why some fold performs better than others.\n\nBut this results can be unreliable because  I am not sure about the hidden private test set (i.e. if the hidden  private test set is as similar, compared to the public, to the train or not)",
          "votes": 2
        },
        {
          "id": 1639964,
          "postDate": "2022-01-06T05:02:46.197Z",
          "content": "<p>\"I have tried some strong augmentation methods such as copy paste but only obtain a little improvement.\"</p>\n<p>you should make a list of variations based on the observation of the images and understanding of data collection process.</p>\n<p>for example:</p>\n<pre><code>1. camera pose (is the camera directly above or at an angle from COTS object)\n- Augmentation: affine, perspective image transform\n\n2. camera distance from object\n- Augmentation: scale transform\n(how much scale to use depends on the object size distribution)\n\n3. deep of water\n- Augmentation:  more bluish,  whitish ..., light scattering, ..\n(hsv change and intensity, color shift)\n\n4. motion blur, focus blur\n\n5. context\n- COTs is on coral or on sands, etc\nAugmentation  : cut and paste\n\n6. water quality\n- bubbles, small fishes, mirky ...\n\n7. occlusions\n\n8. underwater shadows\n\n9. density of COTS\n- occurs in groups. etc ...\n- relation between location and size of box\n\n10. pose of COTS (are their star legs spread out, etc ...)\n\n11. camera noise \n</code></pre>",
          "rawMarkdown": "\"I have tried some strong augmentation methods such as copy paste but only obtain a little improvement.\"\n\nyou should make a list of variations based on the observation of the images and understanding of data collection process.\n\nfor example:\n```\n1. camera pose (is the camera directly above or at an angle from COTS object)\n- Augmentation: affine, perspective image transform\n\n2. camera distance from object\n- Augmentation: scale transform\n(how much scale to use depends on the object size distribution)\n\n3. deep of water\n- Augmentation:  more bluish,  whitish ..., light scattering, ..\n(hsv change and intensity, color shift)\n\n4. motion blur, focus blur\n\n5. context\n- COTs is on coral or on sands, etc\nAugmentation  : cut and paste\n\n6. water quality\n- bubbles, small fishes, mirky ...\n\n7. occlusions\n\n8. underwater shadows\n\n9. density of COTS\n- occurs in groups. etc ...\n- relation between location and size of box\n\n10. pose of COTS (are their star legs spread out, etc ...)\n\n11. camera noise \n\n```",
          "votes": 19
        },
        {
          "id": 1639974,
          "postDate": "2022-01-06T05:11:33.780Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thank you. I never thought about approaching the data that way as a means to optimize the augmentations so I learned something valuable in your work. Very cool to know you are from NUS too! </p>",
          "rawMarkdown": "@hengck23 thank you. I never thought about approaching the data that way as a means to optimize the augmentations so I learned something valuable in your work. Very cool to know you are from NUS too! "
        },
        {
          "id": 1639989,
          "postDate": "2022-01-06T05:21:51.430Z",
          "content": "<p>signs that shows that augmentation is correct:</p>\n<ul>\n<li>you can train from scratch</li>\n<li>not overfitting to validation set when you train with many more epochs</li>\n<li>transformer shows huge improvement (also larger model shows better results)</li>\n<li>results is about the same if you different models</li>\n<li>CV and LB is close</li>\n</ul>",
          "rawMarkdown": "signs that shows that augmentation is correct:\n- you can train from scratch\n- not overfitting to validation set when you train with many more epochs\n- transformer shows huge improvement (also larger model shows better results)\n- results is about the same if you different models\n- CV and LB is close",
          "votes": 2
        },
        {
          "id": 1639990,
          "postDate": "2022-01-06T05:22:23.237Z",
          "content": "<p>there is strong evidence that there are many smaller COTS objects the public test set</p>",
          "rawMarkdown": "there is strong evidence that there are many smaller COTS objects the public test set",
          "votes": 1
        },
        {
          "id": 1640061,
          "postDate": "2022-01-06T06:33:11.363Z",
          "content": "<p>Got it, thanks very much!</p>",
          "rawMarkdown": "Got it, thanks very much!"
        },
        {
          "id": 1640443,
          "postDate": "2022-01-06T13:43:02Z",
          "content": "<p>update of results<br>\n<img src=\"https://i.ibb.co/dm5cQCw/Selection-999-593.png\" alt=\"\"></p>",
          "rawMarkdown": "update of results\n![](https://i.ibb.co/dm5cQCw/Selection-999-593.png)",
          "votes": 5
        },
        {
          "id": 1640559,
          "postDate": "2022-01-06T15:21:14.047Z",
          "content": "<p><a href=\"https://www.kaggle.com/wilbertbhtan\" target=\"_blank\">@wilbertbhtan</a> Did you use only images with labels in your training or add some without labels ?</p>",
          "rawMarkdown": "@wilbertbhtan Did you use only images with labels in your training or add some without labels ?"
        },
        {
          "id": 1640581,
          "postDate": "2022-01-06T15:41:39.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/garvitgarg\" target=\"_blank\">@garvitgarg</a> I used 1% of the training of set with the images without annotations as per Yolov5 documentation. I imagine can add more background images (&lt;10%) to reduce false positives. </p>\n<p><code>Background images. Background images are images with no objects that are added to a dataset to reduce False Positives (FP). We recommend about 0-10% background images to help reduce FPs (COCO has 1000 background images for reference, 1% of the total). No labels are required for background images.</code></p>\n<p>From <a href=\"https://github.com/ultralytics/yolov5/wiki/Tips-for-Best-Training-Results\" target=\"_blank\">Yolov5 </a></p>",
          "rawMarkdown": "@garvitgarg I used 1% of the training of set with the images without annotations as per Yolov5 documentation. I imagine can add more background images (<10%) to reduce false positives. \n\n`Background images. Background images are images with no objects that are added to a dataset to reduce False Positives (FP). We recommend about 0-10% background images to help reduce FPs (COCO has 1000 background images for reference, 1% of the total). No labels are required for background images.`\n\nFrom [Yolov5 ](https://github.com/ultralytics/yolov5/wiki/Tips-for-Best-Training-Results)",
          "votes": 1
        },
        {
          "id": 1640637,
          "postDate": "2022-01-06T16:53:35.757Z",
          "content": "<p>typical miss (i.e. what to augment)</p>\n<p><img src=\"https://i.ibb.co/F8tP6mf/Selection-999-617.png\" alt=\"\"></p>\n<p>it is important is ask yourself this question: \"now I have a miss detection, should I handle it or just ignore it?\"</p>\n<ul>\n<li>it depends on the frequency of occurrences of such difficult samples</li>\n<li>if detecting one difficult case ends up in only 1 or 2 FP, then it is worth it. if it ends up with 10 FP, then it is not worth it.</li>\n</ul>",
          "rawMarkdown": "typical miss (i.e. what to augment)\n\n![](https://i.ibb.co/F8tP6mf/Selection-999-617.png)\n\n\nit is important is ask yourself this question: \"now I have a miss detection, should I handle it or just ignore it?\"\n- it depends on the frequency of occurrences of such difficult samples\n- if detecting one difficult case ends up in only 1 or 2 FP, then it is worth it. if it ends up with 10 FP, then it is not worth it.",
          "votes": 4
        },
        {
          "id": 1640776,
          "postDate": "2022-01-06T19:54:27.727Z",
          "content": "<p>another update:</p>\n<p>this will be the final update of the results. i would like to conclude:</p>\n<ul>\n<li>augmentation is the key  to CV/LB gap</li>\n<li>better split would help, but just by using split by video id could already get the public kernel results</li>\n<li>with better and more augmentation, bigger models benefits</li>\n</ul>\n<p><img src=\"https://i.ibb.co/6yhCsgX/Selection-999-619.png\" alt=\"\"></p>\n<p>the augmentation used to produce the above results is:</p>\n<pre><code>def train_augment6(image, target):\n    image, target = do_random_flip(image, target) # hflip, vflip or both\n    #image, target = do_random_hflip(image, target)\n\n    if np.random.rand() &lt; 0.7:\n        for func in np.random.choice([\n            lambda image, target: do_random_perspective(image, target, m=0.3),\n            lambda image, target: do_random_zoom_small(image, target, m=20), #aug min object size is 20\n            lambda image, target: do_random_zoom_big(image, target, m=120),  #aug max object size is 120\n            lambda image, target: do_random_rotate(image, target, m=30), #up to 30 degree\n        ], 1):\n            image, target = func(image, target)\n            pass\n        pass\n\n    if np.random.rand() &lt; 0.7:\n        for func in np.random.choice([\n            lambda image: do_random_hsv(image, h=20, s=50, v=50),\n            lambda image: do_random_guassian_blur(image, k=[3, 5], s=[0.1, 2.0]),\n            lambda image: do_random_noise(image, m=0.08),\n        ], 1):\n            image = func(image)\n            pass\n\n    return image, target\n</code></pre>\n<p>more augmentations (e.g. mosaic, cut and paste, light scattering, cutout inside bbox, …) will further improve results but are not yet used and shown here</p>",
          "rawMarkdown": "another update:\n\nthis will be the final update of the results. i would like to conclude:\n- augmentation is the key  to CV/LB gap\n- better split would help, but just by using split by video id could already get the public kernel results\n- with better and more augmentation, bigger models benefits\n\n![](https://i.ibb.co/6yhCsgX/Selection-999-619.png)\n\nthe augmentation used to produce the above results is:\n\n```\n\ndef train_augment6(image, target):\n    image, target = do_random_flip(image, target) # hflip, vflip or both\n    #image, target = do_random_hflip(image, target)\n\n    if np.random.rand() < 0.7:\n        for func in np.random.choice([\n            lambda image, target: do_random_perspective(image, target, m=0.3),\n            lambda image, target: do_random_zoom_small(image, target, m=20), #aug min object size is 20\n            lambda image, target: do_random_zoom_big(image, target, m=120),  #aug max object size is 120\n            lambda image, target: do_random_rotate(image, target, m=30), #up to 30 degree\n        ], 1):\n            image, target = func(image, target)\n            pass\n        pass\n\n    if np.random.rand() < 0.7:\n        for func in np.random.choice([\n            lambda image: do_random_hsv(image, h=20, s=50, v=50),\n            lambda image: do_random_guassian_blur(image, k=[3, 5], s=[0.1, 2.0]),\n            lambda image: do_random_noise(image, m=0.08),\n        ], 1):\n            image = func(image)\n            pass\n\n    return image, target\n\n\n```\n\nmore augmentations (e.g. mosaic, cut and paste, light scattering, cutout inside bbox, ...) will further improve results but are not yet used and shown here",
          "votes": 12
        },
        {
          "id": 1640943,
          "postDate": "2022-01-07T01:31:15.540Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for sharing a lot of information. I gained much insight out of this discussion.</p>\n<p>I have a question about your augmentation. Can I ask what <code>do_random_zoom_small</code> and <code>do_random_zoom_big</code> doing? Is this scaling entire image or scaling just for bounding box area?</p>",
          "rawMarkdown": "@hengck23 Thank you for sharing a lot of information. I gained much insight out of this discussion.\n\nI have a question about your augmentation. Can I ask what `do_random_zoom_small` and `do_random_zoom_big` doing? Is this scaling entire image or scaling just for bounding box area?"
        }
      ]
    },
    {
      "id": 1636878,
      "postDate": "2022-01-03T10:56:05.100Z",
      "content": "<p>this is due to the difference in the test and validation data. this is what you can do:</p>\n<ol>\n<li>try different fold of train and validation set</li>\n<li>if you get good validation results by lowering the confidence threshold, e.g. 0.1, it may be likely that test results can be worse. because a low threshold should lead many more FP (it may not happen in your validation set but it can happen to another set)</li>\n</ol>\n<p>the threshold should be selected to give a good FP per 100 images (note that I use per 100 images, meaning that you need many more images in validation to give a good statistics)</p>\n<p>there are 13000 test images, use your train data to estimate num of empty, non-empty images in the test. Also, the distribution of the number of targets per image in test. Then ask yourself: \" do I have enough images in the validation to get a reliable statistics?\"</p>",
      "rawMarkdown": "this is due to the difference in the test and validation data. this is what you can do:\n1. try different fold of train and validation set\n2. if you get good validation results by lowering the confidence threshold, e.g. 0.1, it may be likely that test results can be worse. because a low threshold should lead many more FP (it may not happen in your validation set but it can happen to another set)\n\nthe threshold should be selected to give a good FP per 100 images (note that I use per 100 images, meaning that you need many more images in validation to give a good statistics)\n\nthere are 13000 test images, use your train data to estimate num of empty, non-empty images in the test. Also, the distribution of the number of targets per image in test. Then ask yourself: \" do I have enough images in the validation to get a reliable statistics?\"",
      "votes": 9,
      "replies": [
        {
          "id": 1636884,
          "postDate": "2022-01-03T11:01:57.333Z",
          "content": "<p>tip:  in theory, you can reduce fp by training with any images (e.g. imagenet, cityscape, … just do a bootstrapping to collect hard negative)</p>\n<p>for me, I am downloading underwater, diving video from youtube, etc</p>",
          "rawMarkdown": "tip:  in theory, you can reduce fp by training with any images (e.g. imagenet, cityscape, ... just do a bootstrapping to collect hard negative)\n\nfor me, I am downloading underwater, diving video from youtube, etc",
          "votes": 3
        },
        {
          "id": 1641597,
          "postDate": "2022-01-07T15:13:26.130Z",
          "content": "<p>Hello, I would like to ask when calculating the F2 score for the validation set, do we have to ensure that the percentage of unannotated images in the validation set is equal to the percentage of unannotated images in the entire dataset(i.e., <code>4919/23501</code>)?</p>",
          "rawMarkdown": "Hello, I would like to ask when calculating the F2 score for the validation set, do we have to ensure that the percentage of unannotated images in the validation set is equal to the percentage of unannotated images in the entire dataset(i.e., `4919/23501`)?"
        },
        {
          "id": 1641974,
          "postDate": "2022-01-07T22:13:48.060Z",
          "content": "<p><a href=\"https://www.kaggle.com/shengzheliu\" target=\"_blank\">@shengzheliu</a> F2 score calculated only for the positive dataset is typically higher than that for the entire dataset because of the FPs on the negative dataset. So it tends to cause more CV/LB gaps. To check if you model correctly predicts on the back ground, validating on entire (positive and negative) dataset is necessary.</p>\n<p>One point we have to care is, we can’t know the positive-negative ratio of public and private dataset.<br>\nTherefore, the hyper parameter (like confidence threshold) optimized on your validation data doesn’t necessarily optimal on the test data. E.g. suppose test data has much less annotated data than your validation data and the optimal confidence threshold to your validation set is 0.3. Since test set contain more background data than validation set, your model tend to predict more FPs on test set than validation set. Therefore, the optimal confidence threshold for the test set is typically less than that for validation set.</p>\n<p>For that reason, setting the positive negative ratio equal to train dataset doesn’t necessarily cause optimal performance on the test set. I think we have to design robust model to any positive-negative ratio.</p>",
          "rawMarkdown": "@shengzheliu F2 score calculated only for the positive dataset is typically higher than that for the entire dataset because of the FPs on the negative dataset. So it tends to cause more CV/LB gaps. To check if you model correctly predicts on the back ground, validating on entire (positive and negative) dataset is necessary.\n\nOne point we have to care is, we can’t know the positive-negative ratio of public and private dataset.\nTherefore, the hyper parameter (like confidence threshold) optimized on your validation data doesn’t necessarily optimal on the test data. E.g. suppose test data has much less annotated data than your validation data and the optimal confidence threshold to your validation set is 0.3. Since test set contain more background data than validation set, your model tend to predict more FPs on test set than validation set. Therefore, the optimal confidence threshold for the test set is typically less than that for validation set.\n\nFor that reason, setting the positive negative ratio equal to train dataset doesn’t necessarily cause optimal performance on the test set. I think we have to design robust model to any positive-negative ratio.",
          "votes": 1
        },
        {
          "id": 1646182,
          "postDate": "2022-01-11T14:24:00.860Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Thanks for sharing many helpful tips. Did you trained unlabeled data too? In my case with Yolov5l, object loss is keep increasing and recall is getting lower. If there any tip to  increase true positive rate and reduce false negative rate. Thanks</p>",
          "rawMarkdown": "@hengck23  Thanks for sharing many helpful tips. Did you trained unlabeled data too? In my case with Yolov5l, object loss is keep increasing and recall is getting lower. If there any tip to  increase true positive rate and reduce false negative rate. Thanks"
        }
      ]
    },
    {
      "id": 1637390,
      "postDate": "2022-01-03T20:38:24.133Z",
      "content": "<p>Think I have found the cause for my case guys. For me, it was due to the way I split the train/val dataset. </p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723</a></p>\n<p>The discussion thread above is worth a read. </p>\n<p>In a summary, a few things could attribute to the big gap between val and LB score. </p>\n<ol>\n<li>the way you split train/val data</li>\n<li>number of background images (no startfish images) in the test set. Actually, this may vary between the public test set and private test set too. </li>\n<li>variance between the video sequence in train, val and test. </li>\n</ol>",
      "rawMarkdown": "Think I have found the cause for my case guys. For me, it was due to the way I split the train/val dataset. \n\n[https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723](url)\n\nThe discussion thread above is worth a read. \n\nIn a summary, a few things could attribute to the big gap between val and LB score. \n\n1. the way you split train/val data\n2. number of background images (no startfish images) in the test set. Actually, this may vary between the public test set and private test set too. \n3. variance between the video sequence in train, val and test. \n",
      "votes": 2
    },
    {
      "id": 1636798,
      "postDate": "2022-01-03T09:27:03.770Z",
      "content": "<p>If i'm correct, you trained from other's weight, instead of scratch?</p>",
      "rawMarkdown": "If i'm correct, you trained from other's weight, instead of scratch?",
      "replies": [
        {
          "id": 1636805,
          "postDate": "2022-01-03T09:35:42.373Z",
          "content": "<p>used weights pretrained from coco</p>",
          "rawMarkdown": "used weights pretrained from coco"
        },
        {
          "id": 1637613,
          "postDate": "2022-01-04T04:25:58.770Z",
          "content": "<p>Dose it matters?</p>",
          "rawMarkdown": "Dose it matters?"
        },
        {
          "id": 1638062,
          "postDate": "2022-01-04T12:14:14.447Z",
          "content": "<p>Maybe try from scratch</p>",
          "rawMarkdown": "Maybe try from scratch"
        }
      ]
    },
    {
      "id": 1636741,
      "postDate": "2022-01-03T08:29:20.663Z",
      "content": "<p>Similar observation here. Trained model with some augmentations and got mAP of 61. But my score was only 0.501. My current best single model with mAP = 42 scored  0.510.</p>",
      "rawMarkdown": "Similar observation here. Trained model with some augmentations and got mAP of 61. But my score was only 0.501. My current best single model with mAP = 42 scored  0.510."
    },
    {
      "id": 1636582,
      "postDate": "2022-01-03T04:45:12.367Z",
      "content": "<p>confthre may shift for optimal  LB, but not than much difference.</p>",
      "rawMarkdown": "confthre may shift for optimal  LB, but not than much difference."
    },
    {
      "id": 1636581,
      "postDate": "2022-01-03T04:42:53.620Z",
      "content": "<p>I  met same problem.   After I bought colab pro,   I have more GPU available.  several checkpoints seems better, but <br>\nLB just similar or worse.</p>",
      "rawMarkdown": "I  met same problem.   After I bought colab pro,   I have more GPU available.  several checkpoints seems better, but \nLB just similar or worse."
    },
    {
      "id": 1636576,
      "postDate": "2022-01-03T04:36:55.423Z",
      "content": "<p>I am observing similar results. I trained a yolov5l with some pre-processing and training modifications, higher mAP than public training notebooks using the same fold, but lb is only 0.389. Not really sure what is going on.</p>",
      "rawMarkdown": "I am observing similar results. I trained a yolov5l with some pre-processing and training modifications, higher mAP than public training notebooks using the same fold, but lb is only 0.389. Not really sure what is going on.",
      "replies": [
        {
          "id": 1641852,
          "postDate": "2022-01-07T19:03:16.343Z",
          "content": "<p>Hey I have been trying yolov5l too. I was able to get LB 0.425 but I am struggling to go above this. I got 0.416 without tracking and 0.425 with implementation of Norfair tracking.<br>\nTake a look at video generated this video is generated from 4th fold [Green bounding boxes are yolov5l predictions, Red bounding boxes are original annotations and Black bounding boxes are after tacking]<br>\n<a href=\"url\" target=\"_blank\">https://drive.google.com/file/d/11gp6IDav0X7fNxBYG3C0uL6qOZ6pj0eD/view?usp=sharing</a></p>",
          "rawMarkdown": "Hey I have been trying yolov5l too. I was able to get LB 0.425 but I am struggling to go above this. I got 0.416 without tracking and 0.425 with implementation of Norfair tracking.\nTake a look at video generated this video is generated from 4th fold [Green bounding boxes are yolov5l predictions, Red bounding boxes are original annotations and Black bounding boxes are after tacking]\n[https://drive.google.com/file/d/11gp6IDav0X7fNxBYG3C0uL6qOZ6pj0eD/view?usp=sharing](url)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1638909,
      "postDate": "2022-01-05T06:44:25.243Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1638482,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-04T18:56:28.563000",
      "content": "<p>i think there is going to be video_3, video_4 in the test (while  video_0, video_1, video_2 are the given train). according to the COTS dataset paper, there are 5 locations for data collection. each location has about 6 to 8k images</p>\n<p>you may want to try these, e.g</p>\n<ol>\n<li>if I use train = subset of  video_n only, test = of  video_n only, what is the metric performance?</li>\n<li>if I use train = subset of  video_n only, test = of  video_m only, what is the metric performance?</li>\n</ol>\n<p>(we are trying to measure domain difference between the different video id) pay attention if cross domain will increase number of FP and decrease the confidence score of target object.</p>\n<ol>\n<li>what augmentation/tricks can I use so that my model is robust against the change of video id.</li>\n</ol>\n<p>if you are successful, you should be able to e.g. train on video_0 + video_1, and get good results when testing on video_2</p>",
      "votes": 14,
      "replies": [
        {
          "id": 1639147,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-05T12:30:34.140000",
          "content": "<p>here are the experimental results</p>\n<p><img src=\"https://i.ibb.co/j5Vkb9h/Selection-999-591.png\" alt=\"\"></p>",
          "votes": 18,
          "replies": []
        },
        {
          "id": 1639867,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T02:11:02.437000",
          "content": "<p>effect of better augmentation !!!!<br>\n<img src=\"https://i.ibb.co/NFry0MS/Selection-999-592.png\" alt=\"\"></p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1639904,
          "author_name": "zhengye",
          "author_url": "",
          "post_date": "2022-01-06T03:20:27.413000",
          "content": "<p>Nice ablation experiments,  what is the difference between your weak augmentation and the strong augmentation?  I have tried some strong augmentation methods such as copy paste but only obtain a little improvement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1639938,
          "author_name": "Wilbert Tan",
          "author_url": "",
          "post_date": "2022-01-06T04:31:09.947000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! Agree with you about videos 3 and 4 being in the test. To clarify may I ask if one uses the video-id as the train set, wouldn't that also decrease the number of training samples -  hence the need for data augmentation? For me, that was an aversion to splitting like that, so I split using sequences. I am currently experimenting with ways to expand/augment the training dataset. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1639954,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T04:55:21.243000",
          "content": "<p>\"video-id as the train set, wouldn't that also decrease the number of training samples?\"</p>\n<p>using video id as validation split is for me to design augmentation methods. you can see that the initial CV results range from 0.37 to 0.51. My target is to design augmentation so that this difference is minimized, preferably all above 0.55 regardless of video id split (which I think is possible) and the difference between LB and CV should be smaller than 0.04</p>\n<p>after that, maybe I will change to split by seq or other cv splitting strategy.</p>\n<p>from my experiments, it seems that the public test data is close and smliar to some subset of video0,1,2.<br>\nthat is why some fold performs better than others.</p>\n<p>But this results can be unreliable because  I am not sure about the hidden private test set (i.e. if the hidden  private test set is as similar, compared to the public, to the train or not)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1639964,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T05:02:46.197000",
          "content": "<p>\"I have tried some strong augmentation methods such as copy paste but only obtain a little improvement.\"</p>\n<p>you should make a list of variations based on the observation of the images and understanding of data collection process.</p>\n<p>for example:</p>\n<pre><code>1. camera pose (is the camera directly above or at an angle from COTS object)\n- Augmentation: affine, perspective image transform\n\n2. camera distance from object\n- Augmentation: scale transform\n(how much scale to use depends on the object size distribution)\n\n3. deep of water\n- Augmentation:  more bluish,  whitish ..., light scattering, ..\n(hsv change and intensity, color shift)\n\n4. motion blur, focus blur\n\n5. context\n- COTs is on coral or on sands, etc\nAugmentation  : cut and paste\n\n6. water quality\n- bubbles, small fishes, mirky ...\n\n7. occlusions\n\n8. underwater shadows\n\n9. density of COTS\n- occurs in groups. etc ...\n- relation between location and size of box\n\n10. pose of COTS (are their star legs spread out, etc ...)\n\n11. camera noise \n</code></pre>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 1639974,
          "author_name": "Wilbert Tan",
          "author_url": "",
          "post_date": "2022-01-06T05:11:33.780000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thank you. I never thought about approaching the data that way as a means to optimize the augmentations so I learned something valuable in your work. Very cool to know you are from NUS too! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1639989,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T05:21:51.430000",
          "content": "<p>signs that shows that augmentation is correct:</p>\n<ul>\n<li>you can train from scratch</li>\n<li>not overfitting to validation set when you train with many more epochs</li>\n<li>transformer shows huge improvement (also larger model shows better results)</li>\n<li>results is about the same if you different models</li>\n<li>CV and LB is close</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1639990,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T05:22:23.237000",
          "content": "<p>there is strong evidence that there are many smaller COTS objects the public test set</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1640061,
          "author_name": "zhengye",
          "author_url": "",
          "post_date": "2022-01-06T06:33:11.363000",
          "content": "<p>Got it, thanks very much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1640443,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T13:43:02",
          "content": "<p>update of results<br>\n<img src=\"https://i.ibb.co/dm5cQCw/Selection-999-593.png\" alt=\"\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1640559,
          "author_name": "Garvit Garg",
          "author_url": "",
          "post_date": "2022-01-06T15:21:14.047000",
          "content": "<p><a href=\"https://www.kaggle.com/wilbertbhtan\" target=\"_blank\">@wilbertbhtan</a> Did you use only images with labels in your training or add some without labels ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1640581,
          "author_name": "Wilbert Tan",
          "author_url": "",
          "post_date": "2022-01-06T15:41:39.170000",
          "content": "<p><a href=\"https://www.kaggle.com/garvitgarg\" target=\"_blank\">@garvitgarg</a> I used 1% of the training of set with the images without annotations as per Yolov5 documentation. I imagine can add more background images (&lt;10%) to reduce false positives. </p>\n<p><code>Background images. Background images are images with no objects that are added to a dataset to reduce False Positives (FP). We recommend about 0-10% background images to help reduce FPs (COCO has 1000 background images for reference, 1% of the total). No labels are required for background images.</code></p>\n<p>From <a href=\"https://github.com/ultralytics/yolov5/wiki/Tips-for-Best-Training-Results\" target=\"_blank\">Yolov5 </a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1640637,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T16:53:35.757000",
          "content": "<p>typical miss (i.e. what to augment)</p>\n<p><img src=\"https://i.ibb.co/F8tP6mf/Selection-999-617.png\" alt=\"\"></p>\n<p>it is important is ask yourself this question: \"now I have a miss detection, should I handle it or just ignore it?\"</p>\n<ul>\n<li>it depends on the frequency of occurrences of such difficult samples</li>\n<li>if detecting one difficult case ends up in only 1 or 2 FP, then it is worth it. if it ends up with 10 FP, then it is not worth it.</li>\n</ul>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1640776,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-06T19:54:27.727000",
          "content": "<p>another update:</p>\n<p>this will be the final update of the results. i would like to conclude:</p>\n<ul>\n<li>augmentation is the key  to CV/LB gap</li>\n<li>better split would help, but just by using split by video id could already get the public kernel results</li>\n<li>with better and more augmentation, bigger models benefits</li>\n</ul>\n<p><img src=\"https://i.ibb.co/6yhCsgX/Selection-999-619.png\" alt=\"\"></p>\n<p>the augmentation used to produce the above results is:</p>\n<pre><code>def train_augment6(image, target):\n    image, target = do_random_flip(image, target) # hflip, vflip or both\n    #image, target = do_random_hflip(image, target)\n\n    if np.random.rand() &lt; 0.7:\n        for func in np.random.choice([\n            lambda image, target: do_random_perspective(image, target, m=0.3),\n            lambda image, target: do_random_zoom_small(image, target, m=20), #aug min object size is 20\n            lambda image, target: do_random_zoom_big(image, target, m=120),  #aug max object size is 120\n            lambda image, target: do_random_rotate(image, target, m=30), #up to 30 degree\n        ], 1):\n            image, target = func(image, target)\n            pass\n        pass\n\n    if np.random.rand() &lt; 0.7:\n        for func in np.random.choice([\n            lambda image: do_random_hsv(image, h=20, s=50, v=50),\n            lambda image: do_random_guassian_blur(image, k=[3, 5], s=[0.1, 2.0]),\n            lambda image: do_random_noise(image, m=0.08),\n        ], 1):\n            image = func(image)\n            pass\n\n    return image, target\n</code></pre>\n<p>more augmentations (e.g. mosaic, cut and paste, light scattering, cutout inside bbox, …) will further improve results but are not yet used and shown here</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 1640943,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-07T01:31:15.540000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thank you for sharing a lot of information. I gained much insight out of this discussion.</p>\n<p>I have a question about your augmentation. Can I ask what <code>do_random_zoom_small</code> and <code>do_random_zoom_big</code> doing? Is this scaling entire image or scaling just for bounding box area?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1636878,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-03T10:56:05.100000",
      "content": "<p>this is due to the difference in the test and validation data. this is what you can do:</p>\n<ol>\n<li>try different fold of train and validation set</li>\n<li>if you get good validation results by lowering the confidence threshold, e.g. 0.1, it may be likely that test results can be worse. because a low threshold should lead many more FP (it may not happen in your validation set but it can happen to another set)</li>\n</ol>\n<p>the threshold should be selected to give a good FP per 100 images (note that I use per 100 images, meaning that you need many more images in validation to give a good statistics)</p>\n<p>there are 13000 test images, use your train data to estimate num of empty, non-empty images in the test. Also, the distribution of the number of targets per image in test. Then ask yourself: \" do I have enough images in the validation to get a reliable statistics?\"</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1636884,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-03T11:01:57.333000",
          "content": "<p>tip:  in theory, you can reduce fp by training with any images (e.g. imagenet, cityscape, … just do a bootstrapping to collect hard negative)</p>\n<p>for me, I am downloading underwater, diving video from youtube, etc</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1641597,
          "author_name": "ShengzheLiu",
          "author_url": "",
          "post_date": "2022-01-07T15:13:26.130000",
          "content": "<p>Hello, I would like to ask when calculating the F2 score for the validation set, do we have to ensure that the percentage of unannotated images in the validation set is equal to the percentage of unannotated images in the entire dataset(i.e., <code>4919/23501</code>)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1641974,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-07T22:13:48.060000",
          "content": "<p><a href=\"https://www.kaggle.com/shengzheliu\" target=\"_blank\">@shengzheliu</a> F2 score calculated only for the positive dataset is typically higher than that for the entire dataset because of the FPs on the negative dataset. So it tends to cause more CV/LB gaps. To check if you model correctly predicts on the back ground, validating on entire (positive and negative) dataset is necessary.</p>\n<p>One point we have to care is, we can’t know the positive-negative ratio of public and private dataset.<br>\nTherefore, the hyper parameter (like confidence threshold) optimized on your validation data doesn’t necessarily optimal on the test data. E.g. suppose test data has much less annotated data than your validation data and the optimal confidence threshold to your validation set is 0.3. Since test set contain more background data than validation set, your model tend to predict more FPs on test set than validation set. Therefore, the optimal confidence threshold for the test set is typically less than that for validation set.</p>\n<p>For that reason, setting the positive negative ratio equal to train dataset doesn’t necessarily cause optimal performance on the test set. I think we have to design robust model to any positive-negative ratio.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1646182,
          "author_name": "Ichimaru Gin",
          "author_url": "",
          "post_date": "2022-01-11T14:24:00.860000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Thanks for sharing many helpful tips. Did you trained unlabeled data too? In my case with Yolov5l, object loss is keep increasing and recall is getting lower. If there any tip to  increase true positive rate and reduce false negative rate. Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1637390,
      "author_name": "Early Bird",
      "author_url": "",
      "post_date": "2022-01-03T20:38:24.133000",
      "content": "<p>Think I have found the cause for my case guys. For me, it was due to the way I split the train/val dataset. </p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723</a></p>\n<p>The discussion thread above is worth a read. </p>\n<p>In a summary, a few things could attribute to the big gap between val and LB score. </p>\n<ol>\n<li>the way you split train/val data</li>\n<li>number of background images (no startfish images) in the test set. Actually, this may vary between the public test set and private test set too. </li>\n<li>variance between the video sequence in train, val and test. </li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1636798,
      "author_name": "Owen Xing",
      "author_url": "",
      "post_date": "2022-01-03T09:27:03.770000",
      "content": "<p>If i'm correct, you trained from other's weight, instead of scratch?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1636805,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "2022-01-03T09:35:42.373000",
          "content": "<p>used weights pretrained from coco</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1637613,
          "author_name": "Kevin",
          "author_url": "",
          "post_date": "2022-01-04T04:25:58.770000",
          "content": "<p>Dose it matters?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1638062,
          "author_name": "Owen Xing",
          "author_url": "",
          "post_date": "2022-01-04T12:14:14.447000",
          "content": "<p>Maybe try from scratch</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1636741,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2022-01-03T08:29:20.663000",
      "content": "<p>Similar observation here. Trained model with some augmentations and got mAP of 61. But my score was only 0.501. My current best single model with mAP = 42 scored  0.510.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1636582,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-01-03T04:45:12.367000",
      "content": "<p>confthre may shift for optimal  LB, but not than much difference.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1636581,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-01-03T04:42:53.620000",
      "content": "<p>I  met same problem.   After I bought colab pro,   I have more GPU available.  several checkpoints seems better, but <br>\nLB just similar or worse.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1636576,
      "author_name": "Jebastin Nadar",
      "author_url": "",
      "post_date": "2022-01-03T04:36:55.423000",
      "content": "<p>I am observing similar results. I trained a yolov5l with some pre-processing and training modifications, higher mAP than public training notebooks using the same fold, but lb is only 0.389. Not really sure what is going on.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1641852,
          "author_name": "manas joshi",
          "author_url": "",
          "post_date": "2022-01-07T19:03:16.343000",
          "content": "<p>Hey I have been trying yolov5l too. I was able to get LB 0.425 but I am struggling to go above this. I got 0.416 without tracking and 0.425 with implementation of Norfair tracking.<br>\nTake a look at video generated this video is generated from 4th fold [Green bounding boxes are yolov5l predictions, Red bounding boxes are original annotations and Black bounding boxes are after tacking]<br>\n<a href=\"url\" target=\"_blank\">https://drive.google.com/file/d/11gp6IDav0X7fNxBYG3C0uL6qOZ6pj0eD/view?usp=sharing</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1638909,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-05T06:44:25.243000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1636401": "Hi everyone, \n\nThis is doing my head in... not sure if anyone else experienced this or someone may be able to give me some advice. \n\nI used this model checkpoint in the public dataset to get score of 0.539 just like many people did. [https://www.kaggle.com/dragonzhang/f0-25-yolox-pth](url)\n\nThen, I moved on to train my own yolox model with tuned hyperparameters and more epochs. I was able to get better Precision, Recall and mAP at the end of the training loop. However, when I submitted to the leaderboard, I got much lower results than the previous which used the checkpoint in the public dataset mentioned above. \n\nLeaderboard score\nPublic checkpoint - 0.539\nMy checkpoint - 0.384\n\n**Public checkpoint**\nAverage forward time: 46.69 ms, Average NMS time: 3.08 ms, Average inference time: 49.77 ms\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.262\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.547\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.205\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.120\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.300\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.153\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.318\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.318\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.160\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.361\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = -1.000\n\n2021-12-18 09:57:42.160 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config\n2021-12-18 09:57:43.408 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is **47.58**\n\n**My checkpoint**\nAverage forward time: 46.02 ms, Average NMS time: 2.13 ms, Average inference time: 48.15 ms\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.586\n Average Precision  (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.969\n Average Precision  (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.655\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.479\n Average Precision  (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.607\n Average Precision  (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.668\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=  1 ] = 0.281\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets= 10 ] = 0.612\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.651\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.563\n Average Recall     (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.672\n Average Recall     (AR) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.686\n\n2022-01-02 06:45:52.360 | INFO     | yolox.core.trainer:save_ckpt:318 - Save weights to ./YOLOX_outputs/cots_config\n2022-01-02 06:45:54.322 | INFO     | yolox.core.trainer:after_train:184 - Training of experiment is done and the best AP is **58.64**\n\nAs you can see from the above, the best AP from my checkpoint is 58.64 comparing to 47.58 from the public checkpoint. \n\nI understand we are using the average F2 score which is different to AP (Average Precision) as F2 puts more weight on recall. Nevertheless, I should be getting a better result if both of my precision and recall are better right?\n\n![https://blog.paperspace.com/content/images/2020/09/Fig04-1.jpg]\n\nThe inference notebook I used is very similar to this one:\n[https://www.kaggle.com/parapapapam/yolox-inference-tracking-on-cots-lb-0-539?kernelSessionId=83072655](url)\n\nAnd I change the checkpoint by updating the variable 'CHECKPOINT_FILE'. \n\nThe only thing I can think of at the moment... is maybe I need to tune the thresholds as my model is different now? it shouldn't really make too big of difference tho... \nconfthre = 0.35\nnmsthre = 0.4\n\nAny comments will be much appreciated!!",
    "1638482": "i think there is going to be video\\_3, video\\_4 in the test (while  video\\_0, video\\_1, video\\_2 are the given train). according to the COTS dataset paper, there are 5 locations for data collection. each location has about 6 to 8k images\n\nyou may want to try these, e.g\n1. if I use train = subset of  video\\_n only, test = of  video\\_n only, what is the metric performance?\n2. if I use train = subset of  video\\_n only, test = of  video\\_m only, what is the metric performance?\n\n(we are trying to measure domain difference between the different video id) pay attention if cross domain will increase number of FP and decrease the confidence score of target object.\n\n3. what augmentation/tricks can I use so that my model is robust against the change of video id.\n\nif you are successful, you should be able to e.g. train on video\\_0 + video\\_1, and get good results when testing on video\\_2",
    "1636878": "this is due to the difference in the test and validation data. this is what you can do:\n1. try different fold of train and validation set\n2. if you get good validation results by lowering the confidence threshold, e.g. 0.1, it may be likely that test results can be worse. because a low threshold should lead many more FP (it may not happen in your validation set but it can happen to another set)\n\nthe threshold should be selected to give a good FP per 100 images (note that I use per 100 images, meaning that you need many more images in validation to give a good statistics)\n\nthere are 13000 test images, use your train data to estimate num of empty, non-empty images in the test. Also, the distribution of the number of targets per image in test. Then ask yourself: \" do I have enough images in the validation to get a reliable statistics?\"",
    "1637390": "Think I have found the cause for my case guys. For me, it was due to the way I split the train/val dataset. \n\n[https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/293723](url)\n\nThe discussion thread above is worth a read. \n\nIn a summary, a few things could attribute to the big gap between val and LB score. \n\n1. the way you split train/val data\n2. number of background images (no startfish images) in the test set. Actually, this may vary between the public test set and private test set too. \n3. variance between the video sequence in train, val and test. \n",
    "1636798": "If i'm correct, you trained from other's weight, instead of scratch?",
    "1636741": "Similar observation here. Trained model with some augmentations and got mAP of 61. But my score was only 0.501. My current best single model with mAP = 42 scored  0.510.",
    "1636582": "confthre may shift for optimal  LB, but not than much difference.",
    "1636581": "I  met same problem.   After I bought colab pro,   I have more GPU available.  several checkpoints seems better, but \nLB just similar or worse.",
    "1636576": "I am observing similar results. I trained a yolov5l with some pre-processing and training modifications, higher mAP than public training notebooks using the same fold, but lb is only 0.389. Not really sure what is going on.",
    "1638909": ""
  }
}