{
  "id": 307869,
  "title": "YOLOX on steroids solution (6th on Public LB / 16th on Private LB) [Placeholder]",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307869",
  "author_name": "Alex Wong",
  "post_date": "2022-02-16T00:58:23.355000",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Ladies and gents, this post is a placeholder for my solution using a heavily modified YOLOX. I will post the detailed solution once I've performed some benchmarks to quantify how much each addition adds to the private LB.</p>\n<p>Solution outline:</p>\n<p><strong><em>Model:</em></strong></p>\n<p><strong>YOLOX</strong> (heavily modified)</p>\n<p><strong>Double mosaic mixup</strong></p>\n<ul>\n<li>vanilla YOLOX uses a mixup of two images: a 2x2 mosaic (affine with rotation, scaling, translation) overlaid on top of a non-augmented mixup (with scaling only)</li>\n<li>removing mixup and mosaic impacts mAP</li>\n<li>what if we add more mosaic by overlaying two 2x2 mosaic images (8 images in total)</li>\n</ul>\n<p><strong>Albumentations integration into YOLOX</strong></p>\n<p><strong>Augmentations pre-mosaic</strong></p>\n<ul>\n<li>current implementation only applies augmentation after mosaic, and not to the other mixup image</li>\n</ul>\n<p><strong>Frozen backbone</strong> for final epoch fine-tuning; only train regression head (responsible for obj / bbox outputs). Benefits of this include:</p>\n<ul>\n<li>~80% reduction of GPU memory cost when Darknet53 backbone is frozen. Allows 5x batch size increase, or training at even higher resolutions (I opted for 5x batch size increase)</li>\n<li>Prevents loss of Darknet53 backbone learning from previous epochs that utilise mosaic/mixup.</li>\n</ul>\n<p><strong><em>Post-inference processing:</em></strong></p>\n<p><strong>MemBox</strong> (my tracking solution to video object detection)<br>\nConcept: </p>\n<ul>\n<li>bounding boxes of one frame is likely to overlap that of the next frame</li>\n<li>when the video is panning, bounding boxes have predictable x/y velocities. This velocity can be used to predict the bbox of the next frame</li>\n<li>MemBox matches predicted bboxes against its velocity-based tracked bboxes<br>\nMemBox outputs:</li>\n<li>model-inference / unmatched bboxes: CONF is not changed</li>\n<li>MemBox-predicted tracked / unmatched bboxes: CONF of tracked bbox is reduced by 0.2</li>\n<li>matched bboxes: CONF is increased by 20% CONF from current frame<br>\nOutcomes:</li>\n<li>model inference of the same object across multiple frames will increase its CONF, compared with unmatched inference or tracked bboxes</li>\n</ul>\n<p><strong>WBF Ensemble using max conf</strong> (modified WBF)<br>\nIssue:</p>\n<ul>\n<li>Current WBF uses a weighted average of CONF of individual models</li>\n<li>This may bias against predictions that are missing in some sub-models<br>\nSolution:</li>\n<li>Use \"max\" CONF: this takes the maximum CONF of each sub-model.</li>\n<li>[WIP] it seems that using vanilla WBF benefits private LB whereas Max CONF benefits public LB. Will do some benchmarks to verify this</li>\n</ul>\n<p><strong>BBox scaling</strong></p>\n<ul>\n<li>Using a 0.9x bbox scaling greatly enhances public LB (but damages private LB)</li>\n<li>As discussed in my other post (<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607</a>) this must have something to do with how the COTS are annotated (i.e. tight or loose bboxes around the COTS)</li>\n</ul>\n<p>[ WIP - will post some benchmarks of public / private LB once these are available ]</p>",
  "messages": [
    {
      "id": 1692313,
      "postDate": "2022-02-16T00:58:23.357Z",
      "content": "<p>Ladies and gents, this post is a placeholder for my solution using a heavily modified YOLOX. I will post the detailed solution once I've performed some benchmarks to quantify how much each addition adds to the private LB.</p>\n<p>Solution outline:</p>\n<p><strong><em>Model:</em></strong></p>\n<p><strong>YOLOX</strong> (heavily modified)</p>\n<p><strong>Double mosaic mixup</strong></p>\n<ul>\n<li>vanilla YOLOX uses a mixup of two images: a 2x2 mosaic (affine with rotation, scaling, translation) overlaid on top of a non-augmented mixup (with scaling only)</li>\n<li>removing mixup and mosaic impacts mAP</li>\n<li>what if we add more mosaic by overlaying two 2x2 mosaic images (8 images in total)</li>\n</ul>\n<p><strong>Albumentations integration into YOLOX</strong></p>\n<p><strong>Augmentations pre-mosaic</strong></p>\n<ul>\n<li>current implementation only applies augmentation after mosaic, and not to the other mixup image</li>\n</ul>\n<p><strong>Frozen backbone</strong> for final epoch fine-tuning; only train regression head (responsible for obj / bbox outputs). Benefits of this include:</p>\n<ul>\n<li>~80% reduction of GPU memory cost when Darknet53 backbone is frozen. Allows 5x batch size increase, or training at even higher resolutions (I opted for 5x batch size increase)</li>\n<li>Prevents loss of Darknet53 backbone learning from previous epochs that utilise mosaic/mixup.</li>\n</ul>\n<p><strong><em>Post-inference processing:</em></strong></p>\n<p><strong>MemBox</strong> (my tracking solution to video object detection)<br>\nConcept: </p>\n<ul>\n<li>bounding boxes of one frame is likely to overlap that of the next frame</li>\n<li>when the video is panning, bounding boxes have predictable x/y velocities. This velocity can be used to predict the bbox of the next frame</li>\n<li>MemBox matches predicted bboxes against its velocity-based tracked bboxes<br>\nMemBox outputs:</li>\n<li>model-inference / unmatched bboxes: CONF is not changed</li>\n<li>MemBox-predicted tracked / unmatched bboxes: CONF of tracked bbox is reduced by 0.2</li>\n<li>matched bboxes: CONF is increased by 20% CONF from current frame<br>\nOutcomes:</li>\n<li>model inference of the same object across multiple frames will increase its CONF, compared with unmatched inference or tracked bboxes</li>\n</ul>\n<p><strong>WBF Ensemble using max conf</strong> (modified WBF)<br>\nIssue:</p>\n<ul>\n<li>Current WBF uses a weighted average of CONF of individual models</li>\n<li>This may bias against predictions that are missing in some sub-models<br>\nSolution:</li>\n<li>Use \"max\" CONF: this takes the maximum CONF of each sub-model.</li>\n<li>[WIP] it seems that using vanilla WBF benefits private LB whereas Max CONF benefits public LB. Will do some benchmarks to verify this</li>\n</ul>\n<p><strong>BBox scaling</strong></p>\n<ul>\n<li>Using a 0.9x bbox scaling greatly enhances public LB (but damages private LB)</li>\n<li>As discussed in my other post (<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607</a>) this must have something to do with how the COTS are annotated (i.e. tight or loose bboxes around the COTS)</li>\n</ul>\n<p>[ WIP - will post some benchmarks of public / private LB once these are available ]</p>",
      "rawMarkdown": "Ladies and gents, this post is a placeholder for my solution using a heavily modified YOLOX. I will post the detailed solution once I've performed some benchmarks to quantify how much each addition adds to the private LB.\n\nSolution outline:\n\n***Model:***\n\n**YOLOX** (heavily modified)\n\n**Double mosaic mixup**\n- vanilla YOLOX uses a mixup of two images: a 2x2 mosaic (affine with rotation, scaling, translation) overlaid on top of a non-augmented mixup (with scaling only)\n- removing mixup and mosaic impacts mAP\n- what if we add more mosaic by overlaying two 2x2 mosaic images (8 images in total)\n\n**Albumentations integration into YOLOX**\n\n**Augmentations pre-mosaic**\n- current implementation only applies augmentation after mosaic, and not to the other mixup image\n\n**Frozen backbone** for final epoch fine-tuning; only train regression head (responsible for obj / bbox outputs). Benefits of this include:\n- ~80% reduction of GPU memory cost when Darknet53 backbone is frozen. Allows 5x batch size increase, or training at even higher resolutions (I opted for 5x batch size increase)\n- Prevents loss of Darknet53 backbone learning from previous epochs that utilise mosaic/mixup.\n\n***Post-inference processing:***\n\n**MemBox** (my tracking solution to video object detection)\nConcept: \n- bounding boxes of one frame is likely to overlap that of the next frame\n- when the video is panning, bounding boxes have predictable x/y velocities. This velocity can be used to predict the bbox of the next frame\n- MemBox matches predicted bboxes against its velocity-based tracked bboxes\nMemBox outputs:\n- model-inference / unmatched bboxes: CONF is not changed\n- MemBox-predicted tracked / unmatched bboxes: CONF of tracked bbox is reduced by 0.2\n- matched bboxes: CONF is increased by 20% CONF from current frame\nOutcomes:\n- model inference of the same object across multiple frames will increase its CONF, compared with unmatched inference or tracked bboxes\n\n**WBF Ensemble using max conf** (modified WBF)\nIssue:\n- Current WBF uses a weighted average of CONF of individual models\n- This may bias against predictions that are missing in some sub-models\nSolution:\n- Use \"max\" CONF: this takes the maximum CONF of each sub-model.\n- [WIP] it seems that using vanilla WBF benefits private LB whereas Max CONF benefits public LB. Will do some benchmarks to verify this\n\n**BBox scaling**\n- Using a 0.9x bbox scaling greatly enhances public LB (but damages private LB)\n- As discussed in my other post (https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607) this must have something to do with how the COTS are annotated (i.e. tight or loose bboxes around the COTS)\n\n[ WIP - will post some benchmarks of public / private LB once these are available ]",
      "votes": 9
    },
    {
      "id": 1694943,
      "postDate": "2022-02-17T21:02:42.020Z",
      "content": "<p>14th on the LB already ;)</p>",
      "rawMarkdown": "14th on the LB already ;)",
      "votes": 1
    },
    {
      "id": 1692552,
      "postDate": "2022-02-16T05:42:27.493Z",
      "content": "<p>Great job Alex :) hope you will share MemBox own tracker  :). Really sorry for you that you are 1st runner up to gold medal, you deserved it.</p>",
      "rawMarkdown": "Great job Alex :) hope you will share MemBox own tracker  :). Really sorry for you that you are 1st runner up to gold medal, you deserved it.",
      "votes": 1,
      "replies": [
        {
          "id": 1694921,
          "postDate": "2022-02-17T20:26:53.577Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1695131,
          "postDate": "2022-02-18T01:18:40.273Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> . I intend to write up MemBox as a Github repo so it will be useful to others in future Kaggle competitions.</p>",
          "rawMarkdown": "Thanks @lukaszborecki . I intend to write up MemBox as a Github repo so it will be useful to others in future Kaggle competitions."
        }
      ]
    },
    {
      "id": 1692546,
      "postDate": "2022-02-16T05:33:22.607Z",
      "content": "<p>waiting for your update.</p>\n<p>I change a  little bit multi-scale with yolox,  It is too slow to train with colab , so I gave up trying.</p>",
      "rawMarkdown": "waiting for your update.\n\nI change a  little bit multi-scale with yolox,  It is too slow to train with colab , so I gave up trying."
    }
  ],
  "comments": [
    {
      "id": 1694943,
      "author_name": "danjafish",
      "author_url": "",
      "post_date": "2022-02-17T21:02:42.020000",
      "content": "<p>14th on the LB already ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1692552,
      "author_name": "Lukasz Borecki",
      "author_url": "",
      "post_date": "2022-02-16T05:42:27.493000",
      "content": "<p>Great job Alex :) hope you will share MemBox own tracker  :). Really sorry for you that you are 1st runner up to gold medal, you deserved it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1694921,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-17T20:26:53.577000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1695131,
          "author_name": "Alex Wong",
          "author_url": "",
          "post_date": "2022-02-18T01:18:40.273000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> . I intend to write up MemBox as a Github repo so it will be useful to others in future Kaggle competitions.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1692546,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2022-02-16T05:33:22.607000",
      "content": "<p>waiting for your update.</p>\n<p>I change a  little bit multi-scale with yolox,  It is too slow to train with colab , so I gave up trying.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1692313": "Ladies and gents, this post is a placeholder for my solution using a heavily modified YOLOX. I will post the detailed solution once I've performed some benchmarks to quantify how much each addition adds to the private LB.\n\nSolution outline:\n\n***Model:***\n\n**YOLOX** (heavily modified)\n\n**Double mosaic mixup**\n- vanilla YOLOX uses a mixup of two images: a 2x2 mosaic (affine with rotation, scaling, translation) overlaid on top of a non-augmented mixup (with scaling only)\n- removing mixup and mosaic impacts mAP\n- what if we add more mosaic by overlaying two 2x2 mosaic images (8 images in total)\n\n**Albumentations integration into YOLOX**\n\n**Augmentations pre-mosaic**\n- current implementation only applies augmentation after mosaic, and not to the other mixup image\n\n**Frozen backbone** for final epoch fine-tuning; only train regression head (responsible for obj / bbox outputs). Benefits of this include:\n- ~80% reduction of GPU memory cost when Darknet53 backbone is frozen. Allows 5x batch size increase, or training at even higher resolutions (I opted for 5x batch size increase)\n- Prevents loss of Darknet53 backbone learning from previous epochs that utilise mosaic/mixup.\n\n***Post-inference processing:***\n\n**MemBox** (my tracking solution to video object detection)\nConcept: \n- bounding boxes of one frame is likely to overlap that of the next frame\n- when the video is panning, bounding boxes have predictable x/y velocities. This velocity can be used to predict the bbox of the next frame\n- MemBox matches predicted bboxes against its velocity-based tracked bboxes\nMemBox outputs:\n- model-inference / unmatched bboxes: CONF is not changed\n- MemBox-predicted tracked / unmatched bboxes: CONF of tracked bbox is reduced by 0.2\n- matched bboxes: CONF is increased by 20% CONF from current frame\nOutcomes:\n- model inference of the same object across multiple frames will increase its CONF, compared with unmatched inference or tracked bboxes\n\n**WBF Ensemble using max conf** (modified WBF)\nIssue:\n- Current WBF uses a weighted average of CONF of individual models\n- This may bias against predictions that are missing in some sub-models\nSolution:\n- Use \"max\" CONF: this takes the maximum CONF of each sub-model.\n- [WIP] it seems that using vanilla WBF benefits private LB whereas Max CONF benefits public LB. Will do some benchmarks to verify this\n\n**BBox scaling**\n- Using a 0.9x bbox scaling greatly enhances public LB (but damages private LB)\n- As discussed in my other post (https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/307607) this must have something to do with how the COTS are annotated (i.e. tight or loose bboxes around the COTS)\n\n[ WIP - will post some benchmarks of public / private LB once these are available ]",
    "1694943": "14th on the LB already ;)",
    "1692552": "Great job Alex :) hope you will share MemBox own tracker  :). Really sorry for you that you are 1st runner up to gold medal, you deserved it.",
    "1692546": "waiting for your update.\n\nI change a  little bit multi-scale with yolox,  It is too slow to train with colab , so I gave up trying."
  }
}