{
  "id": 307622,
  "title": "22th place solution - Train4Ever",
  "url": "/competitions/tensorflow-great-barrier-reef/writeups/train4ever-22th-place-solution-train4ever",
  "author_name": "",
  "post_date": "2022-02-18T02:15:02.043Z",
  "votes": 30,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thanks to the hosts for providing an interesting competition. Although not ending up in Gold zone, we would like to share our solution. </p>\n<h3>Our solution consists of 3 stages:</h3>\n<ul>\n<li>Detection</li>\n<li>Tracking</li>\n<li>Classification post processing</li>\n</ul>\n<h3>1. Detection</h3>\n<ul>\n<li>Using some popular models like everyone does. We used CascadeRCNN ResNeSt200, YoloV5, YoloX, YoloR. </li>\n<li>Ensembling: WBF</li>\n</ul>\n<h3>2. Tracking:</h3>\n<ul>\n<li>We did not invest enough time to find a better tracking scheme than the one from the public notebook. As a result, we used that Norfair without any significant modification. </li>\n<li>After Norfair is applied, it is certain that the predictions contain so many FPs though it does increase the recall. It is the work of stage 3 that reduce the effect of the problem. </li>\n</ul>\n<h3>3. Classification:</h3>\n<ul>\n<li>This is important in our pipeline because we do not have too strong detectors and tracking method.</li>\n<li>Flow: detectors predict bboxes -&gt; crop -&gt; classification -&gt; blending</li>\n<li>We have 5 folds. In each fold, each of the 4 detector makes prediction on both train and validation set. We focus on hard false positive (high detection score but wrong) and hard false negative (low detection score but true).  The predictions on train set from all the detectors are concatenated and reduce overlapping by NMS, then used to train the classification models. The same method is used to produce the validation set for the classification models.</li>\n<li>We also crawled some COTS data from the Internet with Creative Commons license to generalize the model more.</li>\n<li>Models: EfficientnetB7, Eca Nfnet L0</li>\n<li>Im size: 128</li>\n<li>Augmentations: heavy augmentations. Cutmix only when training EffB7 </li>\n<li>Training notebooks: <a href=\"https://drive.google.com/file/d/1uD4Wj4pxfScm-NnXuf9jSVXCru-XTGJB/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1uD4Wj4pxfScm-NnXuf9jSVXCru-XTGJB/view?usp=sharing</a> (Eca nfnet l0) , <a href=\"https://drive.google.com/file/d/1pQuk6z5Vl7sfg7v36NVjk3WpbYYglS91/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1pQuk6z5Vl7sfg7v36NVjk3WpbYYglS91/view?usp=sharing</a> (EffB7)</li>\n<li>2 model type x 5 folds = 10 models classification for final submission</li>\n<li>Blending method: Use geometric mean between the classification probability score and the bbox conf score from the detectors.</li>\n<li>We estimated this boosted about 6% on the private LB in our final 0.706 solution. </li>\n</ul>",
  "messages": [
    {
      "id": "1690554",
      "postDate": "02/15/2022 02:11:48",
      "content": "<p>Thanks to the hosts for providing an interesting competition. Although not ending up in Gold zone, we would like to share our solution. </p>\n<h3>Our solution consists of 3 stages:</h3>\n<ul>\n<li>Detection</li>\n<li>Tracking</li>\n<li>Classification post processing</li>\n</ul>\n<h3>1. Detection</h3>\n<ul>\n<li>Using some popular models like everyone does. We used CascadeRCNN ResNeSt200, YoloV5, YoloX, YoloR. </li>\n<li>Ensembling: WBF</li>\n</ul>\n<h3>2. Tracking:</h3>\n<ul>\n<li>We did not invest enough time to find a better tracking scheme than the one from the public notebook. As a result, we used that Norfair without any significant modification. </li>\n<li>After Norfair is applied, it is certain that the predictions contain so many FPs though it does increase the recall. It is the work of stage 3 that reduce the effect of the problem. </li>\n</ul>\n<h3>3. Classification:</h3>\n<ul>\n<li>This is important in our pipeline because we do not have too strong detectors and tracking method.</li>\n<li>Flow: detectors predict bboxes -&gt; crop -&gt; classification -&gt; blending</li>\n<li>We have 5 folds. In each fold, each of the 4 detector makes prediction on both train and validation set. We focus on hard false positive (high detection score but wrong) and hard false negative (low detection score but true).  The predictions on train set from all the detectors are concatenated and reduce overlapping by NMS, then used to train the classification models. The same method is used to produce the validation set for the classification models.</li>\n<li>We also crawled some COTS data from the Internet with Creative Commons license to generalize the model more.</li>\n<li>Models: EfficientnetB7, Eca Nfnet L0</li>\n<li>Im size: 128</li>\n<li>Augmentations: heavy augmentations. Cutmix only when training EffB7 </li>\n<li>Training notebooks: <a href=\"https://drive.google.com/file/d/1uD4Wj4pxfScm-NnXuf9jSVXCru-XTGJB/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1uD4Wj4pxfScm-NnXuf9jSVXCru-XTGJB/view?usp=sharing</a> (Eca nfnet l0) , <a href=\"https://drive.google.com/file/d/1pQuk6z5Vl7sfg7v36NVjk3WpbYYglS91/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1pQuk6z5Vl7sfg7v36NVjk3WpbYYglS91/view?usp=sharing</a> (EffB7)</li>\n<li>2 model type x 5 folds = 10 models classification for final submission</li>\n<li>Blending method: Use geometric mean between the classification probability score and the bbox conf score from the detectors.</li>\n<li>We estimated this boosted about 6% on the private LB in our final 0.706 solution. </li>\n</ul>",
      "rawMarkdown": "Thanks to the hosts for providing an interesting competition. Although not ending up in Gold zone, we would like to share our solution. \n\n### Our solution consists of 3 stages:\n- Detection\n- Tracking\n- Classification post processing\n\n### 1. Detection\n- Using some popular models like everyone does. We used CascadeRCNN ResNeSt200, YoloV5, YoloX, YoloR. \n- Ensembling: WBF\n\n### 2. Tracking:\n- We did not invest enough time to find a better tracking scheme than the one from the public notebook. As a result, we used that Norfair without any significant modification. \n- After Norfair is applied, it is certain that the predictions contain so many FPs though it does increase the recall. It is the work of stage 3 that reduce the effect of the problem. \n\n### 3. Classification:\n- This is important in our pipeline because we do not have too strong detectors and tracking method.\n- Flow: detectors predict bboxes -> crop -> classification -> blending\n- We have 5 folds. In each fold, each of the 4 detector makes prediction on both train and validation set. We focus on hard false positive (high detection score but wrong) and hard false negative (low detection score but true).  The predictions on train set from all the detectors are concatenated and reduce overlapping by NMS, then used to train the classification models. The same method is used to produce the validation set for the classification models.\n- We also crawled some COTS data from the Internet with Creative Commons license to generalize the model more.\n- Models: EfficientnetB7, Eca Nfnet L0\n- Im size: 128\n- Augmentations: heavy augmentations. Cutmix only when training EffB7 \n- Training notebooks: https://drive.google.com/file/d/1uD4Wj4pxfScm-NnXuf9jSVXCru-XTGJB/view?usp=sharing (Eca nfnet l0) , https://drive.google.com/file/d/1pQuk6z5Vl7sfg7v36NVjk3WpbYYglS91/view?usp=sharing (EffB7)\n- 2 model type x 5 folds = 10 models classification for final submission\n- Blending method: Use geometric mean between the classification probability score and the bbox conf score from the detectors.\n- We estimated this boosted about 6% on the private LB in our final 0.706 solution.",
      "votes": null
    },
    {
      "id": "1690568",
      "postDate": "02/15/2022 02:25:41",
      "content": "<ol>\n<li>What is the score of your metric for classifier when training? <br>\n<strong>Update</strong>: I saw them in your notebook, the loss is still non-zero, but It may effectively reduce FP.</li>\n<li>Why not EffNetV2_m?</li>\n</ol>\n<p>Congrats a Nam 💯</p>",
      "rawMarkdown": "1. What is the score of your metric for classifier when training? \n**Update**: I saw them in your notebook, the loss is still non-zero, but It may effectively reduce FP.\n2. Why not EffNetV2_m?\n\nCongrats a Nam 💯",
      "votes": null
    },
    {
      "id": "1690581",
      "postDate": "02/15/2022 02:38:32",
      "content": "<ol>\n<li>We have 2 stage validation for clf.</li>\n</ol>\n<ul>\n<li>First, using AUC, accuracy, sklearn F2 score to evaluate on cropped images.</li>\n<li>Then, blending with detection score and checking the complete pipeline with competition F2 metric</li>\n</ul>\n<ol>\n<li>Too big, I think it is not necessary.</li>\n</ol>",
      "rawMarkdown": "1. We have 2 stage validation for clf.\n- First, using AUC, accuracy, sklearn F2 score to evaluate on cropped images.\n- Then, blending with detection score and checking the complete pipeline with competition F2 metric\n\n\n2. Too big, I think it is not necessary.",
      "votes": null
    },
    {
      "id": "1690612",
      "postDate": "02/15/2022 03:07:28",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> 🎊<br>\nBTW, What data you used for classification? Is this the <a href=\"https://www.kaggle.com/alexteboul/binary-cropped-crown-of-thorns-dataset\" target=\"_blank\">dataset</a> or created by yourself by cropping? taking features from the detectors and adding classification layers afterward. could that be useful? I was hoping someone would do the classification that way. </p>",
      "rawMarkdown": "Congratulations @namgalielei 🎊\nBTW, What data you used for classification? Is this the [dataset](https://www.kaggle.com/alexteboul/binary-cropped-crown-of-thorns-dataset) or created by yourself by cropping? taking features from the detectors and adding classification layers afterward. could that be useful? I was hoping someone would do the classification that way.",
      "votes": null
    },
    {
      "id": "1690622",
      "postDate": "02/15/2022 03:14:35",
      "content": "<p>Congratulations anh Nam và team nhé. </p>",
      "rawMarkdown": "Congratulations anh Nam và team nhé.",
      "votes": null
    },
    {
      "id": "1690627",
      "postDate": "02/15/2022 03:19:27",
      "content": "<p>I did create the crop data myself. And the flow is demonstrated in my attached notebooks. I think it is better to train a separated classification model rather than utilize the feature maps from the detectors (the feature maps from the detectors already carry so many tasks, regression, objectness classify, multi scale training, etc)</p>",
      "rawMarkdown": "I did create the crop data myself. And the flow is demonstrated in my attached notebooks. I think it is better to train a separated classification model rather than utilize the feature maps from the detectors (the feature maps from the detectors already carry so many tasks, regression, objectness classify, multi scale training, etc)",
      "votes": null
    },
    {
      "id": "1690888",
      "postDate": "02/15/2022 06:30:15",
      "content": "<p>great! it's comprehensive.</p>",
      "rawMarkdown": "great! it's comprehensive.",
      "votes": null
    },
    {
      "id": "1690900",
      "postDate": "02/15/2022 06:40:12",
      "content": "<p>Oh I see, understood, thank you for replying ! </p>",
      "rawMarkdown": "Oh I see, understood, thank you for replying !",
      "votes": null
    },
    {
      "id": "1691067",
      "postDate": "02/15/2022 08:12:31",
      "content": "<p>Can you share, what was the accuracy of classifier itself?</p>",
      "rawMarkdown": "Can you share, what was the accuracy of classifier itself?",
      "votes": null
    },
    {
      "id": "1691080",
      "postDate": "02/15/2022 08:21:41",
      "content": "<p>Viewing the output log of my attached notebooks, you'll see</p>",
      "rawMarkdown": "Viewing the output log of my attached notebooks, you'll see",
      "votes": null
    },
    {
      "id": "1691461",
      "postDate": "02/15/2022 12:23:43",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> </p>",
      "rawMarkdown": "Congratulations! @namgalielei",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1690568,
      "author_name": "aengusng",
      "author_url": "",
      "post_date": "02/15/2022 02:25:41",
      "content": "<ol>\n<li>What is the score of your metric for classifier when training? <br>\n<strong>Update</strong>: I saw them in your notebook, the loss is still non-zero, but It may effectively reduce FP.</li>\n<li>Why not EffNetV2_m?</li>\n</ol>\n<p>Congrats a Nam 💯</p>",
      "votes": null,
      "replies": [
        {
          "id": 1690581,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "02/15/2022 02:38:32",
          "content": "<ol>\n<li>We have 2 stage validation for clf.</li>\n</ol>\n<ul>\n<li>First, using AUC, accuracy, sklearn F2 score to evaluate on cropped images.</li>\n<li>Then, blending with detection score and checking the complete pipeline with competition F2 metric</li>\n</ul>\n<ol>\n<li>Too big, I think it is not necessary.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690888,
          "author_name": "aengusng",
          "author_url": "",
          "post_date": "02/15/2022 06:30:15",
          "content": "<p>great! it's comprehensive.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691067,
          "author_name": "bakeryproducts",
          "author_url": "",
          "post_date": "02/15/2022 08:12:31",
          "content": "<p>Can you share, what was the accuracy of classifier itself?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691080,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "02/15/2022 08:21:41",
          "content": "<p>Viewing the output log of my attached notebooks, you'll see</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690612,
      "author_name": "soumya9977",
      "author_url": "",
      "post_date": "02/15/2022 03:07:28",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> 🎊<br>\nBTW, What data you used for classification? Is this the <a href=\"https://www.kaggle.com/alexteboul/binary-cropped-crown-of-thorns-dataset\" target=\"_blank\">dataset</a> or created by yourself by cropping? taking features from the detectors and adding classification layers afterward. could that be useful? I was hoping someone would do the classification that way. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1690627,
          "author_name": "namgalielei",
          "author_url": "",
          "post_date": "02/15/2022 03:19:27",
          "content": "<p>I did create the crop data myself. And the flow is demonstrated in my attached notebooks. I think it is better to train a separated classification model rather than utilize the feature maps from the detectors (the feature maps from the detectors already carry so many tasks, regression, objectness classify, multi scale training, etc)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690900,
          "author_name": "soumya9977",
          "author_url": "",
          "post_date": "02/15/2022 06:40:12",
          "content": "<p>Oh I see, understood, thank you for replying ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690622,
      "author_name": "locbaop",
      "author_url": "",
      "post_date": "02/15/2022 03:14:35",
      "content": "<p>Congratulations anh Nam và team nhé. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1691461,
      "author_name": "deepernet",
      "author_url": "",
      "post_date": "02/15/2022 12:23:43",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/namgalielei\" target=\"_blank\">@namgalielei</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1690554": "Thanks to the hosts for providing an interesting competition. Although not ending up in Gold zone, we would like to share our solution. \n\n### Our solution consists of 3 stages:\n- Detection\n- Tracking\n- Classification post processing\n\n### 1. Detection\n- Using some popular models like everyone does. We used CascadeRCNN ResNeSt200, YoloV5, YoloX, YoloR. \n- Ensembling: WBF\n\n### 2. Tracking:\n- We did not invest enough time to find a better tracking scheme than the one from the public notebook. As a result, we used that Norfair without any significant modification. \n- After Norfair is applied, it is certain that the predictions contain so many FPs though it does increase the recall. It is the work of stage 3 that reduce the effect of the problem. \n\n### 3. Classification:\n- This is important in our pipeline because we do not have too strong detectors and tracking method.\n- Flow: detectors predict bboxes -> crop -> classification -> blending\n- We have 5 folds. In each fold, each of the 4 detector makes prediction on both train and validation set. We focus on hard false positive (high detection score but wrong) and hard false negative (low detection score but true).  The predictions on train set from all the detectors are concatenated and reduce overlapping by NMS, then used to train the classification models. The same method is used to produce the validation set for the classification models.\n- We also crawled some COTS data from the Internet with Creative Commons license to generalize the model more.\n- Models: EfficientnetB7, Eca Nfnet L0\n- Im size: 128\n- Augmentations: heavy augmentations. Cutmix only when training EffB7 \n- Training notebooks: https://drive.google.com/file/d/1uD4Wj4pxfScm-NnXuf9jSVXCru-XTGJB/view?usp=sharing (Eca nfnet l0) , https://drive.google.com/file/d/1pQuk6z5Vl7sfg7v36NVjk3WpbYYglS91/view?usp=sharing (EffB7)\n- 2 model type x 5 folds = 10 models classification for final submission\n- Blending method: Use geometric mean between the classification probability score and the bbox conf score from the detectors.\n- We estimated this boosted about 6% on the private LB in our final 0.706 solution.",
    "1690568": "1. What is the score of your metric for classifier when training? \n**Update**: I saw them in your notebook, the loss is still non-zero, but It may effectively reduce FP.\n2. Why not EffNetV2_m?\n\nCongrats a Nam 💯",
    "1690581": "1. We have 2 stage validation for clf.\n- First, using AUC, accuracy, sklearn F2 score to evaluate on cropped images.\n- Then, blending with detection score and checking the complete pipeline with competition F2 metric\n\n\n2. Too big, I think it is not necessary.",
    "1690612": "Congratulations @namgalielei 🎊\nBTW, What data you used for classification? Is this the [dataset](https://www.kaggle.com/alexteboul/binary-cropped-crown-of-thorns-dataset) or created by yourself by cropping? taking features from the detectors and adding classification layers afterward. could that be useful? I was hoping someone would do the classification that way.",
    "1690622": "Congratulations anh Nam và team nhé.",
    "1690627": "I did create the crop data myself. And the flow is demonstrated in my attached notebooks. I think it is better to train a separated classification model rather than utilize the feature maps from the detectors (the feature maps from the detectors already carry so many tasks, regression, objectness classify, multi scale training, etc)",
    "1690888": "great! it's comprehensive.",
    "1690900": "Oh I see, understood, thank you for replying !",
    "1691067": "Can you share, what was the accuracy of classifier itself?",
    "1691080": "Viewing the output log of my attached notebooks, you'll see",
    "1691461": "Congratulations! @namgalielei"
  },
  "source": "meta"
}