{
  "id": 319145,
  "title": "1st Place Solution",
  "url": "/competitions/ultra-mnist/discussion/319145",
  "author_name": "latent4g",
  "post_date": "2022-04-15T16:52:19.662000",
  "votes": 34,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congrats to all the winners. Code is available here : <a href=\"https://github.com/4g/umnist/tree/master\" target=\"_blank\">https://github.com/4g/umnist/tree/master</a> along with trained models. You can use these to regenerate the results. Instructions to train from scratch are also attached in the readme. Feel free to raise any issues if something doesn't work.    </p>\n<h2>Idea</h2>\n<ul>\n<li>Generate synthetic data that looks like the competition data. </li>\n<li>Use a small object detector like yolov5s and reclassify all labels using a larger network (EffV2B1)</li>\n</ul>\n<p>Input images are large. Any network consuming these directly will be slow. Yolo small gets good box-obj accuracy but low class accuracy. </p>\n<p>Objects in image are not complex. Even at a resolution of 128x128 these can be classified with &gt; 99.9% accuracy using a sufficiently large network. So I train a classifier on digits cropped from synthetic training set and reclassify yolo results with this high accuracy classifier. Yolo alone gives &gt; 95%, adding a classifier gives &gt; 98%. Rest of accuracy comes from post processing magic.  </p>\n<h2>Data generation</h2>\n<h4>Detection</h4>\n<ul>\n<li>2560x2560 input images. Smallest digit size is 6x6. </li>\n<li>Generate checkerboards with boxes, circles and triangles. Apply augmentations like perspective, rotation, zoom to make it look more like competition dataset. </li>\n<li>Generate locations where digits should be put. Make sure digits don't overlap. </li>\n<li>Ensure all 70000 instances from mnist are covered. </li>\n<li>Uniformly distribute sizes of digits. </li>\n<li>Generate empty images with no digits to reduce background confusion. </li>\n<li>Train 34000 with 4000 empty images : val 2000 with 0 empty images.  </li>\n</ul>\n<h4>Classification</h4>\n<ul>\n<li>128x128 images. </li>\n<li>Digit crops from synthetic data. </li>\n<li>Add raw mnist digits </li>\n<li>Train : 230,000 (160k synthetic + 70k raw mnist)</li>\n</ul>\n<h2>Training</h2>\n<h4>Detector</h4>\n<ul>\n<li>Yolov5 small with batch size 2 (max that is supported at this resolution on my gpu)</li>\n<li>light augmentation, because harsh augmentations take digits too much out of the image : rotation, scale, translate, hsv, mosaic. </li>\n<li>SGD 100 epochs. Stop when the object loss reaches .008. Cls loss doesn't matter. Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence. </li>\n</ul>\n<h4>Classifier</h4>\n<ul>\n<li>Effnetv2b1 with 128x128x3 input. Use pretrained imagenet weights for faster convergence. </li>\n<li>augmentations : Shift scale, perspective, invert img, blur, hsv, brightness/contrast. Keep augmentations light as competition data doesn't look augmented. </li>\n<li>Reduce lr every 50 epochs (3e-4, 1e-4, 3e-5, 1e-5). Train for 200. </li>\n</ul>\n<h2>Inference</h2>\n<h4>Detection</h4>\n<ul>\n<li>2560x2560</li>\n<li>Try to increase recall of yolo detector. </li>\n<li>iou_threshold = 0.1. Digits are not intersecting, so any intersecting outputs are wrong </li>\n<li>conf_threshold = 0.1 . To increase recall. </li>\n<li>class agnostic nms</li>\n<li>Maximum 5 digits allowed per image, max_det = 5</li>\n<li>To measure the accuracy, use competition training data. </li>\n<li>&gt; 95%</li>\n</ul>\n<h4>Classification</h4>\n<ul>\n<li>Crop digits detected by yolo. Some of them have skewed aspect ratio, pick the larger side and crop a square. Classifier is trained with square images. Changing the aspect ratio reduces its accuracy. </li>\n<li>TTA: Invertimg, average the predicted probabilties</li>\n<li>&gt; 98%</li>\n</ul>\n<h4>Post processing</h4>\n<ul>\n<li>Some of the large digits have artifacts, which get detected as other digits. IOU threshold cannot remove them because these are very small as compared to the larger digit. Remove any digit that lies completely inside another</li>\n<li>Some of the background is also classified as digits. But the classifier has low confidence for these. Remove any digits where classifier prob &lt; 0.45</li>\n<li>&gt; 99.1% </li>\n</ul>\n<h2>Debugging</h2>\n<p>Since the competition training data is not seen by models, it can be used for validation.  Looking at images that are getting incorrect sum, gives a good idea of augmentations and filters to add.  </p>\n<h2>Tidbits</h2>\n<ul>\n<li>Accuracy on competition train set closely matches the test accuracy</li>\n<li>Yolo took 3 days to train due to large image size</li>\n<li>This solution doesn't help the original problem which authors had proposed of processing large images via networks in any way</li>\n<li>In the beginning I was working towards innovative track using a pretrained unet + blazeblock regressor but then the dataset changed.</li>\n</ul>\n<p>Thanks to organizers for this cool competition. <br>\nAnd thanks Namtran. Your continuous uploads were a big encouragement. </p>",
  "messages": [
    {
      "id": 1756610,
      "postDate": "2022-04-15T16:52:19.663Z",
      "content": "<p>Congrats to all the winners. Code is available here : <a href=\"https://github.com/4g/umnist/tree/master\" target=\"_blank\">https://github.com/4g/umnist/tree/master</a> along with trained models. You can use these to regenerate the results. Instructions to train from scratch are also attached in the readme. Feel free to raise any issues if something doesn't work.    </p>\n<h2>Idea</h2>\n<ul>\n<li>Generate synthetic data that looks like the competition data. </li>\n<li>Use a small object detector like yolov5s and reclassify all labels using a larger network (EffV2B1)</li>\n</ul>\n<p>Input images are large. Any network consuming these directly will be slow. Yolo small gets good box-obj accuracy but low class accuracy. </p>\n<p>Objects in image are not complex. Even at a resolution of 128x128 these can be classified with &gt; 99.9% accuracy using a sufficiently large network. So I train a classifier on digits cropped from synthetic training set and reclassify yolo results with this high accuracy classifier. Yolo alone gives &gt; 95%, adding a classifier gives &gt; 98%. Rest of accuracy comes from post processing magic.  </p>\n<h2>Data generation</h2>\n<h4>Detection</h4>\n<ul>\n<li>2560x2560 input images. Smallest digit size is 6x6. </li>\n<li>Generate checkerboards with boxes, circles and triangles. Apply augmentations like perspective, rotation, zoom to make it look more like competition dataset. </li>\n<li>Generate locations where digits should be put. Make sure digits don't overlap. </li>\n<li>Ensure all 70000 instances from mnist are covered. </li>\n<li>Uniformly distribute sizes of digits. </li>\n<li>Generate empty images with no digits to reduce background confusion. </li>\n<li>Train 34000 with 4000 empty images : val 2000 with 0 empty images.  </li>\n</ul>\n<h4>Classification</h4>\n<ul>\n<li>128x128 images. </li>\n<li>Digit crops from synthetic data. </li>\n<li>Add raw mnist digits </li>\n<li>Train : 230,000 (160k synthetic + 70k raw mnist)</li>\n</ul>\n<h2>Training</h2>\n<h4>Detector</h4>\n<ul>\n<li>Yolov5 small with batch size 2 (max that is supported at this resolution on my gpu)</li>\n<li>light augmentation, because harsh augmentations take digits too much out of the image : rotation, scale, translate, hsv, mosaic. </li>\n<li>SGD 100 epochs. Stop when the object loss reaches .008. Cls loss doesn't matter. Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence. </li>\n</ul>\n<h4>Classifier</h4>\n<ul>\n<li>Effnetv2b1 with 128x128x3 input. Use pretrained imagenet weights for faster convergence. </li>\n<li>augmentations : Shift scale, perspective, invert img, blur, hsv, brightness/contrast. Keep augmentations light as competition data doesn't look augmented. </li>\n<li>Reduce lr every 50 epochs (3e-4, 1e-4, 3e-5, 1e-5). Train for 200. </li>\n</ul>\n<h2>Inference</h2>\n<h4>Detection</h4>\n<ul>\n<li>2560x2560</li>\n<li>Try to increase recall of yolo detector. </li>\n<li>iou_threshold = 0.1. Digits are not intersecting, so any intersecting outputs are wrong </li>\n<li>conf_threshold = 0.1 . To increase recall. </li>\n<li>class agnostic nms</li>\n<li>Maximum 5 digits allowed per image, max_det = 5</li>\n<li>To measure the accuracy, use competition training data. </li>\n<li>&gt; 95%</li>\n</ul>\n<h4>Classification</h4>\n<ul>\n<li>Crop digits detected by yolo. Some of them have skewed aspect ratio, pick the larger side and crop a square. Classifier is trained with square images. Changing the aspect ratio reduces its accuracy. </li>\n<li>TTA: Invertimg, average the predicted probabilties</li>\n<li>&gt; 98%</li>\n</ul>\n<h4>Post processing</h4>\n<ul>\n<li>Some of the large digits have artifacts, which get detected as other digits. IOU threshold cannot remove them because these are very small as compared to the larger digit. Remove any digit that lies completely inside another</li>\n<li>Some of the background is also classified as digits. But the classifier has low confidence for these. Remove any digits where classifier prob &lt; 0.45</li>\n<li>&gt; 99.1% </li>\n</ul>\n<h2>Debugging</h2>\n<p>Since the competition training data is not seen by models, it can be used for validation.  Looking at images that are getting incorrect sum, gives a good idea of augmentations and filters to add.  </p>\n<h2>Tidbits</h2>\n<ul>\n<li>Accuracy on competition train set closely matches the test accuracy</li>\n<li>Yolo took 3 days to train due to large image size</li>\n<li>This solution doesn't help the original problem which authors had proposed of processing large images via networks in any way</li>\n<li>In the beginning I was working towards innovative track using a pretrained unet + blazeblock regressor but then the dataset changed.</li>\n</ul>\n<p>Thanks to organizers for this cool competition. <br>\nAnd thanks Namtran. Your continuous uploads were a big encouragement. </p>",
      "rawMarkdown": "Congrats to all the winners. Code is available here : https://github.com/4g/umnist/tree/master along with trained models. You can use these to regenerate the results. Instructions to train from scratch are also attached in the readme. Feel free to raise any issues if something doesn't work.    \n\n## Idea\n- Generate synthetic data that looks like the competition data. \n- Use a small object detector like yolov5s and reclassify all labels using a larger network (EffV2B1)\n\nInput images are large. Any network consuming these directly will be slow. Yolo small gets good box-obj accuracy but low class accuracy. \n\nObjects in image are not complex. Even at a resolution of 128x128 these can be classified with > 99.9% accuracy using a sufficiently large network. So I train a classifier on digits cropped from synthetic training set and reclassify yolo results with this high accuracy classifier. Yolo alone gives > 95%, adding a classifier gives > 98%. Rest of accuracy comes from post processing magic.  \n\n## Data generation\n#### Detection\n- 2560x2560 input images. Smallest digit size is 6x6. \n- Generate checkerboards with boxes, circles and triangles. Apply augmentations like perspective, rotation, zoom to make it look more like competition dataset. \n- Generate locations where digits should be put. Make sure digits don't overlap. \n- Ensure all 70000 instances from mnist are covered. \n- Uniformly distribute sizes of digits. \n- Generate empty images with no digits to reduce background confusion. \n- Train 34000 with 4000 empty images : val 2000 with 0 empty images.  \n\n#### Classification\n- 128x128 images. \n- Digit crops from synthetic data. \n- Add raw mnist digits \n- Train : 230,000 (160k synthetic + 70k raw mnist)\n\n## Training\n#### Detector\n- Yolov5 small with batch size 2 (max that is supported at this resolution on my gpu)\n- light augmentation, because harsh augmentations take digits too much out of the image : rotation, scale, translate, hsv, mosaic. \n- SGD 100 epochs. Stop when the object loss reaches .008. Cls loss doesn't matter. Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence. \n\n#### Classifier\n- Effnetv2b1 with 128x128x3 input. Use pretrained imagenet weights for faster convergence. \n- augmentations : Shift scale, perspective, invert img, blur, hsv, brightness/contrast. Keep augmentations light as competition data doesn't look augmented. \n- Reduce lr every 50 epochs (3e-4, 1e-4, 3e-5, 1e-5). Train for 200. \n\n## Inference \n#### Detection\n- 2560x2560\n- Try to increase recall of yolo detector. \n- iou_threshold = 0.1. Digits are not intersecting, so any intersecting outputs are wrong \n- conf_threshold = 0.1 . To increase recall. \n- class agnostic nms\n- Maximum 5 digits allowed per image, max_det = 5\n- To measure the accuracy, use competition training data. \n- > 95%\n\n#### Classification\n- Crop digits detected by yolo. Some of them have skewed aspect ratio, pick the larger side and crop a square. Classifier is trained with square images. Changing the aspect ratio reduces its accuracy. \n- TTA: Invertimg, average the predicted probabilties\n- > 98%\n\n#### Post processing\n- Some of the large digits have artifacts, which get detected as other digits. IOU threshold cannot remove them because these are very small as compared to the larger digit. Remove any digit that lies completely inside another\n- Some of the background is also classified as digits. But the classifier has low confidence for these. Remove any digits where classifier prob < 0.45\n- > 99.1% \n\n## Debugging\nSince the competition training data is not seen by models, it can be used for validation.  Looking at images that are getting incorrect sum, gives a good idea of augmentations and filters to add.  \n\n## Tidbits\n- Accuracy on competition train set closely matches the test accuracy\n- Yolo took 3 days to train due to large image size\n- This solution doesn't help the original problem which authors had proposed of processing large images via networks in any way\n- In the beginning I was working towards innovative track using a pretrained unet + blazeblock regressor but then the dataset changed.\n\nThanks to organizers for this cool competition. \nAnd thanks Namtran. Your continuous uploads were a big encouragement. \n",
      "votes": 34
    },
    {
      "id": 1757635,
      "postDate": "2022-04-16T20:41:51.630Z",
      "content": "<p><a href=\"https://www.kaggle.com/marvin42\" target=\"_blank\">@marvin42</a> congratulations! Great result! <br>\nThank you for sharing git with code and exhaustive solution description. Really great!</p>\n<p>Could you help me understand some topics?</p>\n<blockquote>\n  <p>Ensure all 70000 instances from mnist are covered.</p>\n</blockquote>\n<ul>\n<li>What does it mean? Why 70.000?</li>\n</ul>\n<blockquote>\n  <p>Yolo alone gives &gt; 95%, adding a classifier gives &gt; 98%.</p>\n</blockquote>\n<ul>\n<li>As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1? </li>\n</ul>\n<blockquote>\n  <p>Try to increase recall of yolo detector.</p>\n</blockquote>\n<ul>\n<li>How have you tried increase recall? Only using conf_tresh or something else?</li>\n</ul>\n<blockquote>\n  <p>Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence.</p>\n</blockquote>\n<ul>\n<li>Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?</li>\n<li>Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")</li>\n</ul>",
      "rawMarkdown": "@marvin42 congratulations! Great result! \nThank you for sharing git with code and exhaustive solution description. Really great!\n\nCould you help me understand some topics?\n\n> Ensure all 70000 instances from mnist are covered.\n\n- What does it mean? Why 70.000?\n\n>Yolo alone gives > 95%, adding a classifier gives > 98%.\n\n- As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1? \n\n\n>Try to increase recall of yolo detector.\n\n- How have you tried increase recall? Only using conf_tresh or something else?\n\n>Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence.\n\n- Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?\n- Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")\n",
      "votes": 2,
      "replies": [
        {
          "id": 1757766,
          "postDate": "2022-04-17T02:16:01.717Z",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<ul>\n<li><p>What does it mean? Why 70.000?<br>\nMnist dataset has 7000 images per digit. Some of them are hard to detect correctly. So we have to show all of them to the classifier during training. This overfitting makes it easier to detect them when they occur as it is in test set. E.g. there are many 1's that look like 7's.   </p></li>\n<li><p>As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? <br>\nYes you are right. I run yolo with the yolo classifier after training but ignore its class labels. And only use the bbox labels.</p></li>\n<li><p>Have you tried Effnetv1?<br>\nI have used EffNetV2B1  as a backboen for classifier (<a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet_v2/EfficientNetV2B1\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet_v2/EfficientNetV2B1</a>) which is similar to Effnetb1. I have not used effdet anywhere. </p></li>\n<li><p>How have you tried increase recall? Only using conf_tresh or something else?<br>\nOnly using conf_thresh. I tried increasing iou_thresh and max_det but these lead to spurious detections which are harder to remove later.   </p></li>\n<li><p>Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?<br>\nNo</p></li>\n<li><p>Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")<br>\nNo. I ignore the class labels given by yolo. I use the confidence score of large classifier (effnet) to remove low confidence detections. </p></li>\n</ul>\n<p>Btw, I am a yolo noob. Would be happy to know any tips on using it better.    </p>",
          "rawMarkdown": "@remekkinas \n- What does it mean? Why 70.000?\nMnist dataset has 7000 images per digit. Some of them are hard to detect correctly. So we have to show all of them to the classifier during training. This overfitting makes it easier to detect them when they occur as it is in test set. E.g. there are many 1's that look like 7's.   \n\n- As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? \nYes you are right. I run yolo with the yolo classifier after training but ignore its class labels. And only use the bbox labels.\n\n- Have you tried Effnetv1?\nI have used EffNetV2B1  as a backboen for classifier (https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet_v2/EfficientNetV2B1) which is similar to Effnetb1. I have not used effdet anywhere. \n\n- How have you tried increase recall? Only using conf_tresh or something else?\nOnly using conf_thresh. I tried increasing iou_thresh and max_det but these lead to spurious detections which are harder to remove later.   \n\n- Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?\nNo\n\n- Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")\nNo. I ignore the class labels given by yolo. I use the confidence score of large classifier (effnet) to remove low confidence detections. \n\nBtw, I am a yolo noob. Would be happy to know any tips on using it better.    ",
          "votes": 4
        }
      ]
    },
    {
      "id": 1758252,
      "postDate": "2022-04-17T13:52:28.113Z",
      "content": "<p>As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1</p>",
      "rawMarkdown": "As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1"
    },
    {
      "id": 1758620,
      "postDate": "2022-04-17T21:36:09.330Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1757765,
      "postDate": "2022-04-17T02:15:06.747Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1758696,
      "postDate": "2022-04-18T01:19:56.427Z",
      "content": "<p>Thank you, I am learned by your Code.</p>",
      "rawMarkdown": "Thank you, I am learned by your Code."
    }
  ],
  "comments": [
    {
      "id": 1757635,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-04-16T20:41:51.630000",
      "content": "<p><a href=\"https://www.kaggle.com/marvin42\" target=\"_blank\">@marvin42</a> congratulations! Great result! <br>\nThank you for sharing git with code and exhaustive solution description. Really great!</p>\n<p>Could you help me understand some topics?</p>\n<blockquote>\n  <p>Ensure all 70000 instances from mnist are covered.</p>\n</blockquote>\n<ul>\n<li>What does it mean? Why 70.000?</li>\n</ul>\n<blockquote>\n  <p>Yolo alone gives &gt; 95%, adding a classifier gives &gt; 98%.</p>\n</blockquote>\n<ul>\n<li>As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1? </li>\n</ul>\n<blockquote>\n  <p>Try to increase recall of yolo detector.</p>\n</blockquote>\n<ul>\n<li>How have you tried increase recall? Only using conf_tresh or something else?</li>\n</ul>\n<blockquote>\n  <p>Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence.</p>\n</blockquote>\n<ul>\n<li>Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?</li>\n<li>Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 1757766,
          "author_name": "latent4g",
          "author_url": "",
          "post_date": "2022-04-17T02:16:01.717000",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </p>\n<ul>\n<li><p>What does it mean? Why 70.000?<br>\nMnist dataset has 7000 images per digit. Some of them are hard to detect correctly. So we have to show all of them to the classifier during training. This overfitting makes it easier to detect them when they occur as it is in test set. E.g. there are many 1's that look like 7's.   </p></li>\n<li><p>As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? <br>\nYes you are right. I run yolo with the yolo classifier after training but ignore its class labels. And only use the bbox labels.</p></li>\n<li><p>Have you tried Effnetv1?<br>\nI have used EffNetV2B1  as a backboen for classifier (<a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet_v2/EfficientNetV2B1\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet_v2/EfficientNetV2B1</a>) which is similar to Effnetb1. I have not used effdet anywhere. </p></li>\n<li><p>How have you tried increase recall? Only using conf_tresh or something else?<br>\nOnly using conf_thresh. I tried increasing iou_thresh and max_det but these lead to spurious detections which are harder to remove later.   </p></li>\n<li><p>Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?<br>\nNo</p></li>\n<li><p>Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")<br>\nNo. I ignore the class labels given by yolo. I use the confidence score of large classifier (effnet) to remove low confidence detections. </p></li>\n</ul>\n<p>Btw, I am a yolo noob. Would be happy to know any tips on using it better.    </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1758252,
      "author_name": "Puneet_physicist",
      "author_url": "",
      "post_date": "2022-04-17T13:52:28.113000",
      "content": "<p>As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1758620,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-17T21:36:09.330000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1757765,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-17T02:15:06.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1758696,
      "author_name": "qdssad",
      "author_url": "",
      "post_date": "2022-04-18T01:19:56.427000",
      "content": "<p>Thank you, I am learned by your Code.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1756610": "Congrats to all the winners. Code is available here : https://github.com/4g/umnist/tree/master along with trained models. You can use these to regenerate the results. Instructions to train from scratch are also attached in the readme. Feel free to raise any issues if something doesn't work.    \n\n## Idea\n- Generate synthetic data that looks like the competition data. \n- Use a small object detector like yolov5s and reclassify all labels using a larger network (EffV2B1)\n\nInput images are large. Any network consuming these directly will be slow. Yolo small gets good box-obj accuracy but low class accuracy. \n\nObjects in image are not complex. Even at a resolution of 128x128 these can be classified with > 99.9% accuracy using a sufficiently large network. So I train a classifier on digits cropped from synthetic training set and reclassify yolo results with this high accuracy classifier. Yolo alone gives > 95%, adding a classifier gives > 98%. Rest of accuracy comes from post processing magic.  \n\n## Data generation\n#### Detection\n- 2560x2560 input images. Smallest digit size is 6x6. \n- Generate checkerboards with boxes, circles and triangles. Apply augmentations like perspective, rotation, zoom to make it look more like competition dataset. \n- Generate locations where digits should be put. Make sure digits don't overlap. \n- Ensure all 70000 instances from mnist are covered. \n- Uniformly distribute sizes of digits. \n- Generate empty images with no digits to reduce background confusion. \n- Train 34000 with 4000 empty images : val 2000 with 0 empty images.  \n\n#### Classification\n- 128x128 images. \n- Digit crops from synthetic data. \n- Add raw mnist digits \n- Train : 230,000 (160k synthetic + 70k raw mnist)\n\n## Training\n#### Detector\n- Yolov5 small with batch size 2 (max that is supported at this resolution on my gpu)\n- light augmentation, because harsh augmentations take digits too much out of the image : rotation, scale, translate, hsv, mosaic. \n- SGD 100 epochs. Stop when the object loss reaches .008. Cls loss doesn't matter. Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence. \n\n#### Classifier\n- Effnetv2b1 with 128x128x3 input. Use pretrained imagenet weights for faster convergence. \n- augmentations : Shift scale, perspective, invert img, blur, hsv, brightness/contrast. Keep augmentations light as competition data doesn't look augmented. \n- Reduce lr every 50 epochs (3e-4, 1e-4, 3e-5, 1e-5). Train for 200. \n\n## Inference \n#### Detection\n- 2560x2560\n- Try to increase recall of yolo detector. \n- iou_threshold = 0.1. Digits are not intersecting, so any intersecting outputs are wrong \n- conf_threshold = 0.1 . To increase recall. \n- class agnostic nms\n- Maximum 5 digits allowed per image, max_det = 5\n- To measure the accuracy, use competition training data. \n- > 95%\n\n#### Classification\n- Crop digits detected by yolo. Some of them have skewed aspect ratio, pick the larger side and crop a square. Classifier is trained with square images. Changing the aspect ratio reduces its accuracy. \n- TTA: Invertimg, average the predicted probabilties\n- > 98%\n\n#### Post processing\n- Some of the large digits have artifacts, which get detected as other digits. IOU threshold cannot remove them because these are very small as compared to the larger digit. Remove any digit that lies completely inside another\n- Some of the background is also classified as digits. But the classifier has low confidence for these. Remove any digits where classifier prob < 0.45\n- > 99.1% \n\n## Debugging\nSince the competition training data is not seen by models, it can be used for validation.  Looking at images that are getting incorrect sum, gives a good idea of augmentations and filters to add.  \n\n## Tidbits\n- Accuracy on competition train set closely matches the test accuracy\n- Yolo took 3 days to train due to large image size\n- This solution doesn't help the original problem which authors had proposed of processing large images via networks in any way\n- In the beginning I was working towards innovative track using a pretrained unet + blazeblock regressor but then the dataset changed.\n\nThanks to organizers for this cool competition. \nAnd thanks Namtran. Your continuous uploads were a big encouragement. \n",
    "1757635": "@marvin42 congratulations! Great result! \nThank you for sharing git with code and exhaustive solution description. Really great!\n\nCould you help me understand some topics?\n\n> Ensure all 70000 instances from mnist are covered.\n\n- What does it mean? Why 70.000?\n\n>Yolo alone gives > 95%, adding a classifier gives > 98%.\n\n- As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1? \n\n\n>Try to increase recall of yolo detector.\n\n- How have you tried increase recall? Only using conf_tresh or something else?\n\n>Recall should be high as detector needs to retrieve all instances. delete them later if the classifier has low confidence.\n\n- Have you used any strategy to save specific yolov5 checkpoint with higher recall instead of mAP?\n- Have you used label smoothing for model calibration? (since you said \"delete them later if the classifier has low confidence.\")\n",
    "1758252": "As I understand firstly you use classifier from yolov5 (95%) and then class-agnostic detection and effdet for classification (98%)? Have you tried Effnetv1",
    "1758620": "",
    "1757765": "",
    "1758696": "Thank you, I am learned by your Code."
  }
}