{
  "id": 319206,
  "title": "3rd Place --- Multi-scale detection",
  "url": "/competitions/ultra-mnist/discussion/319206",
  "author_name": "NamTran",
  "post_date": "2022-04-16T00:57:27.125000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First and foremost, thank you especially to organizers who created an excited challenge.</p>\n<p>This is my first competition in kaggle playground, i'm very surprise when be 3rd winner :), So I'm honored to share my solution for this competition.</p>\n<p><strong>Key-ideals</strong></p>\n<ul>\n<li>object detection approach: detect all digit in image</li>\n<li>try to generate new dataset as closely as test dataset, and train object detector</li>\n<li>define batch_size_inference mechanism with yolov5 for TTA and speed-up inference time</li>\n<li>Source code for Generating dataset and training model is available here: <a href=\"https://github.com/NamTran11/UltraMNIST_3rd_solution\" target=\"_blank\">https://github.com/NamTran11/UltraMNIST_3rd_solution</a>,<br>\nfeel free to raise any issue for this code :)</li>\n</ul>\n<p><strong>Step By Step</strong></p>\n<ol>\n<li>I generated 60,000 images similar to the training and testing using the MNIST dataset.</li>\n<li>Training single fold yolov5x model with 1024 x 1012 solution (LB 0.88) - <code>normal model</code></li>\n<li>Training another yolov5x <code>tiny</code> model to detect only 'tiny' digit in image (trained with special tiny dataset) - <code>tiny model</code></li>\n<li>Repeat step 1, 2, 3 to get the best <code>normal model</code> and <code>tiny model</code></li>\n<li>Deal with tiny digit in image by TTA (Test time augmentation) for boost accuracy of detectors.<ul>\n<li>with <code>normal model</code>: inference input image with multiple solutions, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536</li>\n<li>with <code>tiny model</code>: split input image in multiple sub-images, detect and merge result in each of sub-image with 9-split + 16-split + 25-split.</li></ul></li>\n<li>Ensemble result of predictions: 9 split, 16 split, 25 split, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536 with list confidence_score [0.8, 0.8, 0.7, 0.95, 0.9, 0.9, 0.95]</li>\n</ol>\n<p>submission result: <a href=\"https://www.kaggle.com/code/namtran1118/ultramnist-3rd-solution\" target=\"_blank\">https://www.kaggle.com/code/namtran1118/ultramnist-3rd-solution</a></p>",
  "messages": [
    {
      "id": 1756842,
      "postDate": "2022-04-16T00:57:27.127Z",
      "content": "<p>First and foremost, thank you especially to organizers who created an excited challenge.</p>\n<p>This is my first competition in kaggle playground, i'm very surprise when be 3rd winner :), So I'm honored to share my solution for this competition.</p>\n<p><strong>Key-ideals</strong></p>\n<ul>\n<li>object detection approach: detect all digit in image</li>\n<li>try to generate new dataset as closely as test dataset, and train object detector</li>\n<li>define batch_size_inference mechanism with yolov5 for TTA and speed-up inference time</li>\n<li>Source code for Generating dataset and training model is available here: <a href=\"https://github.com/NamTran11/UltraMNIST_3rd_solution\" target=\"_blank\">https://github.com/NamTran11/UltraMNIST_3rd_solution</a>,<br>\nfeel free to raise any issue for this code :)</li>\n</ul>\n<p><strong>Step By Step</strong></p>\n<ol>\n<li>I generated 60,000 images similar to the training and testing using the MNIST dataset.</li>\n<li>Training single fold yolov5x model with 1024 x 1012 solution (LB 0.88) - <code>normal model</code></li>\n<li>Training another yolov5x <code>tiny</code> model to detect only 'tiny' digit in image (trained with special tiny dataset) - <code>tiny model</code></li>\n<li>Repeat step 1, 2, 3 to get the best <code>normal model</code> and <code>tiny model</code></li>\n<li>Deal with tiny digit in image by TTA (Test time augmentation) for boost accuracy of detectors.<ul>\n<li>with <code>normal model</code>: inference input image with multiple solutions, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536</li>\n<li>with <code>tiny model</code>: split input image in multiple sub-images, detect and merge result in each of sub-image with 9-split + 16-split + 25-split.</li></ul></li>\n<li>Ensemble result of predictions: 9 split, 16 split, 25 split, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536 with list confidence_score [0.8, 0.8, 0.7, 0.95, 0.9, 0.9, 0.95]</li>\n</ol>\n<p>submission result: <a href=\"https://www.kaggle.com/code/namtran1118/ultramnist-3rd-solution\" target=\"_blank\">https://www.kaggle.com/code/namtran1118/ultramnist-3rd-solution</a></p>",
      "rawMarkdown": "First and foremost, thank you especially to organizers who created an excited challenge.\n\nThis is my first competition in kaggle playground, i'm very surprise when be 3rd winner :), So I'm honored to share my solution for this competition.\n\n**Key-ideals**\n+ object detection approach: detect all digit in image\n+ try to generate new dataset as closely as test dataset, and train object detector\n+ define batch_size_inference mechanism with yolov5 for TTA and speed-up inference time\n+ Source code for Generating dataset and training model is available here: https://github.com/NamTran11/UltraMNIST_3rd_solution,\n feel free to raise any issue for this code :)\n\n\n**Step By Step**\n1. I generated 60,000 images similar to the training and testing using the MNIST dataset.\n2. Training single fold yolov5x model with 1024 x 1012 solution (LB 0.88) - `normal model`\n3. Training another yolov5x `tiny` model to detect only 'tiny' digit in image (trained with special tiny dataset) - `tiny model`\n4. Repeat step 1, 2, 3 to get the best `normal model` and `tiny model`\n5. Deal with tiny digit in image by TTA (Test time augmentation) for boost accuracy of detectors.\n  - with `normal model`: inference input image with multiple solutions, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536\n  - with `tiny model`: split input image in multiple sub-images, detect and merge result in each of sub-image with 9-split + 16-split + 25-split.\n6. Ensemble result of predictions: 9 split, 16 split, 25 split, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536 with list confidence_score [0.8, 0.8, 0.7, 0.95, 0.9, 0.9, 0.95]\n\nsubmission result: https://www.kaggle.com/code/namtran1118/ultramnist-3rd-solution",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1756842": "First and foremost, thank you especially to organizers who created an excited challenge.\n\nThis is my first competition in kaggle playground, i'm very surprise when be 3rd winner :), So I'm honored to share my solution for this competition.\n\n**Key-ideals**\n+ object detection approach: detect all digit in image\n+ try to generate new dataset as closely as test dataset, and train object detector\n+ define batch_size_inference mechanism with yolov5 for TTA and speed-up inference time\n+ Source code for Generating dataset and training model is available here: https://github.com/NamTran11/UltraMNIST_3rd_solution,\n feel free to raise any issue for this code :)\n\n\n**Step By Step**\n1. I generated 60,000 images similar to the training and testing using the MNIST dataset.\n2. Training single fold yolov5x model with 1024 x 1012 solution (LB 0.88) - `normal model`\n3. Training another yolov5x `tiny` model to detect only 'tiny' digit in image (trained with special tiny dataset) - `tiny model`\n4. Repeat step 1, 2, 3 to get the best `normal model` and `tiny model`\n5. Deal with tiny digit in image by TTA (Test time augmentation) for boost accuracy of detectors.\n  - with `normal model`: inference input image with multiple solutions, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536\n  - with `tiny model`: split input image in multiple sub-images, detect and merge result in each of sub-image with 9-split + 16-split + 25-split.\n6. Ensemble result of predictions: 9 split, 16 split, 25 split, 768 x 768, 1024 x 1024, 1280 x 1280, 1536 x 1536 with list confidence_score [0.8, 0.8, 0.7, 0.95, 0.9, 0.9, 0.95]\n\nsubmission result: https://www.kaggle.com/code/namtran1118/ultramnist-3rd-solution"
  }
}