{
  "id": 319110,
  "title": "2nd Place -- Simulate Background",
  "url": "/competitions/ultra-mnist/writeups/wantanewgpu4genshin-2nd-place-simulate-background",
  "author_name": "",
  "post_date": "2022-04-15T11:38:14.026767100Z",
  "votes": 24,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Thanks to the organizers for holding this competition. I've had a 1080Ti in my gaming PC for about 5 years and it's time to retire it ;)</p>\n<p>While I am more than happy to share the methods I used in this competition. But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I <strong>don't</strong> consider the method to be worth learning.</p>\n<h2>Summary</h2>\n<p>To get the new GPU, the end to end model is almost impossible for me. Actually I believe all accuracy &gt; 90% participants either use methods that eliminate the background of the image, or simulate the background of the image.</p>\n<p>My choice was the latter, and I used the MNIST dataset to generate a dataset very similar to the training data. I call it <code>generated data</code> .</p>\n<p>(<a href=\"https://www.kaggle.com/code/haqishen/simulate-background/notebook\" target=\"_blank\">https://www.kaggle.com/code/haqishen/simulate-background/notebook</a>) here is how I generated the data.</p>\n<p>The biggest difference between the generated data and the training data is that the generated data has bounding boxes for all numbers in images.</p>\n<p>So that allows me to train the digital detector by generated data.</p>\n<p>The next thing is relatively simple, I used detector to crop out all the numbers in the images, then use number classifier on them, and finally just add the numbers of each image together to get <code>digit_sum</code>.</p>\n<h2>Step by Step</h2>\n<ul>\n<li><p>In the first step, I observed the training data to find some patterns in the background of the training data and the range of sizes of the number appearing in the training data.</p></li>\n<li><p>By using the observed patterns, I generated 60,000 images very similar to the training data myself using the MNIST dataset. </p></li>\n<li><p>In the third step I use the generated dataset to train the yolov5 detector by 4-fold cross validation manner, which is used to detect the numbers that appear in the images.</p>\n<ul>\n<li>detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.</li></ul></li>\n<li><p>Step 4, I used 4fold yolov5 to make OOF and crop out the box in OOF to make a single-number images dataset.</p></li>\n<li><p>In the 5th step, a 10-class classifier is trained using the single-number dataset generated in the previous step.</p>\n<ul>\n<li>classifier details: backbone=effnetb5, imgsz=224, batch_size=128, only random shift augmentation.</li></ul></li>\n<li><p>Next, I use the trained yolov5 detector and number classifier to predict the original training data. After the prediction is completed I get the sum of the predicted numbers for each image.</p></li>\n<li><p>Step 7, If the sum of the predicted numbers is the same as <code>digit_sum</code>, I consider the image prediction successful (somewhat similar to pseudo here, but the probability of drawing correct predictions is much greater than pseudo). I then add these images to the yolov5 training set to retrain yolov5 and effnetb5. </p></li>\n<li><p>Repeat step 3 to step 7 to refine yolov5 detector for 5 rounds. But except for the 1st round I only train one fold yolov5 model without cross validation for doing experiments faster.</p></li>\n</ul>",
  "messages": [
    {
      "id": "1756296",
      "postDate": "04/15/2022 11:38:14",
      "content": "<p>Thanks to the organizers for holding this competition. I've had a 1080Ti in my gaming PC for about 5 years and it's time to retire it ;)</p>\n<p>While I am more than happy to share the methods I used in this competition. But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I <strong>don't</strong> consider the method to be worth learning.</p>\n<h2>Summary</h2>\n<p>To get the new GPU, the end to end model is almost impossible for me. Actually I believe all accuracy &gt; 90% participants either use methods that eliminate the background of the image, or simulate the background of the image.</p>\n<p>My choice was the latter, and I used the MNIST dataset to generate a dataset very similar to the training data. I call it <code>generated data</code> .</p>\n<p>(<a href=\"https://www.kaggle.com/code/haqishen/simulate-background/notebook\" target=\"_blank\">https://www.kaggle.com/code/haqishen/simulate-background/notebook</a>) here is how I generated the data.</p>\n<p>The biggest difference between the generated data and the training data is that the generated data has bounding boxes for all numbers in images.</p>\n<p>So that allows me to train the digital detector by generated data.</p>\n<p>The next thing is relatively simple, I used detector to crop out all the numbers in the images, then use number classifier on them, and finally just add the numbers of each image together to get <code>digit_sum</code>.</p>\n<h2>Step by Step</h2>\n<ul>\n<li><p>In the first step, I observed the training data to find some patterns in the background of the training data and the range of sizes of the number appearing in the training data.</p></li>\n<li><p>By using the observed patterns, I generated 60,000 images very similar to the training data myself using the MNIST dataset. </p></li>\n<li><p>In the third step I use the generated dataset to train the yolov5 detector by 4-fold cross validation manner, which is used to detect the numbers that appear in the images.</p>\n<ul>\n<li>detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.</li></ul></li>\n<li><p>Step 4, I used 4fold yolov5 to make OOF and crop out the box in OOF to make a single-number images dataset.</p></li>\n<li><p>In the 5th step, a 10-class classifier is trained using the single-number dataset generated in the previous step.</p>\n<ul>\n<li>classifier details: backbone=effnetb5, imgsz=224, batch_size=128, only random shift augmentation.</li></ul></li>\n<li><p>Next, I use the trained yolov5 detector and number classifier to predict the original training data. After the prediction is completed I get the sum of the predicted numbers for each image.</p></li>\n<li><p>Step 7, If the sum of the predicted numbers is the same as <code>digit_sum</code>, I consider the image prediction successful (somewhat similar to pseudo here, but the probability of drawing correct predictions is much greater than pseudo). I then add these images to the yolov5 training set to retrain yolov5 and effnetb5. </p></li>\n<li><p>Repeat step 3 to step 7 to refine yolov5 detector for 5 rounds. But except for the 1st round I only train one fold yolov5 model without cross validation for doing experiments faster.</p></li>\n</ul>",
      "rawMarkdown": "Thanks to the organizers for holding this competition. I've had a 1080Ti in my gaming PC for about 5 years and it's time to retire it ;)\n\nWhile I am more than happy to share the methods I used in this competition. But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I **don't** consider the method to be worth learning.\n\n\n## Summary\n\nTo get the new GPU, the end to end model is almost impossible for me. Actually I believe all accuracy > 90% participants either use methods that eliminate the background of the image, or simulate the background of the image.\n\nMy choice was the latter, and I used the MNIST dataset to generate a dataset very similar to the training data. I call it `generated data` .\n\n(https://www.kaggle.com/code/haqishen/simulate-background/notebook) here is how I generated the data.\n\nThe biggest difference between the generated data and the training data is that the generated data has bounding boxes for all numbers in images.\n\nSo that allows me to train the digital detector by generated data.\n\nThe next thing is relatively simple, I used detector to crop out all the numbers in the images, then use number classifier on them, and finally just add the numbers of each image together to get `digit_sum`.\n\n\n## Step by Step\n\n\n* In the first step, I observed the training data to find some patterns in the background of the training data and the range of sizes of the number appearing in the training data.\n\n* By using the observed patterns, I generated 60,000 images very similar to the training data myself using the MNIST dataset. \n\n* In the third step I use the generated dataset to train the yolov5 detector by 4-fold cross validation manner, which is used to detect the numbers that appear in the images.\n    * detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.\n\n* Step 4, I used 4fold yolov5 to make OOF and crop out the box in OOF to make a single-number images dataset.\n\n* In the 5th step, a 10-class classifier is trained using the single-number dataset generated in the previous step.\n    * classifier details: backbone=effnetb5, imgsz=224, batch_size=128, only random shift augmentation.\n\n* Next, I use the trained yolov5 detector and number classifier to predict the original training data. After the prediction is completed I get the sum of the predicted numbers for each image.\n\n* Step 7, If the sum of the predicted numbers is the same as `digit_sum`, I consider the image prediction successful (somewhat similar to pseudo here, but the probability of drawing correct predictions is much greater than pseudo). I then add these images to the yolov5 training set to retrain yolov5 and effnetb5. \n\n* Repeat step 3 to step 7 to refine yolov5 detector for 5 rounds. But except for the 1st round I only train one fold yolov5 model without cross validation for doing experiments faster.",
      "votes": null
    },
    {
      "id": "1756368",
      "postDate": "04/15/2022 12:53:34",
      "content": "<p>Awesome, thanks for such a detailed write-up!<br>\nCould you please also tell us about how you ran the inference? <br>\nFor me, I just couldn't make it work unless I took the overlapped-tiling approach.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Awesome, thanks for such a detailed write-up!\nCould you please also tell us about how you ran the inference? \nFor me, I just couldn't make it work unless I took the overlapped-tiling approach.\n\nThanks!",
      "votes": null
    },
    {
      "id": "1756476",
      "postDate": "04/15/2022 14:49:25",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Thanks you participating and congratulations on winning the competition. I hope with the new GPU, you would be also helping us to find research solutions in future where we can work without the MNIST :)</p>\n<p>On a different note, we would like to include your submission in the supplementary material of our upcoming paper (not the main paper), and we would acknowledge you for the work. We hope that is fine with you. In case yes, we will be contacting you in the coming days to know more about your solution.</p>",
      "rawMarkdown": "haqishen Thanks you participating and congratulations on winning the competition. I hope with the new GPU, you would be also helping us to find research solutions in future where we can work without the MNIST :)\n\nOn a different note, we would like to include your submission in the supplementary material of our upcoming paper (not the main paper), and we would acknowledge you for the work. We hope that is fine with you. In case yes, we will be contacting you in the coming days to know more about your solution.",
      "votes": null
    },
    {
      "id": "1756526",
      "postDate": "04/15/2022 15:30:23",
      "content": "<p>just resize images from 4000 to 1536 and do train/infer. nothing special</p>",
      "rawMarkdown": "just resize images from 4000 to 1536 and do train/infer. nothing special",
      "votes": null
    },
    {
      "id": "1756529",
      "postDate": "04/15/2022 15:31:37",
      "content": "<p>Sure, waiting for you email ;)</p>",
      "rawMarkdown": "Sure, waiting for you email ;)",
      "votes": null
    },
    {
      "id": "1756655",
      "postDate": "04/15/2022 18:15:12",
      "content": "<p>That's really wonderful. It means that you were able to detect digits at 6x6 resolution!!<br>\nIs there any way of learning things from you? 😄</p>",
      "rawMarkdown": "That's really wonderful. It means that you were able to detect digits at 6x6 resolution!!\nIs there any way of learning things from you? 😄",
      "votes": null
    },
    {
      "id": "1757620",
      "postDate": "04/16/2022 20:09:23",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> - first of all congratulations. I realy enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.</p>\n<blockquote>\n  <p>But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I don't consider the method to be worth learning.</p>\n</blockquote>\n<p>I do not agree with you 😆 - I learned a lot from your solution description 😊 Many CV problems have a problem with data availability (there are very few or very expensive photos eg. SAR images or other much more specific one e.g. damage to products in inspection systems) and then practically 100% of the generated dataset is synthetic. In aerial problems this is common way to generate data using 3D models (even generating 2D images) and then train model to recognize/detect objects. From my experience sometimes it is easy … and sometimes is  difficult to generate data that represent the features of a real scenes (model trained on synthetic DS does not perform well on real data even photos look \"almost\" the same).  The ability to generate a synthetic dataset is then more important than choosing/training the model itself. I like your solution and description. As I said - learned a lot! Thank you for sharing! Enjoy using new GPU 💪💪💪</p>\n<blockquote>\n  <p>detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.</p>\n</blockquote>\n<p>As I can see this was not 1080i training 😁 so … new gaming experience will come soon :)</p>",
      "rawMarkdown": "haqishen - first of all congratulations. I realy enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.\n\n> But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I don't consider the method to be worth learning.\n\nI do not agree with you 😆 - I learned a lot from your solution description 😊 Many CV problems have a problem with data availability (there are very few or very expensive photos eg. SAR images or other much more specific one e.g. damage to products in inspection systems) and then practically 100% of the generated dataset is synthetic. In aerial problems this is common way to generate data using 3D models (even generating 2D images) and then train model to recognize/detect objects. From my experience sometimes it is easy ... and sometimes is  difficult to generate data that represent the features of a real scenes (model trained on synthetic DS does not perform well on real data even photos look \"almost\" the same).  The ability to generate a synthetic dataset is then more important than choosing/training the model itself. I like your solution and description. As I said - learned a lot! Thank you for sharing! Enjoy using new GPU 💪💪💪\n\n>detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.\n\nAs I can see this was not 1080i training 😁 so ... new gaming experience will come soon :)",
      "votes": null
    },
    {
      "id": "1757861",
      "postDate": "04/17/2022 04:55:23",
      "content": "<p>😂not sure, maybe just follow me and read my post and code?</p>",
      "rawMarkdown": "😂not sure, maybe just follow me and read my post and code?",
      "votes": null
    },
    {
      "id": "1757864",
      "postDate": "04/17/2022 04:58:47",
      "content": "<p>Thanks for your kind words! I didn't realize that some places do use synthetic data.</p>\n<blockquote>\n  <p>As I can see this was not 1080i training 😁 so … new gaming experience will come soon :)</p>\n</blockquote>\n<p>There is no doubt that my gaming computer and my workstation are not the same one ;)</p>",
      "rawMarkdown": "Thanks for your kind words! I didn't realize that some places do use synthetic data.\n\n> As I can see this was not 1080i training 😁 so … new gaming experience will come soon :)\n\nThere is no doubt that my gaming computer and my workstation are not the same one ;)",
      "votes": null
    },
    {
      "id": "1758154",
      "postDate": "04/17/2022 11:28:19",
      "content": "<p>yeah, waiting for you to opensource the codes for this competition 😃</p>",
      "rawMarkdown": "yeah, waiting for you to opensource the codes for this competition 😃",
      "votes": null
    },
    {
      "id": "1758159",
      "postDate": "04/17/2022 11:37:22",
      "content": "<p>Yes, yes …. I know what are you using (starfish #1 solution descrpition) - HP desktop with 2xA6000.<br>\nI have almost the same (thanks to HP) …. HP 8series machine but with \"only\" one A6000 and it works perfectly.</p>",
      "rawMarkdown": "Yes, yes .... I know what are you using (starfish #1 solution descrpition) - HP desktop with 2xA6000.\nI have almost the same (thanks to HP) .... HP 8series machine but with \"only\" one A6000 and it works perfectly.",
      "votes": null
    },
    {
      "id": "1758872",
      "postDate": "04/18/2022 05:54:21",
      "content": "<p>I really enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.</p>",
      "rawMarkdown": "I really enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.",
      "votes": null
    },
    {
      "id": "1759716",
      "postDate": "04/18/2022 20:32:26",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1756368,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "04/15/2022 12:53:34",
      "content": "<p>Awesome, thanks for such a detailed write-up!<br>\nCould you please also tell us about how you ran the inference? <br>\nFor me, I just couldn't make it work unless I took the overlapped-tiling approach.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1756526,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "04/15/2022 15:30:23",
          "content": "<p>just resize images from 4000 to 1536 and do train/infer. nothing special</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1756655,
          "author_name": "mightyrains",
          "author_url": "",
          "post_date": "04/15/2022 18:15:12",
          "content": "<p>That's really wonderful. It means that you were able to detect digits at 6x6 resolution!!<br>\nIs there any way of learning things from you? 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1757861,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "04/17/2022 04:55:23",
          "content": "<p>😂not sure, maybe just follow me and read my post and code?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1758154,
          "author_name": "mightyrains",
          "author_url": "",
          "post_date": "04/17/2022 11:28:19",
          "content": "<p>yeah, waiting for you to opensource the codes for this competition 😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1756476,
      "author_name": "dkgupta90",
      "author_url": "",
      "post_date": "04/15/2022 14:49:25",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> Thanks you participating and congratulations on winning the competition. I hope with the new GPU, you would be also helping us to find research solutions in future where we can work without the MNIST :)</p>\n<p>On a different note, we would like to include your submission in the supplementary material of our upcoming paper (not the main paper), and we would acknowledge you for the work. We hope that is fine with you. In case yes, we will be contacting you in the coming days to know more about your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1756529,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "04/15/2022 15:31:37",
          "content": "<p>Sure, waiting for you email ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1757620,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/16/2022 20:09:23",
      "content": "<p><a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> - first of all congratulations. I realy enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.</p>\n<blockquote>\n  <p>But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I don't consider the method to be worth learning.</p>\n</blockquote>\n<p>I do not agree with you 😆 - I learned a lot from your solution description 😊 Many CV problems have a problem with data availability (there are very few or very expensive photos eg. SAR images or other much more specific one e.g. damage to products in inspection systems) and then practically 100% of the generated dataset is synthetic. In aerial problems this is common way to generate data using 3D models (even generating 2D images) and then train model to recognize/detect objects. From my experience sometimes it is easy … and sometimes is  difficult to generate data that represent the features of a real scenes (model trained on synthetic DS does not perform well on real data even photos look \"almost\" the same).  The ability to generate a synthetic dataset is then more important than choosing/training the model itself. I like your solution and description. As I said - learned a lot! Thank you for sharing! Enjoy using new GPU 💪💪💪</p>\n<blockquote>\n  <p>detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.</p>\n</blockquote>\n<p>As I can see this was not 1080i training 😁 so … new gaming experience will come soon :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1757864,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "04/17/2022 04:58:47",
          "content": "<p>Thanks for your kind words! I didn't realize that some places do use synthetic data.</p>\n<blockquote>\n  <p>As I can see this was not 1080i training 😁 so … new gaming experience will come soon :)</p>\n</blockquote>\n<p>There is no doubt that my gaming computer and my workstation are not the same one ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1758159,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "04/17/2022 11:37:22",
          "content": "<p>Yes, yes …. I know what are you using (starfish #1 solution descrpition) - HP desktop with 2xA6000.<br>\nI have almost the same (thanks to HP) …. HP 8series machine but with \"only\" one A6000 and it works perfectly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1758872,
      "author_name": "",
      "author_url": "",
      "post_date": "04/18/2022 05:54:21",
      "content": "<p>I really enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1759716,
      "author_name": "meryemnaciri",
      "author_url": "",
      "post_date": "04/18/2022 20:32:26",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1756296": "Thanks to the organizers for holding this competition. I've had a 1080Ti in my gaming PC for about 5 years and it's time to retire it ;)\n\nWhile I am more than happy to share the methods I used in this competition. But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I **don't** consider the method to be worth learning.\n\n\n## Summary\n\nTo get the new GPU, the end to end model is almost impossible for me. Actually I believe all accuracy > 90% participants either use methods that eliminate the background of the image, or simulate the background of the image.\n\nMy choice was the latter, and I used the MNIST dataset to generate a dataset very similar to the training data. I call it `generated data` .\n\n(https://www.kaggle.com/code/haqishen/simulate-background/notebook) here is how I generated the data.\n\nThe biggest difference between the generated data and the training data is that the generated data has bounding boxes for all numbers in images.\n\nSo that allows me to train the digital detector by generated data.\n\nThe next thing is relatively simple, I used detector to crop out all the numbers in the images, then use number classifier on them, and finally just add the numbers of each image together to get `digit_sum`.\n\n\n## Step by Step\n\n\n* In the first step, I observed the training data to find some patterns in the background of the training data and the range of sizes of the number appearing in the training data.\n\n* By using the observed patterns, I generated 60,000 images very similar to the training data myself using the MNIST dataset. \n\n* In the third step I use the generated dataset to train the yolov5 detector by 4-fold cross validation manner, which is used to detect the numbers that appear in the images.\n    * detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.\n\n* Step 4, I used 4fold yolov5 to make OOF and crop out the box in OOF to make a single-number images dataset.\n\n* In the 5th step, a 10-class classifier is trained using the single-number dataset generated in the previous step.\n    * classifier details: backbone=effnetb5, imgsz=224, batch_size=128, only random shift augmentation.\n\n* Next, I use the trained yolov5 detector and number classifier to predict the original training data. After the prediction is completed I get the sum of the predicted numbers for each image.\n\n* Step 7, If the sum of the predicted numbers is the same as `digit_sum`, I consider the image prediction successful (somewhat similar to pseudo here, but the probability of drawing correct predictions is much greater than pseudo). I then add these images to the yolov5 training set to retrain yolov5 and effnetb5. \n\n* Repeat step 3 to step 7 to refine yolov5 detector for 5 rounds. But except for the 1st round I only train one fold yolov5 model without cross validation for doing experiments faster.",
    "1756368": "Awesome, thanks for such a detailed write-up!\nCould you please also tell us about how you ran the inference? \nFor me, I just couldn't make it work unless I took the overlapped-tiling approach.\n\nThanks!",
    "1756476": "haqishen Thanks you participating and congratulations on winning the competition. I hope with the new GPU, you would be also helping us to find research solutions in future where we can work without the MNIST :)\n\nOn a different note, we would like to include your submission in the supplementary material of our upcoming paper (not the main paper), and we would acknowledge you for the work. We hope that is fine with you. In case yes, we will be contacting you in the coming days to know more about your solution.",
    "1756526": "just resize images from 4000 to 1536 and do train/infer. nothing special",
    "1756529": "Sure, waiting for you email ;)",
    "1756655": "That's really wonderful. It means that you were able to detect digits at 6x6 resolution!!\nIs there any way of learning things from you? 😄",
    "1757620": "haqishen - first of all congratulations. I realy enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.\n\n> But I would like to remind you that this method is only applicable to this competition, and cannot be used in industrial tasks, or other competitions - to be frank, I don't consider the method to be worth learning.\n\nI do not agree with you 😆 - I learned a lot from your solution description 😊 Many CV problems have a problem with data availability (there are very few or very expensive photos eg. SAR images or other much more specific one e.g. damage to products in inspection systems) and then practically 100% of the generated dataset is synthetic. In aerial problems this is common way to generate data using 3D models (even generating 2D images) and then train model to recognize/detect objects. From my experience sometimes it is easy ... and sometimes is  difficult to generate data that represent the features of a real scenes (model trained on synthetic DS does not perform well on real data even photos look \"almost\" the same).  The ability to generate a synthetic dataset is then more important than choosing/training the model itself. I like your solution and description. As I said - learned a lot! Thank you for sharing! Enjoy using new GPU 💪💪💪\n\n>detector details: backbone=yolov5m6, imgsz=1536, batch_size=32, only random shift augmentation.\n\nAs I can see this was not 1080i training 😁 so ... new gaming experience will come soon :)",
    "1757861": "😂not sure, maybe just follow me and read my post and code?",
    "1757864": "Thanks for your kind words! I didn't realize that some places do use synthetic data.\n\n> As I can see this was not 1080i training 😁 so … new gaming experience will come soon :)\n\nThere is no doubt that my gaming computer and my workstation are not the same one ;)",
    "1758154": "yeah, waiting for you to opensource the codes for this competition 😃",
    "1758159": "Yes, yes .... I know what are you using (starfish #1 solution descrpition) - HP desktop with 2xA6000.\nI have almost the same (thanks to HP) .... HP 8series machine but with \"only\" one A6000 and it works perfectly.",
    "1758872": "I really enjoy reading your solutions (not only in this competition). I follow you and read everything you post on Kaggle. I am very impressed how \"easily\" you achieve such great results - respect. You are absolutely AI master.",
    "1759716": "Thank you for sharing"
  },
  "source": "meta"
}