{
  "id": 328593,
  "title": "3rd Place Solution",
  "url": "/competitions/sorghum-id-fgvc-9/discussion/328593",
  "author_name": "MoonFlower",
  "post_date": "2022-06-02T02:32:52.968000",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congrats to all the winners.<br>\nThanks to Kaggle and the hosting team for an interesting competition. <br>\nI appreciate my team member <a href=\"https://www.kaggle.com/SisuoLYU\" target=\"_blank\">@SisuoLYU</a> <a href=\"https://www.kaggle.com/haogood666\" target=\"_blank\">@haogood666</a> !</p>\n<h1><strong>Dataset</strong></h1>\n<p>Through data set analysis, it can be found that the data of this question has the following classification difficulties.<br>\n1 There is a high degree of inter-class visual similarity between different categories, and fine-grained classification is difficult.<br>\n2 Outdoor pictures on different dates and at different times of the day have large differences in lighting, and there are many exposed images.<br>\n3 The training set and the test set are from sorghum grown on two different fields, and there is a problem of domain adaptation.<br>\n4 With the growth time, the state and height of plants are changing.<br>\nFinally, our training size is:1024x1024. Dataset We convert png format to jpeg format, which improves storage space. Thanks to  @<br>\nMithil Salunkhe <br>\n <a href=\"https://www.kaggle.com/competitions/sorghum-id-fgvc-9/discussion/313266\" target=\"_blank\">https://www.kaggle.com/competitions/sorghum-id-fgvc-9/discussion/313266</a> for the open source jpeg data.<br>\nIt is mentioned here that we perform image histogram equalization preprocessing for both training and testing, which improves the problem of exposed images. Here, we are very grateful to@Jun-Ming Chen <br>\n<a href=\"https://www.kaggle.com/code/leoooo333/lb-0-885-sorghum-higer-accuracy\" target=\"_blank\">https://www.kaggle.com/code/leoooo333/lb-0-885-sorghum-higer-accuracy</a> open source notebook. We referenced many of their ideas and proposals.</p>\n<h1><strong>Model</strong></h1>\n<p><img src=\"https://i.postimg.cc/MKPbBNwp/3.jpg\" alt=\"\"></p>\n<p>Baseline: The initial debugging is trained on 512x512 images, resnet50 model, CrossEntropyLoss, cosine restart learning rate plan, AdamW optimizer.<br>\nSummarize:</p>\n<table>\n<thead>\n<tr>\n<th>methods</th>\n<th>scores</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>base(resnet50,512x512)</td>\n<td>0.73</td>\n</tr>\n<tr>\n<td>histogram equalization</td>\n<td>+0.03</td>\n</tr>\n<tr>\n<td>Ibn-resnet50</td>\n<td>+0.05</td>\n</tr>\n<tr>\n<td>bnn-neck</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>The last CNN layer stride is changed to 1</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>Arcface(s=30,m=0.3)</td>\n<td>+0.015</td>\n</tr>\n<tr>\n<td>mixup</td>\n<td>+0.01</td>\n</tr>\n<tr>\n<td>cutmix</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>Awp adversarial training</td>\n<td>+0.01</td>\n</tr>\n<tr>\n<td>512x512-&gt;1024x1024</td>\n<td>+0.04</td>\n</tr>\n<tr>\n<td>fgvc8data</td>\n<td>+0.03</td>\n</tr>\n<tr>\n<td>tta</td>\n<td>+0.02</td>\n</tr>\n<tr>\n<td>pse_udo</td>\n<td>+0.01</td>\n</tr>\n<tr>\n<td>total</td>\n<td>≈0.957</td>\n</tr>\n</tbody>\n</table>\n<p>Ibn-resnet：IBN-Net integrates BN and IN, which can improve the learning ability and generalization ability of the model. The design principle is: use both IN and BN in the shallow layers of the network, and only use BN in the deep layers of the network. In this problem, since the training set and the test set come from two different geographical locations, the addition of IN greatly improves the model performance. <br>\n<img src=\"https://i.postimg.cc/K817Dfcr/20220531110258.png\" alt=\"\"></p>\n<p>Arcface：The topic of this competition is fine-grained visual image classification, with high similarity between classes. Ordinary Softmax Loss does not explicitly optimize feature Embedding to enhance the similarity of intra-class samples and the inconsistency of inter-class samples. ArcFace is an additive angular margin penalty that adds an additive angular margin m between the feature and target weights to enhance both intra-class compactness and inter-class dissimilarity.</p>\n<p>Awp：Adversarial training is used in nlp to improve the performance of the model, such as fgm, pgd, but because it is the disturbance added in the text emb, in the image field, we choose the general awp confrontation training, which can add disturbance to the model weight and input at the same time, This increases the robustness of the model.</p>\n<h1><strong>Others</strong></h1>\n<p>Data augmentation mainly combines the characteristics of this data illumination change. In addition to mixup and cutmix, the following data augmentation:</p>\n<pre><code>train_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.HorizontalFlip(p=0.5),\n         A.VerticalFlip(p=0.5),\n         A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_REFLECT_101, p=0.5),\n         A.OneOf([A.RandomBrightness(limit=0.1, p=1), A.RandomContrast(limit=0.1, p=1)]),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n\nval_test_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n</code></pre>\n<p>Another thing worth noting in this question is that we used the fgvc9 and part of the fgvc8 training dataset. Since the image size of the fgvc8 dataset is 480x480, we spelled 1024x1024. </p>\n<h1><strong>Some failed attempts</strong></h1>\n<p>1 The model is changed to efficient, vit, etc.<br>\n2 The model fusion has not been scored, it may be that the late score is already very high, and the model fusion has not been improved.<br>\n3 The various improvements of loss, such as focalloss and softCrossEntropy, have not changed much.</p>\n<h1><strong>Code Links</strong></h1>\n<p><a href=\"https://github.com/xinchenzju/CV-competition-arsenal/tree/main/CVPR2022-FGVC9-top3\" target=\"_blank\">https://github.com/xinchenzju/CV-competition-arsenal/tree/main/CVPR2022-FGVC9-top3</a></p>",
  "messages": [
    {
      "id": 1808577,
      "postDate": "2022-06-02T02:32:52.970Z",
      "content": "<p>Congrats to all the winners.<br>\nThanks to Kaggle and the hosting team for an interesting competition. <br>\nI appreciate my team member <a href=\"https://www.kaggle.com/SisuoLYU\" target=\"_blank\">@SisuoLYU</a> <a href=\"https://www.kaggle.com/haogood666\" target=\"_blank\">@haogood666</a> !</p>\n<h1><strong>Dataset</strong></h1>\n<p>Through data set analysis, it can be found that the data of this question has the following classification difficulties.<br>\n1 There is a high degree of inter-class visual similarity between different categories, and fine-grained classification is difficult.<br>\n2 Outdoor pictures on different dates and at different times of the day have large differences in lighting, and there are many exposed images.<br>\n3 The training set and the test set are from sorghum grown on two different fields, and there is a problem of domain adaptation.<br>\n4 With the growth time, the state and height of plants are changing.<br>\nFinally, our training size is:1024x1024. Dataset We convert png format to jpeg format, which improves storage space. Thanks to  @<br>\nMithil Salunkhe <br>\n <a href=\"https://www.kaggle.com/competitions/sorghum-id-fgvc-9/discussion/313266\" target=\"_blank\">https://www.kaggle.com/competitions/sorghum-id-fgvc-9/discussion/313266</a> for the open source jpeg data.<br>\nIt is mentioned here that we perform image histogram equalization preprocessing for both training and testing, which improves the problem of exposed images. Here, we are very grateful to@Jun-Ming Chen <br>\n<a href=\"https://www.kaggle.com/code/leoooo333/lb-0-885-sorghum-higer-accuracy\" target=\"_blank\">https://www.kaggle.com/code/leoooo333/lb-0-885-sorghum-higer-accuracy</a> open source notebook. We referenced many of their ideas and proposals.</p>\n<h1><strong>Model</strong></h1>\n<p><img src=\"https://i.postimg.cc/MKPbBNwp/3.jpg\" alt=\"\"></p>\n<p>Baseline: The initial debugging is trained on 512x512 images, resnet50 model, CrossEntropyLoss, cosine restart learning rate plan, AdamW optimizer.<br>\nSummarize:</p>\n<table>\n<thead>\n<tr>\n<th>methods</th>\n<th>scores</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>base(resnet50,512x512)</td>\n<td>0.73</td>\n</tr>\n<tr>\n<td>histogram equalization</td>\n<td>+0.03</td>\n</tr>\n<tr>\n<td>Ibn-resnet50</td>\n<td>+0.05</td>\n</tr>\n<tr>\n<td>bnn-neck</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>The last CNN layer stride is changed to 1</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>Arcface(s=30,m=0.3)</td>\n<td>+0.015</td>\n</tr>\n<tr>\n<td>mixup</td>\n<td>+0.01</td>\n</tr>\n<tr>\n<td>cutmix</td>\n<td>+0.005</td>\n</tr>\n<tr>\n<td>Awp adversarial training</td>\n<td>+0.01</td>\n</tr>\n<tr>\n<td>512x512-&gt;1024x1024</td>\n<td>+0.04</td>\n</tr>\n<tr>\n<td>fgvc8data</td>\n<td>+0.03</td>\n</tr>\n<tr>\n<td>tta</td>\n<td>+0.02</td>\n</tr>\n<tr>\n<td>pse_udo</td>\n<td>+0.01</td>\n</tr>\n<tr>\n<td>total</td>\n<td>≈0.957</td>\n</tr>\n</tbody>\n</table>\n<p>Ibn-resnet：IBN-Net integrates BN and IN, which can improve the learning ability and generalization ability of the model. The design principle is: use both IN and BN in the shallow layers of the network, and only use BN in the deep layers of the network. In this problem, since the training set and the test set come from two different geographical locations, the addition of IN greatly improves the model performance. <br>\n<img src=\"https://i.postimg.cc/K817Dfcr/20220531110258.png\" alt=\"\"></p>\n<p>Arcface：The topic of this competition is fine-grained visual image classification, with high similarity between classes. Ordinary Softmax Loss does not explicitly optimize feature Embedding to enhance the similarity of intra-class samples and the inconsistency of inter-class samples. ArcFace is an additive angular margin penalty that adds an additive angular margin m between the feature and target weights to enhance both intra-class compactness and inter-class dissimilarity.</p>\n<p>Awp：Adversarial training is used in nlp to improve the performance of the model, such as fgm, pgd, but because it is the disturbance added in the text emb, in the image field, we choose the general awp confrontation training, which can add disturbance to the model weight and input at the same time, This increases the robustness of the model.</p>\n<h1><strong>Others</strong></h1>\n<p>Data augmentation mainly combines the characteristics of this data illumination change. In addition to mixup and cutmix, the following data augmentation:</p>\n<pre><code>train_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.HorizontalFlip(p=0.5),\n         A.VerticalFlip(p=0.5),\n         A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_REFLECT_101, p=0.5),\n         A.OneOf([A.RandomBrightness(limit=0.1, p=1), A.RandomContrast(limit=0.1, p=1)]),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n\nval_test_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n</code></pre>\n<p>Another thing worth noting in this question is that we used the fgvc9 and part of the fgvc8 training dataset. Since the image size of the fgvc8 dataset is 480x480, we spelled 1024x1024. </p>\n<h1><strong>Some failed attempts</strong></h1>\n<p>1 The model is changed to efficient, vit, etc.<br>\n2 The model fusion has not been scored, it may be that the late score is already very high, and the model fusion has not been improved.<br>\n3 The various improvements of loss, such as focalloss and softCrossEntropy, have not changed much.</p>\n<h1><strong>Code Links</strong></h1>\n<p><a href=\"https://github.com/xinchenzju/CV-competition-arsenal/tree/main/CVPR2022-FGVC9-top3\" target=\"_blank\">https://github.com/xinchenzju/CV-competition-arsenal/tree/main/CVPR2022-FGVC9-top3</a></p>",
      "rawMarkdown": "Congrats to all the winners.\nThanks to Kaggle and the hosting team for an interesting competition. \nI appreciate my team member @SisuoLYU @haogood666 !\n\n#  **Dataset**\nThrough data set analysis, it can be found that the data of this question has the following classification difficulties.\n1 There is a high degree of inter-class visual similarity between different categories, and fine-grained classification is difficult.\n2 Outdoor pictures on different dates and at different times of the day have large differences in lighting, and there are many exposed images.\n3 The training set and the test set are from sorghum grown on two different fields, and there is a problem of domain adaptation.\n4 With the growth time, the state and height of plants are changing.\nFinally, our training size is:1024x1024. Dataset We convert png format to jpeg format, which improves storage space. Thanks to  @\nMithil Salunkhe \n https://www.kaggle.com/competitions/sorghum-id-fgvc-9/discussion/313266 for the open source jpeg data.\nIt is mentioned here that we perform image histogram equalization preprocessing for both training and testing, which improves the problem of exposed images. Here, we are very grateful to@Jun-Ming Chen \nhttps://www.kaggle.com/code/leoooo333/lb-0-885-sorghum-higer-accuracy open source notebook. We referenced many of their ideas and proposals.\n\n# **Model**\n\n![](https://i.postimg.cc/MKPbBNwp/3.jpg)\n\nBaseline: The initial debugging is trained on 512x512 images, resnet50 model, CrossEntropyLoss, cosine restart learning rate plan, AdamW optimizer.\nSummarize:\n\n| methods |scores |\n| --- | --- |\n|  base(resnet50,512x512)| 0.73 |\n|  histogram equalization| +0.03 |\n|  Ibn-resnet50| +0.05 |\n| bnn-neck| \t+0.005| \n| The last CNN layer stride is changed to 1| \t+0.005| \n| Arcface(s=30,m=0.3)\t| +0.015| \n| mixup\t| +0.01| \n| cutmix\t| +0.005| \n| Awp adversarial training| \t+0.01| \n| 512x512->1024x1024| \t+0.04| \n| fgvc8data| \t+0.03| \n| tta\t| +0.02| \n| pse_udo| \t+0.01| \n| total\t| ≈0.957| \n\nIbn-resnet：IBN-Net integrates BN and IN, which can improve the learning ability and generalization ability of the model. The design principle is: use both IN and BN in the shallow layers of the network, and only use BN in the deep layers of the network. In this problem, since the training set and the test set come from two different geographical locations, the addition of IN greatly improves the model performance. \n![](https://i.postimg.cc/K817Dfcr/20220531110258.png)\n\nArcface：The topic of this competition is fine-grained visual image classification, with high similarity between classes. Ordinary Softmax Loss does not explicitly optimize feature Embedding to enhance the similarity of intra-class samples and the inconsistency of inter-class samples. ArcFace is an additive angular margin penalty that adds an additive angular margin m between the feature and target weights to enhance both intra-class compactness and inter-class dissimilarity.\n\nAwp：Adversarial training is used in nlp to improve the performance of the model, such as fgm, pgd, but because it is the disturbance added in the text emb, in the image field, we choose the general awp confrontation training, which can add disturbance to the model weight and input at the same time, This increases the robustness of the model.\n\n# **Others**\nData augmentation mainly combines the characteristics of this data illumination change. In addition to mixup and cutmix, the following data augmentation:\n```\ntrain_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.HorizontalFlip(p=0.5),\n         A.VerticalFlip(p=0.5),\n         A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_REFLECT_101, p=0.5),\n         A.OneOf([A.RandomBrightness(limit=0.1, p=1), A.RandomContrast(limit=0.1, p=1)]),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n\nval_test_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n```\nAnother thing worth noting in this question is that we used the fgvc9 and part of the fgvc8 training dataset. Since the image size of the fgvc8 dataset is 480x480, we spelled 1024x1024. \n\n# **Some failed attempts**\n\n1 The model is changed to efficient, vit, etc.\n2 The model fusion has not been scored, it may be that the late score is already very high, and the model fusion has not been improved.\n3 The various improvements of loss, such as focalloss and softCrossEntropy, have not changed much.\n\n# **Code Links**\n\nhttps://github.com/xinchenzju/CV-competition-arsenal/tree/main/CVPR2022-FGVC9-top3\n",
      "votes": 12
    },
    {
      "id": 1870765,
      "postDate": "2022-07-25T20:46:20.790Z",
      "content": "<p>Thank you for sharing your solution! Considering the table in which you show a breakdown of the scores per method, how many epochs did you use to get the score for each method? How many epochs did you use before submission?</p>",
      "rawMarkdown": "Thank you for sharing your solution! Considering the table in which you show a breakdown of the scores per method, how many epochs did you use to get the score for each method? How many epochs did you use before submission?",
      "votes": 1,
      "replies": [
        {
          "id": 1871696,
          "postDate": "2022-07-26T12:43:14.973Z",
          "content": "<p>Thank you for following our solutions！ The full experimental setup is 45 epochs. Our final solution is consistent with the table, and most of the methods are gradually improved based on experiments. Because our time is not enough, a small part of the improved methods did not do ablation experiments carefully. </p>",
          "rawMarkdown": "Thank you for following our solutions！ The full experimental setup is 45 epochs. Our final solution is consistent with the table, and most of the methods are gradually improved based on experiments. Because our time is not enough, a small part of the improved methods did not do ablation experiments carefully. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1808593,
      "postDate": "2022-06-02T02:52:12.780Z",
      "content": "<p>Thanks to my team member <a href=\"https://www.kaggle.com/xxxccc333\" target=\"_blank\">@xxxccc333</a> <a href=\"https://www.kaggle.com/haogood666\" target=\"_blank\">@haogood666</a> , this is the Second Cooperation to me and <a href=\"https://www.kaggle.com/xxxccc333\" target=\"_blank\">@xxxccc333</a>,and looking forward to the next cooperation.</p>",
      "rawMarkdown": "Thanks to my team member @xxxccc333 @haogood666 , this is the Second Cooperation to me and @xxxccc333,and looking forward to the next cooperation.",
      "votes": 1
    },
    {
      "id": 1876045,
      "postDate": "2022-07-29T14:43:58.803Z",
      "content": "<p>Could you please tell me how long it took for you to run 45 epochs for the input size of 1024 by 1024?<br>\nIn Kaggle's notebook, it is not possible to load all the 14 GB of jpeg version of the dataset (provided by Mithil Salunkhe ) and if I try to load the images from the disk (I use Pytorch), it takes a very long time to finish one epoch.  I use resnet50 as the base model with crossentropy loss function and Adam optimizer. Could you please let me know if it took so long for you too?</p>",
      "rawMarkdown": "Could you please tell me how long it took for you to run 45 epochs for the input size of 1024 by 1024?\nIn Kaggle's notebook, it is not possible to load all the 14 GB of jpeg version of the dataset (provided by Mithil Salunkhe ) and if I try to load the images from the disk (I use Pytorch), it takes a very long time to finish one epoch.  I use resnet50 as the base model with crossentropy loss function and Adam optimizer. Could you please let me know if it took so long for you too?",
      "replies": [
        {
          "id": 1876064,
          "postDate": "2022-07-29T14:53:26.717Z",
          "content": "<p>Our solution does not use Kaggle's notebook, the gpu we use is 3090. The specific time of a round is forgotten, and the training time is long. Because of the cosine restart learning rate used, you can run 20 epochs and the scores should be similar.</p>",
          "rawMarkdown": "Our solution does not use Kaggle's notebook, the gpu we use is 3090. The specific time of a round is forgotten, and the training time is long. Because of the cosine restart learning rate used, you can run 20 epochs and the scores should be similar.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1810252,
      "postDate": "2022-06-03T11:38:08.323Z",
      "content": "<p>Thanks for this great, thorough writeup and publishing your solution! There's a lot of things to learn from it!<br>\nI'm a bit surprised that you achieved this with only a resnet50!</p>",
      "rawMarkdown": "Thanks for this great, thorough writeup and publishing your solution! There's a lot of things to learn from it!\nI'm a bit surprised that you achieved this with only a resnet50!"
    },
    {
      "id": 1808584,
      "postDate": "2022-06-02T02:43:29.737Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1870765,
      "author_name": "Reza R. Choubeh",
      "author_url": "",
      "post_date": "2022-07-25T20:46:20.790000",
      "content": "<p>Thank you for sharing your solution! Considering the table in which you show a breakdown of the scores per method, how many epochs did you use to get the score for each method? How many epochs did you use before submission?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1871696,
          "author_name": "MoonFlower",
          "author_url": "",
          "post_date": "2022-07-26T12:43:14.973000",
          "content": "<p>Thank you for following our solutions！ The full experimental setup is 45 epochs. Our final solution is consistent with the table, and most of the methods are gradually improved based on experiments. Because our time is not enough, a small part of the improved methods did not do ablation experiments carefully. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1808593,
      "author_name": "SisuoLyu",
      "author_url": "",
      "post_date": "2022-06-02T02:52:12.780000",
      "content": "<p>Thanks to my team member <a href=\"https://www.kaggle.com/xxxccc333\" target=\"_blank\">@xxxccc333</a> <a href=\"https://www.kaggle.com/haogood666\" target=\"_blank\">@haogood666</a> , this is the Second Cooperation to me and <a href=\"https://www.kaggle.com/xxxccc333\" target=\"_blank\">@xxxccc333</a>,and looking forward to the next cooperation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1876045,
      "author_name": "Reza R. Choubeh",
      "author_url": "",
      "post_date": "2022-07-29T14:43:58.803000",
      "content": "<p>Could you please tell me how long it took for you to run 45 epochs for the input size of 1024 by 1024?<br>\nIn Kaggle's notebook, it is not possible to load all the 14 GB of jpeg version of the dataset (provided by Mithil Salunkhe ) and if I try to load the images from the disk (I use Pytorch), it takes a very long time to finish one epoch.  I use resnet50 as the base model with crossentropy loss function and Adam optimizer. Could you please let me know if it took so long for you too?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1876064,
          "author_name": "MoonFlower",
          "author_url": "",
          "post_date": "2022-07-29T14:53:26.717000",
          "content": "<p>Our solution does not use Kaggle's notebook, the gpu we use is 3090. The specific time of a round is forgotten, and the training time is long. Because of the cosine restart learning rate used, you can run 20 epochs and the scores should be similar.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1810252,
      "author_name": "cygn",
      "author_url": "",
      "post_date": "2022-06-03T11:38:08.323000",
      "content": "<p>Thanks for this great, thorough writeup and publishing your solution! There's a lot of things to learn from it!<br>\nI'm a bit surprised that you achieved this with only a resnet50!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1808584,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T02:43:29.737000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1808577": "Congrats to all the winners.\nThanks to Kaggle and the hosting team for an interesting competition. \nI appreciate my team member @SisuoLYU @haogood666 !\n\n#  **Dataset**\nThrough data set analysis, it can be found that the data of this question has the following classification difficulties.\n1 There is a high degree of inter-class visual similarity between different categories, and fine-grained classification is difficult.\n2 Outdoor pictures on different dates and at different times of the day have large differences in lighting, and there are many exposed images.\n3 The training set and the test set are from sorghum grown on two different fields, and there is a problem of domain adaptation.\n4 With the growth time, the state and height of plants are changing.\nFinally, our training size is:1024x1024. Dataset We convert png format to jpeg format, which improves storage space. Thanks to  @\nMithil Salunkhe \n https://www.kaggle.com/competitions/sorghum-id-fgvc-9/discussion/313266 for the open source jpeg data.\nIt is mentioned here that we perform image histogram equalization preprocessing for both training and testing, which improves the problem of exposed images. Here, we are very grateful to@Jun-Ming Chen \nhttps://www.kaggle.com/code/leoooo333/lb-0-885-sorghum-higer-accuracy open source notebook. We referenced many of their ideas and proposals.\n\n# **Model**\n\n![](https://i.postimg.cc/MKPbBNwp/3.jpg)\n\nBaseline: The initial debugging is trained on 512x512 images, resnet50 model, CrossEntropyLoss, cosine restart learning rate plan, AdamW optimizer.\nSummarize:\n\n| methods |scores |\n| --- | --- |\n|  base(resnet50,512x512)| 0.73 |\n|  histogram equalization| +0.03 |\n|  Ibn-resnet50| +0.05 |\n| bnn-neck| \t+0.005| \n| The last CNN layer stride is changed to 1| \t+0.005| \n| Arcface(s=30,m=0.3)\t| +0.015| \n| mixup\t| +0.01| \n| cutmix\t| +0.005| \n| Awp adversarial training| \t+0.01| \n| 512x512->1024x1024| \t+0.04| \n| fgvc8data| \t+0.03| \n| tta\t| +0.02| \n| pse_udo| \t+0.01| \n| total\t| ≈0.957| \n\nIbn-resnet：IBN-Net integrates BN and IN, which can improve the learning ability and generalization ability of the model. The design principle is: use both IN and BN in the shallow layers of the network, and only use BN in the deep layers of the network. In this problem, since the training set and the test set come from two different geographical locations, the addition of IN greatly improves the model performance. \n![](https://i.postimg.cc/K817Dfcr/20220531110258.png)\n\nArcface：The topic of this competition is fine-grained visual image classification, with high similarity between classes. Ordinary Softmax Loss does not explicitly optimize feature Embedding to enhance the similarity of intra-class samples and the inconsistency of inter-class samples. ArcFace is an additive angular margin penalty that adds an additive angular margin m between the feature and target weights to enhance both intra-class compactness and inter-class dissimilarity.\n\nAwp：Adversarial training is used in nlp to improve the performance of the model, such as fgm, pgd, but because it is the disturbance added in the text emb, in the image field, we choose the general awp confrontation training, which can add disturbance to the model weight and input at the same time, This increases the robustness of the model.\n\n# **Others**\nData augmentation mainly combines the characteristics of this data illumination change. In addition to mixup and cutmix, the following data augmentation:\n```\ntrain_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.HorizontalFlip(p=0.5),\n         A.VerticalFlip(p=0.5),\n         A.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, interpolation=cv2.INTER_LINEAR, border_mode=cv2.BORDER_REFLECT_101, p=0.5),\n         A.OneOf([A.RandomBrightness(limit=0.1, p=1), A.RandomContrast(limit=0.1, p=1)]),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n\nval_test_transform = A.Compose([\n         A.CLAHE(clip_limit=40, tile_grid_size=(10, 10),p=1.0),\n         A.Resize(CFG.image_size, CFG.image_size),\n         A.Normalize(),\n         ToTensorV2(p=1.0),\n     ])\n```\nAnother thing worth noting in this question is that we used the fgvc9 and part of the fgvc8 training dataset. Since the image size of the fgvc8 dataset is 480x480, we spelled 1024x1024. \n\n# **Some failed attempts**\n\n1 The model is changed to efficient, vit, etc.\n2 The model fusion has not been scored, it may be that the late score is already very high, and the model fusion has not been improved.\n3 The various improvements of loss, such as focalloss and softCrossEntropy, have not changed much.\n\n# **Code Links**\n\nhttps://github.com/xinchenzju/CV-competition-arsenal/tree/main/CVPR2022-FGVC9-top3\n",
    "1870765": "Thank you for sharing your solution! Considering the table in which you show a breakdown of the scores per method, how many epochs did you use to get the score for each method? How many epochs did you use before submission?",
    "1808593": "Thanks to my team member @xxxccc333 @haogood666 , this is the Second Cooperation to me and @xxxccc333,and looking forward to the next cooperation.",
    "1876045": "Could you please tell me how long it took for you to run 45 epochs for the input size of 1024 by 1024?\nIn Kaggle's notebook, it is not possible to load all the 14 GB of jpeg version of the dataset (provided by Mithil Salunkhe ) and if I try to load the images from the disk (I use Pytorch), it takes a very long time to finish one epoch.  I use resnet50 as the base model with crossentropy loss function and Adam optimizer. Could you please let me know if it took so long for you too?",
    "1810252": "Thanks for this great, thorough writeup and publishing your solution! There's a lot of things to learn from it!\nI'm a bit surprised that you achieved this with only a resnet50!",
    "1808584": ""
  }
}