{
  "id": 412114,
  "title": "2nd place solution",
  "url": "/competitions/ibiohash-2023-fgvc10/discussion/412114",
  "author_name": "Chet",
  "post_date": "2023-05-22T10:14:41.156000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congrats to all winners.</p>\n<p>Thanks to the hosting team for the interesting competition.</p>\n<p>Certainly, thanks my team member <a href=\"https://www.kaggle.com/yuanshuai9574\" target=\"_blank\">@yuanshuai9574</a> and <a href=\"https://www.kaggle.com/blueOatcake\" target=\"_blank\">@blueOatcake</a>.</p>\n<h1>Summary</h1>\n<p>Our solution is simple. We did not take any complicated design into consideration.</p>\n<ol>\n<li>Pertaining to the data, we executed essential data augmentation, which is imperative for visual data that are generally redundant and inefficient. We endeavored to manage the data, i.e., data cleaning. However, this appeared to be ineffective and unnecessary in this competition. This could potentially be attributed to the competition's utilization of public datasets rather than commercial ones, which have already undergone thorough cleansing.</li>\n<li>Regarding the architecture, we initially attempted the simplest baseline, i.e., end-to-end hash learning. Nevertheless, the results were unsatisfactory, leading us to decouple the feature learning and hash components. We independently tested the effectiveness of feature learning and hash learning, with the expectation that superior features would correspond to superior hashes, a hypothesis that is typically valid.</li>\n<li>As for the training algorithm, we made several attempts. The best-performing method turned out to be recall@k[1] + ITQ[2], which respectively accomplished feature learning and hash learning. In our experiments, classification could also learn decent features, but the features obtained from end-to-end hash classification were significantly inferior. Interestingly, margin loss, considered as the most robust baseline in metric learning, performed poorly in our experiments. This could possibly be due to the presence of fine-grained data and label shift, suggesting that we need to meticulously handle feature separability to ensure decent performance on the test set.</li>\n<li>We barely utilized any testing-time techniques, except for TTA (Test-Time Augmentation), despite acknowledging their potential to enhance performance in various scenarios.</li>\n</ol>\n<h1>Code</h1>\n<p>We have open-sourced our main codebase on the following GitHub repository.</p>\n<p><a href=\"https://github.com/codeman34134/iBioHash2023\" target=\"_blank\">https://github.com/codeman34134/iBioHash2023</a></p>\n<h1>Model</h1>\n<p>In this instance, we present the schematic diagram of our finest model. As illustrated, it exudes remarkable simplicity.</p>\n<p>Our swin model is the swinv2_large pretrained on imagenet-21k and fine-tuned on imagenet-1k.<br>\n<a href=\"https://github.com/microsoft/Swin-Transformer\" target=\"_blank\">https://github.com/microsoft/Swin-Transformer</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5116651%2Fbc1f3eae1569f040f52328f30e1326eb%2F11981684748243_.pic.jpg?generation=1684750434024993&amp;alt=media\" alt=\"\"></p>\n<h1>Some failed attempts</h1>\n<ol>\n<li>We attempted to employ self-supervision techniques to identify fine-grained data, but to no avail. It appears that instance discrimination struggles with fine-grained data learning, a finding that is within our expectations.</li>\n<li>We made an attempt at domain adaptation, which unfortunately proved to be unsuccessful. However, we firmly believe that domain adaptation is essential in scenarios involving label shifts.</li>\n</ol>\n<h1>Conclusion</h1>\n<p>If you have any inquiries, wish to engage in further discussions, or seek academic collaboration, please feel free to contact us at tchen@seu.edu.cn.</p>\n<h1><strong>Reference</strong></h1>\n<p>[1] Patel Y, Tolias G, Matas J. Recall@ k surrogate loss with large batches and similarity mixup[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 7502-7511.</p>\n<p>[2] Gong Y, Lazebnik S, Gordo A, et al. Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval[J]. IEEE transactions on pattern analysis and machine intelligence, 2012, 35(12): 2916-2929.</p>",
  "messages": [
    {
      "id": 2269235,
      "postDate": "2023-05-22T10:14:41.157Z",
      "content": "<p>Congrats to all winners.</p>\n<p>Thanks to the hosting team for the interesting competition.</p>\n<p>Certainly, thanks my team member <a href=\"https://www.kaggle.com/yuanshuai9574\" target=\"_blank\">@yuanshuai9574</a> and <a href=\"https://www.kaggle.com/blueOatcake\" target=\"_blank\">@blueOatcake</a>.</p>\n<h1>Summary</h1>\n<p>Our solution is simple. We did not take any complicated design into consideration.</p>\n<ol>\n<li>Pertaining to the data, we executed essential data augmentation, which is imperative for visual data that are generally redundant and inefficient. We endeavored to manage the data, i.e., data cleaning. However, this appeared to be ineffective and unnecessary in this competition. This could potentially be attributed to the competition's utilization of public datasets rather than commercial ones, which have already undergone thorough cleansing.</li>\n<li>Regarding the architecture, we initially attempted the simplest baseline, i.e., end-to-end hash learning. Nevertheless, the results were unsatisfactory, leading us to decouple the feature learning and hash components. We independently tested the effectiveness of feature learning and hash learning, with the expectation that superior features would correspond to superior hashes, a hypothesis that is typically valid.</li>\n<li>As for the training algorithm, we made several attempts. The best-performing method turned out to be recall@k[1] + ITQ[2], which respectively accomplished feature learning and hash learning. In our experiments, classification could also learn decent features, but the features obtained from end-to-end hash classification were significantly inferior. Interestingly, margin loss, considered as the most robust baseline in metric learning, performed poorly in our experiments. This could possibly be due to the presence of fine-grained data and label shift, suggesting that we need to meticulously handle feature separability to ensure decent performance on the test set.</li>\n<li>We barely utilized any testing-time techniques, except for TTA (Test-Time Augmentation), despite acknowledging their potential to enhance performance in various scenarios.</li>\n</ol>\n<h1>Code</h1>\n<p>We have open-sourced our main codebase on the following GitHub repository.</p>\n<p><a href=\"https://github.com/codeman34134/iBioHash2023\" target=\"_blank\">https://github.com/codeman34134/iBioHash2023</a></p>\n<h1>Model</h1>\n<p>In this instance, we present the schematic diagram of our finest model. As illustrated, it exudes remarkable simplicity.</p>\n<p>Our swin model is the swinv2_large pretrained on imagenet-21k and fine-tuned on imagenet-1k.<br>\n<a href=\"https://github.com/microsoft/Swin-Transformer\" target=\"_blank\">https://github.com/microsoft/Swin-Transformer</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5116651%2Fbc1f3eae1569f040f52328f30e1326eb%2F11981684748243_.pic.jpg?generation=1684750434024993&amp;alt=media\" alt=\"\"></p>\n<h1>Some failed attempts</h1>\n<ol>\n<li>We attempted to employ self-supervision techniques to identify fine-grained data, but to no avail. It appears that instance discrimination struggles with fine-grained data learning, a finding that is within our expectations.</li>\n<li>We made an attempt at domain adaptation, which unfortunately proved to be unsuccessful. However, we firmly believe that domain adaptation is essential in scenarios involving label shifts.</li>\n</ol>\n<h1>Conclusion</h1>\n<p>If you have any inquiries, wish to engage in further discussions, or seek academic collaboration, please feel free to contact us at tchen@seu.edu.cn.</p>\n<h1><strong>Reference</strong></h1>\n<p>[1] Patel Y, Tolias G, Matas J. Recall@ k surrogate loss with large batches and similarity mixup[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 7502-7511.</p>\n<p>[2] Gong Y, Lazebnik S, Gordo A, et al. Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval[J]. IEEE transactions on pattern analysis and machine intelligence, 2012, 35(12): 2916-2929.</p>",
      "rawMarkdown": "Congrats to all winners.\n\nThanks to the hosting team for the interesting competition.\n\nCertainly, thanks my team member @yuanshuai9574 and @blueOatcake.\n\n# Summary\n\nOur solution is simple. We did not take any complicated design into consideration.\n\n1.  Pertaining to the data, we executed essential data augmentation, which is imperative for visual data that are generally redundant and inefficient. We endeavored to manage the data, i.e., data cleaning. However, this appeared to be ineffective and unnecessary in this competition. This could potentially be attributed to the competition's utilization of public datasets rather than commercial ones, which have already undergone thorough cleansing.\n2.  Regarding the architecture, we initially attempted the simplest baseline, i.e., end-to-end hash learning. Nevertheless, the results were unsatisfactory, leading us to decouple the feature learning and hash components. We independently tested the effectiveness of feature learning and hash learning, with the expectation that superior features would correspond to superior hashes, a hypothesis that is typically valid.\n3.  As for the training algorithm, we made several attempts. The best-performing method turned out to be recall@k[1] + ITQ[2], which respectively accomplished feature learning and hash learning. In our experiments, classification could also learn decent features, but the features obtained from end-to-end hash classification were significantly inferior. Interestingly, margin loss, considered as the most robust baseline in metric learning, performed poorly in our experiments. This could possibly be due to the presence of fine-grained data and label shift, suggesting that we need to meticulously handle feature separability to ensure decent performance on the test set.\n4.  We barely utilized any testing-time techniques, except for TTA (Test-Time Augmentation), despite acknowledging their potential to enhance performance in various scenarios.\n\n# Code\n\nWe have open-sourced our main codebase on the following GitHub repository.\n\nhttps://github.com/codeman34134/iBioHash2023\n\n# Model\n\nIn this instance, we present the schematic diagram of our finest model. As illustrated, it exudes remarkable simplicity.\n\nOur swin model is the swinv2_large pretrained on imagenet-21k and fine-tuned on imagenet-1k.\nhttps://github.com/microsoft/Swin-Transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5116651%2Fbc1f3eae1569f040f52328f30e1326eb%2F11981684748243_.pic.jpg?generation=1684750434024993&alt=media)\n\n# Some failed attempts\n\n1.  We attempted to employ self-supervision techniques to identify fine-grained data, but to no avail. It appears that instance discrimination struggles with fine-grained data learning, a finding that is within our expectations.\n2.  We made an attempt at domain adaptation, which unfortunately proved to be unsuccessful. However, we firmly believe that domain adaptation is essential in scenarios involving label shifts.\n\n# Conclusion\n\nIf you have any inquiries, wish to engage in further discussions, or seek academic collaboration, please feel free to contact us at tchen@seu.edu.cn.\n\n# **Reference**\n\n[1] Patel Y, Tolias G, Matas J. Recall@ k surrogate loss with large batches and similarity mixup[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 7502-7511.\n\n[2] Gong Y, Lazebnik S, Gordo A, et al. Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval[J]. IEEE transactions on pattern analysis and machine intelligence, 2012, 35(12): 2916-2929.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2269235": "Congrats to all winners.\n\nThanks to the hosting team for the interesting competition.\n\nCertainly, thanks my team member @yuanshuai9574 and @blueOatcake.\n\n# Summary\n\nOur solution is simple. We did not take any complicated design into consideration.\n\n1.  Pertaining to the data, we executed essential data augmentation, which is imperative for visual data that are generally redundant and inefficient. We endeavored to manage the data, i.e., data cleaning. However, this appeared to be ineffective and unnecessary in this competition. This could potentially be attributed to the competition's utilization of public datasets rather than commercial ones, which have already undergone thorough cleansing.\n2.  Regarding the architecture, we initially attempted the simplest baseline, i.e., end-to-end hash learning. Nevertheless, the results were unsatisfactory, leading us to decouple the feature learning and hash components. We independently tested the effectiveness of feature learning and hash learning, with the expectation that superior features would correspond to superior hashes, a hypothesis that is typically valid.\n3.  As for the training algorithm, we made several attempts. The best-performing method turned out to be recall@k[1] + ITQ[2], which respectively accomplished feature learning and hash learning. In our experiments, classification could also learn decent features, but the features obtained from end-to-end hash classification were significantly inferior. Interestingly, margin loss, considered as the most robust baseline in metric learning, performed poorly in our experiments. This could possibly be due to the presence of fine-grained data and label shift, suggesting that we need to meticulously handle feature separability to ensure decent performance on the test set.\n4.  We barely utilized any testing-time techniques, except for TTA (Test-Time Augmentation), despite acknowledging their potential to enhance performance in various scenarios.\n\n# Code\n\nWe have open-sourced our main codebase on the following GitHub repository.\n\nhttps://github.com/codeman34134/iBioHash2023\n\n# Model\n\nIn this instance, we present the schematic diagram of our finest model. As illustrated, it exudes remarkable simplicity.\n\nOur swin model is the swinv2_large pretrained on imagenet-21k and fine-tuned on imagenet-1k.\nhttps://github.com/microsoft/Swin-Transformer\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5116651%2Fbc1f3eae1569f040f52328f30e1326eb%2F11981684748243_.pic.jpg?generation=1684750434024993&alt=media)\n\n# Some failed attempts\n\n1.  We attempted to employ self-supervision techniques to identify fine-grained data, but to no avail. It appears that instance discrimination struggles with fine-grained data learning, a finding that is within our expectations.\n2.  We made an attempt at domain adaptation, which unfortunately proved to be unsuccessful. However, we firmly believe that domain adaptation is essential in scenarios involving label shifts.\n\n# Conclusion\n\nIf you have any inquiries, wish to engage in further discussions, or seek academic collaboration, please feel free to contact us at tchen@seu.edu.cn.\n\n# **Reference**\n\n[1] Patel Y, Tolias G, Matas J. Recall@ k surrogate loss with large batches and similarity mixup[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022: 7502-7511.\n\n[2] Gong Y, Lazebnik S, Gordo A, et al. Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval[J]. IEEE transactions on pattern analysis and machine intelligence, 2012, 35(12): 2916-2929."
  }
}