{
  "id": 328237,
  "title": "Public/Private 3rd Solution",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/writeups/tereka-public-private-3rd-solution",
  "author_name": "",
  "post_date": "2022-05-31T14:17:27.042420900Z",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thank you for host,competitors,team member <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> <br>\nIt is a great competiton!<br>\nhere is our team summary.</p>\n<h1>Summary</h1>\n<ul>\n<li>ArcFaceSubcenter + Dynamic Margin</li>\n<li>Many Backbones(Swin,ConvNeXt, ResNet200D, EfficientNet etc..)</li>\n<li>use External Data(FGVC8)</li>\n<li>use Logits</li>\n</ul>\n<h2>Modeling</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Data</th>\n<th>ImageSize</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Swin Large</td>\n<td>Comp</td>\n<td>512</td>\n<td>0.587</td>\n<td>0.584</td>\n</tr>\n<tr>\n<td>ConvNeXt XLarge</td>\n<td>Comp + External</td>\n<td>384</td>\n<td>0.662</td>\n<td>0.656</td>\n</tr>\n<tr>\n<td>Swin Large</td>\n<td>Comp+ External</td>\n<td>384</td>\n<td>0.659</td>\n<td>0.657</td>\n</tr>\n<tr>\n<td>EfficientNetB7 + DOLG</td>\n<td>Comp + External</td>\n<td>448(stride1)</td>\n<td>0.644</td>\n<td>0.639</td>\n</tr>\n<tr>\n<td>ConvNeXt XLarge</td>\n<td>Comp + External</td>\n<td>512</td>\n<td>0.677</td>\n<td>0.672</td>\n</tr>\n<tr>\n<td>ResNet200D + DOLG</td>\n<td>Comp + External</td>\n<td>640</td>\n<td>0.651</td>\n<td>0.654</td>\n</tr>\n<tr>\n<td>EfficientNetB6 + DOLG</td>\n<td>Comp + External</td>\n<td>640</td>\n<td>0.635</td>\n<td>0.642</td>\n</tr>\n<tr>\n<td>EfficientNetV2S + DOLG</td>\n<td>Comp + External</td>\n<td>1024</td>\n<td>0.645</td>\n<td>0.642</td>\n</tr>\n<tr>\n<td>EfficientNetV2M + DOLG</td>\n<td>Comp + External</td>\n<td>896</td>\n<td>0.666</td>\n<td>0.666</td>\n</tr>\n</tbody>\n</table>\n<h2>Pseudo Labeling</h2>\n<p>We use FGVC9 and FGVC8 competiton data but FGVC8 don't have competiton labels.<br>\nWe annotated label to these data using Pseudo Labeling.</p>\n<p>First, We used KNN matching FGVC9 training and FGVC8 dataset.<br>\nalso if threshold &lt; 0.5, these data used training.</p>\n<p>Next, We adjust labeling. FGVC8 have same hotel list.<br>\nI use mean aggregation method</p>\n<h2>Prediction</h2>\n<p>I use logits because test have a mask, but train do not have a mask.<br>\nI think it is difficult to match training and test so I decide to compare knn vs logits<br>\nLogits is better than knn in my experiments.</p>\n<p>My inference time is about 2hours.</p>\n<h2>did not work</h2>\n<ul>\n<li>used Hotel50K</li>\n<li>gradient checkpointing</li>\n</ul>",
  "messages": [
    {
      "id": "1806873",
      "postDate": "05/31/2022 14:17:27",
      "content": "<p>Thank you for host,competitors,team member <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> <br>\nIt is a great competiton!<br>\nhere is our team summary.</p>\n<h1>Summary</h1>\n<ul>\n<li>ArcFaceSubcenter + Dynamic Margin</li>\n<li>Many Backbones(Swin,ConvNeXt, ResNet200D, EfficientNet etc..)</li>\n<li>use External Data(FGVC8)</li>\n<li>use Logits</li>\n</ul>\n<h2>Modeling</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Data</th>\n<th>ImageSize</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Swin Large</td>\n<td>Comp</td>\n<td>512</td>\n<td>0.587</td>\n<td>0.584</td>\n</tr>\n<tr>\n<td>ConvNeXt XLarge</td>\n<td>Comp + External</td>\n<td>384</td>\n<td>0.662</td>\n<td>0.656</td>\n</tr>\n<tr>\n<td>Swin Large</td>\n<td>Comp+ External</td>\n<td>384</td>\n<td>0.659</td>\n<td>0.657</td>\n</tr>\n<tr>\n<td>EfficientNetB7 + DOLG</td>\n<td>Comp + External</td>\n<td>448(stride1)</td>\n<td>0.644</td>\n<td>0.639</td>\n</tr>\n<tr>\n<td>ConvNeXt XLarge</td>\n<td>Comp + External</td>\n<td>512</td>\n<td>0.677</td>\n<td>0.672</td>\n</tr>\n<tr>\n<td>ResNet200D + DOLG</td>\n<td>Comp + External</td>\n<td>640</td>\n<td>0.651</td>\n<td>0.654</td>\n</tr>\n<tr>\n<td>EfficientNetB6 + DOLG</td>\n<td>Comp + External</td>\n<td>640</td>\n<td>0.635</td>\n<td>0.642</td>\n</tr>\n<tr>\n<td>EfficientNetV2S + DOLG</td>\n<td>Comp + External</td>\n<td>1024</td>\n<td>0.645</td>\n<td>0.642</td>\n</tr>\n<tr>\n<td>EfficientNetV2M + DOLG</td>\n<td>Comp + External</td>\n<td>896</td>\n<td>0.666</td>\n<td>0.666</td>\n</tr>\n</tbody>\n</table>\n<h2>Pseudo Labeling</h2>\n<p>We use FGVC9 and FGVC8 competiton data but FGVC8 don't have competiton labels.<br>\nWe annotated label to these data using Pseudo Labeling.</p>\n<p>First, We used KNN matching FGVC9 training and FGVC8 dataset.<br>\nalso if threshold &lt; 0.5, these data used training.</p>\n<p>Next, We adjust labeling. FGVC8 have same hotel list.<br>\nI use mean aggregation method</p>\n<h2>Prediction</h2>\n<p>I use logits because test have a mask, but train do not have a mask.<br>\nI think it is difficult to match training and test so I decide to compare knn vs logits<br>\nLogits is better than knn in my experiments.</p>\n<p>My inference time is about 2hours.</p>\n<h2>did not work</h2>\n<ul>\n<li>used Hotel50K</li>\n<li>gradient checkpointing</li>\n</ul>",
      "rawMarkdown": "Thank you for host,competitors,team member @ks2019 \nIt is a great competiton!\nhere is our team summary.\n\n# Summary\n- ArcFaceSubcenter + Dynamic Margin\n- Many Backbones(Swin,ConvNeXt, ResNet200D, EfficientNet etc..)\n- use External Data(FGVC8)\n- use Logits\n\n## Modeling\n|Model|Data|ImageSize|Public|Private|\n| --- | --- |\n|Swin Large|Comp|512|0.587|0.584|\n|ConvNeXt XLarge|Comp + External|384|0.662|0.656|\n|Swin Large|Comp+ External|384|0.659|0.657|\n|EfficientNetB7 + DOLG|Comp + External|448(stride1)|0.644|0.639|\n|ConvNeXt XLarge|Comp + External|512|0.677|0.672|\n|ResNet200D + DOLG|Comp + External|640|0.651|0.654|\n|EfficientNetB6 + DOLG|Comp + External|640|0.635|0.642|\n|EfficientNetV2S + DOLG|Comp + External|1024|0.645|0.642|\n|EfficientNetV2M + DOLG|Comp + External|896|0.666|0.666|\n\n## Pseudo Labeling\nWe use FGVC9 and FGVC8 competiton data but FGVC8 don't have competiton labels.\nWe annotated label to these data using Pseudo Labeling.\n\nFirst, We used KNN matching FGVC9 training and FGVC8 dataset.\nalso if threshold < 0.5, these data used training.\n\nNext, We adjust labeling. FGVC8 have same hotel list.\nI use mean aggregation method\n\n## Prediction\nI use logits because test have a mask, but train do not have a mask.\nI think it is difficult to match training and test so I decide to compare knn vs logits\nLogits is better than knn in my experiments.\n\nMy inference time is about 2hours.\n\n## did not work\n- used Hotel50K\n- gradient checkpointing",
      "votes": null
    },
    {
      "id": "1806887",
      "postDate": "05/31/2022 14:33:09",
      "content": "<p>Nice solution, congrats to 3rd place, 0.711 is impressive result.</p>\n<p>Did you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?</p>\n<p>Do you plan to release the source code? I am curious about the ArcFaceSubcenter + Dynamic Margin :-)</p>\n<p>What is DOLG?</p>",
      "rawMarkdown": "Nice solution, congrats to 3rd place, 0.711 is impressive result.\n\nDid you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?\n\nDo you plan to release the source code? I am curious about the ArcFaceSubcenter + Dynamic Margin :-)\n\nWhat is DOLG?",
      "votes": null
    },
    {
      "id": "1806891",
      "postDate": "05/31/2022 14:39:09",
      "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Did you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?</p>\n</blockquote>\n<p>I create mask using albumentations.</p>\n<blockquote>\n  <p>Do you plan to release the source code?</p>\n</blockquote>\n<p>No. </p>\n<blockquote>\n  <p>I am curious about the ArcFaceSubcenter + Dynamic Margin :-)</p>\n</blockquote>\n<p>ArcFaceSubcenter + Dynamic Margin is used for GLR 2020 3rd place.<br>\n<a href=\"https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\" target=\"_blank\">https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757</a><br>\nit have source code.</p>\n<blockquote>\n  <p>What is DOLG?</p>\n</blockquote>\n<p>DOLG is used for GLR2021 1st Solution<br>\nI refer to it code.<br>\n<a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place</a><br>\n<a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/models/ch_mdl_dolg_efficientnet.py#L248\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/models/ch_mdl_dolg_efficientnet.py#L248</a></p>",
      "rawMarkdown": "Thanks!\n\n> Did you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?\n\nI create mask using albumentations.\n\n> Do you plan to release the source code?\n\nNo. \n\n>  I am curious about the ArcFaceSubcenter + Dynamic Margin :-)\n\nArcFaceSubcenter + Dynamic Margin is used for GLR 2020 3rd place.\nhttps://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\nit have source code.\n\n> What is DOLG?\n\nDOLG is used for GLR2021 1st Solution\nI refer to it code.\nhttps://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place\nhttps://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/models/ch_mdl_dolg_efficientnet.py#L248",
      "votes": null
    },
    {
      "id": "1806896",
      "postDate": "05/31/2022 14:43:37",
      "content": "<p>Cool, thanks for the links :-)</p>",
      "rawMarkdown": "Cool, thanks for the links :-)",
      "votes": null
    },
    {
      "id": "1807125",
      "postDate": "05/31/2022 18:24:05",
      "content": "<p>Nice solution!  You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me.  I also like you got subcenter + dynamic margin to work well, for me plain vanilla arcface performed better but intuitively subcenter should work well due to the different scenes within the same hotel (ie bathroom vs bedroom)</p>",
      "rawMarkdown": "Nice solution!  You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me.  I also like you got subcenter + dynamic margin to work well, for me plain vanilla arcface performed better but intuitively subcenter should work well due to the different scenes within the same hotel (ie bathroom vs bedroom)",
      "votes": null
    },
    {
      "id": "1807356",
      "postDate": "06/01/2022 00:05:00",
      "content": "<p>Thank you!<br>\nI read your solution that is also great. (Blind Flip etc..)</p>\n<blockquote>\n  <p>You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me</p>\n</blockquote>\n<p>heavier model is better(ConxNeXt XLarge)<br>\nIf I have more resouces, I tried many models ensemble.</p>\n<blockquote>\n  <p>I also like you got subcenter + dynamic margin </p>\n</blockquote>\n<p>I experiment both ArcFace vanilla and ArcFace Subcenter + Dynamic margin, better is ArcFace Subcenter + Dynamic margin.  </p>\n<blockquote>\n  <p>due to the different scenes within the same hotel (ie bathroom vs bedroom)</p>\n</blockquote>\n<p>I agree, Hotel have difference scenes.</p>",
      "rawMarkdown": "Thank you!\nI read your solution that is also great. (Blind Flip etc..)\n\n> You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me\n\nheavier model is better(ConxNeXt XLarge)\nIf I have more resouces, I tried many models ensemble.\n\n>  I also like you got subcenter + dynamic margin \n\nI experiment both ArcFace vanilla and ArcFace Subcenter + Dynamic margin, better is ArcFace Subcenter + Dynamic margin.  \n  \n> due to the different scenes within the same hotel (ie bathroom vs bedroom)\n\nI agree, Hotel have difference scenes.",
      "votes": null
    },
    {
      "id": "1818044",
      "postDate": "06/12/2022 07:06:41",
      "content": "<p><strong>It looks like your submission is against the rules. Sadly, no one is reading the rules these days.</strong></p>\n<ol>\n<li>COMPETITION DATA.</li>\n</ol>\n<p>\"Competition Data\" means the data or datasets available from the Competition Website for the purpose of use in the Competition, including any prototype or executable code provided on the Competition Website. The Competition Data will contain private and public test sets. Which data belongs to which set will not be made available to participants.</p>\n<p>A. Data Access and Use.</p>\n<p>Competition Use and Non-Commercial &amp; Academic Research: You may access and use the Competition Data for non-commercial purposes only, including for participating in the Competition and on Kaggle.com forums, and for academic research and education. The Competition Sponsor reserves the right to disqualify any participant who uses the Competition Data other than as permitted by the Competition Website and these Rules.</p>\n<p>B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.</p>\n<p>C. External Data. The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset. Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020). Please specify any and all external data and/or models used for training when uploading results.</p>\n<p>Participants are allowed to collect additional annotations on the provided training sets. Participants are not allowed to collect annotations on the test set. Teams should specify that they collected additional annotations when submitting results.</p>",
      "rawMarkdown": "**It looks like your submission is against the rules. Sadly, no one is reading the rules these days.**\n\n7. COMPETITION DATA.\n\n\"Competition Data\" means the data or datasets available from the Competition Website for the purpose of use in the Competition, including any prototype or executable code provided on the Competition Website. The Competition Data will contain private and public test sets. Which data belongs to which set will not be made available to participants.\n\nA. Data Access and Use.\n\nCompetition Use and Non-Commercial & Academic Research: You may access and use the Competition Data for non-commercial purposes only, including for participating in the Competition and on Kaggle.com forums, and for academic research and education. The Competition Sponsor reserves the right to disqualify any participant who uses the Competition Data other than as permitted by the Competition Website and these Rules.\n\nB. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.\n\nC. External Data. The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset. Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020). Please specify any and all external data and/or models used for training when uploading results.\n\nParticipants are allowed to collect additional annotations on the provided training sets. Participants are not allowed to collect annotations on the test set. Teams should specify that they collected additional annotations when submitting results.",
      "votes": null
    },
    {
      "id": "1818121",
      "postDate": "06/12/2022 09:49:33",
      "content": "<p>Technically I agree with you.</p>\n<blockquote>\n  <p>The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset.</p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Solution</th>\n<th>External dataset</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328281#1808228\" target=\"_blank\">1st place</a></td>\n<td>.. very small amount of data from FGVC8, only from classes where &lt;8 samples existed in FGVC9.</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328345\" target=\"_blank\">2nd place</a></td>\n<td>Hotels-50k</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328237\" target=\"_blank\">3rd place</a></td>\n<td>FGVC8 Hotel-ID 2021</td>\n</tr>\n</tbody>\n</table>\n<p>In <a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/317922#1771024\" target=\"_blank\">Use of External Data Sets</a> discussion the competition host said: </p>\n<blockquote>\n  <p>However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>\n</blockquote>\n<p>The rules were not updated though so you are right and none of the solutions is compliant with the current competition rules. The host did not even officially announced that external datasets are allowed the only information was \"<em>If an external dataset fits that goal, then by all means, use it.</em>\". Because the leaderboard is finalized and these solutions were accepted I guess the competition host did not mind.</p>\n<p>I agree that the main goal here is to develop the best approach but even though there are no points and prizes awarded for the competition I think this was not handled well.</p>\n<p><strong>I would encourage everyone to publish their solution and there is a chance you might still be invited to present at the workshop even though you are not in top 3.</strong></p>\n<blockquote>\n  <p>A panel will review the top submissions for the competition based on the description of the methods provided. From this, a subset may be invited to present their results at the workshop.</p>\n</blockquote>",
      "rawMarkdown": "Technically I agree with you.\n> The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset.\n\n| Solution | External dataset |\n| --- | --- |\n| [1st place](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328281#1808228) | .. very small amount of data from FGVC8, only from classes where <8 samples existed in FGVC9. |\n| [2nd place](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328345) | Hotels-50k |\n| [3rd place](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328237) | FGVC8 Hotel-ID 2021 |\n\nIn [Use of External Data Sets](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/317922#1771024) discussion the competition host said: \n>  However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.\n\nThe rules were not updated though so you are right and none of the solutions is compliant with the current competition rules. The host did not even officially announced that external datasets are allowed the only information was \"*If an external dataset fits that goal, then by all means, use it.*\". Because the leaderboard is finalized and these solutions were accepted I guess the competition host did not mind.\n\nI agree that the main goal here is to develop the best approach but even though there are no points and prizes awarded for the competition I think this was not handled well.\n\n**I would encourage everyone to publish their solution and there is a chance you might still be invited to present at the workshop even though you are not in top 3.**\n> A panel will review the top submissions for the competition based on the description of the methods provided. From this, a subset may be invited to present their results at the workshop.",
      "votes": null
    },
    {
      "id": "1818292",
      "postDate": "06/12/2022 13:30:03",
      "content": "<p>Thanks for <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a>  <a href=\"https://www.kaggle.com/michaln\" target=\"_blank\">@michaln</a>. I totally agree with Micael.<br>\nI found the host statement and then I decide to use external data. I understand the host allowed external data.  </p>\n<blockquote>\n  <p>However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>\n</blockquote>",
      "rawMarkdown": "Thanks for @picekl  @michaln. I totally agree with Micael.\nI found the host statement and then I decide to use external data. I understand the host allowed external data.  \n\n> However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1806887,
      "author_name": "michaln",
      "author_url": "",
      "post_date": "05/31/2022 14:33:09",
      "content": "<p>Nice solution, congrats to 3rd place, 0.711 is impressive result.</p>\n<p>Did you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?</p>\n<p>Do you plan to release the source code? I am curious about the ArcFaceSubcenter + Dynamic Margin :-)</p>\n<p>What is DOLG?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1806891,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "05/31/2022 14:39:09",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Did you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?</p>\n</blockquote>\n<p>I create mask using albumentations.</p>\n<blockquote>\n  <p>Do you plan to release the source code?</p>\n</blockquote>\n<p>No. </p>\n<blockquote>\n  <p>I am curious about the ArcFaceSubcenter + Dynamic Margin :-)</p>\n</blockquote>\n<p>ArcFaceSubcenter + Dynamic Margin is used for GLR 2020 3rd place.<br>\n<a href=\"https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\" target=\"_blank\">https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757</a><br>\nit have source code.</p>\n<blockquote>\n  <p>What is DOLG?</p>\n</blockquote>\n<p>DOLG is used for GLR2021 1st Solution<br>\nI refer to it code.<br>\n<a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place</a><br>\n<a href=\"https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/models/ch_mdl_dolg_efficientnet.py#L248\" target=\"_blank\">https://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/models/ch_mdl_dolg_efficientnet.py#L248</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1806896,
          "author_name": "michaln",
          "author_url": "",
          "post_date": "05/31/2022 14:43:37",
          "content": "<p>Cool, thanks for the links :-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1807125,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "05/31/2022 18:24:05",
      "content": "<p>Nice solution!  You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me.  I also like you got subcenter + dynamic margin to work well, for me plain vanilla arcface performed better but intuitively subcenter should work well due to the different scenes within the same hotel (ie bathroom vs bedroom)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1807356,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "06/01/2022 00:05:00",
          "content": "<p>Thank you!<br>\nI read your solution that is also great. (Blind Flip etc..)</p>\n<blockquote>\n  <p>You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me</p>\n</blockquote>\n<p>heavier model is better(ConxNeXt XLarge)<br>\nIf I have more resouces, I tried many models ensemble.</p>\n<blockquote>\n  <p>I also like you got subcenter + dynamic margin </p>\n</blockquote>\n<p>I experiment both ArcFace vanilla and ArcFace Subcenter + Dynamic margin, better is ArcFace Subcenter + Dynamic margin.  </p>\n<blockquote>\n  <p>due to the different scenes within the same hotel (ie bathroom vs bedroom)</p>\n</blockquote>\n<p>I agree, Hotel have difference scenes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1818044,
      "author_name": "picekl",
      "author_url": "",
      "post_date": "06/12/2022 07:06:41",
      "content": "<p><strong>It looks like your submission is against the rules. Sadly, no one is reading the rules these days.</strong></p>\n<ol>\n<li>COMPETITION DATA.</li>\n</ol>\n<p>\"Competition Data\" means the data or datasets available from the Competition Website for the purpose of use in the Competition, including any prototype or executable code provided on the Competition Website. The Competition Data will contain private and public test sets. Which data belongs to which set will not be made available to participants.</p>\n<p>A. Data Access and Use.</p>\n<p>Competition Use and Non-Commercial &amp; Academic Research: You may access and use the Competition Data for non-commercial purposes only, including for participating in the Competition and on Kaggle.com forums, and for academic research and education. The Competition Sponsor reserves the right to disqualify any participant who uses the Competition Data other than as permitted by the Competition Website and these Rules.</p>\n<p>B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.</p>\n<p>C. External Data. The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset. Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020). Please specify any and all external data and/or models used for training when uploading results.</p>\n<p>Participants are allowed to collect additional annotations on the provided training sets. Participants are not allowed to collect annotations on the test set. Teams should specify that they collected additional annotations when submitting results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1818121,
          "author_name": "michaln",
          "author_url": "",
          "post_date": "06/12/2022 09:49:33",
          "content": "<p>Technically I agree with you.</p>\n<blockquote>\n  <p>The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset.</p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Solution</th>\n<th>External dataset</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328281#1808228\" target=\"_blank\">1st place</a></td>\n<td>.. very small amount of data from FGVC8, only from classes where &lt;8 samples existed in FGVC9.</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328345\" target=\"_blank\">2nd place</a></td>\n<td>Hotels-50k</td>\n</tr>\n<tr>\n<td><a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328237\" target=\"_blank\">3rd place</a></td>\n<td>FGVC8 Hotel-ID 2021</td>\n</tr>\n</tbody>\n</table>\n<p>In <a href=\"https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/317922#1771024\" target=\"_blank\">Use of External Data Sets</a> discussion the competition host said: </p>\n<blockquote>\n  <p>However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>\n</blockquote>\n<p>The rules were not updated though so you are right and none of the solutions is compliant with the current competition rules. The host did not even officially announced that external datasets are allowed the only information was \"<em>If an external dataset fits that goal, then by all means, use it.</em>\". Because the leaderboard is finalized and these solutions were accepted I guess the competition host did not mind.</p>\n<p>I agree that the main goal here is to develop the best approach but even though there are no points and prizes awarded for the competition I think this was not handled well.</p>\n<p><strong>I would encourage everyone to publish their solution and there is a chance you might still be invited to present at the workshop even though you are not in top 3.</strong></p>\n<blockquote>\n  <p>A panel will review the top submissions for the competition based on the description of the methods provided. From this, a subset may be invited to present their results at the workshop.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818292,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "06/12/2022 13:30:03",
          "content": "<p>Thanks for <a href=\"https://www.kaggle.com/picekl\" target=\"_blank\">@picekl</a>  <a href=\"https://www.kaggle.com/michaln\" target=\"_blank\">@michaln</a>. I totally agree with Micael.<br>\nI found the host statement and then I decide to use external data. I understand the host allowed external data.  </p>\n<blockquote>\n  <p>However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1806873": "Thank you for host,competitors,team member @ks2019 \nIt is a great competiton!\nhere is our team summary.\n\n# Summary\n- ArcFaceSubcenter + Dynamic Margin\n- Many Backbones(Swin,ConvNeXt, ResNet200D, EfficientNet etc..)\n- use External Data(FGVC8)\n- use Logits\n\n## Modeling\n|Model|Data|ImageSize|Public|Private|\n| --- | --- |\n|Swin Large|Comp|512|0.587|0.584|\n|ConvNeXt XLarge|Comp + External|384|0.662|0.656|\n|Swin Large|Comp+ External|384|0.659|0.657|\n|EfficientNetB7 + DOLG|Comp + External|448(stride1)|0.644|0.639|\n|ConvNeXt XLarge|Comp + External|512|0.677|0.672|\n|ResNet200D + DOLG|Comp + External|640|0.651|0.654|\n|EfficientNetB6 + DOLG|Comp + External|640|0.635|0.642|\n|EfficientNetV2S + DOLG|Comp + External|1024|0.645|0.642|\n|EfficientNetV2M + DOLG|Comp + External|896|0.666|0.666|\n\n## Pseudo Labeling\nWe use FGVC9 and FGVC8 competiton data but FGVC8 don't have competiton labels.\nWe annotated label to these data using Pseudo Labeling.\n\nFirst, We used KNN matching FGVC9 training and FGVC8 dataset.\nalso if threshold < 0.5, these data used training.\n\nNext, We adjust labeling. FGVC8 have same hotel list.\nI use mean aggregation method\n\n## Prediction\nI use logits because test have a mask, but train do not have a mask.\nI think it is difficult to match training and test so I decide to compare knn vs logits\nLogits is better than knn in my experiments.\n\nMy inference time is about 2hours.\n\n## did not work\n- used Hotel50K\n- gradient checkpointing",
    "1806887": "Nice solution, congrats to 3rd place, 0.711 is impressive result.\n\nDid you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?\n\nDo you plan to release the source code? I am curious about the ArcFaceSubcenter + Dynamic Margin :-)\n\nWhat is DOLG?",
    "1806891": "Thanks!\n\n> Did you handle the occlusions in some special way? Did you generate them for training or remove them in test using the provided masks?\n\nI create mask using albumentations.\n\n> Do you plan to release the source code?\n\nNo. \n\n>  I am curious about the ArcFaceSubcenter + Dynamic Margin :-)\n\nArcFaceSubcenter + Dynamic Margin is used for GLR 2020 3rd place.\nhttps://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\nit have source code.\n\n> What is DOLG?\n\nDOLG is used for GLR2021 1st Solution\nI refer to it code.\nhttps://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place\nhttps://github.com/ChristofHenkel/kaggle-landmark-2021-1st-place/blob/main/models/ch_mdl_dolg_efficientnet.py#L248",
    "1806896": "Cool, thanks for the links :-)",
    "1807125": "Nice solution!  You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me.  I also like you got subcenter + dynamic margin to work well, for me plain vanilla arcface performed better but intuitively subcenter should work well due to the different scenes within the same hotel (ie bathroom vs bedroom)",
    "1807356": "Thank you!\nI read your solution that is also great. (Blind Flip etc..)\n\n> You were able to squeeze a lot of performance out of ConvNeXt, I ended up dropping it late because it didn't perform as well as others for me\n\nheavier model is better(ConxNeXt XLarge)\nIf I have more resouces, I tried many models ensemble.\n\n>  I also like you got subcenter + dynamic margin \n\nI experiment both ArcFace vanilla and ArcFace Subcenter + Dynamic margin, better is ArcFace Subcenter + Dynamic margin.  \n  \n> due to the different scenes within the same hotel (ie bathroom vs bedroom)\n\nI agree, Hotel have difference scenes.",
    "1818044": "**It looks like your submission is against the rules. Sadly, no one is reading the rules these days.**\n\n7. COMPETITION DATA.\n\n\"Competition Data\" means the data or datasets available from the Competition Website for the purpose of use in the Competition, including any prototype or executable code provided on the Competition Website. The Competition Data will contain private and public test sets. Which data belongs to which set will not be made available to participants.\n\nA. Data Access and Use.\n\nCompetition Use and Non-Commercial & Academic Research: You may access and use the Competition Data for non-commercial purposes only, including for participating in the Competition and on Kaggle.com forums, and for academic research and education. The Competition Sponsor reserves the right to disqualify any participant who uses the Competition Data other than as permitted by the Competition Website and these Rules.\n\nB. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.\n\nC. External Data. The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset. Pretrained models may be used to construct the algorithms from publicly available academic datasets (e.g. ImageNet, iNaturalist 2017-2018, Herbarium 2020). Please specify any and all external data and/or models used for training when uploading results.\n\nParticipants are allowed to collect additional annotations on the provided training sets. Participants are not allowed to collect annotations on the test set. Teams should specify that they collected additional annotations when submitting results.",
    "1818121": "Technically I agree with you.\n> The general rule is that participants should only use the provided training and validation images for training models to classify the test images. We do not want participants crawling the web in search of additional data or using previous versions of this dataset.\n\n| Solution | External dataset |\n| --- | --- |\n| [1st place](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328281#1808228) | .. very small amount of data from FGVC8, only from classes where <8 samples existed in FGVC9. |\n| [2nd place](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328345) | Hotels-50k |\n| [3rd place](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328237) | FGVC8 Hotel-ID 2021 |\n\nIn [Use of External Data Sets](https://www.kaggle.com/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/317922#1771024) discussion the competition host said: \n>  However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it.\n\nThe rules were not updated though so you are right and none of the solutions is compliant with the current competition rules. The host did not even officially announced that external datasets are allowed the only information was \"*If an external dataset fits that goal, then by all means, use it.*\". Because the leaderboard is finalized and these solutions were accepted I guess the competition host did not mind.\n\nI agree that the main goal here is to develop the best approach but even though there are no points and prizes awarded for the competition I think this was not handled well.\n\n**I would encourage everyone to publish their solution and there is a chance you might still be invited to present at the workshop even though you are not in top 3.**\n> A panel will review the top submissions for the competition based on the description of the methods provided. From this, a subset may be invited to present their results at the workshop.",
    "1818292": "Thanks for @picekl  @michaln. I totally agree with Micael.\nI found the host statement and then I decide to use external data. I understand the host allowed external data.  \n\n> However, I will simply say that there are no points awarded for this competition or monetary awards -- the goal of this competition is to help in the development of the best approaches to recognizing hotels to combat human trafficking. If an external dataset fits that goal, then by all means, use it."
  },
  "source": "meta"
}