{
  "id": 328345,
  "title": "2nd solution",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/328345",
  "author_name": "wanghao",
  "post_date": "2022-06-01T02:50:20.973000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to host for organizing this challenge.</p>\n<h3>Mask</h3>\n<p>Statistics of the mean, variance of the size and position of the mask in the test set. Generate mask with the statistical values in training.</p>\n<h3>data preprocess</h3>\n<p>50k+fgvc9<br></p>\n<ul>\n<li><p>data clean<br><br>\nCalculate the md5 of all the data, and delete the data with the same md5 but different categories, keep data in fgvc9(A total of 45,769 categories remain). </p></li>\n<li><p>direction<br><br>\nUsing the direction model, rotate the image to face upward.</p></li>\n</ul>\n<h3>train step</h3>\n<p>step 1. <br><br>\ntrain all data (10~20 epoch)<br><br>\nstep 2. <br><br>\nFine-tuning the data of 3116 categories in fgvc9 (40epoch) (improve ~0.03)</p>\n<h3>loss function</h3>\n<p>Because the same id of the hotel varies greatly, there are bedrooms, bathrooms, etc. so we use sub-center Arcface(k=3, Dynamic Margin) as the loss function.</p>\n<h3>predict</h3>\n<p>There is little difference between using logits and retrieval in my experiments, so I choose the more convenient logits as the prediction.</p>\n<h3>Model</h3>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>input size</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>swin-base-384</td>\n<td>384</td>\n<td>0.707</td>\n<td>0.688</td>\n</tr>\n<tr>\n<td>swin-large-384</td>\n<td>384</td>\n<td>0.704</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>eca_nfnet_l1</td>\n<td>576</td>\n<td>0.700</td>\n<td>0.688</td>\n</tr>\n<tr>\n<td>efficientNetV2-l</td>\n<td>480</td>\n<td>/</td>\n<td>/</td>\n</tr>\n<tr>\n<td>efficientNet-b5</td>\n<td>576</td>\n<td>/</td>\n<td>/</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>/</td>\n<td>0.732</td>\n<td>0.717</td>\n</tr>\n</tbody>\n</table>\n<h3>dit not work</h3>\n<p>Pseudo Labeling of fgvc8</p>",
  "messages": [
    {
      "id": 1807433,
      "postDate": "2022-06-01T02:50:20.973Z",
      "content": "<p>Thanks to host for organizing this challenge.</p>\n<h3>Mask</h3>\n<p>Statistics of the mean, variance of the size and position of the mask in the test set. Generate mask with the statistical values in training.</p>\n<h3>data preprocess</h3>\n<p>50k+fgvc9<br></p>\n<ul>\n<li><p>data clean<br><br>\nCalculate the md5 of all the data, and delete the data with the same md5 but different categories, keep data in fgvc9(A total of 45,769 categories remain). </p></li>\n<li><p>direction<br><br>\nUsing the direction model, rotate the image to face upward.</p></li>\n</ul>\n<h3>train step</h3>\n<p>step 1. <br><br>\ntrain all data (10~20 epoch)<br><br>\nstep 2. <br><br>\nFine-tuning the data of 3116 categories in fgvc9 (40epoch) (improve ~0.03)</p>\n<h3>loss function</h3>\n<p>Because the same id of the hotel varies greatly, there are bedrooms, bathrooms, etc. so we use sub-center Arcface(k=3, Dynamic Margin) as the loss function.</p>\n<h3>predict</h3>\n<p>There is little difference between using logits and retrieval in my experiments, so I choose the more convenient logits as the prediction.</p>\n<h3>Model</h3>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>input size</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>swin-base-384</td>\n<td>384</td>\n<td>0.707</td>\n<td>0.688</td>\n</tr>\n<tr>\n<td>swin-large-384</td>\n<td>384</td>\n<td>0.704</td>\n<td>0.692</td>\n</tr>\n<tr>\n<td>eca_nfnet_l1</td>\n<td>576</td>\n<td>0.700</td>\n<td>0.688</td>\n</tr>\n<tr>\n<td>efficientNetV2-l</td>\n<td>480</td>\n<td>/</td>\n<td>/</td>\n</tr>\n<tr>\n<td>efficientNet-b5</td>\n<td>576</td>\n<td>/</td>\n<td>/</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>/</td>\n<td>0.732</td>\n<td>0.717</td>\n</tr>\n</tbody>\n</table>\n<h3>dit not work</h3>\n<p>Pseudo Labeling of fgvc8</p>",
      "rawMarkdown": " Thanks to host for organizing this challenge.\n \n ### Mask\n Statistics of the mean, variance of the size and position of the mask in the test set. Generate mask with the statistical values in training.\n \n \n ### data preprocess\n 50k+fgvc9<br/>\n * data clean<br/>\n Calculate the md5 of all the data, and delete the data with the same md5 but different categories, keep data in fgvc9(A total of 45,769 categories remain). \n \n * direction<br/>\n Using the direction model, rotate the image to face upward.\n \n \n ### train step\n step 1. <br/>\n train all data (10~20 epoch)<br/>\n step 2. <br/>\nFine-tuning the data of 3116 categories in fgvc9 (40epoch) (improve ~0.03)\n \n \n \n ### loss function\n Because the same id of the hotel varies greatly, there are bedrooms, bathrooms, etc. so we use sub-center Arcface(k=3, Dynamic Margin) as the loss function.\n \n \n ### predict\n There is little difference between using logits and retrieval in my experiments, so I choose the more convenient logits as the prediction.\n \n \n ### Model\nBackbone | input size | public | private\n---|---|---|---\nswin-base-384|384|0.707|0.688\nswin-large-384|384|0.704|0.692\neca_nfnet_l1|576|0.700|0.688\nefficientNetV2-l|480|/|/\nefficientNet-b5|576|/|/\nensemble|/|0.732|0.717\n\n\n ### dit not work\n Pseudo Labeling of fgvc8\n \n \n \n ",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1807433": " Thanks to host for organizing this challenge.\n \n ### Mask\n Statistics of the mean, variance of the size and position of the mask in the test set. Generate mask with the statistical values in training.\n \n \n ### data preprocess\n 50k+fgvc9<br/>\n * data clean<br/>\n Calculate the md5 of all the data, and delete the data with the same md5 but different categories, keep data in fgvc9(A total of 45,769 categories remain). \n \n * direction<br/>\n Using the direction model, rotate the image to face upward.\n \n \n ### train step\n step 1. <br/>\n train all data (10~20 epoch)<br/>\n step 2. <br/>\nFine-tuning the data of 3116 categories in fgvc9 (40epoch) (improve ~0.03)\n \n \n \n ### loss function\n Because the same id of the hotel varies greatly, there are bedrooms, bathrooms, etc. so we use sub-center Arcface(k=3, Dynamic Margin) as the loss function.\n \n \n ### predict\n There is little difference between using logits and retrieval in my experiments, so I choose the more convenient logits as the prediction.\n \n \n ### Model\nBackbone | input size | public | private\n---|---|---|---\nswin-base-384|384|0.707|0.688\nswin-large-384|384|0.704|0.692\neca_nfnet_l1|576|0.700|0.688\nefficientNetV2-l|480|/|/\nefficientNet-b5|576|/|/\nensemble|/|0.732|0.717\n\n\n ### dit not work\n Pseudo Labeling of fgvc8\n \n \n \n "
  }
}