{
  "id": 96149,
  "title": "2nd place solution [LB 0.667]",
  "url": "/competitions/imet-2019-fgvc6/discussion/96149",
  "author_name": "Pavel Tsai",
  "post_date": "2019-06-18T11:11:40.182000",
  "votes": -34,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks for the competition and congratulations of all participants! Here is a summary of my solution.</p>\n\n<h1>Models</h1>\n\n<p>SE-ResNeXt-50 and SE-DenseNet-161 with PartialConv\nI didn't have much computational power: only 2 x 1080Ti, so I didn't use very large models. I also tried other resnets and densenets, but these two worked best for me.\nCV: 6 folds with iterative stratification.</p>\n\n<h1>Augmentations</h1>\n\n<p>Crop to 640x640 and then scale to 320x320. Random crop on train and center crop on validation and test. I also used random horizontal flip, rotate, gamma, brightness and contrast.\nDuring test time: original + flipped augmentations.</p>\n\n<h1>Training process</h1>\n\n<p>All models were trained in 3 stages: first: with freezed encoder, second: whole network, third: tags only. Batch accumulation was used to achieve batch size 512. I used AdamW optimizer with weight decay 0.01 and lr=1e-4, and then SGD lr=1e-3 (stage 3). And Cosine scheduler with warmup in every stage. Label smoothing and mixup also worked in this competition.</p>\n\n<h1>Threshold</h1>\n\n<p>Separate thresholds for cultures and tags chosen by validation. For tags threshold appeared to be smaller.</p>\n\n<h1>Cleaning and pseudo labelling</h1>\n\n<p>I had two stages of cleaning and pseudo labelling the dataset. Removing high error samples and pseudo labeling greatly improved public LB score, but it seemed to be an overfit. I watched on high error samples and decided to throw them away or not. But anyway, this painstaking dataset cleaning was important.</p>\n\n<h1>Second layer model</h1>\n\n<p>As a second layer model I tried to use LGBM. The technique was quite similar to my quickdraw solution. I put top 30 tags and top 20 cultures to the dataset for LGMB. I mean, most confident classes. Then, predicted <em>is it true, that this class is in answer</em> for every class from this 50. But, unfortunately I didn't have enough time to implemented it in public kernel. But anyway, simple folds and checkpoints averaging worked fine on private here.</p>",
  "messages": [
    {
      "id": 555258,
      "postDate": "2019-06-18T16:32:19.977Z",
      "content": "<p>To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.</p>",
      "rawMarkdown": "To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.",
      "votes": 10
    },
    {
      "id": 556696,
      "postDate": "2019-06-20T14:35:32.977Z",
      "content": "<p>Come on, almost everyone (except some people you know) here in top 10 using V100 or many GPUs to run many experiments to find the best results, and you are using only 2 1080Ti? Come on, if you are so experienced, where are you on all other CV competitions?</p>",
      "rawMarkdown": "Come on, almost everyone (except some people you know) here in top 10 using V100 or many GPUs to run many experiments to find the best results, and you are using only 2 1080Ti? Come on, if you are so experienced, where are you on all other CV competitions?",
      "votes": 5
    },
    {
      "id": 555764,
      "postDate": "2019-06-19T11:42:56.930Z",
      "content": "<p>WOW， 2 * 1080TI !!!\nX5 shared tricks with you, why not DGX?! \nThey are so stingy!\nSad face！</p>",
      "rawMarkdown": "WOW， 2 * 1080TI !!!\nX5 shared tricks with you, why not DGX?! \nThey are so stingy!\nSad face！",
      "votes": 3
    },
    {
      "id": 555491,
      "postDate": "2019-06-19T00:52:37.887Z",
      "content": "<p>Telling a story is too easy.</p>",
      "rawMarkdown": "Telling a story is too easy.",
      "votes": 3
    },
    {
      "id": 555043,
      "postDate": "2019-06-18T11:11:40.183Z",
      "content": "<p>Thanks for the competition and congratulations of all participants! Here is a summary of my solution.</p>\n\n<h1>Models</h1>\n\n<p>SE-ResNeXt-50 and SE-DenseNet-161 with PartialConv\nI didn't have much computational power: only 2 x 1080Ti, so I didn't use very large models. I also tried other resnets and densenets, but these two worked best for me.\nCV: 6 folds with iterative stratification.</p>\n\n<h1>Augmentations</h1>\n\n<p>Crop to 640x640 and then scale to 320x320. Random crop on train and center crop on validation and test. I also used random horizontal flip, rotate, gamma, brightness and contrast.\nDuring test time: original + flipped augmentations.</p>\n\n<h1>Training process</h1>\n\n<p>All models were trained in 3 stages: first: with freezed encoder, second: whole network, third: tags only. Batch accumulation was used to achieve batch size 512. I used AdamW optimizer with weight decay 0.01 and lr=1e-4, and then SGD lr=1e-3 (stage 3). And Cosine scheduler with warmup in every stage. Label smoothing and mixup also worked in this competition.</p>\n\n<h1>Threshold</h1>\n\n<p>Separate thresholds for cultures and tags chosen by validation. For tags threshold appeared to be smaller.</p>\n\n<h1>Cleaning and pseudo labelling</h1>\n\n<p>I had two stages of cleaning and pseudo labelling the dataset. Removing high error samples and pseudo labeling greatly improved public LB score, but it seemed to be an overfit. I watched on high error samples and decided to throw them away or not. But anyway, this painstaking dataset cleaning was important.</p>\n\n<h1>Second layer model</h1>\n\n<p>As a second layer model I tried to use LGBM. The technique was quite similar to my quickdraw solution. I put top 30 tags and top 20 cultures to the dataset for LGMB. I mean, most confident classes. Then, predicted <em>is it true, that this class is in answer</em> for every class from this 50. But, unfortunately I didn't have enough time to implemented it in public kernel. But anyway, simple folds and checkpoints averaging worked fine on private here.</p>",
      "rawMarkdown": "Thanks for the competition and congratulations of all participants! Here is a summary of my solution.\n#Models#\n SE-ResNeXt-50 and SE-DenseNet-161 with PartialConv\nI didn't have much computational power: only 2 x 1080Ti, so I didn't use very large models. I also tried other resnets and densenets, but these two worked best for me.\nCV: 6 folds with iterative stratification.\n#Augmentations#\nCrop to 640x640 and then scale to 320x320. Random crop on train and center crop on validation and test. I also used random horizontal flip, rotate, gamma, brightness and contrast.\nDuring test time: original + flipped augmentations.\n#Training process#\nAll models were trained in 3 stages: first: with freezed encoder, second: whole network, third: tags only. Batch accumulation was used to achieve batch size 512. I used AdamW optimizer with weight decay 0.01 and lr=1e-4, and then SGD lr=1e-3 (stage 3). And Cosine scheduler with warmup in every stage. Label smoothing and mixup also worked in this competition.\n#Threshold#\nSeparate thresholds for cultures and tags chosen by validation. For tags threshold appeared to be smaller.\n#Cleaning and pseudo labelling#\nI had two stages of cleaning and pseudo labelling the dataset. Removing high error samples and pseudo labeling greatly improved public LB score, but it seemed to be an overfit. I watched on high error samples and decided to throw them away or not. But anyway, this painstaking dataset cleaning was important.\n# Second layer model#\nAs a second layer model I tried to use LGBM. The technique was quite similar to my quickdraw solution. I put top 30 tags and top 20 cultures to the dataset for LGMB. I mean, most confident classes. Then, predicted _is it true, that this class is in answer_ for every class from this 50. But, unfortunately I didn't have enough time to implemented it in public kernel. But anyway, simple folds and checkpoints averaging worked fine on private here.",
      "votes": -34
    },
    {
      "id": 562207,
      "postDate": "2019-06-27T04:09:33.193Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 555258,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2019-06-18T16:32:19.977000",
      "content": "<p>To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 556696,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2019-06-20T14:35:32.977000",
      "content": "<p>Come on, almost everyone (except some people you know) here in top 10 using V100 or many GPUs to run many experiments to find the best results, and you are using only 2 1080Ti? Come on, if you are so experienced, where are you on all other CV competitions?</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 555764,
      "author_name": "earhian",
      "author_url": "",
      "post_date": "2019-06-19T11:42:56.930000",
      "content": "<p>WOW， 2 * 1080TI !!!\nX5 shared tricks with you, why not DGX?! \nThey are so stingy!\nSad face！</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 555491,
      "author_name": "qrfaction",
      "author_url": "",
      "post_date": "2019-06-19T00:52:37.887000",
      "content": "<p>Telling a story is too easy.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 562207,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-27T04:09:33.193000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "555258": "To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.",
    "556696": "Come on, almost everyone (except some people you know) here in top 10 using V100 or many GPUs to run many experiments to find the best results, and you are using only 2 1080Ti? Come on, if you are so experienced, where are you on all other CV competitions?",
    "555764": "WOW， 2 * 1080TI !!!\nX5 shared tricks with you, why not DGX?! \nThey are so stingy!\nSad face！",
    "555491": "Telling a story is too easy.",
    "555043": "Thanks for the competition and congratulations of all participants! Here is a summary of my solution.\n#Models#\n SE-ResNeXt-50 and SE-DenseNet-161 with PartialConv\nI didn't have much computational power: only 2 x 1080Ti, so I didn't use very large models. I also tried other resnets and densenets, but these two worked best for me.\nCV: 6 folds with iterative stratification.\n#Augmentations#\nCrop to 640x640 and then scale to 320x320. Random crop on train and center crop on validation and test. I also used random horizontal flip, rotate, gamma, brightness and contrast.\nDuring test time: original + flipped augmentations.\n#Training process#\nAll models were trained in 3 stages: first: with freezed encoder, second: whole network, third: tags only. Batch accumulation was used to achieve batch size 512. I used AdamW optimizer with weight decay 0.01 and lr=1e-4, and then SGD lr=1e-3 (stage 3). And Cosine scheduler with warmup in every stage. Label smoothing and mixup also worked in this competition.\n#Threshold#\nSeparate thresholds for cultures and tags chosen by validation. For tags threshold appeared to be smaller.\n#Cleaning and pseudo labelling#\nI had two stages of cleaning and pseudo labelling the dataset. Removing high error samples and pseudo labeling greatly improved public LB score, but it seemed to be an overfit. I watched on high error samples and decided to throw them away or not. But anyway, this painstaking dataset cleaning was important.\n# Second layer model#\nAs a second layer model I tried to use LGBM. The technique was quite similar to my quickdraw solution. I put top 30 tags and top 20 cultures to the dataset for LGMB. I mean, most confident classes. Then, predicted _is it true, that this class is in answer_ for every class from this 50. But, unfortunately I didn't have enough time to implemented it in public kernel. But anyway, simple folds and checkpoints averaging worked fine on private here.",
    "562207": ""
  }
}