{
  "id": 160724,
  "title": "1st Place Solution to Semi-Supervised Recognition Challenge - FGVC7",
  "url": "/competitions/semi-inat-2020/writeups/kaggle-1st-place-solution-to-semi-supervised-recog",
  "author_name": "",
  "post_date": "2020-06-22T12:23:23.516176100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>(1) Train a base classifier on the labeled data (train and validtion) using ResNest, SeNet and EfficientNet networks with resolution  (448, 520, 640, 720 or 800)</p>\n\n<p>(2) Adopt a progressive reprensentative labeling method to pseudo-label the unlabeled dataset  (+5~7%)\n    2.1 construct a kNN graph by the feature of the labeled and unlabeled data, where the feature is extracted by the classifier \n    2.2 sample data with high indegree and high confidence, which are representative and reliable for the entire data space\n    2.3 single model or ensemble models pseudo-label on the sampled representative data .</p>\n\n<p>(3) Locate object in the image by unsupervised method Class Activation mapping (CAM) by the base classifier </p>\n\n<p>(4) Adopt a two-stream network that takes as input located image (CAM) and origin image and concates their features to classification (+0.5~0.7%)</p>\n\n<p>(5) finetune the model using pseudo-labels trained with the clear train and validation set with consistency regularization. (+1.0~1.5%)</p>\n\n<p>(7) Use GeMPooling and larger-resolution for testing (+0.5~1.0%)</p>\n\n<p>(8) Ensemble differnet models (+0.3%~0.6%) </p>",
  "messages": [
    {
      "id": "896764",
      "postDate": "06/22/2020 12:23:23",
      "content": "<p>(1) Train a base classifier on the labeled data (train and validtion) using ResNest, SeNet and EfficientNet networks with resolution  (448, 520, 640, 720 or 800)</p>\n\n<p>(2) Adopt a progressive reprensentative labeling method to pseudo-label the unlabeled dataset  (+5~7%)\n    2.1 construct a kNN graph by the feature of the labeled and unlabeled data, where the feature is extracted by the classifier \n    2.2 sample data with high indegree and high confidence, which are representative and reliable for the entire data space\n    2.3 single model or ensemble models pseudo-label on the sampled representative data .</p>\n\n<p>(3) Locate object in the image by unsupervised method Class Activation mapping (CAM) by the base classifier </p>\n\n<p>(4) Adopt a two-stream network that takes as input located image (CAM) and origin image and concates their features to classification (+0.5~0.7%)</p>\n\n<p>(5) finetune the model using pseudo-labels trained with the clear train and validation set with consistency regularization. (+1.0~1.5%)</p>\n\n<p>(7) Use GeMPooling and larger-resolution for testing (+0.5~1.0%)</p>\n\n<p>(8) Ensemble differnet models (+0.3%~0.6%) </p>",
      "rawMarkdown": "(1) Train a base classifier on the labeled data (train and validtion) using ResNest, SeNet and EfficientNet networks with resolution  (448, 520, 640, 720 or 800)\n\n(2) Adopt a progressive reprensentative labeling method to pseudo-label the unlabeled dataset  (+5~7%)\n    2.1 construct a kNN graph by the feature of the labeled and unlabeled data, where the feature is extracted by the classifier \n    2.2 sample data with high indegree and high confidence, which are representative and reliable for the entire data space\n    2.3 single model or ensemble models pseudo-label on the sampled representative data .\n\n(3) Locate object in the image by unsupervised method Class Activation mapping (CAM) by the base classifier \n\n(4) Adopt a two-stream network that takes as input located image (CAM) and origin image and concates their features to classification (+0.5~0.7%)\n\n(5) finetune the model using pseudo-labels trained with the clear train and validation set with consistency regularization. (+1.0~1.5%)\n \n(7) Use GeMPooling and larger-resolution for testing (+0.5~1.0%)\n   \n(8) Ensemble differnet models (+0.3%~0.6%)",
      "votes": null
    },
    {
      "id": "896776",
      "postDate": "06/22/2020 12:29:08",
      "content": "<p>Useless method:(1) unsupervised method like MOCO on all data as the pre-trained model does not improve the performance; Consider the out-of-class data as the same class and train a 200+1 classifier to pseudo-label in-class data or as pre-trained model\n(2) bilinear operations (e.g., compact bilinear, factorized bilinear) on labeled data only improves a lot but fail when applying on the pseudo-labeled data\n(3) some data balancing tricks like class-balanced loss and focal loss does not help\n(4) some other semi-supervised methods like fixmatch and mixmatch\n(5) some augmentation techniques (e.g., random erase;cutout；colorjitter) \n(6) multi-scale training does not help\n(7)  pseudo-labels and clear data combine training does not help</p>",
      "rawMarkdown": "Useless method:(1) unsupervised method like MOCO on all data as the pre-trained model does not improve the performance; Consider the out-of-class data as the same class and train a 200+1 classifier to pseudo-label in-class data or as pre-trained model\n(2) bilinear operations (e.g., compact bilinear, factorized bilinear) on labeled data only improves a lot but fail when applying on the pseudo-labeled data\n(3) some data balancing tricks like class-balanced loss and focal loss does not help\n(4) some other semi-supervised methods like fixmatch and mixmatch\n(5) some augmentation techniques (e.g., random erase;cutout；colorjitter) \n(6) multi-scale training does not help\n(7)  pseudo-labels and clear data combine training does not help",
      "votes": null
    },
    {
      "id": "898815",
      "postDate": "06/23/2020 19:13:22",
      "content": "<p>thanks for sharing.\nA few questions:\n- can you elaborate why you used large inputs when images are provided with max size 512 pixels?\n- can you estimate computational budget for training and inference for your winning solution?\n- how many models you used eventually in your winning solution and how many feed forward passes? What would be your best score with a single feed forward pass?</p>",
      "rawMarkdown": "thanks for sharing.\nA few questions:\n- can you elaborate why you used large inputs when images are provided with max size 512 pixels?\n- can you estimate computational budget for training and inference for your winning solution?\n- how many models you used eventually in your winning solution and how many feed forward passes? What would be your best score with a single feed forward pass?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 896776,
      "author_name": "welarn",
      "author_url": "",
      "post_date": "06/22/2020 12:29:08",
      "content": "<p>Useless method:(1) unsupervised method like MOCO on all data as the pre-trained model does not improve the performance; Consider the out-of-class data as the same class and train a 200+1 classifier to pseudo-label in-class data or as pre-trained model\n(2) bilinear operations (e.g., compact bilinear, factorized bilinear) on labeled data only improves a lot but fail when applying on the pseudo-labeled data\n(3) some data balancing tricks like class-balanced loss and focal loss does not help\n(4) some other semi-supervised methods like fixmatch and mixmatch\n(5) some augmentation techniques (e.g., random erase;cutout；colorjitter) \n(6) multi-scale training does not help\n(7)  pseudo-labels and clear data combine training does not help</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 898815,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "06/23/2020 19:13:22",
      "content": "<p>thanks for sharing.\nA few questions:\n- can you elaborate why you used large inputs when images are provided with max size 512 pixels?\n- can you estimate computational budget for training and inference for your winning solution?\n- how many models you used eventually in your winning solution and how many feed forward passes? What would be your best score with a single feed forward pass?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "896764": "(1) Train a base classifier on the labeled data (train and validtion) using ResNest, SeNet and EfficientNet networks with resolution  (448, 520, 640, 720 or 800)\n\n(2) Adopt a progressive reprensentative labeling method to pseudo-label the unlabeled dataset  (+5~7%)\n    2.1 construct a kNN graph by the feature of the labeled and unlabeled data, where the feature is extracted by the classifier \n    2.2 sample data with high indegree and high confidence, which are representative and reliable for the entire data space\n    2.3 single model or ensemble models pseudo-label on the sampled representative data .\n\n(3) Locate object in the image by unsupervised method Class Activation mapping (CAM) by the base classifier \n\n(4) Adopt a two-stream network that takes as input located image (CAM) and origin image and concates their features to classification (+0.5~0.7%)\n\n(5) finetune the model using pseudo-labels trained with the clear train and validation set with consistency regularization. (+1.0~1.5%)\n \n(7) Use GeMPooling and larger-resolution for testing (+0.5~1.0%)\n   \n(8) Ensemble differnet models (+0.3%~0.6%)",
    "896776": "Useless method:(1) unsupervised method like MOCO on all data as the pre-trained model does not improve the performance; Consider the out-of-class data as the same class and train a 200+1 classifier to pseudo-label in-class data or as pre-trained model\n(2) bilinear operations (e.g., compact bilinear, factorized bilinear) on labeled data only improves a lot but fail when applying on the pseudo-labeled data\n(3) some data balancing tricks like class-balanced loss and focal loss does not help\n(4) some other semi-supervised methods like fixmatch and mixmatch\n(5) some augmentation techniques (e.g., random erase;cutout；colorjitter) \n(6) multi-scale training does not help\n(7)  pseudo-labels and clear data combine training does not help",
    "898815": "thanks for sharing.\nA few questions:\n- can you elaborate why you used large inputs when images are provided with max size 512 pixels?\n- can you estimate computational budget for training and inference for your winning solution?\n- how many models you used eventually in your winning solution and how many feed forward passes? What would be your best score with a single feed forward pass?"
  },
  "source": "meta"
}