{
  "id": 71433,
  "title": "Solution description (3rd place)",
  "url": "/competitions/inclusive-images-challenge/writeups/worldwideinclusive-solution-description-3rd-place",
  "author_name": "",
  "post_date": "2018-11-25T20:23:34.020Z",
  "votes": 40,
  "comment_count": 14,
  "views": 0,
  "content": "<h2>Some thoughts</h2>\n\n<p>1) After some analysis of data we noticed that most of differences between train and test (stage 1) was more due to different labelling distributions, than to different geo locations. For example: average number of labels per image in train was 4.083 (maximum labels: 127) while in test only 2.386 (maximum labels: 9). So by assuming stage 1 and stage 2 have similar labeling distributions, we can use very similar, but less risky (potentially less overfitting), thresholds for stage 2 as in stage 1.</p>\n\n<p>2) Since we only have one model to choose in stage 2, in order to balance risk v.s. reward, we chose a slightly more conservative approach than our best LB model in stage 1 – our final selected model scored around 10th position on stage 1 LB, which has a slightly more conservative threshold optimization.</p>\n\n<h2>Models and scores</h2>\n\n<p>Our solution consists of 7 models: <em>ResNet50</em>, <em>Xception</em>, <em>Inception ResNet v2</em> x 5. All neural nets had the same final Dense (FC) layer with ~7K neurons for all available classes with sigmoid activation. However, the five inception-resnet have slightly different training and image augmentation approaches. All nets had almost the same score ~0.5 at stage 1 after threshold tuning. The best one was Inception resnet v2. ResNet50 was trained on 336x336 resolution, others on 299 resolution. I believe single model is enough here to get almost the same score (late submission for random single model gives: 0.337). For example ensemble with majority voting on stage 1: 0.505 + 0.501 + 0.511 LB: 0.516</p>\n\n<h2>Train/validation split</h2>\n\n<ul>\n<li>We used all labels for training: human and computer generated.</li>\n<li>We also used probabilities from computer generated labels rather than fixed 1.0 value.</li>\n<li>We didn't use boxes</li>\n<li>Split was non-uniform stratified by classes (see code below):</li>\n</ul>\n\n<pre>    if count &lt; 50:\n        part = count // 10\n    elif count &lt; 100:\n        part = count // 20\n    elif count &lt; 1000:\n        part = count // 50\n    elif count &lt; 10000:\n        part = count // 100\n    elif count &lt; 100000:\n        part = count // 500\n    else:\n        part = count // 1000\n</pre>\n\n<ul>\n<li>Train: 1728299 images Valid: 14743 images</li>\n</ul>\n\n<h2>Training process</h2>\n\n<ul>\n<li>Nothing really special: random samples from training images</li>\n<li>Training from scratch (not using ImageNet pretrained weights) due to competition rules.</li>\n<li>We used default binary_crossentropy loss for this competition</li>\n<li>For validation during training we used tuning labels from stage_1 and calculate F2 score on each epoch with different thresholds (THRs). At the final epoch for inception resnet v2:</li>\n</ul>\n\n<pre>    loss: 0.0019 - f2beta_loss: -5.6904e-01 - fbeta: 0.7356\n    val_loss: 0.0027 - val_f2beta_loss: -5.6610e-01 - val_fbeta: 0.7038 \n    F2Beta score tuning labels: 0.373804 Optimal THR: 0.7\n</pre>\n\n<ul>\n<li>F2Beta score on tuning labels keep improving with improvement of validation, so it mostly used for checking rather than early stopping.</li>\n<li>Local score on tuning labels was almost the same as on leaderboard.</li>\n<li>Heavy augmentations. Very useful library <a href=\"https://github.com/albu/albumentations\">albumentations</a> helped us here:</li>\n</ul>\n\n<pre>    def strong_aug(p=.5):\n        return Compose([\n            HorizontalFlip(),\n            OneOf([\n                IAAAdditiveGaussianNoise(),\n                GaussNoise(),\n            ], p=0.2),\n            OneOf([\n                MotionBlur(p=.2),\n                MedianBlur(blur_limit=3, p=.1),\n                Blur(blur_limit=3, p=.1),\n            ], p=0.2),\n            ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.2, rotate_limit=10, p=0.1),\n            OneOf([\n                OpticalDistortion(p=0.3),\n                GridDistortion(p=0.1),\n                IAAPiecewiseAffine(p=0.3),\n            ], p=0.2),\n            OneOf([\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomContrast(),\n                RandomBrightness(),\n            ], p=0.3),\n            HueSaturationValue(p=0.3),\n            ToGray(p=0.05),\n            JpegCompression(p=0.2, quality_lower=55, quality_upper=99),\n            ElasticTransform(p=0.1),\n        ], p=p)\n</pre>\n\n<ul>\n<li>One epoch 40000 images. To train model we used 200-500 epochs. Training further continue improving model, but with very small effect.</li>\n<li>For xception and 3 more inception-resnet, we use a different image augmentation strategy. For example, one inception-resnet has image augmentation like this:</li>\n</ul>\n\n<pre>gen = ImageDataGenerator(horizontal_flip = True, vertical_flip = True, width_shift_range = 0.1, height_shift_range = 0.1, channel_shift_range=0.1, shear_range = 0.1, zoom_range = 0.1,  rotation_range = 10, preprocessing_function=inception_resnet_preprocess_input)</pre>\n\n<h2>Threshold tuning</h2>\n\n<p>It was the most important part of the competition. We needed to output discrete set of classes for each image based on probabilities from neural networks. We had wide range of strategies to find optimal thresholds and the F2Beta metric is very sensitive. For example:</p>\n\n<ul>\n<li>F2Beta on validation images was about: ~0.73 </li>\n<li>F2Beta on tuning labels with single threshold: ~0.35 </li>\n<li>LB Score F2Beta with optimized thresholds: ~0.5 LB score</li>\n</ul>\n\n<p>Optimization process was made on tuning labels. It had 4 parameters:</p>\n\n<ul>\n<li>minimum probability for search [MinPS]: 0.01</li>\n<li>maximum probability for search [MaxPS]: 0.99</li>\n<li>default probability: 0.99 - We use this probability if class had no entry in tuning labels. Actually, from 7K classes only ~500 had entry in tuning labels. Default prob 0.99 means that model must be &gt;99% sure to use this class.</li>\n<li>minimum number of entries in class: 1</li>\n</ul>\n\n<p>For conservative models in our final submission we used [0.1-0.9] ranges with 2-3 minimum number entries in class and lower default probability (0.8, 0.9). While on Stage 1 LB it gives worse (less overfitting) result.\nStart from setting all thresholds to 0.5 then iterate all over the classes and find threshold for each class in range [MinPS; MaxPS] which maximize the F2Beta score. Do 2-3 overall iterations until F2Beta score stops increasing.</p>\n\n<h2>Ensembles</h2>\n\n<p>We used simple majority voting for ensemble. If at least 4 (out of 7) models predicted a class then this class goes to the submission file.</p>\n\n<h2>Post processing</h2>\n\n<p>For F2Beta metric it's better to output something if model didn't predict any class. So rows which have empty labels, we will use UNION across all 7 models to regenerate labels. If all 7 models have no predictions, we then will output 3 most common classes found in training set.</p>\n\n<h2>Failed experiments</h2>\n\n<ul>\n<li>We also tried ResNet152 on 448x448 images, but it gave worse results.</li>\n<li>MobileNet on 128x128 images - just to check if it could give comparable result, but it didn't</li>\n<li>SE-ResNext101 for Keras from this repo: <a href=\"https://github.com/titu1994/keras-squeeze-excite-network\">https://github.com/titu1994/keras-squeeze-excite-network</a> - very very slow and almost not converges comparing to other models.</li>\n<li>Label selection/filtering based on “class activation maps” (size of activation field, location, etc.)</li>\n</ul>",
  "messages": [
    {
      "id": "420465",
      "postDate": "11/13/2018 17:03:07",
      "content": "<h2>Some thoughts</h2>\n\n<p>1) After some analysis of data we noticed that most of differences between train and test (stage 1) was more due to different labelling distributions, than to different geo locations. For example: average number of labels per image in train was 4.083 (maximum labels: 127) while in test only 2.386 (maximum labels: 9). So by assuming stage 1 and stage 2 have similar labeling distributions, we can use very similar, but less risky (potentially less overfitting), thresholds for stage 2 as in stage 1.</p>\n\n<p>2) Since we only have one model to choose in stage 2, in order to balance risk v.s. reward, we chose a slightly more conservative approach than our best LB model in stage 1 – our final selected model scored around 10th position on stage 1 LB, which has a slightly more conservative threshold optimization.</p>\n\n<h2>Models and scores</h2>\n\n<p>Our solution consists of 7 models: <em>ResNet50</em>, <em>Xception</em>, <em>Inception ResNet v2</em> x 5. All neural nets had the same final Dense (FC) layer with ~7K neurons for all available classes with sigmoid activation. However, the five inception-resnet have slightly different training and image augmentation approaches. All nets had almost the same score ~0.5 at stage 1 after threshold tuning. The best one was Inception resnet v2. ResNet50 was trained on 336x336 resolution, others on 299 resolution. I believe single model is enough here to get almost the same score (late submission for random single model gives: 0.337). For example ensemble with majority voting on stage 1: 0.505 + 0.501 + 0.511 LB: 0.516</p>\n\n<h2>Train/validation split</h2>\n\n<ul>\n<li>We used all labels for training: human and computer generated.</li>\n<li>We also used probabilities from computer generated labels rather than fixed 1.0 value.</li>\n<li>We didn't use boxes</li>\n<li>Split was non-uniform stratified by classes (see code below):</li>\n</ul>\n\n<pre>    if count &lt; 50:\n        part = count // 10\n    elif count &lt; 100:\n        part = count // 20\n    elif count &lt; 1000:\n        part = count // 50\n    elif count &lt; 10000:\n        part = count // 100\n    elif count &lt; 100000:\n        part = count // 500\n    else:\n        part = count // 1000\n</pre>\n\n<ul>\n<li>Train: 1728299 images Valid: 14743 images</li>\n</ul>\n\n<h2>Training process</h2>\n\n<ul>\n<li>Nothing really special: random samples from training images</li>\n<li>Training from scratch (not using ImageNet pretrained weights) due to competition rules.</li>\n<li>We used default binary_crossentropy loss for this competition</li>\n<li>For validation during training we used tuning labels from stage_1 and calculate F2 score on each epoch with different thresholds (THRs). At the final epoch for inception resnet v2:</li>\n</ul>\n\n<pre>    loss: 0.0019 - f2beta_loss: -5.6904e-01 - fbeta: 0.7356\n    val_loss: 0.0027 - val_f2beta_loss: -5.6610e-01 - val_fbeta: 0.7038 \n    F2Beta score tuning labels: 0.373804 Optimal THR: 0.7\n</pre>\n\n<ul>\n<li>F2Beta score on tuning labels keep improving with improvement of validation, so it mostly used for checking rather than early stopping.</li>\n<li>Local score on tuning labels was almost the same as on leaderboard.</li>\n<li>Heavy augmentations. Very useful library <a href=\"https://github.com/albu/albumentations\">albumentations</a> helped us here:</li>\n</ul>\n\n<pre>    def strong_aug(p=.5):\n        return Compose([\n            HorizontalFlip(),\n            OneOf([\n                IAAAdditiveGaussianNoise(),\n                GaussNoise(),\n            ], p=0.2),\n            OneOf([\n                MotionBlur(p=.2),\n                MedianBlur(blur_limit=3, p=.1),\n                Blur(blur_limit=3, p=.1),\n            ], p=0.2),\n            ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.2, rotate_limit=10, p=0.1),\n            OneOf([\n                OpticalDistortion(p=0.3),\n                GridDistortion(p=0.1),\n                IAAPiecewiseAffine(p=0.3),\n            ], p=0.2),\n            OneOf([\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomContrast(),\n                RandomBrightness(),\n            ], p=0.3),\n            HueSaturationValue(p=0.3),\n            ToGray(p=0.05),\n            JpegCompression(p=0.2, quality_lower=55, quality_upper=99),\n            ElasticTransform(p=0.1),\n        ], p=p)\n</pre>\n\n<ul>\n<li>One epoch 40000 images. To train model we used 200-500 epochs. Training further continue improving model, but with very small effect.</li>\n<li>For xception and 3 more inception-resnet, we use a different image augmentation strategy. For example, one inception-resnet has image augmentation like this:</li>\n</ul>\n\n<pre>gen = ImageDataGenerator(horizontal_flip = True, vertical_flip = True, width_shift_range = 0.1, height_shift_range = 0.1, channel_shift_range=0.1, shear_range = 0.1, zoom_range = 0.1,  rotation_range = 10, preprocessing_function=inception_resnet_preprocess_input)</pre>\n\n<h2>Threshold tuning</h2>\n\n<p>It was the most important part of the competition. We needed to output discrete set of classes for each image based on probabilities from neural networks. We had wide range of strategies to find optimal thresholds and the F2Beta metric is very sensitive. For example:</p>\n\n<ul>\n<li>F2Beta on validation images was about: ~0.73 </li>\n<li>F2Beta on tuning labels with single threshold: ~0.35 </li>\n<li>LB Score F2Beta with optimized thresholds: ~0.5 LB score</li>\n</ul>\n\n<p>Optimization process was made on tuning labels. It had 4 parameters:</p>\n\n<ul>\n<li>minimum probability for search [MinPS]: 0.01</li>\n<li>maximum probability for search [MaxPS]: 0.99</li>\n<li>default probability: 0.99 - We use this probability if class had no entry in tuning labels. Actually, from 7K classes only ~500 had entry in tuning labels. Default prob 0.99 means that model must be &gt;99% sure to use this class.</li>\n<li>minimum number of entries in class: 1</li>\n</ul>\n\n<p>For conservative models in our final submission we used [0.1-0.9] ranges with 2-3 minimum number entries in class and lower default probability (0.8, 0.9). While on Stage 1 LB it gives worse (less overfitting) result.\nStart from setting all thresholds to 0.5 then iterate all over the classes and find threshold for each class in range [MinPS; MaxPS] which maximize the F2Beta score. Do 2-3 overall iterations until F2Beta score stops increasing.</p>\n\n<h2>Ensembles</h2>\n\n<p>We used simple majority voting for ensemble. If at least 4 (out of 7) models predicted a class then this class goes to the submission file.</p>\n\n<h2>Post processing</h2>\n\n<p>For F2Beta metric it's better to output something if model didn't predict any class. So rows which have empty labels, we will use UNION across all 7 models to regenerate labels. If all 7 models have no predictions, we then will output 3 most common classes found in training set.</p>\n\n<h2>Failed experiments</h2>\n\n<ul>\n<li>We also tried ResNet152 on 448x448 images, but it gave worse results.</li>\n<li>MobileNet on 128x128 images - just to check if it could give comparable result, but it didn't</li>\n<li>SE-ResNext101 for Keras from this repo: <a href=\"https://github.com/titu1994/keras-squeeze-excite-network\">https://github.com/titu1994/keras-squeeze-excite-network</a> - very very slow and almost not converges comparing to other models.</li>\n<li>Label selection/filtering based on “class activation maps” (size of activation field, location, etc.)</li>\n</ul>",
      "rawMarkdown": "## Some thoughts ##\n1) After some analysis of data we noticed that most of differences between train and test (stage 1) was more due to different labelling distributions, than to different geo locations. For example: average number of labels per image in train was 4.083 (maximum labels: 127) while in test only 2.386 (maximum labels: 9). So by assuming stage 1 and stage 2 have similar labeling distributions, we can use very similar, but less risky (potentially less overfitting), thresholds for stage 2 as in stage 1.\n\n2) Since we only have one model to choose in stage 2, in order to balance risk v.s. reward, we chose a slightly more conservative approach than our best LB model in stage 1 – our final selected model scored around 10th position on stage 1 LB, which has a slightly more conservative threshold optimization.\n\n## Models and scores ##\nOur solution consists of 7 models: *ResNet50*, *Xception*, *Inception ResNet v2* x 5. All neural nets had the same final Dense (FC) layer with ~7K neurons for all available classes with sigmoid activation. However, the five inception-resnet have slightly different training and image augmentation approaches. All nets had almost the same score ~0.5 at stage 1 after threshold tuning. The best one was Inception resnet v2. ResNet50 was trained on 336x336 resolution, others on 299 resolution. I believe single model is enough here to get almost the same score (late submission for random single model gives: 0.337). For example ensemble with majority voting on stage 1: 0.505 + 0.501 + 0.511 LB: 0.516\n\n## Train/validation split ##\n\n- We used all labels for training: human and computer generated.\n- We also used probabilities from computer generated labels rather than fixed 1.0 value.\n- We didn't use boxes\n- Split was non-uniform stratified by classes (see code below):\n\n<pre>    if count &lt; 50:\n        part = count // 10\n    elif count &lt; 100:\n        part = count // 20\n    elif count &lt; 1000:\n        part = count // 50\n    elif count &lt; 10000:\n        part = count // 100\n    elif count &lt; 100000:\n        part = count // 500\n    else:\n        part = count // 1000\n</pre>\n\n - Train: 1728299 images Valid: 14743 images\n\n## Training process ##\n\n- Nothing really special: random samples from training images\n- Training from scratch (not using ImageNet pretrained weights) due to competition rules.\n- We used default binary_crossentropy loss for this competition\n- For validation during training we used tuning labels from stage_1 and calculate F2 score on each epoch with different thresholds (THRs). At the final epoch for inception resnet v2:\n<pre>    loss: 0.0019 - f2beta_loss: -5.6904e-01 - fbeta: 0.7356\n    val_loss: 0.0027 - val_f2beta_loss: -5.6610e-01 - val_fbeta: 0.7038 \n    F2Beta score tuning labels: 0.373804 Optimal THR: 0.7\n</pre>\n- F2Beta score on tuning labels keep improving with improvement of validation, so it mostly used for checking rather than early stopping.\n- Local score on tuning labels was almost the same as on leaderboard.\n- Heavy augmentations. Very useful library [albumentations][1] helped us here:\n<pre>    def strong_aug(p=.5):\n        return Compose([\n            HorizontalFlip(),\n            OneOf([\n                IAAAdditiveGaussianNoise(),\n                GaussNoise(),\n            ], p=0.2),\n            OneOf([\n                MotionBlur(p=.2),\n                MedianBlur(blur_limit=3, p=.1),\n                Blur(blur_limit=3, p=.1),\n            ], p=0.2),\n            ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.2, rotate_limit=10, p=0.1),\n            OneOf([\n                OpticalDistortion(p=0.3),\n                GridDistortion(p=0.1),\n                IAAPiecewiseAffine(p=0.3),\n            ], p=0.2),\n            OneOf([\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomContrast(),\n                RandomBrightness(),\n            ], p=0.3),\n            HueSaturationValue(p=0.3),\n            ToGray(p=0.05),\n            JpegCompression(p=0.2, quality_lower=55, quality_upper=99),\n            ElasticTransform(p=0.1),\n        ], p=p)\n</pre>\n- One epoch 40000 images. To train model we used 200-500 epochs. Training further continue improving model, but with very small effect.\n- For xception and 3 more inception-resnet, we use a different image augmentation strategy. For example, one inception-resnet has image augmentation like this:\n\n<pre>gen = ImageDataGenerator(horizontal_flip = True, vertical_flip = True, width_shift_range = 0.1, height_shift_range = 0.1, channel_shift_range=0.1, shear_range = 0.1, zoom_range = 0.1,  rotation_range = 10, preprocessing_function=inception_resnet_preprocess_input)</pre>\n\n## Threshold tuning ##\nIt was the most important part of the competition. We needed to output discrete set of classes for each image based on probabilities from neural networks. We had wide range of strategies to find optimal thresholds and the F2Beta metric is very sensitive. For example:\n\n - F2Beta on validation images was about: ~0.73 \n - F2Beta on tuning labels with single threshold: ~0.35 \n - LB Score F2Beta with optimized thresholds: ~0.5 LB score\n\nOptimization process was made on tuning labels. It had 4 parameters:\n\n - minimum probability for search [MinPS]: 0.01\n - maximum probability for search [MaxPS]: 0.99\n - default probability: 0.99 - We use this probability if class had no entry in tuning labels. Actually, from 7K classes only ~500 had entry in tuning labels. Default prob 0.99 means that model must be &gt;99% sure to use this class.\n - minimum number of entries in class: 1\n\nFor conservative models in our final submission we used [0.1-0.9] ranges with 2-3 minimum number entries in class and lower default probability (0.8, 0.9). While on Stage 1 LB it gives worse (less overfitting) result.\nStart from setting all thresholds to 0.5 then iterate all over the classes and find threshold for each class in range [MinPS; MaxPS] which maximize the F2Beta score. Do 2-3 overall iterations until F2Beta score stops increasing.\n\n## Ensembles ##\nWe used simple majority voting for ensemble. If at least 4 (out of 7) models predicted a class then this class goes to the submission file.\n\n## Post processing ##\nFor F2Beta metric it's better to output something if model didn't predict any class. So rows which have empty labels, we will use UNION across all 7 models to regenerate labels. If all 7 models have no predictions, we then will output 3 most common classes found in training set.\n\n## Failed experiments ##\n\n - We also tried ResNet152 on 448x448 images, but it gave worse results.\n - MobileNet on 128x128 images - just to check if it could give comparable result, but it didn't\n - SE-ResNext101 for Keras from this repo: https://github.com/titu1994/keras-squeeze-excite-network - very very slow and almost not converges comparing to other models.\n - Label selection/filtering based on “class activation maps” (size of activation field, location, etc.)\n\n\n  [1]: https://github.com/albu/albumentations",
      "votes": null
    },
    {
      "id": "420485",
      "postDate": "11/13/2018 17:46:04",
      "content": "<p>I like the fact that you points out the real challenge is due to labeling distribution.  So thresholding using the tuning set gives at least .15+ boost in both stages?</p>",
      "rawMarkdown": "I like the fact that you points out the real challenge is due to labeling distribution.  So thresholding using the tuning set gives at least .15+ boost in both stages?",
      "votes": null
    },
    {
      "id": "420495",
      "postDate": "11/13/2018 18:02:49",
      "content": "<blockquote>\n  <p>So thresholding using the tuning set gives at least .15+ boost in both stages?\n  Yes, it was critical to get high score.</p>\n</blockquote>",
      "rawMarkdown": "&gt; So thresholding using the tuning set gives at least .15+ boost in both stages?\nYes, it was critical to get high score.",
      "votes": null
    },
    {
      "id": "420515",
      "postDate": "11/13/2018 18:51:52",
      "content": "<p>Thanks for sharing. May you provide stage1 and stage2 LB score using only single threshold? I trained a inception_v3 based model and achieved 0.3 in stage 1 and 0.18 in stage 2. I just want to check my training process.</p>",
      "rawMarkdown": "Thanks for sharing. May you provide stage1 and stage2 LB score using only single threshold? I trained a inception_v3 based model and achieved 0.3 in stage 1 and 0.18 in stage 2. I just want to check my training process.",
      "votes": null
    },
    {
      "id": "420523",
      "postDate": "11/13/2018 19:06:40",
      "content": "<p>Your result is aligned with what I got. I got .31 from stage1 and .17 on stage2 with ResNet and a single .5 threshold.</p>",
      "rawMarkdown": "Your result is aligned with what I got. I got .31 from stage1 and .17 on stage2 with ResNet and a single .5 threshold.",
      "votes": null
    },
    {
      "id": "420530",
      "postDate": "11/13/2018 19:24:37",
      "content": "<p>Thanks for your info</p>",
      "rawMarkdown": "Thanks for your info",
      "votes": null
    },
    {
      "id": "420549",
      "postDate": "11/13/2018 19:50:58",
      "content": "<p>I have the following for single THR: 0.5</p>\n\n<p>ResNet50: 0.27622</p>\n\n<p>InceptionResnet_v2: 0.30165</p>",
      "rawMarkdown": "I have the following for single THR: 0.5\n\nResNet50: 0.27622\n\nInceptionResnet_v2: 0.30165",
      "votes": null
    },
    {
      "id": "420550",
      "postDate": "11/13/2018 19:59:04",
      "content": "<p>Thanks a lot, that case, my model isn't comparable  to begin with then. </p>",
      "rawMarkdown": "Thanks a lot, that case, my model isn't comparable  to begin with then.",
      "votes": null
    },
    {
      "id": "420554",
      "postDate": "11/13/2018 20:02:39",
      "content": "<p>I have one idea why we could have so different results. We used very heavy augmentations which distort and \"broke\" fotos very much. May be in GEO locations from second stage there are more images with poor quality, and in case of such augmentations poor quality won't reduce performance so much.</p>",
      "rawMarkdown": "I have one idea why we could have so different results. We used very heavy augmentations which distort and \"broke\" fotos very much. May be in GEO locations from second stage there are more images with poor quality, and in case of such augmentations poor quality won't reduce performance so much.",
      "votes": null
    },
    {
      "id": "420555",
      "postDate": "11/13/2018 20:08:21",
      "content": "<p>My augumentation is comparable to yours in that ImageDataGenerator. I would try out the albumentations you mentioned in future.</p>",
      "rawMarkdown": "My augumentation is comparable to yours in that ImageDataGenerator. I would try out the albumentations you mentioned in future.",
      "votes": null
    },
    {
      "id": "420573",
      "postDate": "11/13/2018 20:46:15",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing.",
      "votes": null
    },
    {
      "id": "420577",
      "postDate": "11/13/2018 20:55:26",
      "content": "<p>@ZFTurbo Your 0.5 threshold results are for stage 2?</p>",
      "rawMarkdown": "ZFTurbo Your 0.5 threshold results are for stage 2?",
      "votes": null
    },
    {
      "id": "420585",
      "postDate": "11/13/2018 21:12:13",
      "content": "<p>Yes. I checked through \"late submission\".</p>",
      "rawMarkdown": "Yes. I checked through \"late submission\".",
      "votes": null
    },
    {
      "id": "420612",
      "postDate": "11/13/2018 22:21:21",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "420887",
      "postDate": "11/14/2018 09:29:56",
      "content": "<p>So nice to have this solution and worth it!</p>",
      "rawMarkdown": "So nice to have this solution and worth it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 420485,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "11/13/2018 17:46:04",
      "content": "<p>I like the fact that you points out the real challenge is due to labeling distribution.  So thresholding using the tuning set gives at least .15+ boost in both stages?</p>",
      "votes": null,
      "replies": [
        {
          "id": 420495,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "11/13/2018 18:02:49",
          "content": "<blockquote>\n  <p>So thresholding using the tuning set gives at least .15+ boost in both stages?\n  Yes, it was critical to get high score.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 420515,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "11/13/2018 18:51:52",
      "content": "<p>Thanks for sharing. May you provide stage1 and stage2 LB score using only single threshold? I trained a inception_v3 based model and achieved 0.3 in stage 1 and 0.18 in stage 2. I just want to check my training process.</p>",
      "votes": null,
      "replies": [
        {
          "id": 420523,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "11/13/2018 19:06:40",
          "content": "<p>Your result is aligned with what I got. I got .31 from stage1 and .17 on stage2 with ResNet and a single .5 threshold.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420530,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "11/13/2018 19:24:37",
          "content": "<p>Thanks for your info</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420549,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "11/13/2018 19:50:58",
          "content": "<p>I have the following for single THR: 0.5</p>\n\n<p>ResNet50: 0.27622</p>\n\n<p>InceptionResnet_v2: 0.30165</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420550,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "11/13/2018 19:59:04",
          "content": "<p>Thanks a lot, that case, my model isn't comparable  to begin with then. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420554,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "11/13/2018 20:02:39",
          "content": "<p>I have one idea why we could have so different results. We used very heavy augmentations which distort and \"broke\" fotos very much. May be in GEO locations from second stage there are more images with poor quality, and in case of such augmentations poor quality won't reduce performance so much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420555,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "11/13/2018 20:08:21",
          "content": "<p>My augumentation is comparable to yours in that ImageDataGenerator. I would try out the albumentations you mentioned in future.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420577,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "11/13/2018 20:55:26",
          "content": "<p>@ZFTurbo Your 0.5 threshold results are for stage 2?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420585,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "11/13/2018 21:12:13",
          "content": "<p>Yes. I checked through \"late submission\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 420612,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "11/13/2018 22:21:21",
          "content": "<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 420573,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/13/2018 20:46:15",
      "content": "<p>Congratulations and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420887,
      "author_name": "arunkumarramanan",
      "author_url": "",
      "post_date": "11/14/2018 09:29:56",
      "content": "<p>So nice to have this solution and worth it!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "420465": "## Some thoughts ##\n1) After some analysis of data we noticed that most of differences between train and test (stage 1) was more due to different labelling distributions, than to different geo locations. For example: average number of labels per image in train was 4.083 (maximum labels: 127) while in test only 2.386 (maximum labels: 9). So by assuming stage 1 and stage 2 have similar labeling distributions, we can use very similar, but less risky (potentially less overfitting), thresholds for stage 2 as in stage 1.\n\n2) Since we only have one model to choose in stage 2, in order to balance risk v.s. reward, we chose a slightly more conservative approach than our best LB model in stage 1 – our final selected model scored around 10th position on stage 1 LB, which has a slightly more conservative threshold optimization.\n\n## Models and scores ##\nOur solution consists of 7 models: *ResNet50*, *Xception*, *Inception ResNet v2* x 5. All neural nets had the same final Dense (FC) layer with ~7K neurons for all available classes with sigmoid activation. However, the five inception-resnet have slightly different training and image augmentation approaches. All nets had almost the same score ~0.5 at stage 1 after threshold tuning. The best one was Inception resnet v2. ResNet50 was trained on 336x336 resolution, others on 299 resolution. I believe single model is enough here to get almost the same score (late submission for random single model gives: 0.337). For example ensemble with majority voting on stage 1: 0.505 + 0.501 + 0.511 LB: 0.516\n\n## Train/validation split ##\n\n- We used all labels for training: human and computer generated.\n- We also used probabilities from computer generated labels rather than fixed 1.0 value.\n- We didn't use boxes\n- Split was non-uniform stratified by classes (see code below):\n\n<pre>    if count &lt; 50:\n        part = count // 10\n    elif count &lt; 100:\n        part = count // 20\n    elif count &lt; 1000:\n        part = count // 50\n    elif count &lt; 10000:\n        part = count // 100\n    elif count &lt; 100000:\n        part = count // 500\n    else:\n        part = count // 1000\n</pre>\n\n - Train: 1728299 images Valid: 14743 images\n\n## Training process ##\n\n- Nothing really special: random samples from training images\n- Training from scratch (not using ImageNet pretrained weights) due to competition rules.\n- We used default binary_crossentropy loss for this competition\n- For validation during training we used tuning labels from stage_1 and calculate F2 score on each epoch with different thresholds (THRs). At the final epoch for inception resnet v2:\n<pre>    loss: 0.0019 - f2beta_loss: -5.6904e-01 - fbeta: 0.7356\n    val_loss: 0.0027 - val_f2beta_loss: -5.6610e-01 - val_fbeta: 0.7038 \n    F2Beta score tuning labels: 0.373804 Optimal THR: 0.7\n</pre>\n- F2Beta score on tuning labels keep improving with improvement of validation, so it mostly used for checking rather than early stopping.\n- Local score on tuning labels was almost the same as on leaderboard.\n- Heavy augmentations. Very useful library [albumentations][1] helped us here:\n<pre>    def strong_aug(p=.5):\n        return Compose([\n            HorizontalFlip(),\n            OneOf([\n                IAAAdditiveGaussianNoise(),\n                GaussNoise(),\n            ], p=0.2),\n            OneOf([\n                MotionBlur(p=.2),\n                MedianBlur(blur_limit=3, p=.1),\n                Blur(blur_limit=3, p=.1),\n            ], p=0.2),\n            ShiftScaleRotate(shift_limit=0.0625, scale_limit=0.2, rotate_limit=10, p=0.1),\n            OneOf([\n                OpticalDistortion(p=0.3),\n                GridDistortion(p=0.1),\n                IAAPiecewiseAffine(p=0.3),\n            ], p=0.2),\n            OneOf([\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomContrast(),\n                RandomBrightness(),\n            ], p=0.3),\n            HueSaturationValue(p=0.3),\n            ToGray(p=0.05),\n            JpegCompression(p=0.2, quality_lower=55, quality_upper=99),\n            ElasticTransform(p=0.1),\n        ], p=p)\n</pre>\n- One epoch 40000 images. To train model we used 200-500 epochs. Training further continue improving model, but with very small effect.\n- For xception and 3 more inception-resnet, we use a different image augmentation strategy. For example, one inception-resnet has image augmentation like this:\n\n<pre>gen = ImageDataGenerator(horizontal_flip = True, vertical_flip = True, width_shift_range = 0.1, height_shift_range = 0.1, channel_shift_range=0.1, shear_range = 0.1, zoom_range = 0.1,  rotation_range = 10, preprocessing_function=inception_resnet_preprocess_input)</pre>\n\n## Threshold tuning ##\nIt was the most important part of the competition. We needed to output discrete set of classes for each image based on probabilities from neural networks. We had wide range of strategies to find optimal thresholds and the F2Beta metric is very sensitive. For example:\n\n - F2Beta on validation images was about: ~0.73 \n - F2Beta on tuning labels with single threshold: ~0.35 \n - LB Score F2Beta with optimized thresholds: ~0.5 LB score\n\nOptimization process was made on tuning labels. It had 4 parameters:\n\n - minimum probability for search [MinPS]: 0.01\n - maximum probability for search [MaxPS]: 0.99\n - default probability: 0.99 - We use this probability if class had no entry in tuning labels. Actually, from 7K classes only ~500 had entry in tuning labels. Default prob 0.99 means that model must be &gt;99% sure to use this class.\n - minimum number of entries in class: 1\n\nFor conservative models in our final submission we used [0.1-0.9] ranges with 2-3 minimum number entries in class and lower default probability (0.8, 0.9). While on Stage 1 LB it gives worse (less overfitting) result.\nStart from setting all thresholds to 0.5 then iterate all over the classes and find threshold for each class in range [MinPS; MaxPS] which maximize the F2Beta score. Do 2-3 overall iterations until F2Beta score stops increasing.\n\n## Ensembles ##\nWe used simple majority voting for ensemble. If at least 4 (out of 7) models predicted a class then this class goes to the submission file.\n\n## Post processing ##\nFor F2Beta metric it's better to output something if model didn't predict any class. So rows which have empty labels, we will use UNION across all 7 models to regenerate labels. If all 7 models have no predictions, we then will output 3 most common classes found in training set.\n\n## Failed experiments ##\n\n - We also tried ResNet152 on 448x448 images, but it gave worse results.\n - MobileNet on 128x128 images - just to check if it could give comparable result, but it didn't\n - SE-ResNext101 for Keras from this repo: https://github.com/titu1994/keras-squeeze-excite-network - very very slow and almost not converges comparing to other models.\n - Label selection/filtering based on “class activation maps” (size of activation field, location, etc.)\n\n\n  [1]: https://github.com/albu/albumentations",
    "420485": "I like the fact that you points out the real challenge is due to labeling distribution.  So thresholding using the tuning set gives at least .15+ boost in both stages?",
    "420495": "&gt; So thresholding using the tuning set gives at least .15+ boost in both stages?\nYes, it was critical to get high score.",
    "420515": "Thanks for sharing. May you provide stage1 and stage2 LB score using only single threshold? I trained a inception_v3 based model and achieved 0.3 in stage 1 and 0.18 in stage 2. I just want to check my training process.",
    "420523": "Your result is aligned with what I got. I got .31 from stage1 and .17 on stage2 with ResNet and a single .5 threshold.",
    "420530": "Thanks for your info",
    "420549": "I have the following for single THR: 0.5\n\nResNet50: 0.27622\n\nInceptionResnet_v2: 0.30165",
    "420550": "Thanks a lot, that case, my model isn't comparable  to begin with then.",
    "420554": "I have one idea why we could have so different results. We used very heavy augmentations which distort and \"broke\" fotos very much. May be in GEO locations from second stage there are more images with poor quality, and in case of such augmentations poor quality won't reduce performance so much.",
    "420555": "My augumentation is comparable to yours in that ImageDataGenerator. I would try out the albumentations you mentioned in future.",
    "420573": "Congratulations and thanks for sharing.",
    "420577": "ZFTurbo Your 0.5 threshold results are for stage 2?",
    "420585": "Yes. I checked through \"late submission\".",
    "420612": "Thanks",
    "420887": "So nice to have this solution and worth it!"
  },
  "source": "meta"
}