{
  "id": 108007,
  "title": "58 place solution",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108007",
  "author_name": "Borys Tymchenko",
  "post_date": "2019-09-08T12:33:27.288000",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everybody and congratulations to all of the participants!</p>\n\n<p><strong>Data</strong>\nWe used resized images from 2015 competition to pretrain our models.\nAlso, we added IDRID and MESSIDOR data to cross-validation while finetuning.</p>\n\n<p><strong>Models</strong>\nWe tried a lot of models from the start of competition, but got no luck with either small (ResNet34/DenseNet121) nor very large models (EfficientNet-B6/B7, ResNe(x)t101, DenseNet201)</p>\n\n<p>Best scoring modes were of medium size and medium resolution:\nEfficientNet-B4: 380x380\nEfficientNet-B5: 456x456\nSE-ResNeXt50: 512x512\nSE-ResNeXt50: 380x380\nWe got no improvement after 512x512 images.</p>\n\n<p><strong>Preprocessing</strong>\nAt first, we used circle cropping and Ben's preprocessing fro a while, but then started to experiment with tight cropping and corners masking, which was worse on every model we tried (both CV and LB).\nA week before the competition deadline we realized, that we lose too much information and switched to no preprocessing and hard augmentations instead.</p>\n\n<p><strong>Augmentations</strong>\nWe used a lot of augmentations, all from Albumentations library:</p>\n\n<p>OpticalDistortion, GridDistortion, PiecewiseAffine, CLAHE, HorizontalFlip, VerticalFlip, RandomRotate90, ShiftScaleRotate, RGBShift, RandomBrightnessContrast, AdditiveGaussianNoise, GaussNoise, MotionBlur, MedianBlur, Blur, Sharpen, Emboss, RandomGamma, ToGray, CoarseDropout, Cutout.</p>\n\n<p><strong>Training</strong>\nWe used multi-target learning (classification, ordinal regression and regression), and their result was transformed to regression combined with a linear layer (initialized with 0.3333 weights) to get regression output.</p>\n\n<p>At first, we trained only on the current data, but then switched to pretraining on 2015 dataset.</p>\n\n<p>The flow is following:\nPretrain on 2015 dataset (20 epochs, SGD+Nesterov+CosineLR, Label smoothing)\nTrain 5-fold CV on 2019 data, IDRID and MESSIDOR combined (75 epochs, RAdam, CosineLR, Label smoothing)</p>\n\n<p>Every training stage was made of three substages:\n- Freeze model body and head combiner, train for 5 epochs.\n- Unfreeze body, train N epochs.\n- Freeze everything, unfreeze head combiner, train for 5 epochs.\nIf we trained everything together, we've  always got uneven training of heads.</p>\n\n<p>Also, we added random noise in range (-0.3, 0.3) to labels for regression, that greatly reduced overfitting.</p>\n\n<p><strong>Hardware</strong>\nWe used server from FastGPU.net with 4xV100, that greatly reduced our experiment cycle length. </p>\n\n<p><strong>Ensembling</strong>\nWe picked model by their performance on CV and holdout and did not orient on LB at all.\nOur best performing solution was an ensemble of 20 models (4 models x 5 folds) with 10xTTA (hflip, vflip, transpose, rotate, zoom) combined with 0.25-trimmed mean. This ensemble performed worse both on holdout and LB, but we decided to choose with hope that it will generalize better (what it did!). </p>",
  "messages": [
    {
      "id": 621315,
      "postDate": "2019-09-08T12:33:27.290Z",
      "content": "<p>Hi everybody and congratulations to all of the participants!</p>\n\n<p><strong>Data</strong>\nWe used resized images from 2015 competition to pretrain our models.\nAlso, we added IDRID and MESSIDOR data to cross-validation while finetuning.</p>\n\n<p><strong>Models</strong>\nWe tried a lot of models from the start of competition, but got no luck with either small (ResNet34/DenseNet121) nor very large models (EfficientNet-B6/B7, ResNe(x)t101, DenseNet201)</p>\n\n<p>Best scoring modes were of medium size and medium resolution:\nEfficientNet-B4: 380x380\nEfficientNet-B5: 456x456\nSE-ResNeXt50: 512x512\nSE-ResNeXt50: 380x380\nWe got no improvement after 512x512 images.</p>\n\n<p><strong>Preprocessing</strong>\nAt first, we used circle cropping and Ben's preprocessing fro a while, but then started to experiment with tight cropping and corners masking, which was worse on every model we tried (both CV and LB).\nA week before the competition deadline we realized, that we lose too much information and switched to no preprocessing and hard augmentations instead.</p>\n\n<p><strong>Augmentations</strong>\nWe used a lot of augmentations, all from Albumentations library:</p>\n\n<p>OpticalDistortion, GridDistortion, PiecewiseAffine, CLAHE, HorizontalFlip, VerticalFlip, RandomRotate90, ShiftScaleRotate, RGBShift, RandomBrightnessContrast, AdditiveGaussianNoise, GaussNoise, MotionBlur, MedianBlur, Blur, Sharpen, Emboss, RandomGamma, ToGray, CoarseDropout, Cutout.</p>\n\n<p><strong>Training</strong>\nWe used multi-target learning (classification, ordinal regression and regression), and their result was transformed to regression combined with a linear layer (initialized with 0.3333 weights) to get regression output.</p>\n\n<p>At first, we trained only on the current data, but then switched to pretraining on 2015 dataset.</p>\n\n<p>The flow is following:\nPretrain on 2015 dataset (20 epochs, SGD+Nesterov+CosineLR, Label smoothing)\nTrain 5-fold CV on 2019 data, IDRID and MESSIDOR combined (75 epochs, RAdam, CosineLR, Label smoothing)</p>\n\n<p>Every training stage was made of three substages:\n- Freeze model body and head combiner, train for 5 epochs.\n- Unfreeze body, train N epochs.\n- Freeze everything, unfreeze head combiner, train for 5 epochs.\nIf we trained everything together, we've  always got uneven training of heads.</p>\n\n<p>Also, we added random noise in range (-0.3, 0.3) to labels for regression, that greatly reduced overfitting.</p>\n\n<p><strong>Hardware</strong>\nWe used server from FastGPU.net with 4xV100, that greatly reduced our experiment cycle length. </p>\n\n<p><strong>Ensembling</strong>\nWe picked model by their performance on CV and holdout and did not orient on LB at all.\nOur best performing solution was an ensemble of 20 models (4 models x 5 folds) with 10xTTA (hflip, vflip, transpose, rotate, zoom) combined with 0.25-trimmed mean. This ensemble performed worse both on holdout and LB, but we decided to choose with hope that it will generalize better (what it did!). </p>",
      "rawMarkdown": "Hi everybody and congratulations to all of the participants!\n\n**Data**\nWe used resized images from 2015 competition to pretrain our models.\nAlso, we added IDRID and MESSIDOR data to cross-validation while finetuning.\n\n**Models**\nWe tried a lot of models from the start of competition, but got no luck with either small (ResNet34/DenseNet121) nor very large models (EfficientNet-B6/B7, ResNe(x)t101, DenseNet201)\n\nBest scoring modes were of medium size and medium resolution:\nEfficientNet-B4: 380x380\nEfficientNet-B5: 456x456\nSE-ResNeXt50: 512x512\nSE-ResNeXt50: 380x380\nWe got no improvement after 512x512 images.\n\n**Preprocessing**\nAt first, we used circle cropping and Ben's preprocessing fro a while, but then started to experiment with tight cropping and corners masking, which was worse on every model we tried (both CV and LB).\nA week before the competition deadline we realized, that we lose too much information and switched to no preprocessing and hard augmentations instead.\n\n**Augmentations**\nWe used a lot of augmentations, all from Albumentations library:\n\nOpticalDistortion, GridDistortion, PiecewiseAffine, CLAHE, HorizontalFlip, VerticalFlip, RandomRotate90, ShiftScaleRotate, RGBShift, RandomBrightnessContrast, AdditiveGaussianNoise, GaussNoise, MotionBlur, MedianBlur, Blur, Sharpen, Emboss, RandomGamma, ToGray, CoarseDropout, Cutout.\n\n**Training**\nWe used multi-target learning (classification, ordinal regression and regression), and their result was transformed to regression combined with a linear layer (initialized with 0.3333 weights) to get regression output.\n\nAt first, we trained only on the current data, but then switched to pretraining on 2015 dataset.\n\nThe flow is following:\nPretrain on 2015 dataset (20 epochs, SGD+Nesterov+CosineLR, Label smoothing)\nTrain 5-fold CV on 2019 data, IDRID and MESSIDOR combined (75 epochs, RAdam, CosineLR, Label smoothing)\n\nEvery training stage was made of three substages:\n- Freeze model body and head combiner, train for 5 epochs.\n- Unfreeze body, train N epochs.\n- Freeze everything, unfreeze head combiner, train for 5 epochs.\nIf we trained everything together, we've  always got uneven training of heads.\n\nAlso, we added random noise in range (-0.3, 0.3) to labels for regression, that greatly reduced overfitting.\n\n**Hardware**\nWe used server from FastGPU.net with 4xV100, that greatly reduced our experiment cycle length. \n\n**Ensembling**\nWe picked model by their performance on CV and holdout and did not orient on LB at all.\nOur best performing solution was an ensemble of 20 models (4 models x 5 folds) with 10xTTA (hflip, vflip, transpose, rotate, zoom) combined with 0.25-trimmed mean. This ensemble performed worse both on holdout and LB, but we decided to choose with hope that it will generalize better (what it did!). ",
      "votes": 8
    },
    {
      "id": 621374,
      "postDate": "2019-09-08T13:21:09.107Z",
      "content": "<p>How are you applying to augmentations to training data?  </p>",
      "rawMarkdown": "How are you applying to augmentations to training data?  ",
      "replies": [
        {
          "id": 621400,
          "postDate": "2019-09-08T13:35:28.620Z",
          "content": "<p>We just did it online in the data generator</p>",
          "rawMarkdown": "We just did it online in the data generator",
          "votes": 1
        },
        {
          "id": 621412,
          "postDate": "2019-09-08T13:44:44.020Z",
          "content": "<p>Thanks for replying. I am kind of new to augmentations. \nSo you have applied all(that you have mentioned in the answer) the transformations(with some probability) to each image in the training.</p>\n\n<p>I have applied to very less transformations (to training data) to the data like this \n<code>\nCompose([\n                          Flip(p=0.5),\n                          Rotate(360,p=0.5),\n                          RandomBrightnessContrast(p=0.5),\n])\n</code>\nSo in your case you have done like this (to training data)\n<code>\nCompose([\n                   OpticalDistortion, \n                   GridDistortion, \n                   PiecewiseAffine,\n                   CLAHE, \n                   HorizontalFlip, \n                   VerticalFlip,\n                   .....\n])\n</code></p>\n\n<p>Am I right or I am missing something?</p>",
          "rawMarkdown": "Thanks for replying. I am kind of new to augmentations. \nSo you have applied all(that you have mentioned in the answer) the transformations(with some probability) to each image in the training.\n\nI have applied to very less transformations (to training data) to the data like this \n```\nCompose([\n                          Flip(p=0.5),\n                          Rotate(360,p=0.5),\n                          RandomBrightnessContrast(p=0.5),\n])\n```\nSo in your case you have done like this (to training data)\n```\nCompose([\n                   OpticalDistortion, \n                   GridDistortion, \n                   PiecewiseAffine,\n                   CLAHE, \n                   HorizontalFlip, \n                   VerticalFlip,\n                   .....\n])\n```\n\nAm I right or I am missing something?"
        },
        {
          "id": 621671,
          "postDate": "2019-09-08T20:18:10.710Z",
          "content": "<p>Yes, kinda like this but with different probabilities and grouped by type with OneOf</p>",
          "rawMarkdown": "Yes, kinda like this but with different probabilities and grouped by type with OneOf"
        }
      ]
    },
    {
      "id": 621925,
      "postDate": "2019-09-09T05:13:17.387Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 621374,
      "author_name": "redshellspy",
      "author_url": "",
      "post_date": "2019-09-08T13:21:09.107000",
      "content": "<p>How are you applying to augmentations to training data?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 621400,
          "author_name": "Borys Tymchenko",
          "author_url": "",
          "post_date": "2019-09-08T13:35:28.620000",
          "content": "<p>We just did it online in the data generator</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 621412,
          "author_name": "redshellspy",
          "author_url": "",
          "post_date": "2019-09-08T13:44:44.020000",
          "content": "<p>Thanks for replying. I am kind of new to augmentations. \nSo you have applied all(that you have mentioned in the answer) the transformations(with some probability) to each image in the training.</p>\n\n<p>I have applied to very less transformations (to training data) to the data like this \n<code>\nCompose([\n                          Flip(p=0.5),\n                          Rotate(360,p=0.5),\n                          RandomBrightnessContrast(p=0.5),\n])\n</code>\nSo in your case you have done like this (to training data)\n<code>\nCompose([\n                   OpticalDistortion, \n                   GridDistortion, \n                   PiecewiseAffine,\n                   CLAHE, \n                   HorizontalFlip, \n                   VerticalFlip,\n                   .....\n])\n</code></p>\n\n<p>Am I right or I am missing something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621671,
          "author_name": "Borys Tymchenko",
          "author_url": "",
          "post_date": "2019-09-08T20:18:10.710000",
          "content": "<p>Yes, kinda like this but with different probabilities and grouped by type with OneOf</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621925,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T05:13:17.387000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621315": "Hi everybody and congratulations to all of the participants!\n\n**Data**\nWe used resized images from 2015 competition to pretrain our models.\nAlso, we added IDRID and MESSIDOR data to cross-validation while finetuning.\n\n**Models**\nWe tried a lot of models from the start of competition, but got no luck with either small (ResNet34/DenseNet121) nor very large models (EfficientNet-B6/B7, ResNe(x)t101, DenseNet201)\n\nBest scoring modes were of medium size and medium resolution:\nEfficientNet-B4: 380x380\nEfficientNet-B5: 456x456\nSE-ResNeXt50: 512x512\nSE-ResNeXt50: 380x380\nWe got no improvement after 512x512 images.\n\n**Preprocessing**\nAt first, we used circle cropping and Ben's preprocessing fro a while, but then started to experiment with tight cropping and corners masking, which was worse on every model we tried (both CV and LB).\nA week before the competition deadline we realized, that we lose too much information and switched to no preprocessing and hard augmentations instead.\n\n**Augmentations**\nWe used a lot of augmentations, all from Albumentations library:\n\nOpticalDistortion, GridDistortion, PiecewiseAffine, CLAHE, HorizontalFlip, VerticalFlip, RandomRotate90, ShiftScaleRotate, RGBShift, RandomBrightnessContrast, AdditiveGaussianNoise, GaussNoise, MotionBlur, MedianBlur, Blur, Sharpen, Emboss, RandomGamma, ToGray, CoarseDropout, Cutout.\n\n**Training**\nWe used multi-target learning (classification, ordinal regression and regression), and their result was transformed to regression combined with a linear layer (initialized with 0.3333 weights) to get regression output.\n\nAt first, we trained only on the current data, but then switched to pretraining on 2015 dataset.\n\nThe flow is following:\nPretrain on 2015 dataset (20 epochs, SGD+Nesterov+CosineLR, Label smoothing)\nTrain 5-fold CV on 2019 data, IDRID and MESSIDOR combined (75 epochs, RAdam, CosineLR, Label smoothing)\n\nEvery training stage was made of three substages:\n- Freeze model body and head combiner, train for 5 epochs.\n- Unfreeze body, train N epochs.\n- Freeze everything, unfreeze head combiner, train for 5 epochs.\nIf we trained everything together, we've  always got uneven training of heads.\n\nAlso, we added random noise in range (-0.3, 0.3) to labels for regression, that greatly reduced overfitting.\n\n**Hardware**\nWe used server from FastGPU.net with 4xV100, that greatly reduced our experiment cycle length. \n\n**Ensembling**\nWe picked model by their performance on CV and holdout and did not orient on LB at all.\nOur best performing solution was an ensemble of 20 models (4 models x 5 folds) with 10xTTA (hflip, vflip, transpose, rotate, zoom) combined with 0.25-trimmed mean. This ensemble performed worse both on holdout and LB, but we decided to choose with hope that it will generalize better (what it did!). ",
    "621374": "How are you applying to augmentations to training data?  ",
    "621925": ""
  }
}