{
  "id": 110375,
  "title": "10th place - Pure Magic Solution",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110375",
  "author_name": "Igor Krashenyi",
  "post_date": "2019-09-27T07:35:08.841000",
  "votes": 30,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congrats to all the winners and thanks to the host and Kaggle hosted such an interesting competition. Huge thank to my teammates you are the greatest!</p>\n\n<p><strong>Overview:</strong>\nAll our models were trained with 512x512x6 images. We sampled a random site for each sample per epoch. All our were 2-headed.\nFirst head: classification head had an output with 1139 neurons, \nSecond head: embeddings.</p>\n\n<p><strong>Challenges:</strong>\nValidation \nSpecific experiments like U2OS-4 failed </p>\n\n<p><strong>Validation for CNNs:</strong>\nWe noticed that specific experiment types perform very differently from others. We chose the hardest experiment for our model and built the validation based on it. </p>\n\n<p><strong>Training Augmentations:</strong>\n- flips (horizontal, vertical)\n- rot90 \n- transpose\n- shift (0.25) with 101 reflections \n- rot360 (-180, 180)\n- cutout (32x32, 2 holes)\n- noise (gaussian, localvar, poisson, salt&amp;pepper, speckle)\n- clahe\n- gamma (0.9,1.1)</p>\n\n<p><strong>Data Pre-Processing:</strong>\nper channel (img - img.mean()) / (img.std() + 1e-6)\nWe noticed that due to heavy augmentations we had a true_division warning. This warning caused batchnorms to feel bad, so we did this trick to prevent zero division.</p>\n\n<p><strong>Model training:</strong>\nWe train our models with NVidia Apex and Pytorch 1.2 on 8x1080ti GPU server for 110 epochs. The training takes 1-3 days depending on the model.\nWe did not use oversample.\nBecause of the long iteration, we didn’t use any K-Fold training. We wanted to train it closer to the end of the competition, but we didn’t have much time. </p>\n\n<p><strong>Optimizer:</strong> SGD</p>\n\n<p><strong>Scheduler:</strong> init LR 0.1, \n5 epochs warmup,\ndrop LR every 40 epochs with factor 0.1</p>\n\n<p><strong>Loss Functions:</strong>\nSmooth CrossEntropy for classification, Center Loss for embeddings. </p>\n\n<p><strong>Model:</strong>\nOur final model is Senet154 trained on train data + pseudo-labels.\nTo produce pseudo-labels were using an ensemble SeNet154, SeResnext50, SeResnext101, Polynet, EfficientNet-b6, ResNeXt101-wsl. </p>\n\n<p><strong>Test time augmentations:</strong>\noutput = (model(img1) + model(img2) + model(img1.flip(2)) + model(img1.flip(3)) + model(img2.flip(2)) + model(img2.flip(3)) / 6.</p>\n\n<p><strong>Post-processing:</strong>\nHungry experiment re-calibration and plate re-calibration\n<a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270</a> </p>\n\n<p><strong>What didn’t work or had about the same performance:</strong>\n- Mix-up, manifold mixup\n- ArcFace\n- Last stride 1\n- Focal Loss\n- XGBoost ensembling \n- Metric search based on the embeddings distances\n- GeM\n- Different optimizers like Adam, Radam, Ranger, etc.\n- ResNeXt101-wsl performed worse than expected. </p>",
  "messages": [
    {
      "id": 635137,
      "postDate": "2019-09-27T07:35:08.843Z",
      "content": "<p>Congrats to all the winners and thanks to the host and Kaggle hosted such an interesting competition. Huge thank to my teammates you are the greatest!</p>\n\n<p><strong>Overview:</strong>\nAll our models were trained with 512x512x6 images. We sampled a random site for each sample per epoch. All our were 2-headed.\nFirst head: classification head had an output with 1139 neurons, \nSecond head: embeddings.</p>\n\n<p><strong>Challenges:</strong>\nValidation \nSpecific experiments like U2OS-4 failed </p>\n\n<p><strong>Validation for CNNs:</strong>\nWe noticed that specific experiment types perform very differently from others. We chose the hardest experiment for our model and built the validation based on it. </p>\n\n<p><strong>Training Augmentations:</strong>\n- flips (horizontal, vertical)\n- rot90 \n- transpose\n- shift (0.25) with 101 reflections \n- rot360 (-180, 180)\n- cutout (32x32, 2 holes)\n- noise (gaussian, localvar, poisson, salt&amp;pepper, speckle)\n- clahe\n- gamma (0.9,1.1)</p>\n\n<p><strong>Data Pre-Processing:</strong>\nper channel (img - img.mean()) / (img.std() + 1e-6)\nWe noticed that due to heavy augmentations we had a true_division warning. This warning caused batchnorms to feel bad, so we did this trick to prevent zero division.</p>\n\n<p><strong>Model training:</strong>\nWe train our models with NVidia Apex and Pytorch 1.2 on 8x1080ti GPU server for 110 epochs. The training takes 1-3 days depending on the model.\nWe did not use oversample.\nBecause of the long iteration, we didn’t use any K-Fold training. We wanted to train it closer to the end of the competition, but we didn’t have much time. </p>\n\n<p><strong>Optimizer:</strong> SGD</p>\n\n<p><strong>Scheduler:</strong> init LR 0.1, \n5 epochs warmup,\ndrop LR every 40 epochs with factor 0.1</p>\n\n<p><strong>Loss Functions:</strong>\nSmooth CrossEntropy for classification, Center Loss for embeddings. </p>\n\n<p><strong>Model:</strong>\nOur final model is Senet154 trained on train data + pseudo-labels.\nTo produce pseudo-labels were using an ensemble SeNet154, SeResnext50, SeResnext101, Polynet, EfficientNet-b6, ResNeXt101-wsl. </p>\n\n<p><strong>Test time augmentations:</strong>\noutput = (model(img1) + model(img2) + model(img1.flip(2)) + model(img1.flip(3)) + model(img2.flip(2)) + model(img2.flip(3)) / 6.</p>\n\n<p><strong>Post-processing:</strong>\nHungry experiment re-calibration and plate re-calibration\n<a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270</a> </p>\n\n<p><strong>What didn’t work or had about the same performance:</strong>\n- Mix-up, manifold mixup\n- ArcFace\n- Last stride 1\n- Focal Loss\n- XGBoost ensembling \n- Metric search based on the embeddings distances\n- GeM\n- Different optimizers like Adam, Radam, Ranger, etc.\n- ResNeXt101-wsl performed worse than expected. </p>",
      "rawMarkdown": "Congrats to all the winners and thanks to the host and Kaggle hosted such an interesting competition. Huge thank to my teammates you are the greatest!\n\n**Overview:**\nAll our models were trained with 512x512x6 images. We sampled a random site for each sample per epoch. All our were 2-headed.\nFirst head: classification head had an output with 1139 neurons, \nSecond head: embeddings.\n\n**Challenges:**\nValidation \nSpecific experiments like U2OS-4 failed \n\n**Validation for CNNs:**\nWe noticed that specific experiment types perform very differently from others. We chose the hardest experiment for our model and built the validation based on it. \n\n**Training Augmentations:**\n- flips (horizontal, vertical)\n- rot90 \n- transpose\n- shift (0.25) with 101 reflections \n- rot360 (-180, 180)\n- cutout (32x32, 2 holes)\n- noise (gaussian, localvar, poisson, salt&amp;pepper, speckle)\n- clahe\n- gamma (0.9,1.1)\n\n**Data Pre-Processing:**\nper channel (img - img.mean()) / (img.std() + 1e-6)\nWe noticed that due to heavy augmentations we had a true_division warning. This warning caused batchnorms to feel bad, so we did this trick to prevent zero division.\n\n**Model training:**\nWe train our models with NVidia Apex and Pytorch 1.2 on 8x1080ti GPU server for 110 epochs. The training takes 1-3 days depending on the model.\nWe did not use oversample.\nBecause of the long iteration, we didn’t use any K-Fold training. We wanted to train it closer to the end of the competition, but we didn’t have much time. \n\n**Optimizer:** SGD\n\n**Scheduler:** init LR 0.1, \n5 epochs warmup,\ndrop LR every 40 epochs with factor 0.1\n\n**Loss Functions:**\nSmooth CrossEntropy for classification, Center Loss for embeddings. \n\n**Model:**\nOur final model is Senet154 trained on train data + pseudo-labels.\nTo produce pseudo-labels were using an ensemble SeNet154, SeResnext50, SeResnext101, Polynet, EfficientNet-b6, ResNeXt101-wsl. \n\n**Test time augmentations:**\noutput = (model(img1) + model(img2) + model(img1.flip(2)) + model(img1.flip(3)) + model(img2.flip(2)) + model(img2.flip(3)) / 6.\n\n**Post-processing:**\nHungry experiment re-calibration and plate re-calibration\nhttps://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270 \n\n**What didn’t work or had about the same performance:**\n- Mix-up, manifold mixup\n- ArcFace\n- Last stride 1\n- Focal Loss\n- XGBoost ensembling \n- Metric search based on the embeddings distances\n- GeM\n- Different optimizers like Adam, Radam, Ranger, etc.\n- ResNeXt101-wsl performed worse than expected. \n",
      "votes": 30
    },
    {
      "id": 635197,
      "postDate": "2019-09-27T08:46:49.253Z",
      "content": "<p><code>Smooth CrossEntropy for classification, Center Loss for embeddings.</code>\nSo how do you infer predictions? How do you use those embeddings?</p>\n\n<p>And how do you select pseudo-labels?</p>",
      "rawMarkdown": "`Smooth CrossEntropy for classification, Center Loss for embeddings.`\nSo how do you infer predictions? How do you use those embeddings?\n\nAnd how do you select pseudo-labels?",
      "votes": 1,
      "replies": [
        {
          "id": 635227,
          "postDate": "2019-09-27T09:23:21.570Z",
          "content": "<p>We wanted to find the nearest neighbor based on these embedding (and maybe to combine with classification part), but we didn’t succeed with re-calibration (1 class per experiment) based on the distances. </p>\n\n<p>So center loss was just a regularizer for the model’a embeddings. </p>\n\n<p>We used only logits from model’s linear output. </p>",
          "rawMarkdown": "We wanted to find the nearest neighbor based on these embedding (and maybe to combine with classification part), but we didn’t succeed with re-calibration (1 class per experiment) based on the distances. \n\nSo center loss was just a regularizer for the model’a embeddings. \n\nWe used only logits from model’s linear output. ",
          "votes": 2
        },
        {
          "id": 635301,
          "postDate": "2019-09-27T10:48:13.597Z",
          "content": "<p>Regarding pseudo-labels, we didn't apply any methods to select them. \nWe added all of them after re-calibration to our train set and trained with Smooth-CE loss</p>",
          "rawMarkdown": "Regarding pseudo-labels, we didn't apply any methods to select them. \nWe added all of them after re-calibration to our train set and trained with Smooth-CE loss",
          "votes": 1
        }
      ]
    },
    {
      "id": 635146,
      "postDate": "2019-09-27T07:44:41.583Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 635197,
      "author_name": "Artyom Palvelev",
      "author_url": "",
      "post_date": "2019-09-27T08:46:49.253000",
      "content": "<p><code>Smooth CrossEntropy for classification, Center Loss for embeddings.</code>\nSo how do you infer predictions? How do you use those embeddings?</p>\n\n<p>And how do you select pseudo-labels?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 635227,
          "author_name": "Igor Krashenyi",
          "author_url": "",
          "post_date": "2019-09-27T09:23:21.570000",
          "content": "<p>We wanted to find the nearest neighbor based on these embedding (and maybe to combine with classification part), but we didn’t succeed with re-calibration (1 class per experiment) based on the distances. </p>\n\n<p>So center loss was just a regularizer for the model’a embeddings. </p>\n\n<p>We used only logits from model’s linear output. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 635301,
          "author_name": "Igor Krashenyi",
          "author_url": "",
          "post_date": "2019-09-27T10:48:13.597000",
          "content": "<p>Regarding pseudo-labels, we didn't apply any methods to select them. \nWe added all of them after re-calibration to our train set and trained with Smooth-CE loss</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 635146,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T07:44:41.583000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635137": "Congrats to all the winners and thanks to the host and Kaggle hosted such an interesting competition. Huge thank to my teammates you are the greatest!\n\n**Overview:**\nAll our models were trained with 512x512x6 images. We sampled a random site for each sample per epoch. All our were 2-headed.\nFirst head: classification head had an output with 1139 neurons, \nSecond head: embeddings.\n\n**Challenges:**\nValidation \nSpecific experiments like U2OS-4 failed \n\n**Validation for CNNs:**\nWe noticed that specific experiment types perform very differently from others. We chose the hardest experiment for our model and built the validation based on it. \n\n**Training Augmentations:**\n- flips (horizontal, vertical)\n- rot90 \n- transpose\n- shift (0.25) with 101 reflections \n- rot360 (-180, 180)\n- cutout (32x32, 2 holes)\n- noise (gaussian, localvar, poisson, salt&amp;pepper, speckle)\n- clahe\n- gamma (0.9,1.1)\n\n**Data Pre-Processing:**\nper channel (img - img.mean()) / (img.std() + 1e-6)\nWe noticed that due to heavy augmentations we had a true_division warning. This warning caused batchnorms to feel bad, so we did this trick to prevent zero division.\n\n**Model training:**\nWe train our models with NVidia Apex and Pytorch 1.2 on 8x1080ti GPU server for 110 epochs. The training takes 1-3 days depending on the model.\nWe did not use oversample.\nBecause of the long iteration, we didn’t use any K-Fold training. We wanted to train it closer to the end of the competition, but we didn’t have much time. \n\n**Optimizer:** SGD\n\n**Scheduler:** init LR 0.1, \n5 epochs warmup,\ndrop LR every 40 epochs with factor 0.1\n\n**Loss Functions:**\nSmooth CrossEntropy for classification, Center Loss for embeddings. \n\n**Model:**\nOur final model is Senet154 trained on train data + pseudo-labels.\nTo produce pseudo-labels were using an ensemble SeNet154, SeResnext50, SeResnext101, Polynet, EfficientNet-b6, ResNeXt101-wsl. \n\n**Test time augmentations:**\noutput = (model(img1) + model(img2) + model(img1.flip(2)) + model(img1.flip(3)) + model(img2.flip(2)) + model(img2.flip(3)) / 6.\n\n**Post-processing:**\nHungry experiment re-calibration and plate re-calibration\nhttps://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270 \n\n**What didn’t work or had about the same performance:**\n- Mix-up, manifold mixup\n- ArcFace\n- Last stride 1\n- Focal Loss\n- XGBoost ensembling \n- Metric search based on the embeddings distances\n- GeM\n- Different optimizers like Adam, Radam, Ranger, etc.\n- ResNeXt101-wsl performed worse than expected. \n",
    "635197": "`Smooth CrossEntropy for classification, Center Loss for embeddings.`\nSo how do you infer predictions? How do you use those embeddings?\n\nAnd how do you select pseudo-labels?",
    "635146": ""
  }
}