{
  "id": 108058,
  "title": "7th place solution w/ code [Catalyst, Albumentations]",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108058",
  "author_name": "Eugene Khvedchenya",
  "post_date": "2019-09-08T19:10:08.776000",
  "votes": 102,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I joined this competition on a purpose to secure a solo gold medal, which is required to obtain a GM badge. Solving Kaggle challenges alone is hard, and I'm grateful for all who writes post-morten solutions. It's cool to learn other's approaches and realize we all came to more or less similar solutions. Here's my 5 cents on this competition:</p>\n\n<p>Code is available here: <a href=\"https://github.com/BloodAxe/Kaggle-2019-Blindness-Detection\">https://github.com/BloodAxe/Kaggle-2019-Blindness-Detection</a></p>\n\n<p><strong>TLDR</strong>\n0. SeResNext50, SeResNext101, InceptionV4\n1. Pre-train models on data from 2015, validate on this competition's data and Idrid (test)\n2. Train on this competition's data, Idrid (train) and Messidor\n3. Pseudo-label test dataset, Messidor 2, previous competition's data\n4. Fine-tune models\n5. Predict with horizontal flip TTA</p>\n\n<h2>Preprocessing</h2>\n\n<p>I've tried different preprorcessing methods, including Ben's and some in-house ones based on Clahe, dropping red channel and sperical image unwarping. <em>All of them failed to outperform using images as-is</em>. My explanation to this is the fact, that models we all used here are way more complex (in a number of parameters) and have much bigger capacity and generalization power, so preprocessing has vanishing impact.</p>\n\n<p>My preprocessing pipeline therefore was very simple:</p>\n\n<ul>\n<li>Crop image to get rid of black regions outside of eye's boundary</li>\n<li>Pad image if needed to square size</li>\n<li>Resize to 512x512 pixels (All models were trained in this resolution)</li>\n</ul>\n\n<h2>Models</h2>\n\n<p>I had an impression, that simple average pooling will not be enough to detect decease cases like \"Have hemorrhages in all 4 quadrants\". So I experimented with more advanced designs of heads:</p>\n\n<ul>\n<li>Global weighted average pooling</li>\n<li>Concatenation of average and max pooling</li>\n<li>Coord-Conv + FPN and max-pooling from each FPN layer</li>\n<li>Coord-Conv + LSTM pooling </li>\n</ul>\n\n<p>Most heads performed equally or worse than global average pooling (GAP), so I used GAP in my final models ensemble.</p>\n\n<h2>Task formulation</h2>\n\n<p>Started as classification problem, then switched to classical regression problem and later end up with ordinal-like problem. Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.</p>\n\n<h2>Loss functions</h2>\n\n<p>For classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)\nFor regression I started with MSE, then switched to WingLoss.\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.</p>\n\n<h2>Training</h2>\n\n<p>To increase batch size I used mixed-precision training in fp16 using Apex and 3x1080Ti. </p>\n\n<p>To prevent over-fitting my models and my augmentation pipeline was quite heavy one. Thanks to <a href=\"https://github.com/albu/albumentations\">Albumentations</a> package, it was extremely easy to combine and use them in my training loop:</p>\n\n<p><code>\n    return A.Compose([\n        A.OneOf([\n            A.ShiftScaleRotate(shift_limit=0.05, scale_limit=0.1,\n                               rotate_limit=15,\n                               border_mode=cv2.BORDER_CONSTANT, value=0),\n            A.OpticalDistortion(distort_limit=0.11, shift_limit=0.15,\n                                border_mode=cv2.BORDER_CONSTANT,\n                                value=0),\n            A.NoOp()\n        ]),\n        ZeroTopAndBottom(p=0.3),\n        A.RandomSizedCrop(min_max_height=(int(image_size[0] * 0.75), image_size[0]),\n                          height=image_size[0],\n                          width=image_size[1], p=0.3),\n        A.OneOf([\n            A.RandomBrightnessContrast(brightness_limit=0.5,\n                                       contrast_limit=0.4),\n            IndependentRandomBrightnessContrast(brightness_limit=0.25,\n                                                contrast_limit=0.24),\n            A.RandomGamma(gamma_limit=(50, 150)),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            FancyPCA(alpha_std=4),\n            A.RGBShift(r_shift_limit=20, b_shift_limit=15, g_shift_limit=15),\n            A.HueSaturationValue(hue_shift_limit=5,\n                                 sat_shift_limit=5),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            ChannelIndependentCLAHE(p=0.5),\n            A.CLAHE(),\n            A.NoOp()\n        ]),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5)\n    ])\n</code></p>\n\n<p>For the training, I've used to PyTorch framework and <a href=\"https://github.com/catalyst-team/catalyst\">Catalyst</a> DL framework. With just a bit of extra code for adding Kappa metrics.\nIt may looks like a promotion of these libraries and, in fact, is it, since I contribute to both ;) No, really. I strongly recommend you to try it out. There are many alternatives to fast.ai )</p>\n\n<h2>What didn't work</h2>\n\n<ul>\n<li>I had really high hopes for Unsupervised Data Augmentation (<a href=\"https://arxiv.org/abs/1904.12848\">https://arxiv.org/abs/1904.12848</a>). I even extended their approach for regression problem. For me it didn't work in terms for LB score. However it may be a good contribution to Catalyst.</li>\n<li>Decease augmentation. I tried to artificially generate aneurysms, cotton wools and hemorrhages and change diagnosis for augmented image. This didn't work at all. Probably my fakes were too obvious.</li>\n<li>Mixup / MixMatch. Not worked at all</li>\n<li>EfficientNet. Wasted couple of days on training it from scratch, but without success.</li>\n<li>Heavy models (Senet154, Resnext152, Densenet201). Even with 50% dropout, L2 regularization 1e-3 these monsters were over-fitting badly.</li>\n<li>Stacking. None of KNN, LogisticRegression, Trees didn't work for me at all. Simple averaging of regression scores worked best.</li>\n</ul>\n\n<h2>What did work</h2>\n\n<ul>\n<li>Pre-training on 2015 data</li>\n<li>Heavy augmentations</li>\n<li>Pseudo-labeling AND fine-tuning of head ONLY.</li>\n<li>RAdam. This becomes by default optimized.</li>\n</ul>\n\n<h2>Summary</h2>\n\n<p>Final ensemble was average of predictions of 4 models: SeResNext50, SeResNext101, InceptionV4 + InceptionV4 (finetune on pseudolabeled data) \nAfter revealing private LB, I realized I had submission with optimized thresholds scored at Top 4. Am I disappointed? No. At the time of choosing final submissions I did not consider it as being too risky.</p>\n\n<p>Am I satisfied with a results? Yes.\nCan I do better? Yes.</p>\n\n<p>So see you in the next competitions.</p>",
  "messages": [
    {
      "id": 621633,
      "postDate": "2019-09-08T19:10:08.777Z",
      "content": "<p>I joined this competition on a purpose to secure a solo gold medal, which is required to obtain a GM badge. Solving Kaggle challenges alone is hard, and I'm grateful for all who writes post-morten solutions. It's cool to learn other's approaches and realize we all came to more or less similar solutions. Here's my 5 cents on this competition:</p>\n\n<p>Code is available here: <a href=\"https://github.com/BloodAxe/Kaggle-2019-Blindness-Detection\">https://github.com/BloodAxe/Kaggle-2019-Blindness-Detection</a></p>\n\n<p><strong>TLDR</strong>\n0. SeResNext50, SeResNext101, InceptionV4\n1. Pre-train models on data from 2015, validate on this competition's data and Idrid (test)\n2. Train on this competition's data, Idrid (train) and Messidor\n3. Pseudo-label test dataset, Messidor 2, previous competition's data\n4. Fine-tune models\n5. Predict with horizontal flip TTA</p>\n\n<h2>Preprocessing</h2>\n\n<p>I've tried different preprorcessing methods, including Ben's and some in-house ones based on Clahe, dropping red channel and sperical image unwarping. <em>All of them failed to outperform using images as-is</em>. My explanation to this is the fact, that models we all used here are way more complex (in a number of parameters) and have much bigger capacity and generalization power, so preprocessing has vanishing impact.</p>\n\n<p>My preprocessing pipeline therefore was very simple:</p>\n\n<ul>\n<li>Crop image to get rid of black regions outside of eye's boundary</li>\n<li>Pad image if needed to square size</li>\n<li>Resize to 512x512 pixels (All models were trained in this resolution)</li>\n</ul>\n\n<h2>Models</h2>\n\n<p>I had an impression, that simple average pooling will not be enough to detect decease cases like \"Have hemorrhages in all 4 quadrants\". So I experimented with more advanced designs of heads:</p>\n\n<ul>\n<li>Global weighted average pooling</li>\n<li>Concatenation of average and max pooling</li>\n<li>Coord-Conv + FPN and max-pooling from each FPN layer</li>\n<li>Coord-Conv + LSTM pooling </li>\n</ul>\n\n<p>Most heads performed equally or worse than global average pooling (GAP), so I used GAP in my final models ensemble.</p>\n\n<h2>Task formulation</h2>\n\n<p>Started as classification problem, then switched to classical regression problem and later end up with ordinal-like problem. Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.</p>\n\n<h2>Loss functions</h2>\n\n<p>For classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)\nFor regression I started with MSE, then switched to WingLoss.\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.</p>\n\n<h2>Training</h2>\n\n<p>To increase batch size I used mixed-precision training in fp16 using Apex and 3x1080Ti. </p>\n\n<p>To prevent over-fitting my models and my augmentation pipeline was quite heavy one. Thanks to <a href=\"https://github.com/albu/albumentations\">Albumentations</a> package, it was extremely easy to combine and use them in my training loop:</p>\n\n<p><code>\n    return A.Compose([\n        A.OneOf([\n            A.ShiftScaleRotate(shift_limit=0.05, scale_limit=0.1,\n                               rotate_limit=15,\n                               border_mode=cv2.BORDER_CONSTANT, value=0),\n            A.OpticalDistortion(distort_limit=0.11, shift_limit=0.15,\n                                border_mode=cv2.BORDER_CONSTANT,\n                                value=0),\n            A.NoOp()\n        ]),\n        ZeroTopAndBottom(p=0.3),\n        A.RandomSizedCrop(min_max_height=(int(image_size[0] * 0.75), image_size[0]),\n                          height=image_size[0],\n                          width=image_size[1], p=0.3),\n        A.OneOf([\n            A.RandomBrightnessContrast(brightness_limit=0.5,\n                                       contrast_limit=0.4),\n            IndependentRandomBrightnessContrast(brightness_limit=0.25,\n                                                contrast_limit=0.24),\n            A.RandomGamma(gamma_limit=(50, 150)),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            FancyPCA(alpha_std=4),\n            A.RGBShift(r_shift_limit=20, b_shift_limit=15, g_shift_limit=15),\n            A.HueSaturationValue(hue_shift_limit=5,\n                                 sat_shift_limit=5),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            ChannelIndependentCLAHE(p=0.5),\n            A.CLAHE(),\n            A.NoOp()\n        ]),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5)\n    ])\n</code></p>\n\n<p>For the training, I've used to PyTorch framework and <a href=\"https://github.com/catalyst-team/catalyst\">Catalyst</a> DL framework. With just a bit of extra code for adding Kappa metrics.\nIt may looks like a promotion of these libraries and, in fact, is it, since I contribute to both ;) No, really. I strongly recommend you to try it out. There are many alternatives to fast.ai )</p>\n\n<h2>What didn't work</h2>\n\n<ul>\n<li>I had really high hopes for Unsupervised Data Augmentation (<a href=\"https://arxiv.org/abs/1904.12848\">https://arxiv.org/abs/1904.12848</a>). I even extended their approach for regression problem. For me it didn't work in terms for LB score. However it may be a good contribution to Catalyst.</li>\n<li>Decease augmentation. I tried to artificially generate aneurysms, cotton wools and hemorrhages and change diagnosis for augmented image. This didn't work at all. Probably my fakes were too obvious.</li>\n<li>Mixup / MixMatch. Not worked at all</li>\n<li>EfficientNet. Wasted couple of days on training it from scratch, but without success.</li>\n<li>Heavy models (Senet154, Resnext152, Densenet201). Even with 50% dropout, L2 regularization 1e-3 these monsters were over-fitting badly.</li>\n<li>Stacking. None of KNN, LogisticRegression, Trees didn't work for me at all. Simple averaging of regression scores worked best.</li>\n</ul>\n\n<h2>What did work</h2>\n\n<ul>\n<li>Pre-training on 2015 data</li>\n<li>Heavy augmentations</li>\n<li>Pseudo-labeling AND fine-tuning of head ONLY.</li>\n<li>RAdam. This becomes by default optimized.</li>\n</ul>\n\n<h2>Summary</h2>\n\n<p>Final ensemble was average of predictions of 4 models: SeResNext50, SeResNext101, InceptionV4 + InceptionV4 (finetune on pseudolabeled data) \nAfter revealing private LB, I realized I had submission with optimized thresholds scored at Top 4. Am I disappointed? No. At the time of choosing final submissions I did not consider it as being too risky.</p>\n\n<p>Am I satisfied with a results? Yes.\nCan I do better? Yes.</p>\n\n<p>So see you in the next competitions.</p>",
      "rawMarkdown": "I joined this competition on a purpose to secure a solo gold medal, which is required to obtain a GM badge. Solving Kaggle challenges alone is hard, and I'm grateful for all who writes post-morten solutions. It's cool to learn other's approaches and realize we all came to more or less similar solutions. Here's my 5 cents on this competition:\n\nCode is available here: https://github.com/BloodAxe/Kaggle-2019-Blindness-Detection\n\n**TLDR**\n0. SeResNext50, SeResNext101, InceptionV4\n1. Pre-train models on data from 2015, validate on this competition's data and Idrid (test)\n2. Train on this competition's data, Idrid (train) and Messidor\n3. Pseudo-label test dataset, Messidor 2, previous competition's data\n4. Fine-tune models\n5. Predict with horizontal flip TTA\n\n## Preprocessing\n\nI've tried different preprorcessing methods, including Ben's and some in-house ones based on Clahe, dropping red channel and sperical image unwarping. *All of them failed to outperform using images as-is*. My explanation to this is the fact, that models we all used here are way more complex (in a number of parameters) and have much bigger capacity and generalization power, so preprocessing has vanishing impact.\n\nMy preprocessing pipeline therefore was very simple:\n\n- Crop image to get rid of black regions outside of eye's boundary\n- Pad image if needed to square size\n- Resize to 512x512 pixels (All models were trained in this resolution)\n\n## Models\n\nI had an impression, that simple average pooling will not be enough to detect decease cases like \"Have hemorrhages in all 4 quadrants\". So I experimented with more advanced designs of heads:\n\n- Global weighted average pooling\n- Concatenation of average and max pooling\n- Coord-Conv + FPN and max-pooling from each FPN layer\n- Coord-Conv + LSTM pooling \n\nMost heads performed equally or worse than global average pooling (GAP), so I used GAP in my final models ensemble.\n\n## Task formulation\n\nStarted as classification problem, then switched to classical regression problem and later end up with ordinal-like problem. Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.\n\n## Loss functions\n\nFor classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)\nFor regression I started with MSE, then switched to WingLoss.\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.\n\n## Training\n\nTo increase batch size I used mixed-precision training in fp16 using Apex and 3x1080Ti. \n\nTo prevent over-fitting my models and my augmentation pipeline was quite heavy one. Thanks to [Albumentations](https://github.com/albu/albumentations) package, it was extremely easy to combine and use them in my training loop:\n\n```\n    return A.Compose([\n        A.OneOf([\n            A.ShiftScaleRotate(shift_limit=0.05, scale_limit=0.1,\n                               rotate_limit=15,\n                               border_mode=cv2.BORDER_CONSTANT, value=0),\n            A.OpticalDistortion(distort_limit=0.11, shift_limit=0.15,\n                                border_mode=cv2.BORDER_CONSTANT,\n                                value=0),\n            A.NoOp()\n        ]),\n        ZeroTopAndBottom(p=0.3),\n        A.RandomSizedCrop(min_max_height=(int(image_size[0] * 0.75), image_size[0]),\n                          height=image_size[0],\n                          width=image_size[1], p=0.3),\n        A.OneOf([\n            A.RandomBrightnessContrast(brightness_limit=0.5,\n                                       contrast_limit=0.4),\n            IndependentRandomBrightnessContrast(brightness_limit=0.25,\n                                                contrast_limit=0.24),\n            A.RandomGamma(gamma_limit=(50, 150)),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            FancyPCA(alpha_std=4),\n            A.RGBShift(r_shift_limit=20, b_shift_limit=15, g_shift_limit=15),\n            A.HueSaturationValue(hue_shift_limit=5,\n                                 sat_shift_limit=5),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            ChannelIndependentCLAHE(p=0.5),\n            A.CLAHE(),\n            A.NoOp()\n        ]),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5)\n    ])\n```\n\nFor the training, I've used to PyTorch framework and [Catalyst](https://github.com/catalyst-team/catalyst) DL framework. With just a bit of extra code for adding Kappa metrics.\nIt may looks like a promotion of these libraries and, in fact, is it, since I contribute to both ;) No, really. I strongly recommend you to try it out. There are many alternatives to fast.ai )\n\n## What didn't work\n\n- I had really high hopes for Unsupervised Data Augmentation (https://arxiv.org/abs/1904.12848). I even extended their approach for regression problem. For me it didn't work in terms for LB score. However it may be a good contribution to Catalyst.\n- Decease augmentation. I tried to artificially generate aneurysms, cotton wools and hemorrhages and change diagnosis for augmented image. This didn't work at all. Probably my fakes were too obvious.\n- Mixup / MixMatch. Not worked at all\n- EfficientNet. Wasted couple of days on training it from scratch, but without success.\n- Heavy models (Senet154, Resnext152, Densenet201). Even with 50% dropout, L2 regularization 1e-3 these monsters were over-fitting badly.\n- Stacking. None of KNN, LogisticRegression, Trees didn't work for me at all. Simple averaging of regression scores worked best.\n\n## What did work\n\n- Pre-training on 2015 data\n- Heavy augmentations\n- Pseudo-labeling AND fine-tuning of head ONLY.\n- RAdam. This becomes by default optimized.\n\n## Summary\n\nFinal ensemble was average of predictions of 4 models: SeResNext50, SeResNext101, InceptionV4 + InceptionV4 (finetune on pseudolabeled data) \nAfter revealing private LB, I realized I had submission with optimized thresholds scored at Top 4. Am I disappointed? No. At the time of choosing final submissions I did not consider it as being too risky.\n\nAm I satisfied with a results? Yes.\nCan I do better? Yes.\n\nSo see you in the next competitions.\n\n\n\n",
      "votes": 102
    },
    {
      "id": 622068,
      "postDate": "2019-09-09T08:41:46.480Z",
      "content": "<p>Congrats <a href=\"/bloodaxe\">@bloodaxe</a>! The solo gold is crazy hard work 💪 </p>",
      "rawMarkdown": "Congrats @bloodaxe! The solo gold is crazy hard work 💪 ",
      "votes": 4
    },
    {
      "id": 639466,
      "postDate": "2019-10-03T08:43:55.577Z",
      "content": "<p>Congratulations Eugene. The amazing effort did really paid off. Thank you for sharing.:)</p>",
      "rawMarkdown": "Congratulations Eugene. The amazing effort did really paid off. Thank you for sharing.:)",
      "votes": 1
    },
    {
      "id": 622437,
      "postDate": "2019-09-09T16:51:26.633Z",
      "content": "<p>Greetings! Nice one. Do you have an idea why Effnets didnt work at all?</p>",
      "rawMarkdown": "Greetings! Nice one. Do you have an idea why Effnets didnt work at all?",
      "votes": 1,
      "replies": [
        {
          "id": 622503,
          "postDate": "2019-09-09T18:27:55.500Z",
          "content": "<p>In my case, I think the reason was that I trained them from scratch (using my own porting from pytorch_toolbelt). And I found EffNets are very choosy to learning rate and initialization. I know there are pretrained weights already, but I was too lazy to update <a href=\"https://github.com/BloodAxe/pytorch-toolbelt\">pytorch_toolbelt</a> to support them.</p>",
          "rawMarkdown": "In my case, I think the reason was that I trained them from scratch (using my own porting from pytorch_toolbelt). And I found EffNets are very choosy to learning rate and initialization. I know there are pretrained weights already, but I was too lazy to update [pytorch_toolbelt](https://github.com/BloodAxe/pytorch-toolbelt) to support them.",
          "votes": 3
        }
      ]
    },
    {
      "id": 621898,
      "postDate": "2019-09-09T04:42:25.793Z",
      "content": "<p>Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! <a href=\"/bloodaxe\">@bloodaxe</a> </p>",
      "rawMarkdown": "Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! @bloodaxe ",
      "votes": 1,
      "replies": [
        {
          "id": 622218,
          "postDate": "2019-09-09T11:49:33.050Z",
          "content": "<p>Thanks! It is) The hardest and temping was to decline team-up proposals ;)</p>",
          "rawMarkdown": "Thanks! It is) The hardest and temping was to decline team-up proposals ;)",
          "votes": 2
        }
      ]
    },
    {
      "id": 621783,
      "postDate": "2019-09-09T00:17:37.223Z",
      "content": "<p>Thanks <a href=\"/bloodaxe\">@bloodaxe</a> ! May I ask: what was your best performing model, and what was its score?</p>",
      "rawMarkdown": "Thanks @bloodaxe ! May I ask: what was your best performing model, and what was its score?",
      "votes": 1,
      "replies": [
        {
          "id": 621887,
          "postDate": "2019-09-09T04:24:38.967Z",
          "content": "<p>A best performing model without pseudo labeling was seresnext50. With pseudo labeling it was inceptionv4. Single fold scored 922 on leaderboard.</p>",
          "rawMarkdown": "A best performing model without pseudo labeling was seresnext50. With pseudo labeling it was inceptionv4. Single fold scored 922 on leaderboard.",
          "votes": 2
        }
      ]
    },
    {
      "id": 623801,
      "postDate": "2019-09-11T10:14:10.670Z",
      "content": "<p>Congratulations with your solo gold! And thank you for your summary! So many insights. Amazing idea of task formulation, two new loss functions for me, nice technique of using mixed-precision training to increase the batch size. This all makes your solution somewhat unique. Good luck with your next competitions and reaching the GM badge!💪 </p>",
      "rawMarkdown": "Congratulations with your solo gold! And thank you for your summary! So many insights. Amazing idea of task formulation, two new loss functions for me, nice technique of using mixed-precision training to increase the batch size. This all makes your solution somewhat unique. Good luck with your next competitions and reaching the GM badge!💪 ",
      "votes": 2
    },
    {
      "id": 622887,
      "postDate": "2019-09-10T08:04:11.400Z",
      "content": "<p>Hi Eugene, many congrats to you :) Great solution. I'm happy to see you here on Kaggle. </p>",
      "rawMarkdown": "Hi Eugene, many congrats to you :) Great solution. I'm happy to see you here on Kaggle. ",
      "votes": 2
    },
    {
      "id": 622245,
      "postDate": "2019-09-09T12:28:50.293Z",
      "content": "<p>Congratulations Eugene and thanks for detailed review! Going to use your albumentations library...</p>",
      "rawMarkdown": "Congratulations Eugene and thanks for detailed review! Going to use your albumentations library...",
      "votes": 2
    },
    {
      "id": 622242,
      "postDate": "2019-09-09T12:27:16.617Z",
      "content": "<p>Congrats <a href=\"/bloodaxe\">@bloodaxe</a>  for the solo performance !</p>\n\n<blockquote>\n  <p>Thanks to Albumentations package,</p>\n</blockquote>\n\n<p>BTW I noticed you're one of the main and core contributors to Albumentations.  So you should say \"Thanks to myself\" ;)   </p>\n\n<p>p.s:  This library is amazing and huge time saver for (heavy) augmentation . </p>",
      "rawMarkdown": "Congrats @bloodaxe  for the solo performance !\n\n&gt;  Thanks to Albumentations package,\n\nBTW I noticed you're one of the main and core contributors to Albumentations.  So you should say \"Thanks to myself\" ;)   \n\np.s:  This library is amazing and huge time saver for (heavy) augmentation . ",
      "votes": 2,
      "replies": [
        {
          "id": 622505,
          "postDate": "2019-09-09T18:29:15.400Z",
          "content": "<p>Community who drives this forward, is a key. So thanks to everyone who uses it, sends PR and report issues :)</p>",
          "rawMarkdown": "Community who drives this forward, is a key. So thanks to everyone who uses it, sends PR and report issues :)",
          "votes": 5
        }
      ]
    },
    {
      "id": 621768,
      "postDate": "2019-09-08T23:12:55.320Z",
      "content": "<p>Hi <a href=\"/bloodaxe\">@bloodaxe</a> Eugene, thanks for sharing! I thought that I used a heavy augmentation already, but it seems you are the real heavy one! Congrat for the results and see you next time!</p>\n\n<p>BTW, thanks also for intoducing (and contributing) Catalyst. I am eager to learn and use it!</p>",
      "rawMarkdown": "Hi @bloodaxe Eugene, thanks for sharing! I thought that I used a heavy augmentation already, but it seems you are the real heavy one! Congrat for the results and see you next time!\n\nBTW, thanks also for intoducing (and contributing) Catalyst. I am eager to learn and use it!",
      "votes": 2,
      "replies": [
        {
          "id": 621966,
          "postDate": "2019-09-09T06:05:18.797Z",
          "content": "<p>Hi! Looking forward to see you in Severstal and clouds too! \nPS: Feel free to ask question regarding those libs in ods.ai slack in #tools_catalyst and #tools_albumentations ;)</p>",
          "rawMarkdown": "Hi! Looking forward to see you in Severstal and clouds too! \nPS: Feel free to ask question regarding those libs in ods.ai slack in #tools_catalyst and #tools_albumentations ;)",
          "votes": 4
        }
      ]
    },
    {
      "id": 1224892,
      "postDate": "2021-03-03T06:40:16.360Z",
      "content": "<p>Respected Sir,</p>\n<p>Why you use three loss function??<br>\nWhy u use classification, regression and ordinal regression head??<br>\nPls help me..pls reply me..</p>\n<p>Loss functions</p>\n<p>For classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)<br>\nFor regression I started with MSE, then switched to WingLoss.<br>\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.</p>",
      "rawMarkdown": "Respected Sir,\n\nWhy you use three loss function??\nWhy u use classification, regression and ordinal regression head??\nPls help me..pls reply me..\n\nLoss functions\n\nFor classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)\nFor regression I started with MSE, then switched to WingLoss.\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss."
    },
    {
      "id": 837270,
      "postDate": "2020-05-07T16:46:57.997Z",
      "content": "<p>Hi <a href=\"/bloodaxe\">@bloodaxe</a> \nCould I ask a question?</p>\n\n<p>Did you ass sigmoid activation on ordinal regression or regression or classification?</p>\n\n<p>Because you said: \n\"Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.\"</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hi @bloodaxe \nCould I ask a question?\n\nDid you ass sigmoid activation on ordinal regression or regression or classification?\n\nBecause you said: \n\"Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.\"\n\nThank you\n",
      "replies": [
        {
          "id": 842973,
          "postDate": "2020-05-11T18:24:46.550Z",
          "content": "<p>If I got your question right, the answer would be classical regression, e.g not ordinal.\nThe only trick, is that instead of having last layer as <code>nn.Linear(input, 1)</code> I used <code>nn.Linear(input, 4) + sigmoid + sum</code>. Hope this answers your question.</p>",
          "rawMarkdown": "If I got your question right, the answer would be classical regression, e.g not ordinal.\nThe only trick, is that instead of having last layer as `nn.Linear(input, 1)` I used `nn.Linear(input, 4) + sigmoid + sum`. Hope this answers your question.",
          "votes": 1
        }
      ]
    },
    {
      "id": 625024,
      "postDate": "2019-09-12T16:17:48.380Z",
      "content": "<p>Congratulations for your performance.\nI'm a newbie and if it were possible I'd ask you to explain on your GitHub repository how to execute the code to perform the training and inference that earned you seventh place\nI thank you in advance</p>",
      "rawMarkdown": "Congratulations for your performance.\nI'm a newbie and if it were possible I'd ask you to explain on your GitHub repository how to execute the code to perform the training and inference that earned you seventh place\nI thank you in advance"
    },
    {
      "id": 1110119,
      "postDate": "2020-12-12T12:51:49.983Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1114005,
          "postDate": "2020-12-15T22:05:52.483Z",
          "content": "<p>Mostly Test &amp; Trial and a bit of common sense</p>",
          "rawMarkdown": "Mostly Test & Trial and a bit of common sense"
        }
      ]
    },
    {
      "id": 623274,
      "postDate": "2019-09-10T17:02:02.983Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 622668,
      "postDate": "2019-09-10T00:28:46.677Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 622068,
      "author_name": "Kostya Kravchenko",
      "author_url": "",
      "post_date": "2019-09-09T08:41:46.480000",
      "content": "<p>Congrats <a href=\"/bloodaxe\">@bloodaxe</a>! The solo gold is crazy hard work 💪 </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 639466,
      "author_name": "Manish Kumar",
      "author_url": "",
      "post_date": "2019-10-03T08:43:55.577000",
      "content": "<p>Congratulations Eugene. The amazing effort did really paid off. Thank you for sharing.:)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 622437,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-09-09T16:51:26.633000",
      "content": "<p>Greetings! Nice one. Do you have an idea why Effnets didnt work at all?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 622503,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2019-09-09T18:27:55.500000",
          "content": "<p>In my case, I think the reason was that I trained them from scratch (using my own porting from pytorch_toolbelt). And I found EffNets are very choosy to learning rate and initialization. I know there are pretrained weights already, but I was too lazy to update <a href=\"https://github.com/BloodAxe/pytorch-toolbelt\">pytorch_toolbelt</a> to support them.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 621898,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-09-09T04:42:25.793000",
      "content": "<p>Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! <a href=\"/bloodaxe\">@bloodaxe</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 622218,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2019-09-09T11:49:33.050000",
          "content": "<p>Thanks! It is) The hardest and temping was to decline team-up proposals ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 621783,
      "author_name": "xhlulu",
      "author_url": "",
      "post_date": "2019-09-09T00:17:37.223000",
      "content": "<p>Thanks <a href=\"/bloodaxe\">@bloodaxe</a> ! May I ask: what was your best performing model, and what was its score?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621887,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2019-09-09T04:24:38.967000",
          "content": "<p>A best performing model without pseudo labeling was seresnext50. With pseudo labeling it was inceptionv4. Single fold scored 922 on leaderboard.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 623801,
      "author_name": "Evgeny Kovalev",
      "author_url": "",
      "post_date": "2019-09-11T10:14:10.670000",
      "content": "<p>Congratulations with your solo gold! And thank you for your summary! So many insights. Amazing idea of task formulation, two new loss functions for me, nice technique of using mixed-precision training to increase the batch size. This all makes your solution somewhat unique. Good luck with your next competitions and reaching the GM badge!💪 </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 622887,
      "author_name": "Filemon",
      "author_url": "",
      "post_date": "2019-09-10T08:04:11.400000",
      "content": "<p>Hi Eugene, many congrats to you :) Great solution. I'm happy to see you here on Kaggle. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 622245,
      "author_name": "Anna Novikova",
      "author_url": "",
      "post_date": "2019-09-09T12:28:50.293000",
      "content": "<p>Congratulations Eugene and thanks for detailed review! Going to use your albumentations library...</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 622242,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2019-09-09T12:27:16.617000",
      "content": "<p>Congrats <a href=\"/bloodaxe\">@bloodaxe</a>  for the solo performance !</p>\n\n<blockquote>\n  <p>Thanks to Albumentations package,</p>\n</blockquote>\n\n<p>BTW I noticed you're one of the main and core contributors to Albumentations.  So you should say \"Thanks to myself\" ;)   </p>\n\n<p>p.s:  This library is amazing and huge time saver for (heavy) augmentation . </p>",
      "votes": 2,
      "replies": [
        {
          "id": 622505,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2019-09-09T18:29:15.400000",
          "content": "<p>Community who drives this forward, is a key. So thanks to everyone who uses it, sends PR and report issues :)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 621768,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-08T23:12:55.320000",
      "content": "<p>Hi <a href=\"/bloodaxe\">@bloodaxe</a> Eugene, thanks for sharing! I thought that I used a heavy augmentation already, but it seems you are the real heavy one! Congrat for the results and see you next time!</p>\n\n<p>BTW, thanks also for intoducing (and contributing) Catalyst. I am eager to learn and use it!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 621966,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2019-09-09T06:05:18.797000",
          "content": "<p>Hi! Looking forward to see you in Severstal and clouds too! \nPS: Feel free to ask question regarding those libs in ods.ai slack in #tools_catalyst and #tools_albumentations ;)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1224892,
      "author_name": "dhaval",
      "author_url": "",
      "post_date": "2021-03-03T06:40:16.360000",
      "content": "<p>Respected Sir,</p>\n<p>Why you use three loss function??<br>\nWhy u use classification, regression and ordinal regression head??<br>\nPls help me..pls reply me..</p>\n<p>Loss functions</p>\n<p>For classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)<br>\nFor regression I started with MSE, then switched to WingLoss.<br>\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 837270,
      "author_name": "YS",
      "author_url": "",
      "post_date": "2020-05-07T16:46:57.997000",
      "content": "<p>Hi <a href=\"/bloodaxe\">@bloodaxe</a> \nCould I ask a question?</p>\n\n<p>Did you ass sigmoid activation on ordinal regression or regression or classification?</p>\n\n<p>Because you said: \n\"Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.\"</p>\n\n<p>Thank you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 842973,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-05-11T18:24:46.550000",
          "content": "<p>If I got your question right, the answer would be classical regression, e.g not ordinal.\nThe only trick, is that instead of having last layer as <code>nn.Linear(input, 1)</code> I used <code>nn.Linear(input, 4) + sigmoid + sum</code>. Hope this answers your question.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 625024,
      "author_name": "Andrea Zappacosta",
      "author_url": "",
      "post_date": "2019-09-12T16:17:48.380000",
      "content": "<p>Congratulations for your performance.\nI'm a newbie and if it were possible I'd ask you to explain on your GitHub repository how to execute the code to perform the training and inference that earned you seventh place\nI thank you in advance</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1110119,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-12T12:51:49.983000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1114005,
          "author_name": "Eugene Khvedchenya",
          "author_url": "",
          "post_date": "2020-12-15T22:05:52.483000",
          "content": "<p>Mostly Test &amp; Trial and a bit of common sense</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 623274,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T17:02:02.983000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 622668,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T00:28:46.677000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621633": "I joined this competition on a purpose to secure a solo gold medal, which is required to obtain a GM badge. Solving Kaggle challenges alone is hard, and I'm grateful for all who writes post-morten solutions. It's cool to learn other's approaches and realize we all came to more or less similar solutions. Here's my 5 cents on this competition:\n\nCode is available here: https://github.com/BloodAxe/Kaggle-2019-Blindness-Detection\n\n**TLDR**\n0. SeResNext50, SeResNext101, InceptionV4\n1. Pre-train models on data from 2015, validate on this competition's data and Idrid (test)\n2. Train on this competition's data, Idrid (train) and Messidor\n3. Pseudo-label test dataset, Messidor 2, previous competition's data\n4. Fine-tune models\n5. Predict with horizontal flip TTA\n\n## Preprocessing\n\nI've tried different preprorcessing methods, including Ben's and some in-house ones based on Clahe, dropping red channel and sperical image unwarping. *All of them failed to outperform using images as-is*. My explanation to this is the fact, that models we all used here are way more complex (in a number of parameters) and have much bigger capacity and generalization power, so preprocessing has vanishing impact.\n\nMy preprocessing pipeline therefore was very simple:\n\n- Crop image to get rid of black regions outside of eye's boundary\n- Pad image if needed to square size\n- Resize to 512x512 pixels (All models were trained in this resolution)\n\n## Models\n\nI had an impression, that simple average pooling will not be enough to detect decease cases like \"Have hemorrhages in all 4 quadrants\". So I experimented with more advanced designs of heads:\n\n- Global weighted average pooling\n- Concatenation of average and max pooling\n- Coord-Conv + FPN and max-pooling from each FPN layer\n- Coord-Conv + LSTM pooling \n\nMost heads performed equally or worse than global average pooling (GAP), so I used GAP in my final models ensemble.\n\n## Task formulation\n\nStarted as classification problem, then switched to classical regression problem and later end up with ordinal-like problem. Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.\n\n## Loss functions\n\nFor classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)\nFor regression I started with MSE, then switched to WingLoss.\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.\n\n## Training\n\nTo increase batch size I used mixed-precision training in fp16 using Apex and 3x1080Ti. \n\nTo prevent over-fitting my models and my augmentation pipeline was quite heavy one. Thanks to [Albumentations](https://github.com/albu/albumentations) package, it was extremely easy to combine and use them in my training loop:\n\n```\n    return A.Compose([\n        A.OneOf([\n            A.ShiftScaleRotate(shift_limit=0.05, scale_limit=0.1,\n                               rotate_limit=15,\n                               border_mode=cv2.BORDER_CONSTANT, value=0),\n            A.OpticalDistortion(distort_limit=0.11, shift_limit=0.15,\n                                border_mode=cv2.BORDER_CONSTANT,\n                                value=0),\n            A.NoOp()\n        ]),\n        ZeroTopAndBottom(p=0.3),\n        A.RandomSizedCrop(min_max_height=(int(image_size[0] * 0.75), image_size[0]),\n                          height=image_size[0],\n                          width=image_size[1], p=0.3),\n        A.OneOf([\n            A.RandomBrightnessContrast(brightness_limit=0.5,\n                                       contrast_limit=0.4),\n            IndependentRandomBrightnessContrast(brightness_limit=0.25,\n                                                contrast_limit=0.24),\n            A.RandomGamma(gamma_limit=(50, 150)),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            FancyPCA(alpha_std=4),\n            A.RGBShift(r_shift_limit=20, b_shift_limit=15, g_shift_limit=15),\n            A.HueSaturationValue(hue_shift_limit=5,\n                                 sat_shift_limit=5),\n            A.NoOp()\n        ]),\n        A.OneOf([\n            ChannelIndependentCLAHE(p=0.5),\n            A.CLAHE(),\n            A.NoOp()\n        ]),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5)\n    ])\n```\n\nFor the training, I've used to PyTorch framework and [Catalyst](https://github.com/catalyst-team/catalyst) DL framework. With just a bit of extra code for adding Kappa metrics.\nIt may looks like a promotion of these libraries and, in fact, is it, since I contribute to both ;) No, really. I strongly recommend you to try it out. There are many alternatives to fast.ai )\n\n## What didn't work\n\n- I had really high hopes for Unsupervised Data Augmentation (https://arxiv.org/abs/1904.12848). I even extended their approach for regression problem. For me it didn't work in terms for LB score. However it may be a good contribution to Catalyst.\n- Decease augmentation. I tried to artificially generate aneurysms, cotton wools and hemorrhages and change diagnosis for augmented image. This didn't work at all. Probably my fakes were too obvious.\n- Mixup / MixMatch. Not worked at all\n- EfficientNet. Wasted couple of days on training it from scratch, but without success.\n- Heavy models (Senet154, Resnext152, Densenet201). Even with 50% dropout, L2 regularization 1e-3 these monsters were over-fitting badly.\n- Stacking. None of KNN, LogisticRegression, Trees didn't work for me at all. Simple averaging of regression scores worked best.\n\n## What did work\n\n- Pre-training on 2015 data\n- Heavy augmentations\n- Pseudo-labeling AND fine-tuning of head ONLY.\n- RAdam. This becomes by default optimized.\n\n## Summary\n\nFinal ensemble was average of predictions of 4 models: SeResNext50, SeResNext101, InceptionV4 + InceptionV4 (finetune on pseudolabeled data) \nAfter revealing private LB, I realized I had submission with optimized thresholds scored at Top 4. Am I disappointed? No. At the time of choosing final submissions I did not consider it as being too risky.\n\nAm I satisfied with a results? Yes.\nCan I do better? Yes.\n\nSo see you in the next competitions.\n\n\n\n",
    "622068": "Congrats @bloodaxe! The solo gold is crazy hard work 💪 ",
    "639466": "Congratulations Eugene. The amazing effort did really paid off. Thank you for sharing.:)",
    "622437": "Greetings! Nice one. Do you have an idea why Effnets didnt work at all?",
    "621898": "Congratulations...\nGreat Work...\nThanks for Sharing your Approach &amp; Insights... !! @bloodaxe ",
    "621783": "Thanks @bloodaxe ! May I ask: what was your best performing model, and what was its score?",
    "623801": "Congratulations with your solo gold! And thank you for your summary! So many insights. Amazing idea of task formulation, two new loss functions for me, nice technique of using mixed-precision training to increase the batch size. This all makes your solution somewhat unique. Good luck with your next competitions and reaching the GM badge!💪 ",
    "622887": "Hi Eugene, many congrats to you :) Great solution. I'm happy to see you here on Kaggle. ",
    "622245": "Congratulations Eugene and thanks for detailed review! Going to use your albumentations library...",
    "622242": "Congrats @bloodaxe  for the solo performance !\n\n&gt;  Thanks to Albumentations package,\n\nBTW I noticed you're one of the main and core contributors to Albumentations.  So you should say \"Thanks to myself\" ;)   \n\np.s:  This library is amazing and huge time saver for (heavy) augmentation . ",
    "621768": "Hi @bloodaxe Eugene, thanks for sharing! I thought that I used a heavy augmentation already, but it seems you are the real heavy one! Congrat for the results and see you next time!\n\nBTW, thanks also for intoducing (and contributing) Catalyst. I am eager to learn and use it!",
    "1224892": "Respected Sir,\n\nWhy you use three loss function??\nWhy u use classification, regression and ordinal regression head??\nPls help me..pls reply me..\n\nLoss functions\n\nFor classification, I experimented with soft-CE (label smoothing) (which was better that plain CE), focal + kappa (which was better than soft-CE)\nFor regression I started with MSE, then switched to WingLoss.\nFor ordinal regression I used Huber loss (aka smooth L1 loss) but at the end used Cauchy loss.",
    "837270": "Hi @bloodaxe \nCould I ask a question?\n\nDid you ass sigmoid activation on ordinal regression or regression or classification?\n\nBecause you said: \n\"Unlike some authors suggested to use encoded target vectors, my approach was a bit different - my final linear layer predicted tensor of 4 elements with sigmoid activation followed by summation. By formulating problem like this I was able to use MSE loss and still have predictions bounded at [0;4]. To me, this problem formulation worked best.\"\n\nThank you\n",
    "625024": "Congratulations for your performance.\nI'm a newbie and if it were possible I'd ask you to explain on your GitHub repository how to execute the code to perform the training and inference that earned you seventh place\nI thank you in advance",
    "1110119": "",
    "623274": "",
    "622668": ""
  }
}