{
  "id": 94837,
  "title": "9th place solution (Knowledge Distillation)",
  "url": "/competitions/imet-2019-fgvc6/writeups/appian-9th-place-solution-knowledge-distillation",
  "author_name": "",
  "post_date": "2019-06-16T09:18:32.030Z",
  "votes": 60,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi everyone. This is my solution overview which scored 0.664 on public LB.</p>\n\n<h3>Competition's challenge</h3>\n\n<ol>\n<li>9 hours kernels limit for inference. </li>\n<li>Lack of proper labels. For example if an image has <strong>culture::american</strong>, there also should be <strong>culture::american or british</strong> but in many cases these labels are lacked. </li>\n</ol>\n\n<h3>Approach</h3>\n\n<ol>\n<li>I used <strong>knowledge distillation</strong> to compress the knowledge in an ensemble into student models so that I could use 9 hours more efficiently.</li>\n<li>Knowledge distillation also attends to the lack of proper labels by providing soft targets to training images. Because competition metric favors recall over precision, this should improve lb.</li>\n</ol>\n\n<h3>Hardware</h3>\n\n<p>I used kaggle kernels to train almost all the models. </p>\n\n<h3>Model training</h3>\n\n<p>The model training process can be split into 2 parts. \n- Train <strong>teacher models</strong>\n- Train <strong>student models</strong> using outputs of the teacher</p>\n\n<h3>Teacher models</h3>\n\n<ul>\n<li>2x se_resnext101</li>\n<li>2x se_resnext50</li>\n</ul>\n\n<p>I trained 4 models and averaged their outputs. I shared the performance of them during the competition and you can check that discussion here. <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/92159\">https://www.kaggle.com/c/imet-2019-fgvc6/discussion/92159</a></p>\n\n<p>Averaged output scored <strong>cv 0.6467</strong> and <strong>lb 0.656</strong>. This averaged output(cv 0.6467) on train images were used to train the student models next. </p>\n\n<p>I did not treat tags and cultures separately.</p>\n\n<p><strong>training method</strong>\nI trained only a dense layer for first 2 epochs because their weights are fresh. I got this idea from the kernel Lopuhin kindly shared and result is better with this.\n- 1st epoch (1/10 of baselr, dense layer only)\n- 2nd epoch (5/10 of baselr, dense layer only)\n- 3rd epoch (1/10 of baselr, all layers)\n- 4th epoch (5/10 of baselr, all layers)</p>\n\n<p><strong>se_resnext101</strong>\n- image size: 320\n- optimizer: Adam\n- base lr: 1.2e-4\n- batch size: 36\n- scheduler: ReduceLrOnPlateau or StepLR\n- best epoch: 12 to 15\n- loss: BCEWithLogitsLoss + FBetaLoss\n- augmentations: RandomResizedCrop, Horizontal Flip, Random Erasing\n- TTA: RandomResizedCrop, Horizontal Flip (10 times)\n- Training time: around 13 hours (Kaggle kernel)</p>\n\n<p><strong>performance of 5-fold se_resnext101</strong>\n- cv each: 0.627 | 0.629 | 0.627 | 0.628 | 0.630\n- lb each: 0.630 | ---\n- cv(concat): 0.628\n- lb(mean): <strong>0.652</strong></p>\n\n<p><strong>se_resnext50</strong>\n- base lr: 1.4e-4\n- batch size: 52\n- Training time: around 7 hours</p>\n\n<p><strong>performance of 5-fold se_resnext50</strong>\n- cv each: 0.621 | 0.623 | 0.620 | 0.618 | 0.622\n- lb each: 0.621 | ---\n- cv(concat): 0.621 \n- lb(mean): <strong>0.641</strong></p>\n\n<p><strong>augmentations</strong>\n```\nfrom torchvision import transforms as T</p>\n\n<p>def train_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n        RandomErasing(probability=0.3, sh=0.3),\n    ])</p>\n\n<p>def test_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```</p>\n\n<h3>Student models</h3>\n\n<ul>\n<li>1x se_resnext101</li>\n<li>1x inceptionresnetv2</li>\n</ul>\n\n<p>I trained a new 6-fold se_resnext101 using outputs of the teacher model on training images. I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets. Test images were remained untouched. </p>\n\n<p>Training parameters remained almost the same from teacher models. This 6-fold model scored <strong>cv 0.668</strong> and <strong>lb 0.662</strong> by itself. </p>\n\n<p><strong>performance of 6-fold se_resnext101 (student)</strong>\n- cv each: 0.669 | 0.669 | 0.668 | 0.664 | 0.665 | 0.668\n- lb each: 0.649 | 0.648 | ---\n- cv(concat): 0.668\n- lb(mean): <strong>0.662</strong></p>\n\n<p>I also trained inceptionresnetv2 and averaged with se_resnext101 and scored <strong>lb 0.664</strong>. </p>\n\n<h3>Inference</h3>\n\n<p>The inference part is very simple. I just averaged outputs of the student models. # of TTA is 7. </p>\n\n<h3>Possible Improvements</h3>\n\n<p>I have to admit there are lots of space for improvements. \n- Train more teacher models / student models such as senet152, pnasnet for ensembling. My solution definitely lacks variants for better ensembling. \n- Treat cultures and tags separately for training, thresholding as they have different characteristics. </p>\n\n<p>I'd like to thank Lopuhin for sharing his great kernel where I borrowed some ideas/implementations such as making folds, binarizing outputs. <a href=\"https://www.kaggle.com/lopuhin/imet-2019-submission\">https://www.kaggle.com/lopuhin/imet-2019-submission</a></p>\n\n<p>I also would like to thank Bac Nguyen and his implementation of FbetaLoss. <a href=\"https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss\">https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss</a></p>\n\n<p>Thank you kaggle team and Zhang for hosting such a great competition. </p>\n\n<p>Thanks for reading.</p>",
  "messages": [
    {
      "id": "547098",
      "postDate": "06/07/2019 09:25:26",
      "content": "<p>Hi everyone. This is my solution overview which scored 0.664 on public LB.</p>\n\n<h3>Competition's challenge</h3>\n\n<ol>\n<li>9 hours kernels limit for inference. </li>\n<li>Lack of proper labels. For example if an image has <strong>culture::american</strong>, there also should be <strong>culture::american or british</strong> but in many cases these labels are lacked. </li>\n</ol>\n\n<h3>Approach</h3>\n\n<ol>\n<li>I used <strong>knowledge distillation</strong> to compress the knowledge in an ensemble into student models so that I could use 9 hours more efficiently.</li>\n<li>Knowledge distillation also attends to the lack of proper labels by providing soft targets to training images. Because competition metric favors recall over precision, this should improve lb.</li>\n</ol>\n\n<h3>Hardware</h3>\n\n<p>I used kaggle kernels to train almost all the models. </p>\n\n<h3>Model training</h3>\n\n<p>The model training process can be split into 2 parts. \n- Train <strong>teacher models</strong>\n- Train <strong>student models</strong> using outputs of the teacher</p>\n\n<h3>Teacher models</h3>\n\n<ul>\n<li>2x se_resnext101</li>\n<li>2x se_resnext50</li>\n</ul>\n\n<p>I trained 4 models and averaged their outputs. I shared the performance of them during the competition and you can check that discussion here. <a href=\"https://www.kaggle.com/c/imet-2019-fgvc6/discussion/92159\">https://www.kaggle.com/c/imet-2019-fgvc6/discussion/92159</a></p>\n\n<p>Averaged output scored <strong>cv 0.6467</strong> and <strong>lb 0.656</strong>. This averaged output(cv 0.6467) on train images were used to train the student models next. </p>\n\n<p>I did not treat tags and cultures separately.</p>\n\n<p><strong>training method</strong>\nI trained only a dense layer for first 2 epochs because their weights are fresh. I got this idea from the kernel Lopuhin kindly shared and result is better with this.\n- 1st epoch (1/10 of baselr, dense layer only)\n- 2nd epoch (5/10 of baselr, dense layer only)\n- 3rd epoch (1/10 of baselr, all layers)\n- 4th epoch (5/10 of baselr, all layers)</p>\n\n<p><strong>se_resnext101</strong>\n- image size: 320\n- optimizer: Adam\n- base lr: 1.2e-4\n- batch size: 36\n- scheduler: ReduceLrOnPlateau or StepLR\n- best epoch: 12 to 15\n- loss: BCEWithLogitsLoss + FBetaLoss\n- augmentations: RandomResizedCrop, Horizontal Flip, Random Erasing\n- TTA: RandomResizedCrop, Horizontal Flip (10 times)\n- Training time: around 13 hours (Kaggle kernel)</p>\n\n<p><strong>performance of 5-fold se_resnext101</strong>\n- cv each: 0.627 | 0.629 | 0.627 | 0.628 | 0.630\n- lb each: 0.630 | ---\n- cv(concat): 0.628\n- lb(mean): <strong>0.652</strong></p>\n\n<p><strong>se_resnext50</strong>\n- base lr: 1.4e-4\n- batch size: 52\n- Training time: around 7 hours</p>\n\n<p><strong>performance of 5-fold se_resnext50</strong>\n- cv each: 0.621 | 0.623 | 0.620 | 0.618 | 0.622\n- lb each: 0.621 | ---\n- cv(concat): 0.621 \n- lb(mean): <strong>0.641</strong></p>\n\n<p><strong>augmentations</strong>\n```\nfrom torchvision import transforms as T</p>\n\n<p>def train_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n        RandomErasing(probability=0.3, sh=0.3),\n    ])</p>\n\n<p>def test_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```</p>\n\n<h3>Student models</h3>\n\n<ul>\n<li>1x se_resnext101</li>\n<li>1x inceptionresnetv2</li>\n</ul>\n\n<p>I trained a new 6-fold se_resnext101 using outputs of the teacher model on training images. I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets. Test images were remained untouched. </p>\n\n<p>Training parameters remained almost the same from teacher models. This 6-fold model scored <strong>cv 0.668</strong> and <strong>lb 0.662</strong> by itself. </p>\n\n<p><strong>performance of 6-fold se_resnext101 (student)</strong>\n- cv each: 0.669 | 0.669 | 0.668 | 0.664 | 0.665 | 0.668\n- lb each: 0.649 | 0.648 | ---\n- cv(concat): 0.668\n- lb(mean): <strong>0.662</strong></p>\n\n<p>I also trained inceptionresnetv2 and averaged with se_resnext101 and scored <strong>lb 0.664</strong>. </p>\n\n<h3>Inference</h3>\n\n<p>The inference part is very simple. I just averaged outputs of the student models. # of TTA is 7. </p>\n\n<h3>Possible Improvements</h3>\n\n<p>I have to admit there are lots of space for improvements. \n- Train more teacher models / student models such as senet152, pnasnet for ensembling. My solution definitely lacks variants for better ensembling. \n- Treat cultures and tags separately for training, thresholding as they have different characteristics. </p>\n\n<p>I'd like to thank Lopuhin for sharing his great kernel where I borrowed some ideas/implementations such as making folds, binarizing outputs. <a href=\"https://www.kaggle.com/lopuhin/imet-2019-submission\">https://www.kaggle.com/lopuhin/imet-2019-submission</a></p>\n\n<p>I also would like to thank Bac Nguyen and his implementation of FbetaLoss. <a href=\"https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss\">https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss</a></p>\n\n<p>Thank you kaggle team and Zhang for hosting such a great competition. </p>\n\n<p>Thanks for reading.</p>",
      "rawMarkdown": "Hi everyone. This is my solution overview which scored 0.664 on public LB.\n\n\n### Competition's challenge\n\n1. 9 hours kernels limit for inference. \n2. Lack of proper labels. For example if an image has **culture::american**, there also should be **culture::american or british** but in many cases these labels are lacked. \n\n\n### Approach\n\n1. I used **knowledge distillation** to compress the knowledge in an ensemble into student models so that I could use 9 hours more efficiently.\n2. Knowledge distillation also attends to the lack of proper labels by providing soft targets to training images. Because competition metric favors recall over precision, this should improve lb.\n\n\n### Hardware\n\nI used kaggle kernels to train almost all the models. \n\n\n### Model training\n\nThe model training process can be split into 2 parts. \n- Train **teacher models**\n- Train **student models** using outputs of the teacher\n\n\n### Teacher models\n\n- 2x se\\_resnext101\n- 2x se\\_resnext50\n\nI trained 4 models and averaged their outputs. I shared the performance of them during the competition and you can check that discussion here. https://www.kaggle.com/c/imet-2019-fgvc6/discussion/92159\n\nAveraged output scored **cv 0.6467** and **lb 0.656**. This averaged output(cv 0.6467) on train images were used to train the student models next. \n\nI did not treat tags and cultures separately.\n\n**training method**\nI trained only a dense layer for first 2 epochs because their weights are fresh. I got this idea from the kernel Lopuhin kindly shared and result is better with this.\n- 1st epoch (1/10 of baselr, dense layer only)\n- 2nd epoch (5/10 of baselr, dense layer only)\n- 3rd epoch (1/10 of baselr, all layers)\n- 4th epoch (5/10 of baselr, all layers)\n\n**se\\_resnext101**\n- image size: 320\n- optimizer: Adam\n- base lr: 1.2e-4\n- batch size: 36\n- scheduler: ReduceLrOnPlateau or StepLR\n- best epoch: 12 to 15\n- loss: BCEWithLogitsLoss + FBetaLoss\n- augmentations: RandomResizedCrop, Horizontal Flip, Random Erasing\n- TTA: RandomResizedCrop, Horizontal Flip (10 times)\n- Training time: around 13 hours (Kaggle kernel)\n\n**performance of 5-fold se\\_resnext101**\n- cv each: 0.627 | 0.629 | 0.627 | 0.628 | 0.630\n- lb each: 0.630 | ---\n- cv(concat): 0.628\n- lb(mean): **0.652**\n\n**se\\_resnext50**\n- base lr: 1.4e-4\n- batch size: 52\n- Training time: around 7 hours\n\n**performance of 5-fold se\\_resnext50**\n- cv each: 0.621 | 0.623 | 0.620 | 0.618 | 0.622\n- lb each: 0.621 | ---\n- cv(concat): 0.621 \n- lb(mean): **0.641**\n\n**augmentations**\n```\nfrom torchvision import transforms as T\n\ndef train_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n        RandomErasing(probability=0.3, sh=0.3),\n    ])\n\ndef test_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```\n\n### Student models\n\n- 1x se\\_resnext101\n- 1x inceptionresnetv2\n\nI trained a new 6-fold se\\_resnext101 using outputs of the teacher model on training images. I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets. Test images were remained untouched. \n\nTraining parameters remained almost the same from teacher models. This 6-fold model scored **cv 0.668** and **lb 0.662** by itself. \n\n**performance of 6-fold se\\_resnext101 (student)**\n- cv each: 0.669 | 0.669 | 0.668 | 0.664 | 0.665 | 0.668\n- lb each: 0.649 | 0.648 | ---\n- cv(concat): 0.668\n- lb(mean): **0.662**\n\nI also trained inceptionresnetv2 and averaged with se\\_resnext101 and scored **lb 0.664**. \n\n\n### Inference\n\nThe inference part is very simple. I just averaged outputs of the student models. # of TTA is 7. \n\n\n### Possible Improvements\n\nI have to admit there are lots of space for improvements. \n- Train more teacher models / student models such as senet152, pnasnet for ensembling. My solution definitely lacks variants for better ensembling. \n- Treat cultures and tags separately for training, thresholding as they have different characteristics. \n\n\nI'd like to thank Lopuhin for sharing his great kernel where I borrowed some ideas/implementations such as making folds, binarizing outputs. https://www.kaggle.com/lopuhin/imet-2019-submission\n\nI also would like to thank Bac Nguyen and his implementation of FbetaLoss. https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss\n\nThank you kaggle team and Zhang for hosting such a great competition. \n\nThanks for reading.",
      "votes": null
    },
    {
      "id": "547120",
      "postDate": "06/07/2019 09:54:00",
      "content": "<p>Thanks for sharing, was waiting for your solution.</p>",
      "rawMarkdown": "Thanks for sharing, was waiting for your solution.",
      "votes": null
    },
    {
      "id": "547133",
      "postDate": "06/07/2019 10:14:28",
      "content": "<p>Thank you! I'm happy to share.</p>",
      "rawMarkdown": "Thank you! I'm happy to share.",
      "votes": null
    },
    {
      "id": "547148",
      "postDate": "06/07/2019 10:28:38",
      "content": "<p>Do I understand right that RandomResizedCropV2 was one of the most important configuration of your training process?</p>",
      "rawMarkdown": "Do I understand right that RandomResizedCropV2 was one of the most important configuration of your training process?",
      "votes": null
    },
    {
      "id": "547198",
      "postDate": "06/07/2019 11:47:30",
      "content": "<p>It's a small modification on torchvision's RandomResizedCrop. It does random crop instead of center crop in case of fallback. There are many cases of fallback because of high aspect ratio images on imet and an edge of the image is more important than center of the image in some cases. Actually not very sure how much this alone affect cv/lb compared to vanilla but should be better with this.</p>\n\n<p>```\nclass RandomResizedCropV2(T.RandomResizedCrop):</p>\n\n<pre><code>@staticmethod\ndef get_params(img, scale, ratio):\n\n    # ...\n\n    # fallback\n    w = min(img.size[0], img.size[1])\n    i = random.randint(0, img.size[1] - w)\n    j = random.randint(0, img.size[0] - w)\n\n    return i, j, w, w\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "It's a small modification on torchvision's RandomResizedCrop. It does random crop instead of center crop in case of fallback. There are many cases of fallback because of high aspect ratio images on imet and an edge of the image is more important than center of the image in some cases. Actually not very sure how much this alone affect cv/lb compared to vanilla but should be better with this.\n\n```\nclass RandomResizedCropV2(T.RandomResizedCrop):\n\n    @staticmethod\n    def get_params(img, scale, ratio):\n\n        # ...\n\n        # fallback\n        w = min(img.size[0], img.size[1])\n        i = random.randint(0, img.size[1] - w)\n        j = random.randint(0, img.size[0] - w)\n\n        return i, j, w, w\n\n```",
      "votes": null
    },
    {
      "id": "547440",
      "postDate": "06/07/2019 17:49:31",
      "content": "<p>Nice solution, our experiment on distillation is not very successful, I think the key point is student model cannot trust too much on teacher's model</p>",
      "rawMarkdown": "Nice solution, our experiment on distillation is not very successful, I think the key point is student model cannot trust too much on teacher's model",
      "votes": null
    },
    {
      "id": "547562",
      "postDate": "06/07/2019 21:39:41",
      "content": "<p>Congrats and thanks for sharing!! Your method inspired me a lot with limited computation resources!!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!! Your method inspired me a lot with limited computation resources!!",
      "votes": null
    },
    {
      "id": "547807",
      "postDate": "06/08/2019 10:03:50",
      "content": "<p>&gt; I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets.</p>\n\n<p>Thanks, Appian. Could you explain more about your student targets? I cannot understand this sentence.</p>",
      "rawMarkdown": "&gt; I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets.\n\nThanks, Appian. Could you explain more about your student targets? I cannot understand this sentence.",
      "votes": null
    },
    {
      "id": "547816",
      "postDate": "06/08/2019 10:26:36",
      "content": "<p>Sure thing. </p>\n\n<p><code>\n0.00 0.00 1.00 0.00 (original hard target)\n0.10 0.20 0.70 0.70 (output from parent model)\n0.07 0.14 1.00 0.49 (new targets)\n</code></p>\n\n<p>This is what I wanted to tell. </p>",
      "rawMarkdown": "Sure thing. \n\n```\n0.00 0.00 1.00 0.00 (original hard target)\n0.10 0.20 0.70 0.70 (output from parent model)\n0.07 0.14 1.00 0.49 (new targets)\n```\n\nThis is what I wanted to tell.",
      "votes": null
    },
    {
      "id": "547817",
      "postDate": "06/08/2019 10:28:01",
      "content": "<p>I tried different numbers and 0.7 was the best. 0.9 scored same lb but cv was too good and cv/lb gap was bigger. I got worse lb with 0.5.</p>",
      "rawMarkdown": "I tried different numbers and 0.7 was the best. 0.9 scored same lb but cv was too good and cv/lb gap was bigger. I got worse lb with 0.5.",
      "votes": null
    },
    {
      "id": "547824",
      "postDate": "06/08/2019 10:53:44",
      "content": "<p>Get it, thanks! </p>",
      "rawMarkdown": "Get it, thanks!",
      "votes": null
    },
    {
      "id": "547948",
      "postDate": "06/08/2019 14:55:23",
      "content": "<p>BTW are you still using the original validation data to vaid? Or use the same technique to generate new validation?</p>",
      "rawMarkdown": "BTW are you still using the original validation data to vaid? Or use the same technique to generate new validation?",
      "votes": null
    },
    {
      "id": "548431",
      "postDate": "06/09/2019 11:30:00",
      "content": "<p>Do you mean did I use new targets to evaluate the performace on validation data? If so the answer is no. THe new targets are only used for calculating the loss.</p>",
      "rawMarkdown": "Do you mean did I use new targets to evaluate the performace on validation data? If so the answer is no. THe new targets are only used for calculating the loss.",
      "votes": null
    },
    {
      "id": "548983",
      "postDate": "06/10/2019 06:58:53",
      "content": "<p>Congrats and thanks for sharing! Would you mind telling me the preprocess method you use. I find that the image size of input is 320 as the minimum size of train image is 300.</p>",
      "rawMarkdown": "Congrats and thanks for sharing! Would you mind telling me the preprocess method you use. I find that the image size of input is 320 as the minimum size of train image is 300.",
      "votes": null
    },
    {
      "id": "549279",
      "postDate": "06/10/2019 13:54:17",
      "content": "<p>nice observation of connection between distillation and lb metric</p>",
      "rawMarkdown": "nice observation of connection between distillation and lb metric",
      "votes": null
    },
    {
      "id": "549958",
      "postDate": "06/11/2019 07:10:09",
      "content": "<p>May I ask your loss function of teacher is like this?\n`class FbetaLoss(nn.Module):</p>\n\n<pre><code>def __init__(self, beta=1):\n    super(FbetaLoss, self).__init__()\n    self.small_value = 1e-6\n    self.beta = beta\n\ndef forward(self, logits, labels):\n    beta = self.beta\n    batch_size = logits.size()[0]\n    p = torch.sigmoid(logits)\n    l = labels\n    num_pos = torch.sum(p, 1) + self.small_value\n    num_pos_hat = torch.sum(l, 1) + self.small_value\n    tp = torch.sum(l * p, 1)\n    precise = tp / num_pos\n    recall = tp / num_pos_hat\n    fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + self.small_value)\n    loss = fs.sum() / batch_size\n    return 1 - loss`\n</code></pre>\n\n<p>`class CombineLoss_bce(nn.Module):</p>\n\n<pre><code>def __init__(self):\n    super(CombineLoss_bce, self).__init__()\n    self.fbeta_loss = FbetaLoss(beta=2)\n    self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n\ndef forward(self, logits, labels):\n    loss_beta = self.fbeta_loss(logits, labels)\n    loss_bce = self.bce_loss(logits, labels)\n    return 0.5 * loss_beta + 0.5 * loss_bce\n</code></pre>\n\n<p>`</p>",
      "rawMarkdown": "May I ask your loss function of teacher is like this?\n`class FbetaLoss(nn.Module):\n\n    def __init__(self, beta=1):\n        super(FbetaLoss, self).__init__()\n        self.small_value = 1e-6\n        self.beta = beta\n\n    def forward(self, logits, labels):\n        beta = self.beta\n        batch_size = logits.size()[0]\n        p = torch.sigmoid(logits)\n        l = labels\n        num_pos = torch.sum(p, 1) + self.small_value\n        num_pos_hat = torch.sum(l, 1) + self.small_value\n        tp = torch.sum(l * p, 1)\n        precise = tp / num_pos\n        recall = tp / num_pos_hat\n        fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + self.small_value)\n        loss = fs.sum() / batch_size\n        return 1 - loss`\n`class CombineLoss_bce(nn.Module):\n\n    def __init__(self):\n        super(CombineLoss_bce, self).__init__()\n        self.fbeta_loss = FbetaLoss(beta=2)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n        \n    def forward(self, logits, labels):\n        loss_beta = self.fbeta_loss(logits, labels)\n        loss_bce = self.bce_loss(logits, labels)\n        return 0.5 * loss_beta + 0.5 * loss_bce\n`",
      "votes": null
    },
    {
      "id": "550326",
      "postDate": "06/11/2019 14:01:45",
      "content": "<p>```\nclass FBetaBCE(torch.nn.Module):\n    def <strong>init</strong>(self, fbeta_weight=1.0):\n        super().<strong>init</strong>()\n        self.fbeta_loss = FbetaLoss(beta=2)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n        self.fbeta_weight = fbeta_weight</p>\n\n<pre><code>def forward(self, logits, labels):\n    loss_beta = self.fbeta_loss(logits, labels)\n    loss_bce = self.bce_loss(logits, labels)\n    loss_bce = loss_bce.sum() / loss_bce.shape[0]\n    return loss_beta*self.fbeta_weight + loss_bce\n</code></pre>\n\n<p>```</p>\n\n<p>Almost the same I guess.\nImplementation of fbeta_loss is from <a href=\"https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss\">https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss</a></p>",
      "rawMarkdown": "```\nclass FBetaBCE(torch.nn.Module):\n    def __init__(self, fbeta_weight=1.0):\n        super().__init__()\n        self.fbeta_loss = FbetaLoss(beta=2)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n        self.fbeta_weight = fbeta_weight\n        \n    def forward(self, logits, labels):\n        loss_beta = self.fbeta_loss(logits, labels)\n        loss_bce = self.bce_loss(logits, labels)\n        loss_bce = loss_bce.sum() / loss_bce.shape[0]\n        return loss_beta*self.fbeta_weight + loss_bce\n```\n\nAlmost the same I guess.\nImplementation of fbeta_loss is from https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss",
      "votes": null
    },
    {
      "id": "550329",
      "postDate": "06/11/2019 14:03:03",
      "content": "<p>Thank you!\nThere is no preprocess except augmentations I mentioned. Aspect ratio of randomresizedcrop varies between 4/5 to 5/4 and this is the reason why I used 320px not 300px.</p>",
      "rawMarkdown": "Thank you!\nThere is no preprocess except augmentations I mentioned. Aspect ratio of randomresizedcrop varies between 4/5 to 5/4 and this is the reason why I used 320px not 300px.",
      "votes": null
    },
    {
      "id": "550469",
      "postDate": "06/11/2019 16:49:07",
      "content": "<p>Thanks for sharing your insights <a href=\"/appian\">@appian</a> . If you are able to and willing to, would you share an example script of how you did the distillation step? I think that's a really useful idea for kernel competitions like this. Congrats on your medal!</p>",
      "rawMarkdown": "Thanks for sharing your insights @appian . If you are able to and willing to, would you share an example script of how you did the distillation step? I think that's a really useful idea for kernel competitions like this. Congrats on your medal!",
      "votes": null
    },
    {
      "id": "550728",
      "postDate": "06/12/2019 01:20:50",
      "content": "<p>I see, Thank you!</p>",
      "rawMarkdown": "I see, Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 547120,
      "author_name": "harshthaker",
      "author_url": "",
      "post_date": "06/07/2019 09:54:00",
      "content": "<p>Thanks for sharing, was waiting for your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 547133,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/07/2019 10:14:28",
          "content": "<p>Thank you! I'm happy to share.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 547148,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "06/07/2019 10:28:38",
      "content": "<p>Do I understand right that RandomResizedCropV2 was one of the most important configuration of your training process?</p>",
      "votes": null,
      "replies": [
        {
          "id": 547198,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/07/2019 11:47:30",
          "content": "<p>It's a small modification on torchvision's RandomResizedCrop. It does random crop instead of center crop in case of fallback. There are many cases of fallback because of high aspect ratio images on imet and an edge of the image is more important than center of the image in some cases. Actually not very sure how much this alone affect cv/lb compared to vanilla but should be better with this.</p>\n\n<p>```\nclass RandomResizedCropV2(T.RandomResizedCrop):</p>\n\n<pre><code>@staticmethod\ndef get_params(img, scale, ratio):\n\n    # ...\n\n    # fallback\n    w = min(img.size[0], img.size[1])\n    i = random.randint(0, img.size[1] - w)\n    j = random.randint(0, img.size[0] - w)\n\n    return i, j, w, w\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 547440,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "06/07/2019 17:49:31",
      "content": "<p>Nice solution, our experiment on distillation is not very successful, I think the key point is student model cannot trust too much on teacher's model</p>",
      "votes": null,
      "replies": [
        {
          "id": 547817,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/08/2019 10:28:01",
          "content": "<p>I tried different numbers and 0.7 was the best. 0.9 scored same lb but cv was too good and cv/lb gap was bigger. I got worse lb with 0.5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547948,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "06/08/2019 14:55:23",
          "content": "<p>BTW are you still using the original validation data to vaid? Or use the same technique to generate new validation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 548431,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/09/2019 11:30:00",
          "content": "<p>Do you mean did I use new targets to evaluate the performace on validation data? If so the answer is no. THe new targets are only used for calculating the loss.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 547562,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "06/07/2019 21:39:41",
      "content": "<p>Congrats and thanks for sharing!! Your method inspired me a lot with limited computation resources!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547807,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "06/08/2019 10:03:50",
      "content": "<p>&gt; I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets.</p>\n\n<p>Thanks, Appian. Could you explain more about your student targets? I cannot understand this sentence.</p>",
      "votes": null,
      "replies": [
        {
          "id": 547816,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/08/2019 10:26:36",
          "content": "<p>Sure thing. </p>\n\n<p><code>\n0.00 0.00 1.00 0.00 (original hard target)\n0.10 0.20 0.70 0.70 (output from parent model)\n0.07 0.14 1.00 0.49 (new targets)\n</code></p>\n\n<p>This is what I wanted to tell. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 547824,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "06/08/2019 10:53:44",
          "content": "<p>Get it, thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 548983,
      "author_name": "lanjunyelan",
      "author_url": "",
      "post_date": "06/10/2019 06:58:53",
      "content": "<p>Congrats and thanks for sharing! Would you mind telling me the preprocess method you use. I find that the image size of input is 320 as the minimum size of train image is 300.</p>",
      "votes": null,
      "replies": [
        {
          "id": 550329,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/11/2019 14:03:03",
          "content": "<p>Thank you!\nThere is no preprocess except augmentations I mentioned. Aspect ratio of randomresizedcrop varies between 4/5 to 5/4 and this is the reason why I used 320px not 300px.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550728,
          "author_name": "lanjunyelan",
          "author_url": "",
          "post_date": "06/12/2019 01:20:50",
          "content": "<p>I see, Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 549279,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "06/10/2019 13:54:17",
      "content": "<p>nice observation of connection between distillation and lb metric</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 549958,
      "author_name": "saladjay",
      "author_url": "",
      "post_date": "06/11/2019 07:10:09",
      "content": "<p>May I ask your loss function of teacher is like this?\n`class FbetaLoss(nn.Module):</p>\n\n<pre><code>def __init__(self, beta=1):\n    super(FbetaLoss, self).__init__()\n    self.small_value = 1e-6\n    self.beta = beta\n\ndef forward(self, logits, labels):\n    beta = self.beta\n    batch_size = logits.size()[0]\n    p = torch.sigmoid(logits)\n    l = labels\n    num_pos = torch.sum(p, 1) + self.small_value\n    num_pos_hat = torch.sum(l, 1) + self.small_value\n    tp = torch.sum(l * p, 1)\n    precise = tp / num_pos\n    recall = tp / num_pos_hat\n    fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + self.small_value)\n    loss = fs.sum() / batch_size\n    return 1 - loss`\n</code></pre>\n\n<p>`class CombineLoss_bce(nn.Module):</p>\n\n<pre><code>def __init__(self):\n    super(CombineLoss_bce, self).__init__()\n    self.fbeta_loss = FbetaLoss(beta=2)\n    self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n\ndef forward(self, logits, labels):\n    loss_beta = self.fbeta_loss(logits, labels)\n    loss_bce = self.bce_loss(logits, labels)\n    return 0.5 * loss_beta + 0.5 * loss_bce\n</code></pre>\n\n<p>`</p>",
      "votes": null,
      "replies": [
        {
          "id": 550326,
          "author_name": "appian",
          "author_url": "",
          "post_date": "06/11/2019 14:01:45",
          "content": "<p>```\nclass FBetaBCE(torch.nn.Module):\n    def <strong>init</strong>(self, fbeta_weight=1.0):\n        super().<strong>init</strong>()\n        self.fbeta_loss = FbetaLoss(beta=2)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n        self.fbeta_weight = fbeta_weight</p>\n\n<pre><code>def forward(self, logits, labels):\n    loss_beta = self.fbeta_loss(logits, labels)\n    loss_bce = self.bce_loss(logits, labels)\n    loss_bce = loss_bce.sum() / loss_bce.shape[0]\n    return loss_beta*self.fbeta_weight + loss_bce\n</code></pre>\n\n<p>```</p>\n\n<p>Almost the same I guess.\nImplementation of fbeta_loss is from <a href=\"https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss\">https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550469,
      "author_name": "sairam6087",
      "author_url": "",
      "post_date": "06/11/2019 16:49:07",
      "content": "<p>Thanks for sharing your insights <a href=\"/appian\">@appian</a> . If you are able to and willing to, would you share an example script of how you did the distillation step? I think that's a really useful idea for kernel competitions like this. Congrats on your medal!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "547098": "Hi everyone. This is my solution overview which scored 0.664 on public LB.\n\n\n### Competition's challenge\n\n1. 9 hours kernels limit for inference. \n2. Lack of proper labels. For example if an image has **culture::american**, there also should be **culture::american or british** but in many cases these labels are lacked. \n\n\n### Approach\n\n1. I used **knowledge distillation** to compress the knowledge in an ensemble into student models so that I could use 9 hours more efficiently.\n2. Knowledge distillation also attends to the lack of proper labels by providing soft targets to training images. Because competition metric favors recall over precision, this should improve lb.\n\n\n### Hardware\n\nI used kaggle kernels to train almost all the models. \n\n\n### Model training\n\nThe model training process can be split into 2 parts. \n- Train **teacher models**\n- Train **student models** using outputs of the teacher\n\n\n### Teacher models\n\n- 2x se\\_resnext101\n- 2x se\\_resnext50\n\nI trained 4 models and averaged their outputs. I shared the performance of them during the competition and you can check that discussion here. https://www.kaggle.com/c/imet-2019-fgvc6/discussion/92159\n\nAveraged output scored **cv 0.6467** and **lb 0.656**. This averaged output(cv 0.6467) on train images were used to train the student models next. \n\nI did not treat tags and cultures separately.\n\n**training method**\nI trained only a dense layer for first 2 epochs because their weights are fresh. I got this idea from the kernel Lopuhin kindly shared and result is better with this.\n- 1st epoch (1/10 of baselr, dense layer only)\n- 2nd epoch (5/10 of baselr, dense layer only)\n- 3rd epoch (1/10 of baselr, all layers)\n- 4th epoch (5/10 of baselr, all layers)\n\n**se\\_resnext101**\n- image size: 320\n- optimizer: Adam\n- base lr: 1.2e-4\n- batch size: 36\n- scheduler: ReduceLrOnPlateau or StepLR\n- best epoch: 12 to 15\n- loss: BCEWithLogitsLoss + FBetaLoss\n- augmentations: RandomResizedCrop, Horizontal Flip, Random Erasing\n- TTA: RandomResizedCrop, Horizontal Flip (10 times)\n- Training time: around 13 hours (Kaggle kernel)\n\n**performance of 5-fold se\\_resnext101**\n- cv each: 0.627 | 0.629 | 0.627 | 0.628 | 0.630\n- lb each: 0.630 | ---\n- cv(concat): 0.628\n- lb(mean): **0.652**\n\n**se\\_resnext50**\n- base lr: 1.4e-4\n- batch size: 52\n- Training time: around 7 hours\n\n**performance of 5-fold se\\_resnext50**\n- cv each: 0.621 | 0.623 | 0.620 | 0.618 | 0.622\n- lb each: 0.621 | ---\n- cv(concat): 0.621 \n- lb(mean): **0.641**\n\n**augmentations**\n```\nfrom torchvision import transforms as T\n\ndef train_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n        RandomErasing(probability=0.3, sh=0.3),\n    ])\n\ndef test_transform(size):\n    return T.Compose([\n        RandomResizedCropV2(size, scale=(0.7, 1.0), ratio=(4/5, 5/4)),\n        T.RandomHorizontalFlip(),\n        T.ToTensor(),\n        T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```\n\n### Student models\n\n- 1x se\\_resnext101\n- 1x inceptionresnetv2\n\nI trained a new 6-fold se\\_resnext101 using outputs of the teacher model on training images. I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets. Test images were remained untouched. \n\nTraining parameters remained almost the same from teacher models. This 6-fold model scored **cv 0.668** and **lb 0.662** by itself. \n\n**performance of 6-fold se\\_resnext101 (student)**\n- cv each: 0.669 | 0.669 | 0.668 | 0.664 | 0.665 | 0.668\n- lb each: 0.649 | 0.648 | ---\n- cv(concat): 0.668\n- lb(mean): **0.662**\n\nI also trained inceptionresnetv2 and averaged with se\\_resnext101 and scored **lb 0.664**. \n\n\n### Inference\n\nThe inference part is very simple. I just averaged outputs of the student models. # of TTA is 7. \n\n\n### Possible Improvements\n\nI have to admit there are lots of space for improvements. \n- Train more teacher models / student models such as senet152, pnasnet for ensembling. My solution definitely lacks variants for better ensembling. \n- Treat cultures and tags separately for training, thresholding as they have different characteristics. \n\n\nI'd like to thank Lopuhin for sharing his great kernel where I borrowed some ideas/implementations such as making folds, binarizing outputs. https://www.kaggle.com/lopuhin/imet-2019-submission\n\nI also would like to thank Bac Nguyen and his implementation of FbetaLoss. https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss\n\nThank you kaggle team and Zhang for hosting such a great competition. \n\nThanks for reading.",
    "547120": "Thanks for sharing, was waiting for your solution.",
    "547133": "Thank you! I'm happy to share.",
    "547148": "Do I understand right that RandomResizedCropV2 was one of the most important configuration of your training process?",
    "547198": "It's a small modification on torchvision's RandomResizedCrop. It does random crop instead of center crop in case of fallback. There are many cases of fallback because of high aspect ratio images on imet and an edge of the image is more important than center of the image in some cases. Actually not very sure how much this alone affect cv/lb compared to vanilla but should be better with this.\n\n```\nclass RandomResizedCropV2(T.RandomResizedCrop):\n\n    @staticmethod\n    def get_params(img, scale, ratio):\n\n        # ...\n\n        # fallback\n        w = min(img.size[0], img.size[1])\n        i = random.randint(0, img.size[1] - w)\n        j = random.randint(0, img.size[0] - w)\n\n        return i, j, w, w\n\n```",
    "547440": "Nice solution, our experiment on distillation is not very successful, I think the key point is student model cannot trust too much on teacher's model",
    "547562": "Congrats and thanks for sharing!! Your method inspired me a lot with limited computation resources!!",
    "547807": "&gt; I multiplied it by 0.7 to prevent overtrusting his teacher. I took np.maximum of it and original hard targets to construct new targets.\n\nThanks, Appian. Could you explain more about your student targets? I cannot understand this sentence.",
    "547816": "Sure thing. \n\n```\n0.00 0.00 1.00 0.00 (original hard target)\n0.10 0.20 0.70 0.70 (output from parent model)\n0.07 0.14 1.00 0.49 (new targets)\n```\n\nThis is what I wanted to tell.",
    "547817": "I tried different numbers and 0.7 was the best. 0.9 scored same lb but cv was too good and cv/lb gap was bigger. I got worse lb with 0.5.",
    "547824": "Get it, thanks!",
    "547948": "BTW are you still using the original validation data to vaid? Or use the same technique to generate new validation?",
    "548431": "Do you mean did I use new targets to evaluate the performace on validation data? If so the answer is no. THe new targets are only used for calculating the loss.",
    "548983": "Congrats and thanks for sharing! Would you mind telling me the preprocess method you use. I find that the image size of input is 320 as the minimum size of train image is 300.",
    "549279": "nice observation of connection between distillation and lb metric",
    "549958": "May I ask your loss function of teacher is like this?\n`class FbetaLoss(nn.Module):\n\n    def __init__(self, beta=1):\n        super(FbetaLoss, self).__init__()\n        self.small_value = 1e-6\n        self.beta = beta\n\n    def forward(self, logits, labels):\n        beta = self.beta\n        batch_size = logits.size()[0]\n        p = torch.sigmoid(logits)\n        l = labels\n        num_pos = torch.sum(p, 1) + self.small_value\n        num_pos_hat = torch.sum(l, 1) + self.small_value\n        tp = torch.sum(l * p, 1)\n        precise = tp / num_pos\n        recall = tp / num_pos_hat\n        fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + self.small_value)\n        loss = fs.sum() / batch_size\n        return 1 - loss`\n`class CombineLoss_bce(nn.Module):\n\n    def __init__(self):\n        super(CombineLoss_bce, self).__init__()\n        self.fbeta_loss = FbetaLoss(beta=2)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n        \n    def forward(self, logits, labels):\n        loss_beta = self.fbeta_loss(logits, labels)\n        loss_bce = self.bce_loss(logits, labels)\n        return 0.5 * loss_beta + 0.5 * loss_bce\n`",
    "550326": "```\nclass FBetaBCE(torch.nn.Module):\n    def __init__(self, fbeta_weight=1.0):\n        super().__init__()\n        self.fbeta_loss = FbetaLoss(beta=2)\n        self.bce_loss = nn.BCEWithLogitsLoss(reduction='none')\n        self.fbeta_weight = fbeta_weight\n        \n    def forward(self, logits, labels):\n        loss_beta = self.fbeta_loss(logits, labels)\n        loss_bce = self.bce_loss(logits, labels)\n        loss_bce = loss_bce.sum() / loss_bce.shape[0]\n        return loss_beta*self.fbeta_weight + loss_bce\n```\n\nAlmost the same I guess.\nImplementation of fbeta_loss is from https://www.kaggle.com/backaggle/imet-fastai-starter-focal-and-fbeta-loss",
    "550329": "Thank you!\nThere is no preprocess except augmentations I mentioned. Aspect ratio of randomresizedcrop varies between 4/5 to 5/4 and this is the reason why I used 320px not 300px.",
    "550469": "Thanks for sharing your insights @appian . If you are able to and willing to, would you share an example script of how you did the distillation step? I think that's a really useful idea for kernel competitions like this. Congrats on your medal!",
    "550728": "I see, Thank you!"
  },
  "source": "meta"
}