{
  "id": 82364,
  "title": "57th place solution, SoftTripletLoss, 256x512 image size, fastai v1",
  "url": "/competitions/humpback-whale-identification/writeups/miguel-pinto-57th-place-solution-softtripletloss-2",
  "author_name": "",
  "post_date": "2019-03-01T09:35:22.363Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When I arrived on this competition I had no idea how to solve this kind of problem. It took me a while just to find papers about this problem and understand the intuition behind the concepts involved like episodes, n-way k-shot, metric learning, few-shot learning, meta learning and so on. I implemented all the framework in fastai v1.</p>\n\n<h3>Loss Function</h3>\n\n<p>I have tried several approaches like central loss [1] and prototypical networks [2] but I then found on the discussions this two papers [3], [4] where a variation of triplet loss is used where the hard margin is replaced by a soft margin using the softplus function (that's why I'm referring to it as Soft Triplet Loss). They also use a Batch Hard strategy (BH) in which for each anchor image only the hardest positive and hardest negatives in the mini-batch are used in the loss. This is my implementation of Soft Triplet Loss with Batch Hard and L2 regularization (as suggested by @Iafoss [5]):</p>\n\n<pre><code>class SoftTripletLoss(nn.Module):\n    def __init__(self, fsc, wd=1e-4):\n        super().__init__()\n        self.k_shot = fsc.k_shot\n        self.new_class_number = fsc.new_class_number\n        self.wd = wd\n\n    def forward(self, x, y):\n        # x (64, 128)\n        self.n_way = x.size()[0]//self.k_shot\n        emb_sz = x.size()[-1] # (128)\n        x = x.view(-1, self.k_shot, emb_sz) # (16, 4, 128)\n        L = 0; EPS = 1e-6\n        for i in range(self.n_way-self.new_class_number):\n            for j in range(self.k_shot):\n                I = torch.zeros(self.n_way).long()\n                I[i] = 1\n                J = torch.zeros(self.k_shot).long()\n                J[j] = 1\n                xa = x[I==1, J==1, :].view(1, -1) # (1, 128)\n                xp = x[I==1, J==0, :] # (3, 128)\n                xn = x[I==0, :, :].view(-1, emb_sz) # (15, 4, 128) -&amp;gt; (60, 128)\n                Dp = F.relu((xa-xp).pow_(2).sum(1)+EPS).sqrt_() # (3)\n                Dn = F.relu((xa-xn).pow_(2).sum(1)+EPS).sqrt_() # (60)\n                L += F.softplus(Dp.max(0)[0] - Dn.min(0)[0]) # (1)\n                L += self.wd*((Dp**2).mean() + (Dn**2).mean()) # Regularization\n        return L\n</code></pre>\n\n<h3>Hard samples mining</h3>\n\n<p>Batch hard strategy allowed for a good improvement but it was not enough alone. The next main step was implementing a technique for mining hard samples, hereafter referred as Sample Hard (SH). At the end of each training epoch I compute the distance matrix between all train samples, then to build the mini-batches for the next epoch I follow the steps (I'm using 10-way, 4-shot episodes):</p>\n\n<ol>\n<li>Select an image at random [A1]</li>\n<li>Select the 3 hardest images from the same class (hard positives, largest distances) [A1, A2, A3, A4]</li>\n<li>Select the one hardest image from a different class (hard negatives, closest distance) [A1, A2, A3, A4, B1]</li>\n<li>Repeat from 2. until the mini-batch is constructed [A1, A2, A3, A4, B1, B2, B3, B4, C1, C2, C3, C4, ...] </li>\n<li>For the last 4 images in each mini-batch I select 4 random <em>new whales</em>.</li>\n</ol>\n\n<p>So the distance matrix and mini-batches for the next epoch are only computed at the end of each epoch. This SH strategy is then used together with BH. </p>\n\n<h3>Model</h3>\n\n<p>The model I used is a <strong>Densenet121</strong> with the following head:</p>\n\n<pre><code>class Head(nn.Module):\n    def __init__(self, in_channels=1024, emb_sz=128):\n        super().__init__()\n        self.flat = nn.Sequential(\n            AdaptiveConcatPool2d(1))\n        self.flatten = Flatten()\n        self.bn0 = nn.BatchNorm1d(4*in_channels)\n        self.lin0 = nn.Linear(4*in_channels, in_channels)\n        self.relu = nn.ReLU(inplace=True)\n        self.bn1 = nn.BatchNorm1d(in_channels)\n        self.lin1 = nn.Linear(in_channels, emb_sz)\n\n    def forward(self, x):\n        cut = x.size()[-1]//2\n        x0 = self.flat(x[...,:cut])\n        x1 = self.flat(x[...,cut:])\n        x = torch.cat((self.flatten(x0), self.flatten(x1)), dim=1)\n        x = self.relu(self.lin0(self.bn0(x)))\n        return self.lin1(self.bn1(x))\n</code></pre>\n\n<p>Since I'm using images with ratio 1:2 and there are some horizontal symmetry I divide the images in two (left and right parts) (x0 and x1 in the code), I apply the pooling as usual (using AdaptiveConcatPool2d from fastai) and I finally concatenate the two. The activation maps shown by Heng [6] in another discussion topic show that in some images there are two modes in the activation map, one for the left and other for the right part, that is the intuition why this may help. I didn't check however if the increase in performance is just due to the increase in parameters.</p>\n\n<h3>Image Augmentations</h3>\n\n<p>I cropped the images with bounding boxes and applied the following augmentations (fastai):</p>\n\n<pre><code>from torchvision.transforms import ColorJitter, ToPILImage, ToTensor\nfrom fastai.vision.transform import *\ndef _colorjitter(x, brightness=0, contrast=0, saturation=0, hue=0):\n    topill = ToPILImage()\n    totensor = ToTensor()\n    xmin, xmax = x.min(), x.max()\n    x = (x-xmin)/(xmax-xmin)\n    cj = ColorJitter(brightness, contrast, saturation, hue)\n    x = topill(x)\n    x = cj(x)\n    x = totensor(x)\n    x = x*(xmax-xmin) + xmin\n    return x\ncolorjitter = TfmLighting(_colorjitter)\n\ndef _cutout(x, n_holes:uniform_int=1, length:uniform_int=40):\n    \"Cut out `n_holes` number of square holes of size `length` in image at random locations.\"\n    h,w = x.shape[1:]\n    for n in range(n_holes):\n        h_y = np.random.randint(0, h)\n        h_x = np.random.randint(0, w)\n        y1 = int(np.clip(h_y - length / 2, 0, h))\n        y2 = int(np.clip(h_y + length / 2, 0, h))\n        x1 = int(np.clip(h_x - length / 2, 0, w))\n        x2 = int(np.clip(h_x + length / 2, 0, w))\n        x[:, y1:y2, x1:x2] = 0\n    return x\n\ncutout = TfmPixel(_cutout, order=20)\n\n\ntfms = get_transforms(do_flip=False, \n                 xtra_tfms=[colorjitter(saturation=1.1, hue=0.05), \n                 cutout(n_holes=(1, max(3, int(3*rcutout))),\n                 length=(10*rcutout, 40*rcutout), p=0.5)])\n</code></pre>\n\n<h3>Training</h3>\n\n<p>I train only the Head part of the model with BH for a few epochs (Densenet121 is using imagenet weights) and then I unfreeze the model and train for longer with SH + BH (the first epoch only BH however). Both before and after unfreezing I use one cycle policy (<a href=\"https://sgugger.github.io/the-1cycle-policy.html\">https://sgugger.github.io/the-1cycle-policy.html</a>) implemented in fastai v1. Just one cycle before unfreezing and one cycle after. I found this worked better than multiple cycles in my experiments.</p>\n\n<h3>Best single model + TTA</h3>\n\n<ul>\n<li>Image size; public LB;  approximate train time (1 nvidia 1080)</li>\n<li>64x128 – 0.788 - 3.5h</li>\n<li>128x256 – 0.867 - 6h</li>\n<li>256x512 – 0.904 - 15h</li>\n</ul>\n\n<p>Final score is an ensemble of 13 submissions with public LB &gt; 0.860, reaching the final score of ~ <strong>0.93</strong>.</p>\n\n<h3>References</h3>\n\n<p>[1] <a href=\"https://ydwen.github.io/papers/WenECCV16.pdf\">https://ydwen.github.io/papers/WenECCV16.pdf</a></p>\n\n<p>[2] <a href=\"https://arxiv.org/pdf/1703.05175.pdf\">https://arxiv.org/pdf/1703.05175.pdf</a></p>\n\n<p>[3] <a href=\"https://arxiv.org/pdf/1703.07737.pdf\">https://arxiv.org/pdf/1703.07737.pdf</a></p>\n\n<p>[4] <a href=\"https://arxiv.org/pdf/1901.03662.pdf\">https://arxiv.org/pdf/1901.03662.pdf</a></p>\n\n<p>[5] <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79086\">https://www.kaggle.com/c/humpback-whale-identification/discussion/79086</a></p>\n\n<p>[6] <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79524\">https://www.kaggle.com/c/humpback-whale-identification/discussion/79524</a></p>",
  "messages": [
    {
      "id": "481034",
      "postDate": "03/01/2019 01:40:59",
      "content": "<p>When I arrived on this competition I had no idea how to solve this kind of problem. It took me a while just to find papers about this problem and understand the intuition behind the concepts involved like episodes, n-way k-shot, metric learning, few-shot learning, meta learning and so on. I implemented all the framework in fastai v1.</p>\n\n<h3>Loss Function</h3>\n\n<p>I have tried several approaches like central loss [1] and prototypical networks [2] but I then found on the discussions this two papers [3], [4] where a variation of triplet loss is used where the hard margin is replaced by a soft margin using the softplus function (that's why I'm referring to it as Soft Triplet Loss). They also use a Batch Hard strategy (BH) in which for each anchor image only the hardest positive and hardest negatives in the mini-batch are used in the loss. This is my implementation of Soft Triplet Loss with Batch Hard and L2 regularization (as suggested by @Iafoss [5]):</p>\n\n<pre><code>class SoftTripletLoss(nn.Module):\n    def __init__(self, fsc, wd=1e-4):\n        super().__init__()\n        self.k_shot = fsc.k_shot\n        self.new_class_number = fsc.new_class_number\n        self.wd = wd\n\n    def forward(self, x, y):\n        # x (64, 128)\n        self.n_way = x.size()[0]//self.k_shot\n        emb_sz = x.size()[-1] # (128)\n        x = x.view(-1, self.k_shot, emb_sz) # (16, 4, 128)\n        L = 0; EPS = 1e-6\n        for i in range(self.n_way-self.new_class_number):\n            for j in range(self.k_shot):\n                I = torch.zeros(self.n_way).long()\n                I[i] = 1\n                J = torch.zeros(self.k_shot).long()\n                J[j] = 1\n                xa = x[I==1, J==1, :].view(1, -1) # (1, 128)\n                xp = x[I==1, J==0, :] # (3, 128)\n                xn = x[I==0, :, :].view(-1, emb_sz) # (15, 4, 128) -&amp;gt; (60, 128)\n                Dp = F.relu((xa-xp).pow_(2).sum(1)+EPS).sqrt_() # (3)\n                Dn = F.relu((xa-xn).pow_(2).sum(1)+EPS).sqrt_() # (60)\n                L += F.softplus(Dp.max(0)[0] - Dn.min(0)[0]) # (1)\n                L += self.wd*((Dp**2).mean() + (Dn**2).mean()) # Regularization\n        return L\n</code></pre>\n\n<h3>Hard samples mining</h3>\n\n<p>Batch hard strategy allowed for a good improvement but it was not enough alone. The next main step was implementing a technique for mining hard samples, hereafter referred as Sample Hard (SH). At the end of each training epoch I compute the distance matrix between all train samples, then to build the mini-batches for the next epoch I follow the steps (I'm using 10-way, 4-shot episodes):</p>\n\n<ol>\n<li>Select an image at random [A1]</li>\n<li>Select the 3 hardest images from the same class (hard positives, largest distances) [A1, A2, A3, A4]</li>\n<li>Select the one hardest image from a different class (hard negatives, closest distance) [A1, A2, A3, A4, B1]</li>\n<li>Repeat from 2. until the mini-batch is constructed [A1, A2, A3, A4, B1, B2, B3, B4, C1, C2, C3, C4, ...] </li>\n<li>For the last 4 images in each mini-batch I select 4 random <em>new whales</em>.</li>\n</ol>\n\n<p>So the distance matrix and mini-batches for the next epoch are only computed at the end of each epoch. This SH strategy is then used together with BH. </p>\n\n<h3>Model</h3>\n\n<p>The model I used is a <strong>Densenet121</strong> with the following head:</p>\n\n<pre><code>class Head(nn.Module):\n    def __init__(self, in_channels=1024, emb_sz=128):\n        super().__init__()\n        self.flat = nn.Sequential(\n            AdaptiveConcatPool2d(1))\n        self.flatten = Flatten()\n        self.bn0 = nn.BatchNorm1d(4*in_channels)\n        self.lin0 = nn.Linear(4*in_channels, in_channels)\n        self.relu = nn.ReLU(inplace=True)\n        self.bn1 = nn.BatchNorm1d(in_channels)\n        self.lin1 = nn.Linear(in_channels, emb_sz)\n\n    def forward(self, x):\n        cut = x.size()[-1]//2\n        x0 = self.flat(x[...,:cut])\n        x1 = self.flat(x[...,cut:])\n        x = torch.cat((self.flatten(x0), self.flatten(x1)), dim=1)\n        x = self.relu(self.lin0(self.bn0(x)))\n        return self.lin1(self.bn1(x))\n</code></pre>\n\n<p>Since I'm using images with ratio 1:2 and there are some horizontal symmetry I divide the images in two (left and right parts) (x0 and x1 in the code), I apply the pooling as usual (using AdaptiveConcatPool2d from fastai) and I finally concatenate the two. The activation maps shown by Heng [6] in another discussion topic show that in some images there are two modes in the activation map, one for the left and other for the right part, that is the intuition why this may help. I didn't check however if the increase in performance is just due to the increase in parameters.</p>\n\n<h3>Image Augmentations</h3>\n\n<p>I cropped the images with bounding boxes and applied the following augmentations (fastai):</p>\n\n<pre><code>from torchvision.transforms import ColorJitter, ToPILImage, ToTensor\nfrom fastai.vision.transform import *\ndef _colorjitter(x, brightness=0, contrast=0, saturation=0, hue=0):\n    topill = ToPILImage()\n    totensor = ToTensor()\n    xmin, xmax = x.min(), x.max()\n    x = (x-xmin)/(xmax-xmin)\n    cj = ColorJitter(brightness, contrast, saturation, hue)\n    x = topill(x)\n    x = cj(x)\n    x = totensor(x)\n    x = x*(xmax-xmin) + xmin\n    return x\ncolorjitter = TfmLighting(_colorjitter)\n\ndef _cutout(x, n_holes:uniform_int=1, length:uniform_int=40):\n    \"Cut out `n_holes` number of square holes of size `length` in image at random locations.\"\n    h,w = x.shape[1:]\n    for n in range(n_holes):\n        h_y = np.random.randint(0, h)\n        h_x = np.random.randint(0, w)\n        y1 = int(np.clip(h_y - length / 2, 0, h))\n        y2 = int(np.clip(h_y + length / 2, 0, h))\n        x1 = int(np.clip(h_x - length / 2, 0, w))\n        x2 = int(np.clip(h_x + length / 2, 0, w))\n        x[:, y1:y2, x1:x2] = 0\n    return x\n\ncutout = TfmPixel(_cutout, order=20)\n\n\ntfms = get_transforms(do_flip=False, \n                 xtra_tfms=[colorjitter(saturation=1.1, hue=0.05), \n                 cutout(n_holes=(1, max(3, int(3*rcutout))),\n                 length=(10*rcutout, 40*rcutout), p=0.5)])\n</code></pre>\n\n<h3>Training</h3>\n\n<p>I train only the Head part of the model with BH for a few epochs (Densenet121 is using imagenet weights) and then I unfreeze the model and train for longer with SH + BH (the first epoch only BH however). Both before and after unfreezing I use one cycle policy (<a href=\"https://sgugger.github.io/the-1cycle-policy.html\">https://sgugger.github.io/the-1cycle-policy.html</a>) implemented in fastai v1. Just one cycle before unfreezing and one cycle after. I found this worked better than multiple cycles in my experiments.</p>\n\n<h3>Best single model + TTA</h3>\n\n<ul>\n<li>Image size; public LB;  approximate train time (1 nvidia 1080)</li>\n<li>64x128 – 0.788 - 3.5h</li>\n<li>128x256 – 0.867 - 6h</li>\n<li>256x512 – 0.904 - 15h</li>\n</ul>\n\n<p>Final score is an ensemble of 13 submissions with public LB &gt; 0.860, reaching the final score of ~ <strong>0.93</strong>.</p>\n\n<h3>References</h3>\n\n<p>[1] <a href=\"https://ydwen.github.io/papers/WenECCV16.pdf\">https://ydwen.github.io/papers/WenECCV16.pdf</a></p>\n\n<p>[2] <a href=\"https://arxiv.org/pdf/1703.05175.pdf\">https://arxiv.org/pdf/1703.05175.pdf</a></p>\n\n<p>[3] <a href=\"https://arxiv.org/pdf/1703.07737.pdf\">https://arxiv.org/pdf/1703.07737.pdf</a></p>\n\n<p>[4] <a href=\"https://arxiv.org/pdf/1901.03662.pdf\">https://arxiv.org/pdf/1901.03662.pdf</a></p>\n\n<p>[5] <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79086\">https://www.kaggle.com/c/humpback-whale-identification/discussion/79086</a></p>\n\n<p>[6] <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79524\">https://www.kaggle.com/c/humpback-whale-identification/discussion/79524</a></p>",
      "rawMarkdown": "When I arrived on this competition I had no idea how to solve this kind of problem. It took me a while just to find papers about this problem and understand the intuition behind the concepts involved like episodes, n-way k-shot, metric learning, few-shot learning, meta learning and so on. I implemented all the framework in fastai v1.\n\n\n### Loss Function\nI have tried several approaches like central loss [1] and prototypical networks [2] but I then found on the discussions this two papers [3], [4] where a variation of triplet loss is used where the hard margin is replaced by a soft margin using the softplus function (that's why I'm referring to it as Soft Triplet Loss). They also use a Batch Hard strategy (BH) in which for each anchor image only the hardest positive and hardest negatives in the mini-batch are used in the loss. This is my implementation of Soft Triplet Loss with Batch Hard and L2 regularization (as suggested by @Iafoss [5]):\n\n    \n    class SoftTripletLoss(nn.Module):\n        def __init__(self, fsc, wd=1e-4):\n            super().__init__()\n            self.k_shot = fsc.k_shot\n            self.new_class_number = fsc.new_class_number\n            self.wd = wd\n            \n        def forward(self, x, y):\n            # x (64, 128)\n            self.n_way = x.size()[0]//self.k_shot\n            emb_sz = x.size()[-1] # (128)\n            x = x.view(-1, self.k_shot, emb_sz) # (16, 4, 128)\n            L = 0; EPS = 1e-6\n            for i in range(self.n_way-self.new_class_number):\n                for j in range(self.k_shot):\n                    I = torch.zeros(self.n_way).long()\n                    I[i] = 1\n                    J = torch.zeros(self.k_shot).long()\n                    J[j] = 1\n                    xa = x[I==1, J==1, :].view(1, -1) # (1, 128)\n                    xp = x[I==1, J==0, :] # (3, 128)\n                    xn = x[I==0, :, :].view(-1, emb_sz) # (15, 4, 128) -&gt; (60, 128)\n                    Dp = F.relu((xa-xp).pow_(2).sum(1)+EPS).sqrt_() # (3)\n                    Dn = F.relu((xa-xn).pow_(2).sum(1)+EPS).sqrt_() # (60)\n                    L += F.softplus(Dp.max(0)[0] - Dn.min(0)[0]) # (1)\n                    L += self.wd*((Dp**2).mean() + (Dn**2).mean()) # Regularization\n            return L\n\n\n### Hard samples mining\nBatch hard strategy allowed for a good improvement but it was not enough alone. The next main step was implementing a technique for mining hard samples, hereafter referred as Sample Hard (SH). At the end of each training epoch I compute the distance matrix between all train samples, then to build the mini-batches for the next epoch I follow the steps (I'm using 10-way, 4-shot episodes):\n\n1. Select an image at random [A1]\n2. Select the 3 hardest images from the same class (hard positives, largest distances) [A1, A2, A3, A4]\n3. Select the one hardest image from a different class (hard negatives, closest distance) [A1, A2, A3, A4, B1]\n4. Repeat from 2. until the mini-batch is constructed [A1, A2, A3, A4, B1, B2, B3, B4, C1, C2, C3, C4, ...] \n5. For the last 4 images in each mini-batch I select 4 random *new whales*.\n\nSo the distance matrix and mini-batches for the next epoch are only computed at the end of each epoch. This SH strategy is then used together with BH. \n\n\n### Model\nThe model I used is a **Densenet121** with the following head:\n\n    class Head(nn.Module):\n        def __init__(self, in_channels=1024, emb_sz=128):\n            super().__init__()\n            self.flat = nn.Sequential(\n                AdaptiveConcatPool2d(1))\n            self.flatten = Flatten()\n            self.bn0 = nn.BatchNorm1d(4*in_channels)\n            self.lin0 = nn.Linear(4*in_channels, in_channels)\n            self.relu = nn.ReLU(inplace=True)\n            self.bn1 = nn.BatchNorm1d(in_channels)\n            self.lin1 = nn.Linear(in_channels, emb_sz)\n        \n        def forward(self, x):\n            cut = x.size()[-1]//2\n            x0 = self.flat(x[...,:cut])\n            x1 = self.flat(x[...,cut:])\n            x = torch.cat((self.flatten(x0), self.flatten(x1)), dim=1)\n            x = self.relu(self.lin0(self.bn0(x)))\n            return self.lin1(self.bn1(x))\n\nSince I'm using images with ratio 1:2 and there are some horizontal symmetry I divide the images in two (left and right parts) (x0 and x1 in the code), I apply the pooling as usual (using AdaptiveConcatPool2d from fastai) and I finally concatenate the two. The activation maps shown by Heng [6] in another discussion topic show that in some images there are two modes in the activation map, one for the left and other for the right part, that is the intuition why this may help. I didn't check however if the increase in performance is just due to the increase in parameters.\n\n\n### Image Augmentations\nI cropped the images with bounding boxes and applied the following augmentations (fastai):\n\n    from torchvision.transforms import ColorJitter, ToPILImage, ToTensor\n    from fastai.vision.transform import *\n    def _colorjitter(x, brightness=0, contrast=0, saturation=0, hue=0):\n        topill = ToPILImage()\n        totensor = ToTensor()\n        xmin, xmax = x.min(), x.max()\n        x = (x-xmin)/(xmax-xmin)\n        cj = ColorJitter(brightness, contrast, saturation, hue)\n        x = topill(x)\n        x = cj(x)\n        x = totensor(x)\n        x = x*(xmax-xmin) + xmin\n        return x\n    colorjitter = TfmLighting(_colorjitter)\n    \n    def _cutout(x, n_holes:uniform_int=1, length:uniform_int=40):\n        \"Cut out `n_holes` number of square holes of size `length` in image at random locations.\"\n        h,w = x.shape[1:]\n        for n in range(n_holes):\n            h_y = np.random.randint(0, h)\n            h_x = np.random.randint(0, w)\n            y1 = int(np.clip(h_y - length / 2, 0, h))\n            y2 = int(np.clip(h_y + length / 2, 0, h))\n            x1 = int(np.clip(h_x - length / 2, 0, w))\n            x2 = int(np.clip(h_x + length / 2, 0, w))\n            x[:, y1:y2, x1:x2] = 0\n        return x\n    \n    cutout = TfmPixel(_cutout, order=20)\n\n\n    tfms = get_transforms(do_flip=False, \n                     xtra_tfms=[colorjitter(saturation=1.1, hue=0.05), \n                     cutout(n_holes=(1, max(3, int(3*rcutout))),\n                     length=(10*rcutout, 40*rcutout), p=0.5)])\n\n\n\n### Training\nI train only the Head part of the model with BH for a few epochs (Densenet121 is using imagenet weights) and then I unfreeze the model and train for longer with SH + BH (the first epoch only BH however). Both before and after unfreezing I use one cycle policy (https://sgugger.github.io/the-1cycle-policy.html) implemented in fastai v1. Just one cycle before unfreezing and one cycle after. I found this worked better than multiple cycles in my experiments.\n\n\n### Best single model + TTA\n* Image size; public LB;  approximate train time (1 nvidia 1080)\n* 64x128 – 0.788 - 3.5h\n* 128x256 – 0.867 - 6h\n* 256x512 – 0.904 - 15h\n\nFinal score is an ensemble of 13 submissions with public LB &gt; 0.860, reaching the final score of ~ **0.93**.\n\n### References\n[1] https://ydwen.github.io/papers/WenECCV16.pdf\n\n[2] https://arxiv.org/pdf/1703.05175.pdf\n\n[3] https://arxiv.org/pdf/1703.07737.pdf\n\n[4] https://arxiv.org/pdf/1901.03662.pdf\n\n[5] https://www.kaggle.com/c/humpback-whale-identification/discussion/79086\n\n[6] https://www.kaggle.com/c/humpback-whale-identification/discussion/79524",
      "votes": null
    },
    {
      "id": "481246",
      "postDate": "03/01/2019 07:33:37",
      "content": "<p>Congratulations on getting silver medal <a href=\"/mnpinto\">@mnpinto</a></p>",
      "rawMarkdown": "Congratulations on getting silver medal @mnpinto",
      "votes": null
    },
    {
      "id": "481267",
      "postDate": "03/01/2019 07:53:12",
      "content": "<p>Cool! Thanks for your share! I really learned something form it.</p>",
      "rawMarkdown": "Cool! Thanks for your share! I really learned something form it.",
      "votes": null
    },
    {
      "id": "481479",
      "postDate": "03/01/2019 13:11:05",
      "content": "<p>Congrats <a href=\"/mnpinto\">@mnpinto</a> and thanks for sharing.</p>",
      "rawMarkdown": "Congrats @mnpinto and thanks for sharing.",
      "votes": null
    },
    {
      "id": "481830",
      "postDate": "03/01/2019 22:35:58",
      "content": "<p>simple solution and nice ranking, congrats!</p>",
      "rawMarkdown": "simple solution and nice ranking, congrats!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 481246,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "03/01/2019 07:33:37",
      "content": "<p>Congratulations on getting silver medal <a href=\"/mnpinto\">@mnpinto</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481267,
      "author_name": "benwu232",
      "author_url": "",
      "post_date": "03/01/2019 07:53:12",
      "content": "<p>Cool! Thanks for your share! I really learned something form it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481479,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/01/2019 13:11:05",
      "content": "<p>Congrats <a href=\"/mnpinto\">@mnpinto</a> and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481830,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/01/2019 22:35:58",
      "content": "<p>simple solution and nice ranking, congrats!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "481034": "When I arrived on this competition I had no idea how to solve this kind of problem. It took me a while just to find papers about this problem and understand the intuition behind the concepts involved like episodes, n-way k-shot, metric learning, few-shot learning, meta learning and so on. I implemented all the framework in fastai v1.\n\n\n### Loss Function\nI have tried several approaches like central loss [1] and prototypical networks [2] but I then found on the discussions this two papers [3], [4] where a variation of triplet loss is used where the hard margin is replaced by a soft margin using the softplus function (that's why I'm referring to it as Soft Triplet Loss). They also use a Batch Hard strategy (BH) in which for each anchor image only the hardest positive and hardest negatives in the mini-batch are used in the loss. This is my implementation of Soft Triplet Loss with Batch Hard and L2 regularization (as suggested by @Iafoss [5]):\n\n    \n    class SoftTripletLoss(nn.Module):\n        def __init__(self, fsc, wd=1e-4):\n            super().__init__()\n            self.k_shot = fsc.k_shot\n            self.new_class_number = fsc.new_class_number\n            self.wd = wd\n            \n        def forward(self, x, y):\n            # x (64, 128)\n            self.n_way = x.size()[0]//self.k_shot\n            emb_sz = x.size()[-1] # (128)\n            x = x.view(-1, self.k_shot, emb_sz) # (16, 4, 128)\n            L = 0; EPS = 1e-6\n            for i in range(self.n_way-self.new_class_number):\n                for j in range(self.k_shot):\n                    I = torch.zeros(self.n_way).long()\n                    I[i] = 1\n                    J = torch.zeros(self.k_shot).long()\n                    J[j] = 1\n                    xa = x[I==1, J==1, :].view(1, -1) # (1, 128)\n                    xp = x[I==1, J==0, :] # (3, 128)\n                    xn = x[I==0, :, :].view(-1, emb_sz) # (15, 4, 128) -&gt; (60, 128)\n                    Dp = F.relu((xa-xp).pow_(2).sum(1)+EPS).sqrt_() # (3)\n                    Dn = F.relu((xa-xn).pow_(2).sum(1)+EPS).sqrt_() # (60)\n                    L += F.softplus(Dp.max(0)[0] - Dn.min(0)[0]) # (1)\n                    L += self.wd*((Dp**2).mean() + (Dn**2).mean()) # Regularization\n            return L\n\n\n### Hard samples mining\nBatch hard strategy allowed for a good improvement but it was not enough alone. The next main step was implementing a technique for mining hard samples, hereafter referred as Sample Hard (SH). At the end of each training epoch I compute the distance matrix between all train samples, then to build the mini-batches for the next epoch I follow the steps (I'm using 10-way, 4-shot episodes):\n\n1. Select an image at random [A1]\n2. Select the 3 hardest images from the same class (hard positives, largest distances) [A1, A2, A3, A4]\n3. Select the one hardest image from a different class (hard negatives, closest distance) [A1, A2, A3, A4, B1]\n4. Repeat from 2. until the mini-batch is constructed [A1, A2, A3, A4, B1, B2, B3, B4, C1, C2, C3, C4, ...] \n5. For the last 4 images in each mini-batch I select 4 random *new whales*.\n\nSo the distance matrix and mini-batches for the next epoch are only computed at the end of each epoch. This SH strategy is then used together with BH. \n\n\n### Model\nThe model I used is a **Densenet121** with the following head:\n\n    class Head(nn.Module):\n        def __init__(self, in_channels=1024, emb_sz=128):\n            super().__init__()\n            self.flat = nn.Sequential(\n                AdaptiveConcatPool2d(1))\n            self.flatten = Flatten()\n            self.bn0 = nn.BatchNorm1d(4*in_channels)\n            self.lin0 = nn.Linear(4*in_channels, in_channels)\n            self.relu = nn.ReLU(inplace=True)\n            self.bn1 = nn.BatchNorm1d(in_channels)\n            self.lin1 = nn.Linear(in_channels, emb_sz)\n        \n        def forward(self, x):\n            cut = x.size()[-1]//2\n            x0 = self.flat(x[...,:cut])\n            x1 = self.flat(x[...,cut:])\n            x = torch.cat((self.flatten(x0), self.flatten(x1)), dim=1)\n            x = self.relu(self.lin0(self.bn0(x)))\n            return self.lin1(self.bn1(x))\n\nSince I'm using images with ratio 1:2 and there are some horizontal symmetry I divide the images in two (left and right parts) (x0 and x1 in the code), I apply the pooling as usual (using AdaptiveConcatPool2d from fastai) and I finally concatenate the two. The activation maps shown by Heng [6] in another discussion topic show that in some images there are two modes in the activation map, one for the left and other for the right part, that is the intuition why this may help. I didn't check however if the increase in performance is just due to the increase in parameters.\n\n\n### Image Augmentations\nI cropped the images with bounding boxes and applied the following augmentations (fastai):\n\n    from torchvision.transforms import ColorJitter, ToPILImage, ToTensor\n    from fastai.vision.transform import *\n    def _colorjitter(x, brightness=0, contrast=0, saturation=0, hue=0):\n        topill = ToPILImage()\n        totensor = ToTensor()\n        xmin, xmax = x.min(), x.max()\n        x = (x-xmin)/(xmax-xmin)\n        cj = ColorJitter(brightness, contrast, saturation, hue)\n        x = topill(x)\n        x = cj(x)\n        x = totensor(x)\n        x = x*(xmax-xmin) + xmin\n        return x\n    colorjitter = TfmLighting(_colorjitter)\n    \n    def _cutout(x, n_holes:uniform_int=1, length:uniform_int=40):\n        \"Cut out `n_holes` number of square holes of size `length` in image at random locations.\"\n        h,w = x.shape[1:]\n        for n in range(n_holes):\n            h_y = np.random.randint(0, h)\n            h_x = np.random.randint(0, w)\n            y1 = int(np.clip(h_y - length / 2, 0, h))\n            y2 = int(np.clip(h_y + length / 2, 0, h))\n            x1 = int(np.clip(h_x - length / 2, 0, w))\n            x2 = int(np.clip(h_x + length / 2, 0, w))\n            x[:, y1:y2, x1:x2] = 0\n        return x\n    \n    cutout = TfmPixel(_cutout, order=20)\n\n\n    tfms = get_transforms(do_flip=False, \n                     xtra_tfms=[colorjitter(saturation=1.1, hue=0.05), \n                     cutout(n_holes=(1, max(3, int(3*rcutout))),\n                     length=(10*rcutout, 40*rcutout), p=0.5)])\n\n\n\n### Training\nI train only the Head part of the model with BH for a few epochs (Densenet121 is using imagenet weights) and then I unfreeze the model and train for longer with SH + BH (the first epoch only BH however). Both before and after unfreezing I use one cycle policy (https://sgugger.github.io/the-1cycle-policy.html) implemented in fastai v1. Just one cycle before unfreezing and one cycle after. I found this worked better than multiple cycles in my experiments.\n\n\n### Best single model + TTA\n* Image size; public LB;  approximate train time (1 nvidia 1080)\n* 64x128 – 0.788 - 3.5h\n* 128x256 – 0.867 - 6h\n* 256x512 – 0.904 - 15h\n\nFinal score is an ensemble of 13 submissions with public LB &gt; 0.860, reaching the final score of ~ **0.93**.\n\n### References\n[1] https://ydwen.github.io/papers/WenECCV16.pdf\n\n[2] https://arxiv.org/pdf/1703.05175.pdf\n\n[3] https://arxiv.org/pdf/1703.07737.pdf\n\n[4] https://arxiv.org/pdf/1901.03662.pdf\n\n[5] https://www.kaggle.com/c/humpback-whale-identification/discussion/79086\n\n[6] https://www.kaggle.com/c/humpback-whale-identification/discussion/79524",
    "481246": "Congratulations on getting silver medal @mnpinto",
    "481267": "Cool! Thanks for your share! I really learned something form it.",
    "481479": "Congrats @mnpinto and thanks for sharing.",
    "481830": "simple solution and nice ranking, congrats!"
  },
  "source": "meta"
}