{
  "id": 220446,
  "title": "#6 Solution 🐼Tropic Thunder🐼 🚫 No hand labels🚫",
  "url": "/competitions/rfcx-species-audio-detection/writeups/tropic-thunder-6-solution-tropic-thunder-no-hand-l",
  "author_name": "",
  "post_date": "2021-02-18T21:29:00.530Z",
  "votes": 53,
  "comment_count": 16,
  "views": 0,
  "content": "<p>It was a great competition. I'll document our journey primarily from my perspective, my teammates <a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> <a href=\"https://www.kaggle.com/amezet\" target=\"_blank\">@amezet</a> <a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> may chime in to add extra color… </p>\n<p>Our solution is rank ensemble of two types of models, primarily 97% using the architecture described below, and 3% an ensemble of SED models. The architecture below was made in the last 7 days… I made my first (pretty bad) sub 8 days ago, and I joined team just before team merge deadline and all the ideas were done/implemented in ~7 days…</p>\n<p>A single 5-fold model achieves 0.961 private, 0.956 public; details:</p>\n<p><strong>Input representation.</strong> This is probably the key, we take each TP or FP and after spectrogram (not MEL, b/c MEL was intended to model human audition and the rainforrest species have not evolved like our human audition) crop it time and frequency wise with fixed size on a per-class basis. E.g. for class 0: minimum freq is 5906.25 Hz and max is 8250 Hz, and time-wise we take the longest of time length of TPs, i.e. for class 0: 1.29s. </p>\n<p>With the above sample the spectrogram to yield an image, e.g.:<br>\n<img src=\"https://i.imgur.com/bTw0aeQ.png\" alt=\"\"></p>\n<p>The GT for that TP would be: </p>\n<p><code>tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))</code></p>\n<p>As other competitors, we expand the classes from 24 to 26 b/c to split species that have two songs. </p>\n<p>Other sample:</p>\n<p><img src=\"https://i.imgur.com/JST401j.png\" alt=\"\"></p>\n<p><code>tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))</code></p>\n<p>Note the image size is the same, so time and frequency are effectively strectched wrt to what the net will see, but I believe this is fine as long as the net has enough receptive field (which they have). </p>\n<p>The reason of doing it this way is that we need to inject the time and frequency restrictions as inductive bias somehow, and this looks like an nice way.</p>\n<p><strong>Model architecture.</strong> This is just an image classifier outputing 26 logits, that's it. The only whistle is to add relative positional information in the freq axis (à la <a href=\"https://arxiv.org/abs/1807.03247\" target=\"_blank\">coordconv</a>), so model is embarrasingly simple:</p>\n<pre><code>class TropicModel(Module):\n    def __init__(self):\n        self.trunk = timm.create_model(a.arch,pretrained=True,num_classes=n_species,in_chans=1+a.coord)\n        self.do = nn.Dropout2d(a.do)\n    def forward(self,x):\n        bs,_,freq_bins,time_bins = x.size()\n        coord = torch.linspace(-1,1,freq_bins,dtype=x.dtype,device=x.device).view(1,1,-1,1).expand(bs,1,-1,time_bins)\n        if a.coord: x = torch.cat((x,coord),dim=1)\n        x = self.do(x)\n        return self.trunk(x)\n</code></pre>\n<p><strong>Loss function.</strong> Just masked Focal loss. Actually this was a mistake b/c Focal loss was a renmant of a dead test and I (accidentally) left it there, where I thought (until I checked code now to write writeup) that BCE was being used. Since we are doing balancing BCE should work better.</p>\n<p><strong>Mixover.</strong> Inspired by mixup, mixover takes a bunch of (unbalanced) TP and FPs (which strictly speaking are TNs) and creates combinations of them so that the resulting labels can be supervised (1+NaN=1, 0+NaN=NaN,0+0=0) making sure linear interpolation is not destructive (beta,alpha=4; clip to 0.2,0.8); then it samples from the resulted mixed items computing class distribution so that resulting samples are balanced. Code is a bit tricky, but still:</p>\n<pre><code>class MixOver(MixHandler):\n    \"Inspired by implementation of https://arxiv.org/abs/1710.09412\"\n    def __init__(self, alpha=): super().__init__(alpha)\n    def before_batch(self):\n        ny_dims,nx_dims = len(self.y.size()),len(self.x.size())\n        bs=find_bs(self.xb)\n        all_combinations = L(itertools.combinations(range(find_bs(self.xb)), 2))\n        lam = self.distrib.sample((len(all_combinations),)).squeeze().to(self.x.device).clip(0.2,0.8)\n        lam = torch.stack([lam, 1-lam], 1)\n        self.lam = lam.max(1)[0]\n        comb = all_combinations\n        yb0,yb1 = L(self.yb).itemgot(comb.itemgot(0))[0],L(self.yb).itemgot(comb.itemgot(1))[0]\n        yb_one = torch.full_like(yb0,np.nan)\n        yb_one[yb0&gt;0.5] = yb0[yb0&gt;0.5]\n        yb_one[yb1&gt;0.5] = yb1[yb1&gt;0.5]\n        yb_two = torch.clip(yb0+yb1,0,1.)\n        yb_com = yb_one.clone()\n        yb_com[~torch.isnan(yb_two)] = yb_two[~torch.isnan(yb_two)]\n        n_ones_or_zeros=(~torch.isnan(yb_com)).sum()\n        ones=torch.sum(yb_com&gt;=0.5,dim=1)\n        zeros=torch.sum(yb_com&lt;0.5,dim=1)\n        p_ones = (n_ones_or_zeros/(2*( ones.sum())))/ones\n        p_zeros= (n_ones_or_zeros/(2*(zeros.sum())))/zeros\n        p_zeros[torch.isinf(p_zeros)],p_ones[torch.isinf(p_ones)]=0,0\n        p=(p_ones+p_zeros).cpu().numpy()/(p_ones+p_zeros).sum().item()\n        shuffle=torch.from_numpy(np.random.choice(yb_com.size(0),size=bs,replace=True,p=p)).to(self.x.device)\n        comb = all_combinations[shuffle]\n        xb0,xb1 = tuple(L(self.xb).itemgot(comb.itemgot(0))),tuple(L(self.xb).itemgot(comb.itemgot(1)))\n        self.learn.xb = tuple(L(xb0,xb1).map_zip(torch.lerp,weight=unsqueeze(self.lam[shuffle], n=nx_dims-1)))\n        self.learn.yb = (yb_com[shuffle],)\n</code></pre>\n<p><strong>Augmentations.</strong> Time jitter (10% of time length), white noise (3 dB).</p>\n<p><strong>External data.</strong> We used <a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">Xeno Canto A-M and N-Z</a> recording and since they have 264 species we made the wild assumption that if we have 24 species and they have 10X randomly sampling would give you a \"right\" label (TN) 90% of the time… the goal was to add jungle/rainforest diversity at the expense of a small noisy labels which were accounted for labeling these weak labels as 0.1 (vs 0).</p>\n<p><strong>Pseudolabeling.</strong> We also pseudolabeled using OOF models the non labeled parts of training data to mine more TPs, and manually balanced the resulting pseudo-label TPs. </p>\n<p><strong>Code.</strong> Pytorch, Fastai.</p>\n<p><strong>Final thoughts.</strong> It was a very fun competition and I am very glad of achieving a gold medal with a very short time, I want to thank Kaggle, sponsor and competitors for the competition and discussions; and of course my teammates <a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> <a href=\"https://www.kaggle.com/amezet\" target=\"_blank\">@amezet</a> <a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> for inviting me giving me the small nudge I needed to join 😉</p>\n<p>Edit: I've added no hand labels to the title because this competition was VERY unique in that Kaggle effectively allowed hand labeling and that's very unusual. (Re: hand labeling vs external data, it was also very uncommon not having to disclose which external data you used during competition).</p>",
  "messages": [
    {
      "id": "1208493",
      "postDate": "02/18/2021 09:53:06",
      "content": "<p>It was a great competition. I'll document our journey primarily from my perspective, my teammates <a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> <a href=\"https://www.kaggle.com/amezet\" target=\"_blank\">@amezet</a> <a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> may chime in to add extra color… </p>\n<p>Our solution is rank ensemble of two types of models, primarily 97% using the architecture described below, and 3% an ensemble of SED models. The architecture below was made in the last 7 days… I made my first (pretty bad) sub 8 days ago, and I joined team just before team merge deadline and all the ideas were done/implemented in ~7 days…</p>\n<p>A single 5-fold model achieves 0.961 private, 0.956 public; details:</p>\n<p><strong>Input representation.</strong> This is probably the key, we take each TP or FP and after spectrogram (not MEL, b/c MEL was intended to model human audition and the rainforrest species have not evolved like our human audition) crop it time and frequency wise with fixed size on a per-class basis. E.g. for class 0: minimum freq is 5906.25 Hz and max is 8250 Hz, and time-wise we take the longest of time length of TPs, i.e. for class 0: 1.29s. </p>\n<p>With the above sample the spectrogram to yield an image, e.g.:<br>\n<img src=\"https://i.imgur.com/bTw0aeQ.png\" alt=\"\"></p>\n<p>The GT for that TP would be: </p>\n<p><code>tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))</code></p>\n<p>As other competitors, we expand the classes from 24 to 26 b/c to split species that have two songs. </p>\n<p>Other sample:</p>\n<p><img src=\"https://i.imgur.com/JST401j.png\" alt=\"\"></p>\n<p><code>tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))</code></p>\n<p>Note the image size is the same, so time and frequency are effectively strectched wrt to what the net will see, but I believe this is fine as long as the net has enough receptive field (which they have). </p>\n<p>The reason of doing it this way is that we need to inject the time and frequency restrictions as inductive bias somehow, and this looks like an nice way.</p>\n<p><strong>Model architecture.</strong> This is just an image classifier outputing 26 logits, that's it. The only whistle is to add relative positional information in the freq axis (à la <a href=\"https://arxiv.org/abs/1807.03247\" target=\"_blank\">coordconv</a>), so model is embarrasingly simple:</p>\n<pre><code>class TropicModel(Module):\n    def __init__(self):\n        self.trunk = timm.create_model(a.arch,pretrained=True,num_classes=n_species,in_chans=1+a.coord)\n        self.do = nn.Dropout2d(a.do)\n    def forward(self,x):\n        bs,_,freq_bins,time_bins = x.size()\n        coord = torch.linspace(-1,1,freq_bins,dtype=x.dtype,device=x.device).view(1,1,-1,1).expand(bs,1,-1,time_bins)\n        if a.coord: x = torch.cat((x,coord),dim=1)\n        x = self.do(x)\n        return self.trunk(x)\n</code></pre>\n<p><strong>Loss function.</strong> Just masked Focal loss. Actually this was a mistake b/c Focal loss was a renmant of a dead test and I (accidentally) left it there, where I thought (until I checked code now to write writeup) that BCE was being used. Since we are doing balancing BCE should work better.</p>\n<p><strong>Mixover.</strong> Inspired by mixup, mixover takes a bunch of (unbalanced) TP and FPs (which strictly speaking are TNs) and creates combinations of them so that the resulting labels can be supervised (1+NaN=1, 0+NaN=NaN,0+0=0) making sure linear interpolation is not destructive (beta,alpha=4; clip to 0.2,0.8); then it samples from the resulted mixed items computing class distribution so that resulting samples are balanced. Code is a bit tricky, but still:</p>\n<pre><code>class MixOver(MixHandler):\n    \"Inspired by implementation of https://arxiv.org/abs/1710.09412\"\n    def __init__(self, alpha=): super().__init__(alpha)\n    def before_batch(self):\n        ny_dims,nx_dims = len(self.y.size()),len(self.x.size())\n        bs=find_bs(self.xb)\n        all_combinations = L(itertools.combinations(range(find_bs(self.xb)), 2))\n        lam = self.distrib.sample((len(all_combinations),)).squeeze().to(self.x.device).clip(0.2,0.8)\n        lam = torch.stack([lam, 1-lam], 1)\n        self.lam = lam.max(1)[0]\n        comb = all_combinations\n        yb0,yb1 = L(self.yb).itemgot(comb.itemgot(0))[0],L(self.yb).itemgot(comb.itemgot(1))[0]\n        yb_one = torch.full_like(yb0,np.nan)\n        yb_one[yb0&gt;0.5] = yb0[yb0&gt;0.5]\n        yb_one[yb1&gt;0.5] = yb1[yb1&gt;0.5]\n        yb_two = torch.clip(yb0+yb1,0,1.)\n        yb_com = yb_one.clone()\n        yb_com[~torch.isnan(yb_two)] = yb_two[~torch.isnan(yb_two)]\n        n_ones_or_zeros=(~torch.isnan(yb_com)).sum()\n        ones=torch.sum(yb_com&gt;=0.5,dim=1)\n        zeros=torch.sum(yb_com&lt;0.5,dim=1)\n        p_ones = (n_ones_or_zeros/(2*( ones.sum())))/ones\n        p_zeros= (n_ones_or_zeros/(2*(zeros.sum())))/zeros\n        p_zeros[torch.isinf(p_zeros)],p_ones[torch.isinf(p_ones)]=0,0\n        p=(p_ones+p_zeros).cpu().numpy()/(p_ones+p_zeros).sum().item()\n        shuffle=torch.from_numpy(np.random.choice(yb_com.size(0),size=bs,replace=True,p=p)).to(self.x.device)\n        comb = all_combinations[shuffle]\n        xb0,xb1 = tuple(L(self.xb).itemgot(comb.itemgot(0))),tuple(L(self.xb).itemgot(comb.itemgot(1)))\n        self.learn.xb = tuple(L(xb0,xb1).map_zip(torch.lerp,weight=unsqueeze(self.lam[shuffle], n=nx_dims-1)))\n        self.learn.yb = (yb_com[shuffle],)\n</code></pre>\n<p><strong>Augmentations.</strong> Time jitter (10% of time length), white noise (3 dB).</p>\n<p><strong>External data.</strong> We used <a href=\"https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m\" target=\"_blank\">Xeno Canto A-M and N-Z</a> recording and since they have 264 species we made the wild assumption that if we have 24 species and they have 10X randomly sampling would give you a \"right\" label (TN) 90% of the time… the goal was to add jungle/rainforest diversity at the expense of a small noisy labels which were accounted for labeling these weak labels as 0.1 (vs 0).</p>\n<p><strong>Pseudolabeling.</strong> We also pseudolabeled using OOF models the non labeled parts of training data to mine more TPs, and manually balanced the resulting pseudo-label TPs. </p>\n<p><strong>Code.</strong> Pytorch, Fastai.</p>\n<p><strong>Final thoughts.</strong> It was a very fun competition and I am very glad of achieving a gold medal with a very short time, I want to thank Kaggle, sponsor and competitors for the competition and discussions; and of course my teammates <a href=\"https://www.kaggle.com/jpison\" target=\"_blank\">@jpison</a> <a href=\"https://www.kaggle.com/amezet\" target=\"_blank\">@amezet</a> <a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> for inviting me giving me the small nudge I needed to join 😉</p>\n<p>Edit: I've added no hand labels to the title because this competition was VERY unique in that Kaggle effectively allowed hand labeling and that's very unusual. (Re: hand labeling vs external data, it was also very uncommon not having to disclose which external data you used during competition).</p>",
      "rawMarkdown": "It was a great competition. I'll document our journey primarily from my perspective, my teammates @jpison @amezet @pavelgonchar may chime in to add extra color... \n\nOur solution is rank ensemble of two types of models, primarily 97% using the architecture described below, and 3% an ensemble of SED models. The architecture below was made in the last 7 days... I made my first (pretty bad) sub 8 days ago, and I joined team just before team merge deadline and all the ideas were done/implemented in ~7 days...\n\nA single 5-fold model achieves 0.961 private, 0.956 public; details:\n\n**Input representation.** This is probably the key, we take each TP or FP and after spectrogram (not MEL, b/c MEL was intended to model human audition and the rainforrest species have not evolved like our human audition) crop it time and frequency wise with fixed size on a per-class basis. E.g. for class 0: minimum freq is 5906.25 Hz and max is 8250 Hz, and time-wise we take the longest of time length of TPs, i.e. for class 0: 1.29s. \n\nWith the above sample the spectrogram to yield an image, e.g.:\n![](https://i.imgur.com/bTw0aeQ.png)\n\nThe GT for that TP would be: \n\n```tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))```\n\nAs other competitors, we expand the classes from 24 to 26 b/c to split species that have two songs. \n\nOther sample:\n\n![](https://i.imgur.com/JST401j.png)\n\n``` tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))```\n\nNote the image size is the same, so time and frequency are effectively strectched wrt to what the net will see, but I believe this is fine as long as the net has enough receptive field (which they have). \n\nThe reason of doing it this way is that we need to inject the time and frequency restrictions as inductive bias somehow, and this looks like an nice way.\n\n**Model architecture.** This is just an image classifier outputing 26 logits, that's it. The only whistle is to add relative positional information in the freq axis (à la [coordconv](https://arxiv.org/abs/1807.03247)), so model is embarrasingly simple:\n\n```\nclass TropicModel(Module):\n    def __init__(self):\n        self.trunk = timm.create_model(a.arch,pretrained=True,num_classes=n_species,in_chans=1+a.coord)\n        self.do = nn.Dropout2d(a.do)\n    def forward(self,x):\n        bs,_,freq_bins,time_bins = x.size()\n        coord = torch.linspace(-1,1,freq_bins,dtype=x.dtype,device=x.device).view(1,1,-1,1).expand(bs,1,-1,time_bins)\n        if a.coord: x = torch.cat((x,coord),dim=1)\n        x = self.do(x)\n        return self.trunk(x)\n```\n\n**Loss function.** Just masked Focal loss. Actually this was a mistake b/c Focal loss was a renmant of a dead test and I (accidentally) left it there, where I thought (until I checked code now to write writeup) that BCE was being used. Since we are doing balancing BCE should work better.\n\n**Mixover.** Inspired by mixup, mixover takes a bunch of (unbalanced) TP and FPs (which strictly speaking are TNs) and creates combinations of them so that the resulting labels can be supervised (1+NaN=1, 0+NaN=NaN,0+0=0) making sure linear interpolation is not destructive (beta,alpha=4; clip to 0.2,0.8); then it samples from the resulted mixed items computing class distribution so that resulting samples are balanced. Code is a bit tricky, but still:\n\n```\nclass MixOver(MixHandler):\n    \"Inspired by implementation of https://arxiv.org/abs/1710.09412\"\n    def __init__(self, alpha=): super().__init__(alpha)\n    def before_batch(self):\n        ny_dims,nx_dims = len(self.y.size()),len(self.x.size())\n        bs=find_bs(self.xb)\n        all_combinations = L(itertools.combinations(range(find_bs(self.xb)), 2))\n        lam = self.distrib.sample((len(all_combinations),)).squeeze().to(self.x.device).clip(0.2,0.8)\n        lam = torch.stack([lam, 1-lam], 1)\n        self.lam = lam.max(1)[0]\n        comb = all_combinations\n        yb0,yb1 = L(self.yb).itemgot(comb.itemgot(0))[0],L(self.yb).itemgot(comb.itemgot(1))[0]\n        yb_one = torch.full_like(yb0,np.nan)\n        yb_one[yb0>0.5] = yb0[yb0>0.5]\n        yb_one[yb1>0.5] = yb1[yb1>0.5]\n        yb_two = torch.clip(yb0+yb1,0,1.)\n        yb_com = yb_one.clone()\n        yb_com[~torch.isnan(yb_two)] = yb_two[~torch.isnan(yb_two)]\n        n_ones_or_zeros=(~torch.isnan(yb_com)).sum()\n        ones=torch.sum(yb_com>=0.5,dim=1)\n        zeros=torch.sum(yb_com<0.5,dim=1)\n        p_ones = (n_ones_or_zeros/(2*( ones.sum())))/ones\n        p_zeros= (n_ones_or_zeros/(2*(zeros.sum())))/zeros\n        p_zeros[torch.isinf(p_zeros)],p_ones[torch.isinf(p_ones)]=0,0\n        p=(p_ones+p_zeros).cpu().numpy()/(p_ones+p_zeros).sum().item()\n        shuffle=torch.from_numpy(np.random.choice(yb_com.size(0),size=bs,replace=True,p=p)).to(self.x.device)\n        comb = all_combinations[shuffle]\n        xb0,xb1 = tuple(L(self.xb).itemgot(comb.itemgot(0))),tuple(L(self.xb).itemgot(comb.itemgot(1)))\n        self.learn.xb = tuple(L(xb0,xb1).map_zip(torch.lerp,weight=unsqueeze(self.lam[shuffle], n=nx_dims-1)))\n        self.learn.yb = (yb_com[shuffle],)\n```\n\n**Augmentations.** Time jitter (10% of time length), white noise (3 dB).\n\n**External data.** We used [Xeno Canto A-M and N-Z](https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m) recording and since they have 264 species we made the wild assumption that if we have 24 species and they have 10X randomly sampling would give you a \"right\" label (TN) 90% of the time... the goal was to add jungle/rainforest diversity at the expense of a small noisy labels which were accounted for labeling these weak labels as 0.1 (vs 0).\n\n**Pseudolabeling.** We also pseudolabeled using OOF models the non labeled parts of training data to mine more TPs, and manually balanced the resulting pseudo-label TPs. \n\n**Code.** Pytorch, Fastai.\n\n**Final thoughts.** It was a very fun competition and I am very glad of achieving a gold medal with a very short time, I want to thank Kaggle, sponsor and competitors for the competition and discussions; and of course my teammates @jpison @amezet @pavelgonchar for inviting me giving me the small nudge I needed to join 😉\n\nEdit: I've added no hand labels to the title because this competition was VERY unique in that Kaggle effectively allowed hand labeling and that's very unusual. (Re: hand labeling vs external data, it was also very uncommon not having to disclose which external data you used during competition).",
      "votes": null
    },
    {
      "id": "1208499",
      "postDate": "02/18/2021 09:56:18",
      "content": "<p>Congratz !</p>\n<blockquote>\n  <p>External data. We used Xeno Canto A-M and N-Z recording</p>\n</blockquote>\n<p>Interesting ! Do you know how much it helped your models in the end ?</p>",
      "rawMarkdown": "Congratz !\n\n> External data. We used Xeno Canto A-M and N-Z recording\n\nInteresting ! Do you know how much it helped your models in the end ?",
      "votes": null
    },
    {
      "id": "1208504",
      "postDate": "02/18/2021 10:00:40",
      "content": "<p>Significantly, all other params being equal we burned a sub testing impact:</p>\n<p><img src=\"https://i.imgur.com/PCyPUfA.png\" alt=\"\"></p>",
      "rawMarkdown": "Significantly, all other params being equal we burned a sub testing impact:\n\n![](https://i.imgur.com/PCyPUfA.png)",
      "votes": null
    },
    {
      "id": "1208513",
      "postDate": "02/18/2021 10:07:22",
      "content": "<p>That's a huge boost indeed!</p>",
      "rawMarkdown": "That's a huge boost indeed!",
      "votes": null
    },
    {
      "id": "1208556",
      "postDate": "02/18/2021 10:38:26",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> and team on 7th place</p>",
      "rawMarkdown": "Congrats @antorsae and team on 7th place",
      "votes": null
    },
    {
      "id": "1208657",
      "postDate": "02/18/2021 12:07:46",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": null
    },
    {
      "id": "1209117",
      "postDate": "02/18/2021 17:34:48",
      "content": "<p>congrats! very impressive getting #7 in such a short time, I spent the whole 3 months failing most of the time 😂</p>",
      "rawMarkdown": "congrats! very impressive getting #7 in such a short time, I spent the whole 3 months failing most of the time 😂",
      "votes": null
    },
    {
      "id": "1209340",
      "postDate": "02/18/2021 20:53:35",
      "content": "<p>Thanks!!! Long time since we did Avito!</p>",
      "rawMarkdown": "Thanks!!! Long time since we did Avito!",
      "votes": null
    },
    {
      "id": "1209357",
      "postDate": "02/18/2021 21:08:34",
      "content": "<p>yeah! 3 years and you and <a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> are still killing it</p>",
      "rawMarkdown": "yeah! 3 years and you and @pavelgonchar are still killing it",
      "votes": null
    },
    {
      "id": "1209381",
      "postDate": "02/18/2021 21:24:33",
      "content": "<p>Thanks for sharing.  Your approach is indeed very similar to mine, but you made it work better.  I wasn't aware of coordconv, it is certainly my next read.  </p>\n<p>Do you know how much you get from using external data? </p>\n<p>Congrats on the end result!</p>\n<p>Edit.  I see you answered the external data question already.</p>",
      "rawMarkdown": "Thanks for sharing.  Your approach is indeed very similar to mine, but you made it work better.  I wasn't aware of coordconv, it is certainly my next read.  \n\nDo you know how much you get from using external data? \n\nCongrats on the end result!\n\nEdit.  I see you answered the external data question already.",
      "votes": null
    },
    {
      "id": "1225257",
      "postDate": "03/03/2021 13:31:12",
      "content": "<p>Andres, do you think that using CoordConv based on sines and cosines (like Transformer Positional Encodings) would have improved the linspace CoordConv version?</p>",
      "rawMarkdown": "Andres, do you think that using CoordConv based on sines and cosines (like Transformer Positional Encodings) would have improved the linspace CoordConv version?",
      "votes": null
    },
    {
      "id": "1226258",
      "postDate": "03/04/2021 11:51:46",
      "content": "<p>Very nice question. I have been thinking about that I am not sure about the answer, I would need to test. But one thing for sure is that in that case the injection of coord matrixes should <em>also</em> be tested at the bottleneck, i.e. it makes sense to either insert positional not compressed and narrow (2ch) spatial encodings (similar to coordconv) or positional dense (512/1024/2048ch) encodings right before  final global average pooling. My suspicion the latter would be better but I have not tested (yet).</p>",
      "rawMarkdown": "Very nice question. I have been thinking about that I am not sure about the answer, I would need to test. But one thing for sure is that in that case the injection of coord matrixes should *also* be tested at the bottleneck, i.e. it makes sense to either insert positional not compressed and narrow (2ch) spatial encodings (similar to coordconv) or positional dense (512/1024/2048ch) encodings right before  final global average pooling. My suspicion the latter would be better but I have not tested (yet).",
      "votes": null
    },
    {
      "id": "1226293",
      "postDate": "03/04/2021 12:20:28",
      "content": "<p>Muy buena, Andrés. Enhorabuena al equipo! 👌</p>\n<p>Gracias a estos posts aprendo un poco más de lo mejores día a día! ❤️</p>\n<p>Saludos! 😃</p>",
      "rawMarkdown": "Muy buena, Andrés. Enhorabuena al equipo! 👌\n\nGracias a estos posts aprendo un poco más de lo mejores día a día! ❤️\n\nSaludos! 😃",
      "votes": null
    },
    {
      "id": "1322391",
      "postDate": "05/25/2021 12:30:56",
      "content": "<p>Thanks for the wonderful write-up! I have some rather generic questions (I hope it wont be too late to ask this):</p>\n<ol>\n<li>By masked loss, do u mean setting non-target classes to be <code>nan</code> (i.e. zero weight) as you did here?</li>\n</ol>\n<pre><code>tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))\n</code></pre>\n<ol>\n<li>Did u apply the same masking trick on FP data? (i.e. replace 1. by 0. for the tensor)</li>\n<li>I understand the rationale for applying masking on FP because these FP only implies the absence of a species, but implies no information wrt other species. However, what is the rationale for applying masking on TP data? </li>\n</ol>",
      "rawMarkdown": "Thanks for the wonderful write-up! I have some rather generic questions (I hope it wont be too late to ask this):\n1. By masked loss, do u mean setting non-target classes to be `nan` (i.e. zero weight) as you did here?\n```\ntensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))\n```\n2. Did u apply the same masking trick on FP data? (i.e. replace 1. by 0. for the tensor)\n3. I understand the rationale for applying masking on FP because these FP only implies the absence of a species, but implies no information wrt other species. However, what is the rationale for applying masking on TP data?",
      "votes": null
    },
    {
      "id": "1322611",
      "postDate": "05/25/2021 14:48:58",
      "content": "<p>Hi Alex,</p>\n<ol>\n<li>Yes.</li>\n<li>Yes.</li>\n<li>Rationale is the same: TP means species is present, we don't know anything else. FP means species is not present, we don't know anything else either; so everything is init as nan and populated w/ 1s and 0s with the information we know for sure.</li>\n</ol>",
      "rawMarkdown": "Hi Alex,\n\n1. Yes.\n2. Yes.\n3. Rationale is the same: TP means species is present, we don't know anything else. FP means species is not present, we don't know anything else either; so everything is init as nan and populated w/ 1s and 0s with the information we know for sure.",
      "votes": null
    },
    {
      "id": "1322616",
      "postDate": "05/25/2021 14:51:20",
      "content": "<p>awesome! thanks so much for your reply<br>\nIt helps a lot</p>",
      "rawMarkdown": "awesome! thanks so much for your reply\nIt helps a lot",
      "votes": null
    },
    {
      "id": "1358768",
      "postDate": "06/20/2021 17:58:47",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a><br>\nWould you mind me asking a few more questions?</p>\n<ol>\n<li>what is the image size that you used for model training?</li>\n<li>in pseudo labeling phase, you said you <code>manually balanced the resulting pseudo-label TPs</code>, would like to learn more about your sampling strategy. e.g. What is the sampling prob you set for TP, FP and pseudo labels?</li>\n<li>For pseudo labels, did you apply any penalty (e.g. soft penalty proportional to prediction confidence) on their loss?</li>\n<li>Did u generate pseudo labels by ensemble of ur OOF models?</li>\n<li>For model arch, if I understand correctly, it’s 2 channels, where the 2nd channel encode the positional info of the freq. Did u do any ablation study on the perf gain responsible by this “freq” channel? (Also I noticed the 2nd channel has quite a different scale that doesn’t range from 0 to 1. Do u think rescaling it back to 0 to 1 could further help?)</li>\n</ol>\n<p>Many thanks!</p>",
      "rawMarkdown": "Hi @antorsae\nWould you mind me asking a few more questions?\n1. what is the image size that you used for model training?\n2. in pseudo labeling phase, you said you `manually balanced the resulting pseudo-label TPs`, would like to learn more about your sampling strategy. e.g. What is the sampling prob you set for TP, FP and pseudo labels?\n3. For pseudo labels, did you apply any penalty (e.g. soft penalty proportional to prediction confidence) on their loss?\n4. Did u generate pseudo labels by ensemble of ur OOF models?\n5. For model arch, if I understand correctly, it’s 2 channels, where the 2nd channel encode the positional info of the freq. Did u do any ablation study on the perf gain responsible by this “freq” channel? (Also I noticed the 2nd channel has quite a different scale that doesn’t range from 0 to 1. Do u think rescaling it back to 0 to 1 could further help?)\n\nMany thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208499,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 09:56:18",
      "content": "<p>Congratz !</p>\n<blockquote>\n  <p>External data. We used Xeno Canto A-M and N-Z recording</p>\n</blockquote>\n<p>Interesting ! Do you know how much it helped your models in the end ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1208504,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/18/2021 10:00:40",
          "content": "<p>Significantly, all other params being equal we burned a sub testing impact:</p>\n<p><img src=\"https://i.imgur.com/PCyPUfA.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208513,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "02/18/2021 10:07:22",
          "content": "<p>That's a huge boost indeed!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208556,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 10:38:26",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a> and team on 7th place</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208657,
      "author_name": "riadalmadani",
      "author_url": "",
      "post_date": "02/18/2021 12:07:46",
      "content": "<p>Congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209117,
      "author_name": "dicksonchin93",
      "author_url": "",
      "post_date": "02/18/2021 17:34:48",
      "content": "<p>congrats! very impressive getting #7 in such a short time, I spent the whole 3 months failing most of the time 😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209340,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "02/18/2021 20:53:35",
          "content": "<p>Thanks!!! Long time since we did Avito!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209357,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/18/2021 21:08:34",
          "content": "<p>yeah! 3 years and you and <a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> are still killing it</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209381,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/18/2021 21:24:33",
      "content": "<p>Thanks for sharing.  Your approach is indeed very similar to mine, but you made it work better.  I wasn't aware of coordconv, it is certainly my next read.  </p>\n<p>Do you know how much you get from using external data? </p>\n<p>Congrats on the end result!</p>\n<p>Edit.  I see you answered the external data question already.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1225257,
      "author_name": "javiabellan",
      "author_url": "",
      "post_date": "03/03/2021 13:31:12",
      "content": "<p>Andres, do you think that using CoordConv based on sines and cosines (like Transformer Positional Encodings) would have improved the linspace CoordConv version?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1226258,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "03/04/2021 11:51:46",
          "content": "<p>Very nice question. I have been thinking about that I am not sure about the answer, I would need to test. But one thing for sure is that in that case the injection of coord matrixes should <em>also</em> be tested at the bottleneck, i.e. it makes sense to either insert positional not compressed and narrow (2ch) spatial encodings (similar to coordconv) or positional dense (512/1024/2048ch) encodings right before  final global average pooling. My suspicion the latter would be better but I have not tested (yet).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1226293,
      "author_name": "rodgomrod",
      "author_url": "",
      "post_date": "03/04/2021 12:20:28",
      "content": "<p>Muy buena, Andrés. Enhorabuena al equipo! 👌</p>\n<p>Gracias a estos posts aprendo un poco más de lo mejores día a día! ❤️</p>\n<p>Saludos! 😃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1322391,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "05/25/2021 12:30:56",
      "content": "<p>Thanks for the wonderful write-up! I have some rather generic questions (I hope it wont be too late to ask this):</p>\n<ol>\n<li>By masked loss, do u mean setting non-target classes to be <code>nan</code> (i.e. zero weight) as you did here?</li>\n</ol>\n<pre><code>tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))\n</code></pre>\n<ol>\n<li>Did u apply the same masking trick on FP data? (i.e. replace 1. by 0. for the tensor)</li>\n<li>I understand the rationale for applying masking on FP because these FP only implies the absence of a species, but implies no information wrt other species. However, what is the rationale for applying masking on TP data? </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1322611,
          "author_name": "antorsae",
          "author_url": "",
          "post_date": "05/25/2021 14:48:58",
          "content": "<p>Hi Alex,</p>\n<ol>\n<li>Yes.</li>\n<li>Yes.</li>\n<li>Rationale is the same: TP means species is present, we don't know anything else. FP means species is not present, we don't know anything else either; so everything is init as nan and populated w/ 1s and 0s with the information we know for sure.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1322616,
          "author_name": "alexlwh",
          "author_url": "",
          "post_date": "05/25/2021 14:51:20",
          "content": "<p>awesome! thanks so much for your reply<br>\nIt helps a lot</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1358768,
          "author_name": "alexlwh",
          "author_url": "",
          "post_date": "06/20/2021 17:58:47",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a><br>\nWould you mind me asking a few more questions?</p>\n<ol>\n<li>what is the image size that you used for model training?</li>\n<li>in pseudo labeling phase, you said you <code>manually balanced the resulting pseudo-label TPs</code>, would like to learn more about your sampling strategy. e.g. What is the sampling prob you set for TP, FP and pseudo labels?</li>\n<li>For pseudo labels, did you apply any penalty (e.g. soft penalty proportional to prediction confidence) on their loss?</li>\n<li>Did u generate pseudo labels by ensemble of ur OOF models?</li>\n<li>For model arch, if I understand correctly, it’s 2 channels, where the 2nd channel encode the positional info of the freq. Did u do any ablation study on the perf gain responsible by this “freq” channel? (Also I noticed the 2nd channel has quite a different scale that doesn’t range from 0 to 1. Do u think rescaling it back to 0 to 1 could further help?)</li>\n</ol>\n<p>Many thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1208493": "It was a great competition. I'll document our journey primarily from my perspective, my teammates @jpison @amezet @pavelgonchar may chime in to add extra color... \n\nOur solution is rank ensemble of two types of models, primarily 97% using the architecture described below, and 3% an ensemble of SED models. The architecture below was made in the last 7 days... I made my first (pretty bad) sub 8 days ago, and I joined team just before team merge deadline and all the ideas were done/implemented in ~7 days...\n\nA single 5-fold model achieves 0.961 private, 0.956 public; details:\n\n**Input representation.** This is probably the key, we take each TP or FP and after spectrogram (not MEL, b/c MEL was intended to model human audition and the rainforrest species have not evolved like our human audition) crop it time and frequency wise with fixed size on a per-class basis. E.g. for class 0: minimum freq is 5906.25 Hz and max is 8250 Hz, and time-wise we take the longest of time length of TPs, i.e. for class 0: 1.29s. \n\nWith the above sample the spectrogram to yield an image, e.g.:\n![](https://i.imgur.com/bTw0aeQ.png)\n\nThe GT for that TP would be: \n\n```tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))```\n\nAs other competitors, we expand the classes from 24 to 26 b/c to split species that have two songs. \n\nOther sample:\n\n![](https://i.imgur.com/JST401j.png)\n\n``` tensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))```\n\nNote the image size is the same, so time and frequency are effectively strectched wrt to what the net will see, but I believe this is fine as long as the net has enough receptive field (which they have). \n\nThe reason of doing it this way is that we need to inject the time and frequency restrictions as inductive bias somehow, and this looks like an nice way.\n\n**Model architecture.** This is just an image classifier outputing 26 logits, that's it. The only whistle is to add relative positional information in the freq axis (à la [coordconv](https://arxiv.org/abs/1807.03247)), so model is embarrasingly simple:\n\n```\nclass TropicModel(Module):\n    def __init__(self):\n        self.trunk = timm.create_model(a.arch,pretrained=True,num_classes=n_species,in_chans=1+a.coord)\n        self.do = nn.Dropout2d(a.do)\n    def forward(self,x):\n        bs,_,freq_bins,time_bins = x.size()\n        coord = torch.linspace(-1,1,freq_bins,dtype=x.dtype,device=x.device).view(1,1,-1,1).expand(bs,1,-1,time_bins)\n        if a.coord: x = torch.cat((x,coord),dim=1)\n        x = self.do(x)\n        return self.trunk(x)\n```\n\n**Loss function.** Just masked Focal loss. Actually this was a mistake b/c Focal loss was a renmant of a dead test and I (accidentally) left it there, where I thought (until I checked code now to write writeup) that BCE was being used. Since we are doing balancing BCE should work better.\n\n**Mixover.** Inspired by mixup, mixover takes a bunch of (unbalanced) TP and FPs (which strictly speaking are TNs) and creates combinations of them so that the resulting labels can be supervised (1+NaN=1, 0+NaN=NaN,0+0=0) making sure linear interpolation is not destructive (beta,alpha=4; clip to 0.2,0.8); then it samples from the resulted mixed items computing class distribution so that resulting samples are balanced. Code is a bit tricky, but still:\n\n```\nclass MixOver(MixHandler):\n    \"Inspired by implementation of https://arxiv.org/abs/1710.09412\"\n    def __init__(self, alpha=): super().__init__(alpha)\n    def before_batch(self):\n        ny_dims,nx_dims = len(self.y.size()),len(self.x.size())\n        bs=find_bs(self.xb)\n        all_combinations = L(itertools.combinations(range(find_bs(self.xb)), 2))\n        lam = self.distrib.sample((len(all_combinations),)).squeeze().to(self.x.device).clip(0.2,0.8)\n        lam = torch.stack([lam, 1-lam], 1)\n        self.lam = lam.max(1)[0]\n        comb = all_combinations\n        yb0,yb1 = L(self.yb).itemgot(comb.itemgot(0))[0],L(self.yb).itemgot(comb.itemgot(1))[0]\n        yb_one = torch.full_like(yb0,np.nan)\n        yb_one[yb0>0.5] = yb0[yb0>0.5]\n        yb_one[yb1>0.5] = yb1[yb1>0.5]\n        yb_two = torch.clip(yb0+yb1,0,1.)\n        yb_com = yb_one.clone()\n        yb_com[~torch.isnan(yb_two)] = yb_two[~torch.isnan(yb_two)]\n        n_ones_or_zeros=(~torch.isnan(yb_com)).sum()\n        ones=torch.sum(yb_com>=0.5,dim=1)\n        zeros=torch.sum(yb_com<0.5,dim=1)\n        p_ones = (n_ones_or_zeros/(2*( ones.sum())))/ones\n        p_zeros= (n_ones_or_zeros/(2*(zeros.sum())))/zeros\n        p_zeros[torch.isinf(p_zeros)],p_ones[torch.isinf(p_ones)]=0,0\n        p=(p_ones+p_zeros).cpu().numpy()/(p_ones+p_zeros).sum().item()\n        shuffle=torch.from_numpy(np.random.choice(yb_com.size(0),size=bs,replace=True,p=p)).to(self.x.device)\n        comb = all_combinations[shuffle]\n        xb0,xb1 = tuple(L(self.xb).itemgot(comb.itemgot(0))),tuple(L(self.xb).itemgot(comb.itemgot(1)))\n        self.learn.xb = tuple(L(xb0,xb1).map_zip(torch.lerp,weight=unsqueeze(self.lam[shuffle], n=nx_dims-1)))\n        self.learn.yb = (yb_com[shuffle],)\n```\n\n**Augmentations.** Time jitter (10% of time length), white noise (3 dB).\n\n**External data.** We used [Xeno Canto A-M and N-Z](https://www.kaggle.com/rohanrao/xeno-canto-bird-recordings-extended-a-m) recording and since they have 264 species we made the wild assumption that if we have 24 species and they have 10X randomly sampling would give you a \"right\" label (TN) 90% of the time... the goal was to add jungle/rainforest diversity at the expense of a small noisy labels which were accounted for labeling these weak labels as 0.1 (vs 0).\n\n**Pseudolabeling.** We also pseudolabeled using OOF models the non labeled parts of training data to mine more TPs, and manually balanced the resulting pseudo-label TPs. \n\n**Code.** Pytorch, Fastai.\n\n**Final thoughts.** It was a very fun competition and I am very glad of achieving a gold medal with a very short time, I want to thank Kaggle, sponsor and competitors for the competition and discussions; and of course my teammates @jpison @amezet @pavelgonchar for inviting me giving me the small nudge I needed to join 😉\n\nEdit: I've added no hand labels to the title because this competition was VERY unique in that Kaggle effectively allowed hand labeling and that's very unusual. (Re: hand labeling vs external data, it was also very uncommon not having to disclose which external data you used during competition).",
    "1208499": "Congratz !\n\n> External data. We used Xeno Canto A-M and N-Z recording\n\nInteresting ! Do you know how much it helped your models in the end ?",
    "1208504": "Significantly, all other params being equal we burned a sub testing impact:\n\n![](https://i.imgur.com/PCyPUfA.png)",
    "1208513": "That's a huge boost indeed!",
    "1208556": "Congrats @antorsae and team on 7th place",
    "1208657": "Congratulations",
    "1209117": "congrats! very impressive getting #7 in such a short time, I spent the whole 3 months failing most of the time 😂",
    "1209340": "Thanks!!! Long time since we did Avito!",
    "1209357": "yeah! 3 years and you and @pavelgonchar are still killing it",
    "1209381": "Thanks for sharing.  Your approach is indeed very similar to mine, but you made it work better.  I wasn't aware of coordconv, it is certainly my next read.  \n\nDo you know how much you get from using external data? \n\nCongrats on the end result!\n\nEdit.  I see you answered the external data question already.",
    "1225257": "Andres, do you think that using CoordConv based on sines and cosines (like Transformer Positional Encodings) would have improved the linspace CoordConv version?",
    "1226258": "Very nice question. I have been thinking about that I am not sure about the answer, I would need to test. But one thing for sure is that in that case the injection of coord matrixes should *also* be tested at the bottleneck, i.e. it makes sense to either insert positional not compressed and narrow (2ch) spatial encodings (similar to coordconv) or positional dense (512/1024/2048ch) encodings right before  final global average pooling. My suspicion the latter would be better but I have not tested (yet).",
    "1226293": "Muy buena, Andrés. Enhorabuena al equipo! 👌\n\nGracias a estos posts aprendo un poco más de lo mejores día a día! ❤️\n\nSaludos! 😃",
    "1322391": "Thanks for the wonderful write-up! I have some rather generic questions (I hope it wont be too late to ask this):\n1. By masked loss, do u mean setting non-target classes to be `nan` (i.e. zero weight) as you did here?\n```\ntensor([nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, nan, 1., nan, nan, nan, nan, nan,nan, nan, nan, nan, nan, nan, nan, nan]))\n```\n2. Did u apply the same masking trick on FP data? (i.e. replace 1. by 0. for the tensor)\n3. I understand the rationale for applying masking on FP because these FP only implies the absence of a species, but implies no information wrt other species. However, what is the rationale for applying masking on TP data?",
    "1322611": "Hi Alex,\n\n1. Yes.\n2. Yes.\n3. Rationale is the same: TP means species is present, we don't know anything else. FP means species is not present, we don't know anything else either; so everything is init as nan and populated w/ 1s and 0s with the information we know for sure.",
    "1322616": "awesome! thanks so much for your reply\nIt helps a lot",
    "1358768": "Hi @antorsae\nWould you mind me asking a few more questions?\n1. what is the image size that you used for model training?\n2. in pseudo labeling phase, you said you `manually balanced the resulting pseudo-label TPs`, would like to learn more about your sampling strategy. e.g. What is the sampling prob you set for TP, FP and pseudo labels?\n3. For pseudo labels, did you apply any penalty (e.g. soft penalty proportional to prediction confidence) on their loss?\n4. Did u generate pseudo labels by ensemble of ur OOF models?\n5. For model arch, if I understand correctly, it’s 2 channels, where the 2nd channel encode the positional info of the freq. Did u do any ablation study on the perf gain responsible by this “freq” channel? (Also I noticed the 2nd channel has quite a different scale that doesn’t range from 0 to 1. Do u think rescaling it back to 0 to 1 could further help?)\n\nMany thanks!"
  },
  "source": "meta"
}