{
  "id": 243360,
  "title": "My journey (11th solution, with code)",
  "url": "/competitions/birdclef-2021/discussion/243360",
  "author_name": "CPMP",
  "post_date": "2021-06-02T07:52:49.742000",
  "votes": 80,
  "comment_count": 60,
  "views": 0,
  "content": "<p>I decided to join early this competition because I was frustrated by the previous two bird song competitions.  In Cornell competition I joined too late. In Rainforest competition,  missed some key train/test distribution insights and stagnated in the LB after  a great start.</p>\n<p>The same 5 folds CV model (efficientnet b3 on first and last 5 seconds mel spectrograms) as in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183219\" target=\"_blank\">Cornell competition</a> plus <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">improvements from my Rainforest solution</a> gave 0.69 public LB (0.60 private LB) at my first submission.  This made me very happy as I took the lead on the public LB with it.</p>\n<p>The main difference with Cornell competition is that we were provided train soundscape.  As many I trained models on short audio records and tuned rounding thresholds on soundscape.  For CV I used the same as in my Cornell solution. Compute F1score on no call rows separately from f1 score on bird call rows, then compute final score with:</p>\n<p>score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</p>\n<p>This way a nocall sub CV is the same as the nocall submission LB.</p>\n<p>This made my CV and LB identical for most of my submissions.</p>\n<p>Tuning thresholds improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days.  CV and LB were identical so far.  At the time everyone else score was well below 0.70.</p>\n<p>This great start made me will to hide my score and I stopped submitting until someone matched my score.  It took 2 weeks.  At this time I submitted the same model bagged twice, i.e. the original  model averaged with a second model trained with a different seed.   This scored 0.78 on public LB (0.64 LB).  I knew it was lucky as CV was 0.76, but it looked great still. It took another 2 weeks for someone to beat this public score.  Later, after improving a bit the training procedure I got 0.80 public LB (0.66 private) with a 2 seeds x 5 folds submission of my baseline.</p>\n<p>It means that my baseline alone gives me a top 20 final rank.</p>\n<p>Given single models were so great I assumed ensembling would move me ahead further and I decided to focus on creating a wide range of individual models.  This is my main mistake, I should have worked on ensembling way earlier. </p>\n<p>I only submitted ensembles after the submission outage, 3 days before deadline, and discovered that it was hard to have a blending ensemble that beats all its individual component models.  I beat my best individual model only in my last submission.  It has both a CV and a public LB of 0.80 and is also my best private LB at 0.67.  </p>\n<p>I started a stacking model last day, but this was too late…  So be it.</p>\n<p>While I was holding top public score I decided to not submit and I explored lots of different models.  In particular I explored vision transformers.  I started with ViT and Deit.  First try with 384x384 mel spectrograms were disappointing.  Then I realized that i could use other image dimensions instead of a 24x24 grid of 16x16 patches.  I tried 12x48, i.e. a 192x768 spectrogram, and also a 16x36 grid (256x576 spectrogram).  The only trick is to modify the position embeddings to match the new grid dimension.</p>\n<p>This led to 0.75 public LB (0.64 private).</p>\n<p>But the most interesting one was to forget about square patches altogether.  A 196x576 spectrogram can be seen as 576 time slices of size 196.  Each slice contains 16x16 entries.  It means that I could just use the time slices as input patches.  Here is how this input looks once it is reshaped as a 24x24 grid of 16x16 patches:</p>\n<p><img src=\"https://i.imgur.com/F3zVGVM.pngd\" alt=\"time slices\"></p>\n<p>Maybe surprisingly, vision transformers are happy with this input.  The main advantage is that there is no longer any issue with translation on the frequency axis, which is the main issue with CNNs applied to spectrograms.</p>\n<p>Blending Deit trained on this input with my baseline gave my best sub.  The Deit model alone scores 0.77 on public LB and 0.66 on private LB.</p>\n<p>Although I am disappointed by my final result, I am happy to have explored lots of vision models and devised some new ways to use them.  And being disappointed by a solo gold in a deep learning competition is something I would not have imagined one year ago anyway ;)</p>\n<p>Special thanks to Ross Wightman for his timm package.  It made my model exploration seamless.  </p>\n<p>Edit. I shared code and the paper I submitted to the workshop: <a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">https://github.com/jfpuget/STFT_Transformer</a></p>",
  "messages": [
    {
      "id": 1332598,
      "postDate": "2021-06-02T07:52:49.743Z",
      "content": "<p>I decided to join early this competition because I was frustrated by the previous two bird song competitions.  In Cornell competition I joined too late. In Rainforest competition,  missed some key train/test distribution insights and stagnated in the LB after  a great start.</p>\n<p>The same 5 folds CV model (efficientnet b3 on first and last 5 seconds mel spectrograms) as in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183219\" target=\"_blank\">Cornell competition</a> plus <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">improvements from my Rainforest solution</a> gave 0.69 public LB (0.60 private LB) at my first submission.  This made me very happy as I took the lead on the public LB with it.</p>\n<p>The main difference with Cornell competition is that we were provided train soundscape.  As many I trained models on short audio records and tuned rounding thresholds on soundscape.  For CV I used the same as in my Cornell solution. Compute F1score on no call rows separately from f1 score on bird call rows, then compute final score with:</p>\n<p>score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</p>\n<p>This way a nocall sub CV is the same as the nocall submission LB.</p>\n<p>This made my CV and LB identical for most of my submissions.</p>\n<p>Tuning thresholds improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days.  CV and LB were identical so far.  At the time everyone else score was well below 0.70.</p>\n<p>This great start made me will to hide my score and I stopped submitting until someone matched my score.  It took 2 weeks.  At this time I submitted the same model bagged twice, i.e. the original  model averaged with a second model trained with a different seed.   This scored 0.78 on public LB (0.64 LB).  I knew it was lucky as CV was 0.76, but it looked great still. It took another 2 weeks for someone to beat this public score.  Later, after improving a bit the training procedure I got 0.80 public LB (0.66 private) with a 2 seeds x 5 folds submission of my baseline.</p>\n<p>It means that my baseline alone gives me a top 20 final rank.</p>\n<p>Given single models were so great I assumed ensembling would move me ahead further and I decided to focus on creating a wide range of individual models.  This is my main mistake, I should have worked on ensembling way earlier. </p>\n<p>I only submitted ensembles after the submission outage, 3 days before deadline, and discovered that it was hard to have a blending ensemble that beats all its individual component models.  I beat my best individual model only in my last submission.  It has both a CV and a public LB of 0.80 and is also my best private LB at 0.67.  </p>\n<p>I started a stacking model last day, but this was too late…  So be it.</p>\n<p>While I was holding top public score I decided to not submit and I explored lots of different models.  In particular I explored vision transformers.  I started with ViT and Deit.  First try with 384x384 mel spectrograms were disappointing.  Then I realized that i could use other image dimensions instead of a 24x24 grid of 16x16 patches.  I tried 12x48, i.e. a 192x768 spectrogram, and also a 16x36 grid (256x576 spectrogram).  The only trick is to modify the position embeddings to match the new grid dimension.</p>\n<p>This led to 0.75 public LB (0.64 private).</p>\n<p>But the most interesting one was to forget about square patches altogether.  A 196x576 spectrogram can be seen as 576 time slices of size 196.  Each slice contains 16x16 entries.  It means that I could just use the time slices as input patches.  Here is how this input looks once it is reshaped as a 24x24 grid of 16x16 patches:</p>\n<p><img src=\"https://i.imgur.com/F3zVGVM.pngd\" alt=\"time slices\"></p>\n<p>Maybe surprisingly, vision transformers are happy with this input.  The main advantage is that there is no longer any issue with translation on the frequency axis, which is the main issue with CNNs applied to spectrograms.</p>\n<p>Blending Deit trained on this input with my baseline gave my best sub.  The Deit model alone scores 0.77 on public LB and 0.66 on private LB.</p>\n<p>Although I am disappointed by my final result, I am happy to have explored lots of vision models and devised some new ways to use them.  And being disappointed by a solo gold in a deep learning competition is something I would not have imagined one year ago anyway ;)</p>\n<p>Special thanks to Ross Wightman for his timm package.  It made my model exploration seamless.  </p>\n<p>Edit. I shared code and the paper I submitted to the workshop: <a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">https://github.com/jfpuget/STFT_Transformer</a></p>",
      "rawMarkdown": "I decided to join early this competition because I was frustrated by the previous two bird song competitions.  In Cornell competition I joined too late. In Rainforest competition,  missed some key train/test distribution insights and stagnated in the LB after  a great start.\n\nThe same 5 folds CV model (efficientnet b3 on first and last 5 seconds mel spectrograms) as in [Cornell competition](https://www.kaggle.com/c/birdsong-recognition/discussion/183219) plus [improvements from my Rainforest solution](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304) gave 0.69 public LB (0.60 private LB) at my first submission.  This made me very happy as I took the lead on the public LB with it.\n\nThe main difference with Cornell competition is that we were provided train soundscape.  As many I trained models on short audio records and tuned rounding thresholds on soundscape.  For CV I used the same as in my Cornell solution. Compute F1score on no call rows separately from f1 score on bird call rows, then compute final score with:\n\nscore_all = 0.54 * score_nocall + (1 - 0.54) * score_birds\n\nThis way a nocall sub CV is the same as the nocall submission LB.\n\nThis made my CV and LB identical for most of my submissions.\n\nTuning thresholds improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days.  CV and LB were identical so far.  At the time everyone else score was well below 0.70.\n\nThis great start made me will to hide my score and I stopped submitting until someone matched my score.  It took 2 weeks.  At this time I submitted the same model bagged twice, i.e. the original  model averaged with a second model trained with a different seed.   This scored 0.78 on public LB (0.64 LB).  I knew it was lucky as CV was 0.76, but it looked great still. It took another 2 weeks for someone to beat this public score.  Later, after improving a bit the training procedure I got 0.80 public LB (0.66 private) with a 2 seeds x 5 folds submission of my baseline.\n\nIt means that my baseline alone gives me a top 20 final rank.\n\nGiven single models were so great I assumed ensembling would move me ahead further and I decided to focus on creating a wide range of individual models.  This is my main mistake, I should have worked on ensembling way earlier. \n\nI only submitted ensembles after the submission outage, 3 days before deadline, and discovered that it was hard to have a blending ensemble that beats all its individual component models.  I beat my best individual model only in my last submission.  It has both a CV and a public LB of 0.80 and is also my best private LB at 0.67.  \n\nI started a stacking model last day, but this was too late...  So be it.\n\nWhile I was holding top public score I decided to not submit and I explored lots of different models.  In particular I explored vision transformers.  I started with ViT and Deit.  First try with 384x384 mel spectrograms were disappointing.  Then I realized that i could use other image dimensions instead of a 24x24 grid of 16x16 patches.  I tried 12x48, i.e. a 192x768 spectrogram, and also a 16x36 grid (256x576 spectrogram).  The only trick is to modify the position embeddings to match the new grid dimension.\n\nThis led to 0.75 public LB (0.64 private).\n\nBut the most interesting one was to forget about square patches altogether.  A 196x576 spectrogram can be seen as 576 time slices of size 196.  Each slice contains 16x16 entries.  It means that I could just use the time slices as input patches.  Here is how this input looks once it is reshaped as a 24x24 grid of 16x16 patches:\n\n![time slices](https://i.imgur.com/F3zVGVM.pngd)\n\nMaybe surprisingly, vision transformers are happy with this input.  The main advantage is that there is no longer any issue with translation on the frequency axis, which is the main issue with CNNs applied to spectrograms.\n\nBlending Deit trained on this input with my baseline gave my best sub.  The Deit model alone scores 0.77 on public LB and 0.66 on private LB.\n\nAlthough I am disappointed by my final result, I am happy to have explored lots of vision models and devised some new ways to use them.  And being disappointed by a solo gold in a deep learning competition is something I would not have imagined one year ago anyway ;)\n\nSpecial thanks to Ross Wightman for his timm package.  It made my model exploration seamless.  \n\nEdit. I shared code and the paper I submitted to the workshop: https://github.com/jfpuget/STFT_Transformer",
      "votes": 80
    },
    {
      "id": 1350678,
      "postDate": "2021-06-15T16:36:44.573Z",
      "content": "<p>I shared code and the paper I submitted to the workshop: <a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">https://github.com/jfpuget/STFT_Transformer</a></p>",
      "rawMarkdown": "I shared code and the paper I submitted to the workshop: https://github.com/jfpuget/STFT_Transformer",
      "votes": 6,
      "replies": [
        {
          "id": 3197104,
          "postDate": "2025-05-07T20:15:59.817Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1334968,
      "postDate": "2021-06-03T23:58:06.960Z",
      "content": "<p>Congrats on the strong finish and solo gold. Interesting solution with ViT and Deit.</p>\n<pre><code>This great start made me will to hide my score and I stopped submitting until someone matched my score.\n</code></pre>\n<p>Does withholding submissions in this way really work in practice? I'm curious because I may be finding myself in this position with 2 months left in CommonLit Readability Prize. </p>",
      "rawMarkdown": "Congrats on the strong finish and solo gold. Interesting solution with ViT and Deit.\n\n```\nThis great start made me will to hide my score and I stopped submitting until someone matched my score.\n```\n\nDoes withholding submissions in this way really work in practice? I'm curious because I may be finding myself in this position with 2 months left in CommonLit Readability Prize. ",
      "votes": 1,
      "replies": [
        {
          "id": 1335747,
          "postDate": "2021-06-04T12:17:42.083Z",
          "content": "<p>Thanks.</p>\n<p>It was the first time I did this, and in hindsight I think it was a mistake.  It was a mistake because my real score never went above 0.80, which was the public score.</p>\n<p>The main issue is rather to stay motivated without being pressured by followers on the LB. I did my best but I see I made more progress during last week, when others had closed the LB gap with me, than in the previous 6 weeks.</p>",
          "rawMarkdown": "Thanks.\n\nIt was the first time I did this, and in hindsight I think it was a mistake.  It was a mistake because my real score never went above 0.80, which was the public score.\n\nThe main issue is rather to stay motivated without being pressured by followers on the LB. I did my best but I see I made more progress during last week, when others had closed the LB gap with me, than in the previous 6 weeks.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1334647,
      "postDate": "2021-06-03T16:41:19.640Z",
      "content": "<p>Congrats on solo gold medal. <br>\nIt can be disappointing, but the Solo Gold medal is still great.</p>\n<p>And the part about DieT and ViT is very interesting. </p>\n<p>Thanks for sharing :) </p>\n<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
      "rawMarkdown": "Congrats on solo gold medal. \nIt can be disappointing, but the Solo Gold medal is still great.\n\nAnd the part about DieT and ViT is very interesting. \n\nThanks for sharing :) \n\n@cpmpml ",
      "votes": 1,
      "replies": [
        {
          "id": 1334672,
          "postDate": "2021-06-03T17:17:48.393Z",
          "content": "<p>Thanks, I spent quite some time with vision transformers to make them work here.</p>",
          "rawMarkdown": "Thanks, I spent quite some time with vision transformers to make them work here."
        }
      ]
    },
    {
      "id": 1334542,
      "postDate": "2021-06-03T15:26:02.167Z",
      "content": "<p>Congrats on the great result !! I for one think that the cross validation scheme played a big role in the modeling performance, I used the R blockCV package to generate my folds.</p>\n<p>Curious:</p>\n<p>1) did you use random kfold?<br>\n2) for mixup did you incorporate the secondary labels or did you have to one hot encode the target with just one primary label?</p>\n<p>Thanks for being such an active discussion contributor hope to compete in similar competitions going forward!</p>",
      "rawMarkdown": "Congrats on the great result !! I for one think that the cross validation scheme played a big role in the modeling performance, I used the R blockCV package to generate my folds.\n\nCurious:\n\n1) did you use random kfold?\n2) for mixup did you incorporate the secondary labels or did you have to one hot encode the target with just one primary label?\n\nThanks for being such an active discussion contributor hope to compete in similar competitions going forward!",
      "votes": 1,
      "replies": [
        {
          "id": 1334590,
          "postDate": "2021-06-03T15:59:10.213Z",
          "content": "<p>Thanks.  </p>\n<p>1) I used stratified k fold with shuffle.</p>\n<p>2) For mixup have a look at my Cornell writeup, I provide all details there. </p>",
          "rawMarkdown": "Thanks.  \n\n1) I used stratified k fold with shuffle.\n\n2) For mixup have a look at my Cornell writeup, I provide all details there. "
        }
      ]
    },
    {
      "id": 1333984,
      "postDate": "2021-06-03T07:30:19.500Z",
      "content": "<p>Very interesting solution and insightful use of vision transformers. Congratulations</p>",
      "rawMarkdown": "Very interesting solution and insightful use of vision transformers. Congratulations",
      "votes": 1,
      "replies": [
        {
          "id": 1334335,
          "postDate": "2021-06-03T12:32:51.617Z",
          "content": "<p>Thanks!  We missed you in this comp.</p>",
          "rawMarkdown": "Thanks!  We missed you in this comp."
        }
      ]
    },
    {
      "id": 1333696,
      "postDate": "2021-06-03T01:16:19.123Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> congratulation on solo gold! though the results are a little bit rough( <br>\nThank you for sharing the results about Deit, very interesting insight.</p>",
      "rawMarkdown": "@cpmpml congratulation on solo gold! though the results are a little bit rough( \nThank you for sharing the results about Deit, very interesting insight.",
      "votes": 1,
      "replies": [
        {
          "id": 1333944,
          "postDate": "2021-06-03T07:05:10.423Z",
          "content": "<p>Thanks. Yes, I am quite proud of how I used vision transformers.</p>",
          "rawMarkdown": "Thanks. Yes, I am quite proud of how I used vision transformers.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1333288,
      "postDate": "2021-06-02T15:53:42.343Z",
      "content": "<p>Congrats Oncle!<br>\nI was rooting for you !</p>",
      "rawMarkdown": "Congrats Oncle!\nI was rooting for you !",
      "votes": 1
    },
    {
      "id": 1333254,
      "postDate": "2021-06-02T15:28:35.833Z",
      "content": "<p>Congrats on solo gold! </p>\n<blockquote>\n  <p>0.69 public LB (0.60 private LB) at my first submission</p>\n</blockquote>\n<p>I was overwhelmed by your strong start.</p>\n<blockquote>\n  <p>The Deit model alone scores 0.77 on public LB and 0.66 on private LB.</p>\n</blockquote>\n<p>This is very interesting result. I guess you've found very good techniques you would try in the first place when you face the similar problem statement.</p>",
      "rawMarkdown": "Congrats on solo gold! \n\n> 0.69 public LB (0.60 private LB) at my first submission\n\nI was overwhelmed by your strong start.\n\n> The Deit model alone scores 0.77 on public LB and 0.66 on private LB.\n\nThis is very interesting result. I guess you've found very good techniques you would try in the first place when you face the similar problem statement.",
      "votes": 1,
      "replies": [
        {
          "id": 1333273,
          "postDate": "2021-06-02T15:41:37.187Z",
          "content": "<p>Thanks.  Yes, I spend a LOT of time tuning ViT/Deit training and I am very happy to have succeeded somewhat.  It is why I am not too disappointed by the end result ;)</p>\n<p>I am now convinced that training only on 5 second clips is not the best, and next time I'll certainly have  a look at your SED model ;)</p>\n<p>And congrats on becoming GM soon!</p>",
          "rawMarkdown": "Thanks.  Yes, I spend a LOT of time tuning ViT/Deit training and I am very happy to have succeeded somewhat.  It is why I am not too disappointed by the end result ;)\n\nI am now convinced that training only on 5 second clips is not the best, and next time I'll certainly have  a look at your SED model ;)\n\nAnd congrats on becoming GM soon!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1332776,
      "postDate": "2021-06-02T09:41:07.483Z",
      "content": "<p>Congratulations. Your solution from previous Birdcall competition helped us a lot. Thank you. 😀</p>",
      "rawMarkdown": "Congratulations. Your solution from previous Birdcall competition helped us a lot. Thank you. 😀",
      "votes": 1
    },
    {
      "id": 1332700,
      "postDate": "2021-06-02T08:48:35.060Z",
      "content": "<p>Congrats on solo gold!</p>\n<p>Could you share your baseline model weights and exact mel spectrogram settings? Your solution is almost fully opposite to mine - extremely strong single models, but i do not see any postprocessing.</p>\n<p>It would be really interesting to combine both and see if strong model still benefits from postprocessing as much as weak one.</p>",
      "rawMarkdown": "Congrats on solo gold!\n\nCould you share your baseline model weights and exact mel spectrogram settings? Your solution is almost fully opposite to mine - extremely strong single models, but i do not see any postprocessing.\n\nIt would be really interesting to combine both and see if strong model still benefits from postprocessing as much as weak one.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1332740,
          "postDate": "2021-06-02T09:14:28.727Z",
          "content": "<p>Which is the strong model and the weak model?</p>\n<p>Anyway, I give details about my model in my Cornell writeup, link in the post.</p>",
          "rawMarkdown": "Which is the strong model and the weak model?\n\nAnyway, I give details about my model in my Cornell writeup, link in the post.\n\n"
        },
        {
          "id": 1332762,
          "postDate": "2021-06-02T09:27:52.940Z",
          "content": "<p>Models with a high score would be a strong models. Yours is almost 0.1 higher than mine on private LB.</p>\n<p>Details would sadly not help much, i cant train 300x600 model for 60 epochs on this amount of data in any reasonable time:(</p>",
          "rawMarkdown": "Models with a high score would be a strong models. Yours is almost 0.1 higher than mine on private LB.\n\nDetails would sadly not help much, i cant train 300x600 model for 60 epochs on this amount of data in any reasonable time:(\n\n",
          "votes": 1
        },
        {
          "id": 1332782,
          "postDate": "2021-06-02T09:47:56.803Z",
          "content": "<p>;)  You made me share my postprocessing: <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243378\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/243378</a></p>",
          "rawMarkdown": ";)  You made me share my postprocessing: https://www.kaggle.com/c/birdclef-2021/discussion/243378",
          "votes": 1
        },
        {
          "id": 1332854,
          "postDate": "2021-06-02T10:55:48.667Z",
          "content": "<p>Thank you. I know you are disappointed at the shakeup, but still, congrats on the solo gold, and I'd also like to think you for all the help you provided in discussions throughout the competition.</p>",
          "rawMarkdown": "Thank you. I know you are disappointed at the shakeup, but still, congrats on the solo gold, and I'd also like to think you for all the help you provided in discussions throughout the competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1332629,
      "postDate": "2021-06-02T08:13:06.827Z",
      "content": "<p>Congratulations on the second solo gold! I think you helped me stay competitive. <br>\nI am impressing by your deep consideration of Vision Transformer. I've focused on inference and post-processing, so thank you for sharing!</p>",
      "rawMarkdown": "Congratulations on the second solo gold! I think you helped me stay competitive. \nI am impressing by your deep consideration of Vision Transformer. I've focused on inference and post-processing, so thank you for sharing!",
      "votes": 1
    },
    {
      "id": 1332623,
      "postDate": "2021-06-02T08:11:01.913Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ! I couldn't find it, did you train on whole clips or x seconds clips? If so, how did you set the target for them?</p>\n<blockquote>\n  <p>The main difference with Cornell competition is that we were provided train soundscape. As many I trained models on short audio records and tuned rounding thresholds on soundscape. This improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days. CV and LB were identical so far. At the time everyone else score was well below 0.70.</p>\n</blockquote>\n<p>I think this probably caused your shake-down. I had the same tuning which improved my CV to 0.74 but made my LB 0.64 or so. Then I noticed it is only 17 soundscapes in training set and around 30 in Public LB. This was the reason I expected a shake-up. People can't notice it when they get lucky with the submission.</p>",
      "rawMarkdown": "Congrats @cpmpml ! I couldn't find it, did you train on whole clips or x seconds clips? If so, how did you set the target for them?\n\n> The main difference with Cornell competition is that we were provided train soundscape. As many I trained models on short audio records and tuned rounding thresholds on soundscape. This improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days. CV and LB were identical so far. At the time everyone else score was well below 0.70.\n\nI think this probably caused your shake-down. I had the same tuning which improved my CV to 0.74 but made my LB 0.64 or so. Then I noticed it is only 17 soundscapes in training set and around 30 in Public LB. This was the reason I expected a shake-up. People can't notice it when they get lucky with the submission.",
      "votes": 1,
      "replies": [
        {
          "id": 1332630,
          "postDate": "2021-06-02T08:13:26.750Z",
          "content": "<p>I wrote it ;)</p>\n<blockquote>\n  <p>The same 5 folds CV model (efficientnet b3 on 5 seconds mel spectrograms)</p>\n</blockquote>",
          "rawMarkdown": "I wrote it ;)\n\n> The same 5 folds CV model (efficientnet b3 on 5 seconds mel spectrograms)\n\n",
          "votes": 1
        },
        {
          "id": 1332638,
          "postDate": "2021-06-02T08:17:37.587Z",
          "content": "<p>What was your guys highest CV on all soundscape recordings? Ours was close to 0.84 but we decided to evaluate a bit differently (will write it in the solution post).</p>",
          "rawMarkdown": "What was your guys highest CV on all soundscape recordings? Ours was close to 0.84 but we decided to evaluate a bit differently (will write it in the solution post)."
        },
        {
          "id": 1332639,
          "postDate": "2021-06-02T08:17:50.663Z",
          "content": "<p>Thanks, I somehow missed it. And how did you set the targets for them? Some of the 5 sec clips don't have the primary and secondary labels, right?</p>",
          "rawMarkdown": "Thanks, I somehow missed it. And how did you set the targets for them? Some of the 5 sec clips don't have the primary and secondary labels, right?"
        },
        {
          "id": 1332641,
          "postDate": "2021-06-02T08:20:45.327Z",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> That's what also puzzles me a bit: <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243324#1332617\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/243324#1332617</a></p>",
          "rawMarkdown": "@aerdem4 That's what also puzzles me a bit: https://www.kaggle.com/c/birdclef-2021/discussion/243324#1332617"
        },
        {
          "id": 1332644,
          "postDate": "2021-06-02T08:22:14.700Z",
          "content": "<blockquote>\n  <p>What was your guys highest CV on all soundscape recordings? </p>\n</blockquote>\n<p>0.80, same as LB</p>\n<p>For CV I used the same as in my Cornell solution.  Compute F1score on no call rows separately fromF1 score on bird call rows, then compute final score with:</p>\n<p><code>score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</code></p>\n<p>This way a nocall sub CV is the same as the nocall submission LB.</p>\n<p>This made my CV and LB identical for most of my submissions.</p>\n<p>Updating my post as this is important.</p>",
          "rawMarkdown": "> What was your guys highest CV on all soundscape recordings? \n\n0.80, same as LB\n\nFor CV I used the same as in my Cornell solution.  Compute F1score on no call rows separately fromF1 score on bird call rows, then compute final score with:\n\n`score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds`\n\nThis way a nocall sub CV is the same as the nocall submission LB.\n\nThis made my CV and LB identical for most of my submissions.\n\nUpdating my post as this is important.",
          "votes": 2
        },
        {
          "id": 1332654,
          "postDate": "2021-06-02T08:28:27.453Z",
          "content": "<p>Thanks - got it. Private LB has 0.47 sample submission btw.</p>",
          "rawMarkdown": "Thanks - got it. Private LB has 0.47 sample submission btw.",
          "votes": 1
        },
        {
          "id": 1332673,
          "postDate": "2021-06-02T08:34:43.047Z",
          "content": "<blockquote>\n  <p>Private LB has 0.47 sample submission btw.</p>\n</blockquote>\n<p>I am too lazy to check, but I bet that if I change my formula to this then Cv and private LB match:</p>\n<p><code>score_all = 0.47* score_nocall + (1 - 0.47) * score_birds</code></p>",
          "rawMarkdown": "> Private LB has 0.47 sample submission btw.\n\nI am too lazy to check, but I bet that if I change my formula to this then Cv and private LB match:\n\n`score_all = 0.47* score_nocall + (1 - 0.47) * score_birds`",
          "votes": 1
        },
        {
          "id": 1332676,
          "postDate": "2021-06-02T08:36:54.973Z",
          "content": "<blockquote>\n  <p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ! I couldn't find it, did you train on whole clips or x seconds clips? </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> thanks.  I updated the post to make it clearer.  I used first and last 5 seconds of short train audio.</p>",
          "rawMarkdown": "> Congrats @cpmpml ! I couldn't find it, did you train on whole clips or x seconds clips? \n\n@aerdem4 thanks.  I updated the post to make it clearer.  I used first and last 5 seconds of short train audio.",
          "votes": 1
        },
        {
          "id": 1332921,
          "postDate": "2021-06-02T11:37:17.210Z",
          "content": "<blockquote>\n  <p>What was your guys highest CV on all soundscape recordings? </p>\n</blockquote>\n<p>F1 on SC: <code>0.824</code> / Public <code>0.78x</code> / Private <code>0.66x</code></p>",
          "rawMarkdown": "> What was your guys highest CV on all soundscape recordings? \n\nF1 on SC: `0.824` / Public `0.78x` / Private `0.66x`",
          "votes": 1
        },
        {
          "id": 1333033,
          "postDate": "2021-06-02T12:52:57.923Z",
          "content": "<blockquote>\n  <p>… compute final score with: score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</p>\n</blockquote>\n<p>Used same idea but with: <code>0.56* score_nocall + (1 - 0.56) * score_birds</code></p>\n<p>and up to some range correlated very well with <strong>Public LB</strong>… for eg. </p>\n<p><code>Public 0.78x</code> / <code>Private 0.66x</code> --&gt; est <code>LB: 0.7876</code> | <code>nocall 0.90974</code>, <code>bird 0.6322</code></p>",
          "rawMarkdown": "> ... compute final score with: score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds\n\nUsed same idea but with: `0.56* score_nocall + (1 - 0.56) * score_birds`\n\nand up to some range correlated very well with **Public LB**... for eg. \n\n`Public 0.78x` / `Private 0.66x` --> est `LB: 0.7876` | `nocall 0.90974`, `bird 0.6322`"
        },
        {
          "id": 1333322,
          "postDate": "2021-06-02T16:22:00.950Z",
          "content": "<blockquote>\n  <p>I think this probably caused your shake-down.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> why is that?  Everyone used soundscape to tune rounding thresholds.  What is specific to my case that would be worse than for other teams?</p>\n<p>I am really curious, you maybe right but I don't get why.</p>",
          "rawMarkdown": "> I think this probably caused your shake-down.\n\n@aerdem4 why is that?  Everyone used soundscape to tune rounding thresholds.  What is specific to my case that would be worse than for other teams?\n\nI am really curious, you maybe right but I don't get why."
        },
        {
          "id": 1333324,
          "postDate": "2021-06-02T16:23:26.260Z",
          "content": "<blockquote>\n  <p>Used same idea but with: 0.56* score_nocall + (1 - 0.56) * score_birds</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> why 0.56?  If you predict only no call it would be wrong.</p>",
          "rawMarkdown": "> Used same idea but with: 0.56* score_nocall + (1 - 0.56) * score_birds\n\n@imeintanis why 0.56?  If you predict only no call it would be wrong."
        },
        {
          "id": 1334079,
          "postDate": "2021-06-03T09:18:01.250Z",
          "content": "<blockquote>\n  <p>why is that?  Everyone used soundscape to tune rounding thresholds. What is specific to my case that would be worse than for other teams?</p>\n</blockquote>\n<p>Did you use soundscape level predictions too? For me, I used the prediction for 5 sec clip and total sum for thresholding. By this method, I was improving my local validation a lot but LB was going down. Then I notice thresholds were sensitive to folds etc. Then I thought people who got lucky on LB (probably half of the people) may not notice that there is high variance.</p>",
          "rawMarkdown": "> why is that?  Everyone used soundscape to tune rounding thresholds. What is specific to my case that would be worse than for other teams?\n\nDid you use soundscape level predictions too? For me, I used the prediction for 5 sec clip and total sum for thresholding. By this method, I was improving my local validation a lot but LB was going down. Then I notice thresholds were sensitive to folds etc. Then I thought people who got lucky on LB (probably half of the people) may not notice that there is high variance."
        },
        {
          "id": 1334228,
          "postDate": "2021-06-03T10:50:05.847Z",
          "content": "<blockquote>\n  <p>Did you use soundscape level predictions too? </p>\n</blockquote>\n<p>like everyone else.</p>\n<p>People tuned thresholds using model predictions on train soundscape.</p>\n<p>In your case, the issue is probably that you did not correct for the different proportion of no calls.  Compute F1score on no call rows separately fromF1 score on bird call rows, then compute final score with:</p>\n<p>score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</p>\n<p>If you compute CV this way then CV and public LB are almost identical.</p>\n<p>I still don't get why you say it is luck.</p>",
          "rawMarkdown": "> Did you use soundscape level predictions too? \n\nlike everyone else.\n\nPeople tuned thresholds using model predictions on train soundscape.\n\nIn your case, the issue is probably that you did not correct for the different proportion of no calls.  Compute F1score on no call rows separately fromF1 score on bird call rows, then compute final score with:\n\nscore_all = 0.54 * score_nocall + (1 - 0.54) * score_birds\n\nIf you compute CV this way then CV and public LB are almost identical.\n\nI still don't get why you say it is luck."
        },
        {
          "id": 1334587,
          "postDate": "2021-06-03T15:57:29.053Z",
          "content": "<p>I used 3 soundscapes with no bird for background augmentation and validated my model on 17 landscapes. Simple thresholded model had 0.68 CV and 0.68 LB. When I included soundscape level thresholding, CV improved to 0.74 but LB went down to 0.64 or so. Then I repeat the same experiment with re-training on different folds etc and noticed that the score is very sensitive when soundscape level thresholding is involved. If I got lucky on LB, I wouldn't notice it.</p>",
          "rawMarkdown": "I used 3 soundscapes with no bird for background augmentation and validated my model on 17 landscapes. Simple thresholded model had 0.68 CV and 0.68 LB. When I included soundscape level thresholding, CV improved to 0.74 but LB went down to 0.64 or so. Then I repeat the same experiment with re-training on different folds etc and noticed that the score is very sensitive when soundscape level thresholding is involved. If I got lucky on LB, I wouldn't notice it."
        },
        {
          "id": 1334595,
          "postDate": "2021-06-03T16:02:47.053Z",
          "content": "<p>Still what is specific to me here?  All top team used soundscape to tune thresholds.</p>\n<p>Given you used 3 soundscape data you are fitting to it, and relying on other soundscape looks leaky to me.  </p>",
          "rawMarkdown": "Still what is specific to me here?  All top team used soundscape to tune thresholds.\n\nGiven you used 3 soundscape data you are fitting to it, and relying on other soundscape looks leaky to me.  "
        },
        {
          "id": 1334601,
          "postDate": "2021-06-03T16:06:01.053Z",
          "content": "<blockquote>\n  <p>Given you used 3 soundscape data you are fitting to it, and relying on other soundscape looks leaky to me.</p>\n</blockquote>\n<p>It would be leaky if soundscapes were having overlaps.</p>\n<blockquote>\n  <p>Still what is specific to me here? All top team used soundscape to tune thresholds.</p>\n</blockquote>\n<p>Nothing specific to you:) My guess was that approach was a coin flip and you were one of the many people with positive flip.</p>",
          "rawMarkdown": "> Given you used 3 soundscape data you are fitting to it, and relying on other soundscape looks leaky to me.\n\nIt would be leaky if soundscapes were having overlaps.\n\n> Still what is specific to me here? All top team used soundscape to tune thresholds.\n\nNothing specific to you:) My guess was that approach was a coin flip and you were one of the many people with positive flip."
        },
        {
          "id": 1334659,
          "postDate": "2021-06-03T16:51:00.727Z",
          "content": "<p>Why would all my subs be lucky then?</p>\n<p>How come my best CV is best on both public and private LB?</p>\n<p>I don't think it has to do with luck to be honest.</p>\n<p>Let's agree to disagree ;)</p>\n<p>My hunch is that private LB has less nocalls, hence I probably have optimized for a too high proportion of no calls.  Maybe that's what you want to say afterall?</p>",
          "rawMarkdown": "Why would all my subs be lucky then?\n\nHow come my best CV is best on both public and private LB?\n\nI don't think it has to do with luck to be honest.\n\nLet's agree to disagree ;)\n\nMy hunch is that private LB has less nocalls, hence I probably have optimized for a too high proportion of no calls.  Maybe that's what you want to say afterall?"
        }
      ]
    },
    {
      "id": 1334181,
      "postDate": "2021-06-03T10:00:06.553Z",
      "content": "<p>Congratulations. </p>\n<blockquote>\n  <p>The only trick is to modify the position embeddings to match the new grid dimension.</p>\n</blockquote>\n<p>Could you explain this further?</p>",
      "rawMarkdown": "Congratulations. \n> The only trick is to modify the position embeddings to match the new grid dimension.\n\nCould you explain this further?",
      "votes": 2,
      "replies": [
        {
          "id": 1334339,
          "postDate": "2021-06-03T12:35:03.063Z",
          "content": "<p>ViT and DeiT have position embeddings on a 24x24 grid.  I interpolated them to the new grid, for instance to a 16x36 grid.</p>\n<p>For this I modified timm's implementation of ViT and Deit position embeddings as follows.</p>\n<pre><code>def resize_pos_embed(posemb, new_size, num_tokens=1):\n    # Rescale the grid of position embeddings when loading from state_dict. Adapted from\n    # https://github.com/google-research/vision_transformer/blob/00883dd691c63a6830751563748663526e811cee/vit_jax/checkpoint.py#L224\n    if num_tokens:\n        posemb_tok, posemb_grid = posemb[:, :num_tokens], posemb[0, num_tokens:]\n    else:\n        posemb_tok, posemb_grid = posemb[:, :0], posemb[0]\n    gs_old = int(math.sqrt(len(posemb_grid)))\n    posemb_grid = posemb_grid.reshape(1, gs_old, gs_old, -1).permute(0, 3, 1, 2)\n    posemb_grid = F.interpolate(posemb_grid, size=new_size, mode='bilinear', align_corners=False)\n    posemb_grid = posemb_grid.permute(0, 2, 3, 1).reshape(1, new_size[0] * new_size[1], -1)\n    posemb = torch.cat([posemb_tok, posemb_grid], dim=1)\n    return posemb\n</code></pre>",
          "rawMarkdown": "ViT and DeiT have position embeddings on a 24x24 grid.  I interpolated them to the new grid, for instance to a 16x36 grid.\n\nFor this I modified timm's implementation of ViT and Deit position embeddings as follows.\n\n```\ndef resize_pos_embed(posemb, new_size, num_tokens=1):\n    # Rescale the grid of position embeddings when loading from state_dict. Adapted from\n    # https://github.com/google-research/vision_transformer/blob/00883dd691c63a6830751563748663526e811cee/vit_jax/checkpoint.py#L224\n    if num_tokens:\n        posemb_tok, posemb_grid = posemb[:, :num_tokens], posemb[0, num_tokens:]\n    else:\n        posemb_tok, posemb_grid = posemb[:, :0], posemb[0]\n    gs_old = int(math.sqrt(len(posemb_grid)))\n    posemb_grid = posemb_grid.reshape(1, gs_old, gs_old, -1).permute(0, 3, 1, 2)\n    posemb_grid = F.interpolate(posemb_grid, size=new_size, mode='bilinear', align_corners=False)\n    posemb_grid = posemb_grid.permute(0, 2, 3, 1).reshape(1, new_size[0] * new_size[1], -1)\n    posemb = torch.cat([posemb_tok, posemb_grid], dim=1)\n    return posemb\n\n```",
          "votes": 3
        },
        {
          "id": 1335281,
          "postDate": "2021-06-04T06:06:25.683Z",
          "content": "<p>Thanks for sharing. Congrats! <br>\nI could read your disappointment. 💪</p>",
          "rawMarkdown": "Thanks for sharing. Congrats! \nI could read your disappointment. 💪",
          "votes": 1
        }
      ]
    },
    {
      "id": 1333017,
      "postDate": "2021-06-02T12:39:25.433Z",
      "content": "<p>Congrats ! Birds got a real human buddy now !!</p>",
      "rawMarkdown": "Congrats ! Birds got a real human buddy now !!",
      "votes": 2
    },
    {
      "id": 1332661,
      "postDate": "2021-06-02T08:30:01.480Z",
      "content": "<p>Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work. Thanks for the write-up! Would be nice to see your approach as a working note :) </p>",
      "rawMarkdown": "Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work. Thanks for the write-up! Would be nice to see your approach as a working note :) ",
      "votes": 2,
      "replies": [
        {
          "id": 1332670,
          "postDate": "2021-06-02T08:32:49.410Z",
          "content": "<p>Thanks!.  Training ViT was tricky as they are quite unstable.  Maybe I'll focus on this part in my note as the rest is just fine tuning previous solutions.</p>",
          "rawMarkdown": "Thanks!.  Training ViT was tricky as they are quite unstable.  Maybe I'll focus on this part in my note as the rest is just fine tuning previous solutions.",
          "votes": 2
        },
        {
          "id": 1332684,
          "postDate": "2021-06-02T08:39:49.610Z",
          "content": "<p>I bet! I also tried different patching methods and I think what you did makes most sense for spectrograms that are by default a sequential input. Rearranging the patches to use pre-trained models is what I was missing :)</p>",
          "rawMarkdown": "I bet! I also tried different patching methods and I think what you did makes most sense for spectrograms that are by default a sequential input. Rearranging the patches to use pre-trained models is what I was missing :)",
          "votes": 2
        },
        {
          "id": 1332687,
          "postDate": "2021-06-02T08:41:28.527Z",
          "content": "<p>Yeah, I thought that skipping the stem convolution and directly input the time slices as patch embeddings would be hard to train from scratch because training data is small. Reusing pretrained wieghts was safer.</p>",
          "rawMarkdown": "Yeah, I thought that skipping the stem convolution and directly input the time slices as patch embeddings would be hard to train from scratch because training data is small. Reusing pretrained wieghts was safer."
        },
        {
          "id": 1332800,
          "postDate": "2021-06-02T10:01:19.290Z",
          "content": "<p>I had submitted my deit model actually, alone it scores 0.77/0.66 which is not too bad.</p>",
          "rawMarkdown": "I had submitted my deit model actually, alone it scores 0.77/0.66 which is not too bad.",
          "votes": 1
        },
        {
          "id": 1332852,
          "postDate": "2021-06-02T10:54:07.413Z",
          "content": "<blockquote>\n  <p>Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work.</p>\n</blockquote>\n<p>That's a relief, I thought I was the only one incapable of making a ViT work. My results on the training soundscapes were so poor I ended up not making a submission at all…maybe next year 😄</p>",
          "rawMarkdown": "> Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work.\n\nThat's a relief, I thought I was the only one incapable of making a ViT work. My results on the training soundscapes were so poor I ended up not making a submission at all...maybe next year 😄",
          "votes": 1
        },
        {
          "id": 1341233,
          "postDate": "2021-06-08T14:22:20.427Z",
          "content": "<blockquote>\n  <p>Maybe I'll focus on this part in my note as the rest is just fine tuning previous solutions.</p>\n</blockquote>\n<p>Looking forward to your working notes paper! Really cool that you got transformers to work so well.</p>\n<blockquote>\n  <p>Reusing pretrained wieghts was safer.</p>\n</blockquote>\n<p>I'd be curious how it compares to training from scratch.</p>",
          "rawMarkdown": "> Maybe I'll focus on this part in my note as the rest is just fine tuning previous solutions.\n\nLooking forward to your working notes paper! Really cool that you got transformers to work so well.\n\n> Reusing pretrained wieghts was safer.\n\nI'd be curious how it compares to training from scratch.",
          "votes": 1
        },
        {
          "id": 1341307,
          "postDate": "2021-06-08T15:18:34.057Z",
          "content": "<blockquote>\n  <p>I'd be curious how it compares to training from scratch.</p>\n</blockquote>\n<p>Me too actually, but I didn't felt motivated enough to run the experiment.  I may do it now that dust settled a bit.</p>",
          "rawMarkdown": "> I'd be curious how it compares to training from scratch.\n\nMe too actually, but I didn't felt motivated enough to run the experiment.  I may do it now that dust settled a bit."
        }
      ]
    },
    {
      "id": 1332622,
      "postDate": "2021-06-02T08:09:58.507Z",
      "content": "<p>Great writeup , congrats on solo gold .  It's great to see vision transformers work , we also tried it with square image and it was below par , now I get what was wrong , I didn't know about the different grids we can use. Coming early in the competition was the trick we missed , this made us experiment with Vit as black box only and I was not able to discover it closely (which is something I rarely do) .</p>\n<p>We also tried taking 7 second crops from the start and the end like you did in cornell but somehow it gave worse results , did you do the same in this comp?<br>\nAlso it would be great if you can give us some insights to which models you use and their training strategy</p>",
      "rawMarkdown": "Great writeup , congrats on solo gold .  It's great to see vision transformers work , we also tried it with square image and it was below par , now I get what was wrong , I didn't know about the different grids we can use. Coming early in the competition was the trick we missed , this made us experiment with Vit as black box only and I was not able to discover it closely (which is something I rarely do) .\n\nWe also tried taking 7 second crops from the start and the end like you did in cornell but somehow it gave worse results , did you do the same in this comp?\nAlso it would be great if you can give us some insights to which models you use and their training strategy",
      "votes": 2,
      "replies": [
        {
          "id": 1332633,
          "postDate": "2021-06-02T08:14:39.067Z",
          "content": "<blockquote>\n  <p>you did in cornell but somehow it gave worse results , did you do the same in this comp?</p>\n</blockquote>\n<p>Yes, as I wrote, i used the same approach  as in Cornell.</p>",
          "rawMarkdown": ">  you did in cornell but somehow it gave worse results , did you do the same in this comp?\n\nYes, as I wrote, i used the same approach  as in Cornell.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1392674,
      "postDate": "2021-07-19T01:13:30.050Z",
      "content": "<p>Amazing research on vision transformer. Thank you!</p>",
      "rawMarkdown": "Amazing research on vision transformer. Thank you!"
    },
    {
      "id": 1332779,
      "postDate": "2021-06-02T09:44:26.077Z",
      "content": "<p>Great job! Mind if you share your code? </p>",
      "rawMarkdown": "Great job! Mind if you share your code? ",
      "replies": [
        {
          "id": 1332783,
          "postDate": "2021-06-02T09:48:36.860Z",
          "content": "<p>I won't.  I don't feel motivated to clean and document it.  Sorry.</p>",
          "rawMarkdown": "I won't.  I don't feel motivated to clean and document it.  Sorry.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3197197,
      "postDate": "2025-05-07T21:42:09.373Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1332628,
      "postDate": "2021-06-02T08:12:56.293Z",
      "content": "<p>Great job. Thanks.</p>",
      "rawMarkdown": "Great job. Thanks."
    }
  ],
  "comments": [
    {
      "id": 1350678,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-15T16:36:44.573000",
      "content": "<p>I shared code and the paper I submitted to the workshop: <a href=\"https://github.com/jfpuget/STFT_Transformer\" target=\"_blank\">https://github.com/jfpuget/STFT_Transformer</a></p>",
      "votes": 6,
      "replies": [
        {
          "id": 3197104,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-07T20:15:59.817000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1334968,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2021-06-03T23:58:06.960000",
      "content": "<p>Congrats on the strong finish and solo gold. Interesting solution with ViT and Deit.</p>\n<pre><code>This great start made me will to hide my score and I stopped submitting until someone matched my score.\n</code></pre>\n<p>Does withholding submissions in this way really work in practice? I'm curious because I may be finding myself in this position with 2 months left in CommonLit Readability Prize. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1335747,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-04T12:17:42.083000",
          "content": "<p>Thanks.</p>\n<p>It was the first time I did this, and in hindsight I think it was a mistake.  It was a mistake because my real score never went above 0.80, which was the public score.</p>\n<p>The main issue is rather to stay motivated without being pressured by followers on the LB. I did my best but I see I made more progress during last week, when others had closed the LB gap with me, than in the previous 6 weeks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1334647,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-06-03T16:41:19.640000",
      "content": "<p>Congrats on solo gold medal. <br>\nIt can be disappointing, but the Solo Gold medal is still great.</p>\n<p>And the part about DieT and ViT is very interesting. </p>\n<p>Thanks for sharing :) </p>\n<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1334672,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T17:17:48.393000",
          "content": "<p>Thanks, I spent quite some time with vision transformers to make them work here.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1334542,
      "author_name": "Daniel Furman",
      "author_url": "",
      "post_date": "2021-06-03T15:26:02.167000",
      "content": "<p>Congrats on the great result !! I for one think that the cross validation scheme played a big role in the modeling performance, I used the R blockCV package to generate my folds.</p>\n<p>Curious:</p>\n<p>1) did you use random kfold?<br>\n2) for mixup did you incorporate the secondary labels or did you have to one hot encode the target with just one primary label?</p>\n<p>Thanks for being such an active discussion contributor hope to compete in similar competitions going forward!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1334590,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T15:59:10.213000",
          "content": "<p>Thanks.  </p>\n<p>1) I used stratified k fold with shuffle.</p>\n<p>2) For mixup have a look at my Cornell writeup, I provide all details there. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1333984,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2021-06-03T07:30:19.500000",
      "content": "<p>Very interesting solution and insightful use of vision transformers. Congratulations</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1334335,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T12:32:51.617000",
          "content": "<p>Thanks!  We missed you in this comp.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1333696,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2021-06-03T01:16:19.123000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> congratulation on solo gold! though the results are a little bit rough( <br>\nThank you for sharing the results about Deit, very interesting insight.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1333944,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T07:05:10.423000",
          "content": "<p>Thanks. Yes, I am quite proud of how I used vision transformers.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1333288,
      "author_name": "eagle4",
      "author_url": "",
      "post_date": "2021-06-02T15:53:42.343000",
      "content": "<p>Congrats Oncle!<br>\nI was rooting for you !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1333254,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2021-06-02T15:28:35.833000",
      "content": "<p>Congrats on solo gold! </p>\n<blockquote>\n  <p>0.69 public LB (0.60 private LB) at my first submission</p>\n</blockquote>\n<p>I was overwhelmed by your strong start.</p>\n<blockquote>\n  <p>The Deit model alone scores 0.77 on public LB and 0.66 on private LB.</p>\n</blockquote>\n<p>This is very interesting result. I guess you've found very good techniques you would try in the first place when you face the similar problem statement.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1333273,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T15:41:37.187000",
          "content": "<p>Thanks.  Yes, I spend a LOT of time tuning ViT/Deit training and I am very happy to have succeeded somewhat.  It is why I am not too disappointed by the end result ;)</p>\n<p>I am now convinced that training only on 5 second clips is not the best, and next time I'll certainly have  a look at your SED model ;)</p>\n<p>And congrats on becoming GM soon!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1332776,
      "author_name": "Zaber Ibn Abdul Hakim",
      "author_url": "",
      "post_date": "2021-06-02T09:41:07.483000",
      "content": "<p>Congratulations. Your solution from previous Birdcall competition helped us a lot. Thank you. 😀</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1332700,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "2021-06-02T08:48:35.060000",
      "content": "<p>Congrats on solo gold!</p>\n<p>Could you share your baseline model weights and exact mel spectrogram settings? Your solution is almost fully opposite to mine - extremely strong single models, but i do not see any postprocessing.</p>\n<p>It would be really interesting to combine both and see if strong model still benefits from postprocessing as much as weak one.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1332740,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T09:14:28.727000",
          "content": "<p>Which is the strong model and the weak model?</p>\n<p>Anyway, I give details about my model in my Cornell writeup, link in the post.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332762,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "2021-06-02T09:27:52.940000",
          "content": "<p>Models with a high score would be a strong models. Yours is almost 0.1 higher than mine on private LB.</p>\n<p>Details would sadly not help much, i cant train 300x600 model for 60 epochs on this amount of data in any reasonable time:(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332782,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T09:47:56.803000",
          "content": "<p>;)  You made me share my postprocessing: <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243378\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/243378</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332854,
          "author_name": "Agneev",
          "author_url": "",
          "post_date": "2021-06-02T10:55:48.667000",
          "content": "<p>Thank you. I know you are disappointed at the shakeup, but still, congrats on the solo gold, and I'd also like to think you for all the help you provided in discussions throughout the competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1332629,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2021-06-02T08:13:06.827000",
      "content": "<p>Congratulations on the second solo gold! I think you helped me stay competitive. <br>\nI am impressing by your deep consideration of Vision Transformer. I've focused on inference and post-processing, so thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1332623,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2021-06-02T08:11:01.913000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ! I couldn't find it, did you train on whole clips or x seconds clips? If so, how did you set the target for them?</p>\n<blockquote>\n  <p>The main difference with Cornell competition is that we were provided train soundscape. As many I trained models on short audio records and tuned rounding thresholds on soundscape. This improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days. CV and LB were identical so far. At the time everyone else score was well below 0.70.</p>\n</blockquote>\n<p>I think this probably caused your shake-down. I had the same tuning which improved my CV to 0.74 but made my LB 0.64 or so. Then I noticed it is only 17 soundscapes in training set and around 30 in Public LB. This was the reason I expected a shake-up. People can't notice it when they get lucky with the submission.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1332630,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:13:26.750000",
          "content": "<p>I wrote it ;)</p>\n<blockquote>\n  <p>The same 5 folds CV model (efficientnet b3 on 5 seconds mel spectrograms)</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332638,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-02T08:17:37.587000",
          "content": "<p>What was your guys highest CV on all soundscape recordings? Ours was close to 0.84 but we decided to evaluate a bit differently (will write it in the solution post).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332639,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-06-02T08:17:50.663000",
          "content": "<p>Thanks, I somehow missed it. And how did you set the targets for them? Some of the 5 sec clips don't have the primary and secondary labels, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332641,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-02T08:20:45.327000",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> That's what also puzzles me a bit: <a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243324#1332617\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/243324#1332617</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332644,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:22:14.700000",
          "content": "<blockquote>\n  <p>What was your guys highest CV on all soundscape recordings? </p>\n</blockquote>\n<p>0.80, same as LB</p>\n<p>For CV I used the same as in my Cornell solution.  Compute F1score on no call rows separately fromF1 score on bird call rows, then compute final score with:</p>\n<p><code>score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</code></p>\n<p>This way a nocall sub CV is the same as the nocall submission LB.</p>\n<p>This made my CV and LB identical for most of my submissions.</p>\n<p>Updating my post as this is important.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1332654,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-02T08:28:27.453000",
          "content": "<p>Thanks - got it. Private LB has 0.47 sample submission btw.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332673,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:34:43.047000",
          "content": "<blockquote>\n  <p>Private LB has 0.47 sample submission btw.</p>\n</blockquote>\n<p>I am too lazy to check, but I bet that if I change my formula to this then Cv and private LB match:</p>\n<p><code>score_all = 0.47* score_nocall + (1 - 0.47) * score_birds</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332676,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:36:54.973000",
          "content": "<blockquote>\n  <p>Congrats <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> ! I couldn't find it, did you train on whole clips or x seconds clips? </p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> thanks.  I updated the post to make it clearer.  I used first and last 5 seconds of short train audio.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332921,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2021-06-02T11:37:17.210000",
          "content": "<blockquote>\n  <p>What was your guys highest CV on all soundscape recordings? </p>\n</blockquote>\n<p>F1 on SC: <code>0.824</code> / Public <code>0.78x</code> / Private <code>0.66x</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1333033,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2021-06-02T12:52:57.923000",
          "content": "<blockquote>\n  <p>… compute final score with: score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</p>\n</blockquote>\n<p>Used same idea but with: <code>0.56* score_nocall + (1 - 0.56) * score_birds</code></p>\n<p>and up to some range correlated very well with <strong>Public LB</strong>… for eg. </p>\n<p><code>Public 0.78x</code> / <code>Private 0.66x</code> --&gt; est <code>LB: 0.7876</code> | <code>nocall 0.90974</code>, <code>bird 0.6322</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1333322,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T16:22:00.950000",
          "content": "<blockquote>\n  <p>I think this probably caused your shake-down.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> why is that?  Everyone used soundscape to tune rounding thresholds.  What is specific to my case that would be worse than for other teams?</p>\n<p>I am really curious, you maybe right but I don't get why.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1333324,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T16:23:26.260000",
          "content": "<blockquote>\n  <p>Used same idea but with: 0.56* score_nocall + (1 - 0.56) * score_birds</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> why 0.56?  If you predict only no call it would be wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334079,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-06-03T09:18:01.250000",
          "content": "<blockquote>\n  <p>why is that?  Everyone used soundscape to tune rounding thresholds. What is specific to my case that would be worse than for other teams?</p>\n</blockquote>\n<p>Did you use soundscape level predictions too? For me, I used the prediction for 5 sec clip and total sum for thresholding. By this method, I was improving my local validation a lot but LB was going down. Then I notice thresholds were sensitive to folds etc. Then I thought people who got lucky on LB (probably half of the people) may not notice that there is high variance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334228,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T10:50:05.847000",
          "content": "<blockquote>\n  <p>Did you use soundscape level predictions too? </p>\n</blockquote>\n<p>like everyone else.</p>\n<p>People tuned thresholds using model predictions on train soundscape.</p>\n<p>In your case, the issue is probably that you did not correct for the different proportion of no calls.  Compute F1score on no call rows separately fromF1 score on bird call rows, then compute final score with:</p>\n<p>score_all = 0.54 * score_nocall + (1 - 0.54) * score_birds</p>\n<p>If you compute CV this way then CV and public LB are almost identical.</p>\n<p>I still don't get why you say it is luck.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334587,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-06-03T15:57:29.053000",
          "content": "<p>I used 3 soundscapes with no bird for background augmentation and validated my model on 17 landscapes. Simple thresholded model had 0.68 CV and 0.68 LB. When I included soundscape level thresholding, CV improved to 0.74 but LB went down to 0.64 or so. Then I repeat the same experiment with re-training on different folds etc and noticed that the score is very sensitive when soundscape level thresholding is involved. If I got lucky on LB, I wouldn't notice it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334595,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T16:02:47.053000",
          "content": "<p>Still what is specific to me here?  All top team used soundscape to tune thresholds.</p>\n<p>Given you used 3 soundscape data you are fitting to it, and relying on other soundscape looks leaky to me.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334601,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2021-06-03T16:06:01.053000",
          "content": "<blockquote>\n  <p>Given you used 3 soundscape data you are fitting to it, and relying on other soundscape looks leaky to me.</p>\n</blockquote>\n<p>It would be leaky if soundscapes were having overlaps.</p>\n<blockquote>\n  <p>Still what is specific to me here? All top team used soundscape to tune thresholds.</p>\n</blockquote>\n<p>Nothing specific to you:) My guess was that approach was a coin flip and you were one of the many people with positive flip.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334659,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T16:51:00.727000",
          "content": "<p>Why would all my subs be lucky then?</p>\n<p>How come my best CV is best on both public and private LB?</p>\n<p>I don't think it has to do with luck to be honest.</p>\n<p>Let's agree to disagree ;)</p>\n<p>My hunch is that private LB has less nocalls, hence I probably have optimized for a too high proportion of no calls.  Maybe that's what you want to say afterall?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1334181,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "2021-06-03T10:00:06.553000",
      "content": "<p>Congratulations. </p>\n<blockquote>\n  <p>The only trick is to modify the position embeddings to match the new grid dimension.</p>\n</blockquote>\n<p>Could you explain this further?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1334339,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-03T12:35:03.063000",
          "content": "<p>ViT and DeiT have position embeddings on a 24x24 grid.  I interpolated them to the new grid, for instance to a 16x36 grid.</p>\n<p>For this I modified timm's implementation of ViT and Deit position embeddings as follows.</p>\n<pre><code>def resize_pos_embed(posemb, new_size, num_tokens=1):\n    # Rescale the grid of position embeddings when loading from state_dict. Adapted from\n    # https://github.com/google-research/vision_transformer/blob/00883dd691c63a6830751563748663526e811cee/vit_jax/checkpoint.py#L224\n    if num_tokens:\n        posemb_tok, posemb_grid = posemb[:, :num_tokens], posemb[0, num_tokens:]\n    else:\n        posemb_tok, posemb_grid = posemb[:, :0], posemb[0]\n    gs_old = int(math.sqrt(len(posemb_grid)))\n    posemb_grid = posemb_grid.reshape(1, gs_old, gs_old, -1).permute(0, 3, 1, 2)\n    posemb_grid = F.interpolate(posemb_grid, size=new_size, mode='bilinear', align_corners=False)\n    posemb_grid = posemb_grid.permute(0, 2, 3, 1).reshape(1, new_size[0] * new_size[1], -1)\n    posemb = torch.cat([posemb_tok, posemb_grid], dim=1)\n    return posemb\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1335281,
          "author_name": "MaChaogong",
          "author_url": "",
          "post_date": "2021-06-04T06:06:25.683000",
          "content": "<p>Thanks for sharing. Congrats! <br>\nI could read your disappointment. 💪</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1333017,
      "author_name": "Shai",
      "author_url": "",
      "post_date": "2021-06-02T12:39:25.433000",
      "content": "<p>Congrats ! Birds got a real human buddy now !!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1332661,
      "author_name": "Stefan Kahl",
      "author_url": "",
      "post_date": "2021-06-02T08:30:01.480000",
      "content": "<p>Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work. Thanks for the write-up! Would be nice to see your approach as a working note :) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1332670,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:32:49.410000",
          "content": "<p>Thanks!.  Training ViT was tricky as they are quite unstable.  Maybe I'll focus on this part in my note as the rest is just fine tuning previous solutions.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1332684,
          "author_name": "Stefan Kahl",
          "author_url": "",
          "post_date": "2021-06-02T08:39:49.610000",
          "content": "<p>I bet! I also tried different patching methods and I think what you did makes most sense for spectrograms that are by default a sequential input. Rearranging the patches to use pre-trained models is what I was missing :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1332687,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:41:28.527000",
          "content": "<p>Yeah, I thought that skipping the stem convolution and directly input the time slices as patch embeddings would be hard to train from scratch because training data is small. Reusing pretrained wieghts was safer.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332800,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T10:01:19.290000",
          "content": "<p>I had submitted my deit model actually, alone it scores 0.77/0.66 which is not too bad.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332852,
          "author_name": "Agneev",
          "author_url": "",
          "post_date": "2021-06-02T10:54:07.413000",
          "content": "<blockquote>\n  <p>Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work.</p>\n</blockquote>\n<p>That's a relief, I thought I was the only one incapable of making a ViT work. My results on the training soundscapes were so poor I ended up not making a submission at all…maybe next year 😄</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1341233,
          "author_name": "Jan Schlüter",
          "author_url": "",
          "post_date": "2021-06-08T14:22:20.427000",
          "content": "<blockquote>\n  <p>Maybe I'll focus on this part in my note as the rest is just fine tuning previous solutions.</p>\n</blockquote>\n<p>Looking forward to your working notes paper! Really cool that you got transformers to work so well.</p>\n<blockquote>\n  <p>Reusing pretrained wieghts was safer.</p>\n</blockquote>\n<p>I'd be curious how it compares to training from scratch.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1341307,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-08T15:18:34.057000",
          "content": "<blockquote>\n  <p>I'd be curious how it compares to training from scratch.</p>\n</blockquote>\n<p>Me too actually, but I didn't felt motivated enough to run the experiment.  I may do it now that dust settled a bit.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1332622,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-06-02T08:09:58.507000",
      "content": "<p>Great writeup , congrats on solo gold .  It's great to see vision transformers work , we also tried it with square image and it was below par , now I get what was wrong , I didn't know about the different grids we can use. Coming early in the competition was the trick we missed , this made us experiment with Vit as black box only and I was not able to discover it closely (which is something I rarely do) .</p>\n<p>We also tried taking 7 second crops from the start and the end like you did in cornell but somehow it gave worse results , did you do the same in this comp?<br>\nAlso it would be great if you can give us some insights to which models you use and their training strategy</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1332633,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T08:14:39.067000",
          "content": "<blockquote>\n  <p>you did in cornell but somehow it gave worse results , did you do the same in this comp?</p>\n</blockquote>\n<p>Yes, as I wrote, i used the same approach  as in Cornell.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1392674,
      "author_name": "dan",
      "author_url": "",
      "post_date": "2021-07-19T01:13:30.050000",
      "content": "<p>Amazing research on vision transformer. Thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1332779,
      "author_name": "Stephen Lau",
      "author_url": "",
      "post_date": "2021-06-02T09:44:26.077000",
      "content": "<p>Great job! Mind if you share your code? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1332783,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-02T09:48:36.860000",
          "content": "<p>I won't.  I don't feel motivated to clean and document it.  Sorry.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3197197,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-07T21:42:09.373000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1332628,
      "author_name": "botkop",
      "author_url": "",
      "post_date": "2021-06-02T08:12:56.293000",
      "content": "<p>Great job. Thanks.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1332598": "I decided to join early this competition because I was frustrated by the previous two bird song competitions.  In Cornell competition I joined too late. In Rainforest competition,  missed some key train/test distribution insights and stagnated in the LB after  a great start.\n\nThe same 5 folds CV model (efficientnet b3 on first and last 5 seconds mel spectrograms) as in [Cornell competition](https://www.kaggle.com/c/birdsong-recognition/discussion/183219) plus [improvements from my Rainforest solution](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304) gave 0.69 public LB (0.60 private LB) at my first submission.  This made me very happy as I took the lead on the public LB with it.\n\nThe main difference with Cornell competition is that we were provided train soundscape.  As many I trained models on short audio records and tuned rounding thresholds on soundscape.  For CV I used the same as in my Cornell solution. Compute F1score on no call rows separately from f1 score on bird call rows, then compute final score with:\n\nscore_all = 0.54 * score_nocall + (1 - 0.54) * score_birds\n\nThis way a nocall sub CV is the same as the nocall submission LB.\n\nThis made my CV and LB identical for most of my submissions.\n\nTuning thresholds improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days.  CV and LB were identical so far.  At the time everyone else score was well below 0.70.\n\nThis great start made me will to hide my score and I stopped submitting until someone matched my score.  It took 2 weeks.  At this time I submitted the same model bagged twice, i.e. the original  model averaged with a second model trained with a different seed.   This scored 0.78 on public LB (0.64 LB).  I knew it was lucky as CV was 0.76, but it looked great still. It took another 2 weeks for someone to beat this public score.  Later, after improving a bit the training procedure I got 0.80 public LB (0.66 private) with a 2 seeds x 5 folds submission of my baseline.\n\nIt means that my baseline alone gives me a top 20 final rank.\n\nGiven single models were so great I assumed ensembling would move me ahead further and I decided to focus on creating a wide range of individual models.  This is my main mistake, I should have worked on ensembling way earlier. \n\nI only submitted ensembles after the submission outage, 3 days before deadline, and discovered that it was hard to have a blending ensemble that beats all its individual component models.  I beat my best individual model only in my last submission.  It has both a CV and a public LB of 0.80 and is also my best private LB at 0.67.  \n\nI started a stacking model last day, but this was too late...  So be it.\n\nWhile I was holding top public score I decided to not submit and I explored lots of different models.  In particular I explored vision transformers.  I started with ViT and Deit.  First try with 384x384 mel spectrograms were disappointing.  Then I realized that i could use other image dimensions instead of a 24x24 grid of 16x16 patches.  I tried 12x48, i.e. a 192x768 spectrogram, and also a 16x36 grid (256x576 spectrogram).  The only trick is to modify the position embeddings to match the new grid dimension.\n\nThis led to 0.75 public LB (0.64 private).\n\nBut the most interesting one was to forget about square patches altogether.  A 196x576 spectrogram can be seen as 576 time slices of size 196.  Each slice contains 16x16 entries.  It means that I could just use the time slices as input patches.  Here is how this input looks once it is reshaped as a 24x24 grid of 16x16 patches:\n\n![time slices](https://i.imgur.com/F3zVGVM.pngd)\n\nMaybe surprisingly, vision transformers are happy with this input.  The main advantage is that there is no longer any issue with translation on the frequency axis, which is the main issue with CNNs applied to spectrograms.\n\nBlending Deit trained on this input with my baseline gave my best sub.  The Deit model alone scores 0.77 on public LB and 0.66 on private LB.\n\nAlthough I am disappointed by my final result, I am happy to have explored lots of vision models and devised some new ways to use them.  And being disappointed by a solo gold in a deep learning competition is something I would not have imagined one year ago anyway ;)\n\nSpecial thanks to Ross Wightman for his timm package.  It made my model exploration seamless.  \n\nEdit. I shared code and the paper I submitted to the workshop: https://github.com/jfpuget/STFT_Transformer",
    "1350678": "I shared code and the paper I submitted to the workshop: https://github.com/jfpuget/STFT_Transformer",
    "1334968": "Congrats on the strong finish and solo gold. Interesting solution with ViT and Deit.\n\n```\nThis great start made me will to hide my score and I stopped submitting until someone matched my score.\n```\n\nDoes withholding submissions in this way really work in practice? I'm curious because I may be finding myself in this position with 2 months left in CommonLit Readability Prize. ",
    "1334647": "Congrats on solo gold medal. \nIt can be disappointing, but the Solo Gold medal is still great.\n\nAnd the part about DieT and ViT is very interesting. \n\nThanks for sharing :) \n\n@cpmpml ",
    "1334542": "Congrats on the great result !! I for one think that the cross validation scheme played a big role in the modeling performance, I used the R blockCV package to generate my folds.\n\nCurious:\n\n1) did you use random kfold?\n2) for mixup did you incorporate the secondary labels or did you have to one hot encode the target with just one primary label?\n\nThanks for being such an active discussion contributor hope to compete in similar competitions going forward!",
    "1333984": "Very interesting solution and insightful use of vision transformers. Congratulations",
    "1333696": "@cpmpml congratulation on solo gold! though the results are a little bit rough( \nThank you for sharing the results about Deit, very interesting insight.",
    "1333288": "Congrats Oncle!\nI was rooting for you !",
    "1333254": "Congrats on solo gold! \n\n> 0.69 public LB (0.60 private LB) at my first submission\n\nI was overwhelmed by your strong start.\n\n> The Deit model alone scores 0.77 on public LB and 0.66 on private LB.\n\nThis is very interesting result. I guess you've found very good techniques you would try in the first place when you face the similar problem statement.",
    "1332776": "Congratulations. Your solution from previous Birdcall competition helped us a lot. Thank you. 😀",
    "1332700": "Congrats on solo gold!\n\nCould you share your baseline model weights and exact mel spectrogram settings? Your solution is almost fully opposite to mine - extremely strong single models, but i do not see any postprocessing.\n\nIt would be really interesting to combine both and see if strong model still benefits from postprocessing as much as weak one.\n",
    "1332629": "Congratulations on the second solo gold! I think you helped me stay competitive. \nI am impressing by your deep consideration of Vision Transformer. I've focused on inference and post-processing, so thank you for sharing!",
    "1332623": "Congrats @cpmpml ! I couldn't find it, did you train on whole clips or x seconds clips? If so, how did you set the target for them?\n\n> The main difference with Cornell competition is that we were provided train soundscape. As many I trained models on short audio records and tuned rounding thresholds on soundscape. This improved my public LB to 0.75 (0.64 private) at my 5th submission, after 2 days. CV and LB were identical so far. At the time everyone else score was well below 0.70.\n\nI think this probably caused your shake-down. I had the same tuning which improved my CV to 0.74 but made my LB 0.64 or so. Then I noticed it is only 17 soundscapes in training set and around 30 in Public LB. This was the reason I expected a shake-up. People can't notice it when they get lucky with the submission.",
    "1334181": "Congratulations. \n> The only trick is to modify the position embeddings to match the new grid dimension.\n\nCould you explain this further?",
    "1333017": "Congrats ! Birds got a real human buddy now !!",
    "1332661": "Nice! I suspected that ViT would be capable of scoring high in the leaderboard, but I could never actually make them work. Thanks for the write-up! Would be nice to see your approach as a working note :) ",
    "1332622": "Great writeup , congrats on solo gold .  It's great to see vision transformers work , we also tried it with square image and it was below par , now I get what was wrong , I didn't know about the different grids we can use. Coming early in the competition was the trick we missed , this made us experiment with Vit as black box only and I was not able to discover it closely (which is something I rarely do) .\n\nWe also tried taking 7 second crops from the start and the end like you did in cornell but somehow it gave worse results , did you do the same in this comp?\nAlso it would be great if you can give us some insights to which models you use and their training strategy",
    "1392674": "Amazing research on vision transformer. Thank you!",
    "1332779": "Great job! Mind if you share your code? ",
    "3197197": "",
    "1332628": "Great job. Thanks."
  }
}