{
  "id": 220765,
  "title": "Many mistakes, no regrets",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220765",
  "author_name": "",
  "post_date": "2021-02-19T13:33:40.710895700Z",
  "votes": 15,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Post-RiiiD, when my good friend <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> asked me to join the Jane street competition, I was tempted. It too was a time-series like RiiiD, but then I wanted to wade into something totally new. RFCX with it eco-appeal caught my attention. I was new to both sound processing and CNN… so this seemed like a good place to learn something. Thanks to the wonderful discussion topics by <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> , <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> and my personal fav <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> (really sorry you missed the gold) I was off to a good start. Because I learnt Pytorch for RiiiD, I decided to use the TF dataset pipeline and hence the wonderful notebooks by <a href=\"https://www.kaggle.com/yoshi999\" target=\"_blank\">@yoshi999</a> and <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> (who extended it to GPUs) provided a good start. In hindsight though, I realised going with TF was a wrong decision. Most of the winning entries for Birdclef and prev sound competitions were Pytorch based…I could have reused a lot of stuff from there instead of trying to convert. But no regrets…TF is a gold-standard framework and I am glad to have used it. The official documentation seems to be the weak point though. I had absolutely no trouble quickly learning and implementing Pytorch in RiiiD but here, it was taking far more time. But of course this is a beginner’s perspective and I may be wrong.</p>\n<p>Few things clearly stood out when I examined the problem and the public solutions:</p>\n<ul>\n<li><p>The meta data provided contained the fmin, fmax as shared by the domain experts and I didn’t see any public kernel or discussion on it. I felt this was a key piece of information and mentioned about it in my exploratory kernel, hoping someone might notice and comment on it. My initial thought was to slide along the vertical axis instead of the horizontal axis and then aggregate information..I broke the species into 4 categories based on the f-range. Except one or two species which had huge ranges, all other species fell into neat little freq buckets. So now I could slide along 4-5 buckets of frequency ranges and aggregate species information from all of them (somewhat like what we were doing on the time axis). The question was should I fully invest in building this model because I did have some severe time constraints in these past few weeks. I did an experiment wherein I removed species num 22 (which had max freq of 13K) and then reduced the fmax of the spects from 20K to 12K and re-ran my training. Now my new model should have predicted species 0-21 and 23 with better clarity. I had to blend in species num 22 and I just used the relative rank of species 22 from my original model. So for a particular record, if s22 was ranked 5th in my original model 1, I made it 5th rank in my second model as well. It was a crude blend but I felt it may give me some indication of whether to invest in such a model. Unfortunately, though I reduced fmax from 20K to 12K, there was no improvement in scores. So a big failure.</p></li>\n<li><p>Most discussions and kernels revolved around using log Mel specs. These are more for human speech modelling and I felt there could be merit in also experimenting with plain spects, log vs non-log,  hpss and even ensembling results with mfcc spects. I couldnt get hpss to work with tf. As for the rest, nothing worked. </p></li>\n<li><p>Somewhere in the middle of the competition I read someone announcing that they got 86% with all ‘1’ combination. For a while I toyed with the idea that maybe I was looking at the competition completely incorrectly. Maybe by default each recording contained nearly all birds and a more easier approach to the comp was finding out which species is NOT in the recording and we should be leveraging only the FP data for this. In my excitement I forgot that all ‘1’s is same as all 0’sand unless there was some bug in the competition evaluation, getting 86 was not possible with all 1s. Quickly abandoned the approach</p></li>\n<li><p>Coming to FP, we know that there was nearly 3 times as much data as TP. Definitely it would be a pity to let go of this data. Yet, none of the public kernels/notebooks had given much info on how to leverage this. I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data (Yes I was not ready to let go of that f-range :).  I had earlier abandoned use of 24 diff models - one for each species because this model might not capture inter-dependency between species. But what if I used the output of my first model (with all 24 species ranked) and then augmented that output with the results from the one-model-per-species (simple binary model). I might get the best of both worlds. So assuming I had a 85% score and then I ran my s0 model (levering both s0 TP and s0 FP), I could then compare and augment the original model results with s0 model output. So trained an s0, s3, s13, s19 models to start with and then began blending.  For each model, I would check which are the high probability +ves and -ves and then compare with the original model for those records and make the s’X’ score for that record=highest rank for high +ves and zero for -ves. For e.g. for species 0, I found about 22 records that had high probability of being s0 but in my original model results these 22 records did not contain s0 in top 5 ranks. So I changed the scores of all these 22 records in my original model to be the highest score for that record (thus making sure it got rank 1). Ditto for negatives as well and ditto for several other species. I though species inter-dependencies was anyway captured in the original model and by augmenting it with the confirmed +ves and -ves from the individual 24 models I could take the score higher by leveraging the smaller frequency range associated with each species as well as leveraging the FP data.  Nope..nothing worked and I am still wondering why?</p></li>\n</ul>\n<p>Meanwhile there was significant work pressure and combined with all the above failures, I thought of just taking a break from Kaggle for a few months and logging in later in the year with a fresh perspective. Then 3 things happened:</p>\n<ul>\n<li>Kaggles ML logic kicked in and it decided to show me an old post from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> on CNN ensembling on my home page. Now ensembling was not really my top priority but Chris post was pretty interesting. He had suggested ensembling diff sizes like 255,255 &amp; 555,555 &amp; 1024, 1024 etc and mentioned that it was possible for the model to capture different info from diff sizes and the ensemble almost always scores higher. Extending Chris thoughts further, I thought why don’t I try different length and breadth also and maybe this stretching and pulling might result in even more patterns being discovered and I started experiencing good results</li>\n<li>One of my solo models got 87.6 coupled with some other optimizations</li>\n<li>I had been trying to parallely fine tune the SED model to get better scores as I wanted to use diff model ensemble in my final submission, but was not having much success. <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a> released a good high scoring model using SED which was just what I wanted..</li>\n</ul>\n<p>From 87.6% I moved to my final score of 90.1% with minor other optimisations and ensembling….a beginner’s score. Six months back when I was trying to code the Titanic, I was not sure where I would move on next. I had convinced myself at that time that plunging headlong into real competition was the quickest way to learn but I am not so sure of that now.  I guess I should be happy with 2 bronzes in my first 2 comps but the issue is that in both the competitions I was enjoying the learning process UNTIL I made my first submission. Once I did my first submission, everything changed to just ‘how to improve the score’. So maybe for next few comps, I might just stick to discussions and exploratory notebooks :)</p>\n<p>Once again a big shout out to all those kernels and interesting discussions by all folks. Some of my strategies may sound ridiculous but no regrets. I experimented, learnt something and will try to improve..life moves on…</p>\n<p>Edit: Some things are beginning to make sense now. Approach 1 (sliding along vertical axis and aggregating species) is sort of a hybrid approach where we can try to capture interdependencies between a set of species if any while parallely leveraging the f-range information to cut out all other noise ranges and thereby hope to make a more accurate prediction. However this may not increase score beyond a certain point and almost all winning teams seem to have used FP and freq cropping at species level…So ideally my approach 4 (last approach)…which has the advantage of using FPs and also a much narrower cropping of frequencies should have worked.. </p>\n<p>Why didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains. However there are others who have reported early-nineties score without using pseudo-labelling. So there is definitely something I need to investigate there. Definitely the FP vastly outnumbered the TPs and maybe this was biasing the model..I had thought of reducing FPs randomly such that FP and TP size was similar but didnt get time to do it. Maybe some approach like that would have shown some gains..</p>",
  "messages": [
    {
      "id": "1210490",
      "postDate": "02/19/2021 13:33:40",
      "content": "<p>Post-RiiiD, when my good friend <a href=\"https://www.kaggle.com/carlmcbrideellis\" target=\"_blank\">@carlmcbrideellis</a> asked me to join the Jane street competition, I was tempted. It too was a time-series like RiiiD, but then I wanted to wade into something totally new. RFCX with it eco-appeal caught my attention. I was new to both sound processing and CNN… so this seemed like a good place to learn something. Thanks to the wonderful discussion topics by <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> , <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> and my personal fav <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> (really sorry you missed the gold) I was off to a good start. Because I learnt Pytorch for RiiiD, I decided to use the TF dataset pipeline and hence the wonderful notebooks by <a href=\"https://www.kaggle.com/yoshi999\" target=\"_blank\">@yoshi999</a> and <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a> (who extended it to GPUs) provided a good start. In hindsight though, I realised going with TF was a wrong decision. Most of the winning entries for Birdclef and prev sound competitions were Pytorch based…I could have reused a lot of stuff from there instead of trying to convert. But no regrets…TF is a gold-standard framework and I am glad to have used it. The official documentation seems to be the weak point though. I had absolutely no trouble quickly learning and implementing Pytorch in RiiiD but here, it was taking far more time. But of course this is a beginner’s perspective and I may be wrong.</p>\n<p>Few things clearly stood out when I examined the problem and the public solutions:</p>\n<ul>\n<li><p>The meta data provided contained the fmin, fmax as shared by the domain experts and I didn’t see any public kernel or discussion on it. I felt this was a key piece of information and mentioned about it in my exploratory kernel, hoping someone might notice and comment on it. My initial thought was to slide along the vertical axis instead of the horizontal axis and then aggregate information..I broke the species into 4 categories based on the f-range. Except one or two species which had huge ranges, all other species fell into neat little freq buckets. So now I could slide along 4-5 buckets of frequency ranges and aggregate species information from all of them (somewhat like what we were doing on the time axis). The question was should I fully invest in building this model because I did have some severe time constraints in these past few weeks. I did an experiment wherein I removed species num 22 (which had max freq of 13K) and then reduced the fmax of the spects from 20K to 12K and re-ran my training. Now my new model should have predicted species 0-21 and 23 with better clarity. I had to blend in species num 22 and I just used the relative rank of species 22 from my original model. So for a particular record, if s22 was ranked 5th in my original model 1, I made it 5th rank in my second model as well. It was a crude blend but I felt it may give me some indication of whether to invest in such a model. Unfortunately, though I reduced fmax from 20K to 12K, there was no improvement in scores. So a big failure.</p></li>\n<li><p>Most discussions and kernels revolved around using log Mel specs. These are more for human speech modelling and I felt there could be merit in also experimenting with plain spects, log vs non-log,  hpss and even ensembling results with mfcc spects. I couldnt get hpss to work with tf. As for the rest, nothing worked. </p></li>\n<li><p>Somewhere in the middle of the competition I read someone announcing that they got 86% with all ‘1’ combination. For a while I toyed with the idea that maybe I was looking at the competition completely incorrectly. Maybe by default each recording contained nearly all birds and a more easier approach to the comp was finding out which species is NOT in the recording and we should be leveraging only the FP data for this. In my excitement I forgot that all ‘1’s is same as all 0’sand unless there was some bug in the competition evaluation, getting 86 was not possible with all 1s. Quickly abandoned the approach</p></li>\n<li><p>Coming to FP, we know that there was nearly 3 times as much data as TP. Definitely it would be a pity to let go of this data. Yet, none of the public kernels/notebooks had given much info on how to leverage this. I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data (Yes I was not ready to let go of that f-range :).  I had earlier abandoned use of 24 diff models - one for each species because this model might not capture inter-dependency between species. But what if I used the output of my first model (with all 24 species ranked) and then augmented that output with the results from the one-model-per-species (simple binary model). I might get the best of both worlds. So assuming I had a 85% score and then I ran my s0 model (levering both s0 TP and s0 FP), I could then compare and augment the original model results with s0 model output. So trained an s0, s3, s13, s19 models to start with and then began blending.  For each model, I would check which are the high probability +ves and -ves and then compare with the original model for those records and make the s’X’ score for that record=highest rank for high +ves and zero for -ves. For e.g. for species 0, I found about 22 records that had high probability of being s0 but in my original model results these 22 records did not contain s0 in top 5 ranks. So I changed the scores of all these 22 records in my original model to be the highest score for that record (thus making sure it got rank 1). Ditto for negatives as well and ditto for several other species. I though species inter-dependencies was anyway captured in the original model and by augmenting it with the confirmed +ves and -ves from the individual 24 models I could take the score higher by leveraging the smaller frequency range associated with each species as well as leveraging the FP data.  Nope..nothing worked and I am still wondering why?</p></li>\n</ul>\n<p>Meanwhile there was significant work pressure and combined with all the above failures, I thought of just taking a break from Kaggle for a few months and logging in later in the year with a fresh perspective. Then 3 things happened:</p>\n<ul>\n<li>Kaggles ML logic kicked in and it decided to show me an old post from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> on CNN ensembling on my home page. Now ensembling was not really my top priority but Chris post was pretty interesting. He had suggested ensembling diff sizes like 255,255 &amp; 555,555 &amp; 1024, 1024 etc and mentioned that it was possible for the model to capture different info from diff sizes and the ensemble almost always scores higher. Extending Chris thoughts further, I thought why don’t I try different length and breadth also and maybe this stretching and pulling might result in even more patterns being discovered and I started experiencing good results</li>\n<li>One of my solo models got 87.6 coupled with some other optimizations</li>\n<li>I had been trying to parallely fine tune the SED model to get better scores as I wanted to use diff model ensemble in my final submission, but was not having much success. <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a> released a good high scoring model using SED which was just what I wanted..</li>\n</ul>\n<p>From 87.6% I moved to my final score of 90.1% with minor other optimisations and ensembling….a beginner’s score. Six months back when I was trying to code the Titanic, I was not sure where I would move on next. I had convinced myself at that time that plunging headlong into real competition was the quickest way to learn but I am not so sure of that now.  I guess I should be happy with 2 bronzes in my first 2 comps but the issue is that in both the competitions I was enjoying the learning process UNTIL I made my first submission. Once I did my first submission, everything changed to just ‘how to improve the score’. So maybe for next few comps, I might just stick to discussions and exploratory notebooks :)</p>\n<p>Once again a big shout out to all those kernels and interesting discussions by all folks. Some of my strategies may sound ridiculous but no regrets. I experimented, learnt something and will try to improve..life moves on…</p>\n<p>Edit: Some things are beginning to make sense now. Approach 1 (sliding along vertical axis and aggregating species) is sort of a hybrid approach where we can try to capture interdependencies between a set of species if any while parallely leveraging the f-range information to cut out all other noise ranges and thereby hope to make a more accurate prediction. However this may not increase score beyond a certain point and almost all winning teams seem to have used FP and freq cropping at species level…So ideally my approach 4 (last approach)…which has the advantage of using FPs and also a much narrower cropping of frequencies should have worked.. </p>\n<p>Why didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains. However there are others who have reported early-nineties score without using pseudo-labelling. So there is definitely something I need to investigate there. Definitely the FP vastly outnumbered the TPs and maybe this was biasing the model..I had thought of reducing FPs randomly such that FP and TP size was similar but didnt get time to do it. Maybe some approach like that would have shown some gains..</p>",
      "rawMarkdown": "Post-RiiiD, when my good friend @carlmcbrideellis asked me to join the Jane street competition, I was tempted. It too was a time-series like RiiiD, but then I wanted to wade into something totally new. RFCX with it eco-appeal caught my attention. I was new to both sound processing and CNN… so this seemed like a good place to learn something. Thanks to the wonderful discussion topics by @usharengaraju , @shinmurashinmura and my personal fav @barnwellguy (really sorry you missed the gold) I was off to a good start. Because I learnt Pytorch for RiiiD, I decided to use the TF dataset pipeline and hence the wonderful notebooks by @yoshi999 and @aikhmelnytskyy (who extended it to GPUs) provided a good start. In hindsight though, I realised going with TF was a wrong decision. Most of the winning entries for Birdclef and prev sound competitions were Pytorch based…I could have reused a lot of stuff from there instead of trying to convert. But no regrets…TF is a gold-standard framework and I am glad to have used it. The official documentation seems to be the weak point though. I had absolutely no trouble quickly learning and implementing Pytorch in RiiiD but here, it was taking far more time. But of course this is a beginner’s perspective and I may be wrong.\n\nFew things clearly stood out when I examined the problem and the public solutions:\n- The meta data provided contained the fmin, fmax as shared by the domain experts and I didn’t see any public kernel or discussion on it. I felt this was a key piece of information and mentioned about it in my exploratory kernel, hoping someone might notice and comment on it. My initial thought was to slide along the vertical axis instead of the horizontal axis and then aggregate information..I broke the species into 4 categories based on the f-range. Except one or two species which had huge ranges, all other species fell into neat little freq buckets. So now I could slide along 4-5 buckets of frequency ranges and aggregate species information from all of them (somewhat like what we were doing on the time axis). The question was should I fully invest in building this model because I did have some severe time constraints in these past few weeks. I did an experiment wherein I removed species num 22 (which had max freq of 13K) and then reduced the fmax of the spects from 20K to 12K and re-ran my training. Now my new model should have predicted species 0-21 and 23 with better clarity. I had to blend in species num 22 and I just used the relative rank of species 22 from my original model. So for a particular record, if s22 was ranked 5th in my original model 1, I made it 5th rank in my second model as well. It was a crude blend but I felt it may give me some indication of whether to invest in such a model. Unfortunately, though I reduced fmax from 20K to 12K, there was no improvement in scores. So a big failure.\n\n- Most discussions and kernels revolved around using log Mel specs. These are more for human speech modelling and I felt there could be merit in also experimenting with plain spects, log vs non-log,  hpss and even ensembling results with mfcc spects. I couldnt get hpss to work with tf. As for the rest, nothing worked. \n\n- Somewhere in the middle of the competition I read someone announcing that they got 86% with all ‘1’ combination. For a while I toyed with the idea that maybe I was looking at the competition completely incorrectly. Maybe by default each recording contained nearly all birds and a more easier approach to the comp was finding out which species is NOT in the recording and we should be leveraging only the FP data for this. In my excitement I forgot that all ‘1’s is same as all 0’sand unless there was some bug in the competition evaluation, getting 86 was not possible with all 1s. Quickly abandoned the approach\n\n- Coming to FP, we know that there was nearly 3 times as much data as TP. Definitely it would be a pity to let go of this data. Yet, none of the public kernels/notebooks had given much info on how to leverage this. I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data (Yes I was not ready to let go of that f-range :).  I had earlier abandoned use of 24 diff models - one for each species because this model might not capture inter-dependency between species. But what if I used the output of my first model (with all 24 species ranked) and then augmented that output with the results from the one-model-per-species (simple binary model). I might get the best of both worlds. So assuming I had a 85% score and then I ran my s0 model (levering both s0 TP and s0 FP), I could then compare and augment the original model results with s0 model output. So trained an s0, s3, s13, s19 models to start with and then began blending.  For each model, I would check which are the high probability +ves and -ves and then compare with the original model for those records and make the s’X’ score for that record=highest rank for high +ves and zero for -ves. For e.g. for species 0, I found about 22 records that had high probability of being s0 but in my original model results these 22 records did not contain s0 in top 5 ranks. So I changed the scores of all these 22 records in my original model to be the highest score for that record (thus making sure it got rank 1). Ditto for negatives as well and ditto for several other species. I though species inter-dependencies was anyway captured in the original model and by augmenting it with the confirmed +ves and -ves from the individual 24 models I could take the score higher by leveraging the smaller frequency range associated with each species as well as leveraging the FP data.  Nope..nothing worked and I am still wondering why?\n\nMeanwhile there was significant work pressure and combined with all the above failures, I thought of just taking a break from Kaggle for a few months and logging in later in the year with a fresh perspective. Then 3 things happened:\n- Kaggles ML logic kicked in and it decided to show me an old post from @cdeotte on CNN ensembling on my home page. Now ensembling was not really my top priority but Chris post was pretty interesting. He had suggested ensembling diff sizes like 255,255 & 555,555 & 1024, 1024 etc and mentioned that it was possible for the model to capture different info from diff sizes and the ensemble almost always scores higher. Extending Chris thoughts further, I thought why don’t I try different length and breadth also and maybe this stretching and pulling might result in even more patterns being discovered and I started experiencing good results\n- One of my solo models got 87.6 coupled with some other optimizations\n- I had been trying to parallely fine tune the SED model to get better scores as I wanted to use diff model ensemble in my final submission, but was not having much success. @reppic released a good high scoring model using SED which was just what I wanted..\n\nFrom 87.6% I moved to my final score of 90.1% with minor other optimisations and ensembling….a beginner’s score. Six months back when I was trying to code the Titanic, I was not sure where I would move on next. I had convinced myself at that time that plunging headlong into real competition was the quickest way to learn but I am not so sure of that now.  I guess I should be happy with 2 bronzes in my first 2 comps but the issue is that in both the competitions I was enjoying the learning process UNTIL I made my first submission. Once I did my first submission, everything changed to just ‘how to improve the score’. So maybe for next few comps, I might just stick to discussions and exploratory notebooks :)\n\nOnce again a big shout out to all those kernels and interesting discussions by all folks. Some of my strategies may sound ridiculous but no regrets. I experimented, learnt something and will try to improve..life moves on...\n\nEdit: Some things are beginning to make sense now. Approach 1 (sliding along vertical axis and aggregating species) is sort of a hybrid approach where we can try to capture interdependencies between a set of species if any while parallely leveraging the f-range information to cut out all other noise ranges and thereby hope to make a more accurate prediction. However this may not increase score beyond a certain point and almost all winning teams seem to have used FP and freq cropping at species level...So ideally my approach 4 (last approach)...which has the advantage of using FPs and also a much narrower cropping of frequencies should have worked.. \n\nWhy didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains. However there are others who have reported early-nineties score without using pseudo-labelling. So there is definitely something I need to investigate there. Definitely the FP vastly outnumbered the TPs and maybe this was biasing the model..I had thought of reducing FPs randomly such that FP and TP size was similar but didnt get time to do it. Maybe some approach like that would have shown some gains..",
      "votes": null
    },
    {
      "id": "1210505",
      "postDate": "02/19/2021 13:42:59",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a>,</p>\n<p>Congratulations on the medal! With so many competitions drawing to a close over the last, and in the next, few days I imagine we shall be in for a major new kaggle challenge in the very near future, so please do not  \"<em>…just stick to discussions and exploratory notebooks :)</em>\"</p>\n<p>All the best, and well done,<br>\ncarl</p>",
      "rawMarkdown": "Dear @allohvk,\n\nCongratulations on the medal! With so many competitions drawing to a close over the last, and in the next, few days I imagine we shall be in for a major new kaggle challenge in the very near future, so please do not  \"*...just stick to discussions and exploratory notebooks :)*\"\n\nAll the best, and well done,\ncarl",
      "votes": null
    },
    {
      "id": "1210556",
      "postDate": "02/19/2021 14:10:00",
      "content": "<p>I appreciate your mentioning!<br>\nIn this competition, you are my personal fav too :)</p>\n<blockquote>\n  <p>I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data</p>\n</blockquote>\n<p>That was all I did. </p>\n<blockquote>\n  <p>Once I did my first submission, everything changed to just ‘how to improve the score’.</p>\n</blockquote>\n<p>Very true. The attitude of quietly accumulating knowledge by actually getting to depth of things are threatened by the seemingly omnipresent and omnipotent sentence \"the LB does not improve\" in the forum, which, In my opinion, is a pity, along with many other things.</p>\n<p>EDIT: wish you the best in your path of learning!</p>",
      "rawMarkdown": "I appreciate your mentioning!\nIn this competition, you are my personal fav too :)\n\n>I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data\n\nThat was all I did. \n\n> Once I did my first submission, everything changed to just ‘how to improve the score’.\n\nVery true. The attitude of quietly accumulating knowledge by actually getting to depth of things are threatened by the seemingly omnipresent and omnipotent sentence \"the LB does not improve\" in the forum, which, In my opinion, is a pity, along with many other things.\n\nEDIT: wish you the best in your path of learning!",
      "votes": null
    },
    {
      "id": "1210578",
      "postDate": "02/19/2021 14:29:18",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> ! In my limited wanderings on Kaggle, I did come across couple of Masters (actually just one…) who do this. They are not interested in the submissions or scores but participate heartily in the discussions and notebooks, giving brilliant ideas..they dont go beyond the CV…so there is no temptation and focus doesn't shift from 'solution' to 'score'… I was fortunate enough to have a few discussions in RiiiD. with one such Master</p>",
      "rawMarkdown": "Thanks @barnwellguy ! In my limited wanderings on Kaggle, I did come across couple of Masters (actually just one...) who do this. They are not interested in the submissions or scores but participate heartily in the discussions and notebooks, giving brilliant ideas..they dont go beyond the CV...so there is no temptation and focus doesn't shift from 'solution' to 'score'... I was fortunate enough to have a few discussions in RiiiD. with one such Master",
      "votes": null
    },
    {
      "id": "1210715",
      "postDate": "02/19/2021 16:17:29",
      "content": "<p>Your encouragement in my initial days is one of the strong reasons for me to be on Kaggle, my friend..</p>",
      "rawMarkdown": "Your encouragement in my initial days is one of the strong reasons for me to be on Kaggle, my friend..",
      "votes": null
    },
    {
      "id": "1210814",
      "postDate": "02/19/2021 17:53:39",
      "content": "<p>Congrats on your final score and the medal. Your hard work seems to paid off. I guess Chris blessed all of us with his wisdom :)</p>",
      "rawMarkdown": "Congrats on your final score and the medal. Your hard work seems to paid off. I guess Chris blessed all of us with his wisdom :)",
      "votes": null
    },
    {
      "id": "1211859",
      "postDate": "02/20/2021 16:08:06",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> , u have done tremendously well urself. I think had you cracked how to use FP, you would have moved to Gold..All the best for your next comp!!</p>",
      "rawMarkdown": "Thanks @snnclsr , u have done tremendously well urself. I think had you cracked how to use FP, you would have moved to Gold..All the best for your next comp!!",
      "votes": null
    },
    {
      "id": "1212894",
      "postDate": "02/21/2021 17:45:08",
      "content": "<p>Thanks for sharing</p>\n<p>As I too am on a learning path, but way behind your own curve, I appreciate your sharing of your experience</p>\n<p>Learning from somebody else's choices and reasons is really a kind of one-shot transfer learning- accelerates the reader's learning curve😇</p>",
      "rawMarkdown": "Thanks for sharing\n\nAs I too am on a learning path, but way behind your own curve, I appreciate your sharing of your experience\n\nLearning from somebody else's choices and reasons is really a kind of one-shot transfer learning- accelerates the reader's learning curve😇",
      "votes": null
    },
    {
      "id": "1212925",
      "postDate": "02/21/2021 18:17:49",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> :)</p>\n<p>After reviewing most of the top solutions, it melted down into the pseudo labelling (includes using the FP data) and postprocessing IMHO. </p>\n<p>I wish you the best of luck as well on your journey. Thanks 🙏</p>",
      "rawMarkdown": "Thanks a lot @allohvk :)\n\nAfter reviewing most of the top solutions, it melted down into the pseudo labelling (includes using the FP data) and postprocessing IMHO. \n\nI wish you the best of luck as well on your journey. Thanks 🙏",
      "votes": null
    },
    {
      "id": "1214033",
      "postDate": "02/22/2021 14:49:12",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/robertolofaro\" target=\"_blank\">@robertolofaro</a> I definitely intend to stay back in RFCX for a while and tinker around with the code and the approaches provided by the winning teams. I will definitely try to publish some of my findings…</p>",
      "rawMarkdown": "Thanks @robertolofaro I definitely intend to stay back in RFCX for a while and tinker around with the code and the approaches provided by the winning teams. I will definitely try to publish some of my findings...",
      "votes": null
    },
    {
      "id": "1218405",
      "postDate": "02/25/2021 20:18:02",
      "content": "<p>this public solutions very awesome 💥🔥</p>",
      "rawMarkdown": "this public solutions very awesome 💥🔥",
      "votes": null
    },
    {
      "id": "1220956",
      "postDate": "02/28/2021 14:36:29",
      "content": "<blockquote>\n  <p>Why didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains.</p>\n</blockquote>\n<p>I was also wondering about this, and my hypothesis so far is \"cutting TP regions and pasting randomly\" (from 4th solution) or \"Non-overlap time Cutmix\" (from 23rd solution). This could work similarly as pseudo labeling; of course, we would need to add fake labels that correspond with the TP regions.<br>\nAnyway, I think you did really well, great work!</p>",
      "rawMarkdown": "> Why didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains.\n\nI was also wondering about this, and my hypothesis so far is \"cutting TP regions and pasting randomly\" (from 4th solution) or \"Non-overlap time Cutmix\" (from 23rd solution). This could work similarly as pseudo labeling; of course, we would need to add fake labels that correspond with the TP regions.\nAnyway, I think you did really well, great work!",
      "votes": null
    },
    {
      "id": "1221143",
      "postDate": "02/28/2021 17:38:06",
      "content": "<p>Thanks for your tips <a href=\"https://www.kaggle.com/daisukelab\" target=\"_blank\">@daisukelab</a> I think u may be right. I intend to dig a little more into the approaches in the coming days. </p>",
      "rawMarkdown": "Thanks for your tips @daisukelab I think u may be right. I intend to dig a little more into the approaches in the coming days.",
      "votes": null
    },
    {
      "id": "1251281",
      "postDate": "03/24/2021 16:33:12",
      "content": "<p><a href=\"https://www.kaggle.com/robertolofaro\" target=\"_blank\">@robertolofaro</a> , <a href=\"https://www.kaggle.com/daisukelab\" target=\"_blank\">@daisukelab</a> : In case you are interested, I made some notes at: <a href=\"https://www.kaggle.com/general/226904\" target=\"_blank\">https://www.kaggle.com/general/226904</a></p>\n<p>…with the unbalanced dataset and the masking need etc, it was a good comp to look at the various loss-engineering techniques. Wish I get a chance again</p>",
      "rawMarkdown": "robertolofaro , @daisukelab : In case you are interested, I made some notes at: https://www.kaggle.com/general/226904\n\n...with the unbalanced dataset and the masking need etc, it was a good comp to look at the various loss-engineering techniques. Wish I get a chance again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210505,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "02/19/2021 13:42:59",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a>,</p>\n<p>Congratulations on the medal! With so many competitions drawing to a close over the last, and in the next, few days I imagine we shall be in for a major new kaggle challenge in the very near future, so please do not  \"<em>…just stick to discussions and exploratory notebooks :)</em>\"</p>\n<p>All the best, and well done,<br>\ncarl</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210715,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/19/2021 16:17:29",
          "content": "<p>Your encouragement in my initial days is one of the strong reasons for me to be on Kaggle, my friend..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210556,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "02/19/2021 14:10:00",
      "content": "<p>I appreciate your mentioning!<br>\nIn this competition, you are my personal fav too :)</p>\n<blockquote>\n  <p>I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data</p>\n</blockquote>\n<p>That was all I did. </p>\n<blockquote>\n  <p>Once I did my first submission, everything changed to just ‘how to improve the score’.</p>\n</blockquote>\n<p>Very true. The attitude of quietly accumulating knowledge by actually getting to depth of things are threatened by the seemingly omnipresent and omnipotent sentence \"the LB does not improve\" in the forum, which, In my opinion, is a pity, along with many other things.</p>\n<p>EDIT: wish you the best in your path of learning!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210578,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/19/2021 14:29:18",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/barnwellguy\" target=\"_blank\">@barnwellguy</a> ! In my limited wanderings on Kaggle, I did come across couple of Masters (actually just one…) who do this. They are not interested in the submissions or scores but participate heartily in the discussions and notebooks, giving brilliant ideas..they dont go beyond the CV…so there is no temptation and focus doesn't shift from 'solution' to 'score'… I was fortunate enough to have a few discussions in RiiiD. with one such Master</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210814,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/19/2021 17:53:39",
      "content": "<p>Congrats on your final score and the medal. Your hard work seems to paid off. I guess Chris blessed all of us with his wisdom :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211859,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/20/2021 16:08:06",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> , u have done tremendously well urself. I think had you cracked how to use FP, you would have moved to Gold..All the best for your next comp!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212925,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "02/21/2021 18:17:49",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> :)</p>\n<p>After reviewing most of the top solutions, it melted down into the pseudo labelling (includes using the FP data) and postprocessing IMHO. </p>\n<p>I wish you the best of luck as well on your journey. Thanks 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212894,
      "author_name": "robertolofaro",
      "author_url": "",
      "post_date": "02/21/2021 17:45:08",
      "content": "<p>Thanks for sharing</p>\n<p>As I too am on a learning path, but way behind your own curve, I appreciate your sharing of your experience</p>\n<p>Learning from somebody else's choices and reasons is really a kind of one-shot transfer learning- accelerates the reader's learning curve😇</p>",
      "votes": null,
      "replies": [
        {
          "id": 1214033,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/22/2021 14:49:12",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/robertolofaro\" target=\"_blank\">@robertolofaro</a> I definitely intend to stay back in RFCX for a while and tinker around with the code and the approaches provided by the winning teams. I will definitely try to publish some of my findings…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251281,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "03/24/2021 16:33:12",
          "content": "<p><a href=\"https://www.kaggle.com/robertolofaro\" target=\"_blank\">@robertolofaro</a> , <a href=\"https://www.kaggle.com/daisukelab\" target=\"_blank\">@daisukelab</a> : In case you are interested, I made some notes at: <a href=\"https://www.kaggle.com/general/226904\" target=\"_blank\">https://www.kaggle.com/general/226904</a></p>\n<p>…with the unbalanced dataset and the masking need etc, it was a good comp to look at the various loss-engineering techniques. Wish I get a chance again</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1218405,
      "author_name": "maozragab",
      "author_url": "",
      "post_date": "02/25/2021 20:18:02",
      "content": "<p>this public solutions very awesome 💥🔥</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1220956,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "02/28/2021 14:36:29",
      "content": "<blockquote>\n  <p>Why didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains.</p>\n</blockquote>\n<p>I was also wondering about this, and my hypothesis so far is \"cutting TP regions and pasting randomly\" (from 4th solution) or \"Non-overlap time Cutmix\" (from 23rd solution). This could work similarly as pseudo labeling; of course, we would need to add fake labels that correspond with the TP regions.<br>\nAnyway, I think you did really well, great work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1221143,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/28/2021 17:38:06",
          "content": "<p>Thanks for your tips <a href=\"https://www.kaggle.com/daisukelab\" target=\"_blank\">@daisukelab</a> I think u may be right. I intend to dig a little more into the approaches in the coming days. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210490": "Post-RiiiD, when my good friend @carlmcbrideellis asked me to join the Jane street competition, I was tempted. It too was a time-series like RiiiD, but then I wanted to wade into something totally new. RFCX with it eco-appeal caught my attention. I was new to both sound processing and CNN… so this seemed like a good place to learn something. Thanks to the wonderful discussion topics by @usharengaraju , @shinmurashinmura and my personal fav @barnwellguy (really sorry you missed the gold) I was off to a good start. Because I learnt Pytorch for RiiiD, I decided to use the TF dataset pipeline and hence the wonderful notebooks by @yoshi999 and @aikhmelnytskyy (who extended it to GPUs) provided a good start. In hindsight though, I realised going with TF was a wrong decision. Most of the winning entries for Birdclef and prev sound competitions were Pytorch based…I could have reused a lot of stuff from there instead of trying to convert. But no regrets…TF is a gold-standard framework and I am glad to have used it. The official documentation seems to be the weak point though. I had absolutely no trouble quickly learning and implementing Pytorch in RiiiD but here, it was taking far more time. But of course this is a beginner’s perspective and I may be wrong.\n\nFew things clearly stood out when I examined the problem and the public solutions:\n- The meta data provided contained the fmin, fmax as shared by the domain experts and I didn’t see any public kernel or discussion on it. I felt this was a key piece of information and mentioned about it in my exploratory kernel, hoping someone might notice and comment on it. My initial thought was to slide along the vertical axis instead of the horizontal axis and then aggregate information..I broke the species into 4 categories based on the f-range. Except one or two species which had huge ranges, all other species fell into neat little freq buckets. So now I could slide along 4-5 buckets of frequency ranges and aggregate species information from all of them (somewhat like what we were doing on the time axis). The question was should I fully invest in building this model because I did have some severe time constraints in these past few weeks. I did an experiment wherein I removed species num 22 (which had max freq of 13K) and then reduced the fmax of the spects from 20K to 12K and re-ran my training. Now my new model should have predicted species 0-21 and 23 with better clarity. I had to blend in species num 22 and I just used the relative rank of species 22 from my original model. So for a particular record, if s22 was ranked 5th in my original model 1, I made it 5th rank in my second model as well. It was a crude blend but I felt it may give me some indication of whether to invest in such a model. Unfortunately, though I reduced fmax from 20K to 12K, there was no improvement in scores. So a big failure.\n\n- Most discussions and kernels revolved around using log Mel specs. These are more for human speech modelling and I felt there could be merit in also experimenting with plain spects, log vs non-log,  hpss and even ensembling results with mfcc spects. I couldnt get hpss to work with tf. As for the rest, nothing worked. \n\n- Somewhere in the middle of the competition I read someone announcing that they got 86% with all ‘1’ combination. For a while I toyed with the idea that maybe I was looking at the competition completely incorrectly. Maybe by default each recording contained nearly all birds and a more easier approach to the comp was finding out which species is NOT in the recording and we should be leveraging only the FP data for this. In my excitement I forgot that all ‘1’s is same as all 0’sand unless there was some bug in the competition evaluation, getting 86 was not possible with all 1s. Quickly abandoned the approach\n\n- Coming to FP, we know that there was nearly 3 times as much data as TP. Definitely it would be a pity to let go of this data. Yet, none of the public kernels/notebooks had given much info on how to leverage this. I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data (Yes I was not ready to let go of that f-range :).  I had earlier abandoned use of 24 diff models - one for each species because this model might not capture inter-dependency between species. But what if I used the output of my first model (with all 24 species ranked) and then augmented that output with the results from the one-model-per-species (simple binary model). I might get the best of both worlds. So assuming I had a 85% score and then I ran my s0 model (levering both s0 TP and s0 FP), I could then compare and augment the original model results with s0 model output. So trained an s0, s3, s13, s19 models to start with and then began blending.  For each model, I would check which are the high probability +ves and -ves and then compare with the original model for those records and make the s’X’ score for that record=highest rank for high +ves and zero for -ves. For e.g. for species 0, I found about 22 records that had high probability of being s0 but in my original model results these 22 records did not contain s0 in top 5 ranks. So I changed the scores of all these 22 records in my original model to be the highest score for that record (thus making sure it got rank 1). Ditto for negatives as well and ditto for several other species. I though species inter-dependencies was anyway captured in the original model and by augmenting it with the confirmed +ves and -ves from the individual 24 models I could take the score higher by leveraging the smaller frequency range associated with each species as well as leveraging the FP data.  Nope..nothing worked and I am still wondering why?\n\nMeanwhile there was significant work pressure and combined with all the above failures, I thought of just taking a break from Kaggle for a few months and logging in later in the year with a fresh perspective. Then 3 things happened:\n- Kaggles ML logic kicked in and it decided to show me an old post from @cdeotte on CNN ensembling on my home page. Now ensembling was not really my top priority but Chris post was pretty interesting. He had suggested ensembling diff sizes like 255,255 & 555,555 & 1024, 1024 etc and mentioned that it was possible for the model to capture different info from diff sizes and the ensemble almost always scores higher. Extending Chris thoughts further, I thought why don’t I try different length and breadth also and maybe this stretching and pulling might result in even more patterns being discovered and I started experiencing good results\n- One of my solo models got 87.6 coupled with some other optimizations\n- I had been trying to parallely fine tune the SED model to get better scores as I wanted to use diff model ensemble in my final submission, but was not having much success. @reppic released a good high scoring model using SED which was just what I wanted..\n\nFrom 87.6% I moved to my final score of 90.1% with minor other optimisations and ensembling….a beginner’s score. Six months back when I was trying to code the Titanic, I was not sure where I would move on next. I had convinced myself at that time that plunging headlong into real competition was the quickest way to learn but I am not so sure of that now.  I guess I should be happy with 2 bronzes in my first 2 comps but the issue is that in both the competitions I was enjoying the learning process UNTIL I made my first submission. Once I did my first submission, everything changed to just ‘how to improve the score’. So maybe for next few comps, I might just stick to discussions and exploratory notebooks :)\n\nOnce again a big shout out to all those kernels and interesting discussions by all folks. Some of my strategies may sound ridiculous but no regrets. I experimented, learnt something and will try to improve..life moves on...\n\nEdit: Some things are beginning to make sense now. Approach 1 (sliding along vertical axis and aggregating species) is sort of a hybrid approach where we can try to capture interdependencies between a set of species if any while parallely leveraging the f-range information to cut out all other noise ranges and thereby hope to make a more accurate prediction. However this may not increase score beyond a certain point and almost all winning teams seem to have used FP and freq cropping at species level...So ideally my approach 4 (last approach)...which has the advantage of using FPs and also a much narrower cropping of frequencies should have worked.. \n\nWhy didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains. However there are others who have reported early-nineties score without using pseudo-labelling. So there is definitely something I need to investigate there. Definitely the FP vastly outnumbered the TPs and maybe this was biasing the model..I had thought of reducing FPs randomly such that FP and TP size was similar but didnt get time to do it. Maybe some approach like that would have shown some gains..",
    "1210505": "Dear @allohvk,\n\nCongratulations on the medal! With so many competitions drawing to a close over the last, and in the next, few days I imagine we shall be in for a major new kaggle challenge in the very near future, so please do not  \"*...just stick to discussions and exploratory notebooks :)*\"\n\nAll the best, and well done,\ncarl",
    "1210556": "I appreciate your mentioning!\nIn this competition, you are my personal fav too :)\n\n>I thought of a ‘grand’ plan that could leverage both FP data and the f-range meta data\n\nThat was all I did. \n\n> Once I did my first submission, everything changed to just ‘how to improve the score’.\n\nVery true. The attitude of quietly accumulating knowledge by actually getting to depth of things are threatened by the seemingly omnipresent and omnipotent sentence \"the LB does not improve\" in the forum, which, In my opinion, is a pity, along with many other things.\n\nEDIT: wish you the best in your path of learning!",
    "1210578": "Thanks @barnwellguy ! In my limited wanderings on Kaggle, I did come across couple of Masters (actually just one...) who do this. They are not interested in the submissions or scores but participate heartily in the discussions and notebooks, giving brilliant ideas..they dont go beyond the CV...so there is no temptation and focus doesn't shift from 'solution' to 'score'... I was fortunate enough to have a few discussions in RiiiD. with one such Master",
    "1210715": "Your encouragement in my initial days is one of the strong reasons for me to be on Kaggle, my friend..",
    "1210814": "Congrats on your final score and the medal. Your hard work seems to paid off. I guess Chris blessed all of us with his wisdom :)",
    "1211859": "Thanks @snnclsr , u have done tremendously well urself. I think had you cracked how to use FP, you would have moved to Gold..All the best for your next comp!!",
    "1212894": "Thanks for sharing\n\nAs I too am on a learning path, but way behind your own curve, I appreciate your sharing of your experience\n\nLearning from somebody else's choices and reasons is really a kind of one-shot transfer learning- accelerates the reader's learning curve😇",
    "1212925": "Thanks a lot @allohvk :)\n\nAfter reviewing most of the top solutions, it melted down into the pseudo labelling (includes using the FP data) and postprocessing IMHO. \n\nI wish you the best of luck as well on your journey. Thanks 🙏",
    "1214033": "Thanks @robertolofaro I definitely intend to stay back in RFCX for a while and tinker around with the code and the approaches provided by the winning teams. I will definitely try to publish some of my findings...",
    "1218405": "this public solutions very awesome 💥🔥",
    "1220956": "> Why didnt it? Going by rank number 2's report, it looked like he could leverage TP and freq crops at species level to take score to 85 and then used pseudo labelling and (maybe) other techniques to take it to 95. If so it makes sense because my model was already at 87% or so and maybe that is why my approach didnt make any significant gains.\n\nI was also wondering about this, and my hypothesis so far is \"cutting TP regions and pasting randomly\" (from 4th solution) or \"Non-overlap time Cutmix\" (from 23rd solution). This could work similarly as pseudo labeling; of course, we would need to add fake labels that correspond with the TP regions.\nAnyway, I think you did really well, great work!",
    "1221143": "Thanks for your tips @daisukelab I think u may be right. I intend to dig a little more into the approaches in the coming days.",
    "1251281": "robertolofaro , @daisukelab : In case you are interested, I made some notes at: https://www.kaggle.com/general/226904\n\n...with the unbalanced dataset and the masking need etc, it was a good comp to look at the various loss-engineering techniques. Wish I get a chance again"
  },
  "source": "meta"
}