{
  "id": 92969,
  "title": "What are your best and worst classes?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/92969",
  "author_name": "",
  "post_date": "2019-05-22T05:58:48.927949Z",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've noticed (based on forum posts) that many teams are using the curated data alone.</p>\n\n<p>I'd like to point out a few things in the hope that this might nudge some of you to make better use of the noisy data (which is also a criterion to win the Judges Award, see Eduardo's post in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88060#latest-534872\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88060#latest-534872</a>)</p>\n\n<p>Observations:</p>\n\n<ul>\n<li><p>Lwlrap is designed as a weighted average of per-class lwlrap. So a natural way of analyzing your lwlrap is to look at your per-class lwlraps on whatever hold-out set you are using for evaluation. The Colab notebook and the official baseline both include code for computing per-class lwlraps.</p></li>\n<li><p>Classes vary widely in label quality and vary somewhat in number of instances. Audio datasets are not like ImageNet or MNIST where you have clean labels in equal numbers for all classes. For some audio classes, it is hard to get good examples in sufficient numbers, and some classes are always very confusable with other classes. Thus, your lwlrap will vary widely by class. E.g., the baseline system we provided has per-class lwlraps ranging from 0.894 (Bicycle bell) to 0.127 (Chirp and tweet) (See <a href=\"https://github.com/DCASE-REPO/dcase2019_task2_baseline\">https://github.com/DCASE-REPO/dcase2019_task2_baseline</a> for the full list). Your worst classes will have a lot of room for improvement. Making predictions for your worst classes better could produce a larger reward than blindly trying class-agnostic methods to improve overall lwlrap with diminishing returns.</p></li>\n<li><p>It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset. For some classes, the noisy dataset might actually be better while for others, the noisy dataset will be much worse. Again, let's use our released baseline as an example and consider four options of (1) using curated data only, (2) using noisy data only, (3) using curated and noisy combined, and (4) using curated warmstarted with noisy. For each of these options, we have several classes where that option was better than the others. No option was uniformly the best across all classes. In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.</p></li>\n</ul>\n\n<p>All of this is to suggest that you might want to start doing a class analysis: see what are your best and worst classes (measured by per-class lwlrap on your held-out set), how these vary depending on what dataset you use (and how you combine the datasets), and see if you can focus on improving your worst classes.</p>\n\n<p>So, what are your best and worst classes?</p>",
  "messages": [
    {
      "id": "534970",
      "postDate": "05/22/2019 05:58:48",
      "content": "<p>I've noticed (based on forum posts) that many teams are using the curated data alone.</p>\n\n<p>I'd like to point out a few things in the hope that this might nudge some of you to make better use of the noisy data (which is also a criterion to win the Judges Award, see Eduardo's post in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88060#latest-534872\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88060#latest-534872</a>)</p>\n\n<p>Observations:</p>\n\n<ul>\n<li><p>Lwlrap is designed as a weighted average of per-class lwlrap. So a natural way of analyzing your lwlrap is to look at your per-class lwlraps on whatever hold-out set you are using for evaluation. The Colab notebook and the official baseline both include code for computing per-class lwlraps.</p></li>\n<li><p>Classes vary widely in label quality and vary somewhat in number of instances. Audio datasets are not like ImageNet or MNIST where you have clean labels in equal numbers for all classes. For some audio classes, it is hard to get good examples in sufficient numbers, and some classes are always very confusable with other classes. Thus, your lwlrap will vary widely by class. E.g., the baseline system we provided has per-class lwlraps ranging from 0.894 (Bicycle bell) to 0.127 (Chirp and tweet) (See <a href=\"https://github.com/DCASE-REPO/dcase2019_task2_baseline\">https://github.com/DCASE-REPO/dcase2019_task2_baseline</a> for the full list). Your worst classes will have a lot of room for improvement. Making predictions for your worst classes better could produce a larger reward than blindly trying class-agnostic methods to improve overall lwlrap with diminishing returns.</p></li>\n<li><p>It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset. For some classes, the noisy dataset might actually be better while for others, the noisy dataset will be much worse. Again, let's use our released baseline as an example and consider four options of (1) using curated data only, (2) using noisy data only, (3) using curated and noisy combined, and (4) using curated warmstarted with noisy. For each of these options, we have several classes where that option was better than the others. No option was uniformly the best across all classes. In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.</p></li>\n</ul>\n\n<p>All of this is to suggest that you might want to start doing a class analysis: see what are your best and worst classes (measured by per-class lwlrap on your held-out set), how these vary depending on what dataset you use (and how you combine the datasets), and see if you can focus on improving your worst classes.</p>\n\n<p>So, what are your best and worst classes?</p>",
      "rawMarkdown": "I've noticed (based on forum posts) that many teams are using the curated data alone.\n\nI'd like to point out a few things in the hope that this might nudge some of you to make better use of the noisy data (which is also a criterion to win the Judges Award, see Eduardo's post in https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88060#latest-534872)\n\nObservations:\n\n- Lwlrap is designed as a weighted average of per-class lwlrap. So a natural way of analyzing your lwlrap is to look at your per-class lwlraps on whatever hold-out set you are using for evaluation. The Colab notebook and the official baseline both include code for computing per-class lwlraps.\n\n\n- Classes vary widely in label quality and vary somewhat in number of instances. Audio datasets are not like ImageNet or MNIST where you have clean labels in equal numbers for all classes. For some audio classes, it is hard to get good examples in sufficient numbers, and some classes are always very confusable with other classes. Thus, your lwlrap will vary widely by class. E.g., the baseline system we provided has per-class lwlraps ranging from 0.894 (Bicycle bell) to 0.127 (Chirp and tweet) (See https://github.com/DCASE-REPO/dcase2019_task2_baseline for the full list). Your worst classes will have a lot of room for improvement. Making predictions for your worst classes better could produce a larger reward than blindly trying class-agnostic methods to improve overall lwlrap with diminishing returns.\n\n- It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset. For some classes, the noisy dataset might actually be better while for others, the noisy dataset will be much worse. Again, let's use our released baseline as an example and consider four options of (1) using curated data only, (2) using noisy data only, (3) using curated and noisy combined, and (4) using curated warmstarted with noisy. For each of these options, we have several classes where that option was better than the others. No option was uniformly the best across all classes. In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.\n\nAll of this is to suggest that you might want to start doing a class analysis: see what are your best and worst classes (measured by per-class lwlrap on your held-out set), how these vary depending on what dataset you use (and how you combine the datasets), and see if you can focus on improving your worst classes.\n\nSo, what are your best and worst classes?",
      "votes": null
    },
    {
      "id": "535019",
      "postDate": "05/22/2019 06:49:39",
      "content": "<p>Here is my lwlrap per class (worst to best):</p>\n\n<p>| | lwlrap | weight |\n| ---| --- | --- |\n| Squeak | 0.583473 | 0.013039 |\nFill (with liquid) | 0.611167 | 0.008693\nWalk and footsteps | 0.623984 | 0.013039\nTraffic noise and roadway noise | 0.711915 | 0.013039\nHiss | 0.716003 | 0.013039\nTap | 0.719197 | 0.013039\nChink and clink | 0.736545 | 0.013039\nBuzz | 0.746997 | 0.009736\nCutlery and silverware | 0.762658 | 0.013039\nMechanical fan | 0.766814 | 0.008519\nWater tap and faucet | 0.780234 | 0.013039\nMale speech and man speaking | 0.780767 | 0.013039\nTrickle and dribble | 0.796176 | 0.009214\nFrying (food) | 0.801226 | 0.010953\nSlam | 0.802958 | 0.013039\nAccelerating and revving and vroom | 0.804895 | 0.013039\nBus | 0.805716 | 0.013039\nYell | 0.807863 | 0.013039\nClapping | 0.809714 | 0.013039\nMicrowave oven | 0.811185 | 0.013039\nSink (filling or washing) | 0.813332 | 0.013039\nDishes and pots and pans | 0.817135 | 0.013039\nSneeze | 0.821423 | 0.010953\nScissors | 0.823593 | 0.013039\nStream | 0.823926 | 0.013039\nKnock | 0.836556 | 0.013039\nBathtub (filling or washing) | 0.842000 | 0.013039\nMeow | 0.848685 | 0.013039\nCar passing by | 0.849016 | 0.013039\nMotorcycle | 0.852556 | 0.013039\n... | ... | ...\nZipper (clothing) | 0.911794 | 0.013039\nApplause | 0.913556 | 0.013039\nCrackle | 0.915352 | 0.013039\nElectric guitar | 0.916763 | 0.013039\nShatter | 0.921229 | 0.013039\nRace car and auto racing | 0.921557 | 0.009736\nRaindrop | 0.923667 | 0.013039\nWriting | 0.925614 | 0.013039\nTick-tock | 0.927778 | 0.012517\nFart | 0.928015 | 0.013039\nBark | 0.928307 | 0.013039\nFemale singing | 0.929839 | 0.013039\nBicycle bell | 0.934494 | 0.011648\nSigh | 0.936705 | 0.009910\nChurch bell | 0.939111 | 0.013039\nChild speech and kid speaking | 0.944943 | 0.013039\nMarimba and xylophone | 0.945556 | 0.013039\nAccordion | 0.948582 | 0.008171\nToilet flush | 0.950000 | 0.013039\nHarmonica | 0.950747 | 0.013039\nBass guitar | 0.950794 | 0.013039\nBurping and eructation | 0.955737 | 0.013039\nHi-hat | 0.956902 | 0.013039\nBass drum | 0.958767 | 0.013039\nGlockenspiel | 0.961310 | 0.009736\nAcoustic guitar | 0.961813 | 0.013039\nPurr | 0.963333 | 0.011300\nFinger snapping | 0.967111 | 0.013039\nSkateboard | 0.993333 | 0.013039\nStrum | 1.000000 | 0.013039</p>",
      "rawMarkdown": "Here is my lwlrap per class (worst to best):\n\n| | lwlrap | weight |\n| ---| --- | --- |\n| Squeak | 0.583473 | 0.013039 |\nFill (with liquid) | 0.611167 | 0.008693\nWalk and footsteps | 0.623984 | 0.013039\nTraffic noise and roadway noise | 0.711915 | 0.013039\nHiss | 0.716003 | 0.013039\nTap | 0.719197 | 0.013039\nChink and clink | 0.736545 | 0.013039\nBuzz | 0.746997 | 0.009736\nCutlery and silverware | 0.762658 | 0.013039\nMechanical fan | 0.766814 | 0.008519\nWater tap and faucet | 0.780234 | 0.013039\nMale speech and man speaking | 0.780767 | 0.013039\nTrickle and dribble | 0.796176 | 0.009214\nFrying (food) | 0.801226 | 0.010953\nSlam | 0.802958 | 0.013039\nAccelerating and revving and vroom | 0.804895 | 0.013039\nBus | 0.805716 | 0.013039\nYell | 0.807863 | 0.013039\nClapping | 0.809714 | 0.013039\nMicrowave oven | 0.811185 | 0.013039\nSink (filling or washing) | 0.813332 | 0.013039\nDishes and pots and pans | 0.817135 | 0.013039\nSneeze | 0.821423 | 0.010953\nScissors | 0.823593 | 0.013039\nStream | 0.823926 | 0.013039\nKnock | 0.836556 | 0.013039\nBathtub (filling or washing) | 0.842000 | 0.013039\nMeow | 0.848685 | 0.013039\nCar passing by | 0.849016 | 0.013039\nMotorcycle | 0.852556 | 0.013039\n... | ... | ...\nZipper (clothing) | 0.911794 | 0.013039\nApplause | 0.913556 | 0.013039\nCrackle | 0.915352 | 0.013039\nElectric guitar | 0.916763 | 0.013039\nShatter | 0.921229 | 0.013039\nRace car and auto racing | 0.921557 | 0.009736\nRaindrop | 0.923667 | 0.013039\nWriting | 0.925614 | 0.013039\nTick-tock | 0.927778 | 0.012517\nFart | 0.928015 | 0.013039\nBark | 0.928307 | 0.013039\nFemale singing | 0.929839 | 0.013039\nBicycle bell | 0.934494 | 0.011648\nSigh | 0.936705 | 0.009910\nChurch bell | 0.939111 | 0.013039\nChild speech and kid speaking | 0.944943 | 0.013039\nMarimba and xylophone | 0.945556 | 0.013039\nAccordion | 0.948582 | 0.008171\nToilet flush | 0.950000 | 0.013039\nHarmonica | 0.950747 | 0.013039\nBass guitar | 0.950794 | 0.013039\nBurping and eructation | 0.955737 | 0.013039\nHi-hat | 0.956902 | 0.013039\nBass drum | 0.958767 | 0.013039\nGlockenspiel | 0.961310 | 0.009736\nAcoustic guitar | 0.961813 | 0.013039\nPurr | 0.963333 | 0.011300\nFinger snapping | 0.967111 | 0.013039\nSkateboard | 0.993333 | 0.013039\nStrum | 1.000000 | 0.013039",
      "votes": null
    },
    {
      "id": "535205",
      "postDate": "05/22/2019 13:28:59",
      "content": "<p>I am only using train-curated for now but tried this technique on various models using that data alone. Unsurprising result: Local lwlrap up 0.01+ and PubLB down 0.01+ </p>\n\n<p>I will have to try on noisy but am not optimistic!\nWorst: Fill with liquid and Squeak around 62. Then walk with footsteps 72.\nBest: Strum 100 Accordion 98 Bicycle bell 97</p>",
      "rawMarkdown": "I am only using train-curated for now but tried this technique on various models using that data alone. Unsurprising result: Local lwlrap up 0.01+ and PubLB down 0.01+ \n\nI will have to try on noisy but am not optimistic!\nWorst: Fill with liquid and Squeak around 62. Then walk with footsteps 72.\nBest: Strum 100 Accordion 98 Bicycle bell 97",
      "votes": null
    },
    {
      "id": "535296",
      "postDate": "05/22/2019 16:27:55",
      "content": "<p>Awesome post, Manoj!\nFor those interested in the Judges' Award, rules here:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award\">https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award</a></p>",
      "rawMarkdown": "Awesome post, Manoj!\nFor those interested in the Judges' Award, rules here:\n[https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award](https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award)",
      "votes": null
    },
    {
      "id": "535551",
      "postDate": "05/23/2019 06:00:57",
      "content": "<p>What do you mean exactly when you say you tried \"this technique\" using curated data alone? What I'm trying to encourage is a more judicious use of the noisy data to help out wherever the curated isn't enough, so I'm not sure what you do if you aren't using noisy data.</p>\n\n<p>The per-class lwlraps that you and Eric posted are also helpful in another way: to show that you might be overfitting at the class level to the training set, even if the overall lwlrap is not outrageously high. It looks like both of your models have perfectly memorized the training data for Strum, for example.  You could try increasing regularization to avoid extreme overfitting for certain classes.</p>",
      "rawMarkdown": "What do you mean exactly when you say you tried \"this technique\" using curated data alone? What I'm trying to encourage is a more judicious use of the noisy data to help out wherever the curated isn't enough, so I'm not sure what you do if you aren't using noisy data.\n\nThe per-class lwlraps that you and Eric posted are also helpful in another way: to show that you might be overfitting at the class level to the training set, even if the overall lwlrap is not outrageously high. It looks like both of your models have perfectly memorized the training data for Strum, for example.  You could try increasing regularization to avoid extreme overfitting for certain classes.",
      "votes": null
    },
    {
      "id": "535603",
      "postDate": "05/23/2019 07:57:05",
      "content": "<p>I've been able to get small improvements from incorporating noisy data (0.002~0.01). Not really significant, but at least not worse.</p>\n\n<p>Worst classes: Squeak, Fill (with liquid), Hiss, Walk and footsteps, Bathtub (filling or washing)\nBest classes: Accordion, Skateboard, Finger snapping, Strum</p>\n\n<p>(Not from the current best model, but should be be very similar. )</p>\n\n<blockquote>\n  <p>It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset... In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.</p>\n</blockquote>\n\n<p>This is very interesting. Since the curated labels are not reliable for those classes, I guess we can only find those classes by sampling the audio clips manually?</p>",
      "rawMarkdown": "I've been able to get small improvements from incorporating noisy data (0.002~0.01). Not really significant, but at least not worse.\n\nWorst classes: Squeak, Fill (with liquid), Hiss, Walk and footsteps, Bathtub (filling or washing)\nBest classes: Accordion, Skateboard, Finger snapping, Strum\n\n(Not from the current best model, but should be be very similar. )\n\n&gt; It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset... In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.\n\nThis is very interesting. Since the curated labels are not reliable for those classes, I guess we can only find those classes by sampling the audio clips manually?",
      "votes": null
    },
    {
      "id": "535809",
      "postDate": "05/23/2019 13:33:11",
      "content": "<p>It looks like the host is worried about nobody uses noisy data. I think I can help the host. Our best single model scores LwLRAP 0.848 on OOF data.  Training without noisy data, this model scores LwLRAP 0.832. I won't say how I use noisy data until this competition is over, but yes, I think noisy data is the key to win this competition.</p>",
      "rawMarkdown": "It looks like the host is worried about nobody uses noisy data. I think I can help the host. Our best single model scores LwLRAP 0.848 on OOF data.  Training without noisy data, this model scores LwLRAP 0.832. I won't say how I use noisy data until this competition is over, but yes, I think noisy data is the key to win this competition.",
      "votes": null
    },
    {
      "id": "535899",
      "postDate": "05/23/2019 15:34:29",
      "content": "<p>Nice work :)</p>",
      "rawMarkdown": "Nice work :)",
      "votes": null
    },
    {
      "id": "535968",
      "postDate": "05/23/2019 17:59:32",
      "content": "<p>You shouldn't need to do much manual work. You could see how a model trained on just the noisy data performs on some or all of the curated train set, which would give you a rough picture of per-class performance of pure noisy training. That should be a good starting point for further analysis.</p>\n\n<p>To clarify what I meant in my original post, I am comparing performance of curated vs noisy using per-class lwlrap on the test set. So it's not the case that curated labels are worthless, it's just that for some classes, the noisy labels are better.</p>",
      "rawMarkdown": "You shouldn't need to do much manual work. You could see how a model trained on just the noisy data performs on some or all of the curated train set, which would give you a rough picture of per-class performance of pure noisy training. That should be a good starting point for further analysis.\n\nTo clarify what I meant in my original post, I am comparing performance of curated vs noisy using per-class lwlrap on the test set. So it's not the case that curated labels are worthless, it's just that for some classes, the noisy labels are better.",
      "votes": null
    },
    {
      "id": "536154",
      "postDate": "05/24/2019 03:17:15",
      "content": "<p>Thanks for the clarification. I did not mean the curated labels are worthless, but that they are not more representative for the test set than the noisy ones (for those classes). Now I get that you were actually saying not all noisy labels are worthless.</p>",
      "rawMarkdown": "Thanks for the clarification. I did not mean the curated labels are worthless, but that they are not more representative for the test set than the noisy ones (for those classes). Now I get that you were actually saying not all noisy labels are worthless.",
      "votes": null
    },
    {
      "id": "536416",
      "postDate": "05/24/2019 12:40:35",
      "content": "<p>Exactly, not all the noisy labels are worthless at all. The level of label noise in the noisy set varies highly class-wise. Also, as we explain in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data Section</a>:</p>\n\n<blockquote>\n  <p>The per-class data distribution available for training is, for most of the classes, 300 clips from the noisy subset and 75 clips from the curated subset</p>\n</blockquote>\n\n<p>So, in the classes where the label noise is not large, you may have a good amount of interesting data in the noisy train set. Measures to cope with or mitigate the label noise are a key ingredient to success in this task.</p>\n\n<p>Another issue to consider: the acoustic domain of the noisy set is slightly different to that of the curated train set, and test set. This is another key ingredient as we explained in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data Section</a>:</p>\n\n<blockquote>\n  <p>The test set is used for system evaluation and consists of manually-labeled data from FSD. Since most of the train data come from YFCC, some acoustic domain mismatch between the train and test set can be expected.</p>\n</blockquote>\n\n<p>hope this helps</p>",
      "rawMarkdown": "Exactly, not all the noisy labels are worthless at all. The level of label noise in the noisy set varies highly class-wise. Also, as we explain in the [Data Section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data):\n\n&gt; The per-class data distribution available for training is, for most of the classes, 300 clips from the noisy subset and 75 clips from the curated subset\n\nSo, in the classes where the label noise is not large, you may have a good amount of interesting data in the noisy train set. Measures to cope with or mitigate the label noise are a key ingredient to success in this task.\n\nAnother issue to consider: the acoustic domain of the noisy set is slightly different to that of the curated train set, and test set. This is another key ingredient as we explained in the [Data Section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data):\n\n&gt; The test set is used for system evaluation and consists of manually-labeled data from FSD. Since most of the train data come from YFCC, some acoustic domain mismatch between the train and test set can be expected.\n\nhope this helps",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 535019,
      "author_name": "ebouteillon",
      "author_url": "",
      "post_date": "05/22/2019 06:49:39",
      "content": "<p>Here is my lwlrap per class (worst to best):</p>\n\n<p>| | lwlrap | weight |\n| ---| --- | --- |\n| Squeak | 0.583473 | 0.013039 |\nFill (with liquid) | 0.611167 | 0.008693\nWalk and footsteps | 0.623984 | 0.013039\nTraffic noise and roadway noise | 0.711915 | 0.013039\nHiss | 0.716003 | 0.013039\nTap | 0.719197 | 0.013039\nChink and clink | 0.736545 | 0.013039\nBuzz | 0.746997 | 0.009736\nCutlery and silverware | 0.762658 | 0.013039\nMechanical fan | 0.766814 | 0.008519\nWater tap and faucet | 0.780234 | 0.013039\nMale speech and man speaking | 0.780767 | 0.013039\nTrickle and dribble | 0.796176 | 0.009214\nFrying (food) | 0.801226 | 0.010953\nSlam | 0.802958 | 0.013039\nAccelerating and revving and vroom | 0.804895 | 0.013039\nBus | 0.805716 | 0.013039\nYell | 0.807863 | 0.013039\nClapping | 0.809714 | 0.013039\nMicrowave oven | 0.811185 | 0.013039\nSink (filling or washing) | 0.813332 | 0.013039\nDishes and pots and pans | 0.817135 | 0.013039\nSneeze | 0.821423 | 0.010953\nScissors | 0.823593 | 0.013039\nStream | 0.823926 | 0.013039\nKnock | 0.836556 | 0.013039\nBathtub (filling or washing) | 0.842000 | 0.013039\nMeow | 0.848685 | 0.013039\nCar passing by | 0.849016 | 0.013039\nMotorcycle | 0.852556 | 0.013039\n... | ... | ...\nZipper (clothing) | 0.911794 | 0.013039\nApplause | 0.913556 | 0.013039\nCrackle | 0.915352 | 0.013039\nElectric guitar | 0.916763 | 0.013039\nShatter | 0.921229 | 0.013039\nRace car and auto racing | 0.921557 | 0.009736\nRaindrop | 0.923667 | 0.013039\nWriting | 0.925614 | 0.013039\nTick-tock | 0.927778 | 0.012517\nFart | 0.928015 | 0.013039\nBark | 0.928307 | 0.013039\nFemale singing | 0.929839 | 0.013039\nBicycle bell | 0.934494 | 0.011648\nSigh | 0.936705 | 0.009910\nChurch bell | 0.939111 | 0.013039\nChild speech and kid speaking | 0.944943 | 0.013039\nMarimba and xylophone | 0.945556 | 0.013039\nAccordion | 0.948582 | 0.008171\nToilet flush | 0.950000 | 0.013039\nHarmonica | 0.950747 | 0.013039\nBass guitar | 0.950794 | 0.013039\nBurping and eructation | 0.955737 | 0.013039\nHi-hat | 0.956902 | 0.013039\nBass drum | 0.958767 | 0.013039\nGlockenspiel | 0.961310 | 0.009736\nAcoustic guitar | 0.961813 | 0.013039\nPurr | 0.963333 | 0.011300\nFinger snapping | 0.967111 | 0.013039\nSkateboard | 0.993333 | 0.013039\nStrum | 1.000000 | 0.013039</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 535205,
      "author_name": "robga",
      "author_url": "",
      "post_date": "05/22/2019 13:28:59",
      "content": "<p>I am only using train-curated for now but tried this technique on various models using that data alone. Unsurprising result: Local lwlrap up 0.01+ and PubLB down 0.01+ </p>\n\n<p>I will have to try on noisy but am not optimistic!\nWorst: Fill with liquid and Squeak around 62. Then walk with footsteps 72.\nBest: Strum 100 Accordion 98 Bicycle bell 97</p>",
      "votes": null,
      "replies": [
        {
          "id": 535551,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "05/23/2019 06:00:57",
          "content": "<p>What do you mean exactly when you say you tried \"this technique\" using curated data alone? What I'm trying to encourage is a more judicious use of the noisy data to help out wherever the curated isn't enough, so I'm not sure what you do if you aren't using noisy data.</p>\n\n<p>The per-class lwlraps that you and Eric posted are also helpful in another way: to show that you might be overfitting at the class level to the training set, even if the overall lwlrap is not outrageously high. It looks like both of your models have perfectly memorized the training data for Strum, for example.  You could try increasing regularization to avoid extreme overfitting for certain classes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 535296,
      "author_name": "eduardofonseca",
      "author_url": "",
      "post_date": "05/22/2019 16:27:55",
      "content": "<p>Awesome post, Manoj!\nFor those interested in the Judges' Award, rules here:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award\">https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 535603,
      "author_name": "ceshine",
      "author_url": "",
      "post_date": "05/23/2019 07:57:05",
      "content": "<p>I've been able to get small improvements from incorporating noisy data (0.002~0.01). Not really significant, but at least not worse.</p>\n\n<p>Worst classes: Squeak, Fill (with liquid), Hiss, Walk and footsteps, Bathtub (filling or washing)\nBest classes: Accordion, Skateboard, Finger snapping, Strum</p>\n\n<p>(Not from the current best model, but should be be very similar. )</p>\n\n<blockquote>\n  <p>It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset... In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.</p>\n</blockquote>\n\n<p>This is very interesting. Since the curated labels are not reliable for those classes, I guess we can only find those classes by sampling the audio clips manually?</p>",
      "votes": null,
      "replies": [
        {
          "id": 535968,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "05/23/2019 17:59:32",
          "content": "<p>You shouldn't need to do much manual work. You could see how a model trained on just the noisy data performs on some or all of the curated train set, which would give you a rough picture of per-class performance of pure noisy training. That should be a good starting point for further analysis.</p>\n\n<p>To clarify what I meant in my original post, I am comparing performance of curated vs noisy using per-class lwlrap on the test set. So it's not the case that curated labels are worthless, it's just that for some classes, the noisy labels are better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536154,
          "author_name": "ceshine",
          "author_url": "",
          "post_date": "05/24/2019 03:17:15",
          "content": "<p>Thanks for the clarification. I did not mean the curated labels are worthless, but that they are not more representative for the test set than the noisy ones (for those classes). Now I get that you were actually saying not all noisy labels are worthless.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536416,
          "author_name": "eduardofonseca",
          "author_url": "",
          "post_date": "05/24/2019 12:40:35",
          "content": "<p>Exactly, not all the noisy labels are worthless at all. The level of label noise in the noisy set varies highly class-wise. Also, as we explain in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data Section</a>:</p>\n\n<blockquote>\n  <p>The per-class data distribution available for training is, for most of the classes, 300 clips from the noisy subset and 75 clips from the curated subset</p>\n</blockquote>\n\n<p>So, in the classes where the label noise is not large, you may have a good amount of interesting data in the noisy train set. Measures to cope with or mitigate the label noise are a key ingredient to success in this task.</p>\n\n<p>Another issue to consider: the acoustic domain of the noisy set is slightly different to that of the curated train set, and test set. This is another key ingredient as we explained in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">Data Section</a>:</p>\n\n<blockquote>\n  <p>The test set is used for system evaluation and consists of manually-labeled data from FSD. Since most of the train data come from YFCC, some acoustic domain mismatch between the train and test set can be expected.</p>\n</blockquote>\n\n<p>hope this helps</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 535809,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "05/23/2019 13:33:11",
      "content": "<p>It looks like the host is worried about nobody uses noisy data. I think I can help the host. Our best single model scores LwLRAP 0.848 on OOF data.  Training without noisy data, this model scores LwLRAP 0.832. I won't say how I use noisy data until this competition is over, but yes, I think noisy data is the key to win this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 535899,
          "author_name": "plakal",
          "author_url": "",
          "post_date": "05/23/2019 15:34:29",
          "content": "<p>Nice work :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "534970": "I've noticed (based on forum posts) that many teams are using the curated data alone.\n\nI'd like to point out a few things in the hope that this might nudge some of you to make better use of the noisy data (which is also a criterion to win the Judges Award, see Eduardo's post in https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88060#latest-534872)\n\nObservations:\n\n- Lwlrap is designed as a weighted average of per-class lwlrap. So a natural way of analyzing your lwlrap is to look at your per-class lwlraps on whatever hold-out set you are using for evaluation. The Colab notebook and the official baseline both include code for computing per-class lwlraps.\n\n\n- Classes vary widely in label quality and vary somewhat in number of instances. Audio datasets are not like ImageNet or MNIST where you have clean labels in equal numbers for all classes. For some audio classes, it is hard to get good examples in sufficient numbers, and some classes are always very confusable with other classes. Thus, your lwlrap will vary widely by class. E.g., the baseline system we provided has per-class lwlraps ranging from 0.894 (Bicycle bell) to 0.127 (Chirp and tweet) (See https://github.com/DCASE-REPO/dcase2019_task2_baseline for the full list). Your worst classes will have a lot of room for improvement. Making predictions for your worst classes better could produce a larger reward than blindly trying class-agnostic methods to improve overall lwlrap with diminishing returns.\n\n- It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset. For some classes, the noisy dataset might actually be better while for others, the noisy dataset will be much worse. Again, let's use our released baseline as an example and consider four options of (1) using curated data only, (2) using noisy data only, (3) using curated and noisy combined, and (4) using curated warmstarted with noisy. For each of these options, we have several classes where that option was better than the others. No option was uniformly the best across all classes. In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.\n\nAll of this is to suggest that you might want to start doing a class analysis: see what are your best and worst classes (measured by per-class lwlrap on your held-out set), how these vary depending on what dataset you use (and how you combine the datasets), and see if you can focus on improving your worst classes.\n\nSo, what are your best and worst classes?",
    "535019": "Here is my lwlrap per class (worst to best):\n\n| | lwlrap | weight |\n| ---| --- | --- |\n| Squeak | 0.583473 | 0.013039 |\nFill (with liquid) | 0.611167 | 0.008693\nWalk and footsteps | 0.623984 | 0.013039\nTraffic noise and roadway noise | 0.711915 | 0.013039\nHiss | 0.716003 | 0.013039\nTap | 0.719197 | 0.013039\nChink and clink | 0.736545 | 0.013039\nBuzz | 0.746997 | 0.009736\nCutlery and silverware | 0.762658 | 0.013039\nMechanical fan | 0.766814 | 0.008519\nWater tap and faucet | 0.780234 | 0.013039\nMale speech and man speaking | 0.780767 | 0.013039\nTrickle and dribble | 0.796176 | 0.009214\nFrying (food) | 0.801226 | 0.010953\nSlam | 0.802958 | 0.013039\nAccelerating and revving and vroom | 0.804895 | 0.013039\nBus | 0.805716 | 0.013039\nYell | 0.807863 | 0.013039\nClapping | 0.809714 | 0.013039\nMicrowave oven | 0.811185 | 0.013039\nSink (filling or washing) | 0.813332 | 0.013039\nDishes and pots and pans | 0.817135 | 0.013039\nSneeze | 0.821423 | 0.010953\nScissors | 0.823593 | 0.013039\nStream | 0.823926 | 0.013039\nKnock | 0.836556 | 0.013039\nBathtub (filling or washing) | 0.842000 | 0.013039\nMeow | 0.848685 | 0.013039\nCar passing by | 0.849016 | 0.013039\nMotorcycle | 0.852556 | 0.013039\n... | ... | ...\nZipper (clothing) | 0.911794 | 0.013039\nApplause | 0.913556 | 0.013039\nCrackle | 0.915352 | 0.013039\nElectric guitar | 0.916763 | 0.013039\nShatter | 0.921229 | 0.013039\nRace car and auto racing | 0.921557 | 0.009736\nRaindrop | 0.923667 | 0.013039\nWriting | 0.925614 | 0.013039\nTick-tock | 0.927778 | 0.012517\nFart | 0.928015 | 0.013039\nBark | 0.928307 | 0.013039\nFemale singing | 0.929839 | 0.013039\nBicycle bell | 0.934494 | 0.011648\nSigh | 0.936705 | 0.009910\nChurch bell | 0.939111 | 0.013039\nChild speech and kid speaking | 0.944943 | 0.013039\nMarimba and xylophone | 0.945556 | 0.013039\nAccordion | 0.948582 | 0.008171\nToilet flush | 0.950000 | 0.013039\nHarmonica | 0.950747 | 0.013039\nBass guitar | 0.950794 | 0.013039\nBurping and eructation | 0.955737 | 0.013039\nHi-hat | 0.956902 | 0.013039\nBass drum | 0.958767 | 0.013039\nGlockenspiel | 0.961310 | 0.009736\nAcoustic guitar | 0.961813 | 0.013039\nPurr | 0.963333 | 0.011300\nFinger snapping | 0.967111 | 0.013039\nSkateboard | 0.993333 | 0.013039\nStrum | 1.000000 | 0.013039",
    "535205": "I am only using train-curated for now but tried this technique on various models using that data alone. Unsurprising result: Local lwlrap up 0.01+ and PubLB down 0.01+ \n\nI will have to try on noisy but am not optimistic!\nWorst: Fill with liquid and Squeak around 62. Then walk with footsteps 72.\nBest: Strum 100 Accordion 98 Bicycle bell 97",
    "535296": "Awesome post, Manoj!\nFor those interested in the Judges' Award, rules here:\n[https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award](https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/judges-award)",
    "535551": "What do you mean exactly when you say you tried \"this technique\" using curated data alone? What I'm trying to encourage is a more judicious use of the noisy data to help out wherever the curated isn't enough, so I'm not sure what you do if you aren't using noisy data.\n\nThe per-class lwlraps that you and Eric posted are also helpful in another way: to show that you might be overfitting at the class level to the training set, even if the overall lwlrap is not outrageously high. It looks like both of your models have perfectly memorized the training data for Strum, for example.  You could try increasing regularization to avoid extreme overfitting for certain classes.",
    "535603": "I've been able to get small improvements from incorporating noisy data (0.002~0.01). Not really significant, but at least not worse.\n\nWorst classes: Squeak, Fill (with liquid), Hiss, Walk and footsteps, Bathtub (filling or washing)\nBest classes: Accordion, Skateboard, Finger snapping, Strum\n\n(Not from the current best model, but should be be very similar. )\n\n&gt; It is not uniformly true for all classes that the curated dataset has better labels and training value than the noisy dataset... In particular, there are classes where using curated data only is the worst option and there are classes where using noisy data only is better than even using curated and noisy together.\n\nThis is very interesting. Since the curated labels are not reliable for those classes, I guess we can only find those classes by sampling the audio clips manually?",
    "535809": "It looks like the host is worried about nobody uses noisy data. I think I can help the host. Our best single model scores LwLRAP 0.848 on OOF data.  Training without noisy data, this model scores LwLRAP 0.832. I won't say how I use noisy data until this competition is over, but yes, I think noisy data is the key to win this competition.",
    "535899": "Nice work :)",
    "535968": "You shouldn't need to do much manual work. You could see how a model trained on just the noisy data performs on some or all of the curated train set, which would give you a rough picture of per-class performance of pure noisy training. That should be a good starting point for further analysis.\n\nTo clarify what I meant in my original post, I am comparing performance of curated vs noisy using per-class lwlrap on the test set. So it's not the case that curated labels are worthless, it's just that for some classes, the noisy labels are better.",
    "536154": "Thanks for the clarification. I did not mean the curated labels are worthless, but that they are not more representative for the test set than the noisy ones (for those classes). Now I get that you were actually saying not all noisy labels are worthless.",
    "536416": "Exactly, not all the noisy labels are worthless at all. The level of label noise in the noisy set varies highly class-wise. Also, as we explain in the [Data Section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data):\n\n&gt; The per-class data distribution available for training is, for most of the classes, 300 clips from the noisy subset and 75 clips from the curated subset\n\nSo, in the classes where the label noise is not large, you may have a good amount of interesting data in the noisy train set. Measures to cope with or mitigate the label noise are a key ingredient to success in this task.\n\nAnother issue to consider: the acoustic domain of the noisy set is slightly different to that of the curated train set, and test set. This is another key ingredient as we explained in the [Data Section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data):\n\n&gt; The test set is used for system evaluation and consists of manually-labeled data from FSD. Since most of the train data come from YFCC, some acoustic domain mismatch between the train and test set can be expected.\n\nhope this helps"
  },
  "source": "meta"
}