{
  "id": 95369,
  "title": "How did you handle noisy data?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/95369",
  "author_name": "Darkate",
  "post_date": "2019-06-11T19:46:46.400000",
  "votes": 9,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I am very interested in how everyone deals with noisy data. Using noisy labels? Pseudo-label? MixMatch? UDA? Other SSL approaches?</p>\n\n<p>I was using exactly the same model and training step here except the backbone and the non-linear activation function (PReLU instead of Sigmoid): <a href=\"https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26\">https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26</a></p>\n\n<p>There was a free pdf version of this paper but they removed it from public server.</p>",
  "messages": [
    {
      "id": 1731629,
      "postDate": "2022-03-22T14:50:04.847Z",
      "content": "<p>Anyone trying out this now can check out this audio tagging <a href=\"https://paperswithcode.com/paper/audio-tagging-with-noisy-labels-and-minimal\" target=\"_blank\">paper</a>: </p>\n<p>Also see: most implemented <a href=\"https://paperswithcode.com/task/audio-tagging\" target=\"_blank\">papers</a> on audio tagging. For checking out the performance of different algorithms to see which puts the best results on the table, I’d suggest <a href=\"https://paperswithcode.com/task/audio-tagging\" target=\"_blank\">deepchecks</a> </p>\n<p>Wishing everyone success! </p>",
      "rawMarkdown": "Anyone trying out this now can check out this audio tagging [paper](https://paperswithcode.com/paper/audio-tagging-with-noisy-labels-and-minimal): \n\nAlso see: most implemented [papers](https://paperswithcode.com/task/audio-tagging) on audio tagging. For checking out the performance of different algorithms to see which puts the best results on the table, I’d suggest [deepchecks](https://paperswithcode.com/task/audio-tagging) \n\nWishing everyone success! \n",
      "votes": 3
    },
    {
      "id": 550599,
      "postDate": "2019-06-11T19:46:46.400Z",
      "content": "<p>I am very interested in how everyone deals with noisy data. Using noisy labels? Pseudo-label? MixMatch? UDA? Other SSL approaches?</p>\n\n<p>I was using exactly the same model and training step here except the backbone and the non-linear activation function (PReLU instead of Sigmoid): <a href=\"https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26\">https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26</a></p>\n\n<p>There was a free pdf version of this paper but they removed it from public server.</p>",
      "rawMarkdown": "I am very interested in how everyone deals with noisy data. Using noisy labels? Pseudo-label? MixMatch? UDA? Other SSL approaches?\n\nI was using exactly the same model and training step here except the backbone and the non-linear activation function (PReLU instead of Sigmoid): https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26\n\nThere was a free pdf version of this paper but they removed it from public server.",
      "votes": 8
    },
    {
      "id": 551607,
      "postDate": "2019-06-12T23:06:47.483Z",
      "content": "<p>We added our model additional FC layers for noisy data and did multitask learning.\nSingle task 5-fold CV LwLRAP: 0.829\nMultitask 5-fold CV LwLRAP: 0.849</p>",
      "rawMarkdown": "We added our model additional FC layers for noisy data and did multitask learning.\nSingle task 5-fold CV LwLRAP: 0.829\nMultitask 5-fold CV LwLRAP: 0.849",
      "votes": 4,
      "replies": [
        {
          "id": 552098,
          "postDate": "2019-06-13T14:07:16.700Z",
          "content": "<p>How do you set the multi tasks? Do you treat the multi-label as multi-task or detection and classification tasks?</p>",
          "rawMarkdown": "How do you set the multi tasks? Do you treat the multi-label as multi-task or detection and classification tasks?"
        },
        {
          "id": 552212,
          "postDate": "2019-06-13T16:31:40.453Z",
          "content": "<p>It's like this.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/552212/13515/multitask%20pipeline.png\" alt=\"Multitask pipeline\"></p>",
          "rawMarkdown": "It's like this.\n![Multitask pipeline](https://storage.googleapis.com/kaggle-forum-message-attachments/552212/13515/multitask%20pipeline.png)",
          "votes": 11
        },
        {
          "id": 552418,
          "postDate": "2019-06-13T23:34:17.823Z",
          "content": "<p>Nice!\nSo it would contain two sample per each training step input like a siamese NN and two set of labels to calculate the loss during training?\nBut why? What is the advantage of this method? </p>",
          "rawMarkdown": "Nice!\nSo it would contain two sample per each training step input like a siamese NN and two set of labels to calculate the loss during training?\nBut why? What is the advantage of this method? ",
          "votes": 1
        },
        {
          "id": 552467,
          "postDate": "2019-06-14T02:39:49.617Z",
          "content": "<p><a href=\"/osciiart\">@osciiart</a>, quite interesting idea. I'm curious in which way did u add MixUp to this scheme? Did u mix only noisy labels with noisy and curated with curated or just everything together? </p>",
          "rawMarkdown": "@osciiart, quite interesting idea. I'm curious in which way did u add MixUp to this scheme? Did u mix only noisy labels with noisy and curated with curated or just everything together? "
        },
        {
          "id": 552632,
          "postDate": "2019-06-14T08:36:14.397Z",
          "content": "<p><a href=\"/hwasiti\">@hwasiti</a>, The aim of multitasking learning is to get synergy between 2 tasks without reducing the performance of each task. Multitask learning learns features shared between two tasks and can be expected to achieve higher performance than learning independently. In this competition, the curated data and noisy data are labeled in a very different way, so treating them as the same one will collapse the performance. In our case, ResNet34 architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation.</p>",
          "rawMarkdown": "@hwasiti, The aim of multitasking learning is to get synergy between 2 tasks without reducing the performance of each task. Multitask learning learns features shared between two tasks and can be expected to achieve higher performance than learning independently. In this competition, the curated data and noisy data are labeled in a very different way, so treating them as the same one will collapse the performance. In our case, ResNet34 architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation.",
          "votes": 4
        },
        {
          "id": 552633,
          "postDate": "2019-06-14T08:41:25.657Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a>, MixUp is applied among curated data and among noisy data independently. This is because the main purpose of multitasking learning is to deal with curated labels and noisy labels independently.</p>",
          "rawMarkdown": "@iafoss, MixUp is applied among curated data and among noisy data independently. This is because the main purpose of multitasking learning is to deal with curated labels and noisy labels independently."
        },
        {
          "id": 553587,
          "postDate": "2019-06-16T01:51:42.213Z",
          "content": "<p>It's an exellent idea to use the noisy data. Do you do verification on the noisy data or just use all the noisy data? Weighted loss between the two tasks? How much improvement do you get when applying multi-tasking comparing to single task?</p>",
          "rawMarkdown": "It's an exellent idea to use the noisy data. Do you do verification on the noisy data or just use all the noisy data? Weighted loss between the two tasks? How much improvement do you get when applying multi-tasking comparing to single task?"
        },
        {
          "id": 553650,
          "postDate": "2019-06-16T04:16:35.750Z",
          "content": "<p><a href=\"/smallfade\">@smallfade</a> we used all the noisy data. The ratio of loss-weight is curated: noisy = 1: 1. We have not searched better ratio so much. At least we tried ratio 1:5 and it's not good. We need more experiments. Multitask learning improved score +0.021 on public LB comparing to single-task. </p>",
          "rawMarkdown": "@smallfade we used all the noisy data. The ratio of loss-weight is curated: noisy = 1: 1. We have not searched better ratio so much. At least we tried ratio 1:5 and it's not good. We need more experiments. Multitask learning improved score +0.021 on public LB comparing to single-task. ",
          "votes": 1
        },
        {
          "id": 553652,
          "postDate": "2019-06-16T04:28:36.030Z",
          "content": "<p>Thanks for your sharing and kind explanation.</p>",
          "rawMarkdown": "Thanks for your sharing and kind explanation."
        },
        {
          "id": 554323,
          "postDate": "2019-06-17T09:16:42.857Z",
          "content": "<p><a href=\"/osciiart\">@osciiart</a> How do you handle the prediction at test time for multi-task model. You have multiple heads, so each head has its own prediciton. Do you average them?</p>",
          "rawMarkdown": "@osciiart How do you handle the prediction at test time for multi-task model. You have multiple heads, so each head has its own prediciton. Do you average them?\n"
        },
        {
          "id": 554465,
          "postDate": "2019-06-17T12:49:34.540Z",
          "content": "<p><a href=\"/hwasiti\">@hwasiti</a> Only the curated head is used for prediction.</p>",
          "rawMarkdown": "@hwasiti Only the curated head is used for prediction.",
          "votes": 1
        }
      ]
    },
    {
      "id": 550687,
      "postDate": "2019-06-11T23:30:59.377Z",
      "content": "<p>I used the noisy dataset in a very similar way to Eric, warming up with the noisy set followed by finetuning/semi-supervised with curated + noisy but I also used label smoothing during the \"warmup\" phase. I also used a combination of mixup + specaugment augmentation, but not as sophisticated as SpecMix (nice work <a href=\"/ebouteillon\">@ebouteillon</a> on that!). </p>",
      "rawMarkdown": "I used the noisy dataset in a very similar way to Eric, warming up with the noisy set followed by finetuning/semi-supervised with curated + noisy but I also used label smoothing during the \"warmup\" phase. I also used a combination of mixup + specaugment augmentation, but not as sophisticated as SpecMix (nice work @ebouteillon on that!). ",
      "votes": 1
    },
    {
      "id": 550662,
      "postDate": "2019-06-11T22:22:09.940Z",
      "content": "<p>We also used them for the warm-up stage as <a href=\"/ebouteillon\">@ebouteillon</a> did, although we didn't used them in semi-supervised way, just for warm-up.\nThis pushed our score up around LB0.01, but the effects of this pre-training differ between models.\nWe also tried selecting \"best\" subset of the noisy data but didn't find them useful.</p>",
      "rawMarkdown": "We also used them for the warm-up stage as @ebouteillon did, although we didn't used them in semi-supervised way, just for warm-up.\nThis pushed our score up around LB0.01, but the effects of this pre-training differ between models.\nWe also tried selecting \"best\" subset of the noisy data but didn't find them useful.",
      "votes": 1
    },
    {
      "id": 550645,
      "postDate": "2019-06-11T21:42:31.557Z",
      "content": "<p>Hello <a href=\"/jihangz\">@jihangz</a> ,\nThis paper sounds interesting. What a pity that it is behind a paywall.  😞 \nIf someone could point to some blog or website describing it?</p>",
      "rawMarkdown": "Hello @jihangz ,\nThis paper sounds interesting. What a pity that it is behind a paywall.  😞 \nIf someone could point to some blog or website describing it?",
      "votes": 1,
      "replies": [
        {
          "id": 550650,
          "postDate": "2019-06-11T21:57:45.140Z",
          "content": "<p>I will describe it in my technical report and will talk about my approach in a different post. Just be patient :)</p>",
          "rawMarkdown": "I will describe it in my technical report and will talk about my approach in a different post. Just be patient :)",
          "votes": 1
        },
        {
          "id": 551700,
          "postDate": "2019-06-13T03:01:08.677Z",
          "content": "<p>I found the paper offline, uploading it here if anyone needs it. </p>",
          "rawMarkdown": "I found the paper offline, uploading it here if anyone needs it. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 550638,
      "postDate": "2019-06-11T21:15:31.517Z",
      "content": "<p>Thanks for sharing, how much improvement did you get with this technique?</p>\n\n<p>My approach for noisy data was very basic, I used just the \"best\" ~3500 noisy samples. To select those I predicted the labels for the noisy set based on a model trained with curated data only and selected the ones that the model got correct. </p>\n\n<p>I tried MixMatch and other SSL approaches but none gave me an improvement over using only curated data. But probably it's just that I didn't have enough time to properly test them. </p>",
      "rawMarkdown": "Thanks for sharing, how much improvement did you get with this technique?\n\nMy approach for noisy data was very basic, I used just the \"best\" ~3500 noisy samples. To select those I predicted the labels for the noisy set based on a model trained with curated data only and selected the ones that the model got correct. \n\nI tried MixMatch and other SSL approaches but none gave me an improvement over using only curated data. But probably it's just that I didn't have enough time to properly test them. ",
      "votes": 1,
      "replies": [
        {
          "id": 550642,
          "postDate": "2019-06-11T21:32:22.643Z",
          "content": "<p>Same backbone model (VGGish 9-layer CNN), curated only gave me LB 0.692 with 5-fold CV averaging, while using the approach above gave me LB 0.720 with 5-fold CV averaging, and LB 0.712 with single classifier trained on all data (no train-valid split)</p>",
          "rawMarkdown": "Same backbone model (VGGish 9-layer CNN), curated only gave me LB 0.692 with 5-fold CV averaging, while using the approach above gave me LB 0.720 with 5-fold CV averaging, and LB 0.712 with single classifier trained on all data (no train-valid split)",
          "votes": 1
        }
      ]
    },
    {
      "id": 550655,
      "postDate": "2019-06-11T22:09:16.513Z",
      "content": "<p><a href=\"/jihangz\">@jihangz</a>  Thanks for taking care of myself. Sorry for being impatient. 😄 </p>\n\n<p>To come back to your topic question, I did what I called a \"warm-up pipeline\": I use noisy dataset to warmup training, then fine tune with curated and then use noisy dataset again in a semi-supervised way.\nIf you are interested I put more details in this topic I just created (including github links): <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95382\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95382</a></p>",
      "rawMarkdown": "@jihangz  Thanks for taking care of myself. Sorry for being impatient. 😄 \n\nTo come back to your topic question, I did what I called a \"warm-up pipeline\": I use noisy dataset to warmup training, then fine tune with curated and then use noisy dataset again in a semi-supervised way.\nIf you are interested I put more details in this topic I just created (including github links): https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95382",
      "votes": 2,
      "replies": [
        {
          "id": 550674,
          "postDate": "2019-06-11T22:53:21.023Z",
          "content": "<p>Thanks for sharing. In fact, your pipeline is very similar to what the paper does in general: warmup with noisy, then semi-supervised using both curated and noisy. But I added some sauce on top of the paper</p>",
          "rawMarkdown": "Thanks for sharing. In fact, your pipeline is very similar to what the paper does in general: warmup with noisy, then semi-supervised using both curated and noisy. But I added some sauce on top of the paper"
        }
      ]
    },
    {
      "id": 552868,
      "postDate": "2019-06-14T17:56:36.890Z",
      "content": "<p>I trained on curated first, then fine tuned on curated+noisy, using either label smoothing or sample weight reduction on noisy labels. \nThat gave big improvement on CV( from ~0.82 to 0.9+, measured on curated data only) but didn't generalise that well on LB. Still gave LB improvement about 0.015.</p>",
      "rawMarkdown": "I trained on curated first, then fine tuned on curated+noisy, using either label smoothing or sample weight reduction on noisy labels. \nThat gave big improvement on CV( from ~0.82 to 0.9+, measured on curated data only) but didn't generalise that well on LB. Still gave LB improvement about 0.015."
    },
    {
      "id": 551583,
      "postDate": "2019-06-12T22:07:48.887Z",
      "content": "<p>Hi <a href=\"/jihangz\">@jihangz</a> , thanks for sharing!. Would it be possible to share the pdf ??</p>\n\n<p>Thanks James <a href=\"/jamesrequa\">@jamesrequa</a> , Eric <a href=\"/ebouteillon\">@ebouteillon</a>  , Miguel <a href=\"/mnpinto\">@mnpinto</a>  and <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> as well for all your sharing!</p>",
      "rawMarkdown": "Hi @jihangz , thanks for sharing!. Would it be possible to share the pdf ??\n\nThanks James @jamesrequa , Eric @ebouteillon  , Miguel @mnpinto  and @hidehisaarai1213 as well for all your sharing!",
      "replies": [
        {
          "id": 551731,
          "postDate": "2019-06-13T03:59:54.520Z",
          "content": "<p>I don't think it is proper (or legal) for me to share the pdf after they removed it from public server. You can use any college network to download the paper for free though.</p>",
          "rawMarkdown": "I don't think it is proper (or legal) for me to share the pdf after they removed it from public server. You can use any college network to download the paper for free though."
        }
      ]
    },
    {
      "id": 1731835,
      "postDate": "2022-03-22T18:24:35.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1731629,
      "author_name": "Darnell Obrien",
      "author_url": "",
      "post_date": "2022-03-22T14:50:04.847000",
      "content": "<p>Anyone trying out this now can check out this audio tagging <a href=\"https://paperswithcode.com/paper/audio-tagging-with-noisy-labels-and-minimal\" target=\"_blank\">paper</a>: </p>\n<p>Also see: most implemented <a href=\"https://paperswithcode.com/task/audio-tagging\" target=\"_blank\">papers</a> on audio tagging. For checking out the performance of different algorithms to see which puts the best results on the table, I’d suggest <a href=\"https://paperswithcode.com/task/audio-tagging\" target=\"_blank\">deepchecks</a> </p>\n<p>Wishing everyone success! </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 551607,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2019-06-12T23:06:47.483000",
      "content": "<p>We added our model additional FC layers for noisy data and did multitask learning.\nSingle task 5-fold CV LwLRAP: 0.829\nMultitask 5-fold CV LwLRAP: 0.849</p>",
      "votes": 4,
      "replies": [
        {
          "id": 552098,
          "author_name": "Jing Liu",
          "author_url": "",
          "post_date": "2019-06-13T14:07:16.700000",
          "content": "<p>How do you set the multi tasks? Do you treat the multi-label as multi-task or detection and classification tasks?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 552212,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-13T16:31:40.453000",
          "content": "<p>It's like this.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/552212/13515/multitask%20pipeline.png\" alt=\"Multitask pipeline\"></p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 552418,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-06-13T23:34:17.823000",
          "content": "<p>Nice!\nSo it would contain two sample per each training step input like a siamese NN and two set of labels to calculate the loss during training?\nBut why? What is the advantage of this method? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 552467,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2019-06-14T02:39:49.617000",
          "content": "<p><a href=\"/osciiart\">@osciiart</a>, quite interesting idea. I'm curious in which way did u add MixUp to this scheme? Did u mix only noisy labels with noisy and curated with curated or just everything together? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 552632,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-14T08:36:14.397000",
          "content": "<p><a href=\"/hwasiti\">@hwasiti</a>, The aim of multitasking learning is to get synergy between 2 tasks without reducing the performance of each task. Multitask learning learns features shared between two tasks and can be expected to achieve higher performance than learning independently. In this competition, the curated data and noisy data are labeled in a very different way, so treating them as the same one will collapse the performance. In our case, ResNet34 architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 552633,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-14T08:41:25.657000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a>, MixUp is applied among curated data and among noisy data independently. This is because the main purpose of multitasking learning is to deal with curated labels and noisy labels independently.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 553587,
          "author_name": "Jing Liu",
          "author_url": "",
          "post_date": "2019-06-16T01:51:42.213000",
          "content": "<p>It's an exellent idea to use the noisy data. Do you do verification on the noisy data or just use all the noisy data? Weighted loss between the two tasks? How much improvement do you get when applying multi-tasking comparing to single task?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 553650,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-16T04:16:35.750000",
          "content": "<p><a href=\"/smallfade\">@smallfade</a> we used all the noisy data. The ratio of loss-weight is curated: noisy = 1: 1. We have not searched better ratio so much. At least we tried ratio 1:5 and it's not good. We need more experiments. Multitask learning improved score +0.021 on public LB comparing to single-task. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 553652,
          "author_name": "Jing Liu",
          "author_url": "",
          "post_date": "2019-06-16T04:28:36.030000",
          "content": "<p>Thanks for your sharing and kind explanation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 554323,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-06-17T09:16:42.857000",
          "content": "<p><a href=\"/osciiart\">@osciiart</a> How do you handle the prediction at test time for multi-task model. You have multiple heads, so each head has its own prediciton. Do you average them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 554465,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-17T12:49:34.540000",
          "content": "<p><a href=\"/hwasiti\">@hwasiti</a> Only the curated head is used for prediction.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 550687,
      "author_name": "James Requa",
      "author_url": "",
      "post_date": "2019-06-11T23:30:59.377000",
      "content": "<p>I used the noisy dataset in a very similar way to Eric, warming up with the noisy set followed by finetuning/semi-supervised with curated + noisy but I also used label smoothing during the \"warmup\" phase. I also used a combination of mixup + specaugment augmentation, but not as sophisticated as SpecMix (nice work <a href=\"/ebouteillon\">@ebouteillon</a> on that!). </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 550662,
      "author_name": "Hidehisa Arai",
      "author_url": "",
      "post_date": "2019-06-11T22:22:09.940000",
      "content": "<p>We also used them for the warm-up stage as <a href=\"/ebouteillon\">@ebouteillon</a> did, although we didn't used them in semi-supervised way, just for warm-up.\nThis pushed our score up around LB0.01, but the effects of this pre-training differ between models.\nWe also tried selecting \"best\" subset of the noisy data but didn't find them useful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 550645,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2019-06-11T21:42:31.557000",
      "content": "<p>Hello <a href=\"/jihangz\">@jihangz</a> ,\nThis paper sounds interesting. What a pity that it is behind a paywall.  😞 \nIf someone could point to some blog or website describing it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 550650,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-06-11T21:57:45.140000",
          "content": "<p>I will describe it in my technical report and will talk about my approach in a different post. Just be patient :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 551700,
          "author_name": "yukiya",
          "author_url": "",
          "post_date": "2019-06-13T03:01:08.677000",
          "content": "<p>I found the paper offline, uploading it here if anyone needs it. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 550638,
      "author_name": "Miguel Pinto",
      "author_url": "",
      "post_date": "2019-06-11T21:15:31.517000",
      "content": "<p>Thanks for sharing, how much improvement did you get with this technique?</p>\n\n<p>My approach for noisy data was very basic, I used just the \"best\" ~3500 noisy samples. To select those I predicted the labels for the noisy set based on a model trained with curated data only and selected the ones that the model got correct. </p>\n\n<p>I tried MixMatch and other SSL approaches but none gave me an improvement over using only curated data. But probably it's just that I didn't have enough time to properly test them. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 550642,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-06-11T21:32:22.643000",
          "content": "<p>Same backbone model (VGGish 9-layer CNN), curated only gave me LB 0.692 with 5-fold CV averaging, while using the approach above gave me LB 0.720 with 5-fold CV averaging, and LB 0.712 with single classifier trained on all data (no train-valid split)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 550655,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2019-06-11T22:09:16.513000",
      "content": "<p><a href=\"/jihangz\">@jihangz</a>  Thanks for taking care of myself. Sorry for being impatient. 😄 </p>\n\n<p>To come back to your topic question, I did what I called a \"warm-up pipeline\": I use noisy dataset to warmup training, then fine tune with curated and then use noisy dataset again in a semi-supervised way.\nIf you are interested I put more details in this topic I just created (including github links): <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95382\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95382</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 550674,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-06-11T22:53:21.023000",
          "content": "<p>Thanks for sharing. In fact, your pipeline is very similar to what the paper does in general: warmup with noisy, then semi-supervised using both curated and noisy. But I added some sauce on top of the paper</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 552868,
      "author_name": "Victor Zaguskin",
      "author_url": "",
      "post_date": "2019-06-14T17:56:36.890000",
      "content": "<p>I trained on curated first, then fine tuned on curated+noisy, using either label smoothing or sample weight reduction on noisy labels. \nThat gave big improvement on CV( from ~0.82 to 0.9+, measured on curated data only) but didn't generalise that well on LB. Still gave LB improvement about 0.015.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 551583,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-06-12T22:07:48.887000",
      "content": "<p>Hi <a href=\"/jihangz\">@jihangz</a> , thanks for sharing!. Would it be possible to share the pdf ??</p>\n\n<p>Thanks James <a href=\"/jamesrequa\">@jamesrequa</a> , Eric <a href=\"/ebouteillon\">@ebouteillon</a>  , Miguel <a href=\"/mnpinto\">@mnpinto</a>  and <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> as well for all your sharing!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 551731,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-06-13T03:59:54.520000",
          "content": "<p>I don't think it is proper (or legal) for me to share the pdf after they removed it from public server. You can use any college network to download the paper for free though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1731835,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-22T18:24:35.380000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1731629": "Anyone trying out this now can check out this audio tagging [paper](https://paperswithcode.com/paper/audio-tagging-with-noisy-labels-and-minimal): \n\nAlso see: most implemented [papers](https://paperswithcode.com/task/audio-tagging) on audio tagging. For checking out the performance of different algorithms to see which puts the best results on the table, I’d suggest [deepchecks](https://paperswithcode.com/task/audio-tagging) \n\nWishing everyone success! \n",
    "550599": "I am very interested in how everyone deals with noisy data. Using noisy labels? Pseudo-label? MixMatch? UDA? Other SSL approaches?\n\nI was using exactly the same model and training step here except the backbone and the non-linear activation function (PReLU instead of Sigmoid): https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26\n\nThere was a free pdf version of this paper but they removed it from public server.",
    "551607": "We added our model additional FC layers for noisy data and did multitask learning.\nSingle task 5-fold CV LwLRAP: 0.829\nMultitask 5-fold CV LwLRAP: 0.849",
    "550687": "I used the noisy dataset in a very similar way to Eric, warming up with the noisy set followed by finetuning/semi-supervised with curated + noisy but I also used label smoothing during the \"warmup\" phase. I also used a combination of mixup + specaugment augmentation, but not as sophisticated as SpecMix (nice work @ebouteillon on that!). ",
    "550662": "We also used them for the warm-up stage as @ebouteillon did, although we didn't used them in semi-supervised way, just for warm-up.\nThis pushed our score up around LB0.01, but the effects of this pre-training differ between models.\nWe also tried selecting \"best\" subset of the noisy data but didn't find them useful.",
    "550645": "Hello @jihangz ,\nThis paper sounds interesting. What a pity that it is behind a paywall.  😞 \nIf someone could point to some blog or website describing it?",
    "550638": "Thanks for sharing, how much improvement did you get with this technique?\n\nMy approach for noisy data was very basic, I used just the \"best\" ~3500 noisy samples. To select those I predicted the labels for the noisy set based on a model trained with curated data only and selected the ones that the model got correct. \n\nI tried MixMatch and other SSL approaches but none gave me an improvement over using only curated data. But probably it's just that I didn't have enough time to properly test them. ",
    "550655": "@jihangz  Thanks for taking care of myself. Sorry for being impatient. 😄 \n\nTo come back to your topic question, I did what I called a \"warm-up pipeline\": I use noisy dataset to warmup training, then fine tune with curated and then use noisy dataset again in a semi-supervised way.\nIf you are interested I put more details in this topic I just created (including github links): https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95382",
    "552868": "I trained on curated first, then fine tuned on curated+noisy, using either label smoothing or sample weight reduction on noisy labels. \nThat gave big improvement on CV( from ~0.82 to 0.9+, measured on curated data only) but didn't generalise that well on LB. Still gave LB improvement about 0.015.",
    "551583": "Hi @jihangz , thanks for sharing!. Would it be possible to share the pdf ??\n\nThanks James @jamesrequa , Eric @ebouteillon  , Miguel @mnpinto  and @hidehisaarai1213 as well for all your sharing!",
    "1731835": ""
  }
}