{
  "id": 95344,
  "title": "13th place solution [public LB] - brief summary",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/kuaiyu-13th-place-solution-public-lb-brief-summary",
  "author_name": "",
  "post_date": "2019-07-05T09:50:18.797Z",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The second year to take part in freesound audio tagging, thanks for the organization for the interesting and close to reality audio competition. Thanks for <a href=\"/daisukelab\">@daisukelab</a> <a href=\"/mhiro2\">@mhiro2</a> <a href=\"/ceshine\">@ceshine</a> <a href=\"/jihangz\">@jihangz</a> , and many kagglers for the starter kernels and discussions.</p>\n\n<p>Earlier single model kernel without cv is as below:\n<a href=\"https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\">https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673</a>\nDiscussion about single model LB:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-547036\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-547036</a>\n  <a href=\"https://github.com/sailor88128/dcase2019-task2\">https://github.com/sailor88128/dcase2019-task2</a></p>\n\n<p>Final submission:\n- On curated dataset, 5 fold cv, inception V3, single model LB is 0.691 and 5 model average is 0.713;\n- On curated dataset and noisy dataset, with specific loss function, 5 fold cv, 8 layer cnn, 5 model average 0.678;\n- Geometric average of these two models, 0.73+.</p>\n\n<p>Techniques we ever used:\n1. PCEN spectragram, didnot try many parameters, not sure if it work for this competition; \n2. mixup, about 0.01 increase in LB;\n3. SpecAugment shows no increase, while RandomResizedCrop cannot be explained for audio with a better lb increase;\n4. CV and tta,  about 0.05 increase for the primary single model;\n5. Mixmatch, got a much lower LB with the same model, need to check the code then; \n6. Schedule of lr, cycliclr and cosineannealing do not show much difference.</p>\n\n<p>Others:\n1. While lots of kagglers said shallow cnn works well, we didnot get effective shallow cnn. We ever tried alexnet or vgg, all with low local lwlrap :(\n2. Different cnns were used at first, but we didnot save them well for the final model ensemble;\n3. Dataset with noisy label are not well used. We just used a weight loss function to train the noisy labeled data together with curated data set. The mixmatch we realize for the task maybe wrong in detail and should be checked.</p>",
  "messages": [
    {
      "id": "550408",
      "postDate": "06/11/2019 15:46:03",
      "content": "<p>The second year to take part in freesound audio tagging, thanks for the organization for the interesting and close to reality audio competition. Thanks for <a href=\"/daisukelab\">@daisukelab</a> <a href=\"/mhiro2\">@mhiro2</a> <a href=\"/ceshine\">@ceshine</a> <a href=\"/jihangz\">@jihangz</a> , and many kagglers for the starter kernels and discussions.</p>\n\n<p>Earlier single model kernel without cv is as below:\n<a href=\"https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\">https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673</a>\nDiscussion about single model LB:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-547036\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-547036</a>\n  <a href=\"https://github.com/sailor88128/dcase2019-task2\">https://github.com/sailor88128/dcase2019-task2</a></p>\n\n<p>Final submission:\n- On curated dataset, 5 fold cv, inception V3, single model LB is 0.691 and 5 model average is 0.713;\n- On curated dataset and noisy dataset, with specific loss function, 5 fold cv, 8 layer cnn, 5 model average 0.678;\n- Geometric average of these two models, 0.73+.</p>\n\n<p>Techniques we ever used:\n1. PCEN spectragram, didnot try many parameters, not sure if it work for this competition; \n2. mixup, about 0.01 increase in LB;\n3. SpecAugment shows no increase, while RandomResizedCrop cannot be explained for audio with a better lb increase;\n4. CV and tta,  about 0.05 increase for the primary single model;\n5. Mixmatch, got a much lower LB with the same model, need to check the code then; \n6. Schedule of lr, cycliclr and cosineannealing do not show much difference.</p>\n\n<p>Others:\n1. While lots of kagglers said shallow cnn works well, we didnot get effective shallow cnn. We ever tried alexnet or vgg, all with low local lwlrap :(\n2. Different cnns were used at first, but we didnot save them well for the final model ensemble;\n3. Dataset with noisy label are not well used. We just used a weight loss function to train the noisy labeled data together with curated data set. The mixmatch we realize for the task maybe wrong in detail and should be checked.</p>",
      "rawMarkdown": "The second year to take part in freesound audio tagging, thanks for the organization for the interesting and close to reality audio competition. Thanks for @daisukelab @mhiro2 @ceshine @jihangz , and many kagglers for the starter kernels and discussions.\n\nEarlier single model kernel without cv is as below:\nhttps://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\nDiscussion about single model LB:\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-547036\n  https://github.com/sailor88128/dcase2019-task2\n\n\nFinal submission:\n- On curated dataset, 5 fold cv, inception V3, single model LB is 0.691 and 5 model average is 0.713;\n- On curated dataset and noisy dataset, with specific loss function, 5 fold cv, 8 layer cnn, 5 model average 0.678;\n- Geometric average of these two models, 0.73+.\n\n\nTechniques we ever used:\n1. PCEN spectragram, didnot try many parameters, not sure if it work for this competition; \n2. mixup, about 0.01 increase in LB;\n3. SpecAugment shows no increase, while RandomResizedCrop cannot be explained for audio with a better lb increase;\n4. CV and tta,  about 0.05 increase for the primary single model;\n5. Mixmatch, got a much lower LB with the same model, need to check the code then; \n6. Schedule of lr, cycliclr and cosineannealing do not show much difference.\n\n\nOthers:\n1. While lots of kagglers said shallow cnn works well, we didnot get effective shallow cnn. We ever tried alexnet or vgg, all with low local lwlrap :(\n2. Different cnns were used at first, but we didnot save them well for the final model ensemble;\n3. Dataset with noisy label are not well used. We just used a weight loss function to train the noisy labeled data together with curated data set. The mixmatch we realize for the task maybe wrong in detail and should be checked.",
      "votes": null
    },
    {
      "id": "550435",
      "postDate": "06/11/2019 16:13:53",
      "content": "<p>Thank you for sharing!\nYour idea was very helpful for us. \nCertainly,  RandomResizedCrop is effective in Our Inception v3 model. \nAnd, LB jumps up when 5fold averaging.\nYou said your CV is too low.I found that your CV is low because of RandomResizedCrop.\nThis problem is solved by using tta for validation.</p>\n\n<p>When RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected.</p>\n\n<p>We will post our solution later.\nThank you.</p>",
      "rawMarkdown": "Thank you for sharing!\nYour idea was very helpful for us. \nCertainly,  RandomResizedCrop is effective in Our Inception v3 model. \nAnd, LB jumps up when 5fold averaging.\nYou said your CV is too low.I found that your CV is low because of RandomResizedCrop.\nThis problem is solved by using tta for validation.\n\nWhen RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected.\n\nWe will post our solution later.\nThank you.",
      "votes": null
    },
    {
      "id": "550745",
      "postDate": "06/12/2019 01:49:13",
      "content": "<p>It's good the idea helpful for you. tta for validation was used and local lwlrap increased, but LB almost remained the same.\nLooking forward to your solution.</p>",
      "rawMarkdown": "It's good the idea helpful for you. tta for validation was used and local lwlrap increased, but LB almost remained the same.\nLooking forward to your solution.",
      "votes": null
    },
    {
      "id": "551182",
      "postDate": "06/12/2019 12:33:06",
      "content": "<p>Thanks for sharing <a href=\"/sailorwei\">@sailorwei</a> </p>",
      "rawMarkdown": "Thanks for sharing @sailorwei",
      "votes": null
    },
    {
      "id": "554811",
      "postDate": "06/18/2019 03:36:17",
      "content": "<p>Thanks for sharing, one question about \"Geometric average of these two models, 0.73+.\",\nhave you tried arithmetic mean?\nFor me geometric averaging showed worse result, then I'm curious about the reason (or hints).</p>",
      "rawMarkdown": "Thanks for sharing, one question about \"Geometric average of these two models, 0.73+.\",\nhave you tried arithmetic mean?\nFor me geometric averaging showed worse result, then I'm curious about the reason (or hints).",
      "votes": null
    },
    {
      "id": "554859",
      "postDate": "06/18/2019 05:42:35",
      "content": "<p>Yes, we tried both geometric and arithmetic mean, geometric mean shows better lb for us, about 0.002 higher. I am not sure why either.</p>",
      "rawMarkdown": "Yes, we tried both geometric and arithmetic mean, geometric mean shows better lb for us, about 0.002 higher. I am not sure why either.",
      "votes": null
    },
    {
      "id": "555103",
      "postDate": "06/18/2019 12:38:50",
      "content": "<p>Thanks I see, I might be better revisiting it. I have been using geometric mean, but arithmetic was better in this competition. Thanks again!</p>",
      "rawMarkdown": "Thanks I see, I might be better revisiting it. I have been using geometric mean, but arithmetic was better in this competition. Thanks again!",
      "votes": null
    },
    {
      "id": "555831",
      "postDate": "06/19/2019 13:53:20",
      "content": "<p>At very late stage, I tried time shift in TTA, found it improved a bit (but met memory issue and didn't use it at last). I think RandomResizedCrop may have the similar benefit as long as you didn't resize too much, if you it's not cropped by center, it's like shifting the spectrum temporally. If you can share the parameter you used for RandomResizedCrop or samples of the the augmented spectrum,  we may find an explanation.</p>",
      "rawMarkdown": "At very late stage, I tried time shift in TTA, found it improved a bit (but met memory issue and didn't use it at last). I think RandomResizedCrop may have the similar benefit as long as you didn't resize too much, if you it's not cropped by center, it's like shifting the spectrum temporally. If you can share the parameter you used for RandomResizedCrop or samples of the the augmented spectrum,  we may find an explanation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 550435,
      "author_name": "d1348k",
      "author_url": "",
      "post_date": "06/11/2019 16:13:53",
      "content": "<p>Thank you for sharing!\nYour idea was very helpful for us. \nCertainly,  RandomResizedCrop is effective in Our Inception v3 model. \nAnd, LB jumps up when 5fold averaging.\nYou said your CV is too low.I found that your CV is low because of RandomResizedCrop.\nThis problem is solved by using tta for validation.</p>\n\n<p>When RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected.</p>\n\n<p>We will post our solution later.\nThank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 550745,
          "author_name": "sailorwei",
          "author_url": "",
          "post_date": "06/12/2019 01:49:13",
          "content": "<p>It's good the idea helpful for you. tta for validation was used and local lwlrap increased, but LB almost remained the same.\nLooking forward to your solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551182,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/12/2019 12:33:06",
      "content": "<p>Thanks for sharing <a href=\"/sailorwei\">@sailorwei</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 554811,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/18/2019 03:36:17",
      "content": "<p>Thanks for sharing, one question about \"Geometric average of these two models, 0.73+.\",\nhave you tried arithmetic mean?\nFor me geometric averaging showed worse result, then I'm curious about the reason (or hints).</p>",
      "votes": null,
      "replies": [
        {
          "id": 554859,
          "author_name": "sailorwei",
          "author_url": "",
          "post_date": "06/18/2019 05:42:35",
          "content": "<p>Yes, we tried both geometric and arithmetic mean, geometric mean shows better lb for us, about 0.002 higher. I am not sure why either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 555103,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/18/2019 12:38:50",
          "content": "<p>Thanks I see, I might be better revisiting it. I have been using geometric mean, but arithmetic was better in this competition. Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 555831,
      "author_name": "smallfade",
      "author_url": "",
      "post_date": "06/19/2019 13:53:20",
      "content": "<p>At very late stage, I tried time shift in TTA, found it improved a bit (but met memory issue and didn't use it at last). I think RandomResizedCrop may have the similar benefit as long as you didn't resize too much, if you it's not cropped by center, it's like shifting the spectrum temporally. If you can share the parameter you used for RandomResizedCrop or samples of the the augmented spectrum,  we may find an explanation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "550408": "The second year to take part in freesound audio tagging, thanks for the organization for the interesting and close to reality audio competition. Thanks for @daisukelab @mhiro2 @ceshine @jihangz , and many kagglers for the starter kernels and discussions.\n\nEarlier single model kernel without cv is as below:\nhttps://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\nDiscussion about single model LB:\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-547036\n  https://github.com/sailor88128/dcase2019-task2\n\n\nFinal submission:\n- On curated dataset, 5 fold cv, inception V3, single model LB is 0.691 and 5 model average is 0.713;\n- On curated dataset and noisy dataset, with specific loss function, 5 fold cv, 8 layer cnn, 5 model average 0.678;\n- Geometric average of these two models, 0.73+.\n\n\nTechniques we ever used:\n1. PCEN spectragram, didnot try many parameters, not sure if it work for this competition; \n2. mixup, about 0.01 increase in LB;\n3. SpecAugment shows no increase, while RandomResizedCrop cannot be explained for audio with a better lb increase;\n4. CV and tta,  about 0.05 increase for the primary single model;\n5. Mixmatch, got a much lower LB with the same model, need to check the code then; \n6. Schedule of lr, cycliclr and cosineannealing do not show much difference.\n\n\nOthers:\n1. While lots of kagglers said shallow cnn works well, we didnot get effective shallow cnn. We ever tried alexnet or vgg, all with low local lwlrap :(\n2. Different cnns were used at first, but we didnot save them well for the final model ensemble;\n3. Dataset with noisy label are not well used. We just used a weight loss function to train the noisy labeled data together with curated data set. The mixmatch we realize for the task maybe wrong in detail and should be checked.",
    "550435": "Thank you for sharing!\nYour idea was very helpful for us. \nCertainly,  RandomResizedCrop is effective in Our Inception v3 model. \nAnd, LB jumps up when 5fold averaging.\nYou said your CV is too low.I found that your CV is low because of RandomResizedCrop.\nThis problem is solved by using tta for validation.\n\nWhen RandomResizedCrop is used, val score fluctuate,so if val tta is not used, an appropriate epoch can not be selected.\n\nWe will post our solution later.\nThank you.",
    "550745": "It's good the idea helpful for you. tta for validation was used and local lwlrap increased, but LB almost remained the same.\nLooking forward to your solution.",
    "551182": "Thanks for sharing @sailorwei",
    "554811": "Thanks for sharing, one question about \"Geometric average of these two models, 0.73+.\",\nhave you tried arithmetic mean?\nFor me geometric averaging showed worse result, then I'm curious about the reason (or hints).",
    "554859": "Yes, we tried both geometric and arithmetic mean, geometric mean shows better lb for us, about 0.002 higher. I am not sure why either.",
    "555103": "Thanks I see, I might be better revisiting it. I have been using geometric mean, but arithmetic was better in this competition. Thanks again!",
    "555831": "At very late stage, I tried time shift in TTA, found it improved a bit (but met memory issue and didn't use it at last). I think RandomResizedCrop may have the similar benefit as long as you didn't resize too much, if you it's not cropped by center, it's like shifting the spectrum temporally. If you can share the parameter you used for RandomResizedCrop or samples of the the augmented spectrum,  we may find an explanation."
  },
  "source": "meta"
}