{
  "id": 57051,
  "title": "Beyond 0.9, sharing my everything so far (updated 24 July)",
  "url": "/competitions/freesound-audio-tagging/discussion/57051",
  "author_name": "daisukelab",
  "post_date": "2018-05-18T14:27:30.685000",
  "votes": 39,
  "comment_count": 51,
  "views": 0,
  "content": "<h1>Update 1 August</h1>\n\n<p>Submitted a kernel <a href=\"https://www.kaggle.com/daisukelab/freesound-dataset-kaggle-2018-solution\">here</a>.\nThis was supposed to be my best model, but private score was not unfortunately. My best solution has two more models (4 models in total), but basically the same.</p>\n\n<h1>Update 24 July</h1>\n\n<p>Let me update this post because my final solution has been changed a lot. I simplified after experimenting many different approaches and analyzing their results (over 150), resulting to have following notions:</p>\n\n<ul>\n<li>Test set is composed only with <em>manually verified</em> samples, unlike train set has weak labels. This could show that it is important to select which training examples to feed to the model.</li>\n<li>Public LB score is based on 19% of manually verified test samples, then the 19% is plausible to be slightly or more different with regard to data quantity and quality distribution. We might be better to trust our own validation performance rather than current LB score.</li>\n<li>Audio preprocessing basically should leave information as much as possible. No need to down sampling to lower fs like 24kHz, and feature resolution (n_mels) can be higher like 128 rather than traditional 40.</li>\n<li>Model doesn’t matter as long as it has enough but not too much capacity.</li>\n</ul>\n\n<p>Now I’m using <a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">Simple CNN approach</a> only with SEResNet-50 based model.</p>\n\n<p>BTW I will summarize my solution in a kernel and post right after competition closed.</p>\n\n<p>(And might not have enough time for DCASE paper submission, I'm sorry to organizers...)</p>\n\n<h1>Original post, while having LB score 0.957 or 0.961 (some could be outdated)</h1>\n\n<p>Reaching to score just under 0.9 would be usual kaggle work; preprocess dataset, build models, train them and ensemble predictions.\nThen bringing up models for better performance takes more efforts.\nAs far as I’m working on, building right model was important, but deep diving into audio domain and dataset matters more in this competition so far.</p>\n\n<p>Before I join this competition, I did my best in <a href=\"https://www.kaggle.com/c/acoustic-scene-2018\">TUT Acoustic Scene Classification</a> competition Feb this year, and could earn good lessons.</p>\n\n<p>Then firstly I will share summary of lessons learned in the TUT competition. I’m using almost the same CNN models for this freesound audio tagging competition also. And the augmentations (mixup &amp; cutout/random erasing) which was important also work fine. They are still strong.</p>\n\n<h2>Kernels shared in <a href=\"https://www.kaggle.com/c/acoustic-scene-2018\">TUT Acoustic Scene Classification</a> competition</h2>\n\n<p>These are what I’ve done for getting score 0.9+.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">Simple CNN approach</a> - this is basic kernel.</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/sound-event-based-approach\">Sound event based approach</a> - this is an approach to emphasize temporal change (calling it as sound event here).</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/time-wise-mean-approach\">Time-wise mean approach</a> - this is inverse approach that completely ignore temporal change.</li>\n<li>Regarding augmentation, <a href=\"https://www.kaggle.com/daisukelab/mixup-cutout-or-random-erasing-to-augment\">mixup &amp; cutout/random erasing</a> - Strong augmentation you cannot miss.</li>\n</ul>\n\n<h2>New findings in this competition</h2>\n\n<p>Followings are new findings while working on in this competition, kind of going back to the basics.\nThese have pushed my scores toward 0.95.</p>\n\n<ul>\n<li>When you preprocess, don't stick to 16kHz. The dataset have sound wave with 44.1kHz frequency, and many sounds are not human voice. Then converting to the 16kHz sampling rate which traditional speech recognition system uses is apparently not-adequate. Use higher frequency. For me, it's 24kHz or some more. It depends on model or other preprocessing methods.</li>\n<li>Different audio length might give us different perspective, though I'm not confident yet, but so far so good. Training/ testing audio sample has different length. Models that accepts fixed audio length like my CNNs require preprocessing to trim, repeat or fill blank to samples to make their length unified. A model which handles 2s input audio seems to be different from a model that takes 5s input audio. 2s model focuses on short period. On the other hands, 5s model looks over from higher view. I haven’t try but using CRNN instead of CNN could be the one-size-fits-all.</li>\n</ul>\n\n<h2>Other things for improvement</h2>\n\n<p>This is also basic thing but I needed for improvement.</p>\n\n<ul>\n<li>Fix class imbalance. Some classes have less samples that causes lower performance. I simply oversample fewer number of samples up to maximum number of samples until all the class have the same number of samples.</li>\n<li>Use entire audio wave. Some of implementation seems to use part of audio, but this also causes lower performance. I observed that some long audio has only half or shorter part that matches the class, and other part is nothing to do with the labeled class. I just simply use the entire part as training sample by splitting audio into multiple fixed unit length and put them all back into training set. Randomly choosing part of audio while training would be the same thing, but preprocessing first is faster to train in my environment.</li>\n</ul>\n\n<p>Lastly I should share that I'm not using pseudo labeling so far. So current score is purely done by preprocessing &amp; model engineering.</p>",
  "messages": [
    {
      "id": 330305,
      "postDate": "2018-05-18T14:27:30.687Z",
      "content": "<h1>Update 1 August</h1>\n\n<p>Submitted a kernel <a href=\"https://www.kaggle.com/daisukelab/freesound-dataset-kaggle-2018-solution\">here</a>.\nThis was supposed to be my best model, but private score was not unfortunately. My best solution has two more models (4 models in total), but basically the same.</p>\n\n<h1>Update 24 July</h1>\n\n<p>Let me update this post because my final solution has been changed a lot. I simplified after experimenting many different approaches and analyzing their results (over 150), resulting to have following notions:</p>\n\n<ul>\n<li>Test set is composed only with <em>manually verified</em> samples, unlike train set has weak labels. This could show that it is important to select which training examples to feed to the model.</li>\n<li>Public LB score is based on 19% of manually verified test samples, then the 19% is plausible to be slightly or more different with regard to data quantity and quality distribution. We might be better to trust our own validation performance rather than current LB score.</li>\n<li>Audio preprocessing basically should leave information as much as possible. No need to down sampling to lower fs like 24kHz, and feature resolution (n_mels) can be higher like 128 rather than traditional 40.</li>\n<li>Model doesn’t matter as long as it has enough but not too much capacity.</li>\n</ul>\n\n<p>Now I’m using <a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">Simple CNN approach</a> only with SEResNet-50 based model.</p>\n\n<p>BTW I will summarize my solution in a kernel and post right after competition closed.</p>\n\n<p>(And might not have enough time for DCASE paper submission, I'm sorry to organizers...)</p>\n\n<h1>Original post, while having LB score 0.957 or 0.961 (some could be outdated)</h1>\n\n<p>Reaching to score just under 0.9 would be usual kaggle work; preprocess dataset, build models, train them and ensemble predictions.\nThen bringing up models for better performance takes more efforts.\nAs far as I’m working on, building right model was important, but deep diving into audio domain and dataset matters more in this competition so far.</p>\n\n<p>Before I join this competition, I did my best in <a href=\"https://www.kaggle.com/c/acoustic-scene-2018\">TUT Acoustic Scene Classification</a> competition Feb this year, and could earn good lessons.</p>\n\n<p>Then firstly I will share summary of lessons learned in the TUT competition. I’m using almost the same CNN models for this freesound audio tagging competition also. And the augmentations (mixup &amp; cutout/random erasing) which was important also work fine. They are still strong.</p>\n\n<h2>Kernels shared in <a href=\"https://www.kaggle.com/c/acoustic-scene-2018\">TUT Acoustic Scene Classification</a> competition</h2>\n\n<p>These are what I’ve done for getting score 0.9+.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">Simple CNN approach</a> - this is basic kernel.</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/sound-event-based-approach\">Sound event based approach</a> - this is an approach to emphasize temporal change (calling it as sound event here).</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/time-wise-mean-approach\">Time-wise mean approach</a> - this is inverse approach that completely ignore temporal change.</li>\n<li>Regarding augmentation, <a href=\"https://www.kaggle.com/daisukelab/mixup-cutout-or-random-erasing-to-augment\">mixup &amp; cutout/random erasing</a> - Strong augmentation you cannot miss.</li>\n</ul>\n\n<h2>New findings in this competition</h2>\n\n<p>Followings are new findings while working on in this competition, kind of going back to the basics.\nThese have pushed my scores toward 0.95.</p>\n\n<ul>\n<li>When you preprocess, don't stick to 16kHz. The dataset have sound wave with 44.1kHz frequency, and many sounds are not human voice. Then converting to the 16kHz sampling rate which traditional speech recognition system uses is apparently not-adequate. Use higher frequency. For me, it's 24kHz or some more. It depends on model or other preprocessing methods.</li>\n<li>Different audio length might give us different perspective, though I'm not confident yet, but so far so good. Training/ testing audio sample has different length. Models that accepts fixed audio length like my CNNs require preprocessing to trim, repeat or fill blank to samples to make their length unified. A model which handles 2s input audio seems to be different from a model that takes 5s input audio. 2s model focuses on short period. On the other hands, 5s model looks over from higher view. I haven’t try but using CRNN instead of CNN could be the one-size-fits-all.</li>\n</ul>\n\n<h2>Other things for improvement</h2>\n\n<p>This is also basic thing but I needed for improvement.</p>\n\n<ul>\n<li>Fix class imbalance. Some classes have less samples that causes lower performance. I simply oversample fewer number of samples up to maximum number of samples until all the class have the same number of samples.</li>\n<li>Use entire audio wave. Some of implementation seems to use part of audio, but this also causes lower performance. I observed that some long audio has only half or shorter part that matches the class, and other part is nothing to do with the labeled class. I just simply use the entire part as training sample by splitting audio into multiple fixed unit length and put them all back into training set. Randomly choosing part of audio while training would be the same thing, but preprocessing first is faster to train in my environment.</li>\n</ul>\n\n<p>Lastly I should share that I'm not using pseudo labeling so far. So current score is purely done by preprocessing &amp; model engineering.</p>",
      "rawMarkdown": "# Update 1 August\n\nSubmitted a kernel [here](https://www.kaggle.com/daisukelab/freesound-dataset-kaggle-2018-solution).\nThis was supposed to be my best model, but private score was not unfortunately. My best solution has two more models (4 models in total), but basically the same.\n\n# Update 24 July\n\nLet me update this post because my final solution has been changed a lot. I simplified after experimenting many different approaches and analyzing their results (over 150), resulting to have following notions:\n\n- Test set is composed only with _manually verified_ samples, unlike train set has weak labels. This could show that it is important to select which training examples to feed to the model.\n- Public LB score is based on 19% of manually verified test samples, then the 19% is plausible to be slightly or more different with regard to data quantity and quality distribution. We might be better to trust our own validation performance rather than current LB score.\n- Audio preprocessing basically should leave information as much as possible. No need to down sampling to lower fs like 24kHz, and feature resolution (n_mels) can be higher like 128 rather than traditional 40.\n- Model doesn’t matter as long as it has enough but not too much capacity.\n\nNow I’m using [Simple CNN approach](https://www.kaggle.com/daisukelab/simple-cnn-approach) only with SEResNet-50 based model.\n\nBTW I will summarize my solution in a kernel and post right after competition closed.\n\n(And might not have enough time for DCASE paper submission, I'm sorry to organizers...)\n\n# Original post, while having LB score 0.957 or 0.961 (some could be outdated)\n\nReaching to score just under 0.9 would be usual kaggle work; preprocess dataset, build models, train them and ensemble predictions.\nThen bringing up models for better performance takes more efforts.\nAs far as I’m working on, building right model was important, but deep diving into audio domain and dataset matters more in this competition so far.\n\nBefore I join this competition, I did my best in [TUT Acoustic Scene Classification](https://www.kaggle.com/c/acoustic-scene-2018) competition Feb this year, and could earn good lessons.\n\nThen firstly I will share summary of lessons learned in the TUT competition. I’m using almost the same CNN models for this freesound audio tagging competition also. And the augmentations (mixup &amp; cutout/random erasing) which was important also work fine. They are still strong.\n\n## Kernels shared in [TUT Acoustic Scene Classification](https://www.kaggle.com/c/acoustic-scene-2018) competition\n\nThese are what I’ve done for getting score 0.9+.\n\n- [Simple CNN approach](https://www.kaggle.com/daisukelab/simple-cnn-approach) - this is basic kernel.\n- [Sound event based approach](https://www.kaggle.com/daisukelab/sound-event-based-approach) - this is an approach to emphasize temporal change (calling it as sound event here).\n- [Time-wise mean approach](https://www.kaggle.com/daisukelab/time-wise-mean-approach) - this is inverse approach that completely ignore temporal change.\n- Regarding augmentation, [mixup &amp; cutout/random erasing](https://www.kaggle.com/daisukelab/mixup-cutout-or-random-erasing-to-augment) - Strong augmentation you cannot miss.\n\n## New findings in this competition\n\nFollowings are new findings while working on in this competition, kind of going back to the basics.\nThese have pushed my scores toward 0.95.\n\n- When you preprocess, don't stick to 16kHz. The dataset have sound wave with 44.1kHz frequency, and many sounds are not human voice. Then converting to the 16kHz sampling rate which traditional speech recognition system uses is apparently not-adequate. Use higher frequency. For me, it's 24kHz or some more. It depends on model or other preprocessing methods.\n- Different audio length might give us different perspective, though I'm not confident yet, but so far so good. Training/ testing audio sample has different length. Models that accepts fixed audio length like my CNNs require preprocessing to trim, repeat or fill blank to samples to make their length unified. A model which handles 2s input audio seems to be different from a model that takes 5s input audio. 2s model focuses on short period. On the other hands, 5s model looks over from higher view. I haven’t try but using CRNN instead of CNN could be the one-size-fits-all.\n\n## Other things for improvement\n\nThis is also basic thing but I needed for improvement.\n\n- Fix class imbalance. Some classes have less samples that causes lower performance. I simply oversample fewer number of samples up to maximum number of samples until all the class have the same number of samples.\n- Use entire audio wave. Some of implementation seems to use part of audio, but this also causes lower performance. I observed that some long audio has only half or shorter part that matches the class, and other part is nothing to do with the labeled class. I just simply use the entire part as training sample by splitting audio into multiple fixed unit length and put them all back into training set. Randomly choosing part of audio while training would be the same thing, but preprocessing first is faster to train in my environment.\n\nLastly I should share that I'm not using pseudo labeling so far. So current score is purely done by preprocessing &amp; model engineering.",
      "votes": 39
    },
    {
      "id": 336992,
      "postDate": "2018-06-01T17:44:29.500Z",
      "content": "<p>Thanks Daisuke, some very interesting and useful information. I'm struggling to get past 0.92, you seem so far away at the top! I was just wondering, when you say use the entire audio wave. Do you, for example, split a 15 second training clip into equal portions of 5 seconds, so you have two extra pieces of training data? Have you also done any preprocessing to remove silent parts of the audio?</p>\n\n<p>Also, for anyone who uses Keras, the class imbalance can be fixed by creating a dict of class weights and providing it as an argument to the fit function. It definitely helped boost my score!</p>\n\n<p>&gt; class_weight: Optional dictionary mapping class indices (integers) to a weight (float) value, used for weighting the loss function (during training only). This can be useful to tell the model to \"pay more attention\" to samples from an under-represented class.</p>",
      "rawMarkdown": "Thanks Daisuke, some very interesting and useful information. I'm struggling to get past 0.92, you seem so far away at the top! I was just wondering, when you say use the entire audio wave. Do you, for example, split a 15 second training clip into equal portions of 5 seconds, so you have two extra pieces of training data? Have you also done any preprocessing to remove silent parts of the audio?\n\nAlso, for anyone who uses Keras, the class imbalance can be fixed by creating a dict of class weights and providing it as an argument to the fit function. It definitely helped boost my score!\n\n&gt; class_weight: Optional dictionary mapping class indices (integers) to a weight (float) value, used for weighting the loss function (during training only). This can be useful to tell the model to \"pay more attention\" to samples from an under-represented class.",
      "votes": 3,
      "replies": [
        {
          "id": 337718,
          "postDate": "2018-06-03T15:24:39.333Z",
          "content": "<p>Hi hyphmongo,</p>\n\n<ul>\n<li>When splitting long data, I convert entire wave into melspectrogram (not using MFCC), then cut them into pieces. And I put 10% overlap for cutting them, trying to preserve as many information as possible.</li>\n<li>I use not only removing silence but also repeating too short (or one shot sound), to make entire information density higher. I think, we not only have problem with imbalance number of sample but also imbalance information density among samples. And there could be tendency of sound information sparsity/ density among classes. I haven't measure/ visualize to prove it, but trying to resolve this issue improved my score.</li>\n</ul>",
          "rawMarkdown": "Hi hyphmongo,\n\n- When splitting long data, I convert entire wave into melspectrogram (not using MFCC), then cut them into pieces. And I put 10% overlap for cutting them, trying to preserve as many information as possible.\n- I use not only removing silence but also repeating too short (or one shot sound), to make entire information density higher. I think, we not only have problem with imbalance number of sample but also imbalance information density among samples. And there could be tendency of sound information sparsity/ density among classes. I haven't measure/ visualize to prove it, but trying to resolve this issue improved my score."
        },
        {
          "id": 337728,
          "postDate": "2018-06-03T15:53:14.493Z",
          "content": "<p>Thank you, will give this a try soon! I also found log mel spectrograms to score higher. Not sure if you also tried, but it could be good to also cut the testing samples and make an average of the predictions of each cut. </p>",
          "rawMarkdown": "Thank you, will give this a try soon! I also found log mel spectrograms to score higher. Not sure if you also tried, but it could be good to also cut the testing samples and make an average of the predictions of each cut. "
        },
        {
          "id": 337899,
          "postDate": "2018-06-04T02:13:22.800Z",
          "content": "<p>Yes I 'm taking average of predictions of all cut samples, and it pushed up the score.</p>",
          "rawMarkdown": "Yes I 'm taking average of predictions of all cut samples, and it pushed up the score."
        },
        {
          "id": 338141,
          "postDate": "2018-06-04T13:57:27.747Z",
          "content": "<p>excuse me, <a href=\"/daisukelab\">@daisukelab</a>. I'm still not sure how do you handle long data. After extracting melspectrogram and cutting into pieces. How do you decide to remove which piece? Setting a threshold of of the summation of the melspectrogram to decide silent or not?</p>",
          "rawMarkdown": "excuse me, @daisukelab. I'm still not sure how do you handle long data. After extracting melspectrogram and cutting into pieces. How do you decide to remove which piece? Setting a threshold of of the summation of the melspectrogram to decide silent or not?"
        },
        {
          "id": 338144,
          "postDate": "2018-06-04T14:08:54.163Z",
          "content": "<p>If it's 15s audio and splitting into 5s audio, we will have three 5s log mel spectrogram data.\nNothing to remove. Train them or evaluate for training sample or predict for test sample.\nIf we cut test sample into three pieces like this, calculate geometric mean of them and get final prediction for the original test sample. We are not sure which part have the important information, but I hope informative part would naturally yield higher probability, and it seems to be working.</p>",
          "rawMarkdown": "If it's 15s audio and splitting into 5s audio, we will have three 5s log mel spectrogram data.\nNothing to remove. Train them or evaluate for training sample or predict for test sample.\nIf we cut test sample into three pieces like this, calculate geometric mean of them and get final prediction for the original test sample. We are not sure which part have the important information, but I hope informative part would naturally yield higher probability, and it seems to be working."
        },
        {
          "id": 338155,
          "postDate": "2018-06-04T14:28:37.153Z",
          "content": "<p>Thank you for answering. </p>",
          "rawMarkdown": "Thank you for answering. "
        },
        {
          "id": 340750,
          "postDate": "2018-06-10T07:42:51.317Z",
          "content": "<p>Thanks for your replying! Your work is nice, but I want to know that what is the actual meaning of different models. Does it mean different architectures, different training technique or just different input data? Thanks again.</p>",
          "rawMarkdown": "Thanks for your replying! Your work is nice, but I want to know that what is the actual meaning of different models. Does it mean different architectures, different training technique or just different input data? Thanks again."
        },
        {
          "id": 340851,
          "postDate": "2018-06-10T13:17:20.087Z",
          "content": "<p>Hi @pang, let me get back to my work at <a href=\"https://www.kaggle.com/c/acoustic-scene-2018\">TUT Acoustic Scene Classification</a> competition.\nI used following models, and these three models see different inputs. Model 1 is normal, which sees the input as it is. Model 2 sees input which have only temporal changes. Model 3 sees input <em>without</em> temporal changes. I believe and actually have good result especially with model 2, that's my point.</p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">Simple CNN approach</a> - this is basic kernel.</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/sound-event-based-approach\">Sound event based approach</a> - this is an approach to emphasize temporal change (calling it as sound event here).</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/time-wise-mean-approach\">Time-wise mean approach</a> - this is inverse approach that completely ignore temporal change.</li>\n</ol>",
          "rawMarkdown": "Hi @pang, let me get back to my work at [TUT Acoustic Scene Classification][1] competition.\nI used following models, and these three models see different inputs. Model 1 is normal, which sees the input as it is. Model 2 sees input which have only temporal changes. Model 3 sees input _without_ temporal changes. I believe and actually have good result especially with model 2, that's my point.\n\n1. [Simple CNN approach](https://www.kaggle.com/daisukelab/simple-cnn-approach) - this is basic kernel.\n2. [Sound event based approach](https://www.kaggle.com/daisukelab/sound-event-based-approach) - this is an approach to emphasize temporal change (calling it as sound event here).\n3. [Time-wise mean approach](https://www.kaggle.com/daisukelab/time-wise-mean-approach) - this is inverse approach that completely ignore temporal change.\n\n  [1]: https://www.kaggle.com/c/acoustic-scene-2018"
        },
        {
          "id": 342251,
          "postDate": "2018-06-13T07:12:03.580Z",
          "content": "<p>Thank you! I will follow it.</p>",
          "rawMarkdown": "Thank you! I will follow it."
        }
      ]
    },
    {
      "id": 365720,
      "postDate": "2018-08-03T08:45:44.220Z",
      "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, thank you very much for your discussion and sharing. I started this competition with insights that you have shared. Those mixup, erase/cutout augmentation work pretty well! </p>",
      "rawMarkdown": "Hi @daisukelab, thank you very much for your discussion and sharing. I started this competition with insights that you have shared. Those mixup, erase/cutout augmentation work pretty well! ",
      "votes": 1,
      "replies": [
        {
          "id": 365776,
          "postDate": "2018-08-03T11:29:43.560Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> I agree with <a href=\"/thomeou\">@thomeou</a>. Thank you for sharing your insights on the several data augmentation techniques. They were very helpful indeed.</p>",
          "rawMarkdown": "@daisukelab I agree with @thomeou. Thank you for sharing your insights on the several data augmentation techniques. They were very helpful indeed.",
          "votes": 1
        },
        {
          "id": 367358,
          "postDate": "2018-08-07T15:21:06.767Z",
          "content": "<p>Hi <a href=\"/thomeou\">@thomeou</a>, that sounds good! Thanks for feedback :)</p>\n\n<p>Hi @Gyat, thank you! :)</p>",
          "rawMarkdown": "Hi @thomeou, that sounds good! Thanks for feedback :)\n\nHi @Gyat, thank you! :)"
        }
      ]
    },
    {
      "id": 345698,
      "postDate": "2018-06-20T09:48:05.337Z",
      "content": "<p>Just a question. How long do your models take to complete? Also, what configuration you are on? Like RAM, GPU etc.?</p>",
      "rawMarkdown": "Just a question. How long do your models take to complete? Also, what configuration you are on? Like RAM, GPU etc.?",
      "votes": 1,
      "replies": [
        {
          "id": 346024,
          "postDate": "2018-06-20T23:37:50.747Z",
          "content": "<p>Hi @Gyat,\nIt depends on the model, 4 hours with light model and 16 hours or more with heavy data or bigger model.\nMy PC has 48GB RAM and GTX1080Ti.</p>",
          "rawMarkdown": "Hi @Gyat,\nIt depends on the model, 4 hours with light model and 16 hours or more with heavy data or bigger model.\nMy PC has 48GB RAM and GTX1080Ti."
        }
      ]
    },
    {
      "id": 340706,
      "postDate": "2018-06-10T04:21:14.417Z",
      "content": "<p>Hi, thanks for your nice work! I have a question that my validation acc can not present the real level of the model. As the validation acc increases, it decreases in lb unexpectedly (I used 10-fold verification). Someone said the labels without verification may affect the verification set results. So I want to know how do you choose the verification set to ensure the result is true even in lb or whether you also met this problem? Thank you very much.</p>",
      "rawMarkdown": "Hi, thanks for your nice work! I have a question that my validation acc can not present the real level of the model. As the validation acc increases, it decreases in lb unexpectedly (I used 10-fold verification). Someone said the labels without verification may affect the verification set results. So I want to know how do you choose the verification set to ensure the result is true even in lb or whether you also met this problem? Thank you very much.",
      "votes": 1,
      "replies": [
        {
          "id": 340722,
          "postDate": "2018-06-10T05:52:39.477Z",
          "content": "<p>I have the same question :).\nI use 2 fold basically. And if I try more, validation acc improves very much but LB score degrades... the same thing with you.\nSo I still haven't find answer.</p>\n\n<ul>\n<li>I basically try to use different models or inputs as much as possible, then perform ensemble.\nI <em>guess</em> ensemble of many different perspective seems to avoid overfitting to the weak labels.\nVerifying predictions by looking from many view points would help finding false positives/ negatives.</li>\n<li>But I have to admit that one of my 5-fold single model has LB score of 0.940, this one is strong enough. Though ensembling this with other models doesn't improve... </li>\n<li>Now I'm working on two approach. One is re-labeling of weak train samples.\nThe other is finding best ensemble by choosing some class predictions from a model, then some other class preds from other model and so on.</li>\n</ul>",
          "rawMarkdown": "I have the same question :).\nI use 2 fold basically. And if I try more, validation acc improves very much but LB score degrades... the same thing with you.\nSo I still haven't find answer.\n\n- I basically try to use different models or inputs as much as possible, then perform ensemble.\nI _guess_ ensemble of many different perspective seems to avoid overfitting to the weak labels.\nVerifying predictions by looking from many view points would help finding false positives/ negatives.\n- But I have to admit that one of my 5-fold single model has LB score of 0.940, this one is strong enough. Though ensembling this with other models doesn't improve... \n- Now I'm working on two approach. One is re-labeling of weak train samples.\nThe other is finding best ensemble by choosing some class predictions from a model, then some other class preds from other model and so on."
        }
      ]
    },
    {
      "id": 337890,
      "postDate": "2018-06-04T01:44:34.817Z",
      "content": "<p>Thank you <a href=\"/daisukelab\">@daisukelab</a>, I've been trying to model the variable length data into one model, but my performance is always under 83%. It seems your experiments could be a good guidance.</p>\n\n<p>May I ask a question about mix-up method that: After mix-up process the data in every batch will have more than one incomplete labels, will you use Binary-Cross-Entropy for every outputs as multi-label classification, or still use Cross-Entropy on all outputs?</p>",
      "rawMarkdown": "Thank you @daisukelab, I've been trying to model the variable length data into one model, but my performance is always under 83%. It seems your experiments could be a good guidance.\n\nMay I ask a question about mix-up method that: After mix-up process the data in every batch will have more than one incomplete labels, will you use Binary-Cross-Entropy for every outputs as multi-label classification, or still use Cross-Entropy on all outputs?",
      "votes": 1,
      "replies": [
        {
          "id": 337967,
          "postDate": "2018-06-04T06:25:32.583Z",
          "content": "<p>Hi <a href=\"/idealwhite\">@idealwhite</a>, I just use categorical_crossentropy as usual.\nMy understanding is basically we make the labels as target with one-hot encoded:</p>\n\n<pre><code>[0, 0, 0, ..., 1, 0, 0, ..., 0]\n</code></pre>\n\n<p>Then if we apply mix-up, two targets are mixed with random proportion. It would be like this:</p>\n\n<pre><code>[0, 0, 0.3, 0, 0, ..., 0.7, 0, 0, ..., 0]\n</code></pre>\n\n<p>This doesn't break what is supposed to feed to categorical_crossentropy, and actually work.</p>",
          "rawMarkdown": "Hi @idealwhite, I just use categorical_crossentropy as usual.\nMy understanding is basically we make the labels as target with one-hot encoded:\n\n    [0, 0, 0, ..., 1, 0, 0, ..., 0]\n\nThen if we apply mix-up, two targets are mixed with random proportion. It would be like this:\n\n    [0, 0, 0.3, 0, 0, ..., 0.7, 0, 0, ..., 0]\n\nThis doesn't break what is supposed to feed to categorical_crossentropy, and actually work."
        }
      ]
    },
    {
      "id": 353215,
      "postDate": "2018-07-06T07:45:47.317Z",
      "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>. Thank you very much for sharing your work! As many participants of this challenge (including myself) have benefited from your willingness to share your insights, it would be nice to be able to at least honor your work in the acknowledgements section of our submission for this challenge. Is there any chance we can get your real name for this purpose? :) (In case you have previously published anything on this topic when participating at the TUT Acoustic Scene Classification competition, I would be equally thankful for a source.)</p>",
      "rawMarkdown": "Hi @daisukelab. Thank you very much for sharing your work! As many participants of this challenge (including myself) have benefited from your willingness to share your insights, it would be nice to be able to at least honor your work in the acknowledgements section of our submission for this challenge. Is there any chance we can get your real name for this purpose? :) (In case you have previously published anything on this topic when participating at the TUT Acoustic Scene Classification competition, I would be equally thankful for a source.)",
      "votes": 2,
      "replies": [
        {
          "id": 353553,
          "postDate": "2018-07-07T03:16:15.200Z",
          "content": "<p>Hi @Kevin, it's very nice of you for asking that.\nIt's no problem that sharing ideas also benefit me that I could confirm how effective they are, I'm hoping to make virtuous circle, and it seems to be working. ;)\nI added my Linkedin account link to my profile, then please find it there. Thank you :)</p>",
          "rawMarkdown": "Hi @Kevin, it's very nice of you for asking that.\nIt's no problem that sharing ideas also benefit me that I could confirm how effective they are, I'm hoping to make virtuous circle, and it seems to be working. ;)\nI added my Linkedin account link to my profile, then please find it there. Thank you :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 346927,
      "postDate": "2018-06-22T19:06:57.737Z",
      "content": "<p>Hey <a href=\"/daisukelab\">@daisukelab</a>, I'd also like to thank you for sharing your insights. Some of your data augmentation methods helped to improve my score drastically. I think that the key ingredient is to apply all kinds of data augmentation. I got a decent score with only two 5-fold models. Would you mind telling us how many models you are currently using?</p>",
      "rawMarkdown": "Hey @daisukelab, I'd also like to thank you for sharing your insights. Some of your data augmentation methods helped to improve my score drastically. I think that the key ingredient is to apply all kinds of data augmentation. I got a decent score with only two 5-fold models. Would you mind telling us how many models you are currently using?",
      "replies": [
        {
          "id": 346938,
          "postDate": "2018-06-22T20:26:37.010Z",
          "content": "<p>Hi @Marcel, it sounds very good that augmentation matters.\nAnd it is great that you use only two 5-fold models, sounds better solution than me.\nI'm using many, many, and many models :). Current score is composed of 3 models of all 2-folds, and they have all different CNN, different handling parameter of input.</p>",
          "rawMarkdown": "Hi @Marcel, it sounds very good that augmentation matters.\nAnd it is great that you use only two 5-fold models, sounds better solution than me.\nI'm using many, many, and many models :). Current score is composed of 3 models of all 2-folds, and they have all different CNN, different handling parameter of input."
        },
        {
          "id": 352766,
          "postDate": "2018-07-05T04:13:27.483Z",
          "content": "<p>Hi @Marcel,  you have a good LB score with a simple model. I have the problem that when training with melspectrogram and CNN, overfitting can not be solved, the accuracy for training and validation datasets could be 0.9817 vs 0.7921 after about 100 epochs.\nI have used dropout, regulization, but overfiting is still there.\nI wonder if you have any experience similar?</p>",
          "rawMarkdown": "Hi @Marcel,  you have a good LB score with a simple model. I have the problem that when training with melspectrogram and CNN, overfitting can not be solved, the accuracy for training and validation datasets could be 0.9817 vs 0.7921 after about 100 epochs.\nI have used dropout, regulization, but overfiting is still there.\nI wonder if you have any experience similar?"
        },
        {
          "id": 352772,
          "postDate": "2018-07-05T04:34:21.660Z",
          "content": "<p>Hi Sailor, let me share regarding your problem. Strong augmentation like mixup will solve it, I'm confident about it. My acc always converges to like train:0.85 and valid:0.90 while training.</p>",
          "rawMarkdown": "Hi Sailor, let me share regarding your problem. Strong augmentation like mixup will solve it, I'm confident about it. My acc always converges to like train:0.85 and valid:0.90 while training."
        },
        {
          "id": 352787,
          "postDate": "2018-07-05T05:52:44.167Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> Wow, thanks for your answer. I'll follow your mixup augmentation ~\nAs you wrote in another topic, trim silence, split long samples and mixup can all be effective, right?</p>\n\n<p>In addition, if you have tried to manually verify the train data label? Does that help a lot?</p>",
          "rawMarkdown": "@daisukelab Wow, thanks for your answer. I'll follow your mixup augmentation ~\nAs you wrote in another topic, trim silence, split long samples and mixup can all be effective, right?\n\nIn addition, if you have tried to manually verify the train data label? Does that help a lot?"
        },
        {
          "id": 352808,
          "postDate": "2018-07-05T06:42:50.480Z",
          "content": "<p>Hi Sailor, mixup works fine and you can try other augmentation techniques too. Trimming and splitting both work fine.</p>\n\n<p>And regarding manually verifying train label, you can find discussion here:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052</a></p>",
          "rawMarkdown": "Hi Sailor, mixup works fine and you can try other augmentation techniques too. Trimming and splitting both work fine.\n\nAnd regarding manually verifying train label, you can find discussion here:\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/58052"
        },
        {
          "id": 352835,
          "postDate": "2018-07-05T08:17:09.893Z",
          "content": "<p>Got it, thanks.</p>",
          "rawMarkdown": "Got it, thanks."
        },
        {
          "id": 353220,
          "postDate": "2018-07-06T08:01:39.977Z",
          "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, mixup seems working now. I use 10-folder cross-validation, the accuracy for two folders can be as bellow:</p>\n\n<pre><code>loss: 0.9709 - acc: 0.8900 - val_loss: 0.5129 - val_acc: 0.9137\nloss: 0.7581 - acc: 0.9176 - val_loss: 1.5712 - val_acc: 0.8642\n</code></pre>\n\n<p>I train only 150epochs with early stopping, seems not converge enough now. The generalization of model seems improved a lot.</p>\n\n<p>But map@3 could not be applied after each epoch, I'll see the LB score after the training finished.</p>",
          "rawMarkdown": "Hi @daisukelab, mixup seems working now. I use 10-folder cross-validation, the accuracy for two folders can be as bellow:\n\n    loss: 0.9709 - acc: 0.8900 - val_loss: 0.5129 - val_acc: 0.9137\n    loss: 0.7581 - acc: 0.9176 - val_loss: 1.5712 - val_acc: 0.8642\n\nI train only 150epochs with early stopping, seems not converge enough now. The generalization of model seems improved a lot.\n\nBut map@3 could not be applied after each epoch, I'll see the LB score after the training finished."
        },
        {
          "id": 353555,
          "postDate": "2018-07-07T03:27:14.157Z",
          "content": "<p>Hi @Sailor, that sounds good. Increasing augmentation should make training converge longer, so it seems to be good. I hope you have better results.</p>",
          "rawMarkdown": "Hi @Sailor, that sounds good. Increasing augmentation should make training converge longer, so it seems to be good. I hope you have better results."
        },
        {
          "id": 354663,
          "postDate": "2018-07-10T02:24:08.093Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> ,  I have tried some different parameters to train. The 10folder single model's predictions seems not improved too much. The LB scores for the melspectrogram + cnn model is 0.907 now, far from your 0.94 for single model. I wonder what features and model your best single model use?</p>",
          "rawMarkdown": "@daisukelab ,  I have tried some different parameters to train. The 10folder single model's predictions seems not improved too much. The LB scores for the melspectrogram + cnn model is 0.907 now, far from your 0.94 for single model. I wonder what features and model your best single model use?"
        },
        {
          "id": 354685,
          "postDate": "2018-07-10T04:27:33.750Z",
          "content": "<p>Hi @Sailor</p>\n\n<ul>\n<li>Thanks to @Cocoxili, preprocessing follows almost the same parameter as kernel <a href=\"https://www.kaggle.com/aadan2017/log-mel-features-shown-by-category\">'Log-mel features shown by category.'</a></li>\n<li><p>Model is SE-ResNet-50 with small modification to make it smaller/thinner. I'm using following implementation:</p>\n\n<p><a href=\"https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se_resnet.py\">https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se_resnet.py</a>\n<a href=\"https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se.py\">https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se.py</a></p></li>\n</ul>\n\n<p>But please note that my other light model based on AlexNet which is much smaller than SE-ResNet-50 has LB score 0.933, I think SE-ResNet is too much.</p>",
          "rawMarkdown": "Hi @Sailor\n\n- Thanks to @Cocoxili, preprocessing follows almost the same parameter as kernel ['Log-mel features shown by category.'][1]\n- Model is SE-ResNet-50 with small modification to make it smaller/thinner. I'm using following implementation:\n\n    https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se_resnet.py\n    https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se.py\n\nBut please note that my other light model based on AlexNet which is much smaller than SE-ResNet-50 has LB score 0.933, I think SE-ResNet is too much.\n\n  [1]: https://www.kaggle.com/aadan2017/log-mel-features-shown-by-category"
        },
        {
          "id": 354757,
          "postDate": "2018-07-10T07:36:47.230Z",
          "content": "<p>Thanks, <a href=\"/daisukelab\">@daisukelab</a>. I also use mel-spectrogram as the input feature, it got better score than mfcc or stft. When you say light model based on AlexNet, do you mean this <a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">https://www.kaggle.com/daisukelab/simple-cnn-approach</a> ? I am just tuning the parameter for my 128*176 input, but still converge slow.</p>",
          "rawMarkdown": "Thanks, @daisukelab. I also use mel-spectrogram as the input feature, it got better score than mfcc or stft. When you say light model based on AlexNet, do you mean this https://www.kaggle.com/daisukelab/simple-cnn-approach ? I am just tuning the parameter for my 128*176 input, but still converge slow."
        },
        {
          "id": 354769,
          "postDate": "2018-07-10T08:03:29.900Z",
          "content": "<p>Hi @Sailor,</p>\n\n<p>Yes that's the exact model. :)</p>\n\n<p>And regarding the parameter to convert to mel-spectrogram, I guess this leaves improvement for you. If you follow exact the same n_fft, n_mels and other parameter as same as <a href=\"https://www.kaggle.com/aadan2017/log-mel-features-shown-by-category\"> 'Log-mel features shown by category'</a>, you might see something different. In my case, validation accuracy increased about +0.1% higher than before, though it doesn't directly link to the LB score. And the length of the input might be short, is it 1-2s? I'm using 4 or more seconds with recent attempts.</p>",
          "rawMarkdown": "Hi @Sailor,\n\nYes that's the exact model. :)\n\nAnd regarding the parameter to convert to mel-spectrogram, I guess this leaves improvement for you. If you follow exact the same n_fft, n_mels and other parameter as same as [ 'Log-mel features shown by category'][1], you might see something different. In my case, validation accuracy increased about +0.1% higher than before, though it doesn't directly link to the LB score. And the length of the input might be short, is it 1-2s? I'm using 4 or more seconds with recent attempts.\n\n  [1]: https://www.kaggle.com/aadan2017/log-mel-features-shown-by-category"
        },
        {
          "id": 354782,
          "postDate": "2018-07-10T08:38:12.377Z",
          "content": "<p>Yeah, I will try when the server is available :(. How many epochs your model generally early stops? To get a 10 folds cross validated model, it often spends at least one night or more.\nThis input is 2s. My best score is got when the length is 3s. You got better results using 4s or more? That's may be much slower... </p>",
          "rawMarkdown": "Yeah, I will try when the server is available :(. How many epochs your model generally early stops? To get a 10 folds cross validated model, it often spends at least one night or more.\nThis input is 2s. My best score is got when the length is 3s. You got better results using 4s or more? That's may be much slower... "
        },
        {
          "id": 354934,
          "postDate": "2018-07-10T14:09:56.177Z",
          "content": "<p>Hi, regarding epochs, I don't use early stopping but best val_acc checkpoint. It really depends but usually no later than 300 epochs. It really takes long to train, about 5-10hrs. Running PC all day...</p>",
          "rawMarkdown": "Hi, regarding epochs, I don't use early stopping but best val_acc checkpoint. It really depends but usually no later than 300 epochs. It really takes long to train, about 5-10hrs. Running PC all day..."
        },
        {
          "id": 358000,
          "postDate": "2018-07-17T10:22:51.680Z",
          "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, do you mind telling the performance of crnn? I am trying to use crnn structure based on a dcase2017 paper. The performance is much worse that cnn, which is confused.</p>",
          "rawMarkdown": "Hi @daisukelab, do you mind telling the performance of crnn? I am trying to use crnn structure based on a dcase2017 paper. The performance is much worse that cnn, which is confused."
        },
        {
          "id": 358055,
          "postDate": "2018-07-17T12:50:19.547Z",
          "content": "<p>Hi Sailor, I haven't tried CRNN so much, actually submitted only once. It was 0.913. BTW, recently I'm not looping short samples, and use short ones as is with zero paddings; it was found for me to be better not to repeat... I'm sorry if you are doing that by following my comment.</p>",
          "rawMarkdown": "Hi Sailor, I haven't tried CRNN so much, actually submitted only once. It was 0.913. BTW, recently I'm not looping short samples, and use short ones as is with zero paddings; it was found for me to be better not to repeat... I'm sorry if you are doing that by following my comment."
        },
        {
          "id": 359325,
          "postDate": "2018-07-19T21:48:59.017Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>: Not repeating worked better for you? So you just pad with zeros if the slices are shorter?</p>\n\n<p>Also, recurrence can work but I found it more efficient to slice the samples (with overlap) so that they all have the same size. Training CRNN on the whole samples complicates a lot of stuff (e.g. augmentation) and takes way too much time before you find the right hyperparameters (GRU vs. LSTM, # layers, # units, ...).</p>\n\n<p>Do you still use SENet 50? I tried it too, but it had unnecessarily many parameters.</p>",
          "rawMarkdown": "@daisukelab: Not repeating worked better for you? So you just pad with zeros if the slices are shorter?\n\nAlso, recurrence can work but I found it more efficient to slice the samples (with overlap) so that they all have the same size. Training CRNN on the whole samples complicates a lot of stuff (e.g. augmentation) and takes way too much time before you find the right hyperparameters (GRU vs. LSTM, # layers, # units, ...).\n\nDo you still use SENet 50? I tried it too, but it had unnecessarily many parameters."
        },
        {
          "id": 359871,
          "postDate": "2018-07-21T00:34:06.670Z",
          "content": "<p>Hi @Daniel,</p>\n\n<p>Yes currently I'm not using repeat of short samples. It worked well with other parameters when I started using it once, but now I cannot evaluate if it works well or not.\nI observed that it has side effect that the repeated sample is confused with Hi-hat or other drum-like labels, though it seems not to be frequent.</p>\n\n<p>And regarding models, I agree that it seems to be  difficult to make Recurrent NN work in this task.\nRegarding SENets, I use SEResNet-50 (actually less than 50, modified a little) that has best balance between good modeling and overfitting.</p>\n\n<p>To be honest, I've tried almost all the models listed in <a href=\"https://keras.io/applications/\">Keras applications</a>, as if ordering all the beers on list at a bar. 😅\nMany are too much and overfitted (or would needed more augmentation), here's funny example with Inception v3.</p>\n\n<pre><code>Epoch 8/300\n - 221s - loss: 1.7456 - acc: 0.7004 - val_loss: 2.6311 - val_acc: 0.2956\n</code></pre>\n\n<p>VGG16 worked fine, but I already have SEResNet working better.\nOf course tried ImageNet pre-trained weights, though didn't show big difference. (I'm not using it anymore)</p>\n\n<p>Then for me, any model is ok as long as it performs fine...\nTweaking how to expose data to models; preprocessing parameters and training data selection...</p>",
          "rawMarkdown": "Hi @Daniel,\n\nYes currently I'm not using repeat of short samples. It worked well with other parameters when I started using it once, but now I cannot evaluate if it works well or not.\nI observed that it has side effect that the repeated sample is confused with Hi-hat or other drum-like labels, though it seems not to be frequent.\n\nAnd regarding models, I agree that it seems to be  difficult to make Recurrent NN work in this task.\nRegarding SENets, I use SEResNet-50 (actually less than 50, modified a little) that has best balance between good modeling and overfitting.\n\nTo be honest, I've tried almost all the models listed in [Keras applications][1], as if ordering all the beers on list at a bar. 😅\nMany are too much and overfitted (or would needed more augmentation), here's funny example with Inception v3.\n\n    Epoch 8/300\n     - 221s - loss: 1.7456 - acc: 0.7004 - val_loss: 2.6311 - val_acc: 0.2956\n\nVGG16 worked fine, but I already have SEResNet working better.\nOf course tried ImageNet pre-trained weights, though didn't show big difference. (I'm not using it anymore)\n\nThen for me, any model is ok as long as it performs fine...\nTweaking how to expose data to models; preprocessing parameters and training data selection...\n\n  [1]: https://keras.io/applications/"
        }
      ]
    },
    {
      "id": 346446,
      "postDate": "2018-06-21T17:54:51.503Z",
      "content": "<p>Hi  <a href=\"/daisukelab\">@daisukelab</a> . Thanks you for sharing the nice work.\nI'm trying CRNN to train the model. However, I encounter several problem. I have trouble to get pass 0.85. May I ask some question about this?</p>\n\n<ol>\n<li><p>The CRNN only accept unified length, I've tried setting length as 128, 152, 512.I'm interesting in how you decide the length of Input?</p></li>\n<li><p>I pad the data with zeros after convert it into melspectrogram. Is it better to use other padding method? I've tried other method ,but I'm not sure the accuracy drop because of the padding method or the model. Is it different to cut the data before convert or after convert?</p></li>\n<li><p>While training the CRNN model, my model's Loss dropped rapidly after hundreds of epochs (train acci and val accu is both about 0.7 when this happened). And then the loss stuck in about 1*e-7.  I've tried changing the learning rate  and activate function. However, this problem still happened. Have you ever encounter the same problem while training RNN?</p></li>\n</ol>\n\n<p>My model using keras2.0.8, Tensorflow1.4.0. My CRNN model is  replace the full connection layer of CNN to bidirectional LSTM and add another LSTM for output.</p>",
      "rawMarkdown": "Hi  @daisukelab . Thanks you for sharing the nice work.\nI'm trying CRNN to train the model. However, I encounter several problem. I have trouble to get pass 0.85. May I ask some question about this?\n\n1. The CRNN only accept unified length, I've tried setting length as 128, 152, 512.I'm interesting in how you decide the length of Input?\n\n2. I pad the data with zeros after convert it into melspectrogram. Is it better to use other padding method? I've tried other method ,but I'm not sure the accuracy drop because of the padding method or the model. Is it different to cut the data before convert or after convert?\n\n3. While training the CRNN model, my model's Loss dropped rapidly after hundreds of epochs (train acci and val accu is both about 0.7 when this happened). And then the loss stuck in about 1*e-7.  I've tried changing the learning rate  and activate function. However, this problem still happened. Have you ever encounter the same problem while training RNN?\n\nMy model using keras2.0.8, Tensorflow1.4.0. My CRNN model is  replace the full connection layer of CNN to bidirectional LSTM and add another LSTM for output.",
      "replies": [
        {
          "id": 346937,
          "postDate": "2018-06-22T20:17:26.100Z",
          "content": "<p>Hi @r06921058_&gt;.O,</p>\n\n<ol>\n<li><p>Regarding the length to cut from the original sample, it is one of important parameters as far as I have tried. And cutting from <em>where</em>, as well as <em>disposing or using the other half of cut wave</em> are also important. I’m ensembling models that have different parameters of input sample handling, there seems to be no single answer. Some sample starts from the beginning, some others have sound events in the middle or at the end.</p></li>\n<li><p>Regarding padding zeros, I’ve got better result by <em>repeating</em> the wave until it comes to your desired length rather than padding something. As far as I was checking samples, there seems to be gap of amount of information between longer samples and shorter samples. This seems to be the same thing with imbalance of number of samples among classes, we can improve performance by balancing it. My guess is padding doesn’t solve imbalance of information among different length of samples, but repeating should basically resolve it.</p></li>\n<li><p>I have tried CNN only, so I don’t have exact answer unfortunately.\nBut your model seems to be complicated, then if I would debug it I would confirm CNN is working fine by using CNN part only first. Then checking for the details of entire model.</p></li>\n</ol>",
          "rawMarkdown": "Hi @r06921058_&gt;.O,\n\n1. Regarding the length to cut from the original sample, it is one of important parameters as far as I have tried. And cutting from _where_, as well as _disposing or using the other half of cut wave_ are also important. I’m ensembling models that have different parameters of input sample handling, there seems to be no single answer. Some sample starts from the beginning, some others have sound events in the middle or at the end.\n\n2. Regarding padding zeros, I’ve got better result by _repeating_ the wave until it comes to your desired length rather than padding something. As far as I was checking samples, there seems to be gap of amount of information between longer samples and shorter samples. This seems to be the same thing with imbalance of number of samples among classes, we can improve performance by balancing it. My guess is padding doesn’t solve imbalance of information among different length of samples, but repeating should basically resolve it.\n\n3. I have tried CNN only, so I don’t have exact answer unfortunately.\nBut your model seems to be complicated, then if I would debug it I would confirm CNN is working fine by using CNN part only first. Then checking for the details of entire model."
        },
        {
          "id": 346977,
          "postDate": "2018-06-22T23:35:36.360Z",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>,Thanks for answering these question! I'll try to use repeat to pad, and try other split method. </p>\n\n<p>About 2. If I understand correctly, you repeat the data to the same length before split them into pieces to handle the problem of imbalance? Or you used other method to deal with imbalance problem? Thanks for reply.</p>",
          "rawMarkdown": "@daisukelab,Thanks for answering these question! I'll try to use repeat to pad, and try other split method. \n\nAbout 2. If I understand correctly, you repeat the data to the same length before split them into pieces to handle the problem of imbalance? Or you used other method to deal with imbalance problem? Thanks for reply.\n\n"
        },
        {
          "id": 346992,
          "postDate": "2018-06-23T00:20:29.610Z",
          "content": "<p>Hi, I set a length of one input wave, then apply preprocessing of raw wave as follows:</p>\n\n<ol>\n<li>Trim silence.</li>\n<li>Repeat if short.</li>\n<li>Split if long enough with overlaps.</li>\n</ol>",
          "rawMarkdown": "Hi, I set a length of one input wave, then apply preprocessing of raw wave as follows:\n\n1. Trim silence.\n2. Repeat if short.\n3. Split if long enough with overlaps.",
          "votes": 1
        },
        {
          "id": 347722,
          "postDate": "2018-06-25T07:16:57.387Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        },
        {
          "id": 352764,
          "postDate": "2018-07-05T04:07:10.767Z",
          "content": "<p>Hi daisuke,\nThanks for your share. I wonder when you do silence trim, do you mean trim the silence part in the begining and end of audio? how much is the dB value to judge silence. I just cut the -30dB part, with the same network, the LB score decrease...</p>",
          "rawMarkdown": "Hi daisuke,\nThanks for your share. I wonder when you do silence trim, do you mean trim the silence part in the begining and end of audio? how much is the dB value to judge silence. I just cut the -30dB part, with the same network, the LB score decrease..."
        },
        {
          "id": 352769,
          "postDate": "2018-07-05T04:30:25.240Z",
          "content": "<p>Hi Sailor, my code is simple:</p>\n\n<pre><code>y, _ = librosa.effects.trim(y)\n</code></pre>\n\n<p>It trims both ends.</p>\n\n<p>BTW, r06921058_&gt;.O and all,\nI could be wrong about repeating audio.\nUsing short audio filled with zero padding is now showing good score (though it doesn't exceed current best).</p>",
          "rawMarkdown": "Hi Sailor, my code is simple:\n\n    y, _ = librosa.effects.trim(y)\n\nIt trims both ends.\n\nBTW, r06921058_&gt;.O and all,\nI could be wrong about repeating audio.\nUsing short audio filled with zero padding is now showing good score (though it doesn't exceed current best)."
        },
        {
          "id": 352784,
          "postDate": "2018-07-05T05:43:43.517Z",
          "content": "<p>Thanks, <a href=\"/daisukelab\">@daisukelab</a>.\nSo, trim can help to improve the score?\nI also use this function, but I set the trim_level as 30dB, to remove silence enough. Maybe I cut too lot to get worse score.</p>\n\n<pre><code>data, _ = librosa.effects.trim(data, top_db=trim_level)\n</code></pre>\n\n<p>It seems the default top_db you use is 60dB.</p>\n\n<pre><code>librosa.effects.trim(y, top_db=60, ref=&lt;function amax&gt;, frame_length=2048, hop_length=512)\n</code></pre>\n\n<p><a href=\"http://librosa.github.io/librosa/generated/librosa.effects.trim.html#librosa.effects.trim\">http://librosa.github.io/librosa/generated/librosa.effects.trim.html#librosa.effects.trim</a></p>",
          "rawMarkdown": "Thanks, @daisukelab.\nSo, trim can help to improve the score?\nI also use this function, but I set the trim_level as 30dB, to remove silence enough. Maybe I cut too lot to get worse score.\n\n    data, _ = librosa.effects.trim(data, top_db=trim_level)\n\nIt seems the default top_db you use is 60dB.\n\n    librosa.effects.trim(y, top_db=60, ref="
        },
        {
          "id": 352806,
          "postDate": "2018-07-05T06:37:50.590Z",
          "content": "<p>Hi Sailor, yes I'm using default setting without problem. I haven't compare results for this single parameter - but I think trimming should work better as long as it is used properly, and the default parameter seems to be fine.</p>",
          "rawMarkdown": "Hi Sailor, yes I'm using default setting without problem. I haven't compare results for this single parameter - but I think trimming should work better as long as it is used properly, and the default parameter seems to be fine."
        }
      ]
    },
    {
      "id": 361387,
      "postDate": "2018-07-24T10:59:37.887Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 361390,
          "postDate": "2018-07-24T11:14:40.597Z",
          "content": "<p>Hi rrr, I'm sorry I cannot help with 1-D array solution for that extent, didn't try 1-D...\nIt should work but dunno why, ... did you try 2-D? It works fine for me.</p>",
          "rawMarkdown": "Hi rrr, I'm sorry I cannot help with 1-D array solution for that extent, didn't try 1-D...\nIt should work but dunno why, ... did you try 2-D? It works fine for me."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 336992,
      "author_name": "hyphmongo",
      "author_url": "",
      "post_date": "2018-06-01T17:44:29.500000",
      "content": "<p>Thanks Daisuke, some very interesting and useful information. I'm struggling to get past 0.92, you seem so far away at the top! I was just wondering, when you say use the entire audio wave. Do you, for example, split a 15 second training clip into equal portions of 5 seconds, so you have two extra pieces of training data? Have you also done any preprocessing to remove silent parts of the audio?</p>\n\n<p>Also, for anyone who uses Keras, the class imbalance can be fixed by creating a dict of class weights and providing it as an argument to the fit function. It definitely helped boost my score!</p>\n\n<p>&gt; class_weight: Optional dictionary mapping class indices (integers) to a weight (float) value, used for weighting the loss function (during training only). This can be useful to tell the model to \"pay more attention\" to samples from an under-represented class.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 337718,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-03T15:24:39.333000",
          "content": "<p>Hi hyphmongo,</p>\n\n<ul>\n<li>When splitting long data, I convert entire wave into melspectrogram (not using MFCC), then cut them into pieces. And I put 10% overlap for cutting them, trying to preserve as many information as possible.</li>\n<li>I use not only removing silence but also repeating too short (or one shot sound), to make entire information density higher. I think, we not only have problem with imbalance number of sample but also imbalance information density among samples. And there could be tendency of sound information sparsity/ density among classes. I haven't measure/ visualize to prove it, but trying to resolve this issue improved my score.</li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337728,
          "author_name": "hyphmongo",
          "author_url": "",
          "post_date": "2018-06-03T15:53:14.493000",
          "content": "<p>Thank you, will give this a try soon! I also found log mel spectrograms to score higher. Not sure if you also tried, but it could be good to also cut the testing samples and make an average of the predictions of each cut. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 337899,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-04T02:13:22.800000",
          "content": "<p>Yes I 'm taking average of predictions of all cut samples, and it pushed up the score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338141,
          "author_name": "interstellar",
          "author_url": "",
          "post_date": "2018-06-04T13:57:27.747000",
          "content": "<p>excuse me, <a href=\"/daisukelab\">@daisukelab</a>. I'm still not sure how do you handle long data. After extracting melspectrogram and cutting into pieces. How do you decide to remove which piece? Setting a threshold of of the summation of the melspectrogram to decide silent or not?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338144,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-04T14:08:54.163000",
          "content": "<p>If it's 15s audio and splitting into 5s audio, we will have three 5s log mel spectrogram data.\nNothing to remove. Train them or evaluate for training sample or predict for test sample.\nIf we cut test sample into three pieces like this, calculate geometric mean of them and get final prediction for the original test sample. We are not sure which part have the important information, but I hope informative part would naturally yield higher probability, and it seems to be working.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 338155,
          "author_name": "interstellar",
          "author_url": "",
          "post_date": "2018-06-04T14:28:37.153000",
          "content": "<p>Thank you for answering. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340750,
          "author_name": "pang",
          "author_url": "",
          "post_date": "2018-06-10T07:42:51.317000",
          "content": "<p>Thanks for your replying! Your work is nice, but I want to know that what is the actual meaning of different models. Does it mean different architectures, different training technique or just different input data? Thanks again.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 340851,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-10T13:17:20.087000",
          "content": "<p>Hi @pang, let me get back to my work at <a href=\"https://www.kaggle.com/c/acoustic-scene-2018\">TUT Acoustic Scene Classification</a> competition.\nI used following models, and these three models see different inputs. Model 1 is normal, which sees the input as it is. Model 2 sees input which have only temporal changes. Model 3 sees input <em>without</em> temporal changes. I believe and actually have good result especially with model 2, that's my point.</p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">Simple CNN approach</a> - this is basic kernel.</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/sound-event-based-approach\">Sound event based approach</a> - this is an approach to emphasize temporal change (calling it as sound event here).</li>\n<li><a href=\"https://www.kaggle.com/daisukelab/time-wise-mean-approach\">Time-wise mean approach</a> - this is inverse approach that completely ignore temporal change.</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 342251,
          "author_name": "pang",
          "author_url": "",
          "post_date": "2018-06-13T07:12:03.580000",
          "content": "<p>Thank you! I will follow it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 365720,
      "author_name": "thomeou",
      "author_url": "",
      "post_date": "2018-08-03T08:45:44.220000",
      "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, thank you very much for your discussion and sharing. I started this competition with insights that you have shared. Those mixup, erase/cutout augmentation work pretty well! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 365776,
          "author_name": "Gyat",
          "author_url": "",
          "post_date": "2018-08-03T11:29:43.560000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> I agree with <a href=\"/thomeou\">@thomeou</a>. Thank you for sharing your insights on the several data augmentation techniques. They were very helpful indeed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 367358,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-08-07T15:21:06.767000",
          "content": "<p>Hi <a href=\"/thomeou\">@thomeou</a>, that sounds good! Thanks for feedback :)</p>\n\n<p>Hi @Gyat, thank you! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 345698,
      "author_name": "Gyat",
      "author_url": "",
      "post_date": "2018-06-20T09:48:05.337000",
      "content": "<p>Just a question. How long do your models take to complete? Also, what configuration you are on? Like RAM, GPU etc.?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 346024,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-20T23:37:50.747000",
          "content": "<p>Hi @Gyat,\nIt depends on the model, 4 hours with light model and 16 hours or more with heavy data or bigger model.\nMy PC has 48GB RAM and GTX1080Ti.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 340706,
      "author_name": "pang",
      "author_url": "",
      "post_date": "2018-06-10T04:21:14.417000",
      "content": "<p>Hi, thanks for your nice work! I have a question that my validation acc can not present the real level of the model. As the validation acc increases, it decreases in lb unexpectedly (I used 10-fold verification). Someone said the labels without verification may affect the verification set results. So I want to know how do you choose the verification set to ensure the result is true even in lb or whether you also met this problem? Thank you very much.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 340722,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-10T05:52:39.477000",
          "content": "<p>I have the same question :).\nI use 2 fold basically. And if I try more, validation acc improves very much but LB score degrades... the same thing with you.\nSo I still haven't find answer.</p>\n\n<ul>\n<li>I basically try to use different models or inputs as much as possible, then perform ensemble.\nI <em>guess</em> ensemble of many different perspective seems to avoid overfitting to the weak labels.\nVerifying predictions by looking from many view points would help finding false positives/ negatives.</li>\n<li>But I have to admit that one of my 5-fold single model has LB score of 0.940, this one is strong enough. Though ensembling this with other models doesn't improve... </li>\n<li>Now I'm working on two approach. One is re-labeling of weak train samples.\nThe other is finding best ensemble by choosing some class predictions from a model, then some other class preds from other model and so on.</li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 337890,
      "author_name": "idealwhite",
      "author_url": "",
      "post_date": "2018-06-04T01:44:34.817000",
      "content": "<p>Thank you <a href=\"/daisukelab\">@daisukelab</a>, I've been trying to model the variable length data into one model, but my performance is always under 83%. It seems your experiments could be a good guidance.</p>\n\n<p>May I ask a question about mix-up method that: After mix-up process the data in every batch will have more than one incomplete labels, will you use Binary-Cross-Entropy for every outputs as multi-label classification, or still use Cross-Entropy on all outputs?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 337967,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-04T06:25:32.583000",
          "content": "<p>Hi <a href=\"/idealwhite\">@idealwhite</a>, I just use categorical_crossentropy as usual.\nMy understanding is basically we make the labels as target with one-hot encoded:</p>\n\n<pre><code>[0, 0, 0, ..., 1, 0, 0, ..., 0]\n</code></pre>\n\n<p>Then if we apply mix-up, two targets are mixed with random proportion. It would be like this:</p>\n\n<pre><code>[0, 0, 0.3, 0, 0, ..., 0.7, 0, 0, ..., 0]\n</code></pre>\n\n<p>This doesn't break what is supposed to feed to categorical_crossentropy, and actually work.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 353215,
      "author_name": "Kevin Wilkinghoff",
      "author_url": "",
      "post_date": "2018-07-06T07:45:47.317000",
      "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>. Thank you very much for sharing your work! As many participants of this challenge (including myself) have benefited from your willingness to share your insights, it would be nice to be able to at least honor your work in the acknowledgements section of our submission for this challenge. Is there any chance we can get your real name for this purpose? :) (In case you have previously published anything on this topic when participating at the TUT Acoustic Scene Classification competition, I would be equally thankful for a source.)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 353553,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-07T03:16:15.200000",
          "content": "<p>Hi @Kevin, it's very nice of you for asking that.\nIt's no problem that sharing ideas also benefit me that I could confirm how effective they are, I'm hoping to make virtuous circle, and it seems to be working. ;)\nI added my Linkedin account link to my profile, then please find it there. Thank you :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 346927,
      "author_name": "Marcel Lederle",
      "author_url": "",
      "post_date": "2018-06-22T19:06:57.737000",
      "content": "<p>Hey <a href=\"/daisukelab\">@daisukelab</a>, I'd also like to thank you for sharing your insights. Some of your data augmentation methods helped to improve my score drastically. I think that the key ingredient is to apply all kinds of data augmentation. I got a decent score with only two 5-fold models. Would you mind telling us how many models you are currently using?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 346938,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-22T20:26:37.010000",
          "content": "<p>Hi @Marcel, it sounds very good that augmentation matters.\nAnd it is great that you use only two 5-fold models, sounds better solution than me.\nI'm using many, many, and many models :). Current score is composed of 3 models of all 2-folds, and they have all different CNN, different handling parameter of input.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352766,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-05T04:13:27.483000",
          "content": "<p>Hi @Marcel,  you have a good LB score with a simple model. I have the problem that when training with melspectrogram and CNN, overfitting can not be solved, the accuracy for training and validation datasets could be 0.9817 vs 0.7921 after about 100 epochs.\nI have used dropout, regulization, but overfiting is still there.\nI wonder if you have any experience similar?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352772,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-05T04:34:21.660000",
          "content": "<p>Hi Sailor, let me share regarding your problem. Strong augmentation like mixup will solve it, I'm confident about it. My acc always converges to like train:0.85 and valid:0.90 while training.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352787,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-05T05:52:44.167000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> Wow, thanks for your answer. I'll follow your mixup augmentation ~\nAs you wrote in another topic, trim silence, split long samples and mixup can all be effective, right?</p>\n\n<p>In addition, if you have tried to manually verify the train data label? Does that help a lot?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352808,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-05T06:42:50.480000",
          "content": "<p>Hi Sailor, mixup works fine and you can try other augmentation techniques too. Trimming and splitting both work fine.</p>\n\n<p>And regarding manually verifying train label, you can find discussion here:\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/58052</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352835,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-05T08:17:09.893000",
          "content": "<p>Got it, thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 353220,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-06T08:01:39.977000",
          "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, mixup seems working now. I use 10-folder cross-validation, the accuracy for two folders can be as bellow:</p>\n\n<pre><code>loss: 0.9709 - acc: 0.8900 - val_loss: 0.5129 - val_acc: 0.9137\nloss: 0.7581 - acc: 0.9176 - val_loss: 1.5712 - val_acc: 0.8642\n</code></pre>\n\n<p>I train only 150epochs with early stopping, seems not converge enough now. The generalization of model seems improved a lot.</p>\n\n<p>But map@3 could not be applied after each epoch, I'll see the LB score after the training finished.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 353555,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-07T03:27:14.157000",
          "content": "<p>Hi @Sailor, that sounds good. Increasing augmentation should make training converge longer, so it seems to be good. I hope you have better results.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354663,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-10T02:24:08.093000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> ,  I have tried some different parameters to train. The 10folder single model's predictions seems not improved too much. The LB scores for the melspectrogram + cnn model is 0.907 now, far from your 0.94 for single model. I wonder what features and model your best single model use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354685,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-10T04:27:33.750000",
          "content": "<p>Hi @Sailor</p>\n\n<ul>\n<li>Thanks to @Cocoxili, preprocessing follows almost the same parameter as kernel <a href=\"https://www.kaggle.com/aadan2017/log-mel-features-shown-by-category\">'Log-mel features shown by category.'</a></li>\n<li><p>Model is SE-ResNet-50 with small modification to make it smaller/thinner. I'm using following implementation:</p>\n\n<p><a href=\"https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se_resnet.py\">https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se_resnet.py</a>\n<a href=\"https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se.py\">https://raw.githubusercontent.com/titu1994/keras-squeeze-excite-network/master/se.py</a></p></li>\n</ul>\n\n<p>But please note that my other light model based on AlexNet which is much smaller than SE-ResNet-50 has LB score 0.933, I think SE-ResNet is too much.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354757,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-10T07:36:47.230000",
          "content": "<p>Thanks, <a href=\"/daisukelab\">@daisukelab</a>. I also use mel-spectrogram as the input feature, it got better score than mfcc or stft. When you say light model based on AlexNet, do you mean this <a href=\"https://www.kaggle.com/daisukelab/simple-cnn-approach\">https://www.kaggle.com/daisukelab/simple-cnn-approach</a> ? I am just tuning the parameter for my 128*176 input, but still converge slow.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354769,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-10T08:03:29.900000",
          "content": "<p>Hi @Sailor,</p>\n\n<p>Yes that's the exact model. :)</p>\n\n<p>And regarding the parameter to convert to mel-spectrogram, I guess this leaves improvement for you. If you follow exact the same n_fft, n_mels and other parameter as same as <a href=\"https://www.kaggle.com/aadan2017/log-mel-features-shown-by-category\"> 'Log-mel features shown by category'</a>, you might see something different. In my case, validation accuracy increased about +0.1% higher than before, though it doesn't directly link to the LB score. And the length of the input might be short, is it 1-2s? I'm using 4 or more seconds with recent attempts.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354782,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-10T08:38:12.377000",
          "content": "<p>Yeah, I will try when the server is available :(. How many epochs your model generally early stops? To get a 10 folds cross validated model, it often spends at least one night or more.\nThis input is 2s. My best score is got when the length is 3s. You got better results using 4s or more? That's may be much slower... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 354934,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-10T14:09:56.177000",
          "content": "<p>Hi, regarding epochs, I don't use early stopping but best val_acc checkpoint. It really depends but usually no later than 300 epochs. It really takes long to train, about 5-10hrs. Running PC all day...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 358000,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-17T10:22:51.680000",
          "content": "<p>Hi <a href=\"/daisukelab\">@daisukelab</a>, do you mind telling the performance of crnn? I am trying to use crnn structure based on a dcase2017 paper. The performance is much worse that cnn, which is confused.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 358055,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-17T12:50:19.547000",
          "content": "<p>Hi Sailor, I haven't tried CRNN so much, actually submitted only once. It was 0.913. BTW, recently I'm not looping short samples, and use short ones as is with zero paddings; it was found for me to be better not to repeat... I'm sorry if you are doing that by following my comment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359325,
          "author_name": "Daniel Havir",
          "author_url": "",
          "post_date": "2018-07-19T21:48:59.017000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>: Not repeating worked better for you? So you just pad with zeros if the slices are shorter?</p>\n\n<p>Also, recurrence can work but I found it more efficient to slice the samples (with overlap) so that they all have the same size. Training CRNN on the whole samples complicates a lot of stuff (e.g. augmentation) and takes way too much time before you find the right hyperparameters (GRU vs. LSTM, # layers, # units, ...).</p>\n\n<p>Do you still use SENet 50? I tried it too, but it had unnecessarily many parameters.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 359871,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-21T00:34:06.670000",
          "content": "<p>Hi @Daniel,</p>\n\n<p>Yes currently I'm not using repeat of short samples. It worked well with other parameters when I started using it once, but now I cannot evaluate if it works well or not.\nI observed that it has side effect that the repeated sample is confused with Hi-hat or other drum-like labels, though it seems not to be frequent.</p>\n\n<p>And regarding models, I agree that it seems to be  difficult to make Recurrent NN work in this task.\nRegarding SENets, I use SEResNet-50 (actually less than 50, modified a little) that has best balance between good modeling and overfitting.</p>\n\n<p>To be honest, I've tried almost all the models listed in <a href=\"https://keras.io/applications/\">Keras applications</a>, as if ordering all the beers on list at a bar. 😅\nMany are too much and overfitted (or would needed more augmentation), here's funny example with Inception v3.</p>\n\n<pre><code>Epoch 8/300\n - 221s - loss: 1.7456 - acc: 0.7004 - val_loss: 2.6311 - val_acc: 0.2956\n</code></pre>\n\n<p>VGG16 worked fine, but I already have SEResNet working better.\nOf course tried ImageNet pre-trained weights, though didn't show big difference. (I'm not using it anymore)</p>\n\n<p>Then for me, any model is ok as long as it performs fine...\nTweaking how to expose data to models; preprocessing parameters and training data selection...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 346446,
      "author_name": "r06921058_>.O",
      "author_url": "",
      "post_date": "2018-06-21T17:54:51.503000",
      "content": "<p>Hi  <a href=\"/daisukelab\">@daisukelab</a> . Thanks you for sharing the nice work.\nI'm trying CRNN to train the model. However, I encounter several problem. I have trouble to get pass 0.85. May I ask some question about this?</p>\n\n<ol>\n<li><p>The CRNN only accept unified length, I've tried setting length as 128, 152, 512.I'm interesting in how you decide the length of Input?</p></li>\n<li><p>I pad the data with zeros after convert it into melspectrogram. Is it better to use other padding method? I've tried other method ,but I'm not sure the accuracy drop because of the padding method or the model. Is it different to cut the data before convert or after convert?</p></li>\n<li><p>While training the CRNN model, my model's Loss dropped rapidly after hundreds of epochs (train acci and val accu is both about 0.7 when this happened). And then the loss stuck in about 1*e-7.  I've tried changing the learning rate  and activate function. However, this problem still happened. Have you ever encounter the same problem while training RNN?</p></li>\n</ol>\n\n<p>My model using keras2.0.8, Tensorflow1.4.0. My CRNN model is  replace the full connection layer of CNN to bidirectional LSTM and add another LSTM for output.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 346937,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-22T20:17:26.100000",
          "content": "<p>Hi @r06921058_&gt;.O,</p>\n\n<ol>\n<li><p>Regarding the length to cut from the original sample, it is one of important parameters as far as I have tried. And cutting from <em>where</em>, as well as <em>disposing or using the other half of cut wave</em> are also important. I’m ensembling models that have different parameters of input sample handling, there seems to be no single answer. Some sample starts from the beginning, some others have sound events in the middle or at the end.</p></li>\n<li><p>Regarding padding zeros, I’ve got better result by <em>repeating</em> the wave until it comes to your desired length rather than padding something. As far as I was checking samples, there seems to be gap of amount of information between longer samples and shorter samples. This seems to be the same thing with imbalance of number of samples among classes, we can improve performance by balancing it. My guess is padding doesn’t solve imbalance of information among different length of samples, but repeating should basically resolve it.</p></li>\n<li><p>I have tried CNN only, so I don’t have exact answer unfortunately.\nBut your model seems to be complicated, then if I would debug it I would confirm CNN is working fine by using CNN part only first. Then checking for the details of entire model.</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346977,
          "author_name": "r06921058_>.O",
          "author_url": "",
          "post_date": "2018-06-22T23:35:36.360000",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a>,Thanks for answering these question! I'll try to use repeat to pad, and try other split method. </p>\n\n<p>About 2. If I understand correctly, you repeat the data to the same length before split them into pieces to handle the problem of imbalance? Or you used other method to deal with imbalance problem? Thanks for reply.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 346992,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-06-23T00:20:29.610000",
          "content": "<p>Hi, I set a length of one input wave, then apply preprocessing of raw wave as follows:</p>\n\n<ol>\n<li>Trim silence.</li>\n<li>Repeat if short.</li>\n<li>Split if long enough with overlaps.</li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 347722,
          "author_name": "r06921058_>.O",
          "author_url": "",
          "post_date": "2018-06-25T07:16:57.387000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352764,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-05T04:07:10.767000",
          "content": "<p>Hi daisuke,\nThanks for your share. I wonder when you do silence trim, do you mean trim the silence part in the begining and end of audio? how much is the dB value to judge silence. I just cut the -30dB part, with the same network, the LB score decrease...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352769,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-05T04:30:25.240000",
          "content": "<p>Hi Sailor, my code is simple:</p>\n\n<pre><code>y, _ = librosa.effects.trim(y)\n</code></pre>\n\n<p>It trims both ends.</p>\n\n<p>BTW, r06921058_&gt;.O and all,\nI could be wrong about repeating audio.\nUsing short audio filled with zero padding is now showing good score (though it doesn't exceed current best).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352784,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2018-07-05T05:43:43.517000",
          "content": "<p>Thanks, <a href=\"/daisukelab\">@daisukelab</a>.\nSo, trim can help to improve the score?\nI also use this function, but I set the trim_level as 30dB, to remove silence enough. Maybe I cut too lot to get worse score.</p>\n\n<pre><code>data, _ = librosa.effects.trim(data, top_db=trim_level)\n</code></pre>\n\n<p>It seems the default top_db you use is 60dB.</p>\n\n<pre><code>librosa.effects.trim(y, top_db=60, ref=&lt;function amax&gt;, frame_length=2048, hop_length=512)\n</code></pre>\n\n<p><a href=\"http://librosa.github.io/librosa/generated/librosa.effects.trim.html#librosa.effects.trim\">http://librosa.github.io/librosa/generated/librosa.effects.trim.html#librosa.effects.trim</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 352806,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-05T06:37:50.590000",
          "content": "<p>Hi Sailor, yes I'm using default setting without problem. I haven't compare results for this single parameter - but I think trimming should work better as long as it is used properly, and the default parameter seems to be fine.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 361387,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-24T10:59:37.887000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 361390,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2018-07-24T11:14:40.597000",
          "content": "<p>Hi rrr, I'm sorry I cannot help with 1-D array solution for that extent, didn't try 1-D...\nIt should work but dunno why, ... did you try 2-D? It works fine for me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "330305": "# Update 1 August\n\nSubmitted a kernel [here](https://www.kaggle.com/daisukelab/freesound-dataset-kaggle-2018-solution).\nThis was supposed to be my best model, but private score was not unfortunately. My best solution has two more models (4 models in total), but basically the same.\n\n# Update 24 July\n\nLet me update this post because my final solution has been changed a lot. I simplified after experimenting many different approaches and analyzing their results (over 150), resulting to have following notions:\n\n- Test set is composed only with _manually verified_ samples, unlike train set has weak labels. This could show that it is important to select which training examples to feed to the model.\n- Public LB score is based on 19% of manually verified test samples, then the 19% is plausible to be slightly or more different with regard to data quantity and quality distribution. We might be better to trust our own validation performance rather than current LB score.\n- Audio preprocessing basically should leave information as much as possible. No need to down sampling to lower fs like 24kHz, and feature resolution (n_mels) can be higher like 128 rather than traditional 40.\n- Model doesn’t matter as long as it has enough but not too much capacity.\n\nNow I’m using [Simple CNN approach](https://www.kaggle.com/daisukelab/simple-cnn-approach) only with SEResNet-50 based model.\n\nBTW I will summarize my solution in a kernel and post right after competition closed.\n\n(And might not have enough time for DCASE paper submission, I'm sorry to organizers...)\n\n# Original post, while having LB score 0.957 or 0.961 (some could be outdated)\n\nReaching to score just under 0.9 would be usual kaggle work; preprocess dataset, build models, train them and ensemble predictions.\nThen bringing up models for better performance takes more efforts.\nAs far as I’m working on, building right model was important, but deep diving into audio domain and dataset matters more in this competition so far.\n\nBefore I join this competition, I did my best in [TUT Acoustic Scene Classification](https://www.kaggle.com/c/acoustic-scene-2018) competition Feb this year, and could earn good lessons.\n\nThen firstly I will share summary of lessons learned in the TUT competition. I’m using almost the same CNN models for this freesound audio tagging competition also. And the augmentations (mixup &amp; cutout/random erasing) which was important also work fine. They are still strong.\n\n## Kernels shared in [TUT Acoustic Scene Classification](https://www.kaggle.com/c/acoustic-scene-2018) competition\n\nThese are what I’ve done for getting score 0.9+.\n\n- [Simple CNN approach](https://www.kaggle.com/daisukelab/simple-cnn-approach) - this is basic kernel.\n- [Sound event based approach](https://www.kaggle.com/daisukelab/sound-event-based-approach) - this is an approach to emphasize temporal change (calling it as sound event here).\n- [Time-wise mean approach](https://www.kaggle.com/daisukelab/time-wise-mean-approach) - this is inverse approach that completely ignore temporal change.\n- Regarding augmentation, [mixup &amp; cutout/random erasing](https://www.kaggle.com/daisukelab/mixup-cutout-or-random-erasing-to-augment) - Strong augmentation you cannot miss.\n\n## New findings in this competition\n\nFollowings are new findings while working on in this competition, kind of going back to the basics.\nThese have pushed my scores toward 0.95.\n\n- When you preprocess, don't stick to 16kHz. The dataset have sound wave with 44.1kHz frequency, and many sounds are not human voice. Then converting to the 16kHz sampling rate which traditional speech recognition system uses is apparently not-adequate. Use higher frequency. For me, it's 24kHz or some more. It depends on model or other preprocessing methods.\n- Different audio length might give us different perspective, though I'm not confident yet, but so far so good. Training/ testing audio sample has different length. Models that accepts fixed audio length like my CNNs require preprocessing to trim, repeat or fill blank to samples to make their length unified. A model which handles 2s input audio seems to be different from a model that takes 5s input audio. 2s model focuses on short period. On the other hands, 5s model looks over from higher view. I haven’t try but using CRNN instead of CNN could be the one-size-fits-all.\n\n## Other things for improvement\n\nThis is also basic thing but I needed for improvement.\n\n- Fix class imbalance. Some classes have less samples that causes lower performance. I simply oversample fewer number of samples up to maximum number of samples until all the class have the same number of samples.\n- Use entire audio wave. Some of implementation seems to use part of audio, but this also causes lower performance. I observed that some long audio has only half or shorter part that matches the class, and other part is nothing to do with the labeled class. I just simply use the entire part as training sample by splitting audio into multiple fixed unit length and put them all back into training set. Randomly choosing part of audio while training would be the same thing, but preprocessing first is faster to train in my environment.\n\nLastly I should share that I'm not using pseudo labeling so far. So current score is purely done by preprocessing &amp; model engineering.",
    "336992": "Thanks Daisuke, some very interesting and useful information. I'm struggling to get past 0.92, you seem so far away at the top! I was just wondering, when you say use the entire audio wave. Do you, for example, split a 15 second training clip into equal portions of 5 seconds, so you have two extra pieces of training data? Have you also done any preprocessing to remove silent parts of the audio?\n\nAlso, for anyone who uses Keras, the class imbalance can be fixed by creating a dict of class weights and providing it as an argument to the fit function. It definitely helped boost my score!\n\n&gt; class_weight: Optional dictionary mapping class indices (integers) to a weight (float) value, used for weighting the loss function (during training only). This can be useful to tell the model to \"pay more attention\" to samples from an under-represented class.",
    "365720": "Hi @daisukelab, thank you very much for your discussion and sharing. I started this competition with insights that you have shared. Those mixup, erase/cutout augmentation work pretty well! ",
    "345698": "Just a question. How long do your models take to complete? Also, what configuration you are on? Like RAM, GPU etc.?",
    "340706": "Hi, thanks for your nice work! I have a question that my validation acc can not present the real level of the model. As the validation acc increases, it decreases in lb unexpectedly (I used 10-fold verification). Someone said the labels without verification may affect the verification set results. So I want to know how do you choose the verification set to ensure the result is true even in lb or whether you also met this problem? Thank you very much.",
    "337890": "Thank you @daisukelab, I've been trying to model the variable length data into one model, but my performance is always under 83%. It seems your experiments could be a good guidance.\n\nMay I ask a question about mix-up method that: After mix-up process the data in every batch will have more than one incomplete labels, will you use Binary-Cross-Entropy for every outputs as multi-label classification, or still use Cross-Entropy on all outputs?",
    "353215": "Hi @daisukelab. Thank you very much for sharing your work! As many participants of this challenge (including myself) have benefited from your willingness to share your insights, it would be nice to be able to at least honor your work in the acknowledgements section of our submission for this challenge. Is there any chance we can get your real name for this purpose? :) (In case you have previously published anything on this topic when participating at the TUT Acoustic Scene Classification competition, I would be equally thankful for a source.)",
    "346927": "Hey @daisukelab, I'd also like to thank you for sharing your insights. Some of your data augmentation methods helped to improve my score drastically. I think that the key ingredient is to apply all kinds of data augmentation. I got a decent score with only two 5-fold models. Would you mind telling us how many models you are currently using?",
    "346446": "Hi  @daisukelab . Thanks you for sharing the nice work.\nI'm trying CRNN to train the model. However, I encounter several problem. I have trouble to get pass 0.85. May I ask some question about this?\n\n1. The CRNN only accept unified length, I've tried setting length as 128, 152, 512.I'm interesting in how you decide the length of Input?\n\n2. I pad the data with zeros after convert it into melspectrogram. Is it better to use other padding method? I've tried other method ,but I'm not sure the accuracy drop because of the padding method or the model. Is it different to cut the data before convert or after convert?\n\n3. While training the CRNN model, my model's Loss dropped rapidly after hundreds of epochs (train acci and val accu is both about 0.7 when this happened). And then the loss stuck in about 1*e-7.  I've tried changing the learning rate  and activate function. However, this problem still happened. Have you ever encounter the same problem while training RNN?\n\nMy model using keras2.0.8, Tensorflow1.4.0. My CRNN model is  replace the full connection layer of CNN to bidirectional LSTM and add another LSTM for output.",
    "361387": ""
  }
}