{
  "id": 91881,
  "title": "What's your best single model LB?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/91881",
  "author_name": "wqk",
  "post_date": "2019-05-10T05:25:15.257000",
  "votes": 24,
  "comment_count": 73,
  "views": 0,
  "content": "<p>With only curated dataset and cnn, 20%hold out, my best single model LB is 0.64.</p>\n\n<p>Add mixup, LB to 0.655, with local lwlrap 0.76, almost the same with before.</p>\n\n<p>What's the key point to get better performance in this task?</p>\n\n<hr>\n\n<p>update 20190513</p>\n\n<p>With only cv (\"k-fold averaging\") added, LB to 0.7+, amazing improvement.</p>\n\n<p>The next step may be effective usage of noisy subset.</p>\n\n<hr>\n\n<p>update 20190610</p>\n\n<p>single model with inception V3, reach 0.691, 5 fold cv, 0.71+</p>\n\n<p>single fold kernel here: </p>\n\n<p><a href=\"https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\">https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673</a></p>",
  "messages": [
    {
      "id": 529521,
      "postDate": "2019-05-10T05:25:15.257Z",
      "content": "<p>With only curated dataset and cnn, 20%hold out, my best single model LB is 0.64.</p>\n\n<p>Add mixup, LB to 0.655, with local lwlrap 0.76, almost the same with before.</p>\n\n<p>What's the key point to get better performance in this task?</p>\n\n<hr>\n\n<p>update 20190513</p>\n\n<p>With only cv (\"k-fold averaging\") added, LB to 0.7+, amazing improvement.</p>\n\n<p>The next step may be effective usage of noisy subset.</p>\n\n<hr>\n\n<p>update 20190610</p>\n\n<p>single model with inception V3, reach 0.691, 5 fold cv, 0.71+</p>\n\n<p>single fold kernel here: </p>\n\n<p><a href=\"https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\">https://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673</a></p>",
      "rawMarkdown": "With only curated dataset and cnn, 20%hold out, my best single model LB is 0.64.\n\nAdd mixup, LB to 0.655, with local lwlrap 0.76, almost the same with before.\n\nWhat's the key point to get better performance in this task?\n\n-------------------------\nupdate 20190513\n\nWith only cv (\"k-fold averaging\") added, LB to 0.7+, amazing improvement.\n\nThe next step may be effective usage of noisy subset.\n\n------------------------\nupdate 20190610\n\nsingle model with inception V3, reach 0.691, 5 fold cv, 0.71+\n\nsingle fold kernel here: \n\nhttps://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\n",
      "votes": 23
    },
    {
      "id": 529707,
      "postDate": "2019-05-10T15:01:26Z",
      "content": "<p>My best single (and only) model is a simple CNN with 4 layers trained on the curated data only. It gives me 0.7 LB score.</p>\n\n<p>I wonder if anyone uses something different from CNNs for achieving 0.65+ scores?</p>",
      "rawMarkdown": "My best single (and only) model is a simple CNN with 4 layers trained on the curated data only. It gives me 0.7 LB score.\n\nI wonder if anyone uses something different from CNNs for achieving 0.65+ scores?",
      "votes": 8,
      "replies": [
        {
          "id": 529728,
          "postDate": "2019-05-10T15:40:49.097Z",
          "content": "<p>Thanks for the info. I'm using 30+ convolution layers (LB 0.685 as previously stated). If four layers can already achieve 0.7 LB score, then my 30+ is definitely too many... </p>",
          "rawMarkdown": "Thanks for the info. I'm using 30+ convolution layers (LB 0.685 as previously stated). If four layers can already achieve 0.7 LB score, then my 30+ is definitely too many... ",
          "votes": 1
        },
        {
          "id": 529732,
          "postDate": "2019-05-10T15:47:33.333Z",
          "content": "<p>0.7 LB score with 4 layers! I just tried lots of different cnn structures.</p>",
          "rawMarkdown": "0.7 LB score with 4 layers! I just tried lots of different cnn structures.",
          "votes": 1
        },
        {
          "id": 529734,
          "postDate": "2019-05-10T15:49:44.797Z",
          "content": "<p>Well, I haven't played with the architecture yet, so a deep network might be better than a shallow one eventually. The difference in score may also come from different sources such as data preprocessing, training schedule, etc.</p>",
          "rawMarkdown": "Well, I haven't played with the architecture yet, so a deep network might be better than a shallow one eventually. The difference in score may also come from different sources such as data preprocessing, training schedule, etc.\n",
          "votes": 4
        },
        {
          "id": 529753,
          "postDate": "2019-05-10T16:44:34.717Z",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> from my experience, the architecture rarely matters much until there are other ways to improve, and there are a lot of ways in this particular competition</p>",
          "rawMarkdown": "@sailorwei from my experience, the architecture rarely matters much until there are other ways to improve, and there are a lot of ways in this particular competition",
          "votes": 3
        },
        {
          "id": 529757,
          "postDate": "2019-05-10T16:55:17.130Z",
          "content": "<p>Agreed I think the architecture matters far less in this case.</p>",
          "rawMarkdown": "Agreed I think the architecture matters far less in this case.",
          "votes": 3
        },
        {
          "id": 529771,
          "postDate": "2019-05-10T17:43:56.223Z",
          "content": "<p>Interesting. I have not really explored shallower models yet. The take home message here for me is that I could've saved a lot of training time using a shallower model and sped up the data pipeline iterations.</p>",
          "rawMarkdown": "Interesting. I have not really explored shallower models yet. The take home message here for me is that I could've saved a lot of training time using a shallower model and sped up the data pipeline iterations.",
          "votes": 2
        },
        {
          "id": 529777,
          "postDate": "2019-05-10T18:21:17.900Z",
          "content": "<p>Thanks for the information <a href=\"/ddanevskyi\">@ddanevskyi</a> <a href=\"/jamesrequa\">@jamesrequa</a> . Cannot agree more <a href=\"/ceshine\">@ceshine</a> .</p>",
          "rawMarkdown": "Thanks for the information @ddanevskyi @jamesrequa . Cannot agree more @ceshine .",
          "votes": 2
        },
        {
          "id": 547036,
          "postDate": "2019-06-07T07:19:07.810Z",
          "content": "<p>Agree, I also used a 4 layer network, got 6.5, and did not improve too much with RESNET50.</p>",
          "rawMarkdown": "Agree, I also used a 4 layer network, got 6.5, and did not improve too much with RESNET50."
        }
      ]
    },
    {
      "id": 532617,
      "postDate": "2019-05-17T11:37:05.470Z",
      "content": "<p>To get back to the original topic ;)\nMy single best model, LB: 0.720, CV: 0.86</p>",
      "rawMarkdown": "To get back to the original topic ;)\nMy single best model, LB: 0.720, CV: 0.86",
      "votes": 5,
      "replies": [
        {
          "id": 532687,
          "postDate": "2019-05-17T14:33:18.177Z",
          "content": "<p>Thanks for sharing this info! Did you only use curated data?</p>\n\n<p>From the responses to this thread, there are multiple reports of similar local CV scores with quite different LB scores. I wonder if we are in serious danger of overfitting the leaderboard (by switching between models with similar CV scores but different LB scores).</p>\n\n<p>EDIT: grammar. </p>",
          "rawMarkdown": "Thanks for sharing this info! Did you only use curated data?\n\nFrom the responses to this thread, there are multiple reports of similar local CV scores with quite different LB scores. I wonder if we are in serious danger of overfitting the leaderboard (by switching between models with similar CV scores but different LB scores).\n\nEDIT: grammar. ",
          "votes": 1
        },
        {
          "id": 532692,
          "postDate": "2019-05-17T14:45:37.203Z",
          "content": "<p>Yes, my CV score is computed using only the curated dataset.</p>",
          "rawMarkdown": "Yes, my CV score is computed using only the curated dataset."
        },
        {
          "id": 532730,
          "postDate": "2019-05-17T15:51:38.940Z",
          "content": "<p>That's quite a feat! Thanks for the info.</p>",
          "rawMarkdown": "That's quite a feat! Thanks for the info."
        }
      ]
    },
    {
      "id": 529717,
      "postDate": "2019-05-10T15:25:21.793Z",
      "content": "<p>For the second point, I'm firmly convinced one needs to find a way to combat label noise in the noisy train subset.</p>",
      "rawMarkdown": "For the second point, I'm firmly convinced one needs to find a way to combat label noise in the noisy train subset.",
      "votes": 5
    },
    {
      "id": 529716,
      "postDate": "2019-05-10T15:18:50.477Z",
      "content": "<p>0.659 curated only, simple cnn, 20%hold out, no augmentation. <a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\">this kernel</a> plus some tweeks.\nwonder what will bring k-fold</p>",
      "rawMarkdown": "0.659 curated only, simple cnn, 20%hold out, no augmentation. [this kernel](https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai) plus some tweeks.\nwonder what will bring k-fold",
      "votes": 4,
      "replies": [
        {
          "id": 529733,
          "postDate": "2019-05-10T15:49:07.133Z",
          "content": "<p>mine is so close to yours. I'm trying to bring in k-fold averaging now.</p>",
          "rawMarkdown": "mine is so close to yours. I'm trying to bring in k-fold averaging now.",
          "votes": 1
        },
        {
          "id": 530763,
          "postDate": "2019-05-13T16:10:03.220Z",
          "content": "<p>5 kfold .684 local lwlrap .84-.86</p>",
          "rawMarkdown": "5 kfold .684 local lwlrap .84-.86",
          "votes": 2
        },
        {
          "id": 531562,
          "postDate": "2019-05-15T06:13:46.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/sailorwei\">@wqk</a> are you doing 5 folds or something else.  have you changed anything else but added Kfold (just yes or no - do secret sauce needed). With 5kfold i gain 0.020-.025 and your case is very impressive.</p>",
          "rawMarkdown": "[@wqk](https://www.kaggle.com/sailorwei) are you doing 5 folds or something else.  have you changed anything else but added Kfold (just yes or no - do secret sauce needed). With 5kfold i gain 0.020-.025 and your case is very impressive."
        },
        {
          "id": 531672,
          "postDate": "2019-05-15T10:17:01.670Z",
          "content": "<p>I did not change other things except adding Kfold and increasing epochs, with local lwlraps are around 0.76.</p>",
          "rawMarkdown": "I did not change other things except adding Kfold and increasing epochs, with local lwlraps are around 0.76."
        },
        {
          "id": 532371,
          "postDate": "2019-05-16T18:49:58.077Z",
          "content": "<p>I feel you man... We have the same local lwlrap range and the same boost after 5-fold averaging.</p>",
          "rawMarkdown": "I feel you man... We have the same local lwlrap range and the same boost after 5-fold averaging.",
          "votes": 1
        },
        {
          "id": 533781,
          "postDate": "2019-05-20T02:54:33.080Z",
          "content": "<p>Hi <a href=\"/jihangz\">@jihangz</a> , how is your local lwlrap now? I cannot increase it for several days.</p>",
          "rawMarkdown": "Hi @jihangz , how is your local lwlrap now? I cannot increase it for several days."
        }
      ]
    },
    {
      "id": 529553,
      "postDate": "2019-05-10T07:02:52.113Z",
      "content": "<p>0.689 here with curated, single basic model, k-fold, no augmentation. Local lwlrap 0.860\nI imagine the 'key point' is to use the noisy dataset in some way but haven't done much investigation there yet.</p>",
      "rawMarkdown": "0.689 here with curated, single basic model, k-fold, no augmentation. Local lwlrap 0.860\nI imagine the 'key point' is to use the noisy dataset in some way but haven't done much investigation there yet.",
      "votes": 4,
      "replies": [
        {
          "id": 529566,
          "postDate": "2019-05-10T08:44:19.140Z",
          "content": "<p>Very impressive! </p>\n\n<p>My augmented one achieved 0.685 at best, with local lwlrap of 0.83.</p>",
          "rawMarkdown": "Very impressive! \n\nMy augmented one achieved 0.685 at best, with local lwlrap of 0.83.",
          "votes": 3
        },
        {
          "id": 529709,
          "postDate": "2019-05-10T15:04:05.847Z",
          "content": "<p>That's impressive, your local lwlrap seems to be very high, my best model gives me only 0.83 local lwlrap.</p>",
          "rawMarkdown": "That's impressive, your local lwlrap seems to be very high, my best model gives me only 0.83 local lwlrap.",
          "votes": 2
        },
        {
          "id": 529730,
          "postDate": "2019-05-10T15:43:48.483Z",
          "content": "<p>Really a high local lwlrap!</p>",
          "rawMarkdown": "Really a high local lwlrap!",
          "votes": 1
        },
        {
          "id": 529756,
          "postDate": "2019-05-10T16:54:04.197Z",
          "content": "<p>My best so far is 0.683 with a custom shallow single model using only curated dataset but with this same model my best local lwlrap is much higher over 0.87. Perhaps I need to more carefully select the validation set as there is a large discrepancy with LB score. I am also still looking into how best to utilize the noisy dataset as this is the real purpose of the competition (i.e. SSL). </p>",
          "rawMarkdown": "My best so far is 0.683 with a custom shallow single model using only curated dataset but with this same model my best local lwlrap is much higher over 0.87. Perhaps I need to more carefully select the validation set as there is a large discrepancy with LB score. I am also still looking into how best to utilize the noisy dataset as this is the real purpose of the competition (i.e. SSL). ",
          "votes": 2
        },
        {
          "id": 529808,
          "postDate": "2019-05-10T19:55:32.177Z",
          "content": "<p>Curated only, shallow CNN. I didn't calculate oof CV, but my 5-fold valid lwlraps are between 0.835 and 0.865. Single model gives me LB 0.676. </p>\n\n<p>Based on everyone's info here, I expect a huge shakeup in second stage, like Quora and PetFinder competition with the same inference format (2 stage with different test set).</p>",
          "rawMarkdown": "Curated only, shallow CNN. I didn't calculate oof CV, but my 5-fold valid lwlraps are between 0.835 and 0.865. Single model gives me LB 0.676. \n\nBased on everyone's info here, I expect a huge shakeup in second stage, like Quora and PetFinder competition with the same inference format (2 stage with different test set).",
          "votes": 1
        },
        {
          "id": 529823,
          "postDate": "2019-05-10T21:10:48.990Z",
          "content": "<p><a href=\"/jihangz\">@jihangz</a> may I ask why do you expect a shakeup here? LB metric seems to be pretty stable. There is a gap between a local score and LB score, but as long as my local score improves, LB metric improves too.</p>",
          "rawMarkdown": "@jihangz may I ask why do you expect a shakeup here? LB metric seems to be pretty stable. There is a gap between a local score and LB score, but as long as my local score improves, LB metric improves too.",
          "votes": 1
        },
        {
          "id": 529827,
          "postDate": "2019-05-10T21:37:55.340Z",
          "content": "<p>First, test set is small, and the test set in 2nd stage is 3x the current one. \nSecond, my local score is at the same level as most of yours, but my LB is lower. </p>\n\n<p>Time to trust CV?</p>",
          "rawMarkdown": "First, test set is small, and the test set in 2nd stage is 3x the current one. \nSecond, my local score is at the same level as most of yours, but my LB is lower. \n\nTime to trust CV?"
        },
        {
          "id": 529851,
          "postDate": "2019-05-11T01:02:56.930Z",
          "content": "<p>As far as I have listened to the samples, many of them seems to have strong correlation (coming from similar sound source, music instrument or just applied different EQ/effect to a single source sound...), then I guess it would be difficult to build <em>sound</em> CV folds unless source information or the similar is available to separate train/valid adequately.</p>\n\n<p>I agree with <a href=\"/robga\">@robga</a>, most promising way to go could be to find how to utilize noisy set.</p>",
          "rawMarkdown": "As far as I have listened to the samples, many of them seems to have strong correlation (coming from similar sound source, music instrument or just applied different EQ/effect to a single source sound...), then I guess it would be difficult to build _sound_ CV folds unless source information or the similar is available to separate train/valid adequately.\n\nI agree with @robga, most promising way to go could be to find how to utilize noisy set.",
          "votes": 7
        }
      ]
    },
    {
      "id": 536781,
      "postDate": "2019-05-25T09:10:11.877Z",
      "content": "<p>LB 0.629 with curated only.\nLB 0.680 with both curated and noisy.</p>\n\n<p>I use two models for training. One is to be trained(from scratch to LB 0.680) and the other one(LB 0.629) is to provide soft labels to noisy data just before it's used for training. </p>",
      "rawMarkdown": "LB 0.629 with curated only.\nLB 0.680 with both curated and noisy.\n\nI use two models for training. One is to be trained(from scratch to LB 0.680) and the other one(LB 0.629) is to provide soft labels to noisy data just before it's used for training. ",
      "votes": 3,
      "replies": [
        {
          "id": 536829,
          "postDate": "2019-05-25T11:31:28.790Z",
          "content": "<p>So you use the predicted soft labels in addition to the provided labels of the noisy data? That’s interesting idea!</p>",
          "rawMarkdown": "So you use the predicted soft labels in addition to the provided labels of the noisy data? That’s interesting idea!",
          "votes": 1
        },
        {
          "id": 536950,
          "postDate": "2019-05-25T18:40:41.660Z",
          "content": "<p>Yes, it works well and could be better if I can improve curated only model in the first place. I will try mixup if I can find time.</p>",
          "rawMarkdown": "Yes, it works well and could be better if I can improve curated only model in the first place. I will try mixup if I can find time.",
          "votes": 1
        },
        {
          "id": 536960,
          "postDate": "2019-05-25T20:16:25.560Z",
          "content": "<p>I tried similar stuff but it made my score worse so far.</p>",
          "rawMarkdown": "I tried similar stuff but it made my score worse so far.",
          "votes": 1
        },
        {
          "id": 536972,
          "postDate": "2019-05-25T21:11:20.270Z",
          "content": "<p>I provide soft-labels(sigmoid outputs of the model trained on curated only) to cropped(128x160) noisy mel data every time it's feeded to the model to be trained. I used original labels of noisy data together with soft-labels(used np.maximum) but it could be better just ignoring original labels. </p>\n\n<p>I tried both\n- trained on noisy and then trained on curated in each epoch.\n- trained on noisy for 50-60 epochs and then trained on curated for 20-30 epochs.</p>\n\n<p>The latter was a little better so far.</p>",
          "rawMarkdown": "I provide soft-labels(sigmoid outputs of the model trained on curated only) to cropped(128x160) noisy mel data every time it's feeded to the model to be trained. I used original labels of noisy data together with soft-labels(used np.maximum) but it could be better just ignoring original labels. \n\nI tried both\n- trained on noisy and then trained on curated in each epoch.\n- trained on noisy for 50-60 epochs and then trained on curated for 20-30 epochs.\n\nThe latter was a little better so far.",
          "votes": 3
        },
        {
          "id": 536987,
          "postDate": "2019-05-25T23:02:42.433Z",
          "content": "<p>I tried same approach, but I dont get better result</p>",
          "rawMarkdown": "I tried same approach, but I dont get better result"
        },
        {
          "id": 537062,
          "postDate": "2019-05-26T05:32:13.310Z",
          "content": "<p>Thanks so much <a href=\"/appian\">@appian</a> for this valuable ideas!</p>",
          "rawMarkdown": "Thanks so much @appian for this valuable ideas!"
        }
      ]
    },
    {
      "id": 537853,
      "postDate": "2019-05-27T18:18:17.727Z",
      "content": "<p>Two single models (curated+augmented+noisy) with two different preprocessed datasets.\n&lt;10 minutes with GPU (preprocessing+inference)\nLB M1=0.685\nLB M2=0.680\nBlending    1.05xM1+0.95xM2= 0.696</p>",
      "rawMarkdown": "Two single models (curated+augmented+noisy) with two different preprocessed datasets.\n&lt;10 minutes with GPU (preprocessing+inference)\nLB M1=0.685\nLB M2=0.680\nBlending    1.05xM1+0.95xM2= 0.696",
      "votes": 1
    },
    {
      "id": 536373,
      "postDate": "2019-05-24T11:18:40.813Z",
      "content": "<p>still struggling for 0.67, do not know what's wrong ... sign ...</p>",
      "rawMarkdown": "still struggling for 0.67, do not know what's wrong ... sign ...",
      "votes": 1,
      "replies": [
        {
          "id": 546431,
          "postDate": "2019-06-06T15:30:27.120Z",
          "content": "<p>Looks like you figured it out according to your standing :) , nice work!</p>",
          "rawMarkdown": "Looks like you figured it out according to your standing :) , nice work!",
          "votes": 1
        }
      ]
    },
    {
      "id": 530601,
      "postDate": "2019-05-13T08:48:52.057Z",
      "content": "<p>Hi,\nWould you like to tell us how do you split the validation set? Random or Balanced(Each category has the same quantity)?\nthank you a lot!</p>",
      "rawMarkdown": "Hi,\nWould you like to tell us how do you split the validation set? Random or Balanced(Each category has the same quantity)?\nthank you a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 530919,
          "postDate": "2019-05-14T00:51:47.577Z",
          "content": "<p>KFold is used now, seems StraitifiedKFold cannot accept multi-label dataset. </p>\n\n<p><a href=\"https://www.kaggle.com/osciiart/multilabel-stratifiedkfold-by-randomized-algorithm\">https://www.kaggle.com/osciiart/multilabel-stratifiedkfold-by-randomized-algorithm</a></p>",
          "rawMarkdown": "KFold is used now, seems StraitifiedKFold cannot accept multi-label dataset. \n\nhttps://www.kaggle.com/osciiart/multilabel-stratifiedkfold-by-randomized-algorithm",
          "votes": 1
        },
        {
          "id": 531517,
          "postDate": "2019-05-15T04:07:43.183Z",
          "content": "<p>One way to do it is to treat multi-label as separate classes, but the multi-label cases in this dataset are so rare it's not feasible to do so. My hacky solution is to randomly select one of the label in those cases and make the dataset single-label to creat stratified k-folds. (You need to make sure you pick the same labels when evaluating.)</p>",
          "rawMarkdown": "One way to do it is to treat multi-label as separate classes, but the multi-label cases in this dataset are so rare it's not feasible to do so. My hacky solution is to randomly select one of the label in those cases and make the dataset single-label to creat stratified k-folds. (You need to make sure you pick the same labels when evaluating.)"
        },
        {
          "id": 534924,
          "postDate": "2019-05-22T04:26:19.657Z",
          "content": "<p>@wqk Here's one approach for multilabel stratification: <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a> that seems to work well for me.</p>",
          "rawMarkdown": "@wqk Here's one approach for multilabel stratification: https://github.com/trent-b/iterative-stratification that seems to work well for me.",
          "votes": 1
        }
      ]
    },
    {
      "id": 530976,
      "postDate": "2019-05-14T04:38:40.353Z",
      "content": "<p>Regarding to your 20190503 update <a href=\"/sailorwei\">@sailorwei</a> , are your saying by adding cv, your LB score jumped from 0.655 to 0.7+? I'm not exactly sure what you mean by \"cv\". Can you please elaborate?</p>\n\n<p>EDIT: correction - update previous best score from 0.64 -&gt; 0.655 (the one from mixup).</p>",
      "rawMarkdown": "Regarding to your 20190503 update @sailorwei , are your saying by adding cv, your LB score jumped from 0.655 to 0.7+? I'm not exactly sure what you mean by \"cv\". Can you please elaborate?\n\nEDIT: correction - update previous best score from 0.64 -&gt; 0.655 (the one from mixup).",
      "votes": 2,
      "replies": [
        {
          "id": 530987,
          "postDate": "2019-05-14T04:59:20.667Z",
          "content": "<p>I think it means training N classifiers on N folds of training set and taking the average of N predictions on test set.</p>",
          "rawMarkdown": "I think it means training N classifiers on N folds of training set and taking the average of N predictions on test set.",
          "votes": 2
        },
        {
          "id": 530996,
          "postDate": "2019-05-14T05:12:28.983Z",
          "content": "<p>It's as <a href=\"/jihangz\">@jihangz</a>  said. \"k-fold averaging\" as you mentioned doesnot mean that?</p>",
          "rawMarkdown": "It's as @jihangz  said. \"k-fold averaging\" as you mentioned doesnot mean that?",
          "votes": 1
        },
        {
          "id": 531005,
          "postDate": "2019-05-14T05:35:28.760Z",
          "content": "<p>Yes that what I meant by \"k-fold averaging\". That was a pretty large jump in LB score by just doing averaging, so I just wanted to make sure that's what you meant.</p>",
          "rawMarkdown": "Yes that what I meant by \"k-fold averaging\". That was a pretty large jump in LB score by just doing averaging, so I just wanted to make sure that's what you meant.",
          "votes": 1
        },
        {
          "id": 531470,
          "postDate": "2019-05-15T01:25:18.323Z",
          "content": "<p>Yes, thanks for your previous information. So \"k-fold averaging\" should be stronger than training only one model with all train data? </p>",
          "rawMarkdown": "Yes, thanks for your previous information. So \"k-fold averaging\" should be stronger than training only one model with all train data? ",
          "votes": 1
        },
        {
          "id": 531516,
          "postDate": "2019-05-15T04:00:10.357Z",
          "content": "<p>Yes it usually gives more stable and slightly better results. But it's unusual to have such a big difference in my experience. If you don't mind, can you share your local CV score?</p>",
          "rawMarkdown": "Yes it usually gives more stable and slightly better results. But it's unusual to have such a big difference in my experience. If you don't mind, can you share your local CV score?"
        },
        {
          "id": 532287,
          "postDate": "2019-05-16T15:17:26.980Z",
          "content": "<p>Sure, my local lwlraps are all around 0.76 now. I cannot figure out the reason why it's lower than those in this discussion.</p>",
          "rawMarkdown": "Sure, my local lwlraps are all around 0.76 now. I cannot figure out the reason why it's lower than those in this discussion.",
          "votes": 1
        },
        {
          "id": 532538,
          "postDate": "2019-05-17T07:40:05.103Z",
          "content": "<p>Thanks for sharing. Mine is ~ .85 now (that is for the entire clip, not for a single chunk).  I'm starting to hit the problem of local CV improvements not translating into LB improvements. </p>\n\n<p>The train/test gap of this dataset is a bit bizarre, and might be crucial for getting a better private LB results as some has mentioned (this is just only my own speculation, though). </p>",
          "rawMarkdown": "Thanks for sharing. Mine is ~ .85 now (that is for the entire clip, not for a single chunk).  I'm starting to hit the problem of local CV improvements not translating into LB improvements. \n\nThe train/test gap of this dataset is a bit bizarre, and might be crucial for getting a better private LB results as some has mentioned (this is just only my own speculation, though). "
        },
        {
          "id": 532591,
          "postDate": "2019-05-17T10:41:04.567Z",
          "content": "<p>So high. Maybe overfitting, How many conv layers in your net?</p>",
          "rawMarkdown": "So high. Maybe overfitting, How many conv layers in your net?"
        },
        {
          "id": 532610,
          "postDate": "2019-05-17T11:19:27.387Z",
          "content": "<p><a href=\"/ceshine\">@ceshine</a> \nMy local CV is quite similar as you. I faced the same problem since local CV increases but LB says no. It is pretty weird that my best model has local CV lower than the second one (however the GAP still be high). I am quite sure that It is not randomness. </p>\n\n<p><a href=\"/sailorwei\">@sailorwei</a> \nYour local CV and GAP are so impressive. I just want to ask if you compute local CV in the correct way or not. You said you used mixup, the mixup is only used for training. If you use for validation, the score is low. IMO, it is only the reason that can explain your local CV and GAP. Otherwise, it may come from some <code>secret sauce</code>. </p>",
          "rawMarkdown": "@ceshine \nMy local CV is quite similar as you. I faced the same problem since local CV increases but LB says no. It is pretty weird that my best model has local CV lower than the second one (however the GAP still be high). I am quite sure that It is not randomness. \n\n@sailorwei \nYour local CV and GAP are so impressive. I just want to ask if you compute local CV in the correct way or not. You said you used mixup, the mixup is only used for training. If you use for validation, the score is low. IMO, it is only the reason that can explain your local CV and GAP. Otherwise, it may come from some `secret sauce`. ",
          "votes": 2
        },
        {
          "id": 532657,
          "postDate": "2019-05-17T13:31:41.750Z",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> I wouldn't call it overfitting since the local CV scores were from the validation set. Difference in train and test distribution is more likely the cause IMO. My model is smaller than resnet50.</p>\n\n<p><a href=\"/backaggle\">@backaggle</a> Yeah, most likely not randomness. There seems to be some models or data pipelines that generalize better to the test set. </p>\n\n<p>In case you haven't read it, here's <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89312#latest-523696\">the discussion thread on the topic of train/test discrepancy.</a> </p>\n\n<p>EDIT: reorganized the sentences to make it more clear.</p>",
          "rawMarkdown": "@sailorwei I wouldn't call it overfitting since the local CV scores were from the validation set. Difference in train and test distribution is more likely the cause IMO. My model is smaller than resnet50.\n\n@backaggle Yeah, most likely not randomness. There seems to be some models or data pipelines that generalize better to the test set. \n\nIn case you haven't read it, here's [the discussion thread on the topic of train/test discrepancy.](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89312#latest-523696) \n\nEDIT: reorganized the sentences to make it more clear.",
          "votes": 2
        },
        {
          "id": 533783,
          "postDate": "2019-05-20T03:05:39.477Z",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> mixup is only for the train data each batch. I'm still confused about the low local validation lwlrap. The local validation lwlrap function is from this kernel  <a href=\"https://www.kaggle.com/vinayaks/simple-changes-significant-improvement-in-lb\">https://www.kaggle.com/vinayaks/simple-changes-significant-improvement-in-lb</a> .</p>",
          "rawMarkdown": "@backaggle mixup is only for the train data each batch. I'm still confused about the low local validation lwlrap. The local validation lwlrap function is from this kernel  https://www.kaggle.com/vinayaks/simple-changes-significant-improvement-in-lb ."
        },
        {
          "id": 533795,
          "postDate": "2019-05-20T04:11:53.407Z",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> <br>\nHow do you compute lwlrap ?. In my experiments, there are two cases \n- (1) Split audio into several overlaped segments, then taking average overall \n- (2) Random segments (TTA)&lt; then taking average overall </p>\n\n<p>My results show that <code>(1)</code> gives low <code>lwlrap</code> while <code>(2)</code> gives higher. </p>",
          "rawMarkdown": "@sailorwei  \nHow do you compute lwlrap ?. In my experiments, there are two cases \n- (1) Split audio into several overlaped segments, then taking average overall \n- (2) Random segments (TTA)&lt; then taking average overall \n\nMy results show that `(1)` gives low `lwlrap` while `(2)` gives higher. ",
          "votes": 1
        },
        {
          "id": 533816,
          "postDate": "2019-05-20T05:13:56.247Z",
          "content": "<p>Hi <a href=\"/backaggle\">@backaggle</a>,</p>\n\n<p>I could not clearly understand the difference between (1) and (2), would you mind clarify them further?</p>",
          "rawMarkdown": "Hi @backaggle,\n\nI could not clearly understand the difference between (1) and (2), would you mind clarify them further?"
        },
        {
          "id": 533818,
          "postDate": "2019-05-20T05:26:09.300Z",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> </p>\n\n<ul>\n<li>For <code>(1)</code>. \nOne audio is splitted into several segment parts (let say: window size = 128, stride = 64). \nThe <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89382#latest-517614\">baseline</a> system follows this method with <code>stride = 0</code> (no overlap)</li>\n<li>For <code>(2)</code>.\nWe take randomly <code>n_parts</code>. Your amazing keras kernel already did it (called as TTA). </li>\n</ul>\n\n<p>The final result may be evaluated by taking average among segments. </p>",
          "rawMarkdown": "@ratthachat \n\n- For `(1)`. \nOne audio is splitted into several segment parts (let say: window size = 128, stride = 64). \nThe [baseline](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89382#latest-517614) system follows this method with `stride = 0` (no overlap)\n- For `(2)`.\nWe take randomly `n_parts`. Your amazing keras kernel already did it (called as TTA). \n\nThe final result may be evaluated by taking average among segments. ",
          "votes": 2
        },
        {
          "id": 533869,
          "postDate": "2019-05-20T07:03:15.223Z",
          "content": "<p>Thanks <a href=\"/backaggle\">@backaggle</a> for clarification and praising. My kernel performance is still far beyond yours and the others. Hope to learn more from all of you during this competition!</p>",
          "rawMarkdown": "Thanks @backaggle for clarification and praising. My kernel performance is still far beyond yours and the others. Hope to learn more from all of you during this competition!"
        },
        {
          "id": 535917,
          "postDate": "2019-05-23T16:15:45.117Z",
          "content": "<p><a href=\"/backaggle\">@backaggle</a>  thanks, I missed your message and just got it.</p>",
          "rawMarkdown": "@backaggle  thanks, I missed your message and just got it."
        },
        {
          "id": 538868,
          "postDate": "2019-05-29T08:03:11.440Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 530363,
      "postDate": "2019-05-12T16:21:20.777Z",
      "content": "<p>Hi, just curious if anyone here use Keras and be able to get score more than 0.67+</p>",
      "rawMarkdown": "Hi, just curious if anyone here use Keras and be able to get score more than 0.67+",
      "votes": 2,
      "replies": [
        {
          "id": 530388,
          "postDate": "2019-05-12T17:23:47.327Z",
          "content": "<p>Yes</p>",
          "rawMarkdown": "Yes",
          "votes": 2
        },
        {
          "id": 530691,
          "postDate": "2019-05-13T12:54:56.290Z",
          "content": "<p>Yes.. I have</p>",
          "rawMarkdown": "Yes.. I have",
          "votes": 2
        }
      ]
    },
    {
      "id": 529524,
      "postDate": "2019-05-10T05:41:33.410Z",
      "content": "<p>You can achieve 0.67+ using only the curated dataset, a single model architecture, and k-fold averaging.</p>\n\n<p>I think it's still too early to say which is the key to better performance at this stage of competition. Besides, I'm still lagging behind, so the second question should be left to someone more qualified to answer.</p>",
      "rawMarkdown": "You can achieve 0.67+ using only the curated dataset, a single model architecture, and k-fold averaging.\n\nI think it's still too early to say which is the key to better performance at this stage of competition. Besides, I'm still lagging behind, so the second question should be left to someone more qualified to answer.",
      "votes": 2,
      "replies": [
        {
          "id": 529729,
          "postDate": "2019-05-10T15:42:40.173Z",
          "content": "<p>Great!\nI wonder if k-fold averaging is allowed for the final stage kernel.\nHow many percentage can k-fold averaging bring for the LB than only one model?</p>",
          "rawMarkdown": "Great!\nI wonder if k-fold averaging is allowed for the final stage kernel.\nHow many percentage can k-fold averaging bring for the LB than only one model?",
          "votes": 1
        },
        {
          "id": 529765,
          "postDate": "2019-05-10T17:31:40.847Z",
          "content": "<p>I don't see why k-fold averaging wouldn't be allowed. Can you explain the reason for me?</p>\n\n<p>I always submit averaged predictions because deep models usually have high variations (so LB score can be very unreliable with a just one model).</p>",
          "rawMarkdown": "I don't see why k-fold averaging wouldn't be allowed. Can you explain the reason for me?\n\nI always submit averaged predictions because deep models usually have high variations (so LB score can be very unreliable with a just one model).",
          "votes": 1
        },
        {
          "id": 529802,
          "postDate": "2019-05-10T19:42:29.327Z",
          "content": "<p>The limit of kernel running time may affect, prediction of k models should be done for the stage 2 test set in kaggle kernel.</p>",
          "rawMarkdown": "The limit of kernel running time may affect, prediction of k models should be done for the stage 2 test set in kaggle kernel.",
          "votes": 1
        },
        {
          "id": 529862,
          "postDate": "2019-05-11T02:44:11.847Z",
          "content": "<p>I think as long as you keep the inference time under 20 mins in this stage (using GPU) you should be fine.</p>\n\n<p>EDIT: Correcttion -- 30mins -&gt; 20mins since the stage 2 test set is about three times as large as in stage 1.</p>",
          "rawMarkdown": "I think as long as you keep the inference time under 20 mins in this stage (using GPU) you should be fine.\n\nEDIT: Correcttion -- 30mins -&gt; 20mins since the stage 2 test set is about three times as large as in stage 1.",
          "votes": 2
        }
      ]
    },
    {
      "id": 537017,
      "postDate": "2019-05-26T02:10:05.307Z",
      "content": "<p>How many folds are you using? I have similar single model scores, and using 5-fold I did not get above 0.7</p>",
      "rawMarkdown": "How many folds are you using? I have similar single model scores, and using 5-fold I did not get above 0.7"
    },
    {
      "id": 534871,
      "postDate": "2019-05-22T00:46:23.563Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 535068,
          "postDate": "2019-05-22T08:34:46.073Z",
          "content": "<p>I have the same quriosity. For me, 2D CNN outperforms 1D in both training time and accuracy. I think I could get close to 0.6 if I apply all the bells and whistles I used since then (augmentation, model averaging, etc) but I didn't bother trying.</p>",
          "rawMarkdown": "I have the same quriosity. For me, 2D CNN outperforms 1D in both training time and accuracy. I think I could get close to 0.6 if I apply all the bells and whistles I used since then (augmentation, model averaging, etc) but I didn't bother trying."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 529707,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-05-10T15:01:26",
      "content": "<p>My best single (and only) model is a simple CNN with 4 layers trained on the curated data only. It gives me 0.7 LB score.</p>\n\n<p>I wonder if anyone uses something different from CNNs for achieving 0.65+ scores?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 529728,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-10T15:40:49.097000",
          "content": "<p>Thanks for the info. I'm using 30+ convolution layers (LB 0.685 as previously stated). If four layers can already achieve 0.7 LB score, then my 30+ is definitely too many... </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529732,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T15:47:33.333000",
          "content": "<p>0.7 LB score with 4 layers! I just tried lots of different cnn structures.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529734,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-05-10T15:49:44.797000",
          "content": "<p>Well, I haven't played with the architecture yet, so a deep network might be better than a shallow one eventually. The difference in score may also come from different sources such as data preprocessing, training schedule, etc.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 529753,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-05-10T16:44:34.717000",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> from my experience, the architecture rarely matters much until there are other ways to improve, and there are a lot of ways in this particular competition</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 529757,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-05-10T16:55:17.130000",
          "content": "<p>Agreed I think the architecture matters far less in this case.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 529771,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-10T17:43:56.223000",
          "content": "<p>Interesting. I have not really explored shallower models yet. The take home message here for me is that I could've saved a lot of training time using a shallower model and sped up the data pipeline iterations.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 529777,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T18:21:17.900000",
          "content": "<p>Thanks for the information <a href=\"/ddanevskyi\">@ddanevskyi</a> <a href=\"/jamesrequa\">@jamesrequa</a> . Cannot agree more <a href=\"/ceshine\">@ceshine</a> .</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 547036,
          "author_name": "neoguo",
          "author_url": "",
          "post_date": "2019-06-07T07:19:07.810000",
          "content": "<p>Agree, I also used a 4 layer network, got 6.5, and did not improve too much with RESNET50.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 532617,
      "author_name": "Eric Bouteillon",
      "author_url": "",
      "post_date": "2019-05-17T11:37:05.470000",
      "content": "<p>To get back to the original topic ;)\nMy single best model, LB: 0.720, CV: 0.86</p>",
      "votes": 5,
      "replies": [
        {
          "id": 532687,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-17T14:33:18.177000",
          "content": "<p>Thanks for sharing this info! Did you only use curated data?</p>\n\n<p>From the responses to this thread, there are multiple reports of similar local CV scores with quite different LB scores. I wonder if we are in serious danger of overfitting the leaderboard (by switching between models with similar CV scores but different LB scores).</p>\n\n<p>EDIT: grammar. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532692,
          "author_name": "Eric Bouteillon",
          "author_url": "",
          "post_date": "2019-05-17T14:45:37.203000",
          "content": "<p>Yes, my CV score is computed using only the curated dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532730,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-17T15:51:38.940000",
          "content": "<p>That's quite a feat! Thanks for the info.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 529717,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-05-10T15:25:21.793000",
      "content": "<p>For the second point, I'm firmly convinced one needs to find a way to combat label noise in the noisy train subset.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 529716,
      "author_name": "Miroslav Valan",
      "author_url": "",
      "post_date": "2019-05-10T15:18:50.477000",
      "content": "<p>0.659 curated only, simple cnn, 20%hold out, no augmentation. <a href=\"https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai\">this kernel</a> plus some tweeks.\nwonder what will bring k-fold</p>",
      "votes": 4,
      "replies": [
        {
          "id": 529733,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T15:49:07.133000",
          "content": "<p>mine is so close to yours. I'm trying to bring in k-fold averaging now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 530763,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-05-13T16:10:03.220000",
          "content": "<p>5 kfold .684 local lwlrap .84-.86</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 531562,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-05-15T06:13:46.737000",
          "content": "<p><a href=\"https://www.kaggle.com/sailorwei\">@wqk</a> are you doing 5 folds or something else.  have you changed anything else but added Kfold (just yes or no - do secret sauce needed). With 5kfold i gain 0.020-.025 and your case is very impressive.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 531672,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-15T10:17:01.670000",
          "content": "<p>I did not change other things except adding Kfold and increasing epochs, with local lwlraps are around 0.76.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532371,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-05-16T18:49:58.077000",
          "content": "<p>I feel you man... We have the same local lwlrap range and the same boost after 5-fold averaging.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 533781,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-20T02:54:33.080000",
          "content": "<p>Hi <a href=\"/jihangz\">@jihangz</a> , how is your local lwlrap now? I cannot increase it for several days.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 529553,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2019-05-10T07:02:52.113000",
      "content": "<p>0.689 here with curated, single basic model, k-fold, no augmentation. Local lwlrap 0.860\nI imagine the 'key point' is to use the noisy dataset in some way but haven't done much investigation there yet.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 529566,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-10T08:44:19.140000",
          "content": "<p>Very impressive! </p>\n\n<p>My augmented one achieved 0.685 at best, with local lwlrap of 0.83.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 529709,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-05-10T15:04:05.847000",
          "content": "<p>That's impressive, your local lwlrap seems to be very high, my best model gives me only 0.83 local lwlrap.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 529730,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T15:43:48.483000",
          "content": "<p>Really a high local lwlrap!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529756,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-05-10T16:54:04.197000",
          "content": "<p>My best so far is 0.683 with a custom shallow single model using only curated dataset but with this same model my best local lwlrap is much higher over 0.87. Perhaps I need to more carefully select the validation set as there is a large discrepancy with LB score. I am also still looking into how best to utilize the noisy dataset as this is the real purpose of the competition (i.e. SSL). </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 529808,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-05-10T19:55:32.177000",
          "content": "<p>Curated only, shallow CNN. I didn't calculate oof CV, but my 5-fold valid lwlraps are between 0.835 and 0.865. Single model gives me LB 0.676. </p>\n\n<p>Based on everyone's info here, I expect a huge shakeup in second stage, like Quora and PetFinder competition with the same inference format (2 stage with different test set).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529823,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-05-10T21:10:48.990000",
          "content": "<p><a href=\"/jihangz\">@jihangz</a> may I ask why do you expect a shakeup here? LB metric seems to be pretty stable. There is a gap between a local score and LB score, but as long as my local score improves, LB metric improves too.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529827,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-05-10T21:37:55.340000",
          "content": "<p>First, test set is small, and the test set in 2nd stage is 3x the current one. \nSecond, my local score is at the same level as most of yours, but my LB is lower. </p>\n\n<p>Time to trust CV?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529851,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-05-11T01:02:56.930000",
          "content": "<p>As far as I have listened to the samples, many of them seems to have strong correlation (coming from similar sound source, music instrument or just applied different EQ/effect to a single source sound...), then I guess it would be difficult to build <em>sound</em> CV folds unless source information or the similar is available to separate train/valid adequately.</p>\n\n<p>I agree with <a href=\"/robga\">@robga</a>, most promising way to go could be to find how to utilize noisy set.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 536781,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2019-05-25T09:10:11.877000",
      "content": "<p>LB 0.629 with curated only.\nLB 0.680 with both curated and noisy.</p>\n\n<p>I use two models for training. One is to be trained(from scratch to LB 0.680) and the other one(LB 0.629) is to provide soft labels to noisy data just before it's used for training. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 536829,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-05-25T11:31:28.790000",
          "content": "<p>So you use the predicted soft labels in addition to the provided labels of the noisy data? That’s interesting idea!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 536950,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-05-25T18:40:41.660000",
          "content": "<p>Yes, it works well and could be better if I can improve curated only model in the first place. I will try mixup if I can find time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 536960,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-05-25T20:16:25.560000",
          "content": "<p>I tried similar stuff but it made my score worse so far.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 536972,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2019-05-25T21:11:20.270000",
          "content": "<p>I provide soft-labels(sigmoid outputs of the model trained on curated only) to cropped(128x160) noisy mel data every time it's feeded to the model to be trained. I used original labels of noisy data together with soft-labels(used np.maximum) but it could be better just ignoring original labels. </p>\n\n<p>I tried both\n- trained on noisy and then trained on curated in each epoch.\n- trained on noisy for 50-60 epochs and then trained on curated for 20-30 epochs.</p>\n\n<p>The latter was a little better so far.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 536987,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-05-25T23:02:42.433000",
          "content": "<p>I tried same approach, but I dont get better result</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 537062,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-05-26T05:32:13.310000",
          "content": "<p>Thanks so much <a href=\"/appian\">@appian</a> for this valuable ideas!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 537853,
      "author_name": "F.J.Martinez-de-Pison",
      "author_url": "",
      "post_date": "2019-05-27T18:18:17.727000",
      "content": "<p>Two single models (curated+augmented+noisy) with two different preprocessed datasets.\n&lt;10 minutes with GPU (preprocessing+inference)\nLB M1=0.685\nLB M2=0.680\nBlending    1.05xM1+0.95xM2= 0.696</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 536373,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2019-05-24T11:18:40.813000",
      "content": "<p>still struggling for 0.67, do not know what's wrong ... sign ...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 546431,
          "author_name": "Robert Bracco",
          "author_url": "",
          "post_date": "2019-06-06T15:30:27.120000",
          "content": "<p>Looks like you figured it out according to your standing :) , nice work!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 530601,
      "author_name": "hongxiaofeng",
      "author_url": "",
      "post_date": "2019-05-13T08:48:52.057000",
      "content": "<p>Hi,\nWould you like to tell us how do you split the validation set? Random or Balanced(Each category has the same quantity)?\nthank you a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 530919,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-14T00:51:47.577000",
          "content": "<p>KFold is used now, seems StraitifiedKFold cannot accept multi-label dataset. </p>\n\n<p><a href=\"https://www.kaggle.com/osciiart/multilabel-stratifiedkfold-by-randomized-algorithm\">https://www.kaggle.com/osciiart/multilabel-stratifiedkfold-by-randomized-algorithm</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531517,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-15T04:07:43.183000",
          "content": "<p>One way to do it is to treat multi-label as separate classes, but the multi-label cases in this dataset are so rare it's not feasible to do so. My hacky solution is to randomly select one of the label in those cases and make the dataset single-label to creat stratified k-folds. (You need to make sure you pick the same labels when evaluating.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 534924,
          "author_name": "Josh Varty",
          "author_url": "",
          "post_date": "2019-05-22T04:26:19.657000",
          "content": "<p>@wqk Here's one approach for multilabel stratification: <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a> that seems to work well for me.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 530976,
      "author_name": "Ceshine Lee",
      "author_url": "",
      "post_date": "2019-05-14T04:38:40.353000",
      "content": "<p>Regarding to your 20190503 update <a href=\"/sailorwei\">@sailorwei</a> , are your saying by adding cv, your LB score jumped from 0.655 to 0.7+? I'm not exactly sure what you mean by \"cv\". Can you please elaborate?</p>\n\n<p>EDIT: correction - update previous best score from 0.64 -&gt; 0.655 (the one from mixup).</p>",
      "votes": 2,
      "replies": [
        {
          "id": 530987,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-05-14T04:59:20.667000",
          "content": "<p>I think it means training N classifiers on N folds of training set and taking the average of N predictions on test set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 530996,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-14T05:12:28.983000",
          "content": "<p>It's as <a href=\"/jihangz\">@jihangz</a>  said. \"k-fold averaging\" as you mentioned doesnot mean that?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531005,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-14T05:35:28.760000",
          "content": "<p>Yes that what I meant by \"k-fold averaging\". That was a pretty large jump in LB score by just doing averaging, so I just wanted to make sure that's what you meant.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531470,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-15T01:25:18.323000",
          "content": "<p>Yes, thanks for your previous information. So \"k-fold averaging\" should be stronger than training only one model with all train data? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 531516,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-15T04:00:10.357000",
          "content": "<p>Yes it usually gives more stable and slightly better results. But it's unusual to have such a big difference in my experience. If you don't mind, can you share your local CV score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532287,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-16T15:17:26.980000",
          "content": "<p>Sure, my local lwlraps are all around 0.76 now. I cannot figure out the reason why it's lower than those in this discussion.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 532538,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-17T07:40:05.103000",
          "content": "<p>Thanks for sharing. Mine is ~ .85 now (that is for the entire clip, not for a single chunk).  I'm starting to hit the problem of local CV improvements not translating into LB improvements. </p>\n\n<p>The train/test gap of this dataset is a bit bizarre, and might be crucial for getting a better private LB results as some has mentioned (this is just only my own speculation, though). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532591,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-17T10:41:04.567000",
          "content": "<p>So high. Maybe overfitting, How many conv layers in your net?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 532610,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-05-17T11:19:27.387000",
          "content": "<p><a href=\"/ceshine\">@ceshine</a> \nMy local CV is quite similar as you. I faced the same problem since local CV increases but LB says no. It is pretty weird that my best model has local CV lower than the second one (however the GAP still be high). I am quite sure that It is not randomness. </p>\n\n<p><a href=\"/sailorwei\">@sailorwei</a> \nYour local CV and GAP are so impressive. I just want to ask if you compute local CV in the correct way or not. You said you used mixup, the mixup is only used for training. If you use for validation, the score is low. IMO, it is only the reason that can explain your local CV and GAP. Otherwise, it may come from some <code>secret sauce</code>. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 532657,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-17T13:31:41.750000",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> I wouldn't call it overfitting since the local CV scores were from the validation set. Difference in train and test distribution is more likely the cause IMO. My model is smaller than resnet50.</p>\n\n<p><a href=\"/backaggle\">@backaggle</a> Yeah, most likely not randomness. There seems to be some models or data pipelines that generalize better to the test set. </p>\n\n<p>In case you haven't read it, here's <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89312#latest-523696\">the discussion thread on the topic of train/test discrepancy.</a> </p>\n\n<p>EDIT: reorganized the sentences to make it more clear.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 533783,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-20T03:05:39.477000",
          "content": "<p><a href=\"/backaggle\">@backaggle</a> mixup is only for the train data each batch. I'm still confused about the low local validation lwlrap. The local validation lwlrap function is from this kernel  <a href=\"https://www.kaggle.com/vinayaks/simple-changes-significant-improvement-in-lb\">https://www.kaggle.com/vinayaks/simple-changes-significant-improvement-in-lb</a> .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 533795,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-05-20T04:11:53.407000",
          "content": "<p><a href=\"/sailorwei\">@sailorwei</a> <br>\nHow do you compute lwlrap ?. In my experiments, there are two cases \n- (1) Split audio into several overlaped segments, then taking average overall \n- (2) Random segments (TTA)&lt; then taking average overall </p>\n\n<p>My results show that <code>(1)</code> gives low <code>lwlrap</code> while <code>(2)</code> gives higher. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 533816,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-05-20T05:13:56.247000",
          "content": "<p>Hi <a href=\"/backaggle\">@backaggle</a>,</p>\n\n<p>I could not clearly understand the difference between (1) and (2), would you mind clarify them further?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 533818,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-05-20T05:26:09.300000",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> </p>\n\n<ul>\n<li>For <code>(1)</code>. \nOne audio is splitted into several segment parts (let say: window size = 128, stride = 64). \nThe <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/89382#latest-517614\">baseline</a> system follows this method with <code>stride = 0</code> (no overlap)</li>\n<li>For <code>(2)</code>.\nWe take randomly <code>n_parts</code>. Your amazing keras kernel already did it (called as TTA). </li>\n</ul>\n\n<p>The final result may be evaluated by taking average among segments. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 533869,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-05-20T07:03:15.223000",
          "content": "<p>Thanks <a href=\"/backaggle\">@backaggle</a> for clarification and praising. My kernel performance is still far beyond yours and the others. Hope to learn more from all of you during this competition!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 535917,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-23T16:15:45.117000",
          "content": "<p><a href=\"/backaggle\">@backaggle</a>  thanks, I missed your message and just got it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 538868,
          "author_name": "hongxiaofeng",
          "author_url": "",
          "post_date": "2019-05-29T08:03:11.440000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 530363,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-05-12T16:21:20.777000",
      "content": "<p>Hi, just curious if anyone here use Keras and be able to get score more than 0.67+</p>",
      "votes": 2,
      "replies": [
        {
          "id": 530388,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-05-12T17:23:47.327000",
          "content": "<p>Yes</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 530691,
          "author_name": "Sylvia",
          "author_url": "",
          "post_date": "2019-05-13T12:54:56.290000",
          "content": "<p>Yes.. I have</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 529524,
      "author_name": "Ceshine Lee",
      "author_url": "",
      "post_date": "2019-05-10T05:41:33.410000",
      "content": "<p>You can achieve 0.67+ using only the curated dataset, a single model architecture, and k-fold averaging.</p>\n\n<p>I think it's still too early to say which is the key to better performance at this stage of competition. Besides, I'm still lagging behind, so the second question should be left to someone more qualified to answer.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 529729,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T15:42:40.173000",
          "content": "<p>Great!\nI wonder if k-fold averaging is allowed for the final stage kernel.\nHow many percentage can k-fold averaging bring for the LB than only one model?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529765,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-10T17:31:40.847000",
          "content": "<p>I don't see why k-fold averaging wouldn't be allowed. Can you explain the reason for me?</p>\n\n<p>I always submit averaged predictions because deep models usually have high variations (so LB score can be very unreliable with a just one model).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529802,
          "author_name": "wqk",
          "author_url": "",
          "post_date": "2019-05-10T19:42:29.327000",
          "content": "<p>The limit of kernel running time may affect, prediction of k models should be done for the stage 2 test set in kaggle kernel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 529862,
          "author_name": "Ceshine Lee",
          "author_url": "",
          "post_date": "2019-05-11T02:44:11.847000",
          "content": "<p>I think as long as you keep the inference time under 20 mins in this stage (using GPU) you should be fine.</p>\n\n<p>EDIT: Correcttion -- 30mins -&gt; 20mins since the stage 2 test set is about three times as large as in stage 1.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 537017,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2019-05-26T02:10:05.307000",
      "content": "<p>How many folds are you using? I have similar single model scores, and using 5-fold I did not get above 0.7</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 534871,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-22T00:46:23.563000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 535068,
          "author_name": "Victor Zaguskin",
          "author_url": "",
          "post_date": "2019-05-22T08:34:46.073000",
          "content": "<p>I have the same quriosity. For me, 2D CNN outperforms 1D in both training time and accuracy. I think I could get close to 0.6 if I apply all the bells and whistles I used since then (augmentation, model averaging, etc) but I didn't bother trying.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "529521": "With only curated dataset and cnn, 20%hold out, my best single model LB is 0.64.\n\nAdd mixup, LB to 0.655, with local lwlrap 0.76, almost the same with before.\n\nWhat's the key point to get better performance in this task?\n\n-------------------------\nupdate 20190513\n\nWith only cv (\"k-fold averaging\") added, LB to 0.7+, amazing improvement.\n\nThe next step may be effective usage of noisy subset.\n\n------------------------\nupdate 20190610\n\nsingle model with inception V3, reach 0.691, 5 fold cv, 0.71+\n\nsingle fold kernel here: \n\nhttps://www.kaggle.com/sailorwei/fat2019-2d-cnn-with-mixup-lb-0-673\n",
    "529707": "My best single (and only) model is a simple CNN with 4 layers trained on the curated data only. It gives me 0.7 LB score.\n\nI wonder if anyone uses something different from CNNs for achieving 0.65+ scores?",
    "532617": "To get back to the original topic ;)\nMy single best model, LB: 0.720, CV: 0.86",
    "529717": "For the second point, I'm firmly convinced one needs to find a way to combat label noise in the noisy train subset.",
    "529716": "0.659 curated only, simple cnn, 20%hold out, no augmentation. [this kernel](https://www.kaggle.com/daisukelab/cnn-2d-basic-solution-powered-by-fast-ai) plus some tweeks.\nwonder what will bring k-fold",
    "529553": "0.689 here with curated, single basic model, k-fold, no augmentation. Local lwlrap 0.860\nI imagine the 'key point' is to use the noisy dataset in some way but haven't done much investigation there yet.",
    "536781": "LB 0.629 with curated only.\nLB 0.680 with both curated and noisy.\n\nI use two models for training. One is to be trained(from scratch to LB 0.680) and the other one(LB 0.629) is to provide soft labels to noisy data just before it's used for training. ",
    "537853": "Two single models (curated+augmented+noisy) with two different preprocessed datasets.\n&lt;10 minutes with GPU (preprocessing+inference)\nLB M1=0.685\nLB M2=0.680\nBlending    1.05xM1+0.95xM2= 0.696",
    "536373": "still struggling for 0.67, do not know what's wrong ... sign ...",
    "530601": "Hi,\nWould you like to tell us how do you split the validation set? Random or Balanced(Each category has the same quantity)?\nthank you a lot!",
    "530976": "Regarding to your 20190503 update @sailorwei , are your saying by adding cv, your LB score jumped from 0.655 to 0.7+? I'm not exactly sure what you mean by \"cv\". Can you please elaborate?\n\nEDIT: correction - update previous best score from 0.64 -&gt; 0.655 (the one from mixup).",
    "530363": "Hi, just curious if anyone here use Keras and be able to get score more than 0.67+",
    "529524": "You can achieve 0.67+ using only the curated dataset, a single model architecture, and k-fold averaging.\n\nI think it's still too early to say which is the key to better performance at this stage of competition. Besides, I'm still lagging behind, so the second question should be left to someone more qualified to answer.",
    "537017": "How many folds are you using? I have similar single model scores, and using 5-fold I did not get above 0.7",
    "534871": ""
  }
}