{
  "id": 95890,
  "title": "96th Place Solution: validation of noisy data improve ~0.02 LB score",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/95890",
  "author_name": "Jing Liu",
  "post_date": "2019-06-16T05:31:23.707000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>Background:</strong> First time kaggler, no server  with GPU locally, all computation was on Kaggle kernels. \n<strong>Final Workflow:</strong> \n1. Curated-only 5-fold CNN model got the optimal LB sscore\n2. Preprocess noisy data to get a raw selected noisy dataset, by clustering 7 features (both temporal and spectral features)\n3. Validate the selected noisy dataset by curated-only model\n4. Train the mixed model with the curated and validated noisy data, got around 0.02 score boost.</p>\n\n<p><strong>The detailed workflow pic is too big to display in the post, so here it is</strong> <a href=\"https://drive.google.com/file/d/1AaFDTsedh-3liF7K4V-QAyCWA8l8RtA_/view?usp=sharing\">Final Setup</a>\n- <strong>Features for deep learning:</strong> \nAfter inspecting the distribution of the audio duration of both curated and test dataset,  I've tried clip duration 1, 2, 3 seconds, mel channels 128, 180, 192, and other parameters combination. The best result came with the above config in the final setup. 0.01 and 0.02-3 LB boost from 1s, 3s to 2s respectively.  180 mel channels can improve the result by 0.02 LB, similar to 192 (choose the less feature to save computation time). \n- <strong>Noisy Data Preprocessing:</strong>\n- - This step is mainly for reducing the size of noisy dataset, for it's difficult to deal with all the noisy data in limited kaggle kernels. \n- - Thanks to Essentia: <a href=\"https://essentia.upf.edu/documentation/algorithms_reference.html\">https://essentia.upf.edu/documentation/algorithms_reference.html</a>\n- - Found a critical error during writing this summary: use the wrong window size config(too short) for feature extraction, which may make this step useless and more like random pick of noisy data.\n- <strong>Random pick:</strong>\nThe curated data would overfit if used all frames from the audio files. Tried max = 5 or 10, 5 gave the better result, 0.01 LB boost from max = 10. \n- <strong>Balance the data or not?</strong>\nThere is no time left for me to do the test with balanced dataset, because I chose the wrong balance method at first by both undersampling and upsampling each class to the average count, which produced a worse result because undersampling would reduce the useful data. When I realized this at a pretty late stage, I decided to use the unbalanced dataset, considering the lwlrap from curated-only model didn't show much bias against the imbalance. But this is still an open question and any discussion is welcome. \n- <strong>Network Choice</strong>\nTried several network setups,  CRNN with GRU as well (produced competitive result with curated-only data, but no time to tune the config), 1d RNN network with raw data as input was built and tried but performed badly. So finally chose VGG16 and did some experiments on the last layers. \nKey takeway: \n1. 2 FC layers with 2048 neurons produced worse local lwlrap than 1 FC with 512, and may produce memory error in kaggle kernel. Though not fully tested, it looks in my experiments that the single FC at last model stage produced better result.\n2. Thanks to a newly published paper <em><a href=\"https://arxiv.org/pdf/1905.05928.pdf\">Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks</a></em>,  my LB score boosted 0.03 by changing the unit order accordingly, that is BN-Dropout before the weight layer.\n3. Dropout rate from 0.2 to 0.5 to 0.7, the LB score boosted by 0.031 and 0.001. Though 0.5 and 0.7 produced similar result, chose 0.7 may make it more robust to overfitting.\n4. Batch size didn't show much difference in my experiments (64, 128,256 batch sizes).\n5. The more training fold, the better performance, but it requires more training time. And 10-fold improved only 0.004 compared to 5-fold. So used 5-fold at last for both curated and mixed model. \n6. The CRNN network showed different lwlrap per class result, and may be a complement to  CNN. (But as a beginner in sound scene classification and my future focus is model interpretability, I'd rather not to use ensembling but focus on finding a single optimal network this time)\n- <strong>Data Augmentation:</strong>\nSpecAugment without time warping, and mixup are used.\n- <strong>Noisy Data Validation:</strong>\n*The confidence value = r * label_weight_per_class / lwlrap_per_class*\nThe greater <em>r</em> is, the stricter the validation rule is. The best score is produced when r=5. \n- <strong>Test Setup:</strong>\nFound no improvement with any TTA, so simply use the average of the predicetions from all frames.</p>\n\n<p>Thanks to Kaggle and the competition organization. I've learnt a lot along this competition. Good luck to all in the second stage.p</p>",
  "messages": [
    {
      "id": 553669,
      "postDate": "2019-06-16T05:31:23.707Z",
      "content": "<p><strong>Background:</strong> First time kaggler, no server  with GPU locally, all computation was on Kaggle kernels. \n<strong>Final Workflow:</strong> \n1. Curated-only 5-fold CNN model got the optimal LB sscore\n2. Preprocess noisy data to get a raw selected noisy dataset, by clustering 7 features (both temporal and spectral features)\n3. Validate the selected noisy dataset by curated-only model\n4. Train the mixed model with the curated and validated noisy data, got around 0.02 score boost.</p>\n\n<p><strong>The detailed workflow pic is too big to display in the post, so here it is</strong> <a href=\"https://drive.google.com/file/d/1AaFDTsedh-3liF7K4V-QAyCWA8l8RtA_/view?usp=sharing\">Final Setup</a>\n- <strong>Features for deep learning:</strong> \nAfter inspecting the distribution of the audio duration of both curated and test dataset,  I've tried clip duration 1, 2, 3 seconds, mel channels 128, 180, 192, and other parameters combination. The best result came with the above config in the final setup. 0.01 and 0.02-3 LB boost from 1s, 3s to 2s respectively.  180 mel channels can improve the result by 0.02 LB, similar to 192 (choose the less feature to save computation time). \n- <strong>Noisy Data Preprocessing:</strong>\n- - This step is mainly for reducing the size of noisy dataset, for it's difficult to deal with all the noisy data in limited kaggle kernels. \n- - Thanks to Essentia: <a href=\"https://essentia.upf.edu/documentation/algorithms_reference.html\">https://essentia.upf.edu/documentation/algorithms_reference.html</a>\n- - Found a critical error during writing this summary: use the wrong window size config(too short) for feature extraction, which may make this step useless and more like random pick of noisy data.\n- <strong>Random pick:</strong>\nThe curated data would overfit if used all frames from the audio files. Tried max = 5 or 10, 5 gave the better result, 0.01 LB boost from max = 10. \n- <strong>Balance the data or not?</strong>\nThere is no time left for me to do the test with balanced dataset, because I chose the wrong balance method at first by both undersampling and upsampling each class to the average count, which produced a worse result because undersampling would reduce the useful data. When I realized this at a pretty late stage, I decided to use the unbalanced dataset, considering the lwlrap from curated-only model didn't show much bias against the imbalance. But this is still an open question and any discussion is welcome. \n- <strong>Network Choice</strong>\nTried several network setups,  CRNN with GRU as well (produced competitive result with curated-only data, but no time to tune the config), 1d RNN network with raw data as input was built and tried but performed badly. So finally chose VGG16 and did some experiments on the last layers. \nKey takeway: \n1. 2 FC layers with 2048 neurons produced worse local lwlrap than 1 FC with 512, and may produce memory error in kaggle kernel. Though not fully tested, it looks in my experiments that the single FC at last model stage produced better result.\n2. Thanks to a newly published paper <em><a href=\"https://arxiv.org/pdf/1905.05928.pdf\">Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks</a></em>,  my LB score boosted 0.03 by changing the unit order accordingly, that is BN-Dropout before the weight layer.\n3. Dropout rate from 0.2 to 0.5 to 0.7, the LB score boosted by 0.031 and 0.001. Though 0.5 and 0.7 produced similar result, chose 0.7 may make it more robust to overfitting.\n4. Batch size didn't show much difference in my experiments (64, 128,256 batch sizes).\n5. The more training fold, the better performance, but it requires more training time. And 10-fold improved only 0.004 compared to 5-fold. So used 5-fold at last for both curated and mixed model. \n6. The CRNN network showed different lwlrap per class result, and may be a complement to  CNN. (But as a beginner in sound scene classification and my future focus is model interpretability, I'd rather not to use ensembling but focus on finding a single optimal network this time)\n- <strong>Data Augmentation:</strong>\nSpecAugment without time warping, and mixup are used.\n- <strong>Noisy Data Validation:</strong>\n*The confidence value = r * label_weight_per_class / lwlrap_per_class*\nThe greater <em>r</em> is, the stricter the validation rule is. The best score is produced when r=5. \n- <strong>Test Setup:</strong>\nFound no improvement with any TTA, so simply use the average of the predicetions from all frames.</p>\n\n<p>Thanks to Kaggle and the competition organization. I've learnt a lot along this competition. Good luck to all in the second stage.p</p>",
      "rawMarkdown": "**Background:** First time kaggler, no server  with GPU locally, all computation was on Kaggle kernels. \n**Final Workflow:** \n1. Curated-only 5-fold CNN model got the optimal LB sscore\n2. Preprocess noisy data to get a raw selected noisy dataset, by clustering 7 features (both temporal and spectral features)\n3. Validate the selected noisy dataset by curated-only model\n4. Train the mixed model with the curated and validated noisy data, got around 0.02 score boost.\n\n**The detailed workflow pic is too big to display in the post, so here it is** [Final Setup](https://drive.google.com/file/d/1AaFDTsedh-3liF7K4V-QAyCWA8l8RtA_/view?usp=sharing)\n- **Features for deep learning:** \nAfter inspecting the distribution of the audio duration of both curated and test dataset,  I've tried clip duration 1, 2, 3 seconds, mel channels 128, 180, 192, and other parameters combination. The best result came with the above config in the final setup. 0.01 and 0.02-3 LB boost from 1s, 3s to 2s respectively.  180 mel channels can improve the result by 0.02 LB, similar to 192 (choose the less feature to save computation time). \n- **Noisy Data Preprocessing:**\n- - This step is mainly for reducing the size of noisy dataset, for it's difficult to deal with all the noisy data in limited kaggle kernels. \n- - Thanks to Essentia: https://essentia.upf.edu/documentation/algorithms_reference.html\n- - Found a critical error during writing this summary: use the wrong window size config(too short) for feature extraction, which may make this step useless and more like random pick of noisy data.\n- **Random pick:**\nThe curated data would overfit if used all frames from the audio files. Tried max = 5 or 10, 5 gave the better result, 0.01 LB boost from max = 10. \n- **Balance the data or not?**\nThere is no time left for me to do the test with balanced dataset, because I chose the wrong balance method at first by both undersampling and upsampling each class to the average count, which produced a worse result because undersampling would reduce the useful data. When I realized this at a pretty late stage, I decided to use the unbalanced dataset, considering the lwlrap from curated-only model didn't show much bias against the imbalance. But this is still an open question and any discussion is welcome. \n- **Network Choice**\nTried several network setups,  CRNN with GRU as well (produced competitive result with curated-only data, but no time to tune the config), 1d RNN network with raw data as input was built and tried but performed badly. So finally chose VGG16 and did some experiments on the last layers. \nKey takeway: \n1. 2 FC layers with 2048 neurons produced worse local lwlrap than 1 FC with 512, and may produce memory error in kaggle kernel. Though not fully tested, it looks in my experiments that the single FC at last model stage produced better result.\n2. Thanks to a newly published paper *[Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks](https://arxiv.org/pdf/1905.05928.pdf)*,  my LB score boosted 0.03 by changing the unit order accordingly, that is BN-Dropout before the weight layer.\n3. Dropout rate from 0.2 to 0.5 to 0.7, the LB score boosted by 0.031 and 0.001. Though 0.5 and 0.7 produced similar result, chose 0.7 may make it more robust to overfitting.\n4. Batch size didn't show much difference in my experiments (64, 128,256 batch sizes).\n5. The more training fold, the better performance, but it requires more training time. And 10-fold improved only 0.004 compared to 5-fold. So used 5-fold at last for both curated and mixed model. \n6. The CRNN network showed different lwlrap per class result, and may be a complement to  CNN. (But as a beginner in sound scene classification and my future focus is model interpretability, I'd rather not to use ensembling but focus on finding a single optimal network this time)\n- **Data Augmentation:**\nSpecAugment without time warping, and mixup are used.\n- **Noisy Data Validation:**\n*The confidence value = r \\* label_weight\\_per\\_class / lwlrap\\_per\\_class*\nThe greater *r* is, the stricter the validation rule is. The best score is produced when r=5. \n- **Test Setup:**\nFound no improvement with any TTA, so simply use the average of the predicetions from all frames.\n\nThanks to Kaggle and the competition organization. I've learnt a lot along this competition. Good luck to all in the second stage.p",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "553669": "**Background:** First time kaggler, no server  with GPU locally, all computation was on Kaggle kernels. \n**Final Workflow:** \n1. Curated-only 5-fold CNN model got the optimal LB sscore\n2. Preprocess noisy data to get a raw selected noisy dataset, by clustering 7 features (both temporal and spectral features)\n3. Validate the selected noisy dataset by curated-only model\n4. Train the mixed model with the curated and validated noisy data, got around 0.02 score boost.\n\n**The detailed workflow pic is too big to display in the post, so here it is** [Final Setup](https://drive.google.com/file/d/1AaFDTsedh-3liF7K4V-QAyCWA8l8RtA_/view?usp=sharing)\n- **Features for deep learning:** \nAfter inspecting the distribution of the audio duration of both curated and test dataset,  I've tried clip duration 1, 2, 3 seconds, mel channels 128, 180, 192, and other parameters combination. The best result came with the above config in the final setup. 0.01 and 0.02-3 LB boost from 1s, 3s to 2s respectively.  180 mel channels can improve the result by 0.02 LB, similar to 192 (choose the less feature to save computation time). \n- **Noisy Data Preprocessing:**\n- - This step is mainly for reducing the size of noisy dataset, for it's difficult to deal with all the noisy data in limited kaggle kernels. \n- - Thanks to Essentia: https://essentia.upf.edu/documentation/algorithms_reference.html\n- - Found a critical error during writing this summary: use the wrong window size config(too short) for feature extraction, which may make this step useless and more like random pick of noisy data.\n- **Random pick:**\nThe curated data would overfit if used all frames from the audio files. Tried max = 5 or 10, 5 gave the better result, 0.01 LB boost from max = 10. \n- **Balance the data or not?**\nThere is no time left for me to do the test with balanced dataset, because I chose the wrong balance method at first by both undersampling and upsampling each class to the average count, which produced a worse result because undersampling would reduce the useful data. When I realized this at a pretty late stage, I decided to use the unbalanced dataset, considering the lwlrap from curated-only model didn't show much bias against the imbalance. But this is still an open question and any discussion is welcome. \n- **Network Choice**\nTried several network setups,  CRNN with GRU as well (produced competitive result with curated-only data, but no time to tune the config), 1d RNN network with raw data as input was built and tried but performed badly. So finally chose VGG16 and did some experiments on the last layers. \nKey takeway: \n1. 2 FC layers with 2048 neurons produced worse local lwlrap than 1 FC with 512, and may produce memory error in kaggle kernel. Though not fully tested, it looks in my experiments that the single FC at last model stage produced better result.\n2. Thanks to a newly published paper *[Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks](https://arxiv.org/pdf/1905.05928.pdf)*,  my LB score boosted 0.03 by changing the unit order accordingly, that is BN-Dropout before the weight layer.\n3. Dropout rate from 0.2 to 0.5 to 0.7, the LB score boosted by 0.031 and 0.001. Though 0.5 and 0.7 produced similar result, chose 0.7 may make it more robust to overfitting.\n4. Batch size didn't show much difference in my experiments (64, 128,256 batch sizes).\n5. The more training fold, the better performance, but it requires more training time. And 10-fold improved only 0.004 compared to 5-fold. So used 5-fold at last for both curated and mixed model. \n6. The CRNN network showed different lwlrap per class result, and may be a complement to  CNN. (But as a beginner in sound scene classification and my future focus is model interpretability, I'd rather not to use ensembling but focus on finding a single optimal network this time)\n- **Data Augmentation:**\nSpecAugment without time warping, and mixup are used.\n- **Noisy Data Validation:**\n*The confidence value = r \\* label_weight\\_per\\_class / lwlrap\\_per\\_class*\nThe greater *r* is, the stricter the validation rule is. The best score is produced when r=5. \n- **Test Setup:**\nFound no improvement with any TTA, so simply use the average of the predicetions from all frames.\n\nThanks to Kaggle and the competition organization. I've learnt a lot along this competition. Good luck to all in the second stage.p"
  }
}