{
  "id": 96440,
  "title": "4th solution: Multitask Learning, Semi-supervised Learning and Ensemble",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/kaggler-ja-aims-oumed-4th-solution-multitask-learn",
  "author_name": "",
  "post_date": "2019-07-12T00:18:08.983Z",
  "votes": 60,
  "comment_count": 7,
  "views": 0,
  "content": "<p>We wrote up a technical report of our solution, so you can check our solution in detail when it's published. Here I'd like to describe a brief summary of our solution. Our score is public 0.752 (3rd)/private 0.75787 (4th before pending). <br>\nFor this competition, we used 3 strategies mainly. <br>\n1. Multitask learning with noisy labels. \n2. Semi-supervised learning (SSL) with noisy data. \n3. Averaging models trained with different time windows.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/556685/13619/Fig1_3.png\" alt=\"Multitask pipeline\"></p>\n\n<h2>Models</h2>\n\n<ul>\n<li>ResNet34 with log-mel</li>\n<li>EnvNet-v2 with waveform.</li>\n</ul>\n\n<h2>Preprocessing</h2>\n\n<ul>\n<li>128 mels</li>\n<li>128 Hz (347 STFT hop size)</li>\n</ul>\n\n<p>Log-mel is converted from power to dB after all augmentations applied. Thereafter, it is normalized by the mean and standard deviation of each data.</p>\n\n<h2>Augmentations</h2>\n\n<h3>Log-mel</h3>\n\n<ul>\n<li>Slicing</li>\n<li>MixUp</li>\n<li>Frequency masking</li>\n<li>Gain augmentation</li>\n</ul>\n\n<p>We tried 2, 4 and 8 seconds (256, 512 and 1024 dimensions) as a slicing-length and 4 seconds scores the best. Resizing, warping, time masking and white noise don't work.\n- Additional slicing</p>\n\n<p>Expecting more strong augmentation effect, after basic slicing, we shorten data samples in a range of 25 - 100% of the basic slicing-length by additional slicing and extend to the basic slicing-length by zero paddings. </p>\n\n<h3>Waveform</h3>\n\n<ul>\n<li>Slicing</li>\n<li>MixUp</li>\n<li>Gain augmentaion</li>\n<li>Scaling augmentation</li>\n</ul>\n\n<p>We tried 1.51, 3.02 and 4.54 seconds (66,650, 133,300 and 200,000 dimensions) as a slicing length and 4.54 seconds scores the best. </p>\n\n<h2>Training</h2>\n\n<h3>ResNet</h3>\n\n<ul>\n<li>Adam</li>\n<li>Cyclic cosine annealing (1e-3 -&gt; 1e-6)</li>\n<li>sigmoid and binary crossentropy</li>\n</ul>\n\n<h3>EnvNet</h3>\n\n<ul>\n<li>SGD</li>\n<li>Cyclic cosine annealing (1e-1 -&gt; 1e-6)</li>\n<li>(sigmoid and binary crossentropy) or (SoftMax and KL-divergence)</li>\n</ul>\n\n<h2>Postprocessing</h2>\n\n<p>Prediction using the full length of audio input scores better than prediction using slicing TTA. This may be because important components for classification is concentrated on the beginning part of audio samples. Actually, prediction with slices of the beginning part scores better than prediction with slices of the latter part.  We found padding augmentation is effective TTA. This is an augmentation method that applies zero paddings to both sides of audio samples with various length and averages prediction results. We think this method has an effect to emphasize the start and the end part of audio samples.</p>\n\n<h2>Multitask learning</h2>\n\n<p>In this task, the curated data and noisy data are labeled in a different manner, therefore treating them as the same one makes the model performance worse. To tackle this problem, we used a multitask learning (MTL) approach. The aim of MTL is to get synergy between 2 tasks without reducing the performance of each task. MTL learns features shared between 2 tasks and can be expected to achieve higher performance than learning independently. In our proposal, an encoder architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation. <br>\nMultitask module's components: <br>\n<code>FC(1024)-ReLU-DropOut(0.2)-FC(1024)-ReLU-DropOut(0.1)-FC(80)-sigmoid</code> <br>\nwe used BCE as a loss function. The loss weight ratio of curated and noisy is set as 1:1. By this method, CV lwlrap improved from 0.829 to 0.849 and score on the public LB increased + 0.021.</p>\n\n<h2>Semi-supervised learning</h2>\n\n<p>Applying SSL for this competition is difficult because of 2 reasons.\n1. Difference of data distribution between labeled data and unlabeled data. \n2. Multi-label classification task. </p>\n\n<p>In particular, the method that generates labels online like Mean teacher or MixMatch tends to collapse.  We tried pseudo label, Mean Teacher and MixMatch and all of them are not successful. <br>\nTherefore, we propose an SSL method that is robust to data distribution difference and can handle multi-label data.  For each noisy data sample, we guess the label using the trained model. The guessed label is sharpened by sharpening function proposed by MixMatch (soft pseudo label).  As the temperature of sharpening function, we tried a value of 1, 1.5 or 2 and 2 was the best.  Predictions of the trained model are obtained using snapshot ensemble with all the folds and cycle snapshots of 5-fold CV.  We used MSE as a loss function. we set loss weight of semi-supervised learning as 20. <br>\nBy soft pseudo labeling, The CV lwlrap improved from 0.849 to 0.870. On the other hand, on the public LB, improvement in score was slight (+0.001). We used predictions of all fold models to generate soft pseudo label so that high CV may be because of indirect label leak. However, even if we use labels generated by only the same fold model which has no label leak, CV was improved as compared to one without SSL. The model trained with the soft pseudo label is useful as a component of model averaging (+0.003 on the public LB).</p>\n\n<h2>Ensemble</h2>\n\n<p>For model averaging, we prepared models trained with various conditions. Especially, we found that averaging with models trained with different time window is effective. We used models below. In order to reduce prediction time,  the cycles and padding lengths used for the final submissions and averaging weights were chosen based on CV. \n1. ResNet34 slice=512, MTL\n2. ResNet34 slice=512, MTL, SSL\n3. ResNet34 slice=1024, MTL\n4. EnvNet-v2 slice=133,300, MTL, sigmoid\n5. EnvNet-v2 slice=133,300, MTL, SoftMax\n6. EnvNet-v2 slice=200,000, MTL, SoftMax</p>\n\n<h2>Comparison</h2>\n\n<p>condition/CV lwlrap <br>\nmodel #1 / 0.868 <br>\nmodel #2 / 0.886 <br>\nmodel #3 / 0.862 <br>\nmodel #4 / 0.815 <br>\nmodel #5 / 0.818 <br>\nmodel #6 / 0.820 <br>\n#1 + #3 / 0.876 <br>\n#1 + #2 + #3 / 0.890 <br>\n#4 + #5 + #6 / 0.836 <br>\nsubmission 1 (#1 + #2 + #3 + #4 + #5 + #6) / 0.896 <br>\nsubmission 2 (no pad TTA) / 0.895  </p>\n\n<h2>Code</h2>\n\n<p>I did all model training on kaggle kernels. So I can share all of them. Please note that code readability is low. For example, DataLoader class is named \"MFCCDataset\" but it actually loads log-mel.\n<a href=\"https://www.kaggle.com/osciiart/freesound2019-solution-links?scriptVersionId=16047723\">link to the summary of codes</a> <br>\nThe GitHub repository is available at <a href=\"https://github.com/OsciiArt/Freesound-Audio-Tagging-2019\">here</a>.   </p>\n\n<h2>Public LB history log</h2>\n\n<p>model / CV / Public LB <br>\nResNet18, MixUp, freq masking, slice=256 / 0.819 / 0.669 <br>\n<strong>ResNet34</strong>, MixUp, freq masking, slice=256 / 0.819 / 0.676 <br>\nResNet18, MixUp, freq masking, slice=256, <strong>MTL</strong> / 0.824 / 0.690 <br>\nResNet34, MixUp, freq masking, slice=256, MTL, <strong>snapshot ensemble cycle 1-4</strong> / 0.853 / 0.715 <br>\nResNet34, MixUp, freq masking, gain, slice=256, MTL, snapshot ensemble cycle <strong>1-8</strong> / 0.862 / 0.718 <br>\nResNet34 + <strong>EnvNet</strong> / 0.867 / 0.720 <br>\nResNet34, MixUp, freq masking, gain, <strong>slice=512</strong>, MTL, snapshot ensemble cycle 1-8 (#1)/ 0.860 / 0.731 <br>\n<strong>ResNet34 #1</strong> + EnvNet / 0.866 / 0.735 <br>\nResNet34 #1 + EnvNet, <strong>padding</strong> / 0.874 / 0.737 <br>\nResNet34 #1 + <strong>ResNet34 (SSL, #2)</strong> + <strong>EnvNet #4 + #5 + #6</strong> / 0.887 / 0.743 <br>\nResNet34 #1, #2 + EnvNet #4 + #5 + #6, <strong>padding TTA</strong> / 0.893 / 0.746 <br>\nResNet34 #1, #2 + <strong>ResNet34(slice=1024, #3)</strong> + EnvNet #4 + #5 + #6, padding TTA / 0.896 / 0.752  </p>",
  "messages": [
    {
      "id": "556685",
      "postDate": "06/20/2019 14:17:06",
      "content": "<p>We wrote up a technical report of our solution, so you can check our solution in detail when it's published. Here I'd like to describe a brief summary of our solution. Our score is public 0.752 (3rd)/private 0.75787 (4th before pending). <br>\nFor this competition, we used 3 strategies mainly. <br>\n1. Multitask learning with noisy labels. \n2. Semi-supervised learning (SSL) with noisy data. \n3. Averaging models trained with different time windows.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/556685/13619/Fig1_3.png\" alt=\"Multitask pipeline\"></p>\n\n<h2>Models</h2>\n\n<ul>\n<li>ResNet34 with log-mel</li>\n<li>EnvNet-v2 with waveform.</li>\n</ul>\n\n<h2>Preprocessing</h2>\n\n<ul>\n<li>128 mels</li>\n<li>128 Hz (347 STFT hop size)</li>\n</ul>\n\n<p>Log-mel is converted from power to dB after all augmentations applied. Thereafter, it is normalized by the mean and standard deviation of each data.</p>\n\n<h2>Augmentations</h2>\n\n<h3>Log-mel</h3>\n\n<ul>\n<li>Slicing</li>\n<li>MixUp</li>\n<li>Frequency masking</li>\n<li>Gain augmentation</li>\n</ul>\n\n<p>We tried 2, 4 and 8 seconds (256, 512 and 1024 dimensions) as a slicing-length and 4 seconds scores the best. Resizing, warping, time masking and white noise don't work.\n- Additional slicing</p>\n\n<p>Expecting more strong augmentation effect, after basic slicing, we shorten data samples in a range of 25 - 100% of the basic slicing-length by additional slicing and extend to the basic slicing-length by zero paddings. </p>\n\n<h3>Waveform</h3>\n\n<ul>\n<li>Slicing</li>\n<li>MixUp</li>\n<li>Gain augmentaion</li>\n<li>Scaling augmentation</li>\n</ul>\n\n<p>We tried 1.51, 3.02 and 4.54 seconds (66,650, 133,300 and 200,000 dimensions) as a slicing length and 4.54 seconds scores the best. </p>\n\n<h2>Training</h2>\n\n<h3>ResNet</h3>\n\n<ul>\n<li>Adam</li>\n<li>Cyclic cosine annealing (1e-3 -&gt; 1e-6)</li>\n<li>sigmoid and binary crossentropy</li>\n</ul>\n\n<h3>EnvNet</h3>\n\n<ul>\n<li>SGD</li>\n<li>Cyclic cosine annealing (1e-1 -&gt; 1e-6)</li>\n<li>(sigmoid and binary crossentropy) or (SoftMax and KL-divergence)</li>\n</ul>\n\n<h2>Postprocessing</h2>\n\n<p>Prediction using the full length of audio input scores better than prediction using slicing TTA. This may be because important components for classification is concentrated on the beginning part of audio samples. Actually, prediction with slices of the beginning part scores better than prediction with slices of the latter part.  We found padding augmentation is effective TTA. This is an augmentation method that applies zero paddings to both sides of audio samples with various length and averages prediction results. We think this method has an effect to emphasize the start and the end part of audio samples.</p>\n\n<h2>Multitask learning</h2>\n\n<p>In this task, the curated data and noisy data are labeled in a different manner, therefore treating them as the same one makes the model performance worse. To tackle this problem, we used a multitask learning (MTL) approach. The aim of MTL is to get synergy between 2 tasks without reducing the performance of each task. MTL learns features shared between 2 tasks and can be expected to achieve higher performance than learning independently. In our proposal, an encoder architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation. <br>\nMultitask module's components: <br>\n<code>FC(1024)-ReLU-DropOut(0.2)-FC(1024)-ReLU-DropOut(0.1)-FC(80)-sigmoid</code> <br>\nwe used BCE as a loss function. The loss weight ratio of curated and noisy is set as 1:1. By this method, CV lwlrap improved from 0.829 to 0.849 and score on the public LB increased + 0.021.</p>\n\n<h2>Semi-supervised learning</h2>\n\n<p>Applying SSL for this competition is difficult because of 2 reasons.\n1. Difference of data distribution between labeled data and unlabeled data. \n2. Multi-label classification task. </p>\n\n<p>In particular, the method that generates labels online like Mean teacher or MixMatch tends to collapse.  We tried pseudo label, Mean Teacher and MixMatch and all of them are not successful. <br>\nTherefore, we propose an SSL method that is robust to data distribution difference and can handle multi-label data.  For each noisy data sample, we guess the label using the trained model. The guessed label is sharpened by sharpening function proposed by MixMatch (soft pseudo label).  As the temperature of sharpening function, we tried a value of 1, 1.5 or 2 and 2 was the best.  Predictions of the trained model are obtained using snapshot ensemble with all the folds and cycle snapshots of 5-fold CV.  We used MSE as a loss function. we set loss weight of semi-supervised learning as 20. <br>\nBy soft pseudo labeling, The CV lwlrap improved from 0.849 to 0.870. On the other hand, on the public LB, improvement in score was slight (+0.001). We used predictions of all fold models to generate soft pseudo label so that high CV may be because of indirect label leak. However, even if we use labels generated by only the same fold model which has no label leak, CV was improved as compared to one without SSL. The model trained with the soft pseudo label is useful as a component of model averaging (+0.003 on the public LB).</p>\n\n<h2>Ensemble</h2>\n\n<p>For model averaging, we prepared models trained with various conditions. Especially, we found that averaging with models trained with different time window is effective. We used models below. In order to reduce prediction time,  the cycles and padding lengths used for the final submissions and averaging weights were chosen based on CV. \n1. ResNet34 slice=512, MTL\n2. ResNet34 slice=512, MTL, SSL\n3. ResNet34 slice=1024, MTL\n4. EnvNet-v2 slice=133,300, MTL, sigmoid\n5. EnvNet-v2 slice=133,300, MTL, SoftMax\n6. EnvNet-v2 slice=200,000, MTL, SoftMax</p>\n\n<h2>Comparison</h2>\n\n<p>condition/CV lwlrap <br>\nmodel #1 / 0.868 <br>\nmodel #2 / 0.886 <br>\nmodel #3 / 0.862 <br>\nmodel #4 / 0.815 <br>\nmodel #5 / 0.818 <br>\nmodel #6 / 0.820 <br>\n#1 + #3 / 0.876 <br>\n#1 + #2 + #3 / 0.890 <br>\n#4 + #5 + #6 / 0.836 <br>\nsubmission 1 (#1 + #2 + #3 + #4 + #5 + #6) / 0.896 <br>\nsubmission 2 (no pad TTA) / 0.895  </p>\n\n<h2>Code</h2>\n\n<p>I did all model training on kaggle kernels. So I can share all of them. Please note that code readability is low. For example, DataLoader class is named \"MFCCDataset\" but it actually loads log-mel.\n<a href=\"https://www.kaggle.com/osciiart/freesound2019-solution-links?scriptVersionId=16047723\">link to the summary of codes</a> <br>\nThe GitHub repository is available at <a href=\"https://github.com/OsciiArt/Freesound-Audio-Tagging-2019\">here</a>.   </p>\n\n<h2>Public LB history log</h2>\n\n<p>model / CV / Public LB <br>\nResNet18, MixUp, freq masking, slice=256 / 0.819 / 0.669 <br>\n<strong>ResNet34</strong>, MixUp, freq masking, slice=256 / 0.819 / 0.676 <br>\nResNet18, MixUp, freq masking, slice=256, <strong>MTL</strong> / 0.824 / 0.690 <br>\nResNet34, MixUp, freq masking, slice=256, MTL, <strong>snapshot ensemble cycle 1-4</strong> / 0.853 / 0.715 <br>\nResNet34, MixUp, freq masking, gain, slice=256, MTL, snapshot ensemble cycle <strong>1-8</strong> / 0.862 / 0.718 <br>\nResNet34 + <strong>EnvNet</strong> / 0.867 / 0.720 <br>\nResNet34, MixUp, freq masking, gain, <strong>slice=512</strong>, MTL, snapshot ensemble cycle 1-8 (#1)/ 0.860 / 0.731 <br>\n<strong>ResNet34 #1</strong> + EnvNet / 0.866 / 0.735 <br>\nResNet34 #1 + EnvNet, <strong>padding</strong> / 0.874 / 0.737 <br>\nResNet34 #1 + <strong>ResNet34 (SSL, #2)</strong> + <strong>EnvNet #4 + #5 + #6</strong> / 0.887 / 0.743 <br>\nResNet34 #1, #2 + EnvNet #4 + #5 + #6, <strong>padding TTA</strong> / 0.893 / 0.746 <br>\nResNet34 #1, #2 + <strong>ResNet34(slice=1024, #3)</strong> + EnvNet #4 + #5 + #6, padding TTA / 0.896 / 0.752  </p>",
      "rawMarkdown": "We wrote up a technical report of our solution, so you can check our solution in detail when it's published. Here I'd like to describe a brief summary of our solution. Our score is public 0.752 (3rd)/private 0.75787 (4th before pending).  \nFor this competition, we used 3 strategies mainly.  \n1. Multitask learning with noisy labels. \n2. Semi-supervised learning (SSL) with noisy data. \n3. Averaging models trained with different time windows.\n![Multitask pipeline](https://storage.googleapis.com/kaggle-forum-message-attachments/556685/13619/Fig1_3.png)\n\n## Models\n- ResNet34 with log-mel\n- EnvNet-v2 with waveform.\n\n## Preprocessing\n- 128 mels\n- 128 Hz (347 STFT hop size)\n \nLog-mel is converted from power to dB after all augmentations applied. Thereafter, it is normalized by the mean and standard deviation of each data.\n\n## Augmentations\n### Log-mel\n- Slicing\n- MixUp\n- Frequency masking\n- Gain augmentation\n\nWe tried 2, 4 and 8 seconds (256, 512 and 1024 dimensions) as a slicing-length and 4 seconds scores the best. Resizing, warping, time masking and white noise don't work.\n- Additional slicing\n\nExpecting more strong augmentation effect, after basic slicing, we shorten data samples in a range of 25 - 100% of the basic slicing-length by additional slicing and extend to the basic slicing-length by zero paddings. \n  \n### Waveform\n- Slicing\n- MixUp\n- Gain augmentaion\n- Scaling augmentation\n\nWe tried 1.51, 3.02 and 4.54 seconds (66,650, 133,300 and 200,000 dimensions) as a slicing length and 4.54 seconds scores the best. \n\n## Training\n### ResNet\n- Adam\n- Cyclic cosine annealing (1e-3 -&gt; 1e-6)\n- sigmoid and binary crossentropy\n\n### EnvNet\n- SGD\n- Cyclic cosine annealing (1e-1 -&gt; 1e-6)\n- (sigmoid and binary crossentropy) or (SoftMax and KL-divergence)\n\n## Postprocessing\nPrediction using the full length of audio input scores better than prediction using slicing TTA. This may be because important components for classification is concentrated on the beginning part of audio samples. Actually, prediction with slices of the beginning part scores better than prediction with slices of the latter part.  We found padding augmentation is effective TTA. This is an augmentation method that applies zero paddings to both sides of audio samples with various length and averages prediction results. We think this method has an effect to emphasize the start and the end part of audio samples.\n\n## Multitask learning\nIn this task, the curated data and noisy data are labeled in a different manner, therefore treating them as the same one makes the model performance worse. To tackle this problem, we used a multitask learning (MTL) approach. The aim of MTL is to get synergy between 2 tasks without reducing the performance of each task. MTL learns features shared between 2 tasks and can be expected to achieve higher performance than learning independently. In our proposal, an encoder architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation.   \nMultitask module's components:  \n`FC(1024)-ReLU-DropOut(0.2)-FC(1024)-ReLU-DropOut(0.1)-FC(80)-sigmoid`  \nwe used BCE as a loss function. The loss weight ratio of curated and noisy is set as 1:1. By this method, CV lwlrap improved from 0.829 to 0.849 and score on the public LB increased + 0.021.\n\n## Semi-supervised learning\nApplying SSL for this competition is difficult because of 2 reasons.\n1. Difference of data distribution between labeled data and unlabeled data. \n2. Multi-label classification task. \n\nIn particular, the method that generates labels online like Mean teacher or MixMatch tends to collapse.  We tried pseudo label, Mean Teacher and MixMatch and all of them are not successful.  \nTherefore, we propose an SSL method that is robust to data distribution difference and can handle multi-label data.  For each noisy data sample, we guess the label using the trained model. The guessed label is sharpened by sharpening function proposed by MixMatch (soft pseudo label).  As the temperature of sharpening function, we tried a value of 1, 1.5 or 2 and 2 was the best.  Predictions of the trained model are obtained using snapshot ensemble with all the folds and cycle snapshots of 5-fold CV.  We used MSE as a loss function. we set loss weight of semi-supervised learning as 20.  \nBy soft pseudo labeling, The CV lwlrap improved from 0.849 to 0.870. On the other hand, on the public LB, improvement in score was slight (+0.001). We used predictions of all fold models to generate soft pseudo label so that high CV may be because of indirect label leak. However, even if we use labels generated by only the same fold model which has no label leak, CV was improved as compared to one without SSL. The model trained with the soft pseudo label is useful as a component of model averaging (+0.003 on the public LB).\n\n\n## Ensemble\nFor model averaging, we prepared models trained with various conditions. Especially, we found that averaging with models trained with different time window is effective. We used models below. In order to reduce prediction time,  the cycles and padding lengths used for the final submissions and averaging weights were chosen based on CV. \n1. ResNet34 slice=512, MTL\n2. ResNet34 slice=512, MTL, SSL\n3. ResNet34 slice=1024, MTL\n4. EnvNet-v2 slice=133,300, MTL, sigmoid\n5. EnvNet-v2 slice=133,300, MTL, SoftMax\n6. EnvNet-v2 slice=200,000, MTL, SoftMax\n\n## Comparison\ncondition/CV lwlrap  \nmodel #1 / 0.868  \nmodel #2 / 0.886  \nmodel #3 / 0.862  \nmodel #4 / 0.815  \nmodel #5 / 0.818  \nmodel #6 / 0.820  \n\\#1 + #3 / 0.876  \n\\#1 + #2 + #3 / 0.890  \n\\#4 + #5 + #6 / 0.836  \nsubmission 1 (#1 + #2 + #3 + #4 + #5 + #6) / 0.896  \nsubmission 2 (no pad TTA) / 0.895  \n\n\n## Code\nI did all model training on kaggle kernels. So I can share all of them. Please note that code readability is low. For example, DataLoader class is named \"MFCCDataset\" but it actually loads log-mel.\n[link to the summary of codes](https://www.kaggle.com/osciiart/freesound2019-solution-links?scriptVersionId=16047723)  \nThe GitHub repository is available at [here](https://github.com/OsciiArt/Freesound-Audio-Tagging-2019).   \n\n## Public LB history log\nmodel / CV / Public LB  \nResNet18, MixUp, freq masking, slice=256 / 0.819 / 0.669  \n**ResNet34**, MixUp, freq masking, slice=256 / 0.819 / 0.676  \nResNet18, MixUp, freq masking, slice=256, **MTL** / 0.824 / 0.690  \nResNet34, MixUp, freq masking, slice=256, MTL, **snapshot ensemble cycle 1-4** / 0.853 / 0.715  \nResNet34, MixUp, freq masking, gain, slice=256, MTL, snapshot ensemble cycle **1-8** / 0.862 / 0.718  \nResNet34 + **EnvNet** / 0.867 / 0.720  \nResNet34, MixUp, freq masking, gain, **slice=512**, MTL, snapshot ensemble cycle 1-8 (#1)/ 0.860 / 0.731  \n**ResNet34 #1** + EnvNet / 0.866 / 0.735  \nResNet34 #1 + EnvNet, **padding** / 0.874 / 0.737  \nResNet34 #1 + **ResNet34 (SSL, #2)** + **EnvNet #4 + #5 + #6** / 0.887 / 0.743  \nResNet34 #1, #2 + EnvNet #4 + #5 + #6, **padding TTA** / 0.893 / 0.746  \nResNet34 #1, #2 + **ResNet34(slice=1024, #3)** + EnvNet #4 + #5 + #6, padding TTA / 0.896 / 0.752",
      "votes": null
    },
    {
      "id": "556888",
      "postDate": "06/20/2019 20:09:27",
      "content": "<p>So a modified version of MixMatch works? Just quick glance, in your model #2,  you sharpenened the output and you used softmax to get the multi-class posterior probabilities in the mse part. Thanks for sharing and congrats!</p>\n\n<p>Edit: I see. So the difference is that you used a trained model to generate pseudo-labels.</p>",
      "rawMarkdown": "So a modified version of MixMatch works? Just quick glance, in your model #2,  you sharpenened the output and you used softmax to get the multi-class posterior probabilities in the mse part. Thanks for sharing and congrats!\n\nEdit: I see. So the difference is that you used a trained model to generate pseudo-labels.",
      "votes": null
    },
    {
      "id": "556932",
      "postDate": "06/20/2019 22:32:05",
      "content": "<p>In MixMatch, guessed labels are made by TTA (slicing and flip) of a training model (not a trained model). This is not suit for this competition data because TTA is not so effective like image data. </p>",
      "rawMarkdown": "In MixMatch, guessed labels are made by TTA (slicing and flip) of a training model (not a trained model). This is not suit for this competition data because TTA is not so effective like image data.",
      "votes": null
    },
    {
      "id": "559341",
      "postDate": "06/24/2019 00:14:10",
      "content": "<p>what a great work! Thank you for sharing!</p>",
      "rawMarkdown": "what a great work! Thank you for sharing!",
      "votes": null
    },
    {
      "id": "559385",
      "postDate": "06/24/2019 03:56:12",
      "content": "<p>Thanks <a href=\"/osciiart\">@osciiart</a> -san for this great art of work! (please let me slowly digest your masterpieces)</p>\n\n<p>BTW, in the summary link, <a href=\"https://www.kaggle.com/osciiart/resnet34-mel-ver3-log-multi-hardaug\">https://www.kaggle.com/osciiart/resnet34-mel-ver3-log-multi-hardaug</a> seems not available.</p>",
      "rawMarkdown": "Thanks @osciiart -san for this great art of work! (please let me slowly digest your masterpieces)\n\nBTW, in the summary link, https://www.kaggle.com/osciiart/resnet34-mel-ver3-log-multi-hardaug seems not available.",
      "votes": null
    },
    {
      "id": "559624",
      "postDate": "06/24/2019 10:40:48",
      "content": "<p>Thank you for letting me know. I made it public.</p>",
      "rawMarkdown": "Thank you for letting me know. I made it public.",
      "votes": null
    },
    {
      "id": "559651",
      "postDate": "06/24/2019 11:34:45",
      "content": "<p>Nice write up!</p>",
      "rawMarkdown": "Nice write up!",
      "votes": null
    },
    {
      "id": "561393",
      "postDate": "06/26/2019 11:57:54",
      "content": "<p>Great work, lot of interesting ideas, thanks for sharing! </p>",
      "rawMarkdown": "Great work, lot of interesting ideas, thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 556888,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "06/20/2019 20:09:27",
      "content": "<p>So a modified version of MixMatch works? Just quick glance, in your model #2,  you sharpenened the output and you used softmax to get the multi-class posterior probabilities in the mse part. Thanks for sharing and congrats!</p>\n\n<p>Edit: I see. So the difference is that you used a trained model to generate pseudo-labels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 556932,
          "author_name": "osciiart",
          "author_url": "",
          "post_date": "06/20/2019 22:32:05",
          "content": "<p>In MixMatch, guessed labels are made by TTA (slicing and flip) of a training model (not a trained model). This is not suit for this competition data because TTA is not so effective like image data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 559341,
      "author_name": "junyun1002",
      "author_url": "",
      "post_date": "06/24/2019 00:14:10",
      "content": "<p>what a great work! Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 559385,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "06/24/2019 03:56:12",
      "content": "<p>Thanks <a href=\"/osciiart\">@osciiart</a> -san for this great art of work! (please let me slowly digest your masterpieces)</p>\n\n<p>BTW, in the summary link, <a href=\"https://www.kaggle.com/osciiart/resnet34-mel-ver3-log-multi-hardaug\">https://www.kaggle.com/osciiart/resnet34-mel-ver3-log-multi-hardaug</a> seems not available.</p>",
      "votes": null,
      "replies": [
        {
          "id": 559624,
          "author_name": "osciiart",
          "author_url": "",
          "post_date": "06/24/2019 10:40:48",
          "content": "<p>Thank you for letting me know. I made it public.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 559651,
      "author_name": "klirfan",
      "author_url": "",
      "post_date": "06/24/2019 11:34:45",
      "content": "<p>Nice write up!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 561393,
      "author_name": "hlerogeron",
      "author_url": "",
      "post_date": "06/26/2019 11:57:54",
      "content": "<p>Great work, lot of interesting ideas, thanks for sharing! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "556685": "We wrote up a technical report of our solution, so you can check our solution in detail when it's published. Here I'd like to describe a brief summary of our solution. Our score is public 0.752 (3rd)/private 0.75787 (4th before pending).  \nFor this competition, we used 3 strategies mainly.  \n1. Multitask learning with noisy labels. \n2. Semi-supervised learning (SSL) with noisy data. \n3. Averaging models trained with different time windows.\n![Multitask pipeline](https://storage.googleapis.com/kaggle-forum-message-attachments/556685/13619/Fig1_3.png)\n\n## Models\n- ResNet34 with log-mel\n- EnvNet-v2 with waveform.\n\n## Preprocessing\n- 128 mels\n- 128 Hz (347 STFT hop size)\n \nLog-mel is converted from power to dB after all augmentations applied. Thereafter, it is normalized by the mean and standard deviation of each data.\n\n## Augmentations\n### Log-mel\n- Slicing\n- MixUp\n- Frequency masking\n- Gain augmentation\n\nWe tried 2, 4 and 8 seconds (256, 512 and 1024 dimensions) as a slicing-length and 4 seconds scores the best. Resizing, warping, time masking and white noise don't work.\n- Additional slicing\n\nExpecting more strong augmentation effect, after basic slicing, we shorten data samples in a range of 25 - 100% of the basic slicing-length by additional slicing and extend to the basic slicing-length by zero paddings. \n  \n### Waveform\n- Slicing\n- MixUp\n- Gain augmentaion\n- Scaling augmentation\n\nWe tried 1.51, 3.02 and 4.54 seconds (66,650, 133,300 and 200,000 dimensions) as a slicing length and 4.54 seconds scores the best. \n\n## Training\n### ResNet\n- Adam\n- Cyclic cosine annealing (1e-3 -&gt; 1e-6)\n- sigmoid and binary crossentropy\n\n### EnvNet\n- SGD\n- Cyclic cosine annealing (1e-1 -&gt; 1e-6)\n- (sigmoid and binary crossentropy) or (SoftMax and KL-divergence)\n\n## Postprocessing\nPrediction using the full length of audio input scores better than prediction using slicing TTA. This may be because important components for classification is concentrated on the beginning part of audio samples. Actually, prediction with slices of the beginning part scores better than prediction with slices of the latter part.  We found padding augmentation is effective TTA. This is an augmentation method that applies zero paddings to both sides of audio samples with various length and averages prediction results. We think this method has an effect to emphasize the start and the end part of audio samples.\n\n## Multitask learning\nIn this task, the curated data and noisy data are labeled in a different manner, therefore treating them as the same one makes the model performance worse. To tackle this problem, we used a multitask learning (MTL) approach. The aim of MTL is to get synergy between 2 tasks without reducing the performance of each task. MTL learns features shared between 2 tasks and can be expected to achieve higher performance than learning independently. In our proposal, an encoder architecture learns the features shared between curated and noisy data, and the two separated FC layers learn the difference between the two data. In this way, we can get the advantages of feature learning from noisy data and avoid the disadvantages of noisy label perturbation.   \nMultitask module's components:  \n`FC(1024)-ReLU-DropOut(0.2)-FC(1024)-ReLU-DropOut(0.1)-FC(80)-sigmoid`  \nwe used BCE as a loss function. The loss weight ratio of curated and noisy is set as 1:1. By this method, CV lwlrap improved from 0.829 to 0.849 and score on the public LB increased + 0.021.\n\n## Semi-supervised learning\nApplying SSL for this competition is difficult because of 2 reasons.\n1. Difference of data distribution between labeled data and unlabeled data. \n2. Multi-label classification task. \n\nIn particular, the method that generates labels online like Mean teacher or MixMatch tends to collapse.  We tried pseudo label, Mean Teacher and MixMatch and all of them are not successful.  \nTherefore, we propose an SSL method that is robust to data distribution difference and can handle multi-label data.  For each noisy data sample, we guess the label using the trained model. The guessed label is sharpened by sharpening function proposed by MixMatch (soft pseudo label).  As the temperature of sharpening function, we tried a value of 1, 1.5 or 2 and 2 was the best.  Predictions of the trained model are obtained using snapshot ensemble with all the folds and cycle snapshots of 5-fold CV.  We used MSE as a loss function. we set loss weight of semi-supervised learning as 20.  \nBy soft pseudo labeling, The CV lwlrap improved from 0.849 to 0.870. On the other hand, on the public LB, improvement in score was slight (+0.001). We used predictions of all fold models to generate soft pseudo label so that high CV may be because of indirect label leak. However, even if we use labels generated by only the same fold model which has no label leak, CV was improved as compared to one without SSL. The model trained with the soft pseudo label is useful as a component of model averaging (+0.003 on the public LB).\n\n\n## Ensemble\nFor model averaging, we prepared models trained with various conditions. Especially, we found that averaging with models trained with different time window is effective. We used models below. In order to reduce prediction time,  the cycles and padding lengths used for the final submissions and averaging weights were chosen based on CV. \n1. ResNet34 slice=512, MTL\n2. ResNet34 slice=512, MTL, SSL\n3. ResNet34 slice=1024, MTL\n4. EnvNet-v2 slice=133,300, MTL, sigmoid\n5. EnvNet-v2 slice=133,300, MTL, SoftMax\n6. EnvNet-v2 slice=200,000, MTL, SoftMax\n\n## Comparison\ncondition/CV lwlrap  \nmodel #1 / 0.868  \nmodel #2 / 0.886  \nmodel #3 / 0.862  \nmodel #4 / 0.815  \nmodel #5 / 0.818  \nmodel #6 / 0.820  \n\\#1 + #3 / 0.876  \n\\#1 + #2 + #3 / 0.890  \n\\#4 + #5 + #6 / 0.836  \nsubmission 1 (#1 + #2 + #3 + #4 + #5 + #6) / 0.896  \nsubmission 2 (no pad TTA) / 0.895  \n\n\n## Code\nI did all model training on kaggle kernels. So I can share all of them. Please note that code readability is low. For example, DataLoader class is named \"MFCCDataset\" but it actually loads log-mel.\n[link to the summary of codes](https://www.kaggle.com/osciiart/freesound2019-solution-links?scriptVersionId=16047723)  \nThe GitHub repository is available at [here](https://github.com/OsciiArt/Freesound-Audio-Tagging-2019).   \n\n## Public LB history log\nmodel / CV / Public LB  \nResNet18, MixUp, freq masking, slice=256 / 0.819 / 0.669  \n**ResNet34**, MixUp, freq masking, slice=256 / 0.819 / 0.676  \nResNet18, MixUp, freq masking, slice=256, **MTL** / 0.824 / 0.690  \nResNet34, MixUp, freq masking, slice=256, MTL, **snapshot ensemble cycle 1-4** / 0.853 / 0.715  \nResNet34, MixUp, freq masking, gain, slice=256, MTL, snapshot ensemble cycle **1-8** / 0.862 / 0.718  \nResNet34 + **EnvNet** / 0.867 / 0.720  \nResNet34, MixUp, freq masking, gain, **slice=512**, MTL, snapshot ensemble cycle 1-8 (#1)/ 0.860 / 0.731  \n**ResNet34 #1** + EnvNet / 0.866 / 0.735  \nResNet34 #1 + EnvNet, **padding** / 0.874 / 0.737  \nResNet34 #1 + **ResNet34 (SSL, #2)** + **EnvNet #4 + #5 + #6** / 0.887 / 0.743  \nResNet34 #1, #2 + EnvNet #4 + #5 + #6, **padding TTA** / 0.893 / 0.746  \nResNet34 #1, #2 + **ResNet34(slice=1024, #3)** + EnvNet #4 + #5 + #6, padding TTA / 0.896 / 0.752",
    "556888": "So a modified version of MixMatch works? Just quick glance, in your model #2,  you sharpenened the output and you used softmax to get the multi-class posterior probabilities in the mse part. Thanks for sharing and congrats!\n\nEdit: I see. So the difference is that you used a trained model to generate pseudo-labels.",
    "556932": "In MixMatch, guessed labels are made by TTA (slicing and flip) of a training model (not a trained model). This is not suit for this competition data because TTA is not so effective like image data.",
    "559341": "what a great work! Thank you for sharing!",
    "559385": "Thanks @osciiart -san for this great art of work! (please let me slowly digest your masterpieces)\n\nBTW, in the summary link, https://www.kaggle.com/osciiart/resnet34-mel-ver3-log-multi-hardaug seems not available.",
    "559624": "Thank you for letting me know. I made it public.",
    "559651": "Nice write up!",
    "561393": "Great work, lot of interesting ideas, thanks for sharing!"
  },
  "source": "meta"
}