{
  "id": 97829,
  "title": "Semi-supervised part of 20th solution",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/97829",
  "author_name": "Darkate",
  "post_date": "2019-06-29T05:20:52.428000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First I want to thank researchers at CVSSP, University of Surrey, UK for their great baseline <a href=\"https://github.com/qiuqiangkong/dcase2019_task2\">code</a>. My entire system is built upon it.</p>\n\n<p><strong>Pre-processing</strong>\nNothing special. 32 kHz sampling rate, 500 hop size so that 64 fps. Silence trimming below <code>max - 55 dB</code>, mel size 128. Random excerpt 4-second patches from entire spectrograms. Shorter spectrograms are padded repeatedly.</p>\n\n<p><strong>Data augmentation</strong>\n1. Spec-augment without time warping\n2. Time reversal with 80 more classes (binary target with 160 classes in total)\n3. I'll talk about MixUp later</p>\n\n<p><strong>Semi-supervised learning</strong>\nAlmost the same as <a href=\"https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26\">https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26</a>. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2Fcc2a1c6259b8b72194abef259d524448%2Ffreesound_model.png?generation=1561783764472053&amp;alt=media\" alt=\"\"></p>\n\n<p>Suppose V X y_V is a batch of curated data and W is a batch of noisy data without label, then we have\n![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2F23d2ce6b7bb8c8d60b6e5577b09c32d1%2Ffreesound_algo.png?generation=1561784302344628&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2F23d2ce6b7bb8c8d60b6e5577b09c32d1%2Ffreesound_algo.png?generation=1561784302344628&amp;alt=media</a> =400x320)</p>\n\n<p>The backbone is 9-layer CNN, which is the same as the one in public kernel and CVSSP baseline.</p>\n\n<p><strong>Loss function</strong>\nFocal + ArcFace.\nArcFace is modified for mult-label classification. 2 reasons to use ArcFace here\n1. Enlarge the margin between positive and negative, so that the threshold eta becomes less sensitive.\n2. In the <a href=\"https://arxiv.org/abs/1801.07698\">original ArcFace paper</a>, the authors state that ArcFace enforces less penalty towards samples whose feature vectors have large angle with the weight matrix than other margin-based loss. I suppose it would apply less \"weight\" to wrong pseudo-labels.</p>\n\n<p>In practice using ArcFace makes my NN converge faster and have higher local lwlrap when using only curated data. I don't know what would happen if I trained the semi-supervised model without it because I started using it at a very early stage.</p>\n\n<p><strong>Training Details</strong>\nFirst I warm  up the backbone CNN with noisy set. For the semi-supervised model, I have 6 samples from curated set and 58 samples from noisy set in each batch. Nadam optimizer, cosine annealing learning rate with linear warm-up are used. One can refer to my technical report, Table 2 for details about hyperparameters.</p>\n\n<p><strong>Inference</strong>\nArithmetic mean on 0.25s sliding window. TTA only on clips with short duration by repeating the sliding window with Spec-augment. </p>\n\n<p>Single model trained on entire training set gives me 0.712 lwlrap on stage 1 test.</p>\n\n<p>Again, thank the host for hosting such interesting competition. I am still on my way to my first gold medal :)</p>",
  "messages": [
    {
      "id": 564187,
      "postDate": "2019-06-29T05:20:52.430Z",
      "content": "<p>First I want to thank researchers at CVSSP, University of Surrey, UK for their great baseline <a href=\"https://github.com/qiuqiangkong/dcase2019_task2\">code</a>. My entire system is built upon it.</p>\n\n<p><strong>Pre-processing</strong>\nNothing special. 32 kHz sampling rate, 500 hop size so that 64 fps. Silence trimming below <code>max - 55 dB</code>, mel size 128. Random excerpt 4-second patches from entire spectrograms. Shorter spectrograms are padded repeatedly.</p>\n\n<p><strong>Data augmentation</strong>\n1. Spec-augment without time warping\n2. Time reversal with 80 more classes (binary target with 160 classes in total)\n3. I'll talk about MixUp later</p>\n\n<p><strong>Semi-supervised learning</strong>\nAlmost the same as <a href=\"https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26\">https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26</a>. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2Fcc2a1c6259b8b72194abef259d524448%2Ffreesound_model.png?generation=1561783764472053&amp;alt=media\" alt=\"\"></p>\n\n<p>Suppose V X y_V is a batch of curated data and W is a batch of noisy data without label, then we have\n![](<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2F23d2ce6b7bb8c8d60b6e5577b09c32d1%2Ffreesound_algo.png?generation=1561784302344628&amp;alt=media\">https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2F23d2ce6b7bb8c8d60b6e5577b09c32d1%2Ffreesound_algo.png?generation=1561784302344628&amp;alt=media</a> =400x320)</p>\n\n<p>The backbone is 9-layer CNN, which is the same as the one in public kernel and CVSSP baseline.</p>\n\n<p><strong>Loss function</strong>\nFocal + ArcFace.\nArcFace is modified for mult-label classification. 2 reasons to use ArcFace here\n1. Enlarge the margin between positive and negative, so that the threshold eta becomes less sensitive.\n2. In the <a href=\"https://arxiv.org/abs/1801.07698\">original ArcFace paper</a>, the authors state that ArcFace enforces less penalty towards samples whose feature vectors have large angle with the weight matrix than other margin-based loss. I suppose it would apply less \"weight\" to wrong pseudo-labels.</p>\n\n<p>In practice using ArcFace makes my NN converge faster and have higher local lwlrap when using only curated data. I don't know what would happen if I trained the semi-supervised model without it because I started using it at a very early stage.</p>\n\n<p><strong>Training Details</strong>\nFirst I warm  up the backbone CNN with noisy set. For the semi-supervised model, I have 6 samples from curated set and 58 samples from noisy set in each batch. Nadam optimizer, cosine annealing learning rate with linear warm-up are used. One can refer to my technical report, Table 2 for details about hyperparameters.</p>\n\n<p><strong>Inference</strong>\nArithmetic mean on 0.25s sliding window. TTA only on clips with short duration by repeating the sliding window with Spec-augment. </p>\n\n<p>Single model trained on entire training set gives me 0.712 lwlrap on stage 1 test.</p>\n\n<p>Again, thank the host for hosting such interesting competition. I am still on my way to my first gold medal :)</p>",
      "rawMarkdown": "First I want to thank researchers at CVSSP, University of Surrey, UK for their great baseline [code](https://github.com/qiuqiangkong/dcase2019_task2). My entire system is built upon it.\n\n**Pre-processing**\nNothing special. 32 kHz sampling rate, 500 hop size so that 64 fps. Silence trimming below `max - 55 dB`, mel size 128. Random excerpt 4-second patches from entire spectrograms. Shorter spectrograms are padded repeatedly.\n\n**Data augmentation**\n1. Spec-augment without time warping\n2. Time reversal with 80 more classes (binary target with 160 classes in total)\n3. I'll talk about MixUp later\n\n**Semi-supervised learning**\nAlmost the same as https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2Fcc2a1c6259b8b72194abef259d524448%2Ffreesound_model.png?generation=1561783764472053&amp;alt=media)\n\nSuppose V X y\\_V is a batch of curated data and W is a batch of noisy data without label, then we have\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2F23d2ce6b7bb8c8d60b6e5577b09c32d1%2Ffreesound_algo.png?generation=1561784302344628&amp;alt=media =400x320)\n\nThe backbone is 9-layer CNN, which is the same as the one in public kernel and CVSSP baseline.\n\n**Loss function**\nFocal + ArcFace.\nArcFace is modified for mult-label classification. 2 reasons to use ArcFace here\n1. Enlarge the margin between positive and negative, so that the threshold eta becomes less sensitive.\n2. In the [original ArcFace paper](https://arxiv.org/abs/1801.07698), the authors state that ArcFace enforces less penalty towards samples whose feature vectors have large angle with the weight matrix than other margin-based loss. I suppose it would apply less \"weight\" to wrong pseudo-labels.\n\nIn practice using ArcFace makes my NN converge faster and have higher local lwlrap when using only curated data. I don't know what would happen if I trained the semi-supervised model without it because I started using it at a very early stage.\n\n**Training Details**\nFirst I warm  up the backbone CNN with noisy set. For the semi-supervised model, I have 6 samples from curated set and 58 samples from noisy set in each batch. Nadam optimizer, cosine annealing learning rate with linear warm-up are used. One can refer to my technical report, Table 2 for details about hyperparameters.\n\n**Inference**\nArithmetic mean on 0.25s sliding window. TTA only on clips with short duration by repeating the sliding window with Spec-augment. \n\nSingle model trained on entire training set gives me 0.712 lwlrap on stage 1 test.\n\nAgain, thank the host for hosting such interesting competition. I am still on my way to my first gold medal :)",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "564187": "First I want to thank researchers at CVSSP, University of Surrey, UK for their great baseline [code](https://github.com/qiuqiangkong/dcase2019_task2). My entire system is built upon it.\n\n**Pre-processing**\nNothing special. 32 kHz sampling rate, 500 hop size so that 64 fps. Silence trimming below `max - 55 dB`, mel size 128. Random excerpt 4-second patches from entire spectrograms. Shorter spectrograms are padded repeatedly.\n\n**Data augmentation**\n1. Spec-augment without time warping\n2. Time reversal with 80 more classes (binary target with 160 classes in total)\n3. I'll talk about MixUp later\n\n**Semi-supervised learning**\nAlmost the same as https://link.springer.com/chapter/10.1007/978-3-030-20873-8_26. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2Fcc2a1c6259b8b72194abef259d524448%2Ffreesound_model.png?generation=1561783764472053&amp;alt=media)\n\nSuppose V X y\\_V is a batch of curated data and W is a batch of noisy data without label, then we have\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1605499%2F23d2ce6b7bb8c8d60b6e5577b09c32d1%2Ffreesound_algo.png?generation=1561784302344628&amp;alt=media =400x320)\n\nThe backbone is 9-layer CNN, which is the same as the one in public kernel and CVSSP baseline.\n\n**Loss function**\nFocal + ArcFace.\nArcFace is modified for mult-label classification. 2 reasons to use ArcFace here\n1. Enlarge the margin between positive and negative, so that the threshold eta becomes less sensitive.\n2. In the [original ArcFace paper](https://arxiv.org/abs/1801.07698), the authors state that ArcFace enforces less penalty towards samples whose feature vectors have large angle with the weight matrix than other margin-based loss. I suppose it would apply less \"weight\" to wrong pseudo-labels.\n\nIn practice using ArcFace makes my NN converge faster and have higher local lwlrap when using only curated data. I don't know what would happen if I trained the semi-supervised model without it because I started using it at a very early stage.\n\n**Training Details**\nFirst I warm  up the backbone CNN with noisy set. For the semi-supervised model, I have 6 samples from curated set and 58 samples from noisy set in each batch. Nadam optimizer, cosine annealing learning rate with linear warm-up are used. One can refer to my technical report, Table 2 for details about hyperparameters.\n\n**Inference**\nArithmetic mean on 0.25s sliding window. TTA only on clips with short duration by repeating the sliding window with Spec-augment. \n\nSingle model trained on entire training set gives me 0.712 lwlrap on stage 1 test.\n\nAgain, thank the host for hosting such interesting competition. I am still on my way to my first gold medal :)"
  }
}