{
  "id": 220432,
  "title": "5th place solution (Training Strategy)",
  "url": "/competitions/rfcx-species-audio-detection/writeups/birdcall-revenge-5th-place-solution-training-strat",
  "author_name": "",
  "post_date": "2021-02-18T09:43:58.523689Z",
  "votes": 40,
  "comment_count": 14,
  "views": 0,
  "content": "<h1>5th place solution (Training Strategy)</h1>\n<p>Congratulations to all the participants, and thanks a lot to the organizers for this competition! This has been a very difficult but fun competition:)</p>\n<p>In this thread, I introduce our approach about training strategy.<br>\nAbout ensemble part will be written by my team member.</p>\n<p>Our team ensemble each best model.</p>\n<p>My model is Resnet18 which has a SED header. This model's LWLRAP is Public LB=0.949 /Private LB=0.951, and I trained by google colab using <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198048\" target=\"_blank\">Theo Viel's npz dataset</a>(32 kHz, 128 mels). Thank you, Theo Viel!!</p>\n<p>Our approach has 3 stage, <br>\nOther team members are different in some things likes the base model and hyperparameter, but these default strategies are about the same.</p>\n<h2>1st stage: pre-train</h2>\n<p><em>※I think this part is not important. Team member Ahmet skips this part.</em></p>\n<p>This stage transfers learning from Imagenet to spectograms.</p>\n<p>Theo Viel's npz dataset can be regarded as 128x3751 size image.<br>\nI cut to 512 by this image in sound point from t_min and t_max.<br>\nI train this image by tp_train and 30 sampled fp_train.</p>\n<p>Parameters:</p>\n<ul>\n<li>Adam</li>\n<li>learning_rate=1e-3</li>\n<li>CosineAnnealingLR(max_T=10)</li>\n<li>epoch=50</li>\n</ul>\n<p>Continue 2nd and 3rd stage use this trained weight.</p>\n<h2>2nd stage: pseudo label re-labeling</h2>\n<p>The purpose of stage2 is to improve the model and make pseudo labels by this model.</p>\n<p>Use 1st stage trained weight.</p>\n<p>The key point I think is to calculate gradient loss only labeled frame. The positive labels were sampled from tp_train.csv only and the negative labels were sampled from fp_train.csv only.<br>\nI put 1 to positive label and -1 to negative label.</p>\n<pre><code>tp_dict = {}\nfor recording_id, df in train_tp.groupby(\"recording_id\"):\n    tp_dict[recording_id+\"_posi\"] = df.values[:, [1,3,4,5,6]]\n\nfp_dict = {}\nfor recording_id, df in train_fp.groupby(\"recording_id\"):\n    fp_dict[recording_id+\"_nega\"] = df.values[:, [1,3,4,5,6]]\n\ndef extract_seq_label(label, value):\n    seq_label = np.zeros((24, 3751))  # label, sequence\n    middle = np.ones(24) * -1\n    for species_id, t_min, f_min, t_max, f_max in label:\n        h, t = int(3751*(t_min/60)), int(3751*(t_max/60))\n        m = (t + h)//2\n        middle[species_id] = m\n        seq_label[species_id, h:t] = value\n    return seq_label, middle.astype(int)\n\n# extract positive label and middle point\nfname = \"00204008d\" + \"_posi\"\nposi_label, posi_middle = extract_seq_label(tp_dict[fname], 1) \n\n# extract negative label and middle point\nfname = \"00204008d\" + \"_nega\"\nnega_label, nega_middle = extract_seq_label(fp_dict[fname], -1) \n</code></pre>\n<p>loss function is that:</p>\n<pre><code>def rfcx_2nd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n    posi_label = ((targets == 1).sum(2) &gt; 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) &gt; 0).float().to(device)\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    loss = posi_loss + nega_loss\n    return loss\n</code></pre>\n<p>And image are cut and stack by sliding window.<br>\nI set the window size to 512 and cut out the entire range of 60's audio data by covering it little by little. Cover 49 pixels each, considering that important sounds may be located at the boundaries of the division.</p>\n<pre><code>N_SPLIT_IMG = 8\nWINDOW = 512\nCOVER = 49\n\nslide_img_pos = [[0, WINDOW]]\nfor idx in range(1, N_SPLIT_IMG):\n    h, t = slide_img_pos[idx-1][0], slide_img_pos[idx-1][1]\n    h = t - COVER\n    t = h + WINDOW\n    slide_img_pos.append([h, t])\n\nprint(slide_img_pos)\n# [[0, 512], [463, 975], [926, 1438], [1389, 1901], [1852, 2364], [2315, 2827], [2778, 3290], [3241, 3753]]\n</code></pre>\n<p>I predict each sliding window and put the pseudo label, so I got 8 windows in one 60 sec recording.</p>\n<table>\n<thead>\n<tr>\n<th>patch idx</th>\n<th>pixcel</th>\n<th>time(s)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0〜512</td>\n<td>0〜8</td>\n</tr>\n<tr>\n<td>1</td>\n<td>463〜975</td>\n<td>7〜15</td>\n</tr>\n<tr>\n<td>2</td>\n<td>926〜1438</td>\n<td>14〜23</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1389〜1901</td>\n<td>22〜30</td>\n</tr>\n<tr>\n<td>4</td>\n<td>1852〜2364</td>\n<td>29〜37</td>\n</tr>\n<tr>\n<td>5</td>\n<td>2315〜2827</td>\n<td>37〜45</td>\n</tr>\n<tr>\n<td>6</td>\n<td>2778〜3290</td>\n<td>44〜52</td>\n</tr>\n<tr>\n<td>7</td>\n<td>3241〜3753</td>\n<td>51〜60</td>\n</tr>\n</tbody>\n</table>\n<p>Parameters:</p>\n<ul>\n<li>Adam</li>\n<li>learning_rate=3e-4</li>\n<li>CosineAnnealingLR(max_T=5)</li>\n<li>epoch=5</li>\n</ul>\n<h2>3rd stage: train by label re-labeled</h2>\n<p>This stage trains on the new labels re-labeled by 2nd stage model.</p>\n<p>Use 1st stage trained weight.</p>\n<p>The new label is ensemble by our team output like my 2nd stage.</p>\n<ul>\n<li>our prediction average value is<ul>\n<li><code>&gt;0.5</code>: soft positive = 2</li>\n<li><code>&lt;0.01</code>: soft negative = -2</li></ul></li>\n</ul>\n<p>In this stage, I calculate gradient loss only labeled frame as with 2nd stage. </p>\n<p>Parameters:</p>\n<ul>\n<li>Adam</li>\n<li>learning_rate=3e-4</li>\n<li>CosineAnnealingLR(max_T=5)</li>\n<li>epoch=5</li>\n</ul>\n<p>Some My Tips:</p>\n<ul>\n<li>Don't use soft negative.</li>\n<li>The re-label's loss(soft positive) is weighted 0.5.</li>\n<li>last layer mixup(from <a href=\"https://medium.com/analytics-vidhya/better-result-with-mixup-at-final-layer-e9ba3a4a0c41\" target=\"_blank\">this blog</a>)</li>\n</ul>\n<h2>CV</h2>\n<p>I use <a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a>'s MultilabelStratifiedKFold. Validation data is made from tp_train only and fp_train data is used training in all fold.</p>\n<p>Each stage LWLRAP is that:</p>\n<table>\n<thead>\n<tr>\n<th>stage</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1st</td>\n<td>0.7889</td>\n<td>0.842</td>\n<td>0.865</td>\n</tr>\n<tr>\n<td>2nd</td>\n<td>0.7766</td>\n<td>0.874</td>\n<td>0.878</td>\n</tr>\n<tr>\n<td>3rd</td>\n<td>0.7887</td>\n<td>0.949</td>\n<td>0.951</td>\n</tr>\n</tbody>\n</table>\n<p>3rd stage's re-labeled LWRAP is 0.9621.</p>\n<h2>predict</h2>\n<p>In test time, I increase COVER to 256, so I got 14 windows in one 60 sec recording.<br>\nThe prediction is max pooling in each patch.</p>\n<p>I use clipwise_output in training, and I use framewise_output in prediction. This approach came from <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">shinmura0's discussion thread</a>. Thank you  shinmura0:)</p>\n<h2>did not work for my model</h2>\n<ul>\n<li>TTA</li>\n<li>26 classes (divide song_type)</li>\n<li>label wright loss</li>\n<li>label smoothing(but team member's kuto improved)</li>\n</ul>\n<hr>\n<p>Finally, I would like to thank the team members.<br>\nIf I was alone, I couldn't get these result.<br>\nkuto, Ahmet, thank you very much.</p>\n<p>My code:<br>\n<a href=\"https://github.com/trtd56/RFCX\" target=\"_blank\">https://github.com/trtd56/RFCX</a></p>",
  "messages": [
    {
      "id": "1208444",
      "postDate": "02/18/2021 09:43:58",
      "content": "<h1>5th place solution (Training Strategy)</h1>\n<p>Congratulations to all the participants, and thanks a lot to the organizers for this competition! This has been a very difficult but fun competition:)</p>\n<p>In this thread, I introduce our approach about training strategy.<br>\nAbout ensemble part will be written by my team member.</p>\n<p>Our team ensemble each best model.</p>\n<p>My model is Resnet18 which has a SED header. This model's LWLRAP is Public LB=0.949 /Private LB=0.951, and I trained by google colab using <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198048\" target=\"_blank\">Theo Viel's npz dataset</a>(32 kHz, 128 mels). Thank you, Theo Viel!!</p>\n<p>Our approach has 3 stage, <br>\nOther team members are different in some things likes the base model and hyperparameter, but these default strategies are about the same.</p>\n<h2>1st stage: pre-train</h2>\n<p><em>※I think this part is not important. Team member Ahmet skips this part.</em></p>\n<p>This stage transfers learning from Imagenet to spectograms.</p>\n<p>Theo Viel's npz dataset can be regarded as 128x3751 size image.<br>\nI cut to 512 by this image in sound point from t_min and t_max.<br>\nI train this image by tp_train and 30 sampled fp_train.</p>\n<p>Parameters:</p>\n<ul>\n<li>Adam</li>\n<li>learning_rate=1e-3</li>\n<li>CosineAnnealingLR(max_T=10)</li>\n<li>epoch=50</li>\n</ul>\n<p>Continue 2nd and 3rd stage use this trained weight.</p>\n<h2>2nd stage: pseudo label re-labeling</h2>\n<p>The purpose of stage2 is to improve the model and make pseudo labels by this model.</p>\n<p>Use 1st stage trained weight.</p>\n<p>The key point I think is to calculate gradient loss only labeled frame. The positive labels were sampled from tp_train.csv only and the negative labels were sampled from fp_train.csv only.<br>\nI put 1 to positive label and -1 to negative label.</p>\n<pre><code>tp_dict = {}\nfor recording_id, df in train_tp.groupby(\"recording_id\"):\n    tp_dict[recording_id+\"_posi\"] = df.values[:, [1,3,4,5,6]]\n\nfp_dict = {}\nfor recording_id, df in train_fp.groupby(\"recording_id\"):\n    fp_dict[recording_id+\"_nega\"] = df.values[:, [1,3,4,5,6]]\n\ndef extract_seq_label(label, value):\n    seq_label = np.zeros((24, 3751))  # label, sequence\n    middle = np.ones(24) * -1\n    for species_id, t_min, f_min, t_max, f_max in label:\n        h, t = int(3751*(t_min/60)), int(3751*(t_max/60))\n        m = (t + h)//2\n        middle[species_id] = m\n        seq_label[species_id, h:t] = value\n    return seq_label, middle.astype(int)\n\n# extract positive label and middle point\nfname = \"00204008d\" + \"_posi\"\nposi_label, posi_middle = extract_seq_label(tp_dict[fname], 1) \n\n# extract negative label and middle point\nfname = \"00204008d\" + \"_nega\"\nnega_label, nega_middle = extract_seq_label(fp_dict[fname], -1) \n</code></pre>\n<p>loss function is that:</p>\n<pre><code>def rfcx_2nd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n    posi_label = ((targets == 1).sum(2) &gt; 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) &gt; 0).float().to(device)\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    loss = posi_loss + nega_loss\n    return loss\n</code></pre>\n<p>And image are cut and stack by sliding window.<br>\nI set the window size to 512 and cut out the entire range of 60's audio data by covering it little by little. Cover 49 pixels each, considering that important sounds may be located at the boundaries of the division.</p>\n<pre><code>N_SPLIT_IMG = 8\nWINDOW = 512\nCOVER = 49\n\nslide_img_pos = [[0, WINDOW]]\nfor idx in range(1, N_SPLIT_IMG):\n    h, t = slide_img_pos[idx-1][0], slide_img_pos[idx-1][1]\n    h = t - COVER\n    t = h + WINDOW\n    slide_img_pos.append([h, t])\n\nprint(slide_img_pos)\n# [[0, 512], [463, 975], [926, 1438], [1389, 1901], [1852, 2364], [2315, 2827], [2778, 3290], [3241, 3753]]\n</code></pre>\n<p>I predict each sliding window and put the pseudo label, so I got 8 windows in one 60 sec recording.</p>\n<table>\n<thead>\n<tr>\n<th>patch idx</th>\n<th>pixcel</th>\n<th>time(s)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0〜512</td>\n<td>0〜8</td>\n</tr>\n<tr>\n<td>1</td>\n<td>463〜975</td>\n<td>7〜15</td>\n</tr>\n<tr>\n<td>2</td>\n<td>926〜1438</td>\n<td>14〜23</td>\n</tr>\n<tr>\n<td>3</td>\n<td>1389〜1901</td>\n<td>22〜30</td>\n</tr>\n<tr>\n<td>4</td>\n<td>1852〜2364</td>\n<td>29〜37</td>\n</tr>\n<tr>\n<td>5</td>\n<td>2315〜2827</td>\n<td>37〜45</td>\n</tr>\n<tr>\n<td>6</td>\n<td>2778〜3290</td>\n<td>44〜52</td>\n</tr>\n<tr>\n<td>7</td>\n<td>3241〜3753</td>\n<td>51〜60</td>\n</tr>\n</tbody>\n</table>\n<p>Parameters:</p>\n<ul>\n<li>Adam</li>\n<li>learning_rate=3e-4</li>\n<li>CosineAnnealingLR(max_T=5)</li>\n<li>epoch=5</li>\n</ul>\n<h2>3rd stage: train by label re-labeled</h2>\n<p>This stage trains on the new labels re-labeled by 2nd stage model.</p>\n<p>Use 1st stage trained weight.</p>\n<p>The new label is ensemble by our team output like my 2nd stage.</p>\n<ul>\n<li>our prediction average value is<ul>\n<li><code>&gt;0.5</code>: soft positive = 2</li>\n<li><code>&lt;0.01</code>: soft negative = -2</li></ul></li>\n</ul>\n<p>In this stage, I calculate gradient loss only labeled frame as with 2nd stage. </p>\n<p>Parameters:</p>\n<ul>\n<li>Adam</li>\n<li>learning_rate=3e-4</li>\n<li>CosineAnnealingLR(max_T=5)</li>\n<li>epoch=5</li>\n</ul>\n<p>Some My Tips:</p>\n<ul>\n<li>Don't use soft negative.</li>\n<li>The re-label's loss(soft positive) is weighted 0.5.</li>\n<li>last layer mixup(from <a href=\"https://medium.com/analytics-vidhya/better-result-with-mixup-at-final-layer-e9ba3a4a0c41\" target=\"_blank\">this blog</a>)</li>\n</ul>\n<h2>CV</h2>\n<p>I use <a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a>'s MultilabelStratifiedKFold. Validation data is made from tp_train only and fp_train data is used training in all fold.</p>\n<p>Each stage LWLRAP is that:</p>\n<table>\n<thead>\n<tr>\n<th>stage</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1st</td>\n<td>0.7889</td>\n<td>0.842</td>\n<td>0.865</td>\n</tr>\n<tr>\n<td>2nd</td>\n<td>0.7766</td>\n<td>0.874</td>\n<td>0.878</td>\n</tr>\n<tr>\n<td>3rd</td>\n<td>0.7887</td>\n<td>0.949</td>\n<td>0.951</td>\n</tr>\n</tbody>\n</table>\n<p>3rd stage's re-labeled LWRAP is 0.9621.</p>\n<h2>predict</h2>\n<p>In test time, I increase COVER to 256, so I got 14 windows in one 60 sec recording.<br>\nThe prediction is max pooling in each patch.</p>\n<p>I use clipwise_output in training, and I use framewise_output in prediction. This approach came from <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">shinmura0's discussion thread</a>. Thank you  shinmura0:)</p>\n<h2>did not work for my model</h2>\n<ul>\n<li>TTA</li>\n<li>26 classes (divide song_type)</li>\n<li>label wright loss</li>\n<li>label smoothing(but team member's kuto improved)</li>\n</ul>\n<hr>\n<p>Finally, I would like to thank the team members.<br>\nIf I was alone, I couldn't get these result.<br>\nkuto, Ahmet, thank you very much.</p>\n<p>My code:<br>\n<a href=\"https://github.com/trtd56/RFCX\" target=\"_blank\">https://github.com/trtd56/RFCX</a></p>",
      "rawMarkdown": "# 5th place solution (Training Strategy)\n\nCongratulations to all the participants, and thanks a lot to the organizers for this competition! This has been a very difficult but fun competition:)\n\nIn this thread, I introduce our approach about training strategy.\nAbout ensemble part will be written by my team member.\n\nOur team ensemble each best model.\n\nMy model is Resnet18 which has a SED header. This model's LWLRAP is Public LB=0.949 /Private LB=0.951, and I trained by google colab using [Theo Viel's npz dataset](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198048)(32 kHz, 128 mels). Thank you, Theo Viel!!\n\nOur approach has 3 stage, \nOther team members are different in some things likes the base model and hyperparameter, but these default strategies are about the same.\n\n## 1st stage: pre-train\n\n*※I think this part is not important. Team member Ahmet skips this part.*\n\nThis stage transfers learning from Imagenet to spectograms.\n\nTheo Viel's npz dataset can be regarded as 128x3751 size image.\nI cut to 512 by this image in sound point from t_min and t_max.\nI train this image by tp_train and 30 sampled fp_train.\n\nParameters:\n- Adam\n- learning_rate=1e-3\n- CosineAnnealingLR(max_T=10)\n- epoch=50\n\nContinue 2nd and 3rd stage use this trained weight.\n\n## 2nd stage: pseudo label re-labeling\n\nThe purpose of stage2 is to improve the model and make pseudo labels by this model.\n\nUse 1st stage trained weight.\n\nThe key point I think is to calculate gradient loss only labeled frame. The positive labels were sampled from tp_train.csv only and the negative labels were sampled from fp_train.csv only.\nI put 1 to positive label and -1 to negative label.\n\n```python\ntp_dict = {}\nfor recording_id, df in train_tp.groupby(\"recording_id\"):\n    tp_dict[recording_id+\"_posi\"] = df.values[:, [1,3,4,5,6]]\n\nfp_dict = {}\nfor recording_id, df in train_fp.groupby(\"recording_id\"):\n    fp_dict[recording_id+\"_nega\"] = df.values[:, [1,3,4,5,6]]\n    \ndef extract_seq_label(label, value):\n    seq_label = np.zeros((24, 3751))  # label, sequence\n    middle = np.ones(24) * -1\n    for species_id, t_min, f_min, t_max, f_max in label:\n        h, t = int(3751*(t_min/60)), int(3751*(t_max/60))\n        m = (t + h)//2\n        middle[species_id] = m\n        seq_label[species_id, h:t] = value\n    return seq_label, middle.astype(int)\n\n# extract positive label and middle point\nfname = \"00204008d\" + \"_posi\"\nposi_label, posi_middle = extract_seq_label(tp_dict[fname], 1) \n\n# extract negative label and middle point\nfname = \"00204008d\" + \"_nega\"\nnega_label, nega_middle = extract_seq_label(fp_dict[fname], -1) \n```\n\nloss function is that:\n\n```python\ndef rfcx_2nd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n    posi_label = ((targets == 1).sum(2) > 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) > 0).float().to(device)\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    loss = posi_loss + nega_loss\n    return loss\n```\n\nAnd image are cut and stack by sliding window.\nI set the window size to 512 and cut out the entire range of 60's audio data by covering it little by little. Cover 49 pixels each, considering that important sounds may be located at the boundaries of the division.\n\n```python\nN_SPLIT_IMG = 8\nWINDOW = 512\nCOVER = 49\n\nslide_img_pos = [[0, WINDOW]]\nfor idx in range(1, N_SPLIT_IMG):\n    h, t = slide_img_pos[idx-1][0], slide_img_pos[idx-1][1]\n    h = t - COVER\n    t = h + WINDOW\n    slide_img_pos.append([h, t])\n\nprint(slide_img_pos)\n# [[0, 512], [463, 975], [926, 1438], [1389, 1901], [1852, 2364], [2315, 2827], [2778, 3290], [3241, 3753]]\n```\n\nI predict each sliding window and put the pseudo label, so I got 8 windows in one 60 sec recording.\n\n|patch idx|pixcel|time(s)|\n|--|--|--|\n|0|0〜512|0〜8|\n|1|463〜975|7〜15|\n|2|926〜1438|14〜23|\n|3|1389〜1901|22〜30|\n|4|1852〜2364|29〜37|\n|5|2315〜2827|37〜45|\n|6|2778〜3290|44〜52|\n|7|3241〜3753|51〜60|\n\n\nParameters:\n- Adam\n- learning_rate=3e-4\n- CosineAnnealingLR(max_T=5)\n- epoch=5\n\n## 3rd stage: train by label re-labeled\nThis stage trains on the new labels re-labeled by 2nd stage model.\n\nUse 1st stage trained weight.\n\nThe new label is ensemble by our team output like my 2nd stage.\n- our prediction average value is\n  - `>0.5`: soft positive = 2\n  - `<0.01`: soft negative = -2\n\nIn this stage, I calculate gradient loss only labeled frame as with 2nd stage. \n\nParameters:\n- Adam\n- learning_rate=3e-4\n- CosineAnnealingLR(max_T=5)\n- epoch=5\n\nSome My Tips:\n- Don't use soft negative.\n- The re-label's loss(soft positive) is weighted 0.5.\n- last layer mixup(from [this blog](https://medium.com/analytics-vidhya/better-result-with-mixup-at-final-layer-e9ba3a4a0c41))\n\n## CV\n\nI use [iterative-stratification](https://github.com/trent-b/iterative-stratification)'s MultilabelStratifiedKFold. Validation data is made from tp_train only and fp_train data is used training in all fold.\n\nEach stage LWLRAP is that:\n\n|stage|CV|Public|Private|\n|--|--|--|--|\n|1st|0.7889|0.842|0.865|\n|2nd|0.7766|0.874|0.878|\n|3rd|0.7887|0.949|0.951|\n\n3rd stage's re-labeled LWRAP is 0.9621.\n\n\n## predict\n\nIn test time, I increase COVER to 256, so I got 14 windows in one 60 sec recording.\nThe prediction is max pooling in each patch.\n\nI use clipwise_output in training, and I use framewise_output in prediction. This approach came from [shinmura0's discussion thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684). Thank you  shinmura0:)\n\n## did not work for my model\n- TTA\n- 26 classes (divide song_type)\n- label wright loss\n- label smoothing(but team member's kuto improved)\n\n---\n\nFinally, I would like to thank the team members.\nIf I was alone, I couldn't get these result.\nkuto, Ahmet, thank you very much.\n\nMy code:\nhttps://github.com/trtd56/RFCX",
      "votes": null
    },
    {
      "id": "1208506",
      "postDate": "02/18/2021 10:02:49",
      "content": "<p>I would like to add a bit more details:</p>\n<p>It was important to have robust pseudolabels. Therefore, I have trained a Vision Transformer model and a Wavenet over resnet features for each class independently. Spectograms were generated with respect to min-max frequencies for each class. I have used the same architectures for training the last stage models on pseudolabels as well. Then I have aligned my logits to have the same mean and std as Toda's and got the optimal ensembling weights based on individual class AUC (tp vs not tp). This ensemble later is blended with Kuto's submission. Each model has different bias and training scheme. This way we achieved robustness.</p>",
      "rawMarkdown": "I would like to add a bit more details:\n\nIt was important to have robust pseudolabels. Therefore, I have trained a Vision Transformer model and a Wavenet over resnet features for each class independently. Spectograms were generated with respect to min-max frequencies for each class. I have used the same architectures for training the last stage models on pseudolabels as well. Then I have aligned my logits to have the same mean and std as Toda's and got the optimal ensembling weights based on individual class AUC (tp vs not tp). This ensemble later is blended with Kuto's submission. Each model has different bias and training scheme. This way we achieved robustness.",
      "votes": null
    },
    {
      "id": "1208575",
      "postDate": "02/18/2021 10:46:35",
      "content": "<p>Congrats on 5th place and gold medal <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a> and team, learn alot from your team solution</p>",
      "rawMarkdown": "Congrats on 5th place and gold medal @takamichitoda and team, learn alot from your team solution",
      "votes": null
    },
    {
      "id": "1208709",
      "postDate": "02/18/2021 12:56:49",
      "content": "<p>Great solution! And congrats 5th place.</p>\n<p>I have a question.<br>\nDid you train the model with weak label?</p>",
      "rawMarkdown": "Great solution! And congrats 5th place.\n\nI have a question.\nDid you train the model with weak label?",
      "votes": null
    },
    {
      "id": "1208715",
      "postDate": "02/18/2021 12:59:31",
      "content": "<p>Congratz ! Really glad to see your team made it this far with my melspecs :D </p>",
      "rawMarkdown": "Congratz ! Really glad to see your team made it this far with my melspecs :D",
      "votes": null
    },
    {
      "id": "1209472",
      "postDate": "02/18/2021 23:15:55",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a> <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> and <a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> ! Nice job. </p>\n<p>The way you coded the loss function is pretty similar as the way I coded mine, inclusive I also mapped my labels to -1 and 1.<br>\nThanks for sharing.</p>",
      "rawMarkdown": "Congrats @takamichitoda @aerdem4 and @kuto0633 ! Nice job. \n\nThe way you coded the loss function is pretty similar as the way I coded mine, inclusive I also mapped my labels to -1 and 1.\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "1209473",
      "postDate": "02/18/2021 23:16:33",
      "content": "<p>Your dataset was very useful for someone only having a laptop PC like me. It can make experiments fastly. Thank you!</p>",
      "rawMarkdown": "Your dataset was very useful for someone only having a laptop PC like me. It can make experiments fastly. Thank you!",
      "votes": null
    },
    {
      "id": "1209488",
      "postDate": "02/18/2021 23:30:15",
      "content": "<p>Thank you!</p>\n<p>Yes, I did.<br>\n<code>clipwise_output in training</code> mean is weak label training.<br>\nI train by clipwise label each sliding window.<br>\nI called it soft frame wise training.</p>",
      "rawMarkdown": "Thank you!\n\nYes, I did.\n`clipwise_output in training` mean is weak label training.\nI train by clipwise label each sliding window.\nI called it soft frame wise training.",
      "votes": null
    },
    {
      "id": "1209545",
      "postDate": "02/19/2021 00:23:41",
      "content": "<p>Congrats ! I'm very happy reading your writeup. <br>\nI have a question about CV. </p>\n<p>Why do you think your CV is OK ?</p>\n<p>Because your CV and LB are different and seems not corelated. <br>\nYour CV having value of 0.788 to 0.7887 that is almost constant, but LB varies 0.842 to 0.949 huge gap.</p>",
      "rawMarkdown": "Congrats ! I'm very happy reading your writeup. \nI have a question about CV. \n\nWhy do you think your CV is OK ?\n\nBecause your CV and LB are different and seems not corelated. \nYour CV having value of 0.788 to 0.7887 that is almost constant, but LB varies 0.842 to 0.949 huge gap.",
      "votes": null
    },
    {
      "id": "1209749",
      "postDate": "02/19/2021 03:04:24",
      "content": "<p>Weak label training was strong in competition.<br>\nI'm glad to see that SED is useful.</p>",
      "rawMarkdown": "Weak label training was strong in competition.\nI'm glad to see that SED is useful.",
      "votes": null
    },
    {
      "id": "1209994",
      "postDate": "02/19/2021 06:32:41",
      "content": "<p>Congratulations for your great result!<br>\nI see you've got impressive boost from pseudo training (0.878-&gt;0.951). I have some questions. you said,</p>\n<pre><code>&gt;0.5: soft positive = 2\n&lt;0.01: soft negative = -2\n</code></pre>\n<p>and </p>\n<pre><code>Don't use soft negative.\n</code></pre>\n<p>Do you mean when pseudo labeling, you converted &gt;0.5 predictions to 1 and others to <code>unknown(to mask out during loss computation)</code>?</p>",
      "rawMarkdown": "Congratulations for your great result!\nI see you've got impressive boost from pseudo training (0.878->0.951). I have some questions. you said,\n```\n>0.5: soft positive = 2\n<0.01: soft negative = -2\n```\nand \n```\nDon't use soft negative.\n```\nDo you mean when pseudo labeling, you converted >0.5 predictions to 1 and others to `unknown(to mask out during loss computation)`?",
      "votes": null
    },
    {
      "id": "1212167",
      "postDate": "02/21/2021 00:25:39",
      "content": "<p>Thank you comment and sorry for late.</p>\n<p>I calculated  LWRAP by 3rd stage's re-labeled and it was 0.9621 which near LB.<br>\nI had tried Precision, Recall, and AUC but I could not completely correlate CV and LB.<br>\nSo we trust LB and using various model ensembles to hold robustness.</p>",
      "rawMarkdown": "Thank you comment and sorry for late.\n\nI calculated  LWRAP by 3rd stage's re-labeled and it was 0.9621 which near LB.\nI had tried Precision, Recall, and AUC but I could not completely correlate CV and LB.\nSo we trust LB and using various model ensembles to hold robustness.",
      "votes": null
    },
    {
      "id": "1212169",
      "postDate": "02/21/2021 00:33:25",
      "content": "<p>Thank you comment and sorry for late.</p>\n<p>I treat separately original labels and pseudo labels.</p>\n<p>My first 3rd stage loss function is that:</p>\n<pre><code>def rfcx_3rd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n\n    posi_label = ((targets == 1).sum(2) &gt; 0).float().to(device)\n    soft_posi_label = ((targets == 2).sum(2) &gt; 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) &gt; 0).float().to(device)\n    soft_nega_label = ((targets == -2).sum(2) &gt; 0).float().to(device)\n\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    soft_posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    soft_nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    soft_posi_loss = (soft_posi_loss * soft_posi_label).sum()\n    soft_nega_loss = (soft_nega_loss * soft_nega_label).sum()\n\n    loss = posi_loss + nega_loss + soft_posi_loss*0.5 + soft_nega_loss*0.5\n    return loss\n</code></pre>\n<p>But soft_nega is not good work, so I have removed it.</p>",
      "rawMarkdown": "Thank you comment and sorry for late.\n\nI treat separately original labels and pseudo labels.\n\nMy first 3rd stage loss function is that:\n\n```python\ndef rfcx_3rd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n\n    posi_label = ((targets == 1).sum(2) > 0).float().to(device)\n    soft_posi_label = ((targets == 2).sum(2) > 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) > 0).float().to(device)\n    soft_nega_label = ((targets == -2).sum(2) > 0).float().to(device)\n\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    soft_posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    soft_nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    soft_posi_loss = (soft_posi_loss * soft_posi_label).sum()\n    soft_nega_loss = (soft_nega_loss * soft_nega_label).sum()\n\n    loss = posi_loss + nega_loss + soft_posi_loss*0.5 + soft_nega_loss*0.5\n    return loss\n```\n\nBut soft_nega is not good work, so I have removed it.",
      "votes": null
    },
    {
      "id": "1212217",
      "postDate": "02/21/2021 02:42:48",
      "content": "<p>So you converted confident positive soft predictions to hard labels, then weighted them by 0.5 in loss. Thank you for your reply! </p>",
      "rawMarkdown": "So you converted confident positive soft predictions to hard labels, then weighted them by 0.5 in loss. Thank you for your reply!",
      "votes": null
    },
    {
      "id": "1212522",
      "postDate": "02/21/2021 09:49:05",
      "content": "<p>Congratulations for your great result!</p>",
      "rawMarkdown": "Congratulations for your great result!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1208506,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "02/18/2021 10:02:49",
      "content": "<p>I would like to add a bit more details:</p>\n<p>It was important to have robust pseudolabels. Therefore, I have trained a Vision Transformer model and a Wavenet over resnet features for each class independently. Spectograms were generated with respect to min-max frequencies for each class. I have used the same architectures for training the last stage models on pseudolabels as well. Then I have aligned my logits to have the same mean and std as Toda's and got the optimal ensembling weights based on individual class AUC (tp vs not tp). This ensemble later is blended with Kuto's submission. Each model has different bias and training scheme. This way we achieved robustness.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208575,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "02/18/2021 10:46:35",
      "content": "<p>Congrats on 5th place and gold medal <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a> and team, learn alot from your team solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208709,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "02/18/2021 12:56:49",
      "content": "<p>Great solution! And congrats 5th place.</p>\n<p>I have a question.<br>\nDid you train the model with weak label?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209488,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "02/18/2021 23:30:15",
          "content": "<p>Thank you!</p>\n<p>Yes, I did.<br>\n<code>clipwise_output in training</code> mean is weak label training.<br>\nI train by clipwise label each sliding window.<br>\nI called it soft frame wise training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209749,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "02/19/2021 03:04:24",
          "content": "<p>Weak label training was strong in competition.<br>\nI'm glad to see that SED is useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1208715,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/18/2021 12:59:31",
      "content": "<p>Congratz ! Really glad to see your team made it this far with my melspecs :D </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209473,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "02/18/2021 23:16:33",
          "content": "<p>Your dataset was very useful for someone only having a laptop PC like me. It can make experiments fastly. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209472,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "02/18/2021 23:15:55",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a> <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> and <a href=\"https://www.kaggle.com/kuto0633\" target=\"_blank\">@kuto0633</a> ! Nice job. </p>\n<p>The way you coded the loss function is pretty similar as the way I coded mine, inclusive I also mapped my labels to -1 and 1.<br>\nThanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209545,
      "author_name": "xinyan989",
      "author_url": "",
      "post_date": "02/19/2021 00:23:41",
      "content": "<p>Congrats ! I'm very happy reading your writeup. <br>\nI have a question about CV. </p>\n<p>Why do you think your CV is OK ?</p>\n<p>Because your CV and LB are different and seems not corelated. <br>\nYour CV having value of 0.788 to 0.7887 that is almost constant, but LB varies 0.842 to 0.949 huge gap.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212167,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "02/21/2021 00:25:39",
          "content": "<p>Thank you comment and sorry for late.</p>\n<p>I calculated  LWRAP by 3rd stage's re-labeled and it was 0.9621 which near LB.<br>\nI had tried Precision, Recall, and AUC but I could not completely correlate CV and LB.<br>\nSo we trust LB and using various model ensembles to hold robustness.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209994,
      "author_name": "harangdev",
      "author_url": "",
      "post_date": "02/19/2021 06:32:41",
      "content": "<p>Congratulations for your great result!<br>\nI see you've got impressive boost from pseudo training (0.878-&gt;0.951). I have some questions. you said,</p>\n<pre><code>&gt;0.5: soft positive = 2\n&lt;0.01: soft negative = -2\n</code></pre>\n<p>and </p>\n<pre><code>Don't use soft negative.\n</code></pre>\n<p>Do you mean when pseudo labeling, you converted &gt;0.5 predictions to 1 and others to <code>unknown(to mask out during loss computation)</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212169,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "02/21/2021 00:33:25",
          "content": "<p>Thank you comment and sorry for late.</p>\n<p>I treat separately original labels and pseudo labels.</p>\n<p>My first 3rd stage loss function is that:</p>\n<pre><code>def rfcx_3rd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n\n    posi_label = ((targets == 1).sum(2) &gt; 0).float().to(device)\n    soft_posi_label = ((targets == 2).sum(2) &gt; 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) &gt; 0).float().to(device)\n    soft_nega_label = ((targets == -2).sum(2) &gt; 0).float().to(device)\n\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    soft_posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    soft_nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    soft_posi_loss = (soft_posi_loss * soft_posi_label).sum()\n    soft_nega_loss = (soft_nega_loss * soft_nega_label).sum()\n\n    loss = posi_loss + nega_loss + soft_posi_loss*0.5 + soft_nega_loss*0.5\n    return loss\n</code></pre>\n<p>But soft_nega is not good work, so I have removed it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212217,
          "author_name": "harangdev",
          "author_url": "",
          "post_date": "02/21/2021 02:42:48",
          "content": "<p>So you converted confident positive soft predictions to hard labels, then weighted them by 0.5 in loss. Thank you for your reply! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212522,
      "author_name": "iamdeepak95",
      "author_url": "",
      "post_date": "02/21/2021 09:49:05",
      "content": "<p>Congratulations for your great result!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1208444": "# 5th place solution (Training Strategy)\n\nCongratulations to all the participants, and thanks a lot to the organizers for this competition! This has been a very difficult but fun competition:)\n\nIn this thread, I introduce our approach about training strategy.\nAbout ensemble part will be written by my team member.\n\nOur team ensemble each best model.\n\nMy model is Resnet18 which has a SED header. This model's LWLRAP is Public LB=0.949 /Private LB=0.951, and I trained by google colab using [Theo Viel's npz dataset](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198048)(32 kHz, 128 mels). Thank you, Theo Viel!!\n\nOur approach has 3 stage, \nOther team members are different in some things likes the base model and hyperparameter, but these default strategies are about the same.\n\n## 1st stage: pre-train\n\n*※I think this part is not important. Team member Ahmet skips this part.*\n\nThis stage transfers learning from Imagenet to spectograms.\n\nTheo Viel's npz dataset can be regarded as 128x3751 size image.\nI cut to 512 by this image in sound point from t_min and t_max.\nI train this image by tp_train and 30 sampled fp_train.\n\nParameters:\n- Adam\n- learning_rate=1e-3\n- CosineAnnealingLR(max_T=10)\n- epoch=50\n\nContinue 2nd and 3rd stage use this trained weight.\n\n## 2nd stage: pseudo label re-labeling\n\nThe purpose of stage2 is to improve the model and make pseudo labels by this model.\n\nUse 1st stage trained weight.\n\nThe key point I think is to calculate gradient loss only labeled frame. The positive labels were sampled from tp_train.csv only and the negative labels were sampled from fp_train.csv only.\nI put 1 to positive label and -1 to negative label.\n\n```python\ntp_dict = {}\nfor recording_id, df in train_tp.groupby(\"recording_id\"):\n    tp_dict[recording_id+\"_posi\"] = df.values[:, [1,3,4,5,6]]\n\nfp_dict = {}\nfor recording_id, df in train_fp.groupby(\"recording_id\"):\n    fp_dict[recording_id+\"_nega\"] = df.values[:, [1,3,4,5,6]]\n    \ndef extract_seq_label(label, value):\n    seq_label = np.zeros((24, 3751))  # label, sequence\n    middle = np.ones(24) * -1\n    for species_id, t_min, f_min, t_max, f_max in label:\n        h, t = int(3751*(t_min/60)), int(3751*(t_max/60))\n        m = (t + h)//2\n        middle[species_id] = m\n        seq_label[species_id, h:t] = value\n    return seq_label, middle.astype(int)\n\n# extract positive label and middle point\nfname = \"00204008d\" + \"_posi\"\nposi_label, posi_middle = extract_seq_label(tp_dict[fname], 1) \n\n# extract negative label and middle point\nfname = \"00204008d\" + \"_nega\"\nnega_label, nega_middle = extract_seq_label(fp_dict[fname], -1) \n```\n\nloss function is that:\n\n```python\ndef rfcx_2nd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n    posi_label = ((targets == 1).sum(2) > 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) > 0).float().to(device)\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    loss = posi_loss + nega_loss\n    return loss\n```\n\nAnd image are cut and stack by sliding window.\nI set the window size to 512 and cut out the entire range of 60's audio data by covering it little by little. Cover 49 pixels each, considering that important sounds may be located at the boundaries of the division.\n\n```python\nN_SPLIT_IMG = 8\nWINDOW = 512\nCOVER = 49\n\nslide_img_pos = [[0, WINDOW]]\nfor idx in range(1, N_SPLIT_IMG):\n    h, t = slide_img_pos[idx-1][0], slide_img_pos[idx-1][1]\n    h = t - COVER\n    t = h + WINDOW\n    slide_img_pos.append([h, t])\n\nprint(slide_img_pos)\n# [[0, 512], [463, 975], [926, 1438], [1389, 1901], [1852, 2364], [2315, 2827], [2778, 3290], [3241, 3753]]\n```\n\nI predict each sliding window and put the pseudo label, so I got 8 windows in one 60 sec recording.\n\n|patch idx|pixcel|time(s)|\n|--|--|--|\n|0|0〜512|0〜8|\n|1|463〜975|7〜15|\n|2|926〜1438|14〜23|\n|3|1389〜1901|22〜30|\n|4|1852〜2364|29〜37|\n|5|2315〜2827|37〜45|\n|6|2778〜3290|44〜52|\n|7|3241〜3753|51〜60|\n\n\nParameters:\n- Adam\n- learning_rate=3e-4\n- CosineAnnealingLR(max_T=5)\n- epoch=5\n\n## 3rd stage: train by label re-labeled\nThis stage trains on the new labels re-labeled by 2nd stage model.\n\nUse 1st stage trained weight.\n\nThe new label is ensemble by our team output like my 2nd stage.\n- our prediction average value is\n  - `>0.5`: soft positive = 2\n  - `<0.01`: soft negative = -2\n\nIn this stage, I calculate gradient loss only labeled frame as with 2nd stage. \n\nParameters:\n- Adam\n- learning_rate=3e-4\n- CosineAnnealingLR(max_T=5)\n- epoch=5\n\nSome My Tips:\n- Don't use soft negative.\n- The re-label's loss(soft positive) is weighted 0.5.\n- last layer mixup(from [this blog](https://medium.com/analytics-vidhya/better-result-with-mixup-at-final-layer-e9ba3a4a0c41))\n\n## CV\n\nI use [iterative-stratification](https://github.com/trent-b/iterative-stratification)'s MultilabelStratifiedKFold. Validation data is made from tp_train only and fp_train data is used training in all fold.\n\nEach stage LWLRAP is that:\n\n|stage|CV|Public|Private|\n|--|--|--|--|\n|1st|0.7889|0.842|0.865|\n|2nd|0.7766|0.874|0.878|\n|3rd|0.7887|0.949|0.951|\n\n3rd stage's re-labeled LWRAP is 0.9621.\n\n\n## predict\n\nIn test time, I increase COVER to 256, so I got 14 windows in one 60 sec recording.\nThe prediction is max pooling in each patch.\n\nI use clipwise_output in training, and I use framewise_output in prediction. This approach came from [shinmura0's discussion thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684). Thank you  shinmura0:)\n\n## did not work for my model\n- TTA\n- 26 classes (divide song_type)\n- label wright loss\n- label smoothing(but team member's kuto improved)\n\n---\n\nFinally, I would like to thank the team members.\nIf I was alone, I couldn't get these result.\nkuto, Ahmet, thank you very much.\n\nMy code:\nhttps://github.com/trtd56/RFCX",
    "1208506": "I would like to add a bit more details:\n\nIt was important to have robust pseudolabels. Therefore, I have trained a Vision Transformer model and a Wavenet over resnet features for each class independently. Spectograms were generated with respect to min-max frequencies for each class. I have used the same architectures for training the last stage models on pseudolabels as well. Then I have aligned my logits to have the same mean and std as Toda's and got the optimal ensembling weights based on individual class AUC (tp vs not tp). This ensemble later is blended with Kuto's submission. Each model has different bias and training scheme. This way we achieved robustness.",
    "1208575": "Congrats on 5th place and gold medal @takamichitoda and team, learn alot from your team solution",
    "1208709": "Great solution! And congrats 5th place.\n\nI have a question.\nDid you train the model with weak label?",
    "1208715": "Congratz ! Really glad to see your team made it this far with my melspecs :D",
    "1209472": "Congrats @takamichitoda @aerdem4 and @kuto0633 ! Nice job. \n\nThe way you coded the loss function is pretty similar as the way I coded mine, inclusive I also mapped my labels to -1 and 1.\nThanks for sharing.",
    "1209473": "Your dataset was very useful for someone only having a laptop PC like me. It can make experiments fastly. Thank you!",
    "1209488": "Thank you!\n\nYes, I did.\n`clipwise_output in training` mean is weak label training.\nI train by clipwise label each sliding window.\nI called it soft frame wise training.",
    "1209545": "Congrats ! I'm very happy reading your writeup. \nI have a question about CV. \n\nWhy do you think your CV is OK ?\n\nBecause your CV and LB are different and seems not corelated. \nYour CV having value of 0.788 to 0.7887 that is almost constant, but LB varies 0.842 to 0.949 huge gap.",
    "1209749": "Weak label training was strong in competition.\nI'm glad to see that SED is useful.",
    "1209994": "Congratulations for your great result!\nI see you've got impressive boost from pseudo training (0.878->0.951). I have some questions. you said,\n```\n>0.5: soft positive = 2\n<0.01: soft negative = -2\n```\nand \n```\nDon't use soft negative.\n```\nDo you mean when pseudo labeling, you converted >0.5 predictions to 1 and others to `unknown(to mask out during loss computation)`?",
    "1212167": "Thank you comment and sorry for late.\n\nI calculated  LWRAP by 3rd stage's re-labeled and it was 0.9621 which near LB.\nI had tried Precision, Recall, and AUC but I could not completely correlate CV and LB.\nSo we trust LB and using various model ensembles to hold robustness.",
    "1212169": "Thank you comment and sorry for late.\n\nI treat separately original labels and pseudo labels.\n\nMy first 3rd stage loss function is that:\n\n```python\ndef rfcx_3rd_criterion(outputs, targets):\n    clipwise_preds_att_ti = outputs[\"clipwise_preds_att_ti\"]\n\n    posi_label = ((targets == 1).sum(2) > 0).float().to(device)\n    soft_posi_label = ((targets == 2).sum(2) > 0).float().to(device)\n    nega_label = ((targets == -1).sum(2) > 0).float().to(device)\n    soft_nega_label = ((targets == -2).sum(2) > 0).float().to(device)\n\n    posi_y = torch.ones(clipwise_preds_att_ti.shape).to(device)\n    nega_y = torch.zeros(clipwise_preds_att_ti.shape).to(device)\n\n    posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n    soft_posi_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, posi_y)\n    soft_nega_loss = nn.BCEWithLogitsLoss(reduction=\"none\")(clipwise_preds_att_ti, nega_y)\n\n    posi_loss = (posi_loss * posi_label).sum()\n    nega_loss = (nega_loss * nega_label).sum()\n    soft_posi_loss = (soft_posi_loss * soft_posi_label).sum()\n    soft_nega_loss = (soft_nega_loss * soft_nega_label).sum()\n\n    loss = posi_loss + nega_loss + soft_posi_loss*0.5 + soft_nega_loss*0.5\n    return loss\n```\n\nBut soft_nega is not good work, so I have removed it.",
    "1212217": "So you converted confident positive soft predictions to hard labels, then weighted them by 0.5 in loss. Thank you for your reply!",
    "1212522": "Congratulations for your great result!"
  },
  "source": "meta"
}