{
  "id": 95429,
  "title": "13rd solution summary",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/vfa-13rd-solution-summary",
  "author_name": "",
  "post_date": "2019-08-27T10:25:19.047Z",
  "votes": 22,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>As I have made a slide about my solution, I would like to post it here. Thank organizer and staffs for this great competition.</p>\n\n<p>Here is the brief summary.</p>\n\n<p>Inference:</p>\n\n<ol>\n<li>Raw audio to log mel.</li>\n<li>Apply sliding window to the log mel and predict each windows by CNN.</li>\n<li>Feed temporal prediction given step2 to RNN</li>\n<li>Feed temporal prediction given step2 to Rakeld + ExtraTrees</li>\n<li>Average step3 and step4.</li>\n</ol>\n\n<p>Training:</p>\n\n<ol>\n<li>Train CNN by using only curated data.</li>\n<li>Apply sliding window to the log mel in noisy dataset and predict each windows by CNN.</li>\n<li>Feed temporal prediction given step2 to RNN. Then, I got pseudo labels for noisy dataset.</li>\n<li>Average pseudo labels with original noisy labels by 0.75:0.25.</li>\n<li>Train CNN by using all data.</li>\n<li>Train some models to follow the inference steps.</li>\n</ol>\n\n<p>Details in Training CNN:</p>\n\n<ul>\n<li>Densenet 121</li>\n<li>Spec Aug</li>\n<li>Mixup</li>\n<li>CLR + SWA + Snapshot ensemble.</li>\n</ul>\n\n<p>I intensively tried SSL approach such as ICT, but it did not work.</p>",
  "messages": [
    {
      "id": "550946",
      "postDate": "06/12/2019 06:48:18",
      "content": "<p>Hi,</p>\n\n<p>As I have made a slide about my solution, I would like to post it here. Thank organizer and staffs for this great competition.</p>\n\n<p>Here is the brief summary.</p>\n\n<p>Inference:</p>\n\n<ol>\n<li>Raw audio to log mel.</li>\n<li>Apply sliding window to the log mel and predict each windows by CNN.</li>\n<li>Feed temporal prediction given step2 to RNN</li>\n<li>Feed temporal prediction given step2 to Rakeld + ExtraTrees</li>\n<li>Average step3 and step4.</li>\n</ol>\n\n<p>Training:</p>\n\n<ol>\n<li>Train CNN by using only curated data.</li>\n<li>Apply sliding window to the log mel in noisy dataset and predict each windows by CNN.</li>\n<li>Feed temporal prediction given step2 to RNN. Then, I got pseudo labels for noisy dataset.</li>\n<li>Average pseudo labels with original noisy labels by 0.75:0.25.</li>\n<li>Train CNN by using all data.</li>\n<li>Train some models to follow the inference steps.</li>\n</ol>\n\n<p>Details in Training CNN:</p>\n\n<ul>\n<li>Densenet 121</li>\n<li>Spec Aug</li>\n<li>Mixup</li>\n<li>CLR + SWA + Snapshot ensemble.</li>\n</ul>\n\n<p>I intensively tried SSL approach such as ICT, but it did not work.</p>",
      "rawMarkdown": "Hi,\n\nAs I have made a slide about my solution, I would like to post it here. Thank organizer and staffs for this great competition.\n\nHere is the brief summary.\n\nInference:\n\n1. Raw audio to log mel.\n1. Apply sliding window to the log mel and predict each windows by CNN.\n1. Feed temporal prediction given step2 to RNN\n1. Feed temporal prediction given step2 to Rakeld + ExtraTrees\n1. Average step3 and step4.\n\nTraining:\n\n1. Train CNN by using only curated data.\n1. Apply sliding window to the log mel in noisy dataset and predict each windows by CNN.\n1. Feed temporal prediction given step2 to RNN. Then, I got pseudo labels for noisy dataset.\n1. Average pseudo labels with original noisy labels by 0.75:0.25.\n1. Train CNN by using all data.\n1. Train some models to follow the inference steps.\n\nDetails in Training CNN:\n\n* Densenet 121\n* Spec Aug\n* Mixup\n* CLR + SWA + Snapshot ensemble.\n\nI intensively tried SSL approach such as ICT, but it did not work.",
      "votes": null
    },
    {
      "id": "551041",
      "postDate": "06/12/2019 09:26:30",
      "content": "<p>Your idea on RNN is splendid! Thanks for sharing!</p>",
      "rawMarkdown": "Your idea on RNN is splendid! Thanks for sharing!",
      "votes": null
    },
    {
      "id": "551100",
      "postDate": "06/12/2019 10:44:06",
      "content": "<p>Thanks for sharing <a href=\"/akirasosa\">@akirasosa</a> </p>",
      "rawMarkdown": "Thanks for sharing @akirasosa",
      "votes": null
    },
    {
      "id": "551453",
      "postDate": "06/12/2019 18:33:04",
      "content": "<p>Thank you for sharing, very nice solution <a href=\"/akirasosa\">@akirasosa</a> ! I also got my best results with a CNN + RNN (LSTM) and used it in my final submission so its great to see someone else in this competition did the same :) </p>\n\n<p>For input to the RNN though, did you use the penultimate layer of CNN (i.e. GAP layer extracted feature vector) or something else?</p>\n\n<p>For CNN sliding window, did you do this at fixed intervals or random crops (random starting points with fixed window size).</p>\n\n<p>Rakeld + ExtraTrees is very interesting, I didn't consider adding anything like that.</p>",
      "rawMarkdown": "Thank you for sharing, very nice solution @akirasosa ! I also got my best results with a CNN + RNN (LSTM) and used it in my final submission so its great to see someone else in this competition did the same :) \n\nFor input to the RNN though, did you use the penultimate layer of CNN (i.e. GAP layer extracted feature vector) or something else?\n\nFor CNN sliding window, did you do this at fixed intervals or random crops (random starting points with fixed window size).\n\nRakeld + ExtraTrees is very interesting, I didn't consider adding anything like that.",
      "votes": null
    },
    {
      "id": "551705",
      "postDate": "06/13/2019 03:12:40",
      "content": "<p><a href=\"/jamesrequa\">@jamesrequa</a></p>\n\n<blockquote>\n  <p>did you use the penultimate layer of CNN</p>\n</blockquote>\n\n<p>No. I used the probability predicted by the CNN. I tried some others as an input. But it was the best to just feeding the temporal probability.</p>\n\n<p>After training CNN, I predicted sliding window and save the result which has shape (n_data, n_windows, n_class) as a file.</p>\n\n<blockquote>\n  <p>did you do this at fixed intervals or random crops</p>\n</blockquote>\n\n<p>I used fixed step size 21. The window size is 224, which is equivalent with about 5 sec. If I use smaller step size and more windows, the result becomes better, but the inference time is longer.</p>\n\n<p>I hope you can get good position in final stage. Good luck!</p>",
      "rawMarkdown": "jamesrequa\n&gt; did you use the penultimate layer of CNN\n\nNo. I used the probability predicted by the CNN. I tried some others as an input. But it was the best to just feeding the temporal probability.\n\nAfter training CNN, I predicted sliding window and save the result which has shape (n_data, n_windows, n_class) as a file.\n\n&gt; did you do this at fixed intervals or random crops\n\nI used fixed step size 21. The window size is 224, which is equivalent with about 5 sec. If I use smaller step size and more windows, the result becomes better, but the inference time is longer.\n\nI hope you can get good position in final stage. Good luck!",
      "votes": null
    },
    {
      "id": "551784",
      "postDate": "06/13/2019 05:40:59",
      "content": "<p>Thank you for clarifying on all points, very interesting 👍 Hope you rank high on stage 2 as well!</p>",
      "rawMarkdown": "Thank you for clarifying on all points, very interesting 👍 Hope you rank high on stage 2 as well!",
      "votes": null
    },
    {
      "id": "554812",
      "postDate": "06/18/2019 03:38:41",
      "content": "<p>Thanks for sharing, it's great ideas. One question - \"Average pseudo labels with original noisy labels by 0.75:0.25.\"\nCould I ask how you've come up with this mix ratio? I'm curious about your intuition :)</p>",
      "rawMarkdown": "Thanks for sharing, it's great ideas. One question - \"Average pseudo labels with original noisy labels by 0.75:0.25.\"\nCould I ask how you've come up with this mix ratio? I'm curious about your intuition :)",
      "votes": null
    },
    {
      "id": "554889",
      "postDate": "06/18/2019 06:36:48",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a> It has not so much reason. I tried to search around 1.0, 0.9, 0.8, 0.7.</p>",
      "rawMarkdown": "daisukelab It has not so much reason. I tried to search around 1.0, 0.9, 0.8, 0.7.",
      "votes": null
    },
    {
      "id": "555107",
      "postDate": "06/18/2019 12:44:21",
      "content": "<p>I see, it’s interesting. Thank you! </p>",
      "rawMarkdown": "I see, it’s interesting. Thank you!",
      "votes": null
    },
    {
      "id": "555313",
      "postDate": "06/18/2019 18:18:17",
      "content": "<p>I used the same strategy when I was playing with MixMatch. It was better than vanilla MixMatch, but you know MixMatch in this competition...</p>",
      "rawMarkdown": "I used the same strategy when I was playing with MixMatch. It was better than vanilla MixMatch, but you know MixMatch in this competition...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 551041,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "06/12/2019 09:26:30",
      "content": "<p>Your idea on RNN is splendid! Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 551100,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/12/2019 10:44:06",
      "content": "<p>Thanks for sharing <a href=\"/akirasosa\">@akirasosa</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 551453,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "06/12/2019 18:33:04",
      "content": "<p>Thank you for sharing, very nice solution <a href=\"/akirasosa\">@akirasosa</a> ! I also got my best results with a CNN + RNN (LSTM) and used it in my final submission so its great to see someone else in this competition did the same :) </p>\n\n<p>For input to the RNN though, did you use the penultimate layer of CNN (i.e. GAP layer extracted feature vector) or something else?</p>\n\n<p>For CNN sliding window, did you do this at fixed intervals or random crops (random starting points with fixed window size).</p>\n\n<p>Rakeld + ExtraTrees is very interesting, I didn't consider adding anything like that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 551705,
          "author_name": "akirasosa",
          "author_url": "",
          "post_date": "06/13/2019 03:12:40",
          "content": "<p><a href=\"/jamesrequa\">@jamesrequa</a></p>\n\n<blockquote>\n  <p>did you use the penultimate layer of CNN</p>\n</blockquote>\n\n<p>No. I used the probability predicted by the CNN. I tried some others as an input. But it was the best to just feeding the temporal probability.</p>\n\n<p>After training CNN, I predicted sliding window and save the result which has shape (n_data, n_windows, n_class) as a file.</p>\n\n<blockquote>\n  <p>did you do this at fixed intervals or random crops</p>\n</blockquote>\n\n<p>I used fixed step size 21. The window size is 224, which is equivalent with about 5 sec. If I use smaller step size and more windows, the result becomes better, but the inference time is longer.</p>\n\n<p>I hope you can get good position in final stage. Good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 551784,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "06/13/2019 05:40:59",
          "content": "<p>Thank you for clarifying on all points, very interesting 👍 Hope you rank high on stage 2 as well!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 554812,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/18/2019 03:38:41",
      "content": "<p>Thanks for sharing, it's great ideas. One question - \"Average pseudo labels with original noisy labels by 0.75:0.25.\"\nCould I ask how you've come up with this mix ratio? I'm curious about your intuition :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 554889,
          "author_name": "akirasosa",
          "author_url": "",
          "post_date": "06/18/2019 06:36:48",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> It has not so much reason. I tried to search around 1.0, 0.9, 0.8, 0.7.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 555107,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "06/18/2019 12:44:21",
          "content": "<p>I see, it’s interesting. Thank you! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 555313,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "06/18/2019 18:18:17",
          "content": "<p>I used the same strategy when I was playing with MixMatch. It was better than vanilla MixMatch, but you know MixMatch in this competition...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "550946": "Hi,\n\nAs I have made a slide about my solution, I would like to post it here. Thank organizer and staffs for this great competition.\n\nHere is the brief summary.\n\nInference:\n\n1. Raw audio to log mel.\n1. Apply sliding window to the log mel and predict each windows by CNN.\n1. Feed temporal prediction given step2 to RNN\n1. Feed temporal prediction given step2 to Rakeld + ExtraTrees\n1. Average step3 and step4.\n\nTraining:\n\n1. Train CNN by using only curated data.\n1. Apply sliding window to the log mel in noisy dataset and predict each windows by CNN.\n1. Feed temporal prediction given step2 to RNN. Then, I got pseudo labels for noisy dataset.\n1. Average pseudo labels with original noisy labels by 0.75:0.25.\n1. Train CNN by using all data.\n1. Train some models to follow the inference steps.\n\nDetails in Training CNN:\n\n* Densenet 121\n* Spec Aug\n* Mixup\n* CLR + SWA + Snapshot ensemble.\n\nI intensively tried SSL approach such as ICT, but it did not work.",
    "551041": "Your idea on RNN is splendid! Thanks for sharing!",
    "551100": "Thanks for sharing @akirasosa",
    "551453": "Thank you for sharing, very nice solution @akirasosa ! I also got my best results with a CNN + RNN (LSTM) and used it in my final submission so its great to see someone else in this competition did the same :) \n\nFor input to the RNN though, did you use the penultimate layer of CNN (i.e. GAP layer extracted feature vector) or something else?\n\nFor CNN sliding window, did you do this at fixed intervals or random crops (random starting points with fixed window size).\n\nRakeld + ExtraTrees is very interesting, I didn't consider adding anything like that.",
    "551705": "jamesrequa\n&gt; did you use the penultimate layer of CNN\n\nNo. I used the probability predicted by the CNN. I tried some others as an input. But it was the best to just feeding the temporal probability.\n\nAfter training CNN, I predicted sliding window and save the result which has shape (n_data, n_windows, n_class) as a file.\n\n&gt; did you do this at fixed intervals or random crops\n\nI used fixed step size 21. The window size is 224, which is equivalent with about 5 sec. If I use smaller step size and more windows, the result becomes better, but the inference time is longer.\n\nI hope you can get good position in final stage. Good luck!",
    "551784": "Thank you for clarifying on all points, very interesting 👍 Hope you rank high on stage 2 as well!",
    "554812": "Thanks for sharing, it's great ideas. One question - \"Average pseudo labels with original noisy labels by 0.75:0.25.\"\nCould I ask how you've come up with this mix ratio? I'm curious about your intuition :)",
    "554889": "daisukelab It has not so much reason. I tried to search around 1.0, 0.9, 0.8, 0.7.",
    "555107": "I see, it’s interesting. Thank you!",
    "555313": "I used the same strategy when I was playing with MixMatch. It was better than vanilla MixMatch, but you know MixMatch in this competition..."
  },
  "source": "meta"
}