{
  "id": 95785,
  "title": "My write up.",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/my-write-up",
  "author_name": "",
  "post_date": "2019-06-15T03:55:43.620395400Z",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p>It's finally weekend, that I have time to write a quick post. It might be disappointing if you were come in here to see if there is any magic or tricks.</p>\n\n<p>As many of you know there is quite some difference even between the curated training set and the public test set. For example, those almost duplicated 'Mr. Peanut' clips were very bad and would confuse the model on the test set. For the noisy set, the differences are much larger. The slams seem to be basketball court(slam dunk?), the tap is mostly tip tap dance, the surf and wave clips are much windy. It also has notably much more human speakings in non-human sounds labeled clips. I also tried many methods to use the noisy set, but with limited success reaching only .72.</p>\n\n<p>My .76 model is fine-tuned to overfit the test set. I spent my memorial day holiday weekend, hand made a tuning set with cherry-picked clips from the curated set and noisy set that sounds most similar to the public test set and used this dataset to fine tune my models. It is kind of cheating, for that, it is equivalent to me sneak peek at the exam paper and giving my model a targeted list of past quiz questions to practice right before the exam. It will do well on this particular exam, but might not do well on a new one. </p>\n\n<p>I did not use this method as my final submission, because I don't think it solves the problem the competition is about. I did a similar thing in my previous <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/\">competiton</a> which I missed a top spot for not submitting a fine-tuned model.</p>\n\n<p>So quickly about my solution without the fine-tuning. It is Conv nets trained with mixup and random frequency/time masks on curated set and pseudo-labeled datasets. </p>\n\n<p>I have two pseudo-label sets. Here is how I generated them. For each audio clip, I run my model over sequences of windows on it. For each window, I calculate a ratio of the top k activations vs the sum of all activations. I call this ratio the signal-noise ratio. My first pseudo-label set consists of those windows with the highest ratios and associated most confident labels. My second pseudo-label set consists of full clips, I average activations of windows with highest ratios to be the labels. During training, I am zipping the curated training set and the two pseudo label sets, calculate loss separately and linearly combine them to do backward flow. </p>\n\n<p>On inference time, I do the same scanning through windows, calculating ratios, etc to generate predictions as well. I found it better than averaging random crops. But this may simply due to the fact that my model learned to focus on those strong short signals from the first pseudo label set.</p>",
  "messages": [
    {
      "id": "553050",
      "postDate": "06/15/2019 03:55:43",
      "content": "<p>It's finally weekend, that I have time to write a quick post. It might be disappointing if you were come in here to see if there is any magic or tricks.</p>\n\n<p>As many of you know there is quite some difference even between the curated training set and the public test set. For example, those almost duplicated 'Mr. Peanut' clips were very bad and would confuse the model on the test set. For the noisy set, the differences are much larger. The slams seem to be basketball court(slam dunk?), the tap is mostly tip tap dance, the surf and wave clips are much windy. It also has notably much more human speakings in non-human sounds labeled clips. I also tried many methods to use the noisy set, but with limited success reaching only .72.</p>\n\n<p>My .76 model is fine-tuned to overfit the test set. I spent my memorial day holiday weekend, hand made a tuning set with cherry-picked clips from the curated set and noisy set that sounds most similar to the public test set and used this dataset to fine tune my models. It is kind of cheating, for that, it is equivalent to me sneak peek at the exam paper and giving my model a targeted list of past quiz questions to practice right before the exam. It will do well on this particular exam, but might not do well on a new one. </p>\n\n<p>I did not use this method as my final submission, because I don't think it solves the problem the competition is about. I did a similar thing in my previous <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/\">competiton</a> which I missed a top spot for not submitting a fine-tuned model.</p>\n\n<p>So quickly about my solution without the fine-tuning. It is Conv nets trained with mixup and random frequency/time masks on curated set and pseudo-labeled datasets. </p>\n\n<p>I have two pseudo-label sets. Here is how I generated them. For each audio clip, I run my model over sequences of windows on it. For each window, I calculate a ratio of the top k activations vs the sum of all activations. I call this ratio the signal-noise ratio. My first pseudo-label set consists of those windows with the highest ratios and associated most confident labels. My second pseudo-label set consists of full clips, I average activations of windows with highest ratios to be the labels. During training, I am zipping the curated training set and the two pseudo label sets, calculate loss separately and linearly combine them to do backward flow. </p>\n\n<p>On inference time, I do the same scanning through windows, calculating ratios, etc to generate predictions as well. I found it better than averaging random crops. But this may simply due to the fact that my model learned to focus on those strong short signals from the first pseudo label set.</p>",
      "rawMarkdown": "It's finally weekend, that I have time to write a quick post. It might be disappointing if you were come in here to see if there is any magic or tricks.\n\nAs many of you know there is quite some difference even between the curated training set and the public test set. For example, those almost duplicated 'Mr. Peanut' clips were very bad and would confuse the model on the test set. For the noisy set, the differences are much larger. The slams seem to be basketball court(slam dunk?), the tap is mostly tip tap dance, the surf and wave clips are much windy. It also has notably much more human speakings in non-human sounds labeled clips. I also tried many methods to use the noisy set, but with limited success reaching only .72.\n\nMy .76 model is fine-tuned to overfit the test set. I spent my memorial day holiday weekend, hand made a tuning set with cherry-picked clips from the curated set and noisy set that sounds most similar to the public test set and used this dataset to fine tune my models. It is kind of cheating, for that, it is equivalent to me sneak peek at the exam paper and giving my model a targeted list of past quiz questions to practice right before the exam. It will do well on this particular exam, but might not do well on a new one. \n\nI did not use this method as my final submission, because I don't think it solves the problem the competition is about. I did a similar thing in my previous [competiton](https://www.kaggle.com/c/inclusive-images-challenge/) which I missed a top spot for not submitting a fine-tuned model.\n\nSo quickly about my solution without the fine-tuning. It is Conv nets trained with mixup and random frequency/time masks on curated set and pseudo-labeled datasets. \n\nI have two pseudo-label sets. Here is how I generated them. For each audio clip, I run my model over sequences of windows on it. For each window, I calculate a ratio of the top k activations vs the sum of all activations. I call this ratio the signal-noise ratio. My first pseudo-label set consists of those windows with the highest ratios and associated most confident labels. My second pseudo-label set consists of full clips, I average activations of windows with highest ratios to be the labels. During training, I am zipping the curated training set and the two pseudo label sets, calculate loss separately and linearly combine them to do backward flow. \n\nOn inference time, I do the same scanning through windows, calculating ratios, etc to generate predictions as well. I found it better than averaging random crops. But this may simply due to the fact that my model learned to focus on those strong short signals from the first pseudo label set.",
      "votes": null
    },
    {
      "id": "553067",
      "postDate": "06/15/2019 04:33:03",
      "content": "<p>So..., How much does your final submission score on the public LB? 0.72? Anyway, your public LB score accelerated many teams to get a higher score. GJ.</p>",
      "rawMarkdown": "So..., How much does your final submission score on the public LB? 0.72? Anyway, your public LB score accelerated many teams to get a higher score. GJ.",
      "votes": null
    },
    {
      "id": "553172",
      "postDate": "06/15/2019 08:57:33",
      "content": "<p>thx</p>",
      "rawMarkdown": "thx",
      "votes": null
    },
    {
      "id": "553237",
      "postDate": "06/15/2019 11:05:47",
      "content": "<p>a wise choice.</p>",
      "rawMarkdown": "a wise choice.",
      "votes": null
    },
    {
      "id": "553252",
      "postDate": "06/15/2019 11:37:49",
      "content": "<p>.723 to be accurate. With the limited time I had after work and family, I am happy with silver range scores. </p>",
      "rawMarkdown": ".723 to be accurate. With the limited time I had after work and family, I am happy with silver range scores.",
      "votes": null
    },
    {
      "id": "553255",
      "postDate": "06/15/2019 11:38:49",
      "content": "<p>Yeah, that .76 model probably overfit  like crazy too. </p>",
      "rawMarkdown": "Yeah, that .76 model probably overfit  like crazy too.",
      "votes": null
    },
    {
      "id": "554808",
      "postDate": "06/18/2019 03:28:28",
      "content": "<p>Thank you for sharing your tricks!\nFine tuning to test set will work as workaround of ML system, I'm happy to know that it works greatly.</p>",
      "rawMarkdown": "Thank you for sharing your tricks!\nFine tuning to test set will work as workaround of ML system, I'm happy to know that it works greatly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 553067,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "06/15/2019 04:33:03",
      "content": "<p>So..., How much does your final submission score on the public LB? 0.72? Anyway, your public LB score accelerated many teams to get a higher score. GJ.</p>",
      "votes": null,
      "replies": [
        {
          "id": 553252,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "06/15/2019 11:37:49",
          "content": "<p>.723 to be accurate. With the limited time I had after work and family, I am happy with silver range scores. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 553172,
      "author_name": "action",
      "author_url": "",
      "post_date": "06/15/2019 08:57:33",
      "content": "<p>thx</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 553237,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "06/15/2019 11:05:47",
      "content": "<p>a wise choice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 553255,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "06/15/2019 11:38:49",
          "content": "<p>Yeah, that .76 model probably overfit  like crazy too. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 554808,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/18/2019 03:28:28",
      "content": "<p>Thank you for sharing your tricks!\nFine tuning to test set will work as workaround of ML system, I'm happy to know that it works greatly.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "553050": "It's finally weekend, that I have time to write a quick post. It might be disappointing if you were come in here to see if there is any magic or tricks.\n\nAs many of you know there is quite some difference even between the curated training set and the public test set. For example, those almost duplicated 'Mr. Peanut' clips were very bad and would confuse the model on the test set. For the noisy set, the differences are much larger. The slams seem to be basketball court(slam dunk?), the tap is mostly tip tap dance, the surf and wave clips are much windy. It also has notably much more human speakings in non-human sounds labeled clips. I also tried many methods to use the noisy set, but with limited success reaching only .72.\n\nMy .76 model is fine-tuned to overfit the test set. I spent my memorial day holiday weekend, hand made a tuning set with cherry-picked clips from the curated set and noisy set that sounds most similar to the public test set and used this dataset to fine tune my models. It is kind of cheating, for that, it is equivalent to me sneak peek at the exam paper and giving my model a targeted list of past quiz questions to practice right before the exam. It will do well on this particular exam, but might not do well on a new one. \n\nI did not use this method as my final submission, because I don't think it solves the problem the competition is about. I did a similar thing in my previous [competiton](https://www.kaggle.com/c/inclusive-images-challenge/) which I missed a top spot for not submitting a fine-tuned model.\n\nSo quickly about my solution without the fine-tuning. It is Conv nets trained with mixup and random frequency/time masks on curated set and pseudo-labeled datasets. \n\nI have two pseudo-label sets. Here is how I generated them. For each audio clip, I run my model over sequences of windows on it. For each window, I calculate a ratio of the top k activations vs the sum of all activations. I call this ratio the signal-noise ratio. My first pseudo-label set consists of those windows with the highest ratios and associated most confident labels. My second pseudo-label set consists of full clips, I average activations of windows with highest ratios to be the labels. During training, I am zipping the curated training set and the two pseudo label sets, calculate loss separately and linearly combine them to do backward flow. \n\nOn inference time, I do the same scanning through windows, calculating ratios, etc to generate predictions as well. I found it better than averaging random crops. But this may simply due to the fact that my model learned to focus on those strong short signals from the first pseudo label set.",
    "553067": "So..., How much does your final submission score on the public LB? 0.72? Anyway, your public LB score accelerated many teams to get a higher score. GJ.",
    "553172": "thx",
    "553237": "a wise choice.",
    "553252": ".723 to be accurate. With the limited time I had after work and family, I am happy with silver range scores.",
    "553255": "Yeah, that .76 model probably overfit  like crazy too.",
    "554808": "Thank you for sharing your tricks!\nFine tuning to test set will work as workaround of ML system, I'm happy to know that it works greatly."
  },
  "source": "meta"
}