{
  "id": 167611,
  "title": "Different aggregation over feature map and their effects",
  "url": "/competitions/birdsong-recognition/discussion/167611",
  "author_name": "",
  "post_date": "2020-07-17T09:36:39.135998200Z",
  "votes": 32,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've tried two types of aggregation method on the output of CNN feature extractor: </p>\n\n<ol>\n<li>AvgMaxPool</li>\n<li>Attention</li>\n</ol>\n\n<p>and found that AvgMaxPool has an effect to suppress high probability put on feature maps while Attention doesn't. Therefore, the model using AvgMaxPool aggregation get better score when the global threshold is 0.5 compared to that with threshold 0.8, and the model using Attention aggregation get better score when the global threshold is 0.8 compared to that with threshold 0.5.</p>\n\n<p>Another effect, which I believe, that Attention aggregation has is that Attention aggregation basically has some sort of effect to capture correlation (co-occurence) between different call events. Therefore it may have some positive effect when multiple calls are present in that audio clip. </p>\n\n<p>I'm now trying to combine different aggregation methods in a single model aiming at complementary effect of each method.</p>",
  "messages": [
    {
      "id": "932791",
      "postDate": "07/17/2020 09:36:39",
      "content": "<p>I've tried two types of aggregation method on the output of CNN feature extractor: </p>\n\n<ol>\n<li>AvgMaxPool</li>\n<li>Attention</li>\n</ol>\n\n<p>and found that AvgMaxPool has an effect to suppress high probability put on feature maps while Attention doesn't. Therefore, the model using AvgMaxPool aggregation get better score when the global threshold is 0.5 compared to that with threshold 0.8, and the model using Attention aggregation get better score when the global threshold is 0.8 compared to that with threshold 0.5.</p>\n\n<p>Another effect, which I believe, that Attention aggregation has is that Attention aggregation basically has some sort of effect to capture correlation (co-occurence) between different call events. Therefore it may have some positive effect when multiple calls are present in that audio clip. </p>\n\n<p>I'm now trying to combine different aggregation methods in a single model aiming at complementary effect of each method.</p>",
      "rawMarkdown": "I've tried two types of aggregation method on the output of CNN feature extractor: \n\n1. AvgMaxPool\n2. Attention\n\nand found that AvgMaxPool has an effect to suppress high probability put on feature maps while Attention doesn't. Therefore, the model using AvgMaxPool aggregation get better score when the global threshold is 0.5 compared to that with threshold 0.8, and the model using Attention aggregation get better score when the global threshold is 0.8 compared to that with threshold 0.5.\n\nAnother effect, which I believe, that Attention aggregation has is that Attention aggregation basically has some sort of effect to capture correlation (co-occurence) between different call events. Therefore it may have some positive effect when multiple calls are present in that audio clip. \n\nI'm now trying to combine different aggregation methods in a single model aiming at complementary effect of each method.",
      "votes": null
    },
    {
      "id": "938140",
      "postDate": "07/21/2020 11:19:16",
      "content": "<p>thanks for sharing .it looks interesting. can you share some reference to AvgMaxPool concept. i know about maxpool and avgpool. is AvgMaxPool is a combination of both by applying avg and them taking the max.</p>",
      "rawMarkdown": "thanks for sharing .it looks interesting. can you share some reference to AvgMaxPool concept. i know about maxpool and avgpool. is AvgMaxPool is a combination of both by applying avg and them taking the max.",
      "votes": null
    },
    {
      "id": "939036",
      "postDate": "07/22/2020 01:49:10",
      "content": "<p>Thank you for pointing out, since this was just a typo...I mean <code>AdaptiveMaxPool</code> of PyTorch's layer</p>\n\n<p>But I also think adding <code>Avg</code> result to <code>Max</code> result is either fine or better. I once got improvement with that in FAT2019 competition.</p>",
      "rawMarkdown": "Thank you for pointing out, since this was just a typo...I mean `AdaptiveMaxPool` of PyTorch's layer\n\nBut I also think adding `Avg` result to `Max` result is either fine or better. I once got improvement with that in FAT2019 competition.",
      "votes": null
    },
    {
      "id": "952075",
      "postDate": "07/30/2020 15:39:07",
      "content": "<p>i used column-wise attention, and the results is like this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe08774dd142032705270fa6fb6a9c%2FSelection_045.png?generation=1596123430780582&amp;alt=media\" alt=\"\"></p>\n\n<p>(green: true class attention, red: predicted class attention)</p>\n\n<p>you can further follows this work to \"hand correct attention\" to improve your results\n<a href=\"http://mprg.jp/research/abn_e\">http://mprg.jp/research/abn_e</a>\n<a href=\"https://github.com/machine-perception-robotics-group/attention_branch_network\">https://github.com/machine-perception-robotics-group/attention_branch_network</a>\n<img src=\"http://mprg.jp/wp/wp-content/uploads/2019/06/abn_mitsuhara_1.jpg\" alt=\"\"></p>",
      "rawMarkdown": "i used column-wise attention, and the results is like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe08774dd142032705270fa6fb6a9c%2FSelection_045.png?generation=1596123430780582&amp;alt=media)\n\n(green: true class attention, red: predicted class attention)\n\nyou can further follows this work to \"hand correct attention\" to improve your results\nhttp://mprg.jp/research/abn_e\nhttps://github.com/machine-perception-robotics-group/attention_branch_network\n![](http://mprg.jp/wp/wp-content/uploads/2019/06/abn_mitsuhara_1.jpg)",
      "votes": null
    },
    {
      "id": "952563",
      "postDate": "07/31/2020 03:48:16",
      "content": "<p>Nice visualization!\nI've also tried weakly supervised SED model and visualized the prediction: it outputs class probability for each time segment and confirmed it output high probability around sound event (I just manually checked the correspondence between predicted label and sound events I hear) but sometimes put prediction with long hem that covers no audio region, so manually correcting that seems good (although I don't like the idea of hand-labeling so much)</p>",
      "rawMarkdown": "Nice visualization!\nI've also tried weakly supervised SED model and visualized the prediction: it outputs class probability for each time segment and confirmed it output high probability around sound event (I just manually checked the correspondence between predicted label and sound events I hear) but sometimes put prediction with long hem that covers no audio region, so manually correcting that seems good (although I don't like the idea of hand-labeling so much)",
      "votes": null
    },
    {
      "id": "952574",
      "postDate": "07/31/2020 03:57:22",
      "content": "<p>i am implementing facebook \"contrastive learning using cluster\"\n<a href=\"https://ai.facebook.com/blog/high-performance-self-supervised-image-classification-with-contrastive-clustering/\">https://ai.facebook.com/blog/high-performance-self-supervised-image-classification-with-contrastive-clustering/</a></p>\n\n<p>it shows that just by using 10% label of image net, you get same performance as 100% label.</p>\n\n<p>cluster is good because each bird only as a few key song pattern</p>",
      "rawMarkdown": "i am implementing facebook \"contrastive learning using cluster\"\nhttps://ai.facebook.com/blog/high-performance-self-supervised-image-classification-with-contrastive-clustering/\n\nit shows that just by using 10% label of image net, you get same performance as 100% label.\n\ncluster is good because each bird only as a few key song pattern",
      "votes": null
    },
    {
      "id": "952580",
      "postDate": "07/31/2020 04:06:59",
      "content": "<p>Seems nice and I'm very excited we may see the wave of self-supervised learning here too.</p>",
      "rawMarkdown": "Seems nice and I'm very excited we may see the wave of self-supervised learning here too.",
      "votes": null
    },
    {
      "id": "954091",
      "postDate": "08/01/2020 11:22:22",
      "content": "<blockquote>\n  <p>each bird only as a few key song patterns</p>\n</blockquote>\n\n<p>Actually, bird vocalizations can vary wildly even within the same species. </p>\n\n<p>Check out \"grtgra\" for example (<a href=\"https://en.wikipedia.org/wiki/Great-tailed_grackle\">https://en.wikipedia.org/wiki/Great-tailed_grackle</a>).</p>\n\n<blockquote>\n  <p>Great-tailed grackles have an unusually large repertoire of vocalizations that are used year-round. The sounds range from \"sweet, tinkling notes\" to a \"rusty gate hinge\".[7] Males use a wider variety of vocalization types, while females engage mostly in \"chatter\", however there is a report of a female performing the \"territorial song\"</p>\n</blockquote>",
      "rawMarkdown": "&gt; each bird only as a few key song patterns\n\nActually, bird vocalizations can vary wildly even within the same species. \n\nCheck out \"grtgra\" for example ([https://en.wikipedia.org/wiki/Great-tailed_grackle](https://en.wikipedia.org/wiki/Great-tailed_grackle)).\n\n&gt; Great-tailed grackles have an unusually large repertoire of vocalizations that are used year-round. The sounds range from \"sweet, tinkling notes\" to a \"rusty gate hinge\".[7] Males use a wider variety of vocalization types, while females engage mostly in \"chatter\", however there is a report of a female performing the \"territorial song\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 938140,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "07/21/2020 11:19:16",
      "content": "<p>thanks for sharing .it looks interesting. can you share some reference to AvgMaxPool concept. i know about maxpool and avgpool. is AvgMaxPool is a combination of both by applying avg and them taking the max.</p>",
      "votes": null,
      "replies": [
        {
          "id": 939036,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/22/2020 01:49:10",
          "content": "<p>Thank you for pointing out, since this was just a typo...I mean <code>AdaptiveMaxPool</code> of PyTorch's layer</p>\n\n<p>But I also think adding <code>Avg</code> result to <code>Max</code> result is either fine or better. I once got improvement with that in FAT2019 competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 952075,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/30/2020 15:39:07",
      "content": "<p>i used column-wise attention, and the results is like this:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe08774dd142032705270fa6fb6a9c%2FSelection_045.png?generation=1596123430780582&amp;alt=media\" alt=\"\"></p>\n\n<p>(green: true class attention, red: predicted class attention)</p>\n\n<p>you can further follows this work to \"hand correct attention\" to improve your results\n<a href=\"http://mprg.jp/research/abn_e\">http://mprg.jp/research/abn_e</a>\n<a href=\"https://github.com/machine-perception-robotics-group/attention_branch_network\">https://github.com/machine-perception-robotics-group/attention_branch_network</a>\n<img src=\"http://mprg.jp/wp/wp-content/uploads/2019/06/abn_mitsuhara_1.jpg\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 952563,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/31/2020 03:48:16",
          "content": "<p>Nice visualization!\nI've also tried weakly supervised SED model and visualized the prediction: it outputs class probability for each time segment and confirmed it output high probability around sound event (I just manually checked the correspondence between predicted label and sound events I hear) but sometimes put prediction with long hem that covers no audio region, so manually correcting that seems good (although I don't like the idea of hand-labeling so much)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 952574,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "07/31/2020 03:57:22",
          "content": "<p>i am implementing facebook \"contrastive learning using cluster\"\n<a href=\"https://ai.facebook.com/blog/high-performance-self-supervised-image-classification-with-contrastive-clustering/\">https://ai.facebook.com/blog/high-performance-self-supervised-image-classification-with-contrastive-clustering/</a></p>\n\n<p>it shows that just by using 10% label of image net, you get same performance as 100% label.</p>\n\n<p>cluster is good because each bird only as a few key song pattern</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 952580,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/31/2020 04:06:59",
          "content": "<p>Seems nice and I'm very excited we may see the wave of self-supervised learning here too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 954091,
          "author_name": "ramarlina",
          "author_url": "",
          "post_date": "08/01/2020 11:22:22",
          "content": "<blockquote>\n  <p>each bird only as a few key song patterns</p>\n</blockquote>\n\n<p>Actually, bird vocalizations can vary wildly even within the same species. </p>\n\n<p>Check out \"grtgra\" for example (<a href=\"https://en.wikipedia.org/wiki/Great-tailed_grackle\">https://en.wikipedia.org/wiki/Great-tailed_grackle</a>).</p>\n\n<blockquote>\n  <p>Great-tailed grackles have an unusually large repertoire of vocalizations that are used year-round. The sounds range from \"sweet, tinkling notes\" to a \"rusty gate hinge\".[7] Males use a wider variety of vocalization types, while females engage mostly in \"chatter\", however there is a report of a female performing the \"territorial song\"</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "932791": "I've tried two types of aggregation method on the output of CNN feature extractor: \n\n1. AvgMaxPool\n2. Attention\n\nand found that AvgMaxPool has an effect to suppress high probability put on feature maps while Attention doesn't. Therefore, the model using AvgMaxPool aggregation get better score when the global threshold is 0.5 compared to that with threshold 0.8, and the model using Attention aggregation get better score when the global threshold is 0.8 compared to that with threshold 0.5.\n\nAnother effect, which I believe, that Attention aggregation has is that Attention aggregation basically has some sort of effect to capture correlation (co-occurence) between different call events. Therefore it may have some positive effect when multiple calls are present in that audio clip. \n\nI'm now trying to combine different aggregation methods in a single model aiming at complementary effect of each method.",
    "938140": "thanks for sharing .it looks interesting. can you share some reference to AvgMaxPool concept. i know about maxpool and avgpool. is AvgMaxPool is a combination of both by applying avg and them taking the max.",
    "939036": "Thank you for pointing out, since this was just a typo...I mean `AdaptiveMaxPool` of PyTorch's layer\n\nBut I also think adding `Avg` result to `Max` result is either fine or better. I once got improvement with that in FAT2019 competition.",
    "952075": "i used column-wise attention, and the results is like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc3fe08774dd142032705270fa6fb6a9c%2FSelection_045.png?generation=1596123430780582&amp;alt=media)\n\n(green: true class attention, red: predicted class attention)\n\nyou can further follows this work to \"hand correct attention\" to improve your results\nhttp://mprg.jp/research/abn_e\nhttps://github.com/machine-perception-robotics-group/attention_branch_network\n![](http://mprg.jp/wp/wp-content/uploads/2019/06/abn_mitsuhara_1.jpg)",
    "952563": "Nice visualization!\nI've also tried weakly supervised SED model and visualized the prediction: it outputs class probability for each time segment and confirmed it output high probability around sound event (I just manually checked the correspondence between predicted label and sound events I hear) but sometimes put prediction with long hem that covers no audio region, so manually correcting that seems good (although I don't like the idea of hand-labeling so much)",
    "952574": "i am implementing facebook \"contrastive learning using cluster\"\nhttps://ai.facebook.com/blog/high-performance-self-supervised-image-classification-with-contrastive-clustering/\n\nit shows that just by using 10% label of image net, you get same performance as 100% label.\n\ncluster is good because each bird only as a few key song pattern",
    "952580": "Seems nice and I'm very excited we may see the wave of self-supervised learning here too.",
    "954091": "&gt; each bird only as a few key song patterns\n\nActually, bird vocalizations can vary wildly even within the same species. \n\nCheck out \"grtgra\" for example ([https://en.wikipedia.org/wiki/Great-tailed_grackle](https://en.wikipedia.org/wiki/Great-tailed_grackle)).\n\n&gt; Great-tailed grackles have an unusually large repertoire of vocalizations that are used year-round. The sounds range from \"sweet, tinkling notes\" to a \"rusty gate hinge\".[7] Males use a wider variety of vocalization types, while females engage mostly in \"chatter\", however there is a report of a female performing the \"territorial song\""
  },
  "source": "meta"
}