{
  "id": 216436,
  "title": " Convolution-augmented Transformer for Species Audio Detection",
  "url": "/competitions/rfcx-species-audio-detection/discussion/216436",
  "author_name": "",
  "post_date": "2021-02-02T17:14:27.088057900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Conformer, a recently published paper:<br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/pdf/2005.08100.pdf</a></p>\n<blockquote>\n  <p>Transformer models are good at capturing content-based global interactions, <br>\n  CNNs exploit local features effectively. <br>\n  they combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence.</p>\n</blockquote>\n<p>I haven't tried this method, but it may useful.</p>",
  "messages": [
    {
      "id": "1182998",
      "postDate": "02/02/2021 17:14:27",
      "content": "<p>Conformer, a recently published paper:<br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/pdf/2005.08100.pdf</a></p>\n<blockquote>\n  <p>Transformer models are good at capturing content-based global interactions, <br>\n  CNNs exploit local features effectively. <br>\n  they combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence.</p>\n</blockquote>\n<p>I haven't tried this method, but it may useful.</p>",
      "rawMarkdown": "Conformer, a recently published paper:\n[https://arxiv.org/pdf/2005.08100.pdf](url)\n> Transformer models are good at capturing content-based global interactions, \n> CNNs exploit local features effectively. \nthey combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence.\n\nI haven't tried this method, but it may useful.",
      "votes": null
    },
    {
      "id": "1183418",
      "postDate": "02/03/2021 01:25:34",
      "content": "<p>I tried the Conformer model with the SED method.<br>\nAlthough the score was higher than the method using a simple fully connected layer, training was difficult and the training time required was also increased.<br>\nI referred to the implementation <a href=\"http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Miyazaki_108.pdf\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "I tried the Conformer model with the SED method.\nAlthough the score was higher than the method using a simple fully connected layer, training was difficult and the training time required was also increased.\nI referred to the implementation [here](http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Miyazaki_108.pdf).",
      "votes": null
    },
    {
      "id": "1183532",
      "postDate": "02/03/2021 04:25:30",
      "content": "<p>Thank you for sharing a good paper!</p>",
      "rawMarkdown": "Thank you for sharing a good paper!",
      "votes": null
    },
    {
      "id": "1187071",
      "postDate": "02/05/2021 07:47:15",
      "content": "<p>Since I also believe that time axis information is important, so I am trying to reproduce and implement the winning model of DCASE 2020 Task 4 for use in this competition.<br>\n<a href=\"http://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Miyazaki_92.pdf\" target=\"_blank\">http://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Miyazaki_92.pdf</a></p>\n<p>During training, we perform local feature extraction by convolution in the frequency direction, and global feature extraction by transformer (self-attention), but it is difficult to train well, and LB is not improved much (local lwlrap: 0.87, public lb: 0.66). 669).</p>\n<p>I would like to test the effectiveness of the transformer on sound in this competition, so I will report back here if this method works well :(</p>",
      "rawMarkdown": "Since I also believe that time axis information is important, so I am trying to reproduce and implement the winning model of DCASE 2020 Task 4 for use in this competition.\nhttp://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Miyazaki_92.pdf\n\nDuring training, we perform local feature extraction by convolution in the frequency direction, and global feature extraction by transformer (self-attention), but it is difficult to train well, and LB is not improved much (local lwlrap: 0.87, public lb: 0.66). 669).\n\nI would like to test the effectiveness of the transformer on sound in this competition, so I will report back here if this method works well :(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1183418,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "02/03/2021 01:25:34",
      "content": "<p>I tried the Conformer model with the SED method.<br>\nAlthough the score was higher than the method using a simple fully connected layer, training was difficult and the training time required was also increased.<br>\nI referred to the implementation <a href=\"http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Miyazaki_108.pdf\" target=\"_blank\">here</a>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1183532,
      "author_name": "kbh0287",
      "author_url": "",
      "post_date": "02/03/2021 04:25:30",
      "content": "<p>Thank you for sharing a good paper!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1187071,
      "author_name": "tatsuyatakahashi",
      "author_url": "",
      "post_date": "02/05/2021 07:47:15",
      "content": "<p>Since I also believe that time axis information is important, so I am trying to reproduce and implement the winning model of DCASE 2020 Task 4 for use in this competition.<br>\n<a href=\"http://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Miyazaki_92.pdf\" target=\"_blank\">http://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Miyazaki_92.pdf</a></p>\n<p>During training, we perform local feature extraction by convolution in the frequency direction, and global feature extraction by transformer (self-attention), but it is difficult to train well, and LB is not improved much (local lwlrap: 0.87, public lb: 0.66). 669).</p>\n<p>I would like to test the effectiveness of the transformer on sound in this competition, so I will report back here if this method works well :(</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1182998": "Conformer, a recently published paper:\n[https://arxiv.org/pdf/2005.08100.pdf](url)\n> Transformer models are good at capturing content-based global interactions, \n> CNNs exploit local features effectively. \nthey combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence.\n\nI haven't tried this method, but it may useful.",
    "1183418": "I tried the Conformer model with the SED method.\nAlthough the score was higher than the method using a simple fully connected layer, training was difficult and the training time required was also increased.\nI referred to the implementation [here](http://dcase.community/documents/challenge2020/technical_reports/DCASE2020_Miyazaki_108.pdf).",
    "1183532": "Thank you for sharing a good paper!",
    "1187071": "Since I also believe that time axis information is important, so I am trying to reproduce and implement the winning model of DCASE 2020 Task 4 for use in this competition.\nhttp://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Miyazaki_92.pdf\n\nDuring training, we perform local feature extraction by convolution in the frequency direction, and global feature extraction by transformer (self-attention), but it is difficult to train well, and LB is not improved much (local lwlrap: 0.87, public lb: 0.66). 669).\n\nI would like to test the effectiveness of the transformer on sound in this competition, so I will report back here if this method works well :("
  },
  "source": "meta"
}