{
  "id": 507824,
  "title": "Anybody tried multi-sample an audio file?",
  "url": "/competitions/birdclef-2024/discussion/507824",
  "author_name": "",
  "post_date": "2024-05-27T11:57:28.103086900Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi guys, I'm a kaggle rookie. I'm trying to sample an audio file in such a way:</p>\n<p>For a 10 second audio, I defined sample length to be 5 second and hop length to be 4 second, and it will produce 2 samples (0-5s, 4-9s).</p>\n<p>Does this method help the model to learn better, or other general suggestion? I'm using the following setting and I got 0.64 on LB:</p>\n<p>Backbone: eca_nfnet_l0, GeM pooling<br>\nDataset: </p>\n<ul>\n<li>Xeno-canto (with the over-sampling described above)</li>\n<li>Random split into 70% train, 30% test<br>\nArgumentation: </li>\n<li>Norm mel-spec (with 0-255 scaled, 3 channel)</li>\n<li>Gaussian Noise (uniform 0.0-0.1)</li>\n<li>Random Horizontal Flip (p=0.5)</li>\n<li>XY Mask</li>\n<li>Resize (256*256)<br>\nLoss: FocalBCE (with pos_weight 1/class_count)<br>\nOptimizer: Adam(weight_decay=1e-6)<br>\nOther things:</li>\n<li>torch.optim.swa_utils.update_bn</li>\n<li>Grad clipping norm (max_norm=1)</li>\n<li>Model Averaging</li>\n<li>LR Scheduler (Linear *0.5)</li>\n</ul>\n<p>(I feel I'm lost in the sea of tricks and I'm not sure what is working and what is not)</p>\n<p>Some inspiration from: <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66/notebook\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66/notebook</a></p>\n<p>Any help will be appreciated! XD</p>",
  "messages": [
    {
      "id": "2839139",
      "postDate": "05/27/2024 11:57:28",
      "content": "<p>Hi guys, I'm a kaggle rookie. I'm trying to sample an audio file in such a way:</p>\n<p>For a 10 second audio, I defined sample length to be 5 second and hop length to be 4 second, and it will produce 2 samples (0-5s, 4-9s).</p>\n<p>Does this method help the model to learn better, or other general suggestion? I'm using the following setting and I got 0.64 on LB:</p>\n<p>Backbone: eca_nfnet_l0, GeM pooling<br>\nDataset: </p>\n<ul>\n<li>Xeno-canto (with the over-sampling described above)</li>\n<li>Random split into 70% train, 30% test<br>\nArgumentation: </li>\n<li>Norm mel-spec (with 0-255 scaled, 3 channel)</li>\n<li>Gaussian Noise (uniform 0.0-0.1)</li>\n<li>Random Horizontal Flip (p=0.5)</li>\n<li>XY Mask</li>\n<li>Resize (256*256)<br>\nLoss: FocalBCE (with pos_weight 1/class_count)<br>\nOptimizer: Adam(weight_decay=1e-6)<br>\nOther things:</li>\n<li>torch.optim.swa_utils.update_bn</li>\n<li>Grad clipping norm (max_norm=1)</li>\n<li>Model Averaging</li>\n<li>LR Scheduler (Linear *0.5)</li>\n</ul>\n<p>(I feel I'm lost in the sea of tricks and I'm not sure what is working and what is not)</p>\n<p>Some inspiration from: <a href=\"https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66/notebook\" target=\"_blank\">https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66/notebook</a></p>\n<p>Any help will be appreciated! XD</p>",
      "rawMarkdown": "Hi guys, I'm a kaggle rookie. I'm trying to sample an audio file in such a way:\n\nFor a 10 second audio, I defined sample length to be 5 second and hop length to be 4 second, and it will produce 2 samples (0-5s, 4-9s).\n\nDoes this method help the model to learn better, or other general suggestion? I'm using the following setting and I got 0.64 on LB:\n\nBackbone: eca_nfnet_l0, GeM pooling\nDataset: \n- Xeno-canto (with the over-sampling described above)\n- Random split into 70% train, 30% test\nArgumentation: \n- Norm mel-spec (with 0-255 scaled, 3 channel)\n- Gaussian Noise (uniform 0.0-0.1)\n- Random Horizontal Flip (p=0.5)\n- XY Mask\n- Resize (256*256)\nLoss: FocalBCE (with pos_weight 1/class_count)\nOptimizer: Adam(weight_decay=1e-6)\nOther things:\n- torch.optim.swa_utils.update_bn\n- Grad clipping norm (max_norm=1)\n- Model Averaging\n- LR Scheduler (Linear *0.5)\n\n(I feel I'm lost in the sea of tricks and I'm not sure what is working and what is not)\n\nSome inspiration from: https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66/notebook\n\nAny help will be appreciated! XD",
      "votes": null
    },
    {
      "id": "2840249",
      "postDate": "05/28/2024 03:39:13",
      "content": "<p>oversample may not work well, you can check this discussion <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/502401\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/502401</a>.<br>\nI think it's good to start with a simple baseline and then add those tricks one by one.</p>",
      "rawMarkdown": "oversample may not work well, you can check this discussion https://www.kaggle.com/competitions/birdclef-2024/discussion/502401.\nI think it's good to start with a simple baseline and then add those tricks one by one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2840249,
      "author_name": "befunny",
      "author_url": "",
      "post_date": "05/28/2024 03:39:13",
      "content": "<p>oversample may not work well, you can check this discussion <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/502401\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2024/discussion/502401</a>.<br>\nI think it's good to start with a simple baseline and then add those tricks one by one.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2839139": "Hi guys, I'm a kaggle rookie. I'm trying to sample an audio file in such a way:\n\nFor a 10 second audio, I defined sample length to be 5 second and hop length to be 4 second, and it will produce 2 samples (0-5s, 4-9s).\n\nDoes this method help the model to learn better, or other general suggestion? I'm using the following setting and I got 0.64 on LB:\n\nBackbone: eca_nfnet_l0, GeM pooling\nDataset: \n- Xeno-canto (with the over-sampling described above)\n- Random split into 70% train, 30% test\nArgumentation: \n- Norm mel-spec (with 0-255 scaled, 3 channel)\n- Gaussian Noise (uniform 0.0-0.1)\n- Random Horizontal Flip (p=0.5)\n- XY Mask\n- Resize (256*256)\nLoss: FocalBCE (with pos_weight 1/class_count)\nOptimizer: Adam(weight_decay=1e-6)\nOther things:\n- torch.optim.swa_utils.update_bn\n- Grad clipping norm (max_norm=1)\n- Model Averaging\n- LR Scheduler (Linear *0.5)\n\n(I feel I'm lost in the sea of tricks and I'm not sure what is working and what is not)\n\nSome inspiration from: https://www.kaggle.com/code/salmanahmedtamu/training-0-65-0-66/notebook\n\nAny help will be appreciated! XD",
    "2840249": "oversample may not work well, you can check this discussion https://www.kaggle.com/competitions/birdclef-2024/discussion/502401.\nI think it's good to start with a simple baseline and then add those tricks one by one."
  },
  "source": "meta"
}