{
  "id": 243790,
  "title": "How can I successfully train SED or custom pooling model ?",
  "url": "/competitions/birdclef-2021/discussion/243790",
  "author_name": "",
  "post_date": "2021-06-04T02:36:07.297988Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am sorry for asking very basic question, while others are publishing great solutions now.<br>\nDuring this competition, I sometimes tried to train SED or custom pooling (eg. freq axis only) models for about 30 epochs on ColabPro. However, all of them ended up just all zero f1 score. Did anyone have similar result ?</p>\n<p>Because of this, I used normal CNN models. By using some tricks (rating-based weight loss, background noise from train_soundscape, fine-tuning from non-background noise models) I got bronze model, but I do want to know the key to train SED model for my future work🙏</p>",
  "messages": [
    {
      "id": "1335082",
      "postDate": "06/04/2021 02:36:07",
      "content": "<p>I am sorry for asking very basic question, while others are publishing great solutions now.<br>\nDuring this competition, I sometimes tried to train SED or custom pooling (eg. freq axis only) models for about 30 epochs on ColabPro. However, all of them ended up just all zero f1 score. Did anyone have similar result ?</p>\n<p>Because of this, I used normal CNN models. By using some tricks (rating-based weight loss, background noise from train_soundscape, fine-tuning from non-background noise models) I got bronze model, but I do want to know the key to train SED model for my future work🙏</p>",
      "rawMarkdown": "I am sorry for asking very basic question, while others are publishing great solutions now.\nDuring this competition, I sometimes tried to train SED or custom pooling (eg. freq axis only) models for about 30 epochs on ColabPro. However, all of them ended up just all zero f1 score. Did anyone have similar result ?\n\nBecause of this, I used normal CNN models. By using some tricks (rating-based weight loss, background noise from train_soundscape, fine-tuning from non-background noise models) I got bronze model, but I do want to know the key to train SED model for my future work🙏",
      "votes": null
    },
    {
      "id": "1335113",
      "postDate": "06/04/2021 02:59:00",
      "content": "<p>In my experiments, the SED model was sensitive to learning rate and batch size. <br>\nAlso, some architectures (rexnet150, skresnet-50, etc.) produced only nobird (probably due to lack of tuning) in train_soundscape despite good CV on train_short_audio.</p>",
      "rawMarkdown": "In my experiments, the SED model was sensitive to learning rate and batch size. \nAlso, some architectures (rexnet150, skresnet-50, etc.) produced only nobird (probably due to lack of tuning) in train_soundscape despite good CV on train_short_audio.",
      "votes": null
    },
    {
      "id": "1335915",
      "postDate": "06/04/2021 14:05:48",
      "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Thank you for kind reply 🙌<br>\nI learned a lot from your 4th place solution as well.<br>\nI finally got it and will try to re-train my model after changing some hyper parameters.</p>",
      "rawMarkdown": "tattaka Thank you for kind reply 🙌\nI learned a lot from your 4th place solution as well.\nI finally got it and will try to re-train my model after changing some hyper parameters.",
      "votes": null
    },
    {
      "id": "1336009",
      "postDate": "06/04/2021 15:27:01",
      "content": "<p>Before you start building your network architecture, first thing you need to do is to verify your input data into the network if an input (x) corresponds to a label (y). In case of dense prediction, make sure that the ground-truth label (y) is properly encoded to label indexes (or one-hot encoding). If not, the training won’t work.</p>\n<p>Then, if your dataset in your problem domain is similar to ImageNet dataset, use a pre-trained model on this dataset. The most widely used pre-trained models are VGG net, ResNet, DenseNet or Xception etc. There are many layer architectures, for instance, VGG (19 and 16 layers), ResNet (152, 101, 50 layers or less), DenseNet (201, 169 and 121 layers). Note: Do not try searching hyper-parameters by using more layers nets (e.g. VGG-19, ResNet-152 or DenseNet-201 layers net because it is computationally expensive), use less layers nets instead (e.g. VGG-16, ResNet-50 or DenseNet-121 layers). Pick one pre-trained model that you think it gives the best performance with your hyper-parameters (say ResNet-50 layers). After you obtained the optimal hyper parameters, just select the same but more layers net (say ResNet-101 or ResNet-152 layers) to increase the accuracy.</p>\n<p>Always use normalization layers in your network.  InstanceNormalization gives slightly performance improvements if they use a small batch-size. Or you may also try GroupNormalization.</p>\n<p>Use SpatialDropout after a features concatenation if you have two or more convolution layers (say Li) operate on the same input (say F). Always shuffle your training data, both before training and during training, in case you don’t take benefit from temporal data. This may help improving your network performance. Choose a right optimizer. There are many popular adaptive optimizers such as Adam, Adagrad, Adadelta, or RMSprop etc. SGD+momentum is widely used in various problem domains. <br>\nRead More here: <a href=\"https://towardsdatascience.com/a-bunch-of-tips-and-tricks-for-training-deep-neural-networks-3ca24c31ddc8\" target=\"_blank\">Custom pooling model</a></p>",
      "rawMarkdown": "Before you start building your network architecture, first thing you need to do is to verify your input data into the network if an input (x) corresponds to a label (y). In case of dense prediction, make sure that the ground-truth label (y) is properly encoded to label indexes (or one-hot encoding). If not, the training won’t work.\n\nThen, if your dataset in your problem domain is similar to ImageNet dataset, use a pre-trained model on this dataset. The most widely used pre-trained models are VGG net, ResNet, DenseNet or Xception etc. There are many layer architectures, for instance, VGG (19 and 16 layers), ResNet (152, 101, 50 layers or less), DenseNet (201, 169 and 121 layers). Note: Do not try searching hyper-parameters by using more layers nets (e.g. VGG-19, ResNet-152 or DenseNet-201 layers net because it is computationally expensive), use less layers nets instead (e.g. VGG-16, ResNet-50 or DenseNet-121 layers). Pick one pre-trained model that you think it gives the best performance with your hyper-parameters (say ResNet-50 layers). After you obtained the optimal hyper parameters, just select the same but more layers net (say ResNet-101 or ResNet-152 layers) to increase the accuracy.\n\nAlways use normalization layers in your network.  InstanceNormalization gives slightly performance improvements if they use a small batch-size. Or you may also try GroupNormalization.\n\nUse SpatialDropout after a features concatenation if you have two or more convolution layers (say Li) operate on the same input (say F). Always shuffle your training data, both before training and during training, in case you don’t take benefit from temporal data. This may help improving your network performance. Choose a right optimizer. There are many popular adaptive optimizers such as Adam, Adagrad, Adadelta, or RMSprop etc. SGD+momentum is widely used in various problem domains. \nRead More here: [Custom pooling model](https://towardsdatascience.com/a-bunch-of-tips-and-tricks-for-training-deep-neural-networks-3ca24c31ddc8)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335113,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "06/04/2021 02:59:00",
      "content": "<p>In my experiments, the SED model was sensitive to learning rate and batch size. <br>\nAlso, some architectures (rexnet150, skresnet-50, etc.) produced only nobird (probably due to lack of tuning) in train_soundscape despite good CV on train_short_audio.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335915,
          "author_name": "yutoshibata",
          "author_url": "",
          "post_date": "06/04/2021 14:05:48",
          "content": "<p><a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> Thank you for kind reply 🙌<br>\nI learned a lot from your 4th place solution as well.<br>\nI finally got it and will try to re-train my model after changing some hyper parameters.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336009,
      "author_name": "saimasharleen",
      "author_url": "",
      "post_date": "06/04/2021 15:27:01",
      "content": "<p>Before you start building your network architecture, first thing you need to do is to verify your input data into the network if an input (x) corresponds to a label (y). In case of dense prediction, make sure that the ground-truth label (y) is properly encoded to label indexes (or one-hot encoding). If not, the training won’t work.</p>\n<p>Then, if your dataset in your problem domain is similar to ImageNet dataset, use a pre-trained model on this dataset. The most widely used pre-trained models are VGG net, ResNet, DenseNet or Xception etc. There are many layer architectures, for instance, VGG (19 and 16 layers), ResNet (152, 101, 50 layers or less), DenseNet (201, 169 and 121 layers). Note: Do not try searching hyper-parameters by using more layers nets (e.g. VGG-19, ResNet-152 or DenseNet-201 layers net because it is computationally expensive), use less layers nets instead (e.g. VGG-16, ResNet-50 or DenseNet-121 layers). Pick one pre-trained model that you think it gives the best performance with your hyper-parameters (say ResNet-50 layers). After you obtained the optimal hyper parameters, just select the same but more layers net (say ResNet-101 or ResNet-152 layers) to increase the accuracy.</p>\n<p>Always use normalization layers in your network.  InstanceNormalization gives slightly performance improvements if they use a small batch-size. Or you may also try GroupNormalization.</p>\n<p>Use SpatialDropout after a features concatenation if you have two or more convolution layers (say Li) operate on the same input (say F). Always shuffle your training data, both before training and during training, in case you don’t take benefit from temporal data. This may help improving your network performance. Choose a right optimizer. There are many popular adaptive optimizers such as Adam, Adagrad, Adadelta, or RMSprop etc. SGD+momentum is widely used in various problem domains. <br>\nRead More here: <a href=\"https://towardsdatascience.com/a-bunch-of-tips-and-tricks-for-training-deep-neural-networks-3ca24c31ddc8\" target=\"_blank\">Custom pooling model</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335082": "I am sorry for asking very basic question, while others are publishing great solutions now.\nDuring this competition, I sometimes tried to train SED or custom pooling (eg. freq axis only) models for about 30 epochs on ColabPro. However, all of them ended up just all zero f1 score. Did anyone have similar result ?\n\nBecause of this, I used normal CNN models. By using some tricks (rating-based weight loss, background noise from train_soundscape, fine-tuning from non-background noise models) I got bronze model, but I do want to know the key to train SED model for my future work🙏",
    "1335113": "In my experiments, the SED model was sensitive to learning rate and batch size. \nAlso, some architectures (rexnet150, skresnet-50, etc.) produced only nobird (probably due to lack of tuning) in train_soundscape despite good CV on train_short_audio.",
    "1335915": "tattaka Thank you for kind reply 🙌\nI learned a lot from your 4th place solution as well.\nI finally got it and will try to re-train my model after changing some hyper parameters.",
    "1336009": "Before you start building your network architecture, first thing you need to do is to verify your input data into the network if an input (x) corresponds to a label (y). In case of dense prediction, make sure that the ground-truth label (y) is properly encoded to label indexes (or one-hot encoding). If not, the training won’t work.\n\nThen, if your dataset in your problem domain is similar to ImageNet dataset, use a pre-trained model on this dataset. The most widely used pre-trained models are VGG net, ResNet, DenseNet or Xception etc. There are many layer architectures, for instance, VGG (19 and 16 layers), ResNet (152, 101, 50 layers or less), DenseNet (201, 169 and 121 layers). Note: Do not try searching hyper-parameters by using more layers nets (e.g. VGG-19, ResNet-152 or DenseNet-201 layers net because it is computationally expensive), use less layers nets instead (e.g. VGG-16, ResNet-50 or DenseNet-121 layers). Pick one pre-trained model that you think it gives the best performance with your hyper-parameters (say ResNet-50 layers). After you obtained the optimal hyper parameters, just select the same but more layers net (say ResNet-101 or ResNet-152 layers) to increase the accuracy.\n\nAlways use normalization layers in your network.  InstanceNormalization gives slightly performance improvements if they use a small batch-size. Or you may also try GroupNormalization.\n\nUse SpatialDropout after a features concatenation if you have two or more convolution layers (say Li) operate on the same input (say F). Always shuffle your training data, both before training and during training, in case you don’t take benefit from temporal data. This may help improving your network performance. Choose a right optimizer. There are many popular adaptive optimizers such as Adam, Adagrad, Adadelta, or RMSprop etc. SGD+momentum is widely used in various problem domains. \nRead More here: [Custom pooling model](https://towardsdatascience.com/a-bunch-of-tips-and-tricks-for-training-deep-neural-networks-3ca24c31ddc8)"
  },
  "source": "meta"
}