{
  "id": 176752,
  "title": "Different mix of schedulers and settings",
  "url": "/competitions/birdsong-recognition/discussion/176752",
  "author_name": "",
  "post_date": "2020-08-23T09:27:28.587534100Z",
  "votes": 12,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Have tried some schedulers and settings, all with various epochs, if I see a pattern a stop and run next, maybe not the best but time efficient. Below are some of the results, includes only those that have been shown to work reasonably well.<br>\nSo far StepLR step_size=10 alpha=0.5 or StepLR step_size=5 alpha=0.6 seems to work well,  but will test more combinations.<br>\nI use the new released optimizer AdamP with weight_decay=5e-3.<br>\n<a href=\"https://clovaai.github.io/AdamP/\" target=\"_blank\">https://clovaai.github.io/AdamP/</a></p>\n<p>Model: Efficientnet_b3_ns with dropout 0.5 B64 dim 224x224 and also interpolate 224x224 in the model.</p>\n<p>CosineAnnealingWarmRestarts T_0=20<br>\nLookahead k=10, alpha=0.5<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fb1e80b2c9002fd37c0176d5b7cd72c73%2FLookahead10_AnCosW10.png?generation=1598174400338373&amp;alt=media\" alt=\"\"></p>\n<p>CosineAnnealingWarmRestarts T_0=10<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5f65740f9d1c286263024abce8f31f4%2FNesterov.Ann10ejLookAh.png?generation=1598172974712618&amp;alt=media\" alt=\"\"></p>\n<p>StepLR step_size=20, gamma=0.5<br>\nNesterov<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3d320566bb5fc8e7c3e8db3b68cd8069%2FNesterovStepLR20a5.png?generation=1598172885356940&amp;alt=media\" alt=\"\"></p>\n<p>StepLR step_size=5, gamma=0.6<br>\nNesterov<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F008808f1717804d808aaea66be7ac13d%2FNesterovStepLR-5-a6.png?generation=1598173076202154&amp;alt=media\" alt=\"\"></p>\n<p>StepLR step_size=10, gamma=0.5<br>\nNesterov<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F272ec41021c88ece68520c3f3f3408fd%2FStep10a5Nest.png?generation=1598174591255062&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "982340",
      "postDate": "08/23/2020 09:27:28",
      "content": "<p>Have tried some schedulers and settings, all with various epochs, if I see a pattern a stop and run next, maybe not the best but time efficient. Below are some of the results, includes only those that have been shown to work reasonably well.<br>\nSo far StepLR step_size=10 alpha=0.5 or StepLR step_size=5 alpha=0.6 seems to work well,  but will test more combinations.<br>\nI use the new released optimizer AdamP with weight_decay=5e-3.<br>\n<a href=\"https://clovaai.github.io/AdamP/\" target=\"_blank\">https://clovaai.github.io/AdamP/</a></p>\n<p>Model: Efficientnet_b3_ns with dropout 0.5 B64 dim 224x224 and also interpolate 224x224 in the model.</p>\n<p>CosineAnnealingWarmRestarts T_0=20<br>\nLookahead k=10, alpha=0.5<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fb1e80b2c9002fd37c0176d5b7cd72c73%2FLookahead10_AnCosW10.png?generation=1598174400338373&amp;alt=media\" alt=\"\"></p>\n<p>CosineAnnealingWarmRestarts T_0=10<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5f65740f9d1c286263024abce8f31f4%2FNesterov.Ann10ejLookAh.png?generation=1598172974712618&amp;alt=media\" alt=\"\"></p>\n<p>StepLR step_size=20, gamma=0.5<br>\nNesterov<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3d320566bb5fc8e7c3e8db3b68cd8069%2FNesterovStepLR20a5.png?generation=1598172885356940&amp;alt=media\" alt=\"\"></p>\n<p>StepLR step_size=5, gamma=0.6<br>\nNesterov<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F008808f1717804d808aaea66be7ac13d%2FNesterovStepLR-5-a6.png?generation=1598173076202154&amp;alt=media\" alt=\"\"></p>\n<p>StepLR step_size=10, gamma=0.5<br>\nNesterov<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F272ec41021c88ece68520c3f3f3408fd%2FStep10a5Nest.png?generation=1598174591255062&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Have tried some schedulers and settings, all with various epochs, if I see a pattern a stop and run next, maybe not the best but time efficient. Below are some of the results, includes only those that have been shown to work reasonably well.\nSo far StepLR step_size=10 alpha=0.5 or StepLR step_size=5 alpha=0.6 seems to work well,  but will test more combinations.\nI use the new released optimizer AdamP with weight_decay=5e-3.\nhttps://clovaai.github.io/AdamP/\n\nModel: Efficientnet_b3_ns with dropout 0.5 B64 dim 224x224 and also interpolate 224x224 in the model.\n\nCosineAnnealingWarmRestarts T_0=20\nLookahead k=10, alpha=0.5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fb1e80b2c9002fd37c0176d5b7cd72c73%2FLookahead10_AnCosW10.png?generation=1598174400338373&alt=media)\n\nCosineAnnealingWarmRestarts T_0=10\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5f65740f9d1c286263024abce8f31f4%2FNesterov.Ann10ejLookAh.png?generation=1598172974712618&alt=media)\n\nStepLR step_size=20, gamma=0.5\nNesterov\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3d320566bb5fc8e7c3e8db3b68cd8069%2FNesterovStepLR20a5.png?generation=1598172885356940&alt=media)\n\nStepLR step_size=5, gamma=0.6\nNesterov\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F008808f1717804d808aaea66be7ac13d%2FNesterovStepLR-5-a6.png?generation=1598173076202154&alt=media)\n\nStepLR step_size=10, gamma=0.5\nNesterov\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F272ec41021c88ece68520c3f3f3408fd%2FStep10a5Nest.png?generation=1598174591255062&alt=media)",
      "votes": null
    },
    {
      "id": "982961",
      "postDate": "08/23/2020 21:46:22",
      "content": "<p>Thank you, it was informative 👍</p>",
      "rawMarkdown": "Thank you, it was informative 👍",
      "votes": null
    },
    {
      "id": "984997",
      "postDate": "08/25/2020 12:32:00",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a>, what batch_size  do you use?</p>",
      "rawMarkdown": "Hey @kirderf, what batch_size  do you use?",
      "votes": null
    },
    {
      "id": "985001",
      "postDate": "08/25/2020 12:36:02",
      "content": "<p>64, I wrote B64, maybe a little cryptic</p>",
      "rawMarkdown": "64, I wrote B64, maybe a little cryptic",
      "votes": null
    },
    {
      "id": "985307",
      "postDate": "08/25/2020 16:20:53",
      "content": "<p>Thanks , I also tried same thing enet b2 got almost same losses. But still not able to beat baseline model. Were you able to beat it using efficient net ?. You can see me training in <a href=\"https://www.kaggle.com/rsinda/training-efficientnet-model\" target=\"_blank\">https://www.kaggle.com/rsinda/training-efficientnet-model</a>  . Am I doing it right ?</p>",
      "rawMarkdown": "Thanks , I also tried same thing enet b2 got almost same losses. But still not able to beat baseline model. Were you able to beat it using efficient net ?. You can see me training in https://www.kaggle.com/rsinda/training-efficientnet-model  . Am I doing it right ?",
      "votes": null
    },
    {
      "id": "985395",
      "postDate": "08/25/2020 17:43:57",
      "content": "<p>Yes this is not easy, still testing, can't say I found a good starting point yet. I will try different sound and image quailty/dimensions, see what happens, also other models. For training one can try to use various of Image Augmentations.<br>\nThe time to load training data and the limit of 2h for submit is a challenge. </p>\n<p>Use below to get a better progress view and turn off the valid progress(progress_bar=False), to much validation info in a short time.</p>\n<pre><code>    while not manager.stop_trigger:\n        model.train()\n        progress_bar = tqdm_notebook(train_loader)\n        for batch_idx, (data, target) in enumerate(progress_bar):\n            with manager.run_iteration():\n                data, target = data.to(device), target.to(device)\n                optimizer.zero_grad()\n                output = model(data)\n                loss = loss_func(output, target)\n                progress_bar.set_description(f'train/loss: {loss.item():.6f}')\n                ppe.reporting.report({'train/loss': loss.item()})\n                loss.backward()\n                optimizer.step()\n                scheduler.step()\n</code></pre>",
      "rawMarkdown": "Yes this is not easy, still testing, can't say I found a good starting point yet. I will try different sound and image quailty/dimensions, see what happens, also other models. For training one can try to use various of Image Augmentations.\nThe time to load training data and the limit of 2h for submit is a challenge. \n\nUse below to get a better progress view and turn off the valid progress(progress_bar=False), to much validation info in a short time.\n\n```\n    while not manager.stop_trigger:\n        model.train()\n        progress_bar = tqdm_notebook(train_loader)\n        for batch_idx, (data, target) in enumerate(progress_bar):\n            with manager.run_iteration():\n                data, target = data.to(device), target.to(device)\n                optimizer.zero_grad()\n                output = model(data)\n                loss = loss_func(output, target)\n                progress_bar.set_description(f'train/loss: {loss.item():.6f}')\n                ppe.reporting.report({'train/loss': loss.item()})\n                loss.backward()\n                optimizer.step()\n                scheduler.step()\n\n```",
      "votes": null
    },
    {
      "id": "991679",
      "postDate": "08/30/2020 15:25:32",
      "content": "<p>A better option:</p>\n<p>AdamW<br>\nCosineAnnealingWarmRestarts T_0: 1 T_mult: 2 eta_min: 0.0005<br>\nImage + spectrogram augmentations <br>\n(224x547)<br>\nOnly tried with the smallest EfffnetB0ns so far.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F2cd5c58126e489771c6d52de01a9c96d%2Floss%20(20).png?generation=1598800446474628&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F239729f81bc33f01555aeb6dee855867%2Flr%20(19).png?generation=1598800460018029&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "A better option:\n\nAdamW\nCosineAnnealingWarmRestarts T_0: 1 T_mult: 2 eta_min: 0.0005\nImage + spectrogram augmentations \n(224x547)\nOnly tried with the smallest EfffnetB0ns so far.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F2cd5c58126e489771c6d52de01a9c96d%2Floss%20(20).png?generation=1598800446474628&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F239729f81bc33f01555aeb6dee855867%2Flr%20(19).png?generation=1598800460018029&alt=media)",
      "votes": null
    },
    {
      "id": "991994",
      "postDate": "08/30/2020 19:33:55",
      "content": "<p>I was surprised to see you use B3 on small image size.  Thanks for sharing your experiments.</p>",
      "rawMarkdown": "I was surprised to see you use B3 on small image size.  Thanks for sharing your experiments.",
      "votes": null
    },
    {
      "id": "992078",
      "postDate": "08/30/2020 22:18:19",
      "content": "<p>used B3 as a mid range model for the testing/benchmark, but it gave a boost to the ensemble.<br>\nUsally keep an eye on below</p>\n<pre><code>  # (width_coefficient, depth_coefficient, resolution, dropout_rate)\n  'efficientnet-b0': (1.0, 1.0, 224, 0.2),\n  'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n  'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n  'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n  'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n  'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n  'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n  'efficientnet-b7': (2.0, 3.1, 600, 0.5),\n  'efficientnet-b8': (2.2, 3.6, 672, 0.5),\n  'efficientnet-l2': (4.3, 5.3, 800, 0.5),\n</code></pre>\n<p>There is a lot left to discover and test in this competition. Will see if one has time to reach an OK result before the deadline.</p>",
      "rawMarkdown": "used B3 as a mid range model for the testing/benchmark, but it gave a boost to the ensemble.\nUsally keep an eye on below\n\n      # (width_coefficient, depth_coefficient, resolution, dropout_rate)\n      'efficientnet-b0': (1.0, 1.0, 224, 0.2),\n      'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n      'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n      'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n      'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n      'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n      'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n      'efficientnet-b7': (2.0, 3.1, 600, 0.5),\n      'efficientnet-b8': (2.2, 3.6, 672, 0.5),\n      'efficientnet-l2': (4.3, 5.3, 800, 0.5),\n\nThere is a lot left to discover and test in this competition. Will see if one has time to reach an OK result before the deadline.",
      "votes": null
    },
    {
      "id": "1003783",
      "postDate": "09/09/2020 08:42:22",
      "content": "<p>Better late than never, I will now use this metric instead, seems like a good work environment.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Ffffd3956672398ed985e77fdd231349b%2Floss%20(26).png?generation=1599640672197611&amp;alt=media\" alt=\"\"></p>\n<p>f1_score with sample - competition metric<br>\nrow_wise_f1_score_micro_numpy- same as above, but will stay for a while just to be safe (?)<br>\nf1_score with micro - just to follow the correlation.</p>\n<p>And I'll run a global threshold optimization in every validation step, hopefully also a local optimization per label if I succeed with the coding.</p>",
      "rawMarkdown": "Better late than never, I will now use this metric instead, seems like a good work environment.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Ffffd3956672398ed985e77fdd231349b%2Floss%20(26).png?generation=1599640672197611&alt=media)\n\nf1_score with sample - competition metric\nrow_wise_f1_score_micro_numpy- same as above, but will stay for a while just to be safe (?)\nf1_score with micro - just to follow the correlation.\n\nAnd I'll run a global threshold optimization in every validation step, hopefully also a local optimization per label if I succeed with the coding.",
      "votes": null
    },
    {
      "id": "1005936",
      "postDate": "09/10/2020 22:27:50",
      "content": "<p>EffB3Imagenet<br>\nAdamW<br>\nCosineAnnealingWarmRestarts T_0: batch update instead of epoch and with cycle of 2 epochs and changed lr with 1e-1 when converge.<br>\nData - upsampling and picked only sound rating &gt;=3.<br>\nRandom batchcutout+mixup, dropout 0.3.</p>\n<p>This shows training after 30 epochs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F7fb342fc9375da89407669edc836c570%2Floss%20(31).png?generation=1599776625210733&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F8b7298063c7387fa49483007e36a1f93%2Flr%20(28).png?generation=1599776417747942&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "EffB3Imagenet\nAdamW\nCosineAnnealingWarmRestarts T_0: batch update instead of epoch and with cycle of 2 epochs and changed lr with 1e-1 when converge.\nData - upsampling and picked only sound rating >=3.\nRandom batchcutout+mixup, dropout 0.3.\n\nThis shows training after 30 epochs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F7fb342fc9375da89407669edc836c570%2Floss%20(31).png?generation=1599776625210733&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F8b7298063c7387fa49483007e36a1f93%2Flr%20(28).png?generation=1599776417747942&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 982961,
      "author_name": "",
      "author_url": "",
      "post_date": "08/23/2020 21:46:22",
      "content": "<p>Thank you, it was informative 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 984997,
      "author_name": "rsinda",
      "author_url": "",
      "post_date": "08/25/2020 12:32:00",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a>, what batch_size  do you use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 985001,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/25/2020 12:36:02",
          "content": "<p>64, I wrote B64, maybe a little cryptic</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 985307,
          "author_name": "rsinda",
          "author_url": "",
          "post_date": "08/25/2020 16:20:53",
          "content": "<p>Thanks , I also tried same thing enet b2 got almost same losses. But still not able to beat baseline model. Were you able to beat it using efficient net ?. You can see me training in <a href=\"https://www.kaggle.com/rsinda/training-efficientnet-model\" target=\"_blank\">https://www.kaggle.com/rsinda/training-efficientnet-model</a>  . Am I doing it right ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 985395,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/25/2020 17:43:57",
          "content": "<p>Yes this is not easy, still testing, can't say I found a good starting point yet. I will try different sound and image quailty/dimensions, see what happens, also other models. For training one can try to use various of Image Augmentations.<br>\nThe time to load training data and the limit of 2h for submit is a challenge. </p>\n<p>Use below to get a better progress view and turn off the valid progress(progress_bar=False), to much validation info in a short time.</p>\n<pre><code>    while not manager.stop_trigger:\n        model.train()\n        progress_bar = tqdm_notebook(train_loader)\n        for batch_idx, (data, target) in enumerate(progress_bar):\n            with manager.run_iteration():\n                data, target = data.to(device), target.to(device)\n                optimizer.zero_grad()\n                output = model(data)\n                loss = loss_func(output, target)\n                progress_bar.set_description(f'train/loss: {loss.item():.6f}')\n                ppe.reporting.report({'train/loss': loss.item()})\n                loss.backward()\n                optimizer.step()\n                scheduler.step()\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 991679,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "08/30/2020 15:25:32",
      "content": "<p>A better option:</p>\n<p>AdamW<br>\nCosineAnnealingWarmRestarts T_0: 1 T_mult: 2 eta_min: 0.0005<br>\nImage + spectrogram augmentations <br>\n(224x547)<br>\nOnly tried with the smallest EfffnetB0ns so far.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F2cd5c58126e489771c6d52de01a9c96d%2Floss%20(20).png?generation=1598800446474628&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F239729f81bc33f01555aeb6dee855867%2Flr%20(19).png?generation=1598800460018029&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 991994,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/30/2020 19:33:55",
          "content": "<p>I was surprised to see you use B3 on small image size.  Thanks for sharing your experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 992078,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/30/2020 22:18:19",
          "content": "<p>used B3 as a mid range model for the testing/benchmark, but it gave a boost to the ensemble.<br>\nUsally keep an eye on below</p>\n<pre><code>  # (width_coefficient, depth_coefficient, resolution, dropout_rate)\n  'efficientnet-b0': (1.0, 1.0, 224, 0.2),\n  'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n  'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n  'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n  'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n  'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n  'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n  'efficientnet-b7': (2.0, 3.1, 600, 0.5),\n  'efficientnet-b8': (2.2, 3.6, 672, 0.5),\n  'efficientnet-l2': (4.3, 5.3, 800, 0.5),\n</code></pre>\n<p>There is a lot left to discover and test in this competition. Will see if one has time to reach an OK result before the deadline.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1003783,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "09/09/2020 08:42:22",
      "content": "<p>Better late than never, I will now use this metric instead, seems like a good work environment.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Ffffd3956672398ed985e77fdd231349b%2Floss%20(26).png?generation=1599640672197611&amp;alt=media\" alt=\"\"></p>\n<p>f1_score with sample - competition metric<br>\nrow_wise_f1_score_micro_numpy- same as above, but will stay for a while just to be safe (?)<br>\nf1_score with micro - just to follow the correlation.</p>\n<p>And I'll run a global threshold optimization in every validation step, hopefully also a local optimization per label if I succeed with the coding.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1005936,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "09/10/2020 22:27:50",
      "content": "<p>EffB3Imagenet<br>\nAdamW<br>\nCosineAnnealingWarmRestarts T_0: batch update instead of epoch and with cycle of 2 epochs and changed lr with 1e-1 when converge.<br>\nData - upsampling and picked only sound rating &gt;=3.<br>\nRandom batchcutout+mixup, dropout 0.3.</p>\n<p>This shows training after 30 epochs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F7fb342fc9375da89407669edc836c570%2Floss%20(31).png?generation=1599776625210733&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F8b7298063c7387fa49483007e36a1f93%2Flr%20(28).png?generation=1599776417747942&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "982340": "Have tried some schedulers and settings, all with various epochs, if I see a pattern a stop and run next, maybe not the best but time efficient. Below are some of the results, includes only those that have been shown to work reasonably well.\nSo far StepLR step_size=10 alpha=0.5 or StepLR step_size=5 alpha=0.6 seems to work well,  but will test more combinations.\nI use the new released optimizer AdamP with weight_decay=5e-3.\nhttps://clovaai.github.io/AdamP/\n\nModel: Efficientnet_b3_ns with dropout 0.5 B64 dim 224x224 and also interpolate 224x224 in the model.\n\nCosineAnnealingWarmRestarts T_0=20\nLookahead k=10, alpha=0.5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fb1e80b2c9002fd37c0176d5b7cd72c73%2FLookahead10_AnCosW10.png?generation=1598174400338373&alt=media)\n\nCosineAnnealingWarmRestarts T_0=10\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Fa5f65740f9d1c286263024abce8f31f4%2FNesterov.Ann10ejLookAh.png?generation=1598172974712618&alt=media)\n\nStepLR step_size=20, gamma=0.5\nNesterov\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F3d320566bb5fc8e7c3e8db3b68cd8069%2FNesterovStepLR20a5.png?generation=1598172885356940&alt=media)\n\nStepLR step_size=5, gamma=0.6\nNesterov\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F008808f1717804d808aaea66be7ac13d%2FNesterovStepLR-5-a6.png?generation=1598173076202154&alt=media)\n\nStepLR step_size=10, gamma=0.5\nNesterov\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F272ec41021c88ece68520c3f3f3408fd%2FStep10a5Nest.png?generation=1598174591255062&alt=media)",
    "982961": "Thank you, it was informative 👍",
    "984997": "Hey @kirderf, what batch_size  do you use?",
    "985001": "64, I wrote B64, maybe a little cryptic",
    "985307": "Thanks , I also tried same thing enet b2 got almost same losses. But still not able to beat baseline model. Were you able to beat it using efficient net ?. You can see me training in https://www.kaggle.com/rsinda/training-efficientnet-model  . Am I doing it right ?",
    "985395": "Yes this is not easy, still testing, can't say I found a good starting point yet. I will try different sound and image quailty/dimensions, see what happens, also other models. For training one can try to use various of Image Augmentations.\nThe time to load training data and the limit of 2h for submit is a challenge. \n\nUse below to get a better progress view and turn off the valid progress(progress_bar=False), to much validation info in a short time.\n\n```\n    while not manager.stop_trigger:\n        model.train()\n        progress_bar = tqdm_notebook(train_loader)\n        for batch_idx, (data, target) in enumerate(progress_bar):\n            with manager.run_iteration():\n                data, target = data.to(device), target.to(device)\n                optimizer.zero_grad()\n                output = model(data)\n                loss = loss_func(output, target)\n                progress_bar.set_description(f'train/loss: {loss.item():.6f}')\n                ppe.reporting.report({'train/loss': loss.item()})\n                loss.backward()\n                optimizer.step()\n                scheduler.step()\n\n```",
    "991679": "A better option:\n\nAdamW\nCosineAnnealingWarmRestarts T_0: 1 T_mult: 2 eta_min: 0.0005\nImage + spectrogram augmentations \n(224x547)\nOnly tried with the smallest EfffnetB0ns so far.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F2cd5c58126e489771c6d52de01a9c96d%2Floss%20(20).png?generation=1598800446474628&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F239729f81bc33f01555aeb6dee855867%2Flr%20(19).png?generation=1598800460018029&alt=media)",
    "991994": "I was surprised to see you use B3 on small image size.  Thanks for sharing your experiments.",
    "992078": "used B3 as a mid range model for the testing/benchmark, but it gave a boost to the ensemble.\nUsally keep an eye on below\n\n      # (width_coefficient, depth_coefficient, resolution, dropout_rate)\n      'efficientnet-b0': (1.0, 1.0, 224, 0.2),\n      'efficientnet-b1': (1.0, 1.1, 240, 0.2),\n      'efficientnet-b2': (1.1, 1.2, 260, 0.3),\n      'efficientnet-b3': (1.2, 1.4, 300, 0.3),\n      'efficientnet-b4': (1.4, 1.8, 380, 0.4),\n      'efficientnet-b5': (1.6, 2.2, 456, 0.4),\n      'efficientnet-b6': (1.8, 2.6, 528, 0.5),\n      'efficientnet-b7': (2.0, 3.1, 600, 0.5),\n      'efficientnet-b8': (2.2, 3.6, 672, 0.5),\n      'efficientnet-l2': (4.3, 5.3, 800, 0.5),\n\nThere is a lot left to discover and test in this competition. Will see if one has time to reach an OK result before the deadline.",
    "1003783": "Better late than never, I will now use this metric instead, seems like a good work environment.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2Ffffd3956672398ed985e77fdd231349b%2Floss%20(26).png?generation=1599640672197611&alt=media)\n\nf1_score with sample - competition metric\nrow_wise_f1_score_micro_numpy- same as above, but will stay for a while just to be safe (?)\nf1_score with micro - just to follow the correlation.\n\nAnd I'll run a global threshold optimization in every validation step, hopefully also a local optimization per label if I succeed with the coding.",
    "1005936": "EffB3Imagenet\nAdamW\nCosineAnnealingWarmRestarts T_0: batch update instead of epoch and with cycle of 2 epochs and changed lr with 1e-1 when converge.\nData - upsampling and picked only sound rating >=3.\nRandom batchcutout+mixup, dropout 0.3.\n\nThis shows training after 30 epochs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F7fb342fc9375da89407669edc836c570%2Floss%20(31).png?generation=1599776625210733&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3924325%2F8b7298063c7387fa49483007e36a1f93%2Flr%20(28).png?generation=1599776417747942&alt=media)"
  },
  "source": "meta"
}