{
  "id": 209041,
  "title": "Is a larger model unnecessary?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/209041",
  "author_name": "shinmura0",
  "post_date": "2021-01-06T03:07:49.849000",
  "votes": 40,
  "comment_count": 16,
  "views": 0,
  "content": "<p>In this competition, a smaller model(EfficientNet and MobileNetV2) has good performance.<br>\nBut, is a larger model(like ResNest50) unnecessary? The answer is \"NO\".</p>\n<p>In ensemble phase, a larger model is necessary.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F1720b48aade45ca0e78208fdbfac0383%2Ffig1.png?generation=1609899400061046&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>ResNest50 is not good performance(LB:0.867)</li>\n<li>But \"EfficientNetB0-2(<strong>LB:0.901</strong>) + ResNest50(<strong>LB:0.867</strong>) = <strong>LB:0.903</strong>\"</li>\n<li>ResNest50 is important role</li>\n</ul>\n<p>In this competition, the ensemble is storng. And <strong>mixing a larger model is effective</strong>.<br>\nThe ensemble code is below.</p>\n<pre><code>def ensemble(DataFrame1, DataFrame2, ..., DataFrame5):\n    for i in range(24):\n        name = \"s\" + str(i)\n        DataFrame1[name] = 0.2*(DataFrame1[name]+ ... + DataFrame5[name])\n    return DataFrame1\n</code></pre>\n<p>This is simple average.</p>",
  "messages": [
    {
      "id": 1140473,
      "postDate": "2021-01-06T03:07:49.850Z",
      "content": "<p>In this competition, a smaller model(EfficientNet and MobileNetV2) has good performance.<br>\nBut, is a larger model(like ResNest50) unnecessary? The answer is \"NO\".</p>\n<p>In ensemble phase, a larger model is necessary.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F1720b48aade45ca0e78208fdbfac0383%2Ffig1.png?generation=1609899400061046&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>ResNest50 is not good performance(LB:0.867)</li>\n<li>But \"EfficientNetB0-2(<strong>LB:0.901</strong>) + ResNest50(<strong>LB:0.867</strong>) = <strong>LB:0.903</strong>\"</li>\n<li>ResNest50 is important role</li>\n</ul>\n<p>In this competition, the ensemble is storng. And <strong>mixing a larger model is effective</strong>.<br>\nThe ensemble code is below.</p>\n<pre><code>def ensemble(DataFrame1, DataFrame2, ..., DataFrame5):\n    for i in range(24):\n        name = \"s\" + str(i)\n        DataFrame1[name] = 0.2*(DataFrame1[name]+ ... + DataFrame5[name])\n    return DataFrame1\n</code></pre>\n<p>This is simple average.</p>",
      "rawMarkdown": "In this competition, a smaller model(EfficientNet and MobileNetV2) has good performance.\nBut, is a larger model(like ResNest50) unnecessary? The answer is \"NO\".\n\nIn ensemble phase, a larger model is necessary.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F1720b48aade45ca0e78208fdbfac0383%2Ffig1.png?generation=1609899400061046&alt=media)\n\n+ ResNest50 is not good performance(LB:0.867)\n+ But \"EfficientNetB0-2(**LB:0.901**) + ResNest50(**LB:0.867**) = **LB:0.903**\"\n+ ResNest50 is important role\n\nIn this competition, the ensemble is storng. And **mixing a larger model is effective**.\nThe ensemble code is below.\n\n```\ndef ensemble(DataFrame1, DataFrame2, ..., DataFrame5):\n    for i in range(24):\n        name = \"s\" + str(i)\n        DataFrame1[name] = 0.2*(DataFrame1[name]+ ... + DataFrame5[name])\n    return DataFrame1\n```\n\nThis is simple average.",
      "votes": 40
    },
    {
      "id": 1140942,
      "postDate": "2021-01-06T11:45:23.557Z",
      "content": "<p>I think your ensemble benefits more from diversity (ResNests and EfficientNets are quite different) than model size.</p>\n<p>If you want to compare model size, I would advise checking the performance of EfficientNet-b3 or even b4.</p>",
      "rawMarkdown": "I think your ensemble benefits more from diversity (ResNests and EfficientNets are quite different) than model size.\n\nIf you want to compare model size, I would advise checking the performance of EfficientNet-b3 or even b4.",
      "votes": 7,
      "replies": [
        {
          "id": 1142241,
          "postDate": "2021-01-07T08:48:15.307Z",
          "content": "<p>It seems that the effect of diversity is greater than the size of the model.</p>",
          "rawMarkdown": "It seems that the effect of diversity is greater than the size of the model.",
          "votes": 2
        },
        {
          "id": 1175133,
          "postDate": "2021-01-29T00:05:03.457Z",
          "content": "<p>Maybe I should start looking beyond effnets…</p>",
          "rawMarkdown": "Maybe I should start looking beyond effnets..."
        }
      ]
    },
    {
      "id": 1174090,
      "postDate": "2021-01-28T09:02:16.883Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> , may I ask how did you train your 5 different models? I guess you use StratifiedKFold as 5 folds, and each of your model are trained with all 5 folds, then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?</p>",
      "rawMarkdown": "Thanks for sharing @shinmurashinmura , may I ask how did you train your 5 different models? I guess you use StratifiedKFold as 5 folds, and each of your model are trained with all 5 folds, then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?",
      "replies": [
        {
          "id": 1180362,
          "postDate": "2021-02-01T08:17:59.517Z",
          "content": "<blockquote>\n  <p>then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?</p>\n</blockquote>\n<p>In this case, I trained 25(=5*5) models. And I average each prediction.</p>",
          "rawMarkdown": "> then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?\n\nIn this case, I trained 25(=5*5) models. And I average each prediction."
        }
      ]
    },
    {
      "id": 1143918,
      "postDate": "2021-01-08T06:12:10.113Z",
      "content": "<p>How did you train EfficinetNet-B0 it takes too much RAM</p>",
      "rawMarkdown": "How did you train EfficinetNet-B0 it takes too much RAM",
      "replies": [
        {
          "id": 1144084,
          "postDate": "2021-01-08T08:17:43.197Z",
          "content": "<p>I used batch_size=16. It is a small size.</p>",
          "rawMarkdown": "I used batch_size=16. It is a small size.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1143668,
      "postDate": "2021-01-08T01:56:48.417Z",
      "content": "<p>Hi!<br>\nMay I know how you generated the graph?</p>",
      "rawMarkdown": "Hi!\nMay I know how you generated the graph?",
      "replies": [
        {
          "id": 1143722,
          "postDate": "2021-01-08T03:04:01.187Z",
          "content": "<p>I used \"Google slide\".</p>",
          "rawMarkdown": "I used \"Google slide\".",
          "votes": 1
        }
      ]
    },
    {
      "id": 1141277,
      "postDate": "2021-01-06T15:47:17.507Z",
      "content": "<p>Can you briefly explain the SED part architecture wise? Does it have any specific properties? I still don't understand it properly :( I currently use ResNests backbone with the linear layers and a sigmoid at the end and then use a BCELoss.  Comparing the differences or pointing out to any resource would be too much appreciated. Thanks.</p>",
      "rawMarkdown": "Can you briefly explain the SED part architecture wise? Does it have any specific properties? I still don't understand it properly :( I currently use ResNests backbone with the linear layers and a sigmoid at the end and then use a BCELoss.  Comparing the differences or pointing out to any resource would be too much appreciated. Thanks.",
      "replies": [
        {
          "id": 1142227,
          "postDate": "2021-01-07T08:38:14.967Z",
          "content": "<p>Now, I'm writing new topic \"How to use SED\". Please wait.</p>",
          "rawMarkdown": "Now, I'm writing new topic \"How to use SED\". Please wait.",
          "votes": 5
        },
        {
          "id": 1142810,
          "postDate": "2021-01-07T15:47:18.590Z",
          "content": "<p>Great, I'm waiting for it 🤘</p>",
          "rawMarkdown": "Great, I'm waiting for it 🤘"
        }
      ]
    },
    {
      "id": 1140653,
      "postDate": "2021-01-06T06:59:16.013Z",
      "content": "<p>Thank you i was struggling with  this alot</p>",
      "rawMarkdown": "Thank you i was struggling with  this alot"
    },
    {
      "id": 1144627,
      "postDate": "2021-01-08T15:16:08.727Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1140647,
      "postDate": "2021-01-06T06:54:13.577Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1140646,
      "postDate": "2021-01-06T06:51:29.123Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1140942,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2021-01-06T11:45:23.557000",
      "content": "<p>I think your ensemble benefits more from diversity (ResNests and EfficientNets are quite different) than model size.</p>\n<p>If you want to compare model size, I would advise checking the performance of EfficientNet-b3 or even b4.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1142241,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-07T08:48:15.307000",
          "content": "<p>It seems that the effect of diversity is greater than the size of the model.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1175133,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-01-29T00:05:03.457000",
          "content": "<p>Maybe I should start looking beyond effnets…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1174090,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-01-28T09:02:16.883000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> , may I ask how did you train your 5 different models? I guess you use StratifiedKFold as 5 folds, and each of your model are trained with all 5 folds, then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1180362,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-02-01T08:17:59.517000",
          "content": "<blockquote>\n  <p>then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?</p>\n</blockquote>\n<p>In this case, I trained 25(=5*5) models. And I average each prediction.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1143918,
      "author_name": "Sayantan Mazumdar",
      "author_url": "",
      "post_date": "2021-01-08T06:12:10.113000",
      "content": "<p>How did you train EfficinetNet-B0 it takes too much RAM</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1144084,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-08T08:17:43.197000",
          "content": "<p>I used batch_size=16. It is a small size.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1143668,
      "author_name": "GVKannan",
      "author_url": "",
      "post_date": "2021-01-08T01:56:48.417000",
      "content": "<p>Hi!<br>\nMay I know how you generated the graph?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1143722,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-08T03:04:01.187000",
          "content": "<p>I used \"Google slide\".</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1141277,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2021-01-06T15:47:17.507000",
      "content": "<p>Can you briefly explain the SED part architecture wise? Does it have any specific properties? I still don't understand it properly :( I currently use ResNests backbone with the linear layers and a sigmoid at the end and then use a BCELoss.  Comparing the differences or pointing out to any resource would be too much appreciated. Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1142227,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-07T08:38:14.967000",
          "content": "<p>Now, I'm writing new topic \"How to use SED\". Please wait.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1142810,
          "author_name": "Sinan Calisir",
          "author_url": "",
          "post_date": "2021-01-07T15:47:18.590000",
          "content": "<p>Great, I'm waiting for it 🤘</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1140653,
      "author_name": "Sayantan Mazumdar",
      "author_url": "",
      "post_date": "2021-01-06T06:59:16.013000",
      "content": "<p>Thank you i was struggling with  this alot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1144627,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-08T15:16:08.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1140647,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-06T06:54:13.577000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1140646,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-06T06:51:29.123000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1140473": "In this competition, a smaller model(EfficientNet and MobileNetV2) has good performance.\nBut, is a larger model(like ResNest50) unnecessary? The answer is \"NO\".\n\nIn ensemble phase, a larger model is necessary.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4704212%2F1720b48aade45ca0e78208fdbfac0383%2Ffig1.png?generation=1609899400061046&alt=media)\n\n+ ResNest50 is not good performance(LB:0.867)\n+ But \"EfficientNetB0-2(**LB:0.901**) + ResNest50(**LB:0.867**) = **LB:0.903**\"\n+ ResNest50 is important role\n\nIn this competition, the ensemble is storng. And **mixing a larger model is effective**.\nThe ensemble code is below.\n\n```\ndef ensemble(DataFrame1, DataFrame2, ..., DataFrame5):\n    for i in range(24):\n        name = \"s\" + str(i)\n        DataFrame1[name] = 0.2*(DataFrame1[name]+ ... + DataFrame5[name])\n    return DataFrame1\n```\n\nThis is simple average.",
    "1140942": "I think your ensemble benefits more from diversity (ResNests and EfficientNets are quite different) than model size.\n\nIf you want to compare model size, I would advise checking the performance of EfficientNet-b3 or even b4.",
    "1174090": "Thanks for sharing @shinmurashinmura , may I ask how did you train your 5 different models? I guess you use StratifiedKFold as 5 folds, and each of your model are trained with all 5 folds, then ensemble them (total models are 5*5=25), or you only train each model on only 1 fold(total models are 5)?",
    "1143918": "How did you train EfficinetNet-B0 it takes too much RAM",
    "1143668": "Hi!\nMay I know how you generated the graph?",
    "1141277": "Can you briefly explain the SED part architecture wise? Does it have any specific properties? I still don't understand it properly :( I currently use ResNests backbone with the linear layers and a sigmoid at the end and then use a BCELoss.  Comparing the differences or pointing out to any resource would be too much appreciated. Thanks.",
    "1140653": "Thank you i was struggling with  this alot",
    "1144627": "",
    "1140647": "",
    "1140646": ""
  }
}