{
  "id": 400899,
  "title": "How many models can you ensemble?",
  "url": "/competitions/birdclef-2023/discussion/400899",
  "author_name": "Zhongkai Shangguan",
  "post_date": "2023-04-10T19:22:59.863000",
  "votes": 20,
  "comment_count": 20,
  "views": 0,
  "content": "<p>With the cpu time&lt;120min restriction, may I ask how many models can you ensemble to your final result?<br>\nFor me its 2 to 3 models depend on the model size. Any idea how to accelerate cpu inference time?</p>\n<hr>\n<p>Seems my limit is 3 models. Any idea how to improve?</p>",
  "messages": [
    {
      "id": 2217335,
      "postDate": "2023-04-10T19:22:59.863Z",
      "content": "<p>With the cpu time&lt;120min restriction, may I ask how many models can you ensemble to your final result?<br>\nFor me its 2 to 3 models depend on the model size. Any idea how to accelerate cpu inference time?</p>\n<hr>\n<p>Seems my limit is 3 models. Any idea how to improve?</p>",
      "rawMarkdown": "With the cpu time<120min restriction, may I ask how many models can you ensemble to your final result?\nFor me its 2 to 3 models depend on the model size. Any idea how to accelerate cpu inference time?\n\n----------------------------------------------------------------\nSeems my limit is 3 models. Any idea how to improve?",
      "votes": 20
    },
    {
      "id": 2218775,
      "postDate": "2023-04-12T02:49:27.430Z",
      "content": "<p>I'm considering quantization to speed up inference.<br>\n<a href=\"https://pytorch.org/docs/stable/quantization.html\" target=\"_blank\">Quantization — PyTorch 2.0 documentation</a></p>",
      "rawMarkdown": "I'm considering quantization to speed up inference.\n[Quantization — PyTorch 2.0 documentation](https://pytorch.org/docs/stable/quantization.html)",
      "votes": 1,
      "replies": [
        {
          "id": 2221868,
          "postDate": "2023-04-14T16:39:05.823Z",
          "content": "<p>Could also think about Pruning, Weight clustering, …</p>",
          "rawMarkdown": "Could also think about Pruning, Weight clustering, ..."
        }
      ]
    },
    {
      "id": 2217960,
      "postDate": "2023-04-11T09:46:59.353Z",
      "content": "<p>With torch probably 2 as inference is fast. Struggling to submit a single TF model without using things like onnx. I however haven't tried torch ensemble as the individual models' perf is not great. Only 0.78 for one fold.</p>",
      "rawMarkdown": "With torch probably 2 as inference is fast. Struggling to submit a single TF model without using things like onnx. I however haven't tried torch ensemble as the individual models' perf is not great. Only 0.78 for one fold.",
      "votes": 1,
      "replies": [
        {
          "id": 2218534,
          "postDate": "2023-04-11T19:20:19.993Z",
          "content": "<p>As you mentioned <code>0.78 for one fold</code> and <code>probably 2</code>, do you mean you can ensemble 2*5fold models?</p>",
          "rawMarkdown": "As you mentioned `0.78 for one fold` and `probably 2 `, do you mean you can ensemble 2*5fold models?",
          "votes": 1,
          "replies": [
            {
              "id": 2219082,
              "postDate": "2023-04-12T09:24:01.290Z",
              "content": "<p>I meant that it would be possible to ensemble 2 single fold models</p>",
              "rawMarkdown": "I meant that it would be possible to ensemble 2 single fold models",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2222671,
      "postDate": "2023-04-15T13:19:39.723Z",
      "content": "<p>I am currently working on a Notebook analyzing the impact of different factors on inference time.<br>\n<a href=\"https://www.kaggle.com/code/iamleonie/what-impacts-your-cpu-inference-time\" target=\"_blank\">What impacts your CPU inference time?</a></p>\n<p>The key factors I found so far are:</p>\n<ul>\n<li>Batch size (try the biggest possible with your pipeline)</li>\n<li>Backbone choice</li>\n<li>Spectrogram size (impacted by <code>n_mels</code> and <code>hop_length</code>)</li>\n</ul>\n<p>Note that the inference times will differ slightly for your baseline code, but I think you will get a rough feeling for the tendency.</p>",
      "rawMarkdown": "I am currently working on a Notebook analyzing the impact of different factors on inference time.\n[What impacts your CPU inference time?](https://www.kaggle.com/code/iamleonie/what-impacts-your-cpu-inference-time)\n\nThe key factors I found so far are:\n* Batch size (try the biggest possible with your pipeline)\n* Backbone choice\n* Spectrogram size (impacted by `n_mels` and `hop_length`)\n\nNote that the inference times will differ slightly for your baseline code, but I think you will get a rough feeling for the tendency.",
      "votes": 2
    },
    {
      "id": 2218077,
      "postDate": "2023-04-11T12:11:54.360Z",
      "content": "<p>More than the quantity of models, <br>\nAdding more varied models to the ensemble will boost the performance.<br>\nOn the other hand, adding similar models make less difference to the performance! </p>",
      "rawMarkdown": "More than the quantity of models, \nAdding more varied models to the ensemble will boost the performance.\nOn the other hand, adding similar models make less difference to the performance! "
    },
    {
      "id": 2231823,
      "postDate": "2023-04-23T17:54:31.600Z",
      "content": "<p>depend on several factors, including the size and complexity of the models, the size of the dataset, and the specific hardware you are using</p>",
      "rawMarkdown": "depend on several factors, including the size and complexity of the models, the size of the dataset, and the specific hardware you are using",
      "votes": -4
    },
    {
      "id": 2270606,
      "postDate": "2023-05-23T09:12:41.237Z",
      "content": "<p>hey <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> <br>\nusing your code I'm able to get 5 models on 2 different configs whereas I could only get 3 with my original version, first of all thank you for that and secondly, I think with the way you made your code, ensembling different architectures isn't optimal beacause the loader is called again for every architecture even if it's the same mel-config. <br>\nTo fix that, I call my models this way:</p>\n<pre><code>models_ensemble = []\n\n config  config_enesemble:\n\n    config_models = []\n\n     ckpt_path  config.ckpt_path:\n\n        string = (ckpt_path)\n        filename = os.path.basename(string)\n        filename = os.path.splitext(filename)[]\n        parts = filename.split()\n        timm_model_name = parts[]\n         timm_model_name == :\n            timm_model_name = \n\n        model = timm.create_model(timm_model_name,in_chans=config.in_channels,num_classes=config.num_classes,pretrained=config.pretrained)\n        state_dict = torch.load(f=ckpt_path,map_location=device)\n        model.load_state_dict(state_dict, strict =)\n        model.()\n\n\n        \n        params = (model.parameters())\n         i  ((params)):\n            params[i].data = torch.(params[i].data***) / **\n\n        config_models.append(model)\n\n    models_ensemble.append((config, config_models))\n</code></pre>\n<p>even using your SED architecture you can do something similar I'm sure, <br>\nagain thx for your help and good luck on those 2 final days !</p>",
      "rawMarkdown": "hey @leonshangguan \nusing your code I'm able to get 5 models on 2 different configs whereas I could only get 3 with my original version, first of all thank you for that and secondly, I think with the way you made your code, ensembling different architectures isn't optimal beacause the loader is called again for every architecture even if it's the same mel-config. \nTo fix that, I call my models this way:\n```python\nmodels_ensemble = []\n\nfor config in config_enesemble:\n    \n    config_models = []\n    \n    for ckpt_path in config.ckpt_path:\n        \n        string = str(ckpt_path)\n        filename = os.path.basename(string)\n        filename = os.path.splitext(filename)[0]\n        parts = filename.split('_')\n        timm_model_name = parts[0]\n        if timm_model_name == 'efficientnet':\n            timm_model_name = 'efficientnet_b3'\n        \n        model = timm.create_model(timm_model_name,in_chans=config.in_channels,num_classes=config.num_classes,pretrained=config.pretrained)\n        state_dict = torch.load(f=ckpt_path,map_location=device)\n        model.load_state_dict(state_dict, strict =True)\n        model.eval()\n        \n\n        # this also made the inference much faster without losing cmap (it even increased my score for every comparison)\n        params = list(model.parameters())\n        for i in range(len(params)):\n            params[i].data = torch.round(params[i].data*10**4) / 10**4\n        \n        config_models.append(model)\n        \n    models_ensemble.append((config, config_models))\n```\n\neven using your SED architecture you can do something similar I'm sure, \nagain thx for your help and good luck on those 2 final days !"
    },
    {
      "id": 2228069,
      "postDate": "2023-04-20T08:49:22.257Z",
      "content": "<p>could you share your best lb score  of single model?</p>",
      "rawMarkdown": "could you share your best lb score  of single model?"
    },
    {
      "id": 2221005,
      "postDate": "2023-04-13T21:34:59.230Z",
      "content": "<p>Today I tried to make an ensemble of 2 SED models (single fold each), and sadly exceed the runtime.</p>",
      "rawMarkdown": "Today I tried to make an ensemble of 2 SED models (single fold each), and sadly exceed the runtime."
    },
    {
      "id": 2220210,
      "postDate": "2023-04-13T07:39:36.790Z",
      "content": "<p>Congrats on the jump to 4th. I guess you figured how to ensemble more or you got improvement from a single model?</p>",
      "rawMarkdown": "Congrats on the jump to 4th. I guess you figured how to ensemble more or you got improvement from a single model?",
      "replies": [
        {
          "id": 2220498,
          "postDate": "2023-04-13T12:57:10.333Z",
          "content": "<p>Not really. Just a small step which make my ensemble from 2 models to 3 models. I will post a notebook later this week.</p>",
          "rawMarkdown": "Not really. Just a small step which make my ensemble from 2 models to 3 models. I will post a notebook later this week.",
          "votes": 3,
          "replies": [
            {
              "id": 2220624,
              "postDate": "2023-04-13T14:50:35.817Z",
              "content": "<p>Please refrain from publishing high scoring notebooks that anyone can submit.</p>",
              "rawMarkdown": "Please refrain from publishing high scoring notebooks that anyone can submit.",
              "votes": 2
            },
            {
              "id": 2220637,
              "postDate": "2023-04-13T15:01:50.067Z",
              "content": "<p>I won’t release the trained models</p>",
              "rawMarkdown": "I won’t release the trained models",
              "votes": 1
            }
          ]
        },
        {
          "id": 2221082,
          "postDate": "2023-04-14T01:21:34.333Z",
          "content": "<p>I have the code shared <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "I have the code shared [here](https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference)",
          "votes": 2,
          "replies": [
            {
              "id": 2221387,
              "postDate": "2023-04-14T08:00:22.123Z",
              "content": "<p>Are you ensembling 3 models x 5 folds (=15 models)? </p>",
              "rawMarkdown": "Are you ensembling 3 models x 5 folds (=15 models)? "
            },
            {
              "id": 2221764,
              "postDate": "2023-04-14T14:46:31.303Z",
              "content": "<p>No, the preprocessing for input of the3 models are different which cost additional time, so I have to choose only one fold for each model.</p>",
              "rawMarkdown": "No, the preprocessing for input of the3 models are different which cost additional time, so I have to choose only one fold for each model.",
              "votes": 2
            },
            {
              "id": 2222182,
              "postDate": "2023-04-15T02:15:18.307Z",
              "content": "<p>Thank you for sharing !</p>",
              "rawMarkdown": "Thank you for sharing !"
            }
          ]
        }
      ]
    },
    {
      "id": 2218970,
      "postDate": "2023-04-12T07:05:38.800Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2218775,
      "author_name": "MYSO",
      "author_url": "",
      "post_date": "2023-04-12T02:49:27.430000",
      "content": "<p>I'm considering quantization to speed up inference.<br>\n<a href=\"https://pytorch.org/docs/stable/quantization.html\" target=\"_blank\">Quantization — PyTorch 2.0 documentation</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2221868,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2023-04-14T16:39:05.823000",
          "content": "<p>Could also think about Pruning, Weight clustering, …</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2217960,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-04-11T09:46:59.353000",
      "content": "<p>With torch probably 2 as inference is fast. Struggling to submit a single TF model without using things like onnx. I however haven't tried torch ensemble as the individual models' perf is not great. Only 0.78 for one fold.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2218534,
          "author_name": "Zhongkai Shangguan",
          "author_url": "",
          "post_date": "2023-04-11T19:20:19.993000",
          "content": "<p>As you mentioned <code>0.78 for one fold</code> and <code>probably 2</code>, do you mean you can ensemble 2*5fold models?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2219082,
              "author_name": "nymfree",
              "author_url": "",
              "post_date": "2023-04-12T09:24:01.290000",
              "content": "<p>I meant that it would be possible to ensemble 2 single fold models</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2222671,
      "author_name": "Leonie",
      "author_url": "",
      "post_date": "2023-04-15T13:19:39.723000",
      "content": "<p>I am currently working on a Notebook analyzing the impact of different factors on inference time.<br>\n<a href=\"https://www.kaggle.com/code/iamleonie/what-impacts-your-cpu-inference-time\" target=\"_blank\">What impacts your CPU inference time?</a></p>\n<p>The key factors I found so far are:</p>\n<ul>\n<li>Batch size (try the biggest possible with your pipeline)</li>\n<li>Backbone choice</li>\n<li>Spectrogram size (impacted by <code>n_mels</code> and <code>hop_length</code>)</li>\n</ul>\n<p>Note that the inference times will differ slightly for your baseline code, but I think you will get a rough feeling for the tendency.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2218077,
      "author_name": "‎ Srihari",
      "author_url": "",
      "post_date": "2023-04-11T12:11:54.360000",
      "content": "<p>More than the quantity of models, <br>\nAdding more varied models to the ensemble will boost the performance.<br>\nOn the other hand, adding similar models make less difference to the performance! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2231823,
      "author_name": "Mehrdad Dastouri",
      "author_url": "",
      "post_date": "2023-04-23T17:54:31.600000",
      "content": "<p>depend on several factors, including the size and complexity of the models, the size of the dataset, and the specific hardware you are using</p>",
      "votes": -4,
      "replies": []
    },
    {
      "id": 2270606,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-05-23T09:12:41.237000",
      "content": "<p>hey <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> <br>\nusing your code I'm able to get 5 models on 2 different configs whereas I could only get 3 with my original version, first of all thank you for that and secondly, I think with the way you made your code, ensembling different architectures isn't optimal beacause the loader is called again for every architecture even if it's the same mel-config. <br>\nTo fix that, I call my models this way:</p>\n<pre><code>models_ensemble = []\n\n config  config_enesemble:\n\n    config_models = []\n\n     ckpt_path  config.ckpt_path:\n\n        string = (ckpt_path)\n        filename = os.path.basename(string)\n        filename = os.path.splitext(filename)[]\n        parts = filename.split()\n        timm_model_name = parts[]\n         timm_model_name == :\n            timm_model_name = \n\n        model = timm.create_model(timm_model_name,in_chans=config.in_channels,num_classes=config.num_classes,pretrained=config.pretrained)\n        state_dict = torch.load(f=ckpt_path,map_location=device)\n        model.load_state_dict(state_dict, strict =)\n        model.()\n\n\n        \n        params = (model.parameters())\n         i  ((params)):\n            params[i].data = torch.(params[i].data***) / **\n\n        config_models.append(model)\n\n    models_ensemble.append((config, config_models))\n</code></pre>\n<p>even using your SED architecture you can do something similar I'm sure, <br>\nagain thx for your help and good luck on those 2 final days !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2228069,
      "author_name": "gnannan",
      "author_url": "",
      "post_date": "2023-04-20T08:49:22.257000",
      "content": "<p>could you share your best lb score  of single model?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2221005,
      "author_name": "Maximiliano Diaz Battan",
      "author_url": "",
      "post_date": "2023-04-13T21:34:59.230000",
      "content": "<p>Today I tried to make an ensemble of 2 SED models (single fold each), and sadly exceed the runtime.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2220210,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-04-13T07:39:36.790000",
      "content": "<p>Congrats on the jump to 4th. I guess you figured how to ensemble more or you got improvement from a single model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2220498,
          "author_name": "Zhongkai Shangguan",
          "author_url": "",
          "post_date": "2023-04-13T12:57:10.333000",
          "content": "<p>Not really. Just a small step which make my ensemble from 2 models to 3 models. I will post a notebook later this week.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2220624,
              "author_name": "S. Tomizawa",
              "author_url": "",
              "post_date": "2023-04-13T14:50:35.817000",
              "content": "<p>Please refrain from publishing high scoring notebooks that anyone can submit.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2220637,
              "author_name": "Zhongkai Shangguan",
              "author_url": "",
              "post_date": "2023-04-13T15:01:50.067000",
              "content": "<p>I won’t release the trained models</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2221082,
          "author_name": "Zhongkai Shangguan",
          "author_url": "",
          "post_date": "2023-04-14T01:21:34.333000",
          "content": "<p>I have the code shared <a href=\"https://www.kaggle.com/code/leonshangguan/faster-eb0-sed-model-inference\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2221387,
              "author_name": "S. Tomizawa",
              "author_url": "",
              "post_date": "2023-04-14T08:00:22.123000",
              "content": "<p>Are you ensembling 3 models x 5 folds (=15 models)? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2221764,
              "author_name": "Zhongkai Shangguan",
              "author_url": "",
              "post_date": "2023-04-14T14:46:31.303000",
              "content": "<p>No, the preprocessing for input of the3 models are different which cost additional time, so I have to choose only one fold for each model.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2222182,
              "author_name": "Jiawen9",
              "author_url": "",
              "post_date": "2023-04-15T02:15:18.307000",
              "content": "<p>Thank you for sharing !</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2218970,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-12T07:05:38.800000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2217335": "With the cpu time<120min restriction, may I ask how many models can you ensemble to your final result?\nFor me its 2 to 3 models depend on the model size. Any idea how to accelerate cpu inference time?\n\n----------------------------------------------------------------\nSeems my limit is 3 models. Any idea how to improve?",
    "2218775": "I'm considering quantization to speed up inference.\n[Quantization — PyTorch 2.0 documentation](https://pytorch.org/docs/stable/quantization.html)",
    "2217960": "With torch probably 2 as inference is fast. Struggling to submit a single TF model without using things like onnx. I however haven't tried torch ensemble as the individual models' perf is not great. Only 0.78 for one fold.",
    "2222671": "I am currently working on a Notebook analyzing the impact of different factors on inference time.\n[What impacts your CPU inference time?](https://www.kaggle.com/code/iamleonie/what-impacts-your-cpu-inference-time)\n\nThe key factors I found so far are:\n* Batch size (try the biggest possible with your pipeline)\n* Backbone choice\n* Spectrogram size (impacted by `n_mels` and `hop_length`)\n\nNote that the inference times will differ slightly for your baseline code, but I think you will get a rough feeling for the tendency.",
    "2218077": "More than the quantity of models, \nAdding more varied models to the ensemble will boost the performance.\nOn the other hand, adding similar models make less difference to the performance! ",
    "2231823": "depend on several factors, including the size and complexity of the models, the size of the dataset, and the specific hardware you are using",
    "2270606": "hey @leonshangguan \nusing your code I'm able to get 5 models on 2 different configs whereas I could only get 3 with my original version, first of all thank you for that and secondly, I think with the way you made your code, ensembling different architectures isn't optimal beacause the loader is called again for every architecture even if it's the same mel-config. \nTo fix that, I call my models this way:\n```python\nmodels_ensemble = []\n\nfor config in config_enesemble:\n    \n    config_models = []\n    \n    for ckpt_path in config.ckpt_path:\n        \n        string = str(ckpt_path)\n        filename = os.path.basename(string)\n        filename = os.path.splitext(filename)[0]\n        parts = filename.split('_')\n        timm_model_name = parts[0]\n        if timm_model_name == 'efficientnet':\n            timm_model_name = 'efficientnet_b3'\n        \n        model = timm.create_model(timm_model_name,in_chans=config.in_channels,num_classes=config.num_classes,pretrained=config.pretrained)\n        state_dict = torch.load(f=ckpt_path,map_location=device)\n        model.load_state_dict(state_dict, strict =True)\n        model.eval()\n        \n\n        # this also made the inference much faster without losing cmap (it even increased my score for every comparison)\n        params = list(model.parameters())\n        for i in range(len(params)):\n            params[i].data = torch.round(params[i].data*10**4) / 10**4\n        \n        config_models.append(model)\n        \n    models_ensemble.append((config, config_models))\n```\n\neven using your SED architecture you can do something similar I'm sure, \nagain thx for your help and good luck on those 2 final days !",
    "2228069": "could you share your best lb score  of single model?",
    "2221005": "Today I tried to make an ensemble of 2 SED models (single fold each), and sadly exceed the runtime.",
    "2220210": "Congrats on the jump to 4th. I guess you figured how to ensemble more or you got improvement from a single model?",
    "2218970": ""
  }
}