{
  "id": 202826,
  "title": "How to best ensemble tf models?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202826",
  "author_name": "",
  "post_date": "2020-12-12T07:23:57.630596600Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I used to use PyTorch before and I am trying to get to know TF too.</p>\n<p><strong>Approach 1</strong>: Load each model one-by-one, do a pass through the test set, average the outputs, generate the submission file.</p>\n<p>Pros: Easy to implement<br>\nCons: We need to do N passes for N models.</p>\n<p>Takes about 45 mins for a single effnet-B3 model. So if there are 5 models (maybe folds, maybe different archs), plan for 2 movies to watch while waiting to get the score.</p>\n<p><strong>Approach 2</strong>: Load all the models in memory, make each of their outputs pass through an <code>Average</code> layer to generate the outputs.</p>\n<p>Something like this:</p>\n<pre><code>effnets_funcs = {\n    0: efn.EfficientNetB0,\n    1: efn.EfficientNetB1, \n}\n\n# to get unique name for layers\nimport random, string\ndef randomword(length):\n    letters = string.ascii_lowercase\n    return ''.join(random.choice(letters) for i in range(length))\n</code></pre>\n<p>The model loading function. Here we're not passing in the <code>Input</code> layer, just the <code>input_shape</code>, later on, we use a common <code>Input</code> layer.</p>\n<pre><code>def model_fn(effnet, input_shape, N_CLASSES):\n    input_image = L.Input(shape=input_shape)\n\n    model_func = effnets_funcs[effnet]\n    base_model = model_func(\n        input_shape=(None, None, CHANNELS),\n        include_top=False, \n        weights=None, \n        pooling='avg',\n    )\n    base_model._name = randomword(10)\n\n    model = tf.keras.Sequential([\n        base_model,\n        L.Dropout(.25),\n        L.Dense(N_CLASSES, activation='softmax')\n    ])\n    return model\n\n\n# Takes 2min 17s for 5 B3 + 5 B4 models to load\ndef load_models(mlist):\n    with strategy.scope():\n        K.clear_session()\n        models = []\n        input_shape = (None, None, CHANNELS)\n        for b_num, model_path in mlist:\n            model = model_fn(b_num, input_shape, N_CLASSES)\n            print(model_path)\n            model.load_weights(model_path)\n            models.append(model)\n    return models\n</code></pre>\n<p>Now ensemble them, using a common <code>Input</code> layer:</p>\n<pre><code># Takes 57s for 5 B3 + 5 B4 models\ndef make_ensemble(models):\n    input_image = L.Input(shape=(None, None, CHANNELS))\n    outputs = [model(input_image) for model in tqdm(models)]\n\n    y = L.Average()(outputs)    \n    model = tf.keras.Model(input_image, y)\n    return model\n</code></pre>\n<p>Didn't submit using this approach yet.</p>\n<hr>\n<p>Can the TF pros share how do they do it? I feel there should be a better, efficient way to do this.</p>",
  "messages": [
    {
      "id": "1109884",
      "postDate": "12/12/2020 07:23:57",
      "content": "<p>I used to use PyTorch before and I am trying to get to know TF too.</p>\n<p><strong>Approach 1</strong>: Load each model one-by-one, do a pass through the test set, average the outputs, generate the submission file.</p>\n<p>Pros: Easy to implement<br>\nCons: We need to do N passes for N models.</p>\n<p>Takes about 45 mins for a single effnet-B3 model. So if there are 5 models (maybe folds, maybe different archs), plan for 2 movies to watch while waiting to get the score.</p>\n<p><strong>Approach 2</strong>: Load all the models in memory, make each of their outputs pass through an <code>Average</code> layer to generate the outputs.</p>\n<p>Something like this:</p>\n<pre><code>effnets_funcs = {\n    0: efn.EfficientNetB0,\n    1: efn.EfficientNetB1, \n}\n\n# to get unique name for layers\nimport random, string\ndef randomword(length):\n    letters = string.ascii_lowercase\n    return ''.join(random.choice(letters) for i in range(length))\n</code></pre>\n<p>The model loading function. Here we're not passing in the <code>Input</code> layer, just the <code>input_shape</code>, later on, we use a common <code>Input</code> layer.</p>\n<pre><code>def model_fn(effnet, input_shape, N_CLASSES):\n    input_image = L.Input(shape=input_shape)\n\n    model_func = effnets_funcs[effnet]\n    base_model = model_func(\n        input_shape=(None, None, CHANNELS),\n        include_top=False, \n        weights=None, \n        pooling='avg',\n    )\n    base_model._name = randomword(10)\n\n    model = tf.keras.Sequential([\n        base_model,\n        L.Dropout(.25),\n        L.Dense(N_CLASSES, activation='softmax')\n    ])\n    return model\n\n\n# Takes 2min 17s for 5 B3 + 5 B4 models to load\ndef load_models(mlist):\n    with strategy.scope():\n        K.clear_session()\n        models = []\n        input_shape = (None, None, CHANNELS)\n        for b_num, model_path in mlist:\n            model = model_fn(b_num, input_shape, N_CLASSES)\n            print(model_path)\n            model.load_weights(model_path)\n            models.append(model)\n    return models\n</code></pre>\n<p>Now ensemble them, using a common <code>Input</code> layer:</p>\n<pre><code># Takes 57s for 5 B3 + 5 B4 models\ndef make_ensemble(models):\n    input_image = L.Input(shape=(None, None, CHANNELS))\n    outputs = [model(input_image) for model in tqdm(models)]\n\n    y = L.Average()(outputs)    \n    model = tf.keras.Model(input_image, y)\n    return model\n</code></pre>\n<p>Didn't submit using this approach yet.</p>\n<hr>\n<p>Can the TF pros share how do they do it? I feel there should be a better, efficient way to do this.</p>",
      "rawMarkdown": "I used to use PyTorch before and I am trying to get to know TF too.\n\n**Approach 1**: Load each model one-by-one, do a pass through the test set, average the outputs, generate the submission file.\n\nPros: Easy to implement\nCons: We need to do N passes for N models.\n\nTakes about 45 mins for a single effnet-B3 model. So if there are 5 models (maybe folds, maybe different archs), plan for 2 movies to watch while waiting to get the score.\n\n**Approach 2**: Load all the models in memory, make each of their outputs pass through an `Average` layer to generate the outputs.\n\nSomething like this:\n\n```\neffnets_funcs = {\n    0: efn.EfficientNetB0,\n    1: efn.EfficientNetB1, \n}\n\n# to get unique name for layers\nimport random, string\ndef randomword(length):\n    letters = string.ascii_lowercase\n    return ''.join(random.choice(letters) for i in range(length))\n```\n\nThe model loading function. Here we're not passing in the `Input` layer, just the `input_shape`, later on, we use a common `Input` layer.\n```\ndef model_fn(effnet, input_shape, N_CLASSES):\n    input_image = L.Input(shape=input_shape)\n    \n    model_func = effnets_funcs[effnet]\n    base_model = model_func(\n        input_shape=(None, None, CHANNELS),\n        include_top=False, \n        weights=None, \n        pooling='avg',\n    )\n    base_model._name = randomword(10)\n\n    model = tf.keras.Sequential([\n        base_model,\n        L.Dropout(.25),\n        L.Dense(N_CLASSES, activation='softmax')\n    ])\n    return model\n\n\n# Takes 2min 17s for 5 B3 + 5 B4 models to load\ndef load_models(mlist):\n    with strategy.scope():\n        K.clear_session()\n        models = []\n        input_shape = (None, None, CHANNELS)\n        for b_num, model_path in mlist:\n            model = model_fn(b_num, input_shape, N_CLASSES)\n            print(model_path)\n            model.load_weights(model_path)\n            models.append(model)\n    return models\n```\n\nNow ensemble them, using a common `Input` layer:\n```\n# Takes 57s for 5 B3 + 5 B4 models\ndef make_ensemble(models):\n    input_image = L.Input(shape=(None, None, CHANNELS))\n    outputs = [model(input_image) for model in tqdm(models)]\n    \n    y = L.Average()(outputs)    \n    model = tf.keras.Model(input_image, y)\n    return model\n```\nDidn't submit using this approach yet.\n\n<hr/>\n\nCan the TF pros share how do they do it? I feel there should be a better, efficient way to do this.",
      "votes": null
    },
    {
      "id": "1150400",
      "postDate": "01/12/2021 15:03:38",
      "content": "<p>Try using the functional API.</p>",
      "rawMarkdown": "Try using the functional API.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1150400,
      "author_name": "hypnotu",
      "author_url": "",
      "post_date": "01/12/2021 15:03:38",
      "content": "<p>Try using the functional API.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1109884": "I used to use PyTorch before and I am trying to get to know TF too.\n\n**Approach 1**: Load each model one-by-one, do a pass through the test set, average the outputs, generate the submission file.\n\nPros: Easy to implement\nCons: We need to do N passes for N models.\n\nTakes about 45 mins for a single effnet-B3 model. So if there are 5 models (maybe folds, maybe different archs), plan for 2 movies to watch while waiting to get the score.\n\n**Approach 2**: Load all the models in memory, make each of their outputs pass through an `Average` layer to generate the outputs.\n\nSomething like this:\n\n```\neffnets_funcs = {\n    0: efn.EfficientNetB0,\n    1: efn.EfficientNetB1, \n}\n\n# to get unique name for layers\nimport random, string\ndef randomword(length):\n    letters = string.ascii_lowercase\n    return ''.join(random.choice(letters) for i in range(length))\n```\n\nThe model loading function. Here we're not passing in the `Input` layer, just the `input_shape`, later on, we use a common `Input` layer.\n```\ndef model_fn(effnet, input_shape, N_CLASSES):\n    input_image = L.Input(shape=input_shape)\n    \n    model_func = effnets_funcs[effnet]\n    base_model = model_func(\n        input_shape=(None, None, CHANNELS),\n        include_top=False, \n        weights=None, \n        pooling='avg',\n    )\n    base_model._name = randomword(10)\n\n    model = tf.keras.Sequential([\n        base_model,\n        L.Dropout(.25),\n        L.Dense(N_CLASSES, activation='softmax')\n    ])\n    return model\n\n\n# Takes 2min 17s for 5 B3 + 5 B4 models to load\ndef load_models(mlist):\n    with strategy.scope():\n        K.clear_session()\n        models = []\n        input_shape = (None, None, CHANNELS)\n        for b_num, model_path in mlist:\n            model = model_fn(b_num, input_shape, N_CLASSES)\n            print(model_path)\n            model.load_weights(model_path)\n            models.append(model)\n    return models\n```\n\nNow ensemble them, using a common `Input` layer:\n```\n# Takes 57s for 5 B3 + 5 B4 models\ndef make_ensemble(models):\n    input_image = L.Input(shape=(None, None, CHANNELS))\n    outputs = [model(input_image) for model in tqdm(models)]\n    \n    y = L.Average()(outputs)    \n    model = tf.keras.Model(input_image, y)\n    return model\n```\nDidn't submit using this approach yet.\n\n<hr/>\n\nCan the TF pros share how do they do it? I feel there should be a better, efficient way to do this.",
    "1150400": "Try using the functional API."
  },
  "source": "meta"
}