{
  "id": 348188,
  "title": "Pytorch/TF averaging between folds and within a fold",
  "url": "/competitions/hubmap-organ-segmentation/discussion/348188",
  "author_name": "",
  "post_date": "2022-08-27T12:19:11.148420400Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I’ll start this competition with sharing some techniques I have collected through the years and usually use when it’s time to average some fold/models or within a fold with different saved best steps/epochs.<br>\nMaybe it can help or just give some diversity to someone’s solution.</p>\n<p>Collaboration and teamwork are what Kaggle is all about. Share knowledge and experience, to ultimately create a model that will benefit those in need. 👍<br>\nThey are taken straight from a solution, so of course you need your own modification.<br>\nThe Tensorflow codes I haven’t used for I while, but they should work.</p>\n<p>Averaging during training<br>\n<strong>Keras/Tensorflow</strong></p>\n<pre><code>######################\ndef get_model():\n  inputs = keras.Input(shape=(128,))\n  outputs = layers.Dense(1, activation='sigmoid')(inputs)\n  return keras.Model(inputs, outputs)\n\nmodel1 = get_model()\nmodel2 = get_model()\nmodel3 = get_model()\n\ninputs = keras.Input(shape=(128,))\ny1 = model1(inputs)\ny2 = model2(inputs)\ny3 = model3(inputs)\noutputs = layers.average([y1, y2, y3])\nensemble_model = keras.Model(inputs=inputs, outputs=outputs)\n</code></pre>\n<p>Averaging within saved states in the same fold.<br>\nCan also be used during training checking if no. of states gives better score before saving the model.</p>\n<p><strong>Pytorch</strong></p>\n<pre><code>savemodels = []\n#####per epoch......or can be modified doing per steps etc\nsvname = f\"model_epoch_{e}.pth\"\ntorch.save({'model': model.state_dict()},\n                        svname)\nsavemodels.append(svname)\n\nprint('nomodels for ensemble', len(savemodels))\n\nswa_state_dict = torch.load(savemodels[0], map_location=lambda storage, loc: storage)['model']\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] = torch.zeros_like(v)\nfor mmm in range(len(savemodels)):\n\n    state_dict = torch.load(savemodels[mmm], map_location=lambda storage, loc: storage)['model']\n\n    for k,v in state_dict.items():\n      swa_state_dict[k] += v\n\n    print('ensemble model no.', mmm)\n    print('ensemble model name', savemodels[mmm])\n\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] /= np.array(len(savemodels)).astype(float)\n\n\nmodel.load_state_dict(swa_state_dict)\nprint('ensemble done')\n</code></pre>\n<p><strong>Tensorflow</strong></p>\n<pre><code>all_models = list()\n\nfor i in range(5):\n    model_path = f'…'\n    model = model()\n    model.load_weights(model_path)\n    all_models.append(model)\n\nmembers = all_models\n\ndef model_weight_ensemble(members, weights):\n    n_layers = len(members[0].get_weights())\n    avg_model_weights = list()\n    for layer in range(n_layers):\n        layer_weights = array([model.get_weights()[layer] for model in members])\n        avg_layer_weights = average(layer_weights, axis=0, weights=weights)\n        avg_model_weights.append(avg_layer_weights)\n    model = clone_model(members[0])\n    model.set_weights(avg_model_weights)\n    return model\n\nprint('Loaded %d models' % len(members))\nn_models = len(members)\nweights = [1/n_models for i in range(1, n_models+1)]\nmodel = model_weight_ensemble(members, weights)\nmodel.save_weights('…')\nmodel.summary()\n\n\n##########################\n</code></pre>\n<p>Average between models/folds.</p>\n<p><strong>Pytorch</strong></p>\n<pre><code>#########################\nclass Model:\n    def __init__(self, models):\n        self.models = models\n\n    def __call__(self, x):\n        res = []\n       # x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        res = torch.stack(res)\n        return torch.mean(res, dim=0)\n\nWith TTA\n\nclass Model:\n    def __init__(self, models):\n        self.models = models\n    def __call__(self, x):\n        res = []\n        x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        x = torch.flip(x,dims = [-1])\n        with torch.no_grad():\n            for m in self.models:\n                flipped_mask = m(x)\n                mask = torch.flip(flipped_mask,dims = [-1])\n                res.append(mask)\n        res = torch.stack(res)\n        return torch.sigmoid(torch.mean(res, dim=0))\n\nmodel = Model([model8,model9,model10])\n</code></pre>\n<p>Happy Kaggling! 🙂</p>",
  "messages": [
    {
      "id": "1915833",
      "postDate": "08/27/2022 12:19:11",
      "content": "<p>I’ll start this competition with sharing some techniques I have collected through the years and usually use when it’s time to average some fold/models or within a fold with different saved best steps/epochs.<br>\nMaybe it can help or just give some diversity to someone’s solution.</p>\n<p>Collaboration and teamwork are what Kaggle is all about. Share knowledge and experience, to ultimately create a model that will benefit those in need. 👍<br>\nThey are taken straight from a solution, so of course you need your own modification.<br>\nThe Tensorflow codes I haven’t used for I while, but they should work.</p>\n<p>Averaging during training<br>\n<strong>Keras/Tensorflow</strong></p>\n<pre><code>######################\ndef get_model():\n  inputs = keras.Input(shape=(128,))\n  outputs = layers.Dense(1, activation='sigmoid')(inputs)\n  return keras.Model(inputs, outputs)\n\nmodel1 = get_model()\nmodel2 = get_model()\nmodel3 = get_model()\n\ninputs = keras.Input(shape=(128,))\ny1 = model1(inputs)\ny2 = model2(inputs)\ny3 = model3(inputs)\noutputs = layers.average([y1, y2, y3])\nensemble_model = keras.Model(inputs=inputs, outputs=outputs)\n</code></pre>\n<p>Averaging within saved states in the same fold.<br>\nCan also be used during training checking if no. of states gives better score before saving the model.</p>\n<p><strong>Pytorch</strong></p>\n<pre><code>savemodels = []\n#####per epoch......or can be modified doing per steps etc\nsvname = f\"model_epoch_{e}.pth\"\ntorch.save({'model': model.state_dict()},\n                        svname)\nsavemodels.append(svname)\n\nprint('nomodels for ensemble', len(savemodels))\n\nswa_state_dict = torch.load(savemodels[0], map_location=lambda storage, loc: storage)['model']\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] = torch.zeros_like(v)\nfor mmm in range(len(savemodels)):\n\n    state_dict = torch.load(savemodels[mmm], map_location=lambda storage, loc: storage)['model']\n\n    for k,v in state_dict.items():\n      swa_state_dict[k] += v\n\n    print('ensemble model no.', mmm)\n    print('ensemble model name', savemodels[mmm])\n\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] /= np.array(len(savemodels)).astype(float)\n\n\nmodel.load_state_dict(swa_state_dict)\nprint('ensemble done')\n</code></pre>\n<p><strong>Tensorflow</strong></p>\n<pre><code>all_models = list()\n\nfor i in range(5):\n    model_path = f'…'\n    model = model()\n    model.load_weights(model_path)\n    all_models.append(model)\n\nmembers = all_models\n\ndef model_weight_ensemble(members, weights):\n    n_layers = len(members[0].get_weights())\n    avg_model_weights = list()\n    for layer in range(n_layers):\n        layer_weights = array([model.get_weights()[layer] for model in members])\n        avg_layer_weights = average(layer_weights, axis=0, weights=weights)\n        avg_model_weights.append(avg_layer_weights)\n    model = clone_model(members[0])\n    model.set_weights(avg_model_weights)\n    return model\n\nprint('Loaded %d models' % len(members))\nn_models = len(members)\nweights = [1/n_models for i in range(1, n_models+1)]\nmodel = model_weight_ensemble(members, weights)\nmodel.save_weights('…')\nmodel.summary()\n\n\n##########################\n</code></pre>\n<p>Average between models/folds.</p>\n<p><strong>Pytorch</strong></p>\n<pre><code>#########################\nclass Model:\n    def __init__(self, models):\n        self.models = models\n\n    def __call__(self, x):\n        res = []\n       # x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        res = torch.stack(res)\n        return torch.mean(res, dim=0)\n\nWith TTA\n\nclass Model:\n    def __init__(self, models):\n        self.models = models\n    def __call__(self, x):\n        res = []\n        x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        x = torch.flip(x,dims = [-1])\n        with torch.no_grad():\n            for m in self.models:\n                flipped_mask = m(x)\n                mask = torch.flip(flipped_mask,dims = [-1])\n                res.append(mask)\n        res = torch.stack(res)\n        return torch.sigmoid(torch.mean(res, dim=0))\n\nmodel = Model([model8,model9,model10])\n</code></pre>\n<p>Happy Kaggling! 🙂</p>",
      "rawMarkdown": "I’ll start this competition with sharing some techniques I have collected through the years and usually use when it’s time to average some fold/models or within a fold with different saved best steps/epochs.\nMaybe it can help or just give some diversity to someone’s solution.\n\nCollaboration and teamwork are what Kaggle is all about. Share knowledge and experience, to ultimately create a model that will benefit those in need. 👍\nThey are taken straight from a solution, so of course you need your own modification.\nThe Tensorflow codes I haven’t used for I while, but they should work.\n\nAveraging during training\n**Keras/Tensorflow**\n\n```\n######################\ndef get_model():\n  inputs = keras.Input(shape=(128,))\n  outputs = layers.Dense(1, activation='sigmoid')(inputs)\n  return keras.Model(inputs, outputs)\n\nmodel1 = get_model()\nmodel2 = get_model()\nmodel3 = get_model()\n\ninputs = keras.Input(shape=(128,))\ny1 = model1(inputs)\ny2 = model2(inputs)\ny3 = model3(inputs)\noutputs = layers.average([y1, y2, y3])\nensemble_model = keras.Model(inputs=inputs, outputs=outputs)\n```\n\nAveraging within saved states in the same fold.\nCan also be used during training checking if no. of states gives better score before saving the model.\n\n**Pytorch**\n```\nsavemodels = []\n#####per epoch......or can be modified doing per steps etc\nsvname = f\"model_epoch_{e}.pth\"\ntorch.save({'model': model.state_dict()},\n                        svname)\nsavemodels.append(svname)\n\nprint('nomodels for ensemble', len(savemodels))\n\nswa_state_dict = torch.load(savemodels[0], map_location=lambda storage, loc: storage)['model']\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] = torch.zeros_like(v)\nfor mmm in range(len(savemodels)):\n\n    state_dict = torch.load(savemodels[mmm], map_location=lambda storage, loc: storage)['model']\n\n    for k,v in state_dict.items():\n      swa_state_dict[k] += v\n\n    print('ensemble model no.', mmm)\n    print('ensemble model name', savemodels[mmm])\n\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] /= np.array(len(savemodels)).astype(float)\n\n\nmodel.load_state_dict(swa_state_dict)\nprint('ensemble done')\n\n```\n**Tensorflow**\n```\nall_models = list()\n\nfor i in range(5):\n    model_path = f'…'\n    model = model()\n    model.load_weights(model_path)\n    all_models.append(model)\n\nmembers = all_models\n\ndef model_weight_ensemble(members, weights):\n    n_layers = len(members[0].get_weights())\n    avg_model_weights = list()\n    for layer in range(n_layers):\n        layer_weights = array([model.get_weights()[layer] for model in members])\n        avg_layer_weights = average(layer_weights, axis=0, weights=weights)\n        avg_model_weights.append(avg_layer_weights)\n    model = clone_model(members[0])\n    model.set_weights(avg_model_weights)\n    return model\n\nprint('Loaded %d models' % len(members))\nn_models = len(members)\nweights = [1/n_models for i in range(1, n_models+1)]\nmodel = model_weight_ensemble(members, weights)\nmodel.save_weights('…')\nmodel.summary()\n\n\n##########################\n```\nAverage between models/folds.\n\n**Pytorch**\n```\n#########################\nclass Model:\n    def __init__(self, models):\n        self.models = models\n    \n    def __call__(self, x):\n        res = []\n       # x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        res = torch.stack(res)\n        return torch.mean(res, dim=0)\n\nWith TTA\n\nclass Model:\n    def __init__(self, models):\n        self.models = models\n    def __call__(self, x):\n        res = []\n        x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        x = torch.flip(x,dims = [-1])\n        with torch.no_grad():\n            for m in self.models:\n                flipped_mask = m(x)\n                mask = torch.flip(flipped_mask,dims = [-1])\n                res.append(mask)\n        res = torch.stack(res)\n        return torch.sigmoid(torch.mean(res, dim=0))\n\nmodel = Model([model8,model9,model10])\n```\n\nHappy Kaggling! 🙂",
      "votes": null
    },
    {
      "id": "1918811",
      "postDate": "08/29/2022 21:29:02",
      "content": "<p>Does this improve your score? Averaging seems to make my mine worse than using a single fold </p>",
      "rawMarkdown": "Does this improve your score? Averaging seems to make my mine worse than using a single fold",
      "votes": null
    },
    {
      "id": "1918846",
      "postDate": "08/29/2022 22:51:39",
      "content": "<p>Haven’t reached the first model/fold yet 😉 but in general Average/ensemble do help and also minimize overfitting.</p>",
      "rawMarkdown": "Haven’t reached the first model/fold yet 😉 but in general Average/ensemble do help and also minimize overfitting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1918811,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "08/29/2022 21:29:02",
      "content": "<p>Does this improve your score? Averaging seems to make my mine worse than using a single fold </p>",
      "votes": null,
      "replies": [
        {
          "id": 1918846,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/29/2022 22:51:39",
          "content": "<p>Haven’t reached the first model/fold yet 😉 but in general Average/ensemble do help and also minimize overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1915833": "I’ll start this competition with sharing some techniques I have collected through the years and usually use when it’s time to average some fold/models or within a fold with different saved best steps/epochs.\nMaybe it can help or just give some diversity to someone’s solution.\n\nCollaboration and teamwork are what Kaggle is all about. Share knowledge and experience, to ultimately create a model that will benefit those in need. 👍\nThey are taken straight from a solution, so of course you need your own modification.\nThe Tensorflow codes I haven’t used for I while, but they should work.\n\nAveraging during training\n**Keras/Tensorflow**\n\n```\n######################\ndef get_model():\n  inputs = keras.Input(shape=(128,))\n  outputs = layers.Dense(1, activation='sigmoid')(inputs)\n  return keras.Model(inputs, outputs)\n\nmodel1 = get_model()\nmodel2 = get_model()\nmodel3 = get_model()\n\ninputs = keras.Input(shape=(128,))\ny1 = model1(inputs)\ny2 = model2(inputs)\ny3 = model3(inputs)\noutputs = layers.average([y1, y2, y3])\nensemble_model = keras.Model(inputs=inputs, outputs=outputs)\n```\n\nAveraging within saved states in the same fold.\nCan also be used during training checking if no. of states gives better score before saving the model.\n\n**Pytorch**\n```\nsavemodels = []\n#####per epoch......or can be modified doing per steps etc\nsvname = f\"model_epoch_{e}.pth\"\ntorch.save({'model': model.state_dict()},\n                        svname)\nsavemodels.append(svname)\n\nprint('nomodels for ensemble', len(savemodels))\n\nswa_state_dict = torch.load(savemodels[0], map_location=lambda storage, loc: storage)['model']\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] = torch.zeros_like(v)\nfor mmm in range(len(savemodels)):\n\n    state_dict = torch.load(savemodels[mmm], map_location=lambda storage, loc: storage)['model']\n\n    for k,v in state_dict.items():\n      swa_state_dict[k] += v\n\n    print('ensemble model no.', mmm)\n    print('ensemble model name', savemodels[mmm])\n\nfor k,v in swa_state_dict.items():\n    swa_state_dict[k] /= np.array(len(savemodels)).astype(float)\n\n\nmodel.load_state_dict(swa_state_dict)\nprint('ensemble done')\n\n```\n**Tensorflow**\n```\nall_models = list()\n\nfor i in range(5):\n    model_path = f'…'\n    model = model()\n    model.load_weights(model_path)\n    all_models.append(model)\n\nmembers = all_models\n\ndef model_weight_ensemble(members, weights):\n    n_layers = len(members[0].get_weights())\n    avg_model_weights = list()\n    for layer in range(n_layers):\n        layer_weights = array([model.get_weights()[layer] for model in members])\n        avg_layer_weights = average(layer_weights, axis=0, weights=weights)\n        avg_model_weights.append(avg_layer_weights)\n    model = clone_model(members[0])\n    model.set_weights(avg_model_weights)\n    return model\n\nprint('Loaded %d models' % len(members))\nn_models = len(members)\nweights = [1/n_models for i in range(1, n_models+1)]\nmodel = model_weight_ensemble(members, weights)\nmodel.save_weights('…')\nmodel.summary()\n\n\n##########################\n```\nAverage between models/folds.\n\n**Pytorch**\n```\n#########################\nclass Model:\n    def __init__(self, models):\n        self.models = models\n    \n    def __call__(self, x):\n        res = []\n       # x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        res = torch.stack(res)\n        return torch.mean(res, dim=0)\n\nWith TTA\n\nclass Model:\n    def __init__(self, models):\n        self.models = models\n    def __call__(self, x):\n        res = []\n        x = x.cuda()\n        with torch.no_grad():\n            for m in self.models:\n                res.append(m(x))\n        x = torch.flip(x,dims = [-1])\n        with torch.no_grad():\n            for m in self.models:\n                flipped_mask = m(x)\n                mask = torch.flip(flipped_mask,dims = [-1])\n                res.append(mask)\n        res = torch.stack(res)\n        return torch.sigmoid(torch.mean(res, dim=0))\n\nmodel = Model([model8,model9,model10])\n```\n\nHappy Kaggling! 🙂",
    "1918811": "Does this improve your score? Averaging seems to make my mine worse than using a single fold",
    "1918846": "Haven’t reached the first model/fold yet 😉 but in general Average/ensemble do help and also minimize overfitting."
  },
  "source": "meta"
}