{
  "id": 134427,
  "title": "How to use multi-gpu?",
  "url": "/competitions/deepfake-detection-challenge/discussion/134427",
  "author_name": "",
  "post_date": "2020-03-08T03:38:37.256754200Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>We were trying to train with multi-gpu by the following code:\n<code>\nstrategy=tf.distribute.MirroredStrategy()\nwith strategy.scope():\n    ....#define model/compile\nmodel.fit(X,y)\nmodel.save_weights('model.h5')\n</code>\nbut after loading the saved weights, the results are different from the original model. We are aware of this is because of the weights are saved on different gpus(or something similar like that). Can anyone please help us? Thank you in advance.</p>",
  "messages": [
    {
      "id": "766356",
      "postDate": "03/08/2020 03:38:37",
      "content": "<p>We were trying to train with multi-gpu by the following code:\n<code>\nstrategy=tf.distribute.MirroredStrategy()\nwith strategy.scope():\n    ....#define model/compile\nmodel.fit(X,y)\nmodel.save_weights('model.h5')\n</code>\nbut after loading the saved weights, the results are different from the original model. We are aware of this is because of the weights are saved on different gpus(or something similar like that). Can anyone please help us? Thank you in advance.</p>",
      "rawMarkdown": "We were trying to train with multi-gpu by the following code:\n```\nstrategy=tf.distribute.MirroredStrategy()\nwith strategy.scope():\n    ....#define model/compile\nmodel.fit(X,y)\nmodel.save_weights('model.h5')\n```\nbut after loading the saved weights, the results are different from the original model. We are aware of this is because of the weights are saved on different gpus(or something similar like that). Can anyone please help us? Thank you in advance.",
      "votes": null
    },
    {
      "id": "766443",
      "postDate": "03/08/2020 06:19:57",
      "content": "<p>While using MirroredStrategy(), use callbacks as below.</p>\n\n<p>model_checkpoint = ModelCheckpoint(filepath, monitor = 'val_loss', save_best_only = False)\ncallback_list = [model_checkpoint]\nmodel.fit(......., callbacks = callback_list,....)</p>\n\n<p>This method works fine for me.</p>\n\n<p>However, you can't load this model back and start training with multi gpu again. That is currently broken in tf 2.1</p>",
      "rawMarkdown": "While using MirroredStrategy(), use callbacks as below.\n\nmodel_checkpoint = ModelCheckpoint(filepath, monitor = 'val_loss', save_best_only = False)\ncallback_list = [model_checkpoint]\nmodel.fit(......., callbacks = callback_list,....)\n\nThis method works fine for me.\n\nHowever, you can't load this model back and start training with multi gpu again. That is currently broken in tf 2.1",
      "votes": null
    },
    {
      "id": "766454",
      "postDate": "03/08/2020 06:48:39",
      "content": "<p>We are able to save our models, but in any case, the weights get changed when the same model is loaded on single gpu and makes the model preforms worse and that is no everyday worse performance, makes ~0.05 lower on leaderboard.</p>",
      "rawMarkdown": "We are able to save our models, but in any case, the weights get changed when the same model is loaded on single gpu and makes the model preforms worse and that is no everyday worse performance, makes ~0.05 lower on leaderboard.",
      "votes": null
    },
    {
      "id": "766743",
      "postDate": "03/08/2020 16:21:14",
      "content": "<p>So what <a href=\"/harshitsheoran\">@harshitsheoran</a> means is, the prediction on multi-gpu is different than prediction of weights loaded on single-gpu.</p>",
      "rawMarkdown": "So what @harshitsheoran means is, the prediction on multi-gpu is different than prediction of weights loaded on single-gpu.",
      "votes": null
    },
    {
      "id": "766767",
      "postDate": "03/08/2020 17:03:59",
      "content": "<p>I just verified this on my work system, though on a different work related project. The outputs for the model remains the same when loaded into a single gpu, irrespective of whether it was trained in multi gpu or single gpu. Verified with 4x2080Ti.</p>",
      "rawMarkdown": "I just verified this on my work system, though on a different work related project. The outputs for the model remains the same when loaded into a single gpu, irrespective of whether it was trained in multi gpu or single gpu. Verified with 4x2080Ti.",
      "votes": null
    },
    {
      "id": "766778",
      "postDate": "03/08/2020 17:25:51",
      "content": "<p>Did you use any tricks like <code>tf.device</code>? Simple keras.models.load_model doesn't work and model.load_weights have different results.</p>",
      "rawMarkdown": "Did you use any tricks like `tf.device`? Simple keras.models.load_model doesn't work and model.load_weights have different results.",
      "votes": null
    },
    {
      "id": "767018",
      "postDate": "03/09/2020 03:28:31",
      "content": "<p>I don't save only the weights. I save the entire model as a whole using ModelCheckpoint in callbacks. After training is done, I simply use tf.keras.models.load_model()</p>",
      "rawMarkdown": "I don't save only the weights. I save the entire model as a whole using ModelCheckpoint in callbacks. After training is done, I simply use tf.keras.models.load_model()",
      "votes": null
    },
    {
      "id": "767027",
      "postDate": "03/09/2020 03:58:54",
      "content": "<p>I used <code>model.save('model.h5')</code> then use <code>tf.keras.models.load_model('model.h5)</code>, but it gave a weird error saying <strong>init</strong> doesn't have <strong>*</strong>. I'll try ModelCheckpoint. Thank you for your solution for us!</p>",
      "rawMarkdown": "I used `model.save('model.h5')` then use `tf.keras.models.load_model('model.h5)`, but it gave a weird error saying __init__ doesn't have *****. I'll try ModelCheckpoint. Thank you for your solution for us!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 766443,
      "author_name": "akashnandi",
      "author_url": "",
      "post_date": "03/08/2020 06:19:57",
      "content": "<p>While using MirroredStrategy(), use callbacks as below.</p>\n\n<p>model_checkpoint = ModelCheckpoint(filepath, monitor = 'val_loss', save_best_only = False)\ncallback_list = [model_checkpoint]\nmodel.fit(......., callbacks = callback_list,....)</p>\n\n<p>This method works fine for me.</p>\n\n<p>However, you can't load this model back and start training with multi gpu again. That is currently broken in tf 2.1</p>",
      "votes": null,
      "replies": [
        {
          "id": 766454,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/08/2020 06:48:39",
          "content": "<p>We are able to save our models, but in any case, the weights get changed when the same model is loaded on single gpu and makes the model preforms worse and that is no everyday worse performance, makes ~0.05 lower on leaderboard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766743,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/08/2020 16:21:14",
          "content": "<p>So what <a href=\"/harshitsheoran\">@harshitsheoran</a> means is, the prediction on multi-gpu is different than prediction of weights loaded on single-gpu.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766767,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "03/08/2020 17:03:59",
          "content": "<p>I just verified this on my work system, though on a different work related project. The outputs for the model remains the same when loaded into a single gpu, irrespective of whether it was trained in multi gpu or single gpu. Verified with 4x2080Ti.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766778,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/08/2020 17:25:51",
          "content": "<p>Did you use any tricks like <code>tf.device</code>? Simple keras.models.load_model doesn't work and model.load_weights have different results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767018,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "03/09/2020 03:28:31",
          "content": "<p>I don't save only the weights. I save the entire model as a whole using ModelCheckpoint in callbacks. After training is done, I simply use tf.keras.models.load_model()</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767027,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/09/2020 03:58:54",
          "content": "<p>I used <code>model.save('model.h5')</code> then use <code>tf.keras.models.load_model('model.h5)</code>, but it gave a weird error saying <strong>init</strong> doesn't have <strong>*</strong>. I'll try ModelCheckpoint. Thank you for your solution for us!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "766356": "We were trying to train with multi-gpu by the following code:\n```\nstrategy=tf.distribute.MirroredStrategy()\nwith strategy.scope():\n    ....#define model/compile\nmodel.fit(X,y)\nmodel.save_weights('model.h5')\n```\nbut after loading the saved weights, the results are different from the original model. We are aware of this is because of the weights are saved on different gpus(or something similar like that). Can anyone please help us? Thank you in advance.",
    "766443": "While using MirroredStrategy(), use callbacks as below.\n\nmodel_checkpoint = ModelCheckpoint(filepath, monitor = 'val_loss', save_best_only = False)\ncallback_list = [model_checkpoint]\nmodel.fit(......., callbacks = callback_list,....)\n\nThis method works fine for me.\n\nHowever, you can't load this model back and start training with multi gpu again. That is currently broken in tf 2.1",
    "766454": "We are able to save our models, but in any case, the weights get changed when the same model is loaded on single gpu and makes the model preforms worse and that is no everyday worse performance, makes ~0.05 lower on leaderboard.",
    "766743": "So what @harshitsheoran means is, the prediction on multi-gpu is different than prediction of weights loaded on single-gpu.",
    "766767": "I just verified this on my work system, though on a different work related project. The outputs for the model remains the same when loaded into a single gpu, irrespective of whether it was trained in multi gpu or single gpu. Verified with 4x2080Ti.",
    "766778": "Did you use any tricks like `tf.device`? Simple keras.models.load_model doesn't work and model.load_weights have different results.",
    "767018": "I don't save only the weights. I save the entire model as a whole using ModelCheckpoint in callbacks. After training is done, I simply use tf.keras.models.load_model()",
    "767027": "I used `model.save('model.h5')` then use `tf.keras.models.load_model('model.h5)`, but it gave a weird error saying __init__ doesn't have *****. I'll try ModelCheckpoint. Thank you for your solution for us!"
  },
  "source": "meta"
}