{
  "id": 103503,
  "title": "Unable to reproduce result on GCP",
  "url": "/competitions/recursion-cellular-image-classification/discussion/103503",
  "author_name": "",
  "post_date": "2019-08-09T16:47:48.177900100Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys, I have been working on Kaggle kernels until we received the GCP credits. Ever since, I have been trying to move my work to GCP. Yesterday I converted my <code>notebook.ipynb</code> to <code>notebook.py</code> to run on 4 GPUs on GCP. </p>\n\n<p>The only code I added was \n```\nbatch_size = 144 # instead of 32</p>\n\n<p>if torch.cuda.is_available() and torch.cuda.device_count() &gt; 1: \n    model = torch.nn.DataParallel(model)\nmodel.to(device)\n```\nHowever, the validation result I have been getting on GCP is significantly lower than running on Kaggle kernels. </p>\n\n<p>Does anyone have any idea why this is the case? Thank you! </p>",
  "messages": [
    {
      "id": "595775",
      "postDate": "08/09/2019 16:47:48",
      "content": "<p>Hi guys, I have been working on Kaggle kernels until we received the GCP credits. Ever since, I have been trying to move my work to GCP. Yesterday I converted my <code>notebook.ipynb</code> to <code>notebook.py</code> to run on 4 GPUs on GCP. </p>\n\n<p>The only code I added was \n```\nbatch_size = 144 # instead of 32</p>\n\n<p>if torch.cuda.is_available() and torch.cuda.device_count() &gt; 1: \n    model = torch.nn.DataParallel(model)\nmodel.to(device)\n```\nHowever, the validation result I have been getting on GCP is significantly lower than running on Kaggle kernels. </p>\n\n<p>Does anyone have any idea why this is the case? Thank you! </p>",
      "rawMarkdown": "Hi guys, I have been working on Kaggle kernels until we received the GCP credits. Ever since, I have been trying to move my work to GCP. Yesterday I converted my `notebook.ipynb` to `notebook.py` to run on 4 GPUs on GCP. \n\nThe only code I added was \n```\nbatch_size = 144 # instead of 32\n\nif torch.cuda.is_available() and torch.cuda.device_count() &gt; 1: \n    model = torch.nn.DataParallel(model)\nmodel.to(device)\n```\nHowever, the validation result I have been getting on GCP is significantly lower than running on Kaggle kernels. \n\nDoes anyone have any idea why this is the case? Thank you!",
      "votes": null
    },
    {
      "id": "595822",
      "postDate": "08/09/2019 18:24:45",
      "content": "<p>Just increasing the batch size can lead to worse results for the same number of epochs. Try increasing the learning rate to compensate for the bigger batch size (yes, they work in kind of different directions).</p>",
      "rawMarkdown": "Just increasing the batch size can lead to worse results for the same number of epochs. Try increasing the learning rate to compensate for the bigger batch size (yes, they work in kind of different directions).",
      "votes": null
    },
    {
      "id": "595925",
      "postDate": "08/09/2019 21:13:40",
      "content": "<p>thanks! i will try working on the learning rate. and thank you for emphasizing they work in different directions - i was going to asking if you did not mention it. Pretty counter-intuitive!</p>",
      "rawMarkdown": "thanks! i will try working on the learning rate. and thank you for emphasizing they work in different directions - i was going to asking if you did not mention it. Pretty counter-intuitive!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 595822,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/09/2019 18:24:45",
      "content": "<p>Just increasing the batch size can lead to worse results for the same number of epochs. Try increasing the learning rate to compensate for the bigger batch size (yes, they work in kind of different directions).</p>",
      "votes": null,
      "replies": [
        {
          "id": 595925,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "08/09/2019 21:13:40",
          "content": "<p>thanks! i will try working on the learning rate. and thank you for emphasizing they work in different directions - i was going to asking if you did not mention it. Pretty counter-intuitive!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "595775": "Hi guys, I have been working on Kaggle kernels until we received the GCP credits. Ever since, I have been trying to move my work to GCP. Yesterday I converted my `notebook.ipynb` to `notebook.py` to run on 4 GPUs on GCP. \n\nThe only code I added was \n```\nbatch_size = 144 # instead of 32\n\nif torch.cuda.is_available() and torch.cuda.device_count() &gt; 1: \n    model = torch.nn.DataParallel(model)\nmodel.to(device)\n```\nHowever, the validation result I have been getting on GCP is significantly lower than running on Kaggle kernels. \n\nDoes anyone have any idea why this is the case? Thank you!",
    "595822": "Just increasing the batch size can lead to worse results for the same number of epochs. Try increasing the learning rate to compensate for the bigger batch size (yes, they work in kind of different directions).",
    "595925": "thanks! i will try working on the learning rate. and thank you for emphasizing they work in different directions - i was going to asking if you did not mention it. Pretty counter-intuitive!"
  },
  "source": "meta"
}