{
  "id": 268124,
  "title": "Pytorch Lightning 1.4.2 does not converge, use 1.3.8",
  "url": "/competitions/landmark-recognition-2021/discussion/268124",
  "author_name": "",
  "post_date": "2021-08-26T05:50:34.766450300Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So I discovered this after feeling like a madman for 2 days</p>\n<ul>\n<li>Kaggle notebooks have the default PyTorch lightning version of 1.3.8. Models do converge</li>\n<li>On my local machine, I had the PyTorch lightning version of 1.4.2. Models <strong>did not</strong> converge.  </li>\n</ul>\n<p>I found this by trying to fit an Efficient Net B0 on 2 images. On Kaggle, it took 20 epochs and loss converged to almost zero. On my local machine, after 50 epoch the loss was still in a very weird place. After I changed the version from 1.4.2 to 1.3.8, everything started to work as expected. This then translated into convergence on the entire dataset</p>",
  "messages": [
    {
      "id": "1491057",
      "postDate": "08/26/2021 05:50:34",
      "content": "<p>So I discovered this after feeling like a madman for 2 days</p>\n<ul>\n<li>Kaggle notebooks have the default PyTorch lightning version of 1.3.8. Models do converge</li>\n<li>On my local machine, I had the PyTorch lightning version of 1.4.2. Models <strong>did not</strong> converge.  </li>\n</ul>\n<p>I found this by trying to fit an Efficient Net B0 on 2 images. On Kaggle, it took 20 epochs and loss converged to almost zero. On my local machine, after 50 epoch the loss was still in a very weird place. After I changed the version from 1.4.2 to 1.3.8, everything started to work as expected. This then translated into convergence on the entire dataset</p>",
      "rawMarkdown": "So I discovered this after feeling like a madman for 2 days\n\n- Kaggle notebooks have the default PyTorch lightning version of 1.3.8. Models do converge\n- On my local machine, I had the PyTorch lightning version of 1.4.2. Models **did not** converge.  \n\nI found this by trying to fit an Efficient Net B0 on 2 images. On Kaggle, it took 20 epochs and loss converged to almost zero. On my local machine, after 50 epoch the loss was still in a very weird place. After I changed the version from 1.4.2 to 1.3.8, everything started to work as expected. This then translated into convergence on the entire dataset",
      "votes": null
    },
    {
      "id": "1496293",
      "postDate": "08/30/2021 09:02:13",
      "content": "<p>Thank u for efforts, thank u for sharing =))</p>",
      "rawMarkdown": "Thank u for efforts, thank u for sharing =))",
      "votes": null
    },
    {
      "id": "1500748",
      "postDate": "09/02/2021 15:40:13",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>. Would you mind to open an issue on PyTorch Lightning and share the notebooks you used, so we can reproduce the non convergence and bisect the commit which broke convergence. </p>",
      "rawMarkdown": "Hey @narsil. Would you mind to open an issue on PyTorch Lightning and share the notebooks you used, so we can reproduce the non convergence and bisect the commit which broke convergence.",
      "votes": null
    },
    {
      "id": "1500881",
      "postDate": "09/02/2021 17:56:21",
      "content": "<p>Hey, I will do that - if you permit - after the competition end. Right now I spend every bit of free time experimenting to improve our scores. But after the competition closes I commit I will open the issue. Bt w is it something other people also report or I am alone?</p>",
      "rawMarkdown": "Hey, I will do that - if you permit - after the competition end. Right now I spend every bit of free time experimenting to improve our scores. But after the competition closes I commit I will open the issue. Bt w is it something other people also report or I am alone?",
      "votes": null
    },
    {
      "id": "1510466",
      "postDate": "09/12/2021 12:51:31",
      "content": "<p>thank you nice hint. I'll try to do that.😄</p>",
      "rawMarkdown": "thank you nice hint. I'll try to do that.😄",
      "votes": null
    },
    {
      "id": "1515105",
      "postDate": "09/16/2021 17:58:15",
      "content": "<p>thanks for sharing.</p>",
      "rawMarkdown": "thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1496293,
      "author_name": "firefliesqn",
      "author_url": "",
      "post_date": "08/30/2021 09:02:13",
      "content": "<p>Thank u for efforts, thank u for sharing =))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1500748,
      "author_name": "thomaschaton",
      "author_url": "",
      "post_date": "09/02/2021 15:40:13",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>. Would you mind to open an issue on PyTorch Lightning and share the notebooks you used, so we can reproduce the non convergence and bisect the commit which broke convergence. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1500881,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "09/02/2021 17:56:21",
          "content": "<p>Hey, I will do that - if you permit - after the competition end. Right now I spend every bit of free time experimenting to improve our scores. But after the competition closes I commit I will open the issue. Bt w is it something other people also report or I am alone?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1510466,
      "author_name": "tensorchoko",
      "author_url": "",
      "post_date": "09/12/2021 12:51:31",
      "content": "<p>thank you nice hint. I'll try to do that.😄</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1515105,
      "author_name": "waddlegassa",
      "author_url": "",
      "post_date": "09/16/2021 17:58:15",
      "content": "<p>thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1491057": "So I discovered this after feeling like a madman for 2 days\n\n- Kaggle notebooks have the default PyTorch lightning version of 1.3.8. Models do converge\n- On my local machine, I had the PyTorch lightning version of 1.4.2. Models **did not** converge.  \n\nI found this by trying to fit an Efficient Net B0 on 2 images. On Kaggle, it took 20 epochs and loss converged to almost zero. On my local machine, after 50 epoch the loss was still in a very weird place. After I changed the version from 1.4.2 to 1.3.8, everything started to work as expected. This then translated into convergence on the entire dataset",
    "1496293": "Thank u for efforts, thank u for sharing =))",
    "1500748": "Hey @narsil. Would you mind to open an issue on PyTorch Lightning and share the notebooks you used, so we can reproduce the non convergence and bisect the commit which broke convergence.",
    "1500881": "Hey, I will do that - if you permit - after the competition end. Right now I spend every bit of free time experimenting to improve our scores. But after the competition closes I commit I will open the issue. Bt w is it something other people also report or I am alone?",
    "1510466": "thank you nice hint. I'll try to do that.😄",
    "1515105": "thanks for sharing."
  },
  "source": "meta"
}