{
  "id": 482900,
  "title": "Trying to solve Timm vs Torchvision mystery (Hardcore)(Partially solved)",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/482900",
  "author_name": "",
  "post_date": "2024-03-10T00:49:33.388060400Z",
  "votes": 21,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n<h1>TL:DR</h1>\n<p>Lot's of code, don't read!</p>\n<h1>Intro</h1>\n<p>Imagine you have two absolutely identical models, the same weights, same required_grad=True modules (nothing is frozen), the same data  and you get completely different training results.</p>\n<h2>Let's look at this code snippet:</h2>\n<pre><code> timm\n torch\n torch  nn\n torchvision.models  efficientnet_b0\n\nWEIGHTS_FILE = \nmodel1 = timm.create_model(, pretrained=, in_chans=, num_classes=)\n\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(),\n     nn.Linear(model1.classifier.in_features, out_features=)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nmodel2.classifier[] = nn.Linear(model2.classifier[].in_features, out_features=)\n</code></pre>\n<p>I have initialized two models, model1 is initialized with timm and model2 - with torchvision.</p>\n<h2>Let's compare state dicts:</h2>\n<pre><code>state_dict_model1 = model1.state_dict()\nstate_dict_model2 = model2.state_dict()\n\n key1, key2  (state_dict_model1.keys(), state_dict_model2.keys()):\n    are_state_dicts_layers_equal = torch.allclose(state_dict_model1[key1], state_dict_model2[key2])\n      are_state_dicts_layers_equal:\n        (key1, key2)\n    ()\n</code></pre>\n<h2>Only two weights are different:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F12b9a6b6e58c6a847d4e1376e8075e83%2FScreenshot%20from%202024-03-09%2019-14-02.png?generation=1710029662047382&amp;alt=media\"></p>\n<h2>Let's fix it since the classifier Linear layer was initialized randomly:</h2>\n<pre><code> pytorch_lightning  seed_everything\n\nWEIGHTS_FILE = \nmodel1 = timm.create_model(, pretrained=, in_chans=, num_classes=)\n\nseed_everything()\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(),\n     nn.Linear(model1.classifier.in_features, out_features=)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nseed_everything()\nmodel2.classifier[] = nn.Linear(model2.classifier[].in_features, out_features=)\n</code></pre>\n<h2>Now weights are the same:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fcc7d5259f37445266210846c1db688c1%2FScreenshot%20from%202024-03-09%2019-17-56.png?generation=1710029896038494&amp;alt=media\"> </p>\n<h2>It seems everything is in order but wait a second, what's that, we got <code>stochastic_depth</code> in the TorchVision model:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F66daffd24cbc1f0ba7832129641c9db2%2FScreenshot%20from%202024-03-09%2019-19-39.png?generation=1710030126476729&amp;alt=media\"></p>\n<p>Though there is <code>drop_path</code> in timm which is <code>nn.Identity</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc328e873723df832eb07458a2efc16ad%2FScreenshot%20from%202024-03-09%2019-19-58.png?generation=1710030197652612&amp;alt=media\"></p>\n<h2>Here is a link which leads to the first major difference, <a href=\"https://github.com/huggingface/pytorch-image-models/blob/2ec2f1aa73e3976553b1ddcb4245a42052a59138/timm/models/efficientnet.py#L1557\" target=\"_blank\">click</a>.</h2>\n<pre><code> () -&gt; EfficientNet:\n    \n    \n</code></pre>\n<p>For those who are hardcore enough, here is a refresher what <code>stochastic_depth</code> is <a href=\"https://arxiv.org/pdf/1603.09382.pdf\" target=\"_blank\">(link to the paper, click)</a>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fb6e23a1b3edef510529b19af2555e223%2FScreenshot%20from%202024-03-09%2019-25-56.png?generation=1710030413014612&amp;alt=media\"></p>\n<h2>Let's fix the timm's <code>drop_path</code> (a.k.a <code>stochastic_depth</code>) as the timm recommends during the training:</h2>\n<pre><code>model1 = timm.create_model(, pretrained=, in_chans=, num_classes=, drop_path_rate=)\n</code></pre>\n<h2>Ok, we fixed it, but timm still rounds that up, (don't worry about <code>mode=row</code>, timm does the same):</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F30a93e2ac52b61157361eb6cbb671313%2FScreenshot%20from%202024-03-09%2019-31-49.png?generation=1710030722421349&amp;alt=media\"></p>\n<h2>And after all dances we still did not fix the issue:</h2>\n<p>The TorchVision model training run:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fd5512b2ac22a11cf3a96375d90cb52d2%2FTorchVisionRun.png?generation=1710032356348560&amp;alt=media\"></p>\n<p>The timm model training run:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe24a5948c15526a29fc1ceb7e3616ae7%2FTimmRun.png?generation=1710032368083592&amp;alt=media\"></p>\n<h2>Outro:</h2>\n<p>Thank you for making that to the end!</p>\n<p>On the screenshots above we can observe training of two identical models sharing the same weights, the same modules and the same data but the results after the three epochs are completely different, not even close.<br>\nI am now more interested in figuring this thing out than in the competition itself, what's going on?</p>\n<p>I was only able to find a relative issue with timm model initialization <a href=\"https://blog.problemsolversguild.com/fastai/technical/exploration/2022/05/02/zero_init_last_resnet18_performance.html\" target=\"_blank\">here</a>. Though the author compared <code>pretrained=False</code> case. Timm is full of surprises…</p>",
  "messages": [
    {
      "id": "2689571",
      "postDate": "03/10/2024 00:49:33",
      "content": "<p>Hi,</p>\n<h1>TL:DR</h1>\n<p>Lot's of code, don't read!</p>\n<h1>Intro</h1>\n<p>Imagine you have two absolutely identical models, the same weights, same required_grad=True modules (nothing is frozen), the same data  and you get completely different training results.</p>\n<h2>Let's look at this code snippet:</h2>\n<pre><code> timm\n torch\n torch  nn\n torchvision.models  efficientnet_b0\n\nWEIGHTS_FILE = \nmodel1 = timm.create_model(, pretrained=, in_chans=, num_classes=)\n\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(),\n     nn.Linear(model1.classifier.in_features, out_features=)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nmodel2.classifier[] = nn.Linear(model2.classifier[].in_features, out_features=)\n</code></pre>\n<p>I have initialized two models, model1 is initialized with timm and model2 - with torchvision.</p>\n<h2>Let's compare state dicts:</h2>\n<pre><code>state_dict_model1 = model1.state_dict()\nstate_dict_model2 = model2.state_dict()\n\n key1, key2  (state_dict_model1.keys(), state_dict_model2.keys()):\n    are_state_dicts_layers_equal = torch.allclose(state_dict_model1[key1], state_dict_model2[key2])\n      are_state_dicts_layers_equal:\n        (key1, key2)\n    ()\n</code></pre>\n<h2>Only two weights are different:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F12b9a6b6e58c6a847d4e1376e8075e83%2FScreenshot%20from%202024-03-09%2019-14-02.png?generation=1710029662047382&amp;alt=media\"></p>\n<h2>Let's fix it since the classifier Linear layer was initialized randomly:</h2>\n<pre><code> pytorch_lightning  seed_everything\n\nWEIGHTS_FILE = \nmodel1 = timm.create_model(, pretrained=, in_chans=, num_classes=)\n\nseed_everything()\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(),\n     nn.Linear(model1.classifier.in_features, out_features=)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nseed_everything()\nmodel2.classifier[] = nn.Linear(model2.classifier[].in_features, out_features=)\n</code></pre>\n<h2>Now weights are the same:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fcc7d5259f37445266210846c1db688c1%2FScreenshot%20from%202024-03-09%2019-17-56.png?generation=1710029896038494&amp;alt=media\"> </p>\n<h2>It seems everything is in order but wait a second, what's that, we got <code>stochastic_depth</code> in the TorchVision model:</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F66daffd24cbc1f0ba7832129641c9db2%2FScreenshot%20from%202024-03-09%2019-19-39.png?generation=1710030126476729&amp;alt=media\"></p>\n<p>Though there is <code>drop_path</code> in timm which is <code>nn.Identity</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc328e873723df832eb07458a2efc16ad%2FScreenshot%20from%202024-03-09%2019-19-58.png?generation=1710030197652612&amp;alt=media\"></p>\n<h2>Here is a link which leads to the first major difference, <a href=\"https://github.com/huggingface/pytorch-image-models/blob/2ec2f1aa73e3976553b1ddcb4245a42052a59138/timm/models/efficientnet.py#L1557\" target=\"_blank\">click</a>.</h2>\n<pre><code> () -&gt; EfficientNet:\n    \n    \n</code></pre>\n<p>For those who are hardcore enough, here is a refresher what <code>stochastic_depth</code> is <a href=\"https://arxiv.org/pdf/1603.09382.pdf\" target=\"_blank\">(link to the paper, click)</a>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fb6e23a1b3edef510529b19af2555e223%2FScreenshot%20from%202024-03-09%2019-25-56.png?generation=1710030413014612&amp;alt=media\"></p>\n<h2>Let's fix the timm's <code>drop_path</code> (a.k.a <code>stochastic_depth</code>) as the timm recommends during the training:</h2>\n<pre><code>model1 = timm.create_model(, pretrained=, in_chans=, num_classes=, drop_path_rate=)\n</code></pre>\n<h2>Ok, we fixed it, but timm still rounds that up, (don't worry about <code>mode=row</code>, timm does the same):</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F30a93e2ac52b61157361eb6cbb671313%2FScreenshot%20from%202024-03-09%2019-31-49.png?generation=1710030722421349&amp;alt=media\"></p>\n<h2>And after all dances we still did not fix the issue:</h2>\n<p>The TorchVision model training run:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fd5512b2ac22a11cf3a96375d90cb52d2%2FTorchVisionRun.png?generation=1710032356348560&amp;alt=media\"></p>\n<p>The timm model training run:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe24a5948c15526a29fc1ceb7e3616ae7%2FTimmRun.png?generation=1710032368083592&amp;alt=media\"></p>\n<h2>Outro:</h2>\n<p>Thank you for making that to the end!</p>\n<p>On the screenshots above we can observe training of two identical models sharing the same weights, the same modules and the same data but the results after the three epochs are completely different, not even close.<br>\nI am now more interested in figuring this thing out than in the competition itself, what's going on?</p>\n<p>I was only able to find a relative issue with timm model initialization <a href=\"https://blog.problemsolversguild.com/fastai/technical/exploration/2022/05/02/zero_init_last_resnet18_performance.html\" target=\"_blank\">here</a>. Though the author compared <code>pretrained=False</code> case. Timm is full of surprises…</p>",
      "rawMarkdown": "Hi,\n\n#TL:DR \nLot's of code, don't read!\n\n#Intro\nImagine you have two absolutely identical models, the same weights, same required_grad=True modules (nothing is frozen), the same data  and you get completely different training results.\n\n## Let's look at this code snippet:\n\n```python\nimport timm\nimport torch\nfrom torch import nn\nfrom torchvision.models import efficientnet_b0\n\nWEIGHTS_FILE = 'saved_models/efficientnet_b0_rwightman-7f5810bc.pth'\nmodel1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6)\n\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(0.2),\n     nn.Linear(model1.classifier.in_features, out_features=6)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nmodel2.classifier[1] = nn.Linear(model2.classifier[1].in_features, out_features=6)\n```\n\nI have initialized two models, model1 is initialized with timm and model2 - with torchvision.\n\n## Let's compare state dicts:\n\n```python\nstate_dict_model1 = model1.state_dict()\nstate_dict_model2 = model2.state_dict()\n\nfor key1, key2 in zip(state_dict_model1.keys(), state_dict_model2.keys()):\n    are_state_dicts_layers_equal = torch.allclose(state_dict_model1[key1], state_dict_model2[key2])\n    if not are_state_dicts_layers_equal:\n        print(key1, key2)\n    print(f\"State Dicts are {'Equal' if are_state_dicts_layers_equal else 'Not Equal'}\")\n```\n\n## Only two weights are different:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F12b9a6b6e58c6a847d4e1376e8075e83%2FScreenshot%20from%202024-03-09%2019-14-02.png?generation=1710029662047382&alt=media)\n\n## Let's fix it since the classifier Linear layer was initialized randomly:\n\n```python\nfrom pytorch_lightning import seed_everything\n\nWEIGHTS_FILE = 'saved_models/efficientnet_b0_rwightman-7f5810bc.pth'\nmodel1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6)\n\nseed_everything(360)\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(0.2),\n     nn.Linear(model1.classifier.in_features, out_features=6)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nseed_everything(360)\nmodel2.classifier[1] = nn.Linear(model2.classifier[1].in_features, out_features=6)\n```\n\n## Now weights are the same:\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fcc7d5259f37445266210846c1db688c1%2FScreenshot%20from%202024-03-09%2019-17-56.png?generation=1710029896038494&alt=media) \n\n## It seems everything is in order but wait a second, what's that, we got `stochastic_depth` in the TorchVision model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F66daffd24cbc1f0ba7832129641c9db2%2FScreenshot%20from%202024-03-09%2019-19-39.png?generation=1710030126476729&alt=media)\n\nThough there is `drop_path` in timm which is `nn.Identity`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc328e873723df832eb07458a2efc16ad%2FScreenshot%20from%202024-03-09%2019-19-58.png?generation=1710030197652612&alt=media)\n\n## Here is a link which leads to the first major difference, [click](https://github.com/huggingface/pytorch-image-models/blob/2ec2f1aa73e3976553b1ddcb4245a42052a59138/timm/models/efficientnet.py#L1557).\n\n```python\ndef efficientnet_b0(pretrained=False, **kwargs) -> EfficientNet:\n    \"\"\" EfficientNet-B0 \"\"\"\n    # NOTE for train, drop_rate should be 0.2, drop_path_rate should be 0.2\n```\n\nFor those who are hardcore enough, here is a refresher what `stochastic_depth` is [(link to the paper, click)](https://arxiv.org/pdf/1603.09382.pdf):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fb6e23a1b3edef510529b19af2555e223%2FScreenshot%20from%202024-03-09%2019-25-56.png?generation=1710030413014612&alt=media)\n\n## Let's fix the timm's `drop_path` (a.k.a `stochastic_depth`) as the timm recommends during the training:\n```python\nmodel1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6, drop_path_rate=0.2)\n\n```\n\n## Ok, we fixed it, but timm still rounds that up, (don't worry about `mode=row`, timm does the same):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F30a93e2ac52b61157361eb6cbb671313%2FScreenshot%20from%202024-03-09%2019-31-49.png?generation=1710030722421349&alt=media)\n\n## And after all dances we still did not fix the issue:\n\nThe TorchVision model training run:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fd5512b2ac22a11cf3a96375d90cb52d2%2FTorchVisionRun.png?generation=1710032356348560&alt=media)\n\nThe timm model training run:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe24a5948c15526a29fc1ceb7e3616ae7%2FTimmRun.png?generation=1710032368083592&alt=media)\n\n## Outro:\nThank you for making that to the end!\n\nOn the screenshots above we can observe training of two identical models sharing the same weights, the same modules and the same data but the results after the three epochs are completely different, not even close.\nI am now more interested in figuring this thing out than in the competition itself, what's going on?\n\nI was only able to find a relative issue with timm model initialization [here](https://blog.problemsolversguild.com/fastai/technical/exploration/2022/05/02/zero_init_last_resnet18_performance.html). Though the author compared `pretrained=False` case. Timm is full of surprises...",
      "votes": null
    },
    {
      "id": "2689592",
      "postDate": "03/10/2024 01:48:44",
      "content": "<p>There are many random processes during training such as random batches, random dropout, random augmentations, etc. Additionally GPUs process in parallel and sometimes the order that processes finish affect the result. Have you removed all this randomness?</p>",
      "rawMarkdown": "There are many random processes during training such as random batches, random dropout, random augmentations, etc. Additionally GPUs process in parallel and sometimes the order that processes finish affect the result. Have you removed all this randomness?",
      "votes": null
    },
    {
      "id": "2689644",
      "postDate": "03/10/2024 03:17:55",
      "content": "<p>Hi Chris, by the time you wrote your comment, I figured out to get the same results for random tensor by removing <code>stochastic_depth</code> from TorchVision model, on top of that I removed the only Dropout in the classifier for TorchVision. </p>\n<p>Though having started the training with new initialized TorchVision model, the results still held the same.<br>\nThen, I switched to shuffle=False in the loader -&gt; Still the same. <br>\nThen, I removed <code>ddp</code> and trained on single GPU and \"voila\" magically Timm started showing exactly the same good results as the TorchVision Model. Now I narrowed down the issue to the <code>ddp</code> processes. Now I am curious what is going on there, lol.</p>\n<pre><code>In []:  timm\n     ...:  torch\n     ...:  torch  nn\n     ...:  torchvision.models  efficientnet_b0\n     ...:  pytorch_lightning  seed_everything\n     ...: \n     ...: WEIGHTS_FILE = \n     ...: model1 = timm.create_model(, pretrained=, in_chans=, num_classes=)\n     ...: \n     ...: seed_everything()\n     ...: model1.classifier = nn.Sequential(\n     ...:      nn.Linear(model1.classifier.in_features, out_features=)\n     ...:  )\n     ...: \n     ...: model2 = efficientnet_b0(stochastic_depth_prob=)\n     ...: model2.load_state_dict(torch.load(WEIGHTS_FILE))\n     ...: \n     ...: seed_everything()\n     ...: model2.classifier[] = nn.Identity()\n     ...: model2.classifier[] = nn.Linear(model2.classifier[].in_features, out_features=)\n     ...: \n     ...: x = torch.randn(, , , )\n     ...: torch.allclose(model1(x), model2(x))\nSeed  to \nSeed  to \nOut[]: \n</code></pre>",
      "rawMarkdown": "Hi Chris, by the time you wrote your comment, I figured out to get the same results for random tensor by removing `stochastic_depth` from TorchVision model, on top of that I removed the only Dropout in the classifier for TorchVision. \n\nThough having started the training with new initialized TorchVision model, the results still held the same.\nThen, I switched to shuffle=False in the loader -> Still the same. \nThen, I removed `ddp` and trained on single GPU and \"voila\" magically Timm started showing exactly the same good results as the TorchVision Model. Now I narrowed down the issue to the `ddp` processes. Now I am curious what is going on there, lol.\n\n```python\nIn [426]: import timm\n     ...: import torch\n     ...: from torch import nn\n     ...: from torchvision.models import efficientnet_b0\n     ...: from pytorch_lightning import seed_everything\n     ...: \n     ...: WEIGHTS_FILE = 'saved_models/efficientnet_b0_rwightman-7f5810bc.pth'\n     ...: model1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6)\n     ...: \n     ...: seed_everything(360)\n     ...: model1.classifier = nn.Sequential(\n     ...:      nn.Linear(model1.classifier.in_features, out_features=6)\n     ...:  )\n     ...: \n     ...: model2 = efficientnet_b0(stochastic_depth_prob=0.0)\n     ...: model2.load_state_dict(torch.load(WEIGHTS_FILE))\n     ...: \n     ...: seed_everything(360)\n     ...: model2.classifier[0] = nn.Identity()\n     ...: model2.classifier[1] = nn.Linear(model2.classifier[1].in_features, out_features=6)\n     ...: \n     ...: x = torch.randn(32, 3, 256, 256)\n     ...: torch.allclose(model1(x), model2(x))\nSeed set to 360\nSeed set to 360\nOut[426]: True\n```",
      "votes": null
    },
    {
      "id": "2694310",
      "postDate": "03/13/2024 01:45:32",
      "content": "<p>Thanks for sharing this idea🙇</p>",
      "rawMarkdown": "Thanks for sharing this idea🙇",
      "votes": null
    },
    {
      "id": "2704002",
      "postDate": "03/18/2024 14:10:48",
      "content": "<p>Same issue on my side. Using Timm + DDP made my score worse … have no idea though</p>",
      "rawMarkdown": "Same issue on my side. Using Timm + DDP made my score worse ... have no idea though",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2689592,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/10/2024 01:48:44",
      "content": "<p>There are many random processes during training such as random batches, random dropout, random augmentations, etc. Additionally GPUs process in parallel and sometimes the order that processes finish affect the result. Have you removed all this randomness?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2689644,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "03/10/2024 03:17:55",
          "content": "<p>Hi Chris, by the time you wrote your comment, I figured out to get the same results for random tensor by removing <code>stochastic_depth</code> from TorchVision model, on top of that I removed the only Dropout in the classifier for TorchVision. </p>\n<p>Though having started the training with new initialized TorchVision model, the results still held the same.<br>\nThen, I switched to shuffle=False in the loader -&gt; Still the same. <br>\nThen, I removed <code>ddp</code> and trained on single GPU and \"voila\" magically Timm started showing exactly the same good results as the TorchVision Model. Now I narrowed down the issue to the <code>ddp</code> processes. Now I am curious what is going on there, lol.</p>\n<pre><code>In []:  timm\n     ...:  torch\n     ...:  torch  nn\n     ...:  torchvision.models  efficientnet_b0\n     ...:  pytorch_lightning  seed_everything\n     ...: \n     ...: WEIGHTS_FILE = \n     ...: model1 = timm.create_model(, pretrained=, in_chans=, num_classes=)\n     ...: \n     ...: seed_everything()\n     ...: model1.classifier = nn.Sequential(\n     ...:      nn.Linear(model1.classifier.in_features, out_features=)\n     ...:  )\n     ...: \n     ...: model2 = efficientnet_b0(stochastic_depth_prob=)\n     ...: model2.load_state_dict(torch.load(WEIGHTS_FILE))\n     ...: \n     ...: seed_everything()\n     ...: model2.classifier[] = nn.Identity()\n     ...: model2.classifier[] = nn.Linear(model2.classifier[].in_features, out_features=)\n     ...: \n     ...: x = torch.randn(, , , )\n     ...: torch.allclose(model1(x), model2(x))\nSeed  to \nSeed  to \nOut[]: \n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2704002,
              "author_name": "fuckvenkatraman",
              "author_url": "",
              "post_date": "03/18/2024 14:10:48",
              "content": "<p>Same issue on my side. Using Timm + DDP made my score worse … have no idea though</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2694310,
      "author_name": "complexai",
      "author_url": "",
      "post_date": "03/13/2024 01:45:32",
      "content": "<p>Thanks for sharing this idea🙇</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2689571": "Hi,\n\n#TL:DR \nLot's of code, don't read!\n\n#Intro\nImagine you have two absolutely identical models, the same weights, same required_grad=True modules (nothing is frozen), the same data  and you get completely different training results.\n\n## Let's look at this code snippet:\n\n```python\nimport timm\nimport torch\nfrom torch import nn\nfrom torchvision.models import efficientnet_b0\n\nWEIGHTS_FILE = 'saved_models/efficientnet_b0_rwightman-7f5810bc.pth'\nmodel1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6)\n\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(0.2),\n     nn.Linear(model1.classifier.in_features, out_features=6)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nmodel2.classifier[1] = nn.Linear(model2.classifier[1].in_features, out_features=6)\n```\n\nI have initialized two models, model1 is initialized with timm and model2 - with torchvision.\n\n## Let's compare state dicts:\n\n```python\nstate_dict_model1 = model1.state_dict()\nstate_dict_model2 = model2.state_dict()\n\nfor key1, key2 in zip(state_dict_model1.keys(), state_dict_model2.keys()):\n    are_state_dicts_layers_equal = torch.allclose(state_dict_model1[key1], state_dict_model2[key2])\n    if not are_state_dicts_layers_equal:\n        print(key1, key2)\n    print(f\"State Dicts are {'Equal' if are_state_dicts_layers_equal else 'Not Equal'}\")\n```\n\n## Only two weights are different:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F12b9a6b6e58c6a847d4e1376e8075e83%2FScreenshot%20from%202024-03-09%2019-14-02.png?generation=1710029662047382&alt=media)\n\n## Let's fix it since the classifier Linear layer was initialized randomly:\n\n```python\nfrom pytorch_lightning import seed_everything\n\nWEIGHTS_FILE = 'saved_models/efficientnet_b0_rwightman-7f5810bc.pth'\nmodel1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6)\n\nseed_everything(360)\nmodel1.classifier = nn.Sequential(\n     nn.Dropout(0.2),\n     nn.Linear(model1.classifier.in_features, out_features=6)\n )\n\nmodel2 = efficientnet_b0()\nmodel2.load_state_dict(torch.load(WEIGHTS_FILE))\n\nseed_everything(360)\nmodel2.classifier[1] = nn.Linear(model2.classifier[1].in_features, out_features=6)\n```\n\n## Now weights are the same:\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fcc7d5259f37445266210846c1db688c1%2FScreenshot%20from%202024-03-09%2019-17-56.png?generation=1710029896038494&alt=media) \n\n## It seems everything is in order but wait a second, what's that, we got `stochastic_depth` in the TorchVision model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F66daffd24cbc1f0ba7832129641c9db2%2FScreenshot%20from%202024-03-09%2019-19-39.png?generation=1710030126476729&alt=media)\n\nThough there is `drop_path` in timm which is `nn.Identity`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc328e873723df832eb07458a2efc16ad%2FScreenshot%20from%202024-03-09%2019-19-58.png?generation=1710030197652612&alt=media)\n\n## Here is a link which leads to the first major difference, [click](https://github.com/huggingface/pytorch-image-models/blob/2ec2f1aa73e3976553b1ddcb4245a42052a59138/timm/models/efficientnet.py#L1557).\n\n```python\ndef efficientnet_b0(pretrained=False, **kwargs) -> EfficientNet:\n    \"\"\" EfficientNet-B0 \"\"\"\n    # NOTE for train, drop_rate should be 0.2, drop_path_rate should be 0.2\n```\n\nFor those who are hardcore enough, here is a refresher what `stochastic_depth` is [(link to the paper, click)](https://arxiv.org/pdf/1603.09382.pdf):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fb6e23a1b3edef510529b19af2555e223%2FScreenshot%20from%202024-03-09%2019-25-56.png?generation=1710030413014612&alt=media)\n\n## Let's fix the timm's `drop_path` (a.k.a `stochastic_depth`) as the timm recommends during the training:\n```python\nmodel1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6, drop_path_rate=0.2)\n\n```\n\n## Ok, we fixed it, but timm still rounds that up, (don't worry about `mode=row`, timm does the same):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F30a93e2ac52b61157361eb6cbb671313%2FScreenshot%20from%202024-03-09%2019-31-49.png?generation=1710030722421349&alt=media)\n\n## And after all dances we still did not fix the issue:\n\nThe TorchVision model training run:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fd5512b2ac22a11cf3a96375d90cb52d2%2FTorchVisionRun.png?generation=1710032356348560&alt=media)\n\nThe timm model training run:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe24a5948c15526a29fc1ceb7e3616ae7%2FTimmRun.png?generation=1710032368083592&alt=media)\n\n## Outro:\nThank you for making that to the end!\n\nOn the screenshots above we can observe training of two identical models sharing the same weights, the same modules and the same data but the results after the three epochs are completely different, not even close.\nI am now more interested in figuring this thing out than in the competition itself, what's going on?\n\nI was only able to find a relative issue with timm model initialization [here](https://blog.problemsolversguild.com/fastai/technical/exploration/2022/05/02/zero_init_last_resnet18_performance.html). Though the author compared `pretrained=False` case. Timm is full of surprises...",
    "2689592": "There are many random processes during training such as random batches, random dropout, random augmentations, etc. Additionally GPUs process in parallel and sometimes the order that processes finish affect the result. Have you removed all this randomness?",
    "2689644": "Hi Chris, by the time you wrote your comment, I figured out to get the same results for random tensor by removing `stochastic_depth` from TorchVision model, on top of that I removed the only Dropout in the classifier for TorchVision. \n\nThough having started the training with new initialized TorchVision model, the results still held the same.\nThen, I switched to shuffle=False in the loader -> Still the same. \nThen, I removed `ddp` and trained on single GPU and \"voila\" magically Timm started showing exactly the same good results as the TorchVision Model. Now I narrowed down the issue to the `ddp` processes. Now I am curious what is going on there, lol.\n\n```python\nIn [426]: import timm\n     ...: import torch\n     ...: from torch import nn\n     ...: from torchvision.models import efficientnet_b0\n     ...: from pytorch_lightning import seed_everything\n     ...: \n     ...: WEIGHTS_FILE = 'saved_models/efficientnet_b0_rwightman-7f5810bc.pth'\n     ...: model1 = timm.create_model('timm/efficientnet_b0', pretrained=True, in_chans=3, num_classes=6)\n     ...: \n     ...: seed_everything(360)\n     ...: model1.classifier = nn.Sequential(\n     ...:      nn.Linear(model1.classifier.in_features, out_features=6)\n     ...:  )\n     ...: \n     ...: model2 = efficientnet_b0(stochastic_depth_prob=0.0)\n     ...: model2.load_state_dict(torch.load(WEIGHTS_FILE))\n     ...: \n     ...: seed_everything(360)\n     ...: model2.classifier[0] = nn.Identity()\n     ...: model2.classifier[1] = nn.Linear(model2.classifier[1].in_features, out_features=6)\n     ...: \n     ...: x = torch.randn(32, 3, 256, 256)\n     ...: torch.allclose(model1(x), model2(x))\nSeed set to 360\nSeed set to 360\nOut[426]: True\n```",
    "2694310": "Thanks for sharing this idea🙇",
    "2704002": "Same issue on my side. Using Timm + DDP made my score worse ... have no idea though"
  },
  "source": "meta"
}