{
  "id": 268022,
  "title": "Pytorch \"model.load_state_dict\"  cause weird \"Notebook Threw Exception\".",
  "url": "/competitions/landmark-recognition-2021/discussion/268022",
  "author_name": "Lilin Chen",
  "post_date": "2021-08-25T16:25:06.044000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, guys.<br>\nI encountered a very weird submission error even when I successfully running on the notebook. <br>\n<img src=\"https://i.imgur.com/L46aM7O.png\" alt=\"\"></p>\n<hr>\n<p>After debugging by sending the sampe_submission file, I found the key function resulting in the submission error. <a href=\"https://www.kaggle.com/lilinchen/lanmark-inference-test\" target=\"_blank\">my notebook</a></p>\n<pre><code>model.load_state_dict(model_state_dict)\n</code></pre>\n<p><strong># model_state_dict = torch.load(os.path.join(\"../input/model-vgg19\",'model_final.pth'))  is ok I already checked.</strong></p>\n<hr>\n<h3>Here are a few points about the model.</h3>\n<ul>\n<li>I trained from my local device and uploaded it to the notebook.</li>\n<li>Pytorch version is \"'1.9.0+cu111'\" in my device, and it's \"1.7.0\" in the notebook.</li>\n<li>I have tried both \"model.cpu()\" and \"model.cuda()\". Each of them will be failed.</li>\n</ul>\n<h3>Weird things</h3>\n<ul>\n<li>I didn't encounter any error when I was running the code in the notebook. </li>\n<li>Submission is ok even my prediction and \"submission.csv\" had been created on ../output workspace. But it shows \"Notebook Threw Exception\" and \"Public Score Error\". How come?</li>\n</ul>",
  "messages": [
    {
      "id": 1490453,
      "postDate": "2021-08-25T16:25:06.043Z",
      "content": "<p>Hi, guys.<br>\nI encountered a very weird submission error even when I successfully running on the notebook. <br>\n<img src=\"https://i.imgur.com/L46aM7O.png\" alt=\"\"></p>\n<hr>\n<p>After debugging by sending the sampe_submission file, I found the key function resulting in the submission error. <a href=\"https://www.kaggle.com/lilinchen/lanmark-inference-test\" target=\"_blank\">my notebook</a></p>\n<pre><code>model.load_state_dict(model_state_dict)\n</code></pre>\n<p><strong># model_state_dict = torch.load(os.path.join(\"../input/model-vgg19\",'model_final.pth'))  is ok I already checked.</strong></p>\n<hr>\n<h3>Here are a few points about the model.</h3>\n<ul>\n<li>I trained from my local device and uploaded it to the notebook.</li>\n<li>Pytorch version is \"'1.9.0+cu111'\" in my device, and it's \"1.7.0\" in the notebook.</li>\n<li>I have tried both \"model.cpu()\" and \"model.cuda()\". Each of them will be failed.</li>\n</ul>\n<h3>Weird things</h3>\n<ul>\n<li>I didn't encounter any error when I was running the code in the notebook. </li>\n<li>Submission is ok even my prediction and \"submission.csv\" had been created on ../output workspace. But it shows \"Notebook Threw Exception\" and \"Public Score Error\". How come?</li>\n</ul>",
      "rawMarkdown": "Hi, guys.\nI encountered a very weird submission error even when I successfully running on the notebook. \n![](https://i.imgur.com/L46aM7O.png)\n***\nAfter debugging by sending the sampe_submission file, I found the key function resulting in the submission error. [my notebook](https://www.kaggle.com/lilinchen/lanmark-inference-test)\n```\nmodel.load_state_dict(model_state_dict)\n```\n **# model_state_dict = torch.load(os.path.join(\"../input/model-vgg19\",'model_final.pth'))  is ok I already checked.**\n***\n### Here are a few points about the model.\n* I trained from my local device and uploaded it to the notebook.\n* Pytorch version is \"'1.9.0+cu111'\" in my device, and it's \"1.7.0\" in the notebook.\n* I have tried both \"model.cpu()\" and \"model.cuda()\". Each of them will be failed.\n\n### Weird things\n* I didn't encounter any error when I was running the code in the notebook. \n* Submission is ok even my prediction and \"submission.csv\" had been created on ../output workspace. But it shows \"Notebook Threw Exception\" and \"Public Score Error\". How come?\n\n\n\n\n\n\n",
      "votes": 4
    },
    {
      "id": 1493830,
      "postDate": "2021-08-28T07:29:21.060Z",
      "content": "<p>[Sovled] A author answer my problem in the comment.<br>\n<a href=\"https://www.kaggle.com/hdsk38/pytorch-starter-inference-efficientnet/comments#1493754\" target=\"_blank\">https://www.kaggle.com/hdsk38/pytorch-starter-inference-efficientnet/comments#1493754</a></p>\n<p>The shortly conclusion is <strong><em>train_df.landmark_id.nunique()</em></strong> seems to be different from it was in private rerun. I modify it to constant and everything is no problem. (I mistakenly blamed the pytorch model.load_state_dict 😂</p>",
      "rawMarkdown": "[Sovled] A author answer my problem in the comment.\nhttps://www.kaggle.com/hdsk38/pytorch-starter-inference-efficientnet/comments#1493754\n\nThe shortly conclusion is ***train_df.landmark_id.nunique()*** seems to be different from it was in private rerun. I modify it to constant and everything is no problem. (I mistakenly blamed the pytorch model.load_state_dict 😂",
      "votes": 1
    },
    {
      "id": 1492464,
      "postDate": "2021-08-27T07:37:59.317Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1493830,
      "author_name": "Lilin Chen",
      "author_url": "",
      "post_date": "2021-08-28T07:29:21.060000",
      "content": "<p>[Sovled] A author answer my problem in the comment.<br>\n<a href=\"https://www.kaggle.com/hdsk38/pytorch-starter-inference-efficientnet/comments#1493754\" target=\"_blank\">https://www.kaggle.com/hdsk38/pytorch-starter-inference-efficientnet/comments#1493754</a></p>\n<p>The shortly conclusion is <strong><em>train_df.landmark_id.nunique()</em></strong> seems to be different from it was in private rerun. I modify it to constant and everything is no problem. (I mistakenly blamed the pytorch model.load_state_dict 😂</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1492464,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-27T07:37:59.317000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1490453": "Hi, guys.\nI encountered a very weird submission error even when I successfully running on the notebook. \n![](https://i.imgur.com/L46aM7O.png)\n***\nAfter debugging by sending the sampe_submission file, I found the key function resulting in the submission error. [my notebook](https://www.kaggle.com/lilinchen/lanmark-inference-test)\n```\nmodel.load_state_dict(model_state_dict)\n```\n **# model_state_dict = torch.load(os.path.join(\"../input/model-vgg19\",'model_final.pth'))  is ok I already checked.**\n***\n### Here are a few points about the model.\n* I trained from my local device and uploaded it to the notebook.\n* Pytorch version is \"'1.9.0+cu111'\" in my device, and it's \"1.7.0\" in the notebook.\n* I have tried both \"model.cpu()\" and \"model.cuda()\". Each of them will be failed.\n\n### Weird things\n* I didn't encounter any error when I was running the code in the notebook. \n* Submission is ok even my prediction and \"submission.csv\" had been created on ../output workspace. But it shows \"Notebook Threw Exception\" and \"Public Score Error\". How come?\n\n\n\n\n\n\n",
    "1493830": "[Sovled] A author answer my problem in the comment.\nhttps://www.kaggle.com/hdsk38/pytorch-starter-inference-efficientnet/comments#1493754\n\nThe shortly conclusion is ***train_df.landmark_id.nunique()*** seems to be different from it was in private rerun. I modify it to constant and everything is no problem. (I mistakenly blamed the pytorch model.load_state_dict 😂",
    "1492464": ""
  }
}