{
  "id": 125639,
  "title": "Yet another \"Submission CSV Not Found\"",
  "url": "/competitions/bengaliai-cv19/discussion/125639",
  "author_name": "",
  "post_date": "2020-01-12T11:39:57.836480900Z",
  "votes": 5,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everybody, \nI would like to report a strange behaviour of my submission kernel that leads to a frustrating \"Submission CSV Not Found\" error for some cases. Let me explain:\nI have created the following submission pipeline:\n1. Loop along test parquet files and extract the contents as \"png\" files in the working directory</p>\n\n<ol>\n<li><p>Define my model (I am using @iafoss great <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">example</a>) with several back-bones. I have modified the original code to accept several model backbones (from torchvision densnet and resnet and from pretrainedmodels inceptionresnetv2, se_resnext50_x32x4d ). </p></li>\n<li><p>Loop along the test png files and along my folding (k=4) much like what @iafoss proposes <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference\">here</a>.</p></li>\n</ol>\n\n<p>When I use the original <code>fastai.vision.models</code> eg. <code>models.densenet121</code> everything goes gracefully and I have my result. \nWhen I use the locally loaded <code>pretrainedmodels</code>\n<code>\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nfrom pretrainedmodels import se_resnext50_32x4d\n</code>\nThen I get the following strange behaviour:\n1. Small testing runs successfully and create the 36 lines  submission file\n2. When I use the training parquet files as input the kernel runs for 20 minutes, finishes gracefully and creates a much bigger submission file as expected.\n3. When I try to submit for the competition I get the \"Submission CSV Not Found\"</p>\n\n<p>I have the same pipeline, just changing the baseline model. Just some final notes: \n* I never use pretrained support since network access is not allowed.\n* I reduced the batch size just in case it was a GPU memory issue\n* when feeding with train parquet files the whole pipeline finishes gracefully\n* the pipeline works using all densenet /resnet variants from torchvision/fast.ai but not from any of the <code>pretrainedmodels</code> package.\n* there are public kernels that use <code>pretrainedmodels</code> library successfully</p>\n\n<p>Did anyone faced such an issue? Thank you in advanced!!!</p>\n\n<p>PS. Locally I have no issue training/inferencing the models!!!</p>",
  "messages": [
    {
      "id": "716869",
      "postDate": "01/12/2020 11:39:57",
      "content": "<p>Hello everybody, \nI would like to report a strange behaviour of my submission kernel that leads to a frustrating \"Submission CSV Not Found\" error for some cases. Let me explain:\nI have created the following submission pipeline:\n1. Loop along test parquet files and extract the contents as \"png\" files in the working directory</p>\n\n<ol>\n<li><p>Define my model (I am using @iafoss great <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">example</a>) with several back-bones. I have modified the original code to accept several model backbones (from torchvision densnet and resnet and from pretrainedmodels inceptionresnetv2, se_resnext50_x32x4d ). </p></li>\n<li><p>Loop along the test png files and along my folding (k=4) much like what @iafoss proposes <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference\">here</a>.</p></li>\n</ol>\n\n<p>When I use the original <code>fastai.vision.models</code> eg. <code>models.densenet121</code> everything goes gracefully and I have my result. \nWhen I use the locally loaded <code>pretrainedmodels</code>\n<code>\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nfrom pretrainedmodels import se_resnext50_32x4d\n</code>\nThen I get the following strange behaviour:\n1. Small testing runs successfully and create the 36 lines  submission file\n2. When I use the training parquet files as input the kernel runs for 20 minutes, finishes gracefully and creates a much bigger submission file as expected.\n3. When I try to submit for the competition I get the \"Submission CSV Not Found\"</p>\n\n<p>I have the same pipeline, just changing the baseline model. Just some final notes: \n* I never use pretrained support since network access is not allowed.\n* I reduced the batch size just in case it was a GPU memory issue\n* when feeding with train parquet files the whole pipeline finishes gracefully\n* the pipeline works using all densenet /resnet variants from torchvision/fast.ai but not from any of the <code>pretrainedmodels</code> package.\n* there are public kernels that use <code>pretrainedmodels</code> library successfully</p>\n\n<p>Did anyone faced such an issue? Thank you in advanced!!!</p>\n\n<p>PS. Locally I have no issue training/inferencing the models!!!</p>",
      "rawMarkdown": "Hello everybody, \nI would like to report a strange behaviour of my submission kernel that leads to a frustrating \"Submission CSV Not Found\" error for some cases. Let me explain:\nI have created the following submission pipeline:\n1. Loop along test parquet files and extract the contents as \"png\" files in the working directory\n\n2. Define my model (I am using @iafoss great [example](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964)) with several back-bones. I have modified the original code to accept several model backbones (from torchvision densnet and resnet and from pretrainedmodels inceptionresnetv2, se_resnext50_x32x4d ). \n\n3. Loop along the test png files and along my folding (k=4) much like what @iafoss proposes [here](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference).\n\nWhen I use the original `fastai.vision.models` eg. `models.densenet121` everything goes gracefully and I have my result. \nWhen I use the locally loaded `pretrainedmodels`\n```\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nfrom pretrainedmodels import se_resnext50_32x4d\n```\nThen I get the following strange behaviour:\n1. Small testing runs successfully and create the 36 lines  submission file\n2. When I use the training parquet files as input the kernel runs for 20 minutes, finishes gracefully and creates a much bigger submission file as expected.\n3. When I try to submit for the competition I get the \"Submission CSV Not Found\"\n\nI have the same pipeline, just changing the baseline model. Just some final notes: \n* I never use pretrained support since network access is not allowed.\n* I reduced the batch size just in case it was a GPU memory issue\n* when feeding with train parquet files the whole pipeline finishes gracefully\n* the pipeline works using all densenet /resnet variants from torchvision/fast.ai but not from any of the `pretrainedmodels` package.\n* there are public kernels that use `pretrainedmodels` library successfully\n\nDid anyone faced such an issue? Thank you in advanced!!!\n\nPS. Locally I have no issue training/inferencing the models!!!",
      "votes": null
    },
    {
      "id": "716875",
      "postDate": "01/12/2020 11:55:02",
      "content": "<p>This could be an OOM issue as <a href=\"/haqishen\">@haqishen</a> suggested <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125529\">here</a>.\nHow many models are you loading at the same time? Maybe you're getting an OOM because of loading too many models? You could try an ablation test here.</p>",
      "rawMarkdown": "This could be an OOM issue as @haqishen suggested [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/125529).\nHow many models are you loading at the same time? Maybe you're getting an OOM because of loading too many models? You could try an ablation test here.",
      "votes": null
    },
    {
      "id": "716877",
      "postDate": "01/12/2020 12:00:16",
      "content": "<p>Thank you for your answer. I use 4 folds for training and hence 4 models for inference. I load all 4 models in a list like this:\n<code>\nmodels = []\nfor i in range(4):\n        model = Dnet_1ch().cuda()\n        model.load_state_dict(torch.load(..., map_location=torch.device('cpu')));\n        model.eval();\n        models.append(model)\n</code></p>",
      "rawMarkdown": "Thank you for your answer. I use 4 folds for training and hence 4 models for inference. I load all 4 models in a list like this:\n```\nmodels = []\nfor i in range(4):\n        model = Dnet_1ch().cuda()\n        model.load_state_dict(torch.load(..., map_location=torch.device('cpu')));\n        model.eval();\n        models.append(model)\n```",
      "votes": null
    },
    {
      "id": "716889",
      "postDate": "01/12/2020 12:24:03",
      "content": "<p>I am also using <code>pretrainedmodels</code> but in a different way:\n```\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null</p>\n\n<p>import pretrainedmodels</p>\n\n<p>base_model = pretrainedmodels.<strong>dict</strong>'se_resnext50_32x4d'\n```</p>\n\n<p>this works fine for me without any problem</p>",
      "rawMarkdown": "I am also using `pretrainedmodels` but in a different way:\n```\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null\n\nimport pretrainedmodels\n\nbase_model = pretrainedmodels.__dict__['se_resnext50_32x4d'](pretrained='imagenet')\n```\n\nthis works fine for me without any problem",
      "votes": null
    },
    {
      "id": "716895",
      "postDate": "01/12/2020 12:35:58",
      "content": "<p>I'm having the same problem but only since this morning, my pipe has not changed but now I end up with \"Submission CSV Not Found\" error. (I even tried a smaller pipeline that the previous one that ran successfully but still got the error.)</p>\n\n<p>I tried to look for OOM error, but I'm able to run my script without problem in the kaggle environment with the training set (which is approximately the same size as the final testing set).</p>\n\n<p>I don't understand what is going on either, any help would be much appreciated.</p>",
      "rawMarkdown": "I'm having the same problem but only since this morning, my pipe has not changed but now I end up with \"Submission CSV Not Found\" error. (I even tried a smaller pipeline that the previous one that ran successfully but still got the error.)\n\nI tried to look for OOM error, but I'm able to run my script without problem in the kaggle environment with the training set (which is approximately the same size as the final testing set).\n\nI don't understand what is going on either, any help would be much appreciated.",
      "votes": null
    },
    {
      "id": "716897",
      "postDate": "01/12/2020 12:36:15",
      "content": "<p>Interesting thank you very much for your suggestion. If it is not a memory issue when I load 4 models at a time, then probably is the installation.. However that fails to explain why the whole pipeline works with the training input.</p>\n\n<p>Thank you <a href=\"/bibek777\">@bibek777</a> </p>",
      "rawMarkdown": "Interesting thank you very much for your suggestion. If it is not a memory issue when I load 4 models at a time, then probably is the installation.. However that fails to explain why the whole pipeline works with the training input.\n\nThank you @bibek777",
      "votes": null
    },
    {
      "id": "717013",
      "postDate": "01/12/2020 15:55:57",
      "content": "<p>Thanks to <a href=\"/bibek777\">@bibek777</a> the problem is solved. To whom it may concern in order to use <code>pretrainedmodels</code>. Please use the following lines:\n<code>\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null\nimport pretrainedmodels\n</code></p>\n\n<p>Just importing the folder may work on interactive mode. The only gray zone left is why simple import:\n<code>\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nimport pretrainedmodels\n</code>\nwas working when inferencing train data!</p>",
      "rawMarkdown": "Thanks to @bibek777 the problem is solved. To whom it may concern in order to use `pretrainedmodels`. Please use the following lines:\n```\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null\nimport pretrainedmodels\n```\n\nJust importing the folder may work on interactive mode. The only gray zone left is why simple import:\n```\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nimport pretrainedmodels\n```\nwas working when inferencing train data!",
      "votes": null
    },
    {
      "id": "717031",
      "postDate": "01/12/2020 16:24:02",
      "content": "<p>I'm not using <code>pretrainedmodels</code> ... my pipe did not change but I spoiled my 5 submissions for today, still don't know what to do!</p>",
      "rawMarkdown": "I'm not using `pretrainedmodels` ... my pipe did not change but I spoiled my 5 submissions for today, still don't know what to do!",
      "votes": null
    },
    {
      "id": "717634",
      "postDate": "01/13/2020 11:51:38",
      "content": "<p>I have exactly the same behaviour with the only difference I load my localy trained .h5 model (build in keras). When I use the public notebook pipelines everything run smooth, but when I load and make inference on my model fails during submission. \nI should also mention that I have tested the code against the training set and is OK (both in time and memory).  </p>",
      "rawMarkdown": "I have exactly the same behaviour with the only difference I load my localy trained .h5 model (build in keras). When I use the public notebook pipelines everything run smooth, but when I load and make inference on my model fails during submission. \nI should also mention that I have tested the code against the training set and is OK (both in time and memory).",
      "votes": null
    },
    {
      "id": "717700",
      "postDate": "01/13/2020 13:44:13",
      "content": "<p>Hey Ioannis,</p>\n\n<p>If you are using any external python package, please make sure you are installing it correctly. Just adding it to path and import it may not suffice. This was my problem with the pytorch <code>pretrainedmodels</code> package. Once I installed it properly using <code>pip</code> from a local dataset the problem was solved.  </p>",
      "rawMarkdown": "Hey Ioannis,\n\nIf you are using any external python package, please make sure you are installing it correctly. Just adding it to path and import it may not suffice. This was my problem with the pytorch `pretrainedmodels` package. Once I installed it properly using `pip` from a local dataset the problem was solved.",
      "votes": null
    },
    {
      "id": "719502",
      "postDate": "01/15/2020 15:00:04",
      "content": "<p>Thanks Kostas for your suggestion. It's not the case for my error but I'll have it in mind. </p>",
      "rawMarkdown": "Thanks Kostas for your suggestion. It's not the case for my error but I'll have it in mind.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 716875,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "01/12/2020 11:55:02",
      "content": "<p>This could be an OOM issue as <a href=\"/haqishen\">@haqishen</a> suggested <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/125529\">here</a>.\nHow many models are you loading at the same time? Maybe you're getting an OOM because of loading too many models? You could try an ablation test here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 716877,
          "author_name": "voglinio",
          "author_url": "",
          "post_date": "01/12/2020 12:00:16",
          "content": "<p>Thank you for your answer. I use 4 folds for training and hence 4 models for inference. I load all 4 models in a list like this:\n<code>\nmodels = []\nfor i in range(4):\n        model = Dnet_1ch().cuda()\n        model.load_state_dict(torch.load(..., map_location=torch.device('cpu')));\n        model.eval();\n        models.append(model)\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 716889,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "01/12/2020 12:24:03",
      "content": "<p>I am also using <code>pretrainedmodels</code> but in a different way:\n```\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null</p>\n\n<p>import pretrainedmodels</p>\n\n<p>base_model = pretrainedmodels.<strong>dict</strong>'se_resnext50_32x4d'\n```</p>\n\n<p>this works fine for me without any problem</p>",
      "votes": null,
      "replies": [
        {
          "id": 716897,
          "author_name": "voglinio",
          "author_url": "",
          "post_date": "01/12/2020 12:36:15",
          "content": "<p>Interesting thank you very much for your suggestion. If it is not a memory issue when I load 4 models at a time, then probably is the installation.. However that fails to explain why the whole pipeline works with the training input.</p>\n\n<p>Thank you <a href=\"/bibek777\">@bibek777</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 716895,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "01/12/2020 12:35:58",
      "content": "<p>I'm having the same problem but only since this morning, my pipe has not changed but now I end up with \"Submission CSV Not Found\" error. (I even tried a smaller pipeline that the previous one that ran successfully but still got the error.)</p>\n\n<p>I tried to look for OOM error, but I'm able to run my script without problem in the kaggle environment with the training set (which is approximately the same size as the final testing set).</p>\n\n<p>I don't understand what is going on either, any help would be much appreciated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 717013,
      "author_name": "voglinio",
      "author_url": "",
      "post_date": "01/12/2020 15:55:57",
      "content": "<p>Thanks to <a href=\"/bibek777\">@bibek777</a> the problem is solved. To whom it may concern in order to use <code>pretrainedmodels</code>. Please use the following lines:\n<code>\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null\nimport pretrainedmodels\n</code></p>\n\n<p>Just importing the folder may work on interactive mode. The only gray zone left is why simple import:\n<code>\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nimport pretrainedmodels\n</code>\nwas working when inferencing train data!</p>",
      "votes": null,
      "replies": [
        {
          "id": 717031,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "01/12/2020 16:24:02",
          "content": "<p>I'm not using <code>pretrainedmodels</code> ... my pipe did not change but I spoiled my 5 submissions for today, still don't know what to do!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 717634,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "01/13/2020 11:51:38",
      "content": "<p>I have exactly the same behaviour with the only difference I load my localy trained .h5 model (build in keras). When I use the public notebook pipelines everything run smooth, but when I load and make inference on my model fails during submission. \nI should also mention that I have tested the code against the training set and is OK (both in time and memory).  </p>",
      "votes": null,
      "replies": [
        {
          "id": 717700,
          "author_name": "voglinio",
          "author_url": "",
          "post_date": "01/13/2020 13:44:13",
          "content": "<p>Hey Ioannis,</p>\n\n<p>If you are using any external python package, please make sure you are installing it correctly. Just adding it to path and import it may not suffice. This was my problem with the pytorch <code>pretrainedmodels</code> package. Once I installed it properly using <code>pip</code> from a local dataset the problem was solved.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719502,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "01/15/2020 15:00:04",
          "content": "<p>Thanks Kostas for your suggestion. It's not the case for my error but I'll have it in mind. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "716869": "Hello everybody, \nI would like to report a strange behaviour of my submission kernel that leads to a frustrating \"Submission CSV Not Found\" error for some cases. Let me explain:\nI have created the following submission pipeline:\n1. Loop along test parquet files and extract the contents as \"png\" files in the working directory\n\n2. Define my model (I am using @iafoss great [example](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964)) with several back-bones. I have modified the original code to accept several model backbones (from torchvision densnet and resnet and from pretrainedmodels inceptionresnetv2, se_resnext50_x32x4d ). \n\n3. Loop along the test png files and along my folding (k=4) much like what @iafoss proposes [here](https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-inference).\n\nWhen I use the original `fastai.vision.models` eg. `models.densenet121` everything goes gracefully and I have my result. \nWhen I use the locally loaded `pretrainedmodels`\n```\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nfrom pretrainedmodels import se_resnext50_32x4d\n```\nThen I get the following strange behaviour:\n1. Small testing runs successfully and create the 36 lines  submission file\n2. When I use the training parquet files as input the kernel runs for 20 minutes, finishes gracefully and creates a much bigger submission file as expected.\n3. When I try to submit for the competition I get the \"Submission CSV Not Found\"\n\nI have the same pipeline, just changing the baseline model. Just some final notes: \n* I never use pretrained support since network access is not allowed.\n* I reduced the batch size just in case it was a GPU memory issue\n* when feeding with train parquet files the whole pipeline finishes gracefully\n* the pipeline works using all densenet /resnet variants from torchvision/fast.ai but not from any of the `pretrainedmodels` package.\n* there are public kernels that use `pretrainedmodels` library successfully\n\nDid anyone faced such an issue? Thank you in advanced!!!\n\nPS. Locally I have no issue training/inferencing the models!!!",
    "716875": "This could be an OOM issue as @haqishen suggested [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/125529).\nHow many models are you loading at the same time? Maybe you're getting an OOM because of loading too many models? You could try an ablation test here.",
    "716877": "Thank you for your answer. I use 4 folds for training and hence 4 models for inference. I load all 4 models in a list like this:\n```\nmodels = []\nfor i in range(4):\n        model = Dnet_1ch().cuda()\n        model.load_state_dict(torch.load(..., map_location=torch.device('cpu')));\n        model.eval();\n        models.append(model)\n```",
    "716889": "I am also using `pretrainedmodels` but in a different way:\n```\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null\n\nimport pretrainedmodels\n\nbase_model = pretrainedmodels.__dict__['se_resnext50_32x4d'](pretrained='imagenet')\n```\n\nthis works fine for me without any problem",
    "716895": "I'm having the same problem but only since this morning, my pipe has not changed but now I end up with \"Submission CSV Not Found\" error. (I even tried a smaller pipeline that the previous one that ran successfully but still got the error.)\n\nI tried to look for OOM error, but I'm able to run my script without problem in the kaggle environment with the training set (which is approximately the same size as the final testing set).\n\nI don't understand what is going on either, any help would be much appreciated.",
    "716897": "Interesting thank you very much for your suggestion. If it is not a memory issue when I load 4 models at a time, then probably is the installation.. However that fails to explain why the whole pipeline works with the training input.\n\nThank you @bibek777",
    "717013": "Thanks to @bibek777 the problem is solved. To whom it may concern in order to use `pretrainedmodels`. Please use the following lines:\n```\n!pip install ../input/pretrainedmodels/pretrainedmodels-0.7.4/pretrainedmodels-0.7.4/ &gt; /dev/null\nimport pretrainedmodels\n```\n\nJust importing the folder may work on interactive mode. The only gray zone left is why simple import:\n```\nimport sys\nsys.path.append('/kaggle/input/pretrainedmodels/pretrainedmodels-0.7.4/')\nimport pretrainedmodels\n```\nwas working when inferencing train data!",
    "717031": "I'm not using `pretrainedmodels` ... my pipe did not change but I spoiled my 5 submissions for today, still don't know what to do!",
    "717634": "I have exactly the same behaviour with the only difference I load my localy trained .h5 model (build in keras). When I use the public notebook pipelines everything run smooth, but when I load and make inference on my model fails during submission. \nI should also mention that I have tested the code against the training set and is OK (both in time and memory).",
    "717700": "Hey Ioannis,\n\nIf you are using any external python package, please make sure you are installing it correctly. Just adding it to path and import it may not suffice. This was my problem with the pytorch `pretrainedmodels` package. Once I installed it properly using `pip` from a local dataset the problem was solved.",
    "719502": "Thanks Kostas for your suggestion. It's not the case for my error but I'll have it in mind."
  },
  "source": "meta"
}