{
  "id": 207674,
  "title": "Submission csv not found - unsure why",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/207674",
  "author_name": "",
  "post_date": "2020-12-30T19:52:20.550501100Z",
  "votes": -1,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hello.</p>\n<p>I have tried twice to get my inference notebook to save the csv correctly, but the competition does not find it. It does save to the output folder, and as you can see when I display the df it does have content, and is the same layout as the sample submission. Alas, there is apparently still a problem.</p>\n<p>Notebook: <a href=\"https://www.kaggle.com/blueturtle/cassava-inference-2\" target=\"_blank\">https://www.kaggle.com/blueturtle/cassava-inference-2</a></p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1132936",
      "postDate": "12/30/2020 19:52:20",
      "content": "<p>Hello.</p>\n<p>I have tried twice to get my inference notebook to save the csv correctly, but the competition does not find it. It does save to the output folder, and as you can see when I display the df it does have content, and is the same layout as the sample submission. Alas, there is apparently still a problem.</p>\n<p>Notebook: <a href=\"https://www.kaggle.com/blueturtle/cassava-inference-2\" target=\"_blank\">https://www.kaggle.com/blueturtle/cassava-inference-2</a></p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hello.\n\nI have tried twice to get my inference notebook to save the csv correctly, but the competition does not find it. It does save to the output folder, and as you can see when I display the df it does have content, and is the same layout as the sample submission. Alas, there is apparently still a problem.\n\nNotebook: https://www.kaggle.com/blueturtle/cassava-inference-2\n\nThanks!",
      "votes": null
    },
    {
      "id": "1133034",
      "postDate": "12/30/2020 21:38:54",
      "content": "<p>Is 32 too big for batch size on your test_loader?   You may be getting memory error when the full 15K private test set runs  that would not exist when you do the initial submission with the public test of only a single image.    Try 4 to 8 for the batch size.</p>",
      "rawMarkdown": "Is 32 too big for batch size on your test_loader?   You may be getting memory error when the full 15K private test set runs  that would not exist when you do the initial submission with the public test of only a single image.    Try 4 to 8 for the batch size.",
      "votes": null
    },
    {
      "id": "1133055",
      "postDate": "12/30/2020 22:18:06",
      "content": "<p>There must be some error. Try running the inference one on the training dataset and you will know what the problem is. Most likely a memory issue. </p>",
      "rawMarkdown": "There must be some error. Try running the inference one on the training dataset and you will know what the problem is. Most likely a memory issue.",
      "votes": null
    },
    {
      "id": "1133423",
      "postDate": "12/31/2020 07:46:27",
      "content": "<p>Thank you both. I will try to reduce the batch_size. It is 32 for both training and validation though with no errors so my intuition would say it should be ok? But I am still new to all of the intricacies of this so if not I appreciate the feedback as to why. I am assuming that the memory accessible to use for training is the same that would be used for testing. Is this too big of an asssumption?</p>",
      "rawMarkdown": "Thank you both. I will try to reduce the batch_size. It is 32 for both training and validation though with no errors so my intuition would say it should be ok? But I am still new to all of the intricacies of this so if not I appreciate the feedback as to why. I am assuming that the memory accessible to use for training is the same that would be used for testing. Is this too big of an asssumption?",
      "votes": null
    },
    {
      "id": "1133468",
      "postDate": "12/31/2020 08:39:07",
      "content": "<p>I tried both batch_size 4 and 8 and still got the same problem - Submission csv not found.</p>",
      "rawMarkdown": "I tried both batch_size 4 and 8 and still got the same problem - Submission csv not found.",
      "votes": null
    },
    {
      "id": "1133696",
      "postDate": "12/31/2020 12:48:57",
      "content": "<p><a href=\"https://www.kaggle.com/blueturtle\" target=\"_blank\">@blueturtle</a>  Reducing batch size is a way but it will not work always ,In tensorflow there are concept like mixed_precision and few other optimization technique 'https://www.tensorflow.org/guide/mixed_precision'.<br>\nI have less idea on pytorch try something similar .</p>",
      "rawMarkdown": "blueturtle  Reducing batch size is a way but it will not work always ,In tensorflow there are concept like mixed_precision and few other optimization technique 'https://www.tensorflow.org/guide/mixed_precision'.\nI have less idea on pytorch try something similar .",
      "votes": null
    },
    {
      "id": "1133840",
      "postDate": "12/31/2020 15:21:24",
      "content": "<p>Since your model is in private data set(s) cannot run your kernel and I don't know Pytorch that well to trouble shoot just reading the code.  I would replace the test folder with the training folder and RunAll to see what error codes or what the submission file looks like when processing thousands of images rather than the single image currently in public test set.</p>",
      "rawMarkdown": "Since your model is in private data set(s) cannot run your kernel and I don't know Pytorch that well to trouble shoot just reading the code.  I would replace the test folder with the training folder and RunAll to see what error codes or what the submission file looks like when processing thousands of images rather than the single image currently in public test set.",
      "votes": null
    },
    {
      "id": "1134977",
      "postDate": "01/01/2021 19:01:02",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4714993%2Fe72f98952637dced121944cd9daf6667%2FScreenshot%202020-12-31%20at%2018.07.32.png?generation=1609527468521638&amp;alt=media\" alt=\"\"></p>\n<p>Doing this gave the included error.</p>\n<p>It found the 21397 training images in the folder ok but I am assuming it exprienced some problem passing ~75% of them through the model as the output was only 5350, and hence the lengths didn't match up so I got the error?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4714993%2Fe72f98952637dced121944cd9daf6667%2FScreenshot%202020-12-31%20at%2018.07.32.png?generation=1609527468521638&alt=media)\n\nDoing this gave the included error.\n\nIt found the 21397 training images in the folder ok but I am assuming it exprienced some problem passing ~75% of them through the model as the output was only 5350, and hence the lengths didn't match up so I got the error?",
      "votes": null
    },
    {
      "id": "1135018",
      "postDate": "01/01/2021 19:48:24",
      "content": "<p>So when the same/similar error occurs when running the private test you get the same crash but the nature of the kaggle top secret hidden private stuff kills the process leaving no error messages - it than looks for the csv file which of course has not been created because of the error.</p>\n<p>My first guess would be that your model makes a inf or nan set of predictions for image number 5350 - but not sure thats a good first guess.   </p>\n<p>Since your running this in the kernel with the Run All - add a print statement to print the image name and the yhat value before you do the argmax step.</p>\n<p>If you not worried about your model being top secret and a winner - make the dataset with the pth file(s) public and I can run your script on my local machines - the quota for GPU time really slows down our efforts to find and fix.  Back a year ago before the quota's I often had 4 kernels running on Kaggle - with the quota I invested in local machines so I could keep having fun :)</p>",
      "rawMarkdown": "So when the same/similar error occurs when running the private test you get the same crash but the nature of the kaggle top secret hidden private stuff kills the process leaving no error messages - it than looks for the csv file which of course has not been created because of the error.\n\nMy first guess would be that your model makes a inf or nan set of predictions for image number 5350 - but not sure thats a good first guess.   \n\nSince your running this in the kernel with the Run All - add a print statement to print the image name and the yhat value before you do the argmax step.\n\nIf you not worried about your model being top secret and a winner - make the dataset with the pth file(s) public and I can run your script on my local machines - the quota for GPU time really slows down our efforts to find and fix.  Back a year ago before the quota's I often had 4 kernels running on Kaggle - with the quota I invested in local machines so I could keep having fun :)",
      "votes": null
    },
    {
      "id": "1136242",
      "postDate": "01/02/2021 21:37:25",
      "content": "<p>I have made the training notebook public now: <a href=\"https://www.kaggle.com/blueturtle/cassava-training-resnet50\" target=\"_blank\">https://www.kaggle.com/blueturtle/cassava-training-resnet50</a></p>\n<p>Many thanks! My local machine is a low-mid tier iMac with default GPU so sadly Kaggle's GPUs are the best I have access to.</p>",
      "rawMarkdown": "I have made the training notebook public now: https://www.kaggle.com/blueturtle/cassava-training-resnet50\n\nMany thanks! My local machine is a low-mid tier iMac with default GPU so sadly Kaggle's GPUs are the best I have access to.",
      "votes": null
    },
    {
      "id": "1137505",
      "postDate": "01/04/2021 02:16:21",
      "content": "<p>OK - got the training running on Kaggle to generate the model and will probably not get chance to run your submission script until tomorrow.</p>\n<p>I lost my wife two years ago to lung cancer.  Used her insurance money to build what eventually ended up being 4 machines for Kaggle competitions.  Kaggle GPU's pretty decent but the hours limit on their usage makes it hard to participate in more than one competition on kind of a low/slow development.   </p>",
      "rawMarkdown": "OK - got the training running on Kaggle to generate the model and will probably not get chance to run your submission script until tomorrow.\n\nI lost my wife two years ago to lung cancer.  Used her insurance money to build what eventually ended up being 4 machines for Kaggle competitions.  Kaggle GPU's pretty decent but the hours limit on their usage makes it hard to participate in more than one competition on kind of a low/slow development.",
      "votes": null
    },
    {
      "id": "1139905",
      "postDate": "01/05/2021 17:33:08",
      "content": "<p>Awesome, thank you.</p>\n<p>I lost my grandad to it too. I hope you're doing ok (Y)</p>",
      "rawMarkdown": "Awesome, thank you.\n\nI lost my grandad to it too. I hope you're doing ok (Y)",
      "votes": null
    },
    {
      "id": "1140294",
      "postDate": "01/05/2021 23:07:17",
      "content": "<p>No luck running on local machines - Kaggle library versions and my local versions not same for Pytorch and making them the same would result in some tensorflow versions changes for some reason - since I don't use pytorch and need to keep all four machines the same I am just using Kaggle to check out things.</p>\n<p>Have reproduced your error when trying to use the training images - just now added a bit of error trapping and did a Save - if I did it wrong the kernel should die in less than 20 minutes - if I did it right than probably 4 hours run time.</p>",
      "rawMarkdown": "No luck running on local machines - Kaggle library versions and my local versions not same for Pytorch and making them the same would result in some tensorflow versions changes for some reason - since I don't use pytorch and need to keep all four machines the same I am just using Kaggle to check out things.\n\nHave reproduced your error when trying to use the training images - just now added a bit of error trapping and did a Save - if I did it wrong the kernel should die in less than 20 minutes - if I did it right than probably 4 hours run time.",
      "votes": null
    },
    {
      "id": "1144766",
      "postDate": "01/08/2021 16:54:06",
      "content": "<p>Not having any luck.  Burned through 1/2 my quota of GPU time - don't understand pytorch enough to find the error.   Try to run for more than 1 epoch - that might take the issue where the model seems to be predicting values that kill the kernel.</p>",
      "rawMarkdown": "Not having any luck.  Burned through 1/2 my quota of GPU time - don't understand pytorch enough to find the error.   Try to run for more than 1 epoch - that might take the issue where the model seems to be predicting values that kill the kernel.",
      "votes": null
    },
    {
      "id": "1144770",
      "postDate": "01/08/2021 16:58:25",
      "content": "<p>Thanks for trying though :) I will try and cross-reference mine with others' Notebooks to see if there's any obvious reasons.</p>",
      "rawMarkdown": "Thanks for trying though :) I will try and cross-reference mine with others' Notebooks to see if there's any obvious reasons.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1133034,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "12/30/2020 21:38:54",
      "content": "<p>Is 32 too big for batch size on your test_loader?   You may be getting memory error when the full 15K private test set runs  that would not exist when you do the initial submission with the public test of only a single image.    Try 4 to 8 for the batch size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1133468,
          "author_name": "blueturtle",
          "author_url": "",
          "post_date": "12/31/2020 08:39:07",
          "content": "<p>I tried both batch_size 4 and 8 and still got the same problem - Submission csv not found.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133840,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/31/2020 15:21:24",
          "content": "<p>Since your model is in private data set(s) cannot run your kernel and I don't know Pytorch that well to trouble shoot just reading the code.  I would replace the test folder with the training folder and RunAll to see what error codes or what the submission file looks like when processing thousands of images rather than the single image currently in public test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134977,
          "author_name": "blueturtle",
          "author_url": "",
          "post_date": "01/01/2021 19:01:02",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4714993%2Fe72f98952637dced121944cd9daf6667%2FScreenshot%202020-12-31%20at%2018.07.32.png?generation=1609527468521638&amp;alt=media\" alt=\"\"></p>\n<p>Doing this gave the included error.</p>\n<p>It found the 21397 training images in the folder ok but I am assuming it exprienced some problem passing ~75% of them through the model as the output was only 5350, and hence the lengths didn't match up so I got the error?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1135018,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/01/2021 19:48:24",
          "content": "<p>So when the same/similar error occurs when running the private test you get the same crash but the nature of the kaggle top secret hidden private stuff kills the process leaving no error messages - it than looks for the csv file which of course has not been created because of the error.</p>\n<p>My first guess would be that your model makes a inf or nan set of predictions for image number 5350 - but not sure thats a good first guess.   </p>\n<p>Since your running this in the kernel with the Run All - add a print statement to print the image name and the yhat value before you do the argmax step.</p>\n<p>If you not worried about your model being top secret and a winner - make the dataset with the pth file(s) public and I can run your script on my local machines - the quota for GPU time really slows down our efforts to find and fix.  Back a year ago before the quota's I often had 4 kernels running on Kaggle - with the quota I invested in local machines so I could keep having fun :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136242,
          "author_name": "blueturtle",
          "author_url": "",
          "post_date": "01/02/2021 21:37:25",
          "content": "<p>I have made the training notebook public now: <a href=\"https://www.kaggle.com/blueturtle/cassava-training-resnet50\" target=\"_blank\">https://www.kaggle.com/blueturtle/cassava-training-resnet50</a></p>\n<p>Many thanks! My local machine is a low-mid tier iMac with default GPU so sadly Kaggle's GPUs are the best I have access to.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137505,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/04/2021 02:16:21",
          "content": "<p>OK - got the training running on Kaggle to generate the model and will probably not get chance to run your submission script until tomorrow.</p>\n<p>I lost my wife two years ago to lung cancer.  Used her insurance money to build what eventually ended up being 4 machines for Kaggle competitions.  Kaggle GPU's pretty decent but the hours limit on their usage makes it hard to participate in more than one competition on kind of a low/slow development.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139905,
          "author_name": "blueturtle",
          "author_url": "",
          "post_date": "01/05/2021 17:33:08",
          "content": "<p>Awesome, thank you.</p>\n<p>I lost my grandad to it too. I hope you're doing ok (Y)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140294,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2021 23:07:17",
          "content": "<p>No luck running on local machines - Kaggle library versions and my local versions not same for Pytorch and making them the same would result in some tensorflow versions changes for some reason - since I don't use pytorch and need to keep all four machines the same I am just using Kaggle to check out things.</p>\n<p>Have reproduced your error when trying to use the training images - just now added a bit of error trapping and did a Save - if I did it wrong the kernel should die in less than 20 minutes - if I did it right than probably 4 hours run time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144766,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/08/2021 16:54:06",
          "content": "<p>Not having any luck.  Burned through 1/2 my quota of GPU time - don't understand pytorch enough to find the error.   Try to run for more than 1 epoch - that might take the issue where the model seems to be predicting values that kill the kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144770,
          "author_name": "blueturtle",
          "author_url": "",
          "post_date": "01/08/2021 16:58:25",
          "content": "<p>Thanks for trying though :) I will try and cross-reference mine with others' Notebooks to see if there's any obvious reasons.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1133055,
      "author_name": "sauravchat",
      "author_url": "",
      "post_date": "12/30/2020 22:18:06",
      "content": "<p>There must be some error. Try running the inference one on the training dataset and you will know what the problem is. Most likely a memory issue. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133423,
      "author_name": "blueturtle",
      "author_url": "",
      "post_date": "12/31/2020 07:46:27",
      "content": "<p>Thank you both. I will try to reduce the batch_size. It is 32 for both training and validation though with no errors so my intuition would say it should be ok? But I am still new to all of the intricacies of this so if not I appreciate the feedback as to why. I am assuming that the memory accessible to use for training is the same that would be used for testing. Is this too big of an asssumption?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133696,
      "author_name": "ssarkar445",
      "author_url": "",
      "post_date": "12/31/2020 12:48:57",
      "content": "<p><a href=\"https://www.kaggle.com/blueturtle\" target=\"_blank\">@blueturtle</a>  Reducing batch size is a way but it will not work always ,In tensorflow there are concept like mixed_precision and few other optimization technique 'https://www.tensorflow.org/guide/mixed_precision'.<br>\nI have less idea on pytorch try something similar .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1132936": "Hello.\n\nI have tried twice to get my inference notebook to save the csv correctly, but the competition does not find it. It does save to the output folder, and as you can see when I display the df it does have content, and is the same layout as the sample submission. Alas, there is apparently still a problem.\n\nNotebook: https://www.kaggle.com/blueturtle/cassava-inference-2\n\nThanks!",
    "1133034": "Is 32 too big for batch size on your test_loader?   You may be getting memory error when the full 15K private test set runs  that would not exist when you do the initial submission with the public test of only a single image.    Try 4 to 8 for the batch size.",
    "1133055": "There must be some error. Try running the inference one on the training dataset and you will know what the problem is. Most likely a memory issue.",
    "1133423": "Thank you both. I will try to reduce the batch_size. It is 32 for both training and validation though with no errors so my intuition would say it should be ok? But I am still new to all of the intricacies of this so if not I appreciate the feedback as to why. I am assuming that the memory accessible to use for training is the same that would be used for testing. Is this too big of an asssumption?",
    "1133468": "I tried both batch_size 4 and 8 and still got the same problem - Submission csv not found.",
    "1133696": "blueturtle  Reducing batch size is a way but it will not work always ,In tensorflow there are concept like mixed_precision and few other optimization technique 'https://www.tensorflow.org/guide/mixed_precision'.\nI have less idea on pytorch try something similar .",
    "1133840": "Since your model is in private data set(s) cannot run your kernel and I don't know Pytorch that well to trouble shoot just reading the code.  I would replace the test folder with the training folder and RunAll to see what error codes or what the submission file looks like when processing thousands of images rather than the single image currently in public test set.",
    "1134977": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4714993%2Fe72f98952637dced121944cd9daf6667%2FScreenshot%202020-12-31%20at%2018.07.32.png?generation=1609527468521638&alt=media)\n\nDoing this gave the included error.\n\nIt found the 21397 training images in the folder ok but I am assuming it exprienced some problem passing ~75% of them through the model as the output was only 5350, and hence the lengths didn't match up so I got the error?",
    "1135018": "So when the same/similar error occurs when running the private test you get the same crash but the nature of the kaggle top secret hidden private stuff kills the process leaving no error messages - it than looks for the csv file which of course has not been created because of the error.\n\nMy first guess would be that your model makes a inf or nan set of predictions for image number 5350 - but not sure thats a good first guess.   \n\nSince your running this in the kernel with the Run All - add a print statement to print the image name and the yhat value before you do the argmax step.\n\nIf you not worried about your model being top secret and a winner - make the dataset with the pth file(s) public and I can run your script on my local machines - the quota for GPU time really slows down our efforts to find and fix.  Back a year ago before the quota's I often had 4 kernels running on Kaggle - with the quota I invested in local machines so I could keep having fun :)",
    "1136242": "I have made the training notebook public now: https://www.kaggle.com/blueturtle/cassava-training-resnet50\n\nMany thanks! My local machine is a low-mid tier iMac with default GPU so sadly Kaggle's GPUs are the best I have access to.",
    "1137505": "OK - got the training running on Kaggle to generate the model and will probably not get chance to run your submission script until tomorrow.\n\nI lost my wife two years ago to lung cancer.  Used her insurance money to build what eventually ended up being 4 machines for Kaggle competitions.  Kaggle GPU's pretty decent but the hours limit on their usage makes it hard to participate in more than one competition on kind of a low/slow development.",
    "1139905": "Awesome, thank you.\n\nI lost my grandad to it too. I hope you're doing ok (Y)",
    "1140294": "No luck running on local machines - Kaggle library versions and my local versions not same for Pytorch and making them the same would result in some tensorflow versions changes for some reason - since I don't use pytorch and need to keep all four machines the same I am just using Kaggle to check out things.\n\nHave reproduced your error when trying to use the training images - just now added a bit of error trapping and did a Save - if I did it wrong the kernel should die in less than 20 minutes - if I did it right than probably 4 hours run time.",
    "1144766": "Not having any luck.  Burned through 1/2 my quota of GPU time - don't understand pytorch enough to find the error.   Try to run for more than 1 epoch - that might take the issue where the model seems to be predicting values that kill the kernel.",
    "1144770": "Thanks for trying though :) I will try and cross-reference mine with others' Notebooks to see if there's any obvious reasons."
  },
  "source": "meta"
}