{
  "id": 102588,
  "title": "Copying from input data to working dir fails when submitted?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102588",
  "author_name": "Rocking Tree",
  "post_date": "2019-08-03T01:24:57.008000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>I've really been struggling with this for the last week or so and it has made this competition feel like I'm fighting the computer instead of using it to help people :(</p>\n\n<p>I have managed to make successful submissions based on reading the provided input data and doing inference and writing out a submission.csv. Then I decided to transform the source data before training, and thus need to do the same transformation on the test data for Kaggle.</p>\n\n<p>Once I started doing this I started receiving generic \"Status: Submission Error. Error. No additional details provided for this error.\". It has taken me a great deal of tries trying to isolate which part of the script was failing when run on the hidden test data after submission and it seems to have come down to the section below. I stubbed out the submission.csv a long time ago, it simply writes each image in the test_images directory to the output csv with a rating of 0. If I add the code block below, without changing that, it starts failing on Kernel submission.</p>\n\n<p>```\nsource_files = os.listdir(source_data_path_test)</p>\n\n<p>for file in source_files:\n    file_from = os.path.join(source_data_path_test/file)\n    print(\"copying {} to {}\".format(file_from, data_path_test_sized))\n    shutil.copy(file_from, data_path_test_sized)\n```</p>\n\n<p>(Variable declarations below). </p>\n\n<p>I am out of submissions for today (again), but when I don't try to copy the data at all the submission goes through, when I try to copy the data the submission generically errors out. This works perfectly fine when run in the notebook which makes this especially frustrating - I'm just copying from the input to the working dir so I struggle to see how that might change when the Kernel is submitted.</p>\n\n<p>```\n%reload_ext autoreload\n%autoreload 2\n%matplotlib inline\nfrom fastai.vision import *\nfrom fastai.metrics import error_rate\nfrom pathlib import Path\nimport gc\nimport os\nimport multiprocessing\nimport cv2\nimport numpy as np\nimport datetime</p>\n\n<h1>Copy our Python script from the input data to our working directory to import it.</h1>\n\n<h1>Forum Note: This isn't actually being used to pre-process the data, I took that out of the equation a long time ago.</h1>\n\n<p>import shutil\nshutil.copyfile(src= \"../input/testdataset2/combined_processing_kernel.py\", dst=\"../working/combined_processing_kernel.py\")</p>\n\n<p>custom_input_data_path = Path(\"../input/testdataset2\")</p>\n\n<h1>working_path = Path(\"../working/\")</h1>\n\n<p>working_path = Path(\".\")\n```</p>\n\n<p>```\nfinal_image_size = 448\ndata_path = Path(\"../input/aptos2019-blindness-detection\") #\nsource_data_path_training = data_path/\"train_images\"\nsource_data_path_test = data_path/\"test_images</p>\n\n<p>test_folder_name_sized = \"test_images_{}\".format(final_image_size)\ndata_path_test_sized = working_path/test_folder_name_sized #</p>\n\n<p>os.makedirs(data_path_training_sized, exist_ok=True)\nos.makedirs(data_path_test_sized, exist_ok=True)</p>\n\n<h1>After this we do the copy block posted above, and then eventually we simply iterate through source_data_path_test, get the filename out and write that with a zero to the submission.csv to rule out data format errors.</h1>\n\n<p>```</p>\n\n<p>Is there an obvious mistake I'm making here? I'm afraid to wrap the whole file I/O block in a try/except for fear that it'll just silently fail then and I'll start missing images.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 590986,
      "postDate": "2019-08-03T01:24:57.007Z",
      "content": "<p>Hi all,</p>\n\n<p>I've really been struggling with this for the last week or so and it has made this competition feel like I'm fighting the computer instead of using it to help people :(</p>\n\n<p>I have managed to make successful submissions based on reading the provided input data and doing inference and writing out a submission.csv. Then I decided to transform the source data before training, and thus need to do the same transformation on the test data for Kaggle.</p>\n\n<p>Once I started doing this I started receiving generic \"Status: Submission Error. Error. No additional details provided for this error.\". It has taken me a great deal of tries trying to isolate which part of the script was failing when run on the hidden test data after submission and it seems to have come down to the section below. I stubbed out the submission.csv a long time ago, it simply writes each image in the test_images directory to the output csv with a rating of 0. If I add the code block below, without changing that, it starts failing on Kernel submission.</p>\n\n<p>```\nsource_files = os.listdir(source_data_path_test)</p>\n\n<p>for file in source_files:\n    file_from = os.path.join(source_data_path_test/file)\n    print(\"copying {} to {}\".format(file_from, data_path_test_sized))\n    shutil.copy(file_from, data_path_test_sized)\n```</p>\n\n<p>(Variable declarations below). </p>\n\n<p>I am out of submissions for today (again), but when I don't try to copy the data at all the submission goes through, when I try to copy the data the submission generically errors out. This works perfectly fine when run in the notebook which makes this especially frustrating - I'm just copying from the input to the working dir so I struggle to see how that might change when the Kernel is submitted.</p>\n\n<p>```\n%reload_ext autoreload\n%autoreload 2\n%matplotlib inline\nfrom fastai.vision import *\nfrom fastai.metrics import error_rate\nfrom pathlib import Path\nimport gc\nimport os\nimport multiprocessing\nimport cv2\nimport numpy as np\nimport datetime</p>\n\n<h1>Copy our Python script from the input data to our working directory to import it.</h1>\n\n<h1>Forum Note: This isn't actually being used to pre-process the data, I took that out of the equation a long time ago.</h1>\n\n<p>import shutil\nshutil.copyfile(src= \"../input/testdataset2/combined_processing_kernel.py\", dst=\"../working/combined_processing_kernel.py\")</p>\n\n<p>custom_input_data_path = Path(\"../input/testdataset2\")</p>\n\n<h1>working_path = Path(\"../working/\")</h1>\n\n<p>working_path = Path(\".\")\n```</p>\n\n<p>```\nfinal_image_size = 448\ndata_path = Path(\"../input/aptos2019-blindness-detection\") #\nsource_data_path_training = data_path/\"train_images\"\nsource_data_path_test = data_path/\"test_images</p>\n\n<p>test_folder_name_sized = \"test_images_{}\".format(final_image_size)\ndata_path_test_sized = working_path/test_folder_name_sized #</p>\n\n<p>os.makedirs(data_path_training_sized, exist_ok=True)\nos.makedirs(data_path_test_sized, exist_ok=True)</p>\n\n<h1>After this we do the copy block posted above, and then eventually we simply iterate through source_data_path_test, get the filename out and write that with a zero to the submission.csv to rule out data format errors.</h1>\n\n<p>```</p>\n\n<p>Is there an obvious mistake I'm making here? I'm afraid to wrap the whole file I/O block in a try/except for fear that it'll just silently fail then and I'll start missing images.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi all,\n\nI've really been struggling with this for the last week or so and it has made this competition feel like I'm fighting the computer instead of using it to help people :(\n\nI have managed to make successful submissions based on reading the provided input data and doing inference and writing out a submission.csv. Then I decided to transform the source data before training, and thus need to do the same transformation on the test data for Kaggle.\n\nOnce I started doing this I started receiving generic \"Status: Submission Error. Error. No additional details provided for this error.\". It has taken me a great deal of tries trying to isolate which part of the script was failing when run on the hidden test data after submission and it seems to have come down to the section below. I stubbed out the submission.csv a long time ago, it simply writes each image in the test_images directory to the output csv with a rating of 0. If I add the code block below, without changing that, it starts failing on Kernel submission.\n\n```\nsource_files = os.listdir(source_data_path_test)\n        \nfor file in source_files:\n    file_from = os.path.join(source_data_path_test/file)\n    print(\"copying {} to {}\".format(file_from, data_path_test_sized))\n    shutil.copy(file_from, data_path_test_sized)\n```\n\n(Variable declarations below). \n\nI am out of submissions for today (again), but when I don't try to copy the data at all the submission goes through, when I try to copy the data the submission generically errors out. This works perfectly fine when run in the notebook which makes this especially frustrating - I'm just copying from the input to the working dir so I struggle to see how that might change when the Kernel is submitted.\n\n```\n%reload_ext autoreload\n%autoreload 2\n%matplotlib inline\nfrom fastai.vision import *\nfrom fastai.metrics import error_rate\nfrom pathlib import Path\nimport gc\nimport os\nimport multiprocessing\nimport cv2\nimport numpy as np\nimport datetime\n\n# Copy our Python script from the input data to our working directory to import it.\n# Forum Note: This isn't actually being used to pre-process the data, I took that out of the equation a long time ago. \nimport shutil\nshutil.copyfile(src= \"../input/testdataset2/combined_processing_kernel.py\", dst=\"../working/combined_processing_kernel.py\")\n\ncustom_input_data_path = Path(\"../input/testdataset2\")\n#working_path = Path(\"../working/\")\nworking_path = Path(\".\")\n```\n\n```\nfinal_image_size = 448\ndata_path = Path(\"../input/aptos2019-blindness-detection\") #\nsource_data_path_training = data_path/\"train_images\"\nsource_data_path_test = data_path/\"test_images\n\ntest_folder_name_sized = \"test_images_{}\".format(final_image_size)\ndata_path_test_sized = working_path/test_folder_name_sized #\n\nos.makedirs(data_path_training_sized, exist_ok=True)\nos.makedirs(data_path_test_sized, exist_ok=True)\n\n# After this we do the copy block posted above, and then eventually we simply iterate through source_data_path_test, get the filename out and write that with a zero to the submission.csv to rule out data format errors.\n```\n\nIs there an obvious mistake I'm making here? I'm afraid to wrap the whole file I/O block in a try/except for fear that it'll just silently fail then and I'll start missing images.\n\nThanks!"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "590986": "Hi all,\n\nI've really been struggling with this for the last week or so and it has made this competition feel like I'm fighting the computer instead of using it to help people :(\n\nI have managed to make successful submissions based on reading the provided input data and doing inference and writing out a submission.csv. Then I decided to transform the source data before training, and thus need to do the same transformation on the test data for Kaggle.\n\nOnce I started doing this I started receiving generic \"Status: Submission Error. Error. No additional details provided for this error.\". It has taken me a great deal of tries trying to isolate which part of the script was failing when run on the hidden test data after submission and it seems to have come down to the section below. I stubbed out the submission.csv a long time ago, it simply writes each image in the test_images directory to the output csv with a rating of 0. If I add the code block below, without changing that, it starts failing on Kernel submission.\n\n```\nsource_files = os.listdir(source_data_path_test)\n        \nfor file in source_files:\n    file_from = os.path.join(source_data_path_test/file)\n    print(\"copying {} to {}\".format(file_from, data_path_test_sized))\n    shutil.copy(file_from, data_path_test_sized)\n```\n\n(Variable declarations below). \n\nI am out of submissions for today (again), but when I don't try to copy the data at all the submission goes through, when I try to copy the data the submission generically errors out. This works perfectly fine when run in the notebook which makes this especially frustrating - I'm just copying from the input to the working dir so I struggle to see how that might change when the Kernel is submitted.\n\n```\n%reload_ext autoreload\n%autoreload 2\n%matplotlib inline\nfrom fastai.vision import *\nfrom fastai.metrics import error_rate\nfrom pathlib import Path\nimport gc\nimport os\nimport multiprocessing\nimport cv2\nimport numpy as np\nimport datetime\n\n# Copy our Python script from the input data to our working directory to import it.\n# Forum Note: This isn't actually being used to pre-process the data, I took that out of the equation a long time ago. \nimport shutil\nshutil.copyfile(src= \"../input/testdataset2/combined_processing_kernel.py\", dst=\"../working/combined_processing_kernel.py\")\n\ncustom_input_data_path = Path(\"../input/testdataset2\")\n#working_path = Path(\"../working/\")\nworking_path = Path(\".\")\n```\n\n```\nfinal_image_size = 448\ndata_path = Path(\"../input/aptos2019-blindness-detection\") #\nsource_data_path_training = data_path/\"train_images\"\nsource_data_path_test = data_path/\"test_images\n\ntest_folder_name_sized = \"test_images_{}\".format(final_image_size)\ndata_path_test_sized = working_path/test_folder_name_sized #\n\nos.makedirs(data_path_training_sized, exist_ok=True)\nos.makedirs(data_path_test_sized, exist_ok=True)\n\n# After this we do the copy block posted above, and then eventually we simply iterate through source_data_path_test, get the filename out and write that with a zero to the submission.csv to rule out data format errors.\n```\n\nIs there an obvious mistake I'm making here? I'm afraid to wrap the whole file I/O block in a try/except for fear that it'll just silently fail then and I'll start missing images.\n\nThanks!"
  }
}