{
  "id": 130895,
  "title": "Training on Google Colab Pro: Issues faced",
  "url": "/competitions/bengaliai-cv19/discussion/130895",
  "author_name": "",
  "post_date": "2020-02-17T02:48:49.246113100Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Unable to train my model on kaggle due to runtime issues, I decided to take up <a href=\"https://colab.research.google.com/signup\">google colab pro</a> subscription. I was able to replicate my kaggle setup on colab and I started the training, to be faced with <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/130256\">Getting error OSError: [Errno 5] Input/output error: '../input/grapheme-imgs-224x224'</a>. The simple diagnosis of the issue is that if too many (1,000+) files are present in a single directory in colab, it crashes. </p>\n\n<p>Fair enough. We have 2,00,000 images to train from. And if they can't be present in a single directory, then nested directories is the way to go. So, I wrote a nifty script to split these into 500 folders, each folder having 400 images. And I modified the dataloader in my training code to incorporate this change in file structure. </p>\n\n<p>Fair enough. Now the code runs without any error. Just one issue. It takes 8 hrs for one epoch to train. And I'm talking about P-100 with 27.4 GB ram used to treat inputs of size 224x224 on the model densenet161 with a batch size of 32. I am confident that it's because of the change in file structure because the one time I had managed to get the training started in the flat file structure, the training was much much faster.</p>\n\n<p>Any ideas how to resolve this issue?</p>\n\n<p>Anyone using Google Colab Pro for training? How does your pipeline look?</p>\n\n<p>I'm a newbie and I continue to run into one brick wall after another, looking for a lifeline here.</p>",
  "messages": [
    {
      "id": "747926",
      "postDate": "02/17/2020 02:48:49",
      "content": "<p>Unable to train my model on kaggle due to runtime issues, I decided to take up <a href=\"https://colab.research.google.com/signup\">google colab pro</a> subscription. I was able to replicate my kaggle setup on colab and I started the training, to be faced with <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/130256\">Getting error OSError: [Errno 5] Input/output error: '../input/grapheme-imgs-224x224'</a>. The simple diagnosis of the issue is that if too many (1,000+) files are present in a single directory in colab, it crashes. </p>\n\n<p>Fair enough. We have 2,00,000 images to train from. And if they can't be present in a single directory, then nested directories is the way to go. So, I wrote a nifty script to split these into 500 folders, each folder having 400 images. And I modified the dataloader in my training code to incorporate this change in file structure. </p>\n\n<p>Fair enough. Now the code runs without any error. Just one issue. It takes 8 hrs for one epoch to train. And I'm talking about P-100 with 27.4 GB ram used to treat inputs of size 224x224 on the model densenet161 with a batch size of 32. I am confident that it's because of the change in file structure because the one time I had managed to get the training started in the flat file structure, the training was much much faster.</p>\n\n<p>Any ideas how to resolve this issue?</p>\n\n<p>Anyone using Google Colab Pro for training? How does your pipeline look?</p>\n\n<p>I'm a newbie and I continue to run into one brick wall after another, looking for a lifeline here.</p>",
      "rawMarkdown": "Unable to train my model on kaggle due to runtime issues, I decided to take up [google colab pro](https://colab.research.google.com/signup) subscription. I was able to replicate my kaggle setup on colab and I started the training, to be faced with [Getting error OSError: [Errno 5] Input/output error: '../input/grapheme-imgs-224x224'](https://www.kaggle.com/c/bengaliai-cv19/discussion/130256). The simple diagnosis of the issue is that if too many (1,000+) files are present in a single directory in colab, it crashes. \n\nFair enough. We have 2,00,000 images to train from. And if they can't be present in a single directory, then nested directories is the way to go. So, I wrote a nifty script to split these into 500 folders, each folder having 400 images. And I modified the dataloader in my training code to incorporate this change in file structure. \n\nFair enough. Now the code runs without any error. Just one issue. It takes 8 hrs for one epoch to train. And I'm talking about P-100 with 27.4 GB ram used to treat inputs of size 224x224 on the model densenet161 with a batch size of 32. I am confident that it's because of the change in file structure because the one time I had managed to get the training started in the flat file structure, the training was much much faster.\n\nAny ideas how to resolve this issue?\n\nAnyone using Google Colab Pro for training? How does your pipeline look?\n\nI'm a newbie and I continue to run into one brick wall after another, looking for a lifeline here.",
      "votes": null
    },
    {
      "id": "747947",
      "postDate": "02/17/2020 03:13:22",
      "content": "<p>I'm using normal google colab and I've never had any issues loading and using all the files.</p>",
      "rawMarkdown": "I'm using normal google colab and I've never had any issues loading and using all the files.",
      "votes": null
    },
    {
      "id": "747964",
      "postDate": "02/17/2020 03:51:11",
      "content": "<p>FAQ says that “For now, Colab Pro is only available in the US.” May be that is the reason?</p>",
      "rawMarkdown": "FAQ says that “For now, Colab Pro is only available in the US.” May be that is the reason?",
      "votes": null
    },
    {
      "id": "747992",
      "postDate": "02/17/2020 05:05:55",
      "content": "<p>Don't place files in mount folder (this folder is shared between drive server and colab server). Read file from this too slow (Wait time for sync, the more files the more slower, read big single file is fast). Let's copy all yours file to another folder.\n- Zip all images to a zip file.\n- Upload zip file to drive.\n- Mount drive to colab.\n- Using '!cp' command to copy zip file to another folder in hard disk of colab server.\n- Using '!unzip --qq' command to unzip file.</p>",
      "rawMarkdown": "Don't place files in mount folder (this folder is shared between drive server and colab server). Read file from this too slow (Wait time for sync, the more files the more slower, read big single file is fast). Let's copy all yours file to another folder.\n- Zip all images to a zip file.\n- Upload zip file to drive.\n- Mount drive to colab.\n- Using '!cp' command to copy zip file to another folder in hard disk of colab server.\n- Using '!unzip --qq' command to unzip file.",
      "votes": null
    },
    {
      "id": "748020",
      "postDate": "02/17/2020 05:41:34",
      "content": "<p>I don't think so. The Input/Output error is independent of geography. Seems like the country check is only a very loose check - it just asks for the ZIP code.</p>",
      "rawMarkdown": "I don't think so. The Input/Output error is independent of geography. Seems like the country check is only a very loose check - it just asks for the ZIP code.",
      "votes": null
    },
    {
      "id": "748021",
      "postDate": "02/17/2020 05:42:07",
      "content": "<p>Did you mount google drive onto google colab?</p>",
      "rawMarkdown": "Did you mount google drive onto google colab?",
      "votes": null
    },
    {
      "id": "748022",
      "postDate": "02/17/2020 05:44:48",
      "content": "<p>That's an excellent idea <a href=\"/quan0095\">@quan0095</a> . Thanks for the correct and quick diagnosis. I will surely try this. </p>\n\n<p>You're a life-saver.</p>",
      "rawMarkdown": "That's an excellent idea @quan0095 . Thanks for the correct and quick diagnosis. I will surely try this. \n\nYou're a life-saver.",
      "votes": null
    },
    {
      "id": "748204",
      "postDate": "02/17/2020 09:24:07",
      "content": "<p>Whenever I use colab to train I always download the data from Kaggle using Kaggle's command line API. Using a network drive to load data while training is not a good idea since you will wait significantly long for input/output operations.</p>",
      "rawMarkdown": "Whenever I use colab to train I always download the data from Kaggle using Kaggle's command line API. Using a network drive to load data while training is not a good idea since you will wait significantly long for input/output operations.",
      "votes": null
    },
    {
      "id": "748221",
      "postDate": "02/17/2020 09:43:47",
      "content": "<p>It seems you're right! On free colab I do the same as <a href=\"/quan0095\">@quan0095</a>:</p>\n\n<p>```\nimport os\nimport errno</p>\n\n<p>from google.colab import drive\ndrive.mount('/content/drive')</p>\n\n<p>INPUT_DIR ='/content/drive/My Drive/kaggle/bengali/input/'\nTRAIN_DIR = './train/'\nTRAIN_IMG_DIR=TRAIN_DIR+'imgs'\nDATASET= 'bengali.zip'</p>\n\n<p>try:\n  os.mkdir(TRAIN_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_DIR+' already exists')\n    else:\n        raise\ntry:\n  os.mkdir(TRAIN_IMG_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_IMG_DIR+' already exists')\n    else:\n        raise</p>\n\n<p>os.system('cp '+ '\"'+INPUT_DIR+DATASET+'\" ' + TRAIN_DIR)\nos.system('unzip -q '+TRAIN_DIR+DATASET+ ' -d '+ TRAIN_IMG_DIR)\n```</p>",
      "rawMarkdown": "It seems you're right! On free colab I do the same as @quan0095:\n\n```\nimport os\nimport errno\n\nfrom google.colab import drive\ndrive.mount('/content/drive')\n\nINPUT_DIR ='/content/drive/My Drive/kaggle/bengali/input/'\nTRAIN_DIR = './train/'\nTRAIN_IMG_DIR=TRAIN_DIR+'imgs'\nDATASET= 'bengali.zip'\n\ntry:\n  os.mkdir(TRAIN_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_DIR+' already exists')\n    else:\n        raise\ntry:\n  os.mkdir(TRAIN_IMG_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_IMG_DIR+' already exists')\n    else:\n        raise\n\nos.system('cp '+ '\"'+INPUT_DIR+DATASET+'\" ' + TRAIN_DIR)\nos.system('unzip -q '+TRAIN_DIR+DATASET+ ' -d '+ TRAIN_IMG_DIR)\n```",
      "votes": null
    },
    {
      "id": "748455",
      "postDate": "02/17/2020 14:36:03",
      "content": "<p>I only mount to save models, I use the Kaggle API to download the data</p>",
      "rawMarkdown": "I only mount to save models, I use the Kaggle API to download the data",
      "votes": null
    },
    {
      "id": "759875",
      "postDate": "02/29/2020 15:09:35",
      "content": "<p>How do you do that to avoid the \"Getting error OSError: [Errno 5]\" error later on the process?</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "How do you do that to avoid the \"Getting error OSError: [Errno 5]\" error later on the process?\n\nThank you",
      "votes": null
    },
    {
      "id": "911279",
      "postDate": "07/01/2020 16:34:16",
      "content": "<p>I think Colab cannot solve your problem since the runtimes will always be limited. No way your algo with several epochs can be trained under that much limited time frame.\nIf you don't want to go on cloud networks to keep your pocket intact then I'd suggest you have a look at <a href=\"https://www.qblocks.cloud\">Q Blocks</a>. It's a peer to peer computing platform that offers GPU instances starting at just $0.15 per hour. For the cost of a CPU, you get a GPU.</p>\n\n<p>The instances can be pre-configured with AI frameworks such as Tensorflow and Pytorch and Jupyter notebooks come pre-configured so you don't have to waste time in setting up anything. </p>\n\n<p>Could you share any other problem you faced with Colab?</p>",
      "rawMarkdown": "I think Colab cannot solve your problem since the runtimes will always be limited. No way your algo with several epochs can be trained under that much limited time frame.\nIf you don't want to go on cloud networks to keep your pocket intact then I'd suggest you have a look at [Q Blocks](https://www.qblocks.cloud). It's a peer to peer computing platform that offers GPU instances starting at just $0.15 per hour. For the cost of a CPU, you get a GPU.\n\nThe instances can be pre-configured with AI frameworks such as Tensorflow and Pytorch and Jupyter notebooks come pre-configured so you don't have to waste time in setting up anything. \n\nCould you share any other problem you faced with Colab?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 747947,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "02/17/2020 03:13:22",
      "content": "<p>I'm using normal google colab and I've never had any issues loading and using all the files.</p>",
      "votes": null,
      "replies": [
        {
          "id": 748021,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "02/17/2020 05:42:07",
          "content": "<p>Did you mount google drive onto google colab?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748455,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/17/2020 14:36:03",
          "content": "<p>I only mount to save models, I use the Kaggle API to download the data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 747964,
      "author_name": "andreyzotov",
      "author_url": "",
      "post_date": "02/17/2020 03:51:11",
      "content": "<p>FAQ says that “For now, Colab Pro is only available in the US.” May be that is the reason?</p>",
      "votes": null,
      "replies": [
        {
          "id": 748020,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "02/17/2020 05:41:34",
          "content": "<p>I don't think so. The Input/Output error is independent of geography. Seems like the country check is only a very loose check - it just asks for the ZIP code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748221,
          "author_name": "andreyzotov",
          "author_url": "",
          "post_date": "02/17/2020 09:43:47",
          "content": "<p>It seems you're right! On free colab I do the same as <a href=\"/quan0095\">@quan0095</a>:</p>\n\n<p>```\nimport os\nimport errno</p>\n\n<p>from google.colab import drive\ndrive.mount('/content/drive')</p>\n\n<p>INPUT_DIR ='/content/drive/My Drive/kaggle/bengali/input/'\nTRAIN_DIR = './train/'\nTRAIN_IMG_DIR=TRAIN_DIR+'imgs'\nDATASET= 'bengali.zip'</p>\n\n<p>try:\n  os.mkdir(TRAIN_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_DIR+' already exists')\n    else:\n        raise\ntry:\n  os.mkdir(TRAIN_IMG_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_IMG_DIR+' already exists')\n    else:\n        raise</p>\n\n<p>os.system('cp '+ '\"'+INPUT_DIR+DATASET+'\" ' + TRAIN_DIR)\nos.system('unzip -q '+TRAIN_DIR+DATASET+ ' -d '+ TRAIN_IMG_DIR)\n```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 747992,
      "author_name": "quan0095",
      "author_url": "",
      "post_date": "02/17/2020 05:05:55",
      "content": "<p>Don't place files in mount folder (this folder is shared between drive server and colab server). Read file from this too slow (Wait time for sync, the more files the more slower, read big single file is fast). Let's copy all yours file to another folder.\n- Zip all images to a zip file.\n- Upload zip file to drive.\n- Mount drive to colab.\n- Using '!cp' command to copy zip file to another folder in hard disk of colab server.\n- Using '!unzip --qq' command to unzip file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 748022,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "02/17/2020 05:44:48",
          "content": "<p>That's an excellent idea <a href=\"/quan0095\">@quan0095</a> . Thanks for the correct and quick diagnosis. I will surely try this. </p>\n\n<p>You're a life-saver.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 748204,
      "author_name": "ibraheemmoosa",
      "author_url": "",
      "post_date": "02/17/2020 09:24:07",
      "content": "<p>Whenever I use colab to train I always download the data from Kaggle using Kaggle's command line API. Using a network drive to load data while training is not a good idea since you will wait significantly long for input/output operations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 759875,
          "author_name": "joovasco",
          "author_url": "",
          "post_date": "02/29/2020 15:09:35",
          "content": "<p>How do you do that to avoid the \"Getting error OSError: [Errno 5]\" error later on the process?</p>\n\n<p>Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 911279,
      "author_name": "genesis96839",
      "author_url": "",
      "post_date": "07/01/2020 16:34:16",
      "content": "<p>I think Colab cannot solve your problem since the runtimes will always be limited. No way your algo with several epochs can be trained under that much limited time frame.\nIf you don't want to go on cloud networks to keep your pocket intact then I'd suggest you have a look at <a href=\"https://www.qblocks.cloud\">Q Blocks</a>. It's a peer to peer computing platform that offers GPU instances starting at just $0.15 per hour. For the cost of a CPU, you get a GPU.</p>\n\n<p>The instances can be pre-configured with AI frameworks such as Tensorflow and Pytorch and Jupyter notebooks come pre-configured so you don't have to waste time in setting up anything. </p>\n\n<p>Could you share any other problem you faced with Colab?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "747926": "Unable to train my model on kaggle due to runtime issues, I decided to take up [google colab pro](https://colab.research.google.com/signup) subscription. I was able to replicate my kaggle setup on colab and I started the training, to be faced with [Getting error OSError: [Errno 5] Input/output error: '../input/grapheme-imgs-224x224'](https://www.kaggle.com/c/bengaliai-cv19/discussion/130256). The simple diagnosis of the issue is that if too many (1,000+) files are present in a single directory in colab, it crashes. \n\nFair enough. We have 2,00,000 images to train from. And if they can't be present in a single directory, then nested directories is the way to go. So, I wrote a nifty script to split these into 500 folders, each folder having 400 images. And I modified the dataloader in my training code to incorporate this change in file structure. \n\nFair enough. Now the code runs without any error. Just one issue. It takes 8 hrs for one epoch to train. And I'm talking about P-100 with 27.4 GB ram used to treat inputs of size 224x224 on the model densenet161 with a batch size of 32. I am confident that it's because of the change in file structure because the one time I had managed to get the training started in the flat file structure, the training was much much faster.\n\nAny ideas how to resolve this issue?\n\nAnyone using Google Colab Pro for training? How does your pipeline look?\n\nI'm a newbie and I continue to run into one brick wall after another, looking for a lifeline here.",
    "747947": "I'm using normal google colab and I've never had any issues loading and using all the files.",
    "747964": "FAQ says that “For now, Colab Pro is only available in the US.” May be that is the reason?",
    "747992": "Don't place files in mount folder (this folder is shared between drive server and colab server). Read file from this too slow (Wait time for sync, the more files the more slower, read big single file is fast). Let's copy all yours file to another folder.\n- Zip all images to a zip file.\n- Upload zip file to drive.\n- Mount drive to colab.\n- Using '!cp' command to copy zip file to another folder in hard disk of colab server.\n- Using '!unzip --qq' command to unzip file.",
    "748020": "I don't think so. The Input/Output error is independent of geography. Seems like the country check is only a very loose check - it just asks for the ZIP code.",
    "748021": "Did you mount google drive onto google colab?",
    "748022": "That's an excellent idea @quan0095 . Thanks for the correct and quick diagnosis. I will surely try this. \n\nYou're a life-saver.",
    "748204": "Whenever I use colab to train I always download the data from Kaggle using Kaggle's command line API. Using a network drive to load data while training is not a good idea since you will wait significantly long for input/output operations.",
    "748221": "It seems you're right! On free colab I do the same as @quan0095:\n\n```\nimport os\nimport errno\n\nfrom google.colab import drive\ndrive.mount('/content/drive')\n\nINPUT_DIR ='/content/drive/My Drive/kaggle/bengali/input/'\nTRAIN_DIR = './train/'\nTRAIN_IMG_DIR=TRAIN_DIR+'imgs'\nDATASET= 'bengali.zip'\n\ntry:\n  os.mkdir(TRAIN_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_DIR+' already exists')\n    else:\n        raise\ntry:\n  os.mkdir(TRAIN_IMG_DIR)\nexcept OSError as e:\n    if e.errno == errno.EEXIST:\n        print(TRAIN_IMG_DIR+' already exists')\n    else:\n        raise\n\nos.system('cp '+ '\"'+INPUT_DIR+DATASET+'\" ' + TRAIN_DIR)\nos.system('unzip -q '+TRAIN_DIR+DATASET+ ' -d '+ TRAIN_IMG_DIR)\n```",
    "748455": "I only mount to save models, I use the Kaggle API to download the data",
    "759875": "How do you do that to avoid the \"Getting error OSError: [Errno 5]\" error later on the process?\n\nThank you",
    "911279": "I think Colab cannot solve your problem since the runtimes will always be limited. No way your algo with several epochs can be trained under that much limited time frame.\nIf you don't want to go on cloud networks to keep your pocket intact then I'd suggest you have a look at [Q Blocks](https://www.qblocks.cloud). It's a peer to peer computing platform that offers GPU instances starting at just $0.15 per hour. For the cost of a CPU, you get a GPU.\n\nThe instances can be pre-configured with AI frameworks such as Tensorflow and Pytorch and Jupyter notebooks come pre-configured so you don't have to waste time in setting up anything. \n\nCould you share any other problem you faced with Colab?"
  },
  "source": "meta"
}