{
  "id": 134808,
  "title": "Cannot load data/train on colab",
  "url": "/competitions/bengaliai-cv19/discussion/134808",
  "author_name": "",
  "post_date": "2020-03-10T14:06:41.770859200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Been trying to run this thing and i couldn't get this one thing is it even possible to train from parquet files? or we are suppose to convert them to images or some arrays(which would be memory extensive)? I even tried loading the feather files but gets error 404-not found.\nCommand i ran: !kaggle datasets download corochann/bengaliai-cv19-feather.\nI don't know what if someone can help me, it'll just get easier.\nBTW i have a 1050ti how long will it take for say 10 epochs?\nTLDR;  SOMEONE SHARE THIER RECIPE FOR TRAINING ON COLAB.</p>",
  "messages": [
    {
      "id": "768160",
      "postDate": "03/10/2020 14:06:41",
      "content": "<p>Been trying to run this thing and i couldn't get this one thing is it even possible to train from parquet files? or we are suppose to convert them to images or some arrays(which would be memory extensive)? I even tried loading the feather files but gets error 404-not found.\nCommand i ran: !kaggle datasets download corochann/bengaliai-cv19-feather.\nI don't know what if someone can help me, it'll just get easier.\nBTW i have a 1050ti how long will it take for say 10 epochs?\nTLDR;  SOMEONE SHARE THIER RECIPE FOR TRAINING ON COLAB.</p>",
      "rawMarkdown": "Been trying to run this thing and i couldn't get this one thing is it even possible to train from parquet files? or we are suppose to convert them to images or some arrays(which would be memory extensive)? I even tried loading the feather files but gets error 404-not found.\nCommand i ran: !kaggle datasets download corochann/bengaliai-cv19-feather.\nI don't know what if someone can help me, it'll just get easier.\nBTW i have a 1050ti how long will it take for say 10 epochs?\nTLDR;  SOMEONE SHARE THIER RECIPE FOR TRAINING ON COLAB.",
      "votes": null
    },
    {
      "id": "768186",
      "postDate": "03/10/2020 14:18:32",
      "content": "<p>Convert the images along with the preprocessing you want and save it into a numpy array with numpy.savez_compressed , this takes less memory than the parquet files. Upload these to your google drive and use it for training.</p>\n\n<p>Test all non GPU dependant code on local machine before trying colab, this will save time in case of simple bugs.</p>\n\n<p>P.S - It also helps to save the one hot encoded labels as numpy arrays and loading before training. 1 min of initial load time saves you many hours (for very long runs) of CPU data preparation time.</p>\n\n<p>P.P.S - This is my experience with TensorFlow and Keras, I don't know efficiency of Pytorch dataloaders but I imagine it won't too different.</p>",
      "rawMarkdown": "Convert the images along with the preprocessing you want and save it into a numpy array with numpy.savez_compressed , this takes less memory than the parquet files. Upload these to your google drive and use it for training.\n\nTest all non GPU dependant code on local machine before trying colab, this will save time in case of simple bugs.\n\nP.S - It also helps to save the one hot encoded labels as numpy arrays and loading before training. 1 min of initial load time saves you many hours (for very long runs) of CPU data preparation time.\n\nP.P.S - This is my experience with TensorFlow and Keras, I don't know efficiency of Pytorch dataloaders but I imagine it won't too different.",
      "votes": null
    },
    {
      "id": "772041",
      "postDate": "03/15/2020 00:19:29",
      "content": "<p>You can mount your Google Drive on your Colab script; place the Kaggle training data in Google drive, and load to your Colab script.</p>\n\n<p>Something like:\n<code>\nfrom google.colab import drive\ndrive.mount('/content/gdrive', force_remount=True)\n</code>\nand you'll be able to use\n<code>\n!ls \"/content/gdrive/My Drive\"\n</code>\nif the mount is successful</p>",
      "rawMarkdown": "You can mount your Google Drive on your Colab script; place the Kaggle training data in Google drive, and load to your Colab script.\n\nSomething like:\n```\nfrom google.colab import drive\ndrive.mount('/content/gdrive', force_remount=True)\n```\nand you'll be able to use\n```\n!ls \"/content/gdrive/My Drive\"\n```\nif the mount is successful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768186,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "03/10/2020 14:18:32",
      "content": "<p>Convert the images along with the preprocessing you want and save it into a numpy array with numpy.savez_compressed , this takes less memory than the parquet files. Upload these to your google drive and use it for training.</p>\n\n<p>Test all non GPU dependant code on local machine before trying colab, this will save time in case of simple bugs.</p>\n\n<p>P.S - It also helps to save the one hot encoded labels as numpy arrays and loading before training. 1 min of initial load time saves you many hours (for very long runs) of CPU data preparation time.</p>\n\n<p>P.P.S - This is my experience with TensorFlow and Keras, I don't know efficiency of Pytorch dataloaders but I imagine it won't too different.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 772041,
      "author_name": "axiostpc",
      "author_url": "",
      "post_date": "03/15/2020 00:19:29",
      "content": "<p>You can mount your Google Drive on your Colab script; place the Kaggle training data in Google drive, and load to your Colab script.</p>\n\n<p>Something like:\n<code>\nfrom google.colab import drive\ndrive.mount('/content/gdrive', force_remount=True)\n</code>\nand you'll be able to use\n<code>\n!ls \"/content/gdrive/My Drive\"\n</code>\nif the mount is successful</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "768160": "Been trying to run this thing and i couldn't get this one thing is it even possible to train from parquet files? or we are suppose to convert them to images or some arrays(which would be memory extensive)? I even tried loading the feather files but gets error 404-not found.\nCommand i ran: !kaggle datasets download corochann/bengaliai-cv19-feather.\nI don't know what if someone can help me, it'll just get easier.\nBTW i have a 1050ti how long will it take for say 10 epochs?\nTLDR;  SOMEONE SHARE THIER RECIPE FOR TRAINING ON COLAB.",
    "768186": "Convert the images along with the preprocessing you want and save it into a numpy array with numpy.savez_compressed , this takes less memory than the parquet files. Upload these to your google drive and use it for training.\n\nTest all non GPU dependant code on local machine before trying colab, this will save time in case of simple bugs.\n\nP.S - It also helps to save the one hot encoded labels as numpy arrays and loading before training. 1 min of initial load time saves you many hours (for very long runs) of CPU data preparation time.\n\nP.P.S - This is my experience with TensorFlow and Keras, I don't know efficiency of Pytorch dataloaders but I imagine it won't too different.",
    "772041": "You can mount your Google Drive on your Colab script; place the Kaggle training data in Google drive, and load to your Colab script.\n\nSomething like:\n```\nfrom google.colab import drive\ndrive.mount('/content/gdrive', force_remount=True)\n```\nand you'll be able to use\n```\n!ls \"/content/gdrive/My Drive\"\n```\nif the mount is successful"
  },
  "source": "meta"
}