{
  "id": 180316,
  "title": "Colab training",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/180316",
  "author_name": "",
  "post_date": "2020-09-04T15:38:39.137145600Z",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi,<br>\nThis competition looks quite interesting however I do not have my own GPU so that I use Colab pro to run my kaggle scripts. From what I read it seems a quite huge dataset and I'm wondering if some people are using colab for this competition ? Is it a viable option ? </p>",
  "messages": [
    {
      "id": "998241",
      "postDate": "09/04/2020 15:38:39",
      "content": "<p>Hi,<br>\nThis competition looks quite interesting however I do not have my own GPU so that I use Colab pro to run my kaggle scripts. From what I read it seems a quite huge dataset and I'm wondering if some people are using colab for this competition ? Is it a viable option ? </p>",
      "rawMarkdown": "Hi,\nThis competition looks quite interesting however I do not have my own GPU so that I use Colab pro to run my kaggle scripts. From what I read it seems a quite huge dataset and I'm wondering if some people are using colab for this competition ? Is it a viable option ?",
      "votes": null
    },
    {
      "id": "999256",
      "postDate": "09/05/2020 14:03:44",
      "content": "<p>I am trying the same thing (using colab) but facing some problems. The dataset is too large so we cannot put the dataset directly on the colab. However, loading them from google drive seems very slow.</p>\n<p>The following code is part of my training script.</p>\n<pre><code>train_cfg = cfg[\"train_data_loader\"] #training data (not the complete version)\nrasterizer = build_rasterizer(cfg, dm)\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\ntrain_dataloader = DataLoader(train_dataset, shuffle=train_cfg[\"shuffle\"], batch_size=train_cfg[\"batch_size\"], \nnum_workers=train_cfg[\"num_workers\"]) \n</code></pre>\n<p>It won't take a long time by using kaggle notebook. However, for colab, it needs hours to load it (I have not succeed to loading it yet). BTW, I put my dataset on google drive now.</p>",
      "rawMarkdown": "I am trying the same thing (using colab) but facing some problems. The dataset is too large so we cannot put the dataset directly on the colab. However, loading them from google drive seems very slow.\n\nThe following code is part of my training script.\n\n```\ntrain_cfg = cfg[\"train_data_loader\"] #training data (not the complete version)\nrasterizer = build_rasterizer(cfg, dm)\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\ntrain_dataloader = DataLoader(train_dataset, shuffle=train_cfg[\"shuffle\"], batch_size=train_cfg[\"batch_size\"], \nnum_workers=train_cfg[\"num_workers\"]) \n```\n\nIt won't take a long time by using kaggle notebook. However, for colab, it needs hours to load it (I have not succeed to loading it yet). BTW, I put my dataset on google drive now.",
      "votes": null
    },
    {
      "id": "999261",
      "postDate": "09/05/2020 14:06:49",
      "content": "<p>Transfer the dataset into *.tfrecord (TPU) may helps, but the size may be really large.</p>",
      "rawMarkdown": "Transfer the dataset into *.tfrecord (TPU) may helps, but the size may be really large.",
      "votes": null
    },
    {
      "id": "1001258",
      "postDate": "09/07/2020 07:13:51",
      "content": "<p>Thanks for you reply. For the previous challenges the datasets were on my gdrive so that when I started a training I first copied and extracted the images on colab temporary drive. It took about 5/10 min, something acceptable. Did you get any success since your reply ?</p>",
      "rawMarkdown": "Thanks for you reply. For the previous challenges the datasets were on my gdrive so that when I started a training I first copied and extracted the images on colab temporary drive. It took about 5/10 min, something acceptable. Did you get any success since your reply ?",
      "votes": null
    },
    {
      "id": "1001444",
      "postDate": "09/07/2020 10:36:30",
      "content": "<p>You can use GCS to load, don't need download local.</p>",
      "rawMarkdown": "You can use GCS to load, don't need download local.",
      "votes": null
    },
    {
      "id": "1001496",
      "postDate": "09/07/2020 11:17:20",
      "content": "<p>I heard something similar but did not pay enough attention. So all I have to do is to request access to GCS in my script and then access data as a kaggle kernel without copying the whole dataset?</p>",
      "rawMarkdown": "I heard something similar but did not pay enough attention. So all I have to do is to request access to GCS in my script and then access data as a kaggle kernel without copying the whole dataset?",
      "votes": null
    },
    {
      "id": "1001505",
      "postDate": "09/07/2020 11:22:21",
      "content": "<p>yep. can get GCS adress, it's public dataset.  </p>",
      "rawMarkdown": "yep. can get GCS adress, it's public dataset.",
      "votes": null
    },
    {
      "id": "1001512",
      "postDate": "09/07/2020 11:26:48",
      "content": "<p>Ok thanks I'll definitely have a look</p>",
      "rawMarkdown": "Ok thanks I'll definitely have a look",
      "votes": null
    },
    {
      "id": "1001893",
      "postDate": "09/07/2020 16:47:14",
      "content": "<p>I think if you want to load data through GCS address, you need to create *.tfrecord (TPU) first, and access the data through tensorflow API.<br>\nTrivial functions like cv2.imread, glob.glob, etc. will not work when using GCS address as data path. (I have tried but I failed).<br>\nPlease tell me whether I am wrong, Thanks. I will be appreciated if you share the methods or documents about how to load data from GCS address without *tfrecord.</p>",
      "rawMarkdown": "I think if you want to load data through GCS address, you need to create *.tfrecord (TPU) first, and access the data through tensorflow API.\nTrivial functions like cv2.imread, glob.glob, etc. will not work when using GCS address as data path. (I have tried but I failed).\nPlease tell me whether I am wrong, Thanks. I will be appreciated if you share the methods or documents about how to load data from GCS address without *tfrecord.",
      "votes": null
    },
    {
      "id": "1001896",
      "postDate": "09/07/2020 16:50:08",
      "content": "<p>It seems that the tfrecords are provided officially. You may try these.<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180008\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180008</a></p>",
      "rawMarkdown": "It seems that the tfrecords are provided officially. You may try these.\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180008",
      "votes": null
    },
    {
      "id": "1002019",
      "postDate": "09/07/2020 18:36:20",
      "content": "<p>I'm looking for a solution without using tfrecord.. haha.  <br>\nI agree, I'll be grateful if someone can share a sample using GCS with PyTorch for example. </p>",
      "rawMarkdown": "I'm looking for a solution without using tfrecord.. haha.  \nI agree, I'll be grateful if someone can share a sample using GCS with PyTorch for example.",
      "votes": null
    },
    {
      "id": "1005905",
      "postDate": "09/10/2020 21:21:12",
      "content": "<p><a href=\"https://github.com/vahidk/tfrecord\" target=\"_blank\">https://github.com/vahidk/tfrecord</a> It can be read tfrecord for pytorch</p>",
      "rawMarkdown": "https://github.com/vahidk/tfrecord It can be read tfrecord for pytorch",
      "votes": null
    },
    {
      "id": "1006775",
      "postDate": "09/11/2020 14:43:18",
      "content": "<p>Thanks I didn't know that, I'll have a look as soon as OSIC challenge ends </p>",
      "rawMarkdown": "Thanks I didn't know that, I'll have a look as soon as OSIC challenge ends",
      "votes": null
    },
    {
      "id": "1008537",
      "postDate": "09/13/2020 07:41:41",
      "content": "<p>Hi, You can check my notebook for the same. I am training using Google Colab. I am training with a batch size of 64 using Colab (Not Pro). 😎 <br>\n<a href=\"https://www.kaggle.com/kool777/ultimate-google-colab-training-batch-size-64\" target=\"_blank\">https://www.kaggle.com/kool777/ultimate-google-colab-training-batch-size-64</a></p>",
      "rawMarkdown": "Hi, You can check my notebook for the same. I am training using Google Colab. I am training with a batch size of 64 using Colab (Not Pro). 😎 \nhttps://www.kaggle.com/kool777/ultimate-google-colab-training-batch-size-64",
      "votes": null
    },
    {
      "id": "1009697",
      "postDate": "09/14/2020 07:15:02",
      "content": "<p>Hi many thanks for your kernel, it looks well detailed ! How long does it take to make the notebook ready (data DL + extraction) ?</p>",
      "rawMarkdown": "Hi many thanks for your kernel, it looks well detailed ! How long does it take to make the notebook ready (data DL + extraction) ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 999256,
      "author_name": "hardworkingkaggler",
      "author_url": "",
      "post_date": "09/05/2020 14:03:44",
      "content": "<p>I am trying the same thing (using colab) but facing some problems. The dataset is too large so we cannot put the dataset directly on the colab. However, loading them from google drive seems very slow.</p>\n<p>The following code is part of my training script.</p>\n<pre><code>train_cfg = cfg[\"train_data_loader\"] #training data (not the complete version)\nrasterizer = build_rasterizer(cfg, dm)\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\ntrain_dataloader = DataLoader(train_dataset, shuffle=train_cfg[\"shuffle\"], batch_size=train_cfg[\"batch_size\"], \nnum_workers=train_cfg[\"num_workers\"]) \n</code></pre>\n<p>It won't take a long time by using kaggle notebook. However, for colab, it needs hours to load it (I have not succeed to loading it yet). BTW, I put my dataset on google drive now.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1001258,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/07/2020 07:13:51",
          "content": "<p>Thanks for you reply. For the previous challenges the datasets were on my gdrive so that when I started a training I first copied and extracted the images on colab temporary drive. It took about 5/10 min, something acceptable. Did you get any success since your reply ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001896,
          "author_name": "hardworkingkaggler",
          "author_url": "",
          "post_date": "09/07/2020 16:50:08",
          "content": "<p>It seems that the tfrecords are provided officially. You may try these.<br>\n<a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180008\" target=\"_blank\">https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180008</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 999261,
      "author_name": "hardworkingkaggler",
      "author_url": "",
      "post_date": "09/05/2020 14:06:49",
      "content": "<p>Transfer the dataset into *.tfrecord (TPU) may helps, but the size may be really large.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1001444,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "09/07/2020 10:36:30",
      "content": "<p>You can use GCS to load, don't need download local.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1001496,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/07/2020 11:17:20",
          "content": "<p>I heard something similar but did not pay enough attention. So all I have to do is to request access to GCS in my script and then access data as a kaggle kernel without copying the whole dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001505,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "09/07/2020 11:22:21",
          "content": "<p>yep. can get GCS adress, it's public dataset.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001512,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/07/2020 11:26:48",
          "content": "<p>Ok thanks I'll definitely have a look</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001893,
          "author_name": "hardworkingkaggler",
          "author_url": "",
          "post_date": "09/07/2020 16:47:14",
          "content": "<p>I think if you want to load data through GCS address, you need to create *.tfrecord (TPU) first, and access the data through tensorflow API.<br>\nTrivial functions like cv2.imread, glob.glob, etc. will not work when using GCS address as data path. (I have tried but I failed).<br>\nPlease tell me whether I am wrong, Thanks. I will be appreciated if you share the methods or documents about how to load data from GCS address without *tfrecord.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1002019,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/07/2020 18:36:20",
          "content": "<p>I'm looking for a solution without using tfrecord.. haha.  <br>\nI agree, I'll be grateful if someone can share a sample using GCS with PyTorch for example. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1005905,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "09/10/2020 21:21:12",
          "content": "<p><a href=\"https://github.com/vahidk/tfrecord\" target=\"_blank\">https://github.com/vahidk/tfrecord</a> It can be read tfrecord for pytorch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1006775,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/11/2020 14:43:18",
          "content": "<p>Thanks I didn't know that, I'll have a look as soon as OSIC challenge ends </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1008537,
      "author_name": "kool777",
      "author_url": "",
      "post_date": "09/13/2020 07:41:41",
      "content": "<p>Hi, You can check my notebook for the same. I am training using Google Colab. I am training with a batch size of 64 using Colab (Not Pro). 😎 <br>\n<a href=\"https://www.kaggle.com/kool777/ultimate-google-colab-training-batch-size-64\" target=\"_blank\">https://www.kaggle.com/kool777/ultimate-google-colab-training-batch-size-64</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1009697,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "09/14/2020 07:15:02",
          "content": "<p>Hi many thanks for your kernel, it looks well detailed ! How long does it take to make the notebook ready (data DL + extraction) ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "998241": "Hi,\nThis competition looks quite interesting however I do not have my own GPU so that I use Colab pro to run my kaggle scripts. From what I read it seems a quite huge dataset and I'm wondering if some people are using colab for this competition ? Is it a viable option ?",
    "999256": "I am trying the same thing (using colab) but facing some problems. The dataset is too large so we cannot put the dataset directly on the colab. However, loading them from google drive seems very slow.\n\nThe following code is part of my training script.\n\n```\ntrain_cfg = cfg[\"train_data_loader\"] #training data (not the complete version)\nrasterizer = build_rasterizer(cfg, dm)\ntrain_zarr = ChunkedDataset(dm.require(train_cfg[\"key\"])).open()\ntrain_dataset = AgentDataset(cfg, train_zarr, rasterizer)\ntrain_dataloader = DataLoader(train_dataset, shuffle=train_cfg[\"shuffle\"], batch_size=train_cfg[\"batch_size\"], \nnum_workers=train_cfg[\"num_workers\"]) \n```\n\nIt won't take a long time by using kaggle notebook. However, for colab, it needs hours to load it (I have not succeed to loading it yet). BTW, I put my dataset on google drive now.",
    "999261": "Transfer the dataset into *.tfrecord (TPU) may helps, but the size may be really large.",
    "1001258": "Thanks for you reply. For the previous challenges the datasets were on my gdrive so that when I started a training I first copied and extracted the images on colab temporary drive. It took about 5/10 min, something acceptable. Did you get any success since your reply ?",
    "1001444": "You can use GCS to load, don't need download local.",
    "1001496": "I heard something similar but did not pay enough attention. So all I have to do is to request access to GCS in my script and then access data as a kaggle kernel without copying the whole dataset?",
    "1001505": "yep. can get GCS adress, it's public dataset.",
    "1001512": "Ok thanks I'll definitely have a look",
    "1001893": "I think if you want to load data through GCS address, you need to create *.tfrecord (TPU) first, and access the data through tensorflow API.\nTrivial functions like cv2.imread, glob.glob, etc. will not work when using GCS address as data path. (I have tried but I failed).\nPlease tell me whether I am wrong, Thanks. I will be appreciated if you share the methods or documents about how to load data from GCS address without *tfrecord.",
    "1001896": "It seems that the tfrecords are provided officially. You may try these.\nhttps://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/180008",
    "1002019": "I'm looking for a solution without using tfrecord.. haha.  \nI agree, I'll be grateful if someone can share a sample using GCS with PyTorch for example.",
    "1005905": "https://github.com/vahidk/tfrecord It can be read tfrecord for pytorch",
    "1006775": "Thanks I didn't know that, I'll have a look as soon as OSIC challenge ends",
    "1008537": "Hi, You can check my notebook for the same. I am training using Google Colab. I am training with a batch size of 64 using Colab (Not Pro). 😎 \nhttps://www.kaggle.com/kool777/ultimate-google-colab-training-batch-size-64",
    "1009697": "Hi many thanks for your kernel, it looks well detailed ! How long does it take to make the notebook ready (data DL + extraction) ?"
  },
  "source": "meta"
}