{
  "id": 229086,
  "title": "Better to train model in Google Colab than Kaggle",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/229086",
  "author_name": "Old Monk",
  "post_date": "2021-03-28T08:35:52.718000",
  "votes": 11,
  "comment_count": 14,
  "views": 0,
  "content": "<p>This is a huge dataset and usually few GPU hours are going in training the model. If we are participating in multiple competition, we know how precious GPU hours are on Kaggle.<br>\nIt is better to train the model in Colab and use the submission file in Kaggle.<br>\nFew helpful links below</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/general/74235\" target=\"_blank\">https://www.kaggle.com/general/74235</a></li>\n<li><a href=\"https://www.kaggle.com/learn-forum/82587\" target=\"_blank\">https://www.kaggle.com/learn-forum/82587</a></li>\n</ul>\n<p>Another useful way is to use parquet dataset prepared by Saurabh (@saurabhshahane) below. Thanks Saurabh!</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/saurabhshahane/herbarium-2021-parquet-train-and-test?select=Train.parquet\" target=\"_blank\">https://www.kaggle.com/saurabhshahane/herbarium-2021-parquet-train-and-test?select=Train.parquet</a></li>\n</ul>",
  "messages": [
    {
      "id": 1254924,
      "postDate": "2021-03-28T08:35:52.717Z",
      "content": "<p>This is a huge dataset and usually few GPU hours are going in training the model. If we are participating in multiple competition, we know how precious GPU hours are on Kaggle.<br>\nIt is better to train the model in Colab and use the submission file in Kaggle.<br>\nFew helpful links below</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/general/74235\" target=\"_blank\">https://www.kaggle.com/general/74235</a></li>\n<li><a href=\"https://www.kaggle.com/learn-forum/82587\" target=\"_blank\">https://www.kaggle.com/learn-forum/82587</a></li>\n</ul>\n<p>Another useful way is to use parquet dataset prepared by Saurabh (@saurabhshahane) below. Thanks Saurabh!</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/saurabhshahane/herbarium-2021-parquet-train-and-test?select=Train.parquet\" target=\"_blank\">https://www.kaggle.com/saurabhshahane/herbarium-2021-parquet-train-and-test?select=Train.parquet</a></li>\n</ul>",
      "rawMarkdown": "This is a huge dataset and usually few GPU hours are going in training the model. If we are participating in multiple competition, we know how precious GPU hours are on Kaggle.\nIt is better to train the model in Colab and use the submission file in Kaggle.\nFew helpful links below\n\n- https://www.kaggle.com/general/74235\n- https://www.kaggle.com/learn-forum/82587\n\nAnother useful way is to use parquet dataset prepared by Saurabh (@saurabhshahane) below. Thanks Saurabh!\n\n- https://www.kaggle.com/saurabhshahane/herbarium-2021-parquet-train-and-test?select=Train.parquet",
      "votes": 11
    },
    {
      "id": 1284958,
      "postDate": "2021-04-26T12:59:31.203Z",
      "content": "<p>How are guys going about Colab's small disk size? I see 30 GB available, not enough to train on the entire (compressed) dataset. For example, <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>'s TFRecords take 41 GB.<br>\nI tried to use my google drive, but it does not handle larger number of files. The files unzipped to gdrive are somewhere buffered (os.path.exist() == True), but never arrive and are lost on a new VM.</p>",
      "rawMarkdown": "How are guys going about Colab's small disk size? I see 30 GB available, not enough to train on the entire (compressed) dataset. For example, [@luigisaetta](https://www.kaggle.com/luigisaetta)'s TFRecords take 41 GB.\nI tried to use my google drive, but it does not handle larger number of files. The files unzipped to gdrive are somewhere buffered (os.path.exist() == True), but never arrive and are lost on a new VM.",
      "votes": 1
    },
    {
      "id": 1265050,
      "postDate": "2021-04-06T15:31:47.170Z",
      "content": "<p>Thank you very much. This is very helpful!</p>",
      "rawMarkdown": "Thank you very much. This is very helpful!",
      "votes": 1,
      "replies": [
        {
          "id": 1265272,
          "postDate": "2021-04-06T18:40:54.057Z",
          "content": "<p>Sure thing!</p>",
          "rawMarkdown": "Sure thing!"
        }
      ]
    },
    {
      "id": 1263413,
      "postDate": "2021-04-05T11:54:47.463Z",
      "content": "<p>You can directly submit a submission from colab itself using the API token as far as I know. <br>\nI personally prefer Colab over Kaggle Notebooks too.</p>",
      "rawMarkdown": "You can directly submit a submission from colab itself using the API token as far as I know. \nI personally prefer Colab over Kaggle Notebooks too.",
      "votes": 1,
      "replies": [
        {
          "id": 1265271,
          "postDate": "2021-04-06T18:40:00.387Z",
          "content": "<p>Ok that is great, I knew that we can get the data and run python code to get final submission file but was not aware that we can make direct submission from colab to kaggle. Thanks for sharing!</p>",
          "rawMarkdown": "Ok that is great, I knew that we can get the data and run python code to get final submission file but was not aware that we can make direct submission from colab to kaggle. Thanks for sharing!"
        }
      ]
    },
    {
      "id": 1256830,
      "postDate": "2021-03-30T09:29:30.590Z",
      "content": "<p>This is helpful. Thank you <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> for sharing.</p>",
      "rawMarkdown": "This is helpful. Thank you @saurabhbagchi for sharing.",
      "votes": 1
    },
    {
      "id": 1255163,
      "postDate": "2021-03-28T14:11:14.907Z",
      "content": "<p>I wonder how all you guys succeed using GPU. For me, with TPU I'm still not able to train on the entire dataset (exceeding TPU session time). In my team we have defined a strategy to split the training in several step, and should start working on the entire dataset by today evening.</p>\n<p>In the meantime, since it is not off topic, I remember that I have reduced the images to 256x256 and packed in TFRecords. The dataset is: <a href=\"https://www.kaggle.com/luigisaetta/herb2021-256\" target=\"_blank\">Herb-256x256-TfRecords</a>.</p>",
      "rawMarkdown": "I wonder how all you guys succeed using GPU. For me, with TPU I'm still not able to train on the entire dataset (exceeding TPU session time). In my team we have defined a strategy to split the training in several step, and should start working on the entire dataset by today evening.\n\nIn the meantime, since it is not off topic, I remember that I have reduced the images to 256x256 and packed in TFRecords. The dataset is: [Herb-256x256-TfRecords](https://www.kaggle.com/luigisaetta/herb2021-256).",
      "votes": 1,
      "replies": [
        {
          "id": 1255164,
          "postDate": "2021-03-28T14:14:52.593Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, this is very helpful!</p>",
          "rawMarkdown": "Thanks @luigisaetta, this is very helpful!"
        }
      ]
    },
    {
      "id": 1255060,
      "postDate": "2021-03-28T11:40:06.917Z",
      "content": "<p>This is helpful. Thank you <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> for sharing. This is an important toolset.</p>",
      "rawMarkdown": "This is helpful. Thank you @saurabhbagchi for sharing. This is an important toolset.",
      "votes": 1,
      "replies": [
        {
          "id": 1255162,
          "postDate": "2021-03-28T14:11:12.283Z",
          "content": "<p>Most welcome <a href=\"https://www.kaggle.com/olusesiadebisi\" target=\"_blank\">@olusesiadebisi</a> !</p>",
          "rawMarkdown": "Most welcome @olusesiadebisi !"
        }
      ]
    },
    {
      "id": 1311377,
      "postDate": "2021-05-17T11:12:38.113Z",
      "content": "<p>thanks for the info!</p>",
      "rawMarkdown": "thanks for the info!",
      "votes": 1
    },
    {
      "id": 1288017,
      "postDate": "2021-04-29T15:52:07.757Z",
      "content": "<p>Thanks for giving alternative option</p>",
      "rawMarkdown": "Thanks for giving alternative option",
      "votes": 1
    },
    {
      "id": 1285238,
      "postDate": "2021-04-26T17:53:45.873Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1270124,
      "postDate": "2021-04-11T10:11:01.690Z",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1284958,
      "author_name": "Marius Wanko",
      "author_url": "",
      "post_date": "2021-04-26T12:59:31.203000",
      "content": "<p>How are guys going about Colab's small disk size? I see 30 GB available, not enough to train on the entire (compressed) dataset. For example, <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>'s TFRecords take 41 GB.<br>\nI tried to use my google drive, but it does not handle larger number of files. The files unzipped to gdrive are somewhere buffered (os.path.exist() == True), but never arrive and are lost on a new VM.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1265050,
      "author_name": "Mohamed Bakrey Mahmoud",
      "author_url": "",
      "post_date": "2021-04-06T15:31:47.170000",
      "content": "<p>Thank you very much. This is very helpful!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1265272,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-04-06T18:40:54.057000",
          "content": "<p>Sure thing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1263413,
      "author_name": "Manas Vardhan",
      "author_url": "",
      "post_date": "2021-04-05T11:54:47.463000",
      "content": "<p>You can directly submit a submission from colab itself using the API token as far as I know. <br>\nI personally prefer Colab over Kaggle Notebooks too.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1265271,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-04-06T18:40:00.387000",
          "content": "<p>Ok that is great, I knew that we can get the data and run python code to get final submission file but was not aware that we can make direct submission from colab to kaggle. Thanks for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1256830,
      "author_name": "Sumit Gahlot",
      "author_url": "",
      "post_date": "2021-03-30T09:29:30.590000",
      "content": "<p>This is helpful. Thank you <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1255163,
      "author_name": "Luigi Saetta",
      "author_url": "",
      "post_date": "2021-03-28T14:11:14.907000",
      "content": "<p>I wonder how all you guys succeed using GPU. For me, with TPU I'm still not able to train on the entire dataset (exceeding TPU session time). In my team we have defined a strategy to split the training in several step, and should start working on the entire dataset by today evening.</p>\n<p>In the meantime, since it is not off topic, I remember that I have reduced the images to 256x256 and packed in TFRecords. The dataset is: <a href=\"https://www.kaggle.com/luigisaetta/herb2021-256\" target=\"_blank\">Herb-256x256-TfRecords</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1255164,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-03-28T14:14:52.593000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/luigisaetta\" target=\"_blank\">@luigisaetta</a>, this is very helpful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1255060,
      "author_name": "Olusesi Adebisi",
      "author_url": "",
      "post_date": "2021-03-28T11:40:06.917000",
      "content": "<p>This is helpful. Thank you <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> for sharing. This is an important toolset.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1255162,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-03-28T14:11:12.283000",
          "content": "<p>Most welcome <a href=\"https://www.kaggle.com/olusesiadebisi\" target=\"_blank\">@olusesiadebisi</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1311377,
      "author_name": "Kushagra Singh",
      "author_url": "",
      "post_date": "2021-05-17T11:12:38.113000",
      "content": "<p>thanks for the info!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1288017,
      "author_name": "Shubh Patel",
      "author_url": "",
      "post_date": "2021-04-29T15:52:07.757000",
      "content": "<p>Thanks for giving alternative option</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1285238,
      "author_name": "Miguel Xavier Santos",
      "author_url": "",
      "post_date": "2021-04-26T17:53:45.873000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1270124,
      "author_name": "Prince Bubezi",
      "author_url": "",
      "post_date": "2021-04-11T10:11:01.690000",
      "content": "<p>Thank you.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1254924": "This is a huge dataset and usually few GPU hours are going in training the model. If we are participating in multiple competition, we know how precious GPU hours are on Kaggle.\nIt is better to train the model in Colab and use the submission file in Kaggle.\nFew helpful links below\n\n- https://www.kaggle.com/general/74235\n- https://www.kaggle.com/learn-forum/82587\n\nAnother useful way is to use parquet dataset prepared by Saurabh (@saurabhshahane) below. Thanks Saurabh!\n\n- https://www.kaggle.com/saurabhshahane/herbarium-2021-parquet-train-and-test?select=Train.parquet",
    "1284958": "How are guys going about Colab's small disk size? I see 30 GB available, not enough to train on the entire (compressed) dataset. For example, [@luigisaetta](https://www.kaggle.com/luigisaetta)'s TFRecords take 41 GB.\nI tried to use my google drive, but it does not handle larger number of files. The files unzipped to gdrive are somewhere buffered (os.path.exist() == True), but never arrive and are lost on a new VM.",
    "1265050": "Thank you very much. This is very helpful!",
    "1263413": "You can directly submit a submission from colab itself using the API token as far as I know. \nI personally prefer Colab over Kaggle Notebooks too.",
    "1256830": "This is helpful. Thank you @saurabhbagchi for sharing.",
    "1255163": "I wonder how all you guys succeed using GPU. For me, with TPU I'm still not able to train on the entire dataset (exceeding TPU session time). In my team we have defined a strategy to split the training in several step, and should start working on the entire dataset by today evening.\n\nIn the meantime, since it is not off topic, I remember that I have reduced the images to 256x256 and packed in TFRecords. The dataset is: [Herb-256x256-TfRecords](https://www.kaggle.com/luigisaetta/herb2021-256).",
    "1255060": "This is helpful. Thank you @saurabhbagchi for sharing. This is an important toolset.",
    "1311377": "thanks for the info!",
    "1288017": "Thanks for giving alternative option",
    "1285238": "Thanks for sharing!",
    "1270124": "Thank you."
  }
}