{
  "id": 171916,
  "title": "TF Records Training data",
  "url": "/competitions/landmark-retrieval-2020/discussion/171916",
  "author_name": "",
  "post_date": "2020-08-02T22:46:23.383891800Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>it appears that for efficient training, we should be using TF-Records to maximize disk read speed. Most people use a script in the [DELF package] (<a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training\">https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training</a>) to generate these records which lead to files with a total size of ~100GB.</p>\n\n<p>The current maximum supported Dataset size on Kaggle however is 20GB and only offline submissions are valid, so you can't really include all TF-Records as an external dataset for model training and submissions.</p>\n\n<p>How do you guys deal with this problem? Is it even possible to train models on the entirety of the dataset or do you guys just resort to a subset (~20GB) of the data?</p>",
  "messages": [
    {
      "id": "955762",
      "postDate": "08/02/2020 22:46:23",
      "content": "<p>Hello,</p>\n\n<p>it appears that for efficient training, we should be using TF-Records to maximize disk read speed. Most people use a script in the [DELF package] (<a href=\"https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training\">https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training</a>) to generate these records which lead to files with a total size of ~100GB.</p>\n\n<p>The current maximum supported Dataset size on Kaggle however is 20GB and only offline submissions are valid, so you can't really include all TF-Records as an external dataset for model training and submissions.</p>\n\n<p>How do you guys deal with this problem? Is it even possible to train models on the entirety of the dataset or do you guys just resort to a subset (~20GB) of the data?</p>",
      "rawMarkdown": "Hello,\n\nit appears that for efficient training, we should be using TF-Records to maximize disk read speed. Most people use a script in the [DELF package] (https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training) to generate these records which lead to files with a total size of ~100GB.\n\nThe current maximum supported Dataset size on Kaggle however is 20GB and only offline submissions are valid, so you can't really include all TF-Records as an external dataset for model training and submissions.\n\nHow do you guys deal with this problem? Is it even possible to train models on the entirety of the dataset or do you guys just resort to a subset (~20GB) of the data?",
      "votes": null
    },
    {
      "id": "956903",
      "postDate": "08/03/2020 22:25:32",
      "content": "<p>Hey just bumping up this topic, I am currently trying to generate TFrecords dataset (multiple public 20GB datasets) in a kaggle notebook however disk space is limited to 5GB. Is there any way to circumvent this or do I have to do this locally and then re-upload the records again?</p>",
      "rawMarkdown": "Hey just bumping up this topic, I am currently trying to generate TFrecords dataset (multiple public 20GB datasets) in a kaggle notebook however disk space is limited to 5GB. Is there any way to circumvent this or do I have to do this locally and then re-upload the records again?",
      "votes": null
    },
    {
      "id": "959935",
      "postDate": "08/06/2020 03:17:17",
      "content": "<p>Does this comp have the same data as the Google Recognition Comp? If so, TFRecords have been uploaded <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231\">here</a></p>",
      "rawMarkdown": "Does this comp have the same data as the Google Recognition Comp? If so, TFRecords have been uploaded [here][1]\n\n[1]: https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 956903,
      "author_name": "nawidsayed",
      "author_url": "",
      "post_date": "08/03/2020 22:25:32",
      "content": "<p>Hey just bumping up this topic, I am currently trying to generate TFrecords dataset (multiple public 20GB datasets) in a kaggle notebook however disk space is limited to 5GB. Is there any way to circumvent this or do I have to do this locally and then re-upload the records again?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959935,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/06/2020 03:17:17",
      "content": "<p>Does this comp have the same data as the Google Recognition Comp? If so, TFRecords have been uploaded <a href=\"https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231\">here</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "955762": "Hello,\n\nit appears that for efficient training, we should be using TF-Records to maximize disk read speed. Most people use a script in the [DELF package] (https://github.com/tensorflow/models/tree/master/research/delf/delf/python/training#prepare-the-data-for-training) to generate these records which lead to files with a total size of ~100GB.\n\nThe current maximum supported Dataset size on Kaggle however is 20GB and only offline submissions are valid, so you can't really include all TF-Records as an external dataset for model training and submissions.\n\nHow do you guys deal with this problem? Is it even possible to train models on the entirety of the dataset or do you guys just resort to a subset (~20GB) of the data?",
    "956903": "Hey just bumping up this topic, I am currently trying to generate TFrecords dataset (multiple public 20GB datasets) in a kaggle notebook however disk space is limited to 5GB. Is there any way to circumvent this or do I have to do this locally and then re-upload the records again?",
    "959935": "Does this comp have the same data as the Google Recognition Comp? If so, TFRecords have been uploaded [here][1]\n\n[1]: https://www.kaggle.com/c/landmark-recognition-2020/discussion/172231"
  },
  "source": "meta"
}