{
  "id": 396267,
  "title": "Data size is huge",
  "url": "/competitions/asl-signs/discussion/396267",
  "author_name": "",
  "post_date": "2023-03-21T01:06:55.776574500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>How do you guys train model on such a huge dataset that is 30 GB.Any clues?</p>",
  "messages": [
    {
      "id": "2190020",
      "postDate": "03/21/2023 01:06:55",
      "content": "<p>How do you guys train model on such a huge dataset that is 30 GB.Any clues?</p>",
      "rawMarkdown": "How do you guys train model on such a huge dataset that is 30 GB.Any clues?",
      "votes": null
    },
    {
      "id": "2190067",
      "postDate": "03/21/2023 02:23:03",
      "content": "<p>in fact, the dataset size is relatively small, compared to the other competitions. Moreover, you can iterate the whole dataset within 5s if you use proper pipeline. for me, training 1 epoch only takes about 30 secs using .tfrecords and tf.data pipeline.</p>",
      "rawMarkdown": "in fact, the dataset size is relatively small, compared to the other competitions. Moreover, you can iterate the whole dataset within 5s if you use proper pipeline. for me, training 1 epoch only takes about 30 secs using .tfrecords and tf.data pipeline.",
      "votes": null
    },
    {
      "id": "2190114",
      "postDate": "03/21/2023 03:43:25",
      "content": "<p>can you show me the code to import data using tf.records and tf.data pipeline</p>",
      "rawMarkdown": "can you show me the code to import data using tf.records and tf.data pipeline",
      "votes": null
    },
    {
      "id": "2190582",
      "postDate": "03/21/2023 11:05:51",
      "content": "<p>there already exist many public notebooks you can refer to. for example, if you want to use tfrecords+tf.data, you can try this notebook.<br>\n<a href=\"https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-convlstm1d\" target=\"_blank\">https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-convlstm1d</a></p>",
      "rawMarkdown": "there already exist many public notebooks you can refer to. for example, if you want to use tfrecords+tf.data, you can try this notebook.\nhttps://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-convlstm1d",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2190067,
      "author_name": "hoyso48",
      "author_url": "",
      "post_date": "03/21/2023 02:23:03",
      "content": "<p>in fact, the dataset size is relatively small, compared to the other competitions. Moreover, you can iterate the whole dataset within 5s if you use proper pipeline. for me, training 1 epoch only takes about 30 secs using .tfrecords and tf.data pipeline.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2190114,
          "author_name": "johnny77kozman0",
          "author_url": "",
          "post_date": "03/21/2023 03:43:25",
          "content": "<p>can you show me the code to import data using tf.records and tf.data pipeline</p>",
          "votes": null,
          "replies": [
            {
              "id": 2190582,
              "author_name": "hoyso48",
              "author_url": "",
              "post_date": "03/21/2023 11:05:51",
              "content": "<p>there already exist many public notebooks you can refer to. for example, if you want to use tfrecords+tf.data, you can try this notebook.<br>\n<a href=\"https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-convlstm1d\" target=\"_blank\">https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-convlstm1d</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2190020": "How do you guys train model on such a huge dataset that is 30 GB.Any clues?",
    "2190067": "in fact, the dataset size is relatively small, compared to the other competitions. Moreover, you can iterate the whole dataset within 5s if you use proper pipeline. for me, training 1 epoch only takes about 30 secs using .tfrecords and tf.data pipeline.",
    "2190114": "can you show me the code to import data using tf.records and tf.data pipeline",
    "2190582": "there already exist many public notebooks you can refer to. for example, if you want to use tfrecords+tf.data, you can try this notebook.\nhttps://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-convlstm1d"
  },
  "source": "meta"
}