{
  "id": 394560,
  "title": "Novice Doubt",
  "url": "/competitions/asl-signs/discussion/394560",
  "author_name": "",
  "post_date": "2023-03-14T00:28:10.891777300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I am new to the Tensorflow Lite modules and am still learning about it.</p>\n<p>I was curious if there is a better way to use the Training Parquet File which is approximately 38 GB without downloading it locally for the scope of this project.</p>\n<p>Is there a way where I can store it on cloud and use it with decent performance or is storing the file locally going to be the most optimal approach.</p>\n<p>Any advice or suggestion will be very helpful.</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "2180576",
      "postDate": "03/14/2023 00:28:10",
      "content": "<p>Hi everyone,</p>\n<p>I am new to the Tensorflow Lite modules and am still learning about it.</p>\n<p>I was curious if there is a better way to use the Training Parquet File which is approximately 38 GB without downloading it locally for the scope of this project.</p>\n<p>Is there a way where I can store it on cloud and use it with decent performance or is storing the file locally going to be the most optimal approach.</p>\n<p>Any advice or suggestion will be very helpful.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hi everyone,\n\nI am new to the Tensorflow Lite modules and am still learning about it.\n\nI was curious if there is a better way to use the Training Parquet File which is approximately 38 GB without downloading it locally for the scope of this project.\n\nIs there a way where I can store it on cloud and use it with decent performance or is storing the file locally going to be the most optimal approach.\n\nAny advice or suggestion will be very helpful.\n\nThank you!",
      "votes": null
    },
    {
      "id": "2180679",
      "postDate": "03/14/2023 03:00:31",
      "content": "<p>Welcome to the world of Tensorflow Lite! Regarding your question, there are definitely ways to work with large files without downloading them locally. One option is to store the file in the cloud, such as on a cloud storage service like Google Cloud Storage or Amazon S3. This would allow you to access the file remotely without downloading it to your local machine.</p>\n<p>In terms of performance, accessing the file remotely can sometimes be slower than accessing it locally, but this will depend on a number of factors such as the speed of your internet connection and the amount of data you need to transfer.</p>\n<p>If you do decide to store the file remotely, you'll need to use an appropriate library or API to access it from your code. This will depend on where you decide to store the file and what tools you're using to work with the data.</p>",
      "rawMarkdown": "Welcome to the world of Tensorflow Lite! Regarding your question, there are definitely ways to work with large files without downloading them locally. One option is to store the file in the cloud, such as on a cloud storage service like Google Cloud Storage or Amazon S3. This would allow you to access the file remotely without downloading it to your local machine.\n\nIn terms of performance, accessing the file remotely can sometimes be slower than accessing it locally, but this will depend on a number of factors such as the speed of your internet connection and the amount of data you need to transfer.\n\nIf you do decide to store the file remotely, you'll need to use an appropriate library or API to access it from your code. This will depend on where you decide to store the file and what tools you're using to work with the data.",
      "votes": null
    },
    {
      "id": "2181045",
      "postDate": "03/14/2023 09:14:48",
      "content": "<p>Note, that parquet files contain a lot of redundant information, like the name of each landmark per frame, etc. You could significantly reduce the size of the dataset by preprocessing it. For instance, the raw dataset converted to TFRecords format takes approximately 21GB. If you reduce the time dimension to 12 points, it would take ~6GB. Having that, you probably could download it locally. The drawback of such an approach is the need to recreate the dataset each time you decide to try new features. <br>\nAt the same time working with a full dataset requires a lot of effort to make the data pipeline efficient enough. It took me a few days to achieve 6s per dataset itertion</p>",
      "rawMarkdown": "Note, that parquet files contain a lot of redundant information, like the name of each landmark per frame, etc. You could significantly reduce the size of the dataset by preprocessing it. For instance, the raw dataset converted to TFRecords format takes approximately 21GB. If you reduce the time dimension to 12 points, it would take ~6GB. Having that, you probably could download it locally. The drawback of such an approach is the need to recreate the dataset each time you decide to try new features. \nAt the same time working with a full dataset requires a lot of effort to make the data pipeline efficient enough. It took me a few days to achieve 6s per dataset itertion",
      "votes": null
    },
    {
      "id": "2181185",
      "postDate": "03/14/2023 11:26:26",
      "content": "<p>You can also remove most face landmarks like eyes, nose, etc, and also probably use float16 for saving the data. My tar archived dataset of npy files (without compression) with full time duration, left_hand, right_hand and face lip landmarks in float16 is like 1.8 GB.</p>",
      "rawMarkdown": "You can also remove most face landmarks like eyes, nose, etc, and also probably use float16 for saving the data. My tar archived dataset of npy files (without compression) with full time duration, left_hand, right_hand and face lip landmarks in float16 is like 1.8 GB.",
      "votes": null
    },
    {
      "id": "2181343",
      "postDate": "03/14/2023 13:25:18",
      "content": "<p>Yeah, you are right. There is only one drawback of such an approach. Once you decide to use pose, for example, you would require to regenerate your dataset from scratch, which takes about half an hour.</p>\n<p>Thank you for your idea of using float16, I think I would migrate my dataset to this precision</p>",
      "rawMarkdown": "Yeah, you are right. There is only one drawback of such an approach. Once you decide to use pose, for example, you would require to regenerate your dataset from scratch, which takes about half an hour.\n\nThank you for your idea of using float16, I think I would migrate my dataset to this precision",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2180679,
      "author_name": "siddharthkumarsah",
      "author_url": "",
      "post_date": "03/14/2023 03:00:31",
      "content": "<p>Welcome to the world of Tensorflow Lite! Regarding your question, there are definitely ways to work with large files without downloading them locally. One option is to store the file in the cloud, such as on a cloud storage service like Google Cloud Storage or Amazon S3. This would allow you to access the file remotely without downloading it to your local machine.</p>\n<p>In terms of performance, accessing the file remotely can sometimes be slower than accessing it locally, but this will depend on a number of factors such as the speed of your internet connection and the amount of data you need to transfer.</p>\n<p>If you do decide to store the file remotely, you'll need to use an appropriate library or API to access it from your code. This will depend on where you decide to store the file and what tools you're using to work with the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2181045,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "03/14/2023 09:14:48",
      "content": "<p>Note, that parquet files contain a lot of redundant information, like the name of each landmark per frame, etc. You could significantly reduce the size of the dataset by preprocessing it. For instance, the raw dataset converted to TFRecords format takes approximately 21GB. If you reduce the time dimension to 12 points, it would take ~6GB. Having that, you probably could download it locally. The drawback of such an approach is the need to recreate the dataset each time you decide to try new features. <br>\nAt the same time working with a full dataset requires a lot of effort to make the data pipeline efficient enough. It took me a few days to achieve 6s per dataset itertion</p>",
      "votes": null,
      "replies": [
        {
          "id": 2181185,
          "author_name": "vbogach",
          "author_url": "",
          "post_date": "03/14/2023 11:26:26",
          "content": "<p>You can also remove most face landmarks like eyes, nose, etc, and also probably use float16 for saving the data. My tar archived dataset of npy files (without compression) with full time duration, left_hand, right_hand and face lip landmarks in float16 is like 1.8 GB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2181343,
              "author_name": "meowmeowmeowmeowmeow",
              "author_url": "",
              "post_date": "03/14/2023 13:25:18",
              "content": "<p>Yeah, you are right. There is only one drawback of such an approach. Once you decide to use pose, for example, you would require to regenerate your dataset from scratch, which takes about half an hour.</p>\n<p>Thank you for your idea of using float16, I think I would migrate my dataset to this precision</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2180576": "Hi everyone,\n\nI am new to the Tensorflow Lite modules and am still learning about it.\n\nI was curious if there is a better way to use the Training Parquet File which is approximately 38 GB without downloading it locally for the scope of this project.\n\nIs there a way where I can store it on cloud and use it with decent performance or is storing the file locally going to be the most optimal approach.\n\nAny advice or suggestion will be very helpful.\n\nThank you!",
    "2180679": "Welcome to the world of Tensorflow Lite! Regarding your question, there are definitely ways to work with large files without downloading them locally. One option is to store the file in the cloud, such as on a cloud storage service like Google Cloud Storage or Amazon S3. This would allow you to access the file remotely without downloading it to your local machine.\n\nIn terms of performance, accessing the file remotely can sometimes be slower than accessing it locally, but this will depend on a number of factors such as the speed of your internet connection and the amount of data you need to transfer.\n\nIf you do decide to store the file remotely, you'll need to use an appropriate library or API to access it from your code. This will depend on where you decide to store the file and what tools you're using to work with the data.",
    "2181045": "Note, that parquet files contain a lot of redundant information, like the name of each landmark per frame, etc. You could significantly reduce the size of the dataset by preprocessing it. For instance, the raw dataset converted to TFRecords format takes approximately 21GB. If you reduce the time dimension to 12 points, it would take ~6GB. Having that, you probably could download it locally. The drawback of such an approach is the need to recreate the dataset each time you decide to try new features. \nAt the same time working with a full dataset requires a lot of effort to make the data pipeline efficient enough. It took me a few days to achieve 6s per dataset itertion",
    "2181185": "You can also remove most face landmarks like eyes, nose, etc, and also probably use float16 for saving the data. My tar archived dataset of npy files (without compression) with full time duration, left_hand, right_hand and face lip landmarks in float16 is like 1.8 GB.",
    "2181343": "Yeah, you are right. There is only one drawback of such an approach. Once you decide to use pose, for example, you would require to regenerate your dataset from scratch, which takes about half an hour.\n\nThank you for your idea of using float16, I think I would migrate my dataset to this precision"
  },
  "source": "meta"
}