{
  "id": 309994,
  "title": "How to read the data quickly and not re-read again in new session?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/309994",
  "author_name": "",
  "post_date": "2022-02-26T23:16:21.994610400Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Kagglers,</p>\n<p>Does anyone know how to read the data quickly? I am currently using <code>cudf</code>.</p>\n<p>Sometimes the session ends and then I need to re-read the data again(takes 4+ mins). Is there a way to skip the step and return back to the same session?</p>",
  "messages": [
    {
      "id": "1705877",
      "postDate": "02/26/2022 23:16:21",
      "content": "<p>Hello Kagglers,</p>\n<p>Does anyone know how to read the data quickly? I am currently using <code>cudf</code>.</p>\n<p>Sometimes the session ends and then I need to re-read the data again(takes 4+ mins). Is there a way to skip the step and return back to the same session?</p>",
      "rawMarkdown": "Hello Kagglers,\n\nDoes anyone know how to read the data quickly? I am currently using `cudf`.\n\nSometimes the session ends and then I need to re-read the data again(takes 4+ mins). Is there a way to skip the step and return back to the same session?",
      "votes": null
    },
    {
      "id": "1705925",
      "postDate": "02/27/2022 01:38:36",
      "content": "<p>Hello! </p>\n<p>You can use parquet files for it. Someone has <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306572\" target=\"_blank\">kindly shared</a> a dataset. </p>",
      "rawMarkdown": "Hello! \n\nYou can use parquet files for it. Someone has [kindly shared](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306572) a dataset.",
      "votes": null
    },
    {
      "id": "1706197",
      "postDate": "02/27/2022 09:46:41",
      "content": "<p>You can either use a more efficient format like Parquet or Feather, or you could start a new VM on GCP or AWS that way you won't have any limits (and probably more computing power).</p>",
      "rawMarkdown": "You can either use a more efficient format like Parquet or Feather, or you could start a new VM on GCP or AWS that way you won't have any limits (and probably more computing power).",
      "votes": null
    },
    {
      "id": "1706856",
      "postDate": "02/27/2022 22:24:36",
      "content": "<ol>\n<li>Use any binary data format for storing your files (e.g. parquet).</li>\n<li>Try to use numerical data type instead of string (e.g. encode the customer_id to int32), avoid saving the string value.</li>\n</ol>",
      "rawMarkdown": "1. Use any binary data format for storing your files (e.g. parquet).\n2. Try to use numerical data type instead of string (e.g. encode the customer_id to int32), avoid saving the string value.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1705925,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/27/2022 01:38:36",
      "content": "<p>Hello! </p>\n<p>You can use parquet files for it. Someone has <a href=\"https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306572\" target=\"_blank\">kindly shared</a> a dataset. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1706197,
      "author_name": "souamesannis",
      "author_url": "",
      "post_date": "02/27/2022 09:46:41",
      "content": "<p>You can either use a more efficient format like Parquet or Feather, or you could start a new VM on GCP or AWS that way you won't have any limits (and probably more computing power).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1706856,
      "author_name": "thariqnugrohotomo",
      "author_url": "",
      "post_date": "02/27/2022 22:24:36",
      "content": "<ol>\n<li>Use any binary data format for storing your files (e.g. parquet).</li>\n<li>Try to use numerical data type instead of string (e.g. encode the customer_id to int32), avoid saving the string value.</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1705877": "Hello Kagglers,\n\nDoes anyone know how to read the data quickly? I am currently using `cudf`.\n\nSometimes the session ends and then I need to re-read the data again(takes 4+ mins). Is there a way to skip the step and return back to the same session?",
    "1705925": "Hello! \n\nYou can use parquet files for it. Someone has [kindly shared](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/306572) a dataset.",
    "1706197": "You can either use a more efficient format like Parquet or Feather, or you could start a new VM on GCP or AWS that way you won't have any limits (and probably more computing power).",
    "1706856": "1. Use any binary data format for storing your files (e.g. parquet).\n2. Try to use numerical data type instead of string (e.g. encode the customer_id to int32), avoid saving the string value."
  },
  "source": "meta"
}