{
  "id": 200872,
  "title": "How to create the pkl.gzip file?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/200872",
  "author_name": "",
  "post_date": "2020-12-02T08:09:34.439738300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Can anyone share the code of creating riiid_train.pkl.gzip? Thanks!</p>",
  "messages": [
    {
      "id": "1099257",
      "postDate": "12/02/2020 08:09:34",
      "content": "<p>Can anyone share the code of creating riiid_train.pkl.gzip? Thanks!</p>",
      "rawMarkdown": "Can anyone share the code of creating riiid_train.pkl.gzip? Thanks!",
      "votes": null
    },
    {
      "id": "1099280",
      "postDate": "12/02/2020 08:30:35",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets#File-Formats\" target=\"_blank\">https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets#File-Formats</a></p>",
      "rawMarkdown": "https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets#File-Formats",
      "votes": null
    },
    {
      "id": "1099981",
      "postDate": "12/02/2020 18:31:07",
      "content": "<p>Yeah, it seems that you read the data riiid_train.pkl.gzi as input. </p>\n<p>I wonder how to create it in Python. Did you just use to_pickle to convert the csv data to pkl.gzi data? Does that require a lot of memory as well?</p>",
      "rawMarkdown": "Yeah, it seems that you read the data riiid_train.pkl.gzi as input. \n\nI wonder how to create it in Python. Did you just use to_pickle to convert the csv data to pkl.gzi data? Does that require a lot of memory as well?",
      "votes": null
    },
    {
      "id": "1103285",
      "postDate": "12/05/2020 19:35:44",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/ifashion\" target=\"_blank\">@ifashion</a>.</p>\n<p>I did this conversion in this notebook <a href=\"https://www.kaggle.com/pedrocouto39/fast-reading-w-pickle-feather-parquet-jay\" target=\"_blank\">here</a>. It took a bit less memory than train.csv</p>",
      "rawMarkdown": "Hi, @ifashion.\n\nI did this conversion in this notebook [here](https://www.kaggle.com/pedrocouto39/fast-reading-w-pickle-feather-parquet-jay). It took a bit less memory than train.csv",
      "votes": null
    },
    {
      "id": "1103624",
      "postDate": "12/06/2020 04:45:25",
      "content": "<p>Thank you, I checked that out.</p>",
      "rawMarkdown": "Thank you, I checked that out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1099280,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "12/02/2020 08:30:35",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets#File-Formats\" target=\"_blank\">https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets#File-Formats</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1099981,
          "author_name": "ifashion",
          "author_url": "",
          "post_date": "12/02/2020 18:31:07",
          "content": "<p>Yeah, it seems that you read the data riiid_train.pkl.gzi as input. </p>\n<p>I wonder how to create it in Python. Did you just use to_pickle to convert the csv data to pkl.gzi data? Does that require a lot of memory as well?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103285,
      "author_name": "pedrocouto39",
      "author_url": "",
      "post_date": "12/05/2020 19:35:44",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/ifashion\" target=\"_blank\">@ifashion</a>.</p>\n<p>I did this conversion in this notebook <a href=\"https://www.kaggle.com/pedrocouto39/fast-reading-w-pickle-feather-parquet-jay\" target=\"_blank\">here</a>. It took a bit less memory than train.csv</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103624,
          "author_name": "ifashion",
          "author_url": "",
          "post_date": "12/06/2020 04:45:25",
          "content": "<p>Thank you, I checked that out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1099257": "Can anyone share the code of creating riiid_train.pkl.gzip? Thanks!",
    "1099280": "https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets#File-Formats",
    "1099981": "Yeah, it seems that you read the data riiid_train.pkl.gzi as input. \n\nI wonder how to create it in Python. Did you just use to_pickle to convert the csv data to pkl.gzi data? Does that require a lot of memory as well?",
    "1103285": "Hi, @ifashion.\n\nI did this conversion in this notebook [here](https://www.kaggle.com/pedrocouto39/fast-reading-w-pickle-feather-parquet-jay). It took a bit less memory than train.csv",
    "1103624": "Thank you, I checked that out."
  },
  "source": "meta"
}