{
  "id": 272531,
  "title": "Methods and Formats: Rapids, Dask, Datatable, Feather, HDF5, Jay, Parquet, Pickle",
  "url": "/competitions/wikipedia-image-caption/discussion/272531",
  "author_name": "",
  "post_date": "2021-09-15T23:53:08.279560600Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Reading Large Datasets:</p>\n<p>METHODS: Dask, Datatable, Pandas, Rapids.</p>\n<h1>Dask</h1>\n<p>Nain (aakashnain) made an excellent Kaggle Notebook:</p>\n<p><a href=\"https://www.kaggle.com/aakashnain/can-we-read-faster\" target=\"_blank\">https://www.kaggle.com/aakashnain/can-we-read-faster</a></p>\n<h1>Datatable</h1>\n<p>Udbhav Pangotra showed full with charm how to perform with Datatable:</p>\n<p><a href=\"https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm\" target=\"_blank\">https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm</a></p>\n<p>And Darek Kleczek applied Datatable in a crystal clear way<br>\n<a href=\"https://www.kaggle.com/thedrcat/wiki-image-caption-eda-and-baseline\" target=\"_blank\">https://www.kaggle.com/thedrcat/wiki-image-caption-eda-and-baseline</a></p>\n<p>However, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.</p>\n<p><a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html</a></p>\n<h1>Rapids</h1>\n<p>Turn on GPU and type import cudf and/or import cuml and Kaggle notebooks already have RAPIDS installed.  </p>\n<p>Though I got that memory error below.  Hence, I will wait till anybody works with Rapids.  <br>\nMemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:69: cudaErrorMemoryAllocation out of memory</p>\n<p>FORMATS: Feather, hdf5, Parquet, Pickle, Jay.</p>\n<h1>Feather</h1>\n<p>Store data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.</p>\n<p>Msafi04 made a feather Dataset to help us train the data.<br>\n<a href=\"https://www.kaggle.com/msafi04/train-tsv-file-to-feather-files\" target=\"_blank\">https://www.kaggle.com/msafi04/train-tsv-file-to-feather-files</a></p>\n<h1>hdf5, Parquet, Pickle, Jay</h1>\n<p>All these formats above require that somebody make a Dataset like Vopani did in this Brilliant Notebook below:<br>\nCode by Vopani <a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets</a></p>\n<p>Format: hdf5</p>\n<p>\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"</p>\n<p>Format: Jay <br>\n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"</p>\n<p>Format: Parquet<br>\n\"Parquet now extensively used with Spark.\"</p>\n<p>Gaurav Rawat made a very interessant \"10_folds_stratified_parquet&amp;feather\"  in tabular September 2021.<br>\n<a href=\"https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather\" target=\"_blank\">https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather</a></p>\n<p>Format: Pickle<br>\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"</p>\n<p>Thanks to Vopani for having inpired this topic:<br>\n<a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets</a></p>\n<p>And all the others that are sharing their work in this competition.</p>",
  "messages": [
    {
      "id": "1514308",
      "postDate": "09/15/2021 23:53:08",
      "content": "<p>Reading Large Datasets:</p>\n<p>METHODS: Dask, Datatable, Pandas, Rapids.</p>\n<h1>Dask</h1>\n<p>Nain (aakashnain) made an excellent Kaggle Notebook:</p>\n<p><a href=\"https://www.kaggle.com/aakashnain/can-we-read-faster\" target=\"_blank\">https://www.kaggle.com/aakashnain/can-we-read-faster</a></p>\n<h1>Datatable</h1>\n<p>Udbhav Pangotra showed full with charm how to perform with Datatable:</p>\n<p><a href=\"https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm\" target=\"_blank\">https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm</a></p>\n<p>And Darek Kleczek applied Datatable in a crystal clear way<br>\n<a href=\"https://www.kaggle.com/thedrcat/wiki-image-caption-eda-and-baseline\" target=\"_blank\">https://www.kaggle.com/thedrcat/wiki-image-caption-eda-and-baseline</a></p>\n<p>However, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.</p>\n<p><a href=\"https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\" target=\"_blank\">https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html</a></p>\n<h1>Rapids</h1>\n<p>Turn on GPU and type import cudf and/or import cuml and Kaggle notebooks already have RAPIDS installed.  </p>\n<p>Though I got that memory error below.  Hence, I will wait till anybody works with Rapids.  <br>\nMemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:69: cudaErrorMemoryAllocation out of memory</p>\n<p>FORMATS: Feather, hdf5, Parquet, Pickle, Jay.</p>\n<h1>Feather</h1>\n<p>Store data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.</p>\n<p>Msafi04 made a feather Dataset to help us train the data.<br>\n<a href=\"https://www.kaggle.com/msafi04/train-tsv-file-to-feather-files\" target=\"_blank\">https://www.kaggle.com/msafi04/train-tsv-file-to-feather-files</a></p>\n<h1>hdf5, Parquet, Pickle, Jay</h1>\n<p>All these formats above require that somebody make a Dataset like Vopani did in this Brilliant Notebook below:<br>\nCode by Vopani <a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets</a></p>\n<p>Format: hdf5</p>\n<p>\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"</p>\n<p>Format: Jay <br>\n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"</p>\n<p>Format: Parquet<br>\n\"Parquet now extensively used with Spark.\"</p>\n<p>Gaurav Rawat made a very interessant \"10_folds_stratified_parquet&amp;feather\"  in tabular September 2021.<br>\n<a href=\"https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather\" target=\"_blank\">https://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather</a></p>\n<p>Format: Pickle<br>\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"</p>\n<p>Thanks to Vopani for having inpired this topic:<br>\n<a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets</a></p>\n<p>And all the others that are sharing their work in this competition.</p>",
      "rawMarkdown": "Reading Large Datasets:\n\nMETHODS: Dask, Datatable, Pandas, Rapids.\n\n#Dask\nNain (aakashnain) made an excellent Kaggle Notebook:\n\nhttps://www.kaggle.com/aakashnain/can-we-read-faster\n\n#Datatable\nUdbhav Pangotra showed full with charm how to perform with Datatable:\n\nhttps://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm\n\nAnd Darek Kleczek applied Datatable in a crystal clear way\nhttps://www.kaggle.com/thedrcat/wiki-image-caption-eda-and-baseline\n\nHowever, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.\n\nhttps://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\n\n#Rapids\n\nTurn on GPU and type import cudf and/or import cuml and Kaggle notebooks already have RAPIDS installed.  \n\nThough I got that memory error below.  Hence, I will wait till anybody works with Rapids.  \nMemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:69: cudaErrorMemoryAllocation out of memory\n\nFORMATS: Feather, hdf5, Parquet, Pickle, Jay.\n\n#Feather\n\nStore data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.\n\nMsafi04 made a feather Dataset to help us train the data.\nhttps://www.kaggle.com/msafi04/train-tsv-file-to-feather-files\n\n#hdf5, Parquet, Pickle, Jay\n\nAll these formats above require that somebody make a Dataset like Vopani did in this Brilliant Notebook below:\nCode by Vopani https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\n\nFormat: hdf5\n\n\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"\n\nFormat: Jay \n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"\n\nFormat: Parquet\n\"Parquet now extensively used with Spark.\"\n\nGaurav Rawat made a very interessant \"10_folds_stratified_parquet&feather\"  in tabular September 2021.\nhttps://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather\n\nFormat: Pickle\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"\n\nThanks to Vopani for having inpired this topic:\nhttps://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\n\nAnd all the others that are sharing their work in this competition.",
      "votes": null
    },
    {
      "id": "1514351",
      "postDate": "09/16/2021 02:25:38",
      "content": "<p>Thank you for sharing !<br>\nKaggle is also using these methods more and more when using huge data.<br>\nI will be studying these.</p>",
      "rawMarkdown": "Thank you for sharing !\nKaggle is also using these methods more and more when using huge data.\nI will be studying these.",
      "votes": null
    },
    {
      "id": "1514355",
      "postDate": "09/16/2021 02:32:35",
      "content": "<p>Thank you for answering and the support too Crinoid.</p>",
      "rawMarkdown": "Thank you for answering and the support too Crinoid.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1514351,
      "author_name": "takemi",
      "author_url": "",
      "post_date": "09/16/2021 02:25:38",
      "content": "<p>Thank you for sharing !<br>\nKaggle is also using these methods more and more when using huge data.<br>\nI will be studying these.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1514355,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/16/2021 02:32:35",
          "content": "<p>Thank you for answering and the support too Crinoid.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1514308": "Reading Large Datasets:\n\nMETHODS: Dask, Datatable, Pandas, Rapids.\n\n#Dask\nNain (aakashnain) made an excellent Kaggle Notebook:\n\nhttps://www.kaggle.com/aakashnain/can-we-read-faster\n\n#Datatable\nUdbhav Pangotra showed full with charm how to perform with Datatable:\n\nhttps://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm\n\nAnd Darek Kleczek applied Datatable in a crystal clear way\nhttps://www.kaggle.com/thedrcat/wiki-image-caption-eda-and-baseline\n\nHowever, Datatable has some missing functionalities. Which means that there are some functions in pandas that do not have an equivalent in datatable yet, and are likely to be implemented.\n\nhttps://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html\n\n#Rapids\n\nTurn on GPU and type import cudf and/or import cuml and Kaggle notebooks already have RAPIDS installed.  \n\nThough I got that memory error below.  Hence, I will wait till anybody works with Rapids.  \nMemoryError: std::bad_alloc: CUDA error at: /opt/conda/include/rmm/mr/device/cuda_memory_resource.hpp:69: cudaErrorMemoryAllocation out of memory\n\nFORMATS: Feather, hdf5, Parquet, Pickle, Jay.\n\n#Feather\n\nStore data in feather (binary) format specifically for pandas. It significantly improves reading speed of datasets.\n\nMsafi04 made a feather Dataset to help us train the data.\nhttps://www.kaggle.com/msafi04/train-tsv-file-to-feather-files\n\n#hdf5, Parquet, Pickle, Jay\n\nAll these formats above require that somebody make a Dataset like Vopani did in this Brilliant Notebook below:\nCode by Vopani https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\n\nFormat: hdf5\n\n\"HDF5 is a high-performance data management suite to store, manage and process large and complex data.\"\n\nFormat: Jay \n\"Datatable uses .jay (binary) format which makes reading datasets blazing fast.\"\n\nFormat: Parquet\n\"Parquet now extensively used with Spark.\"\n\nGaurav Rawat made a very interessant \"10_folds_stratified_parquet&feather\"  in tabular September 2021.\nhttps://www.kaggle.com/gauravbrills/10-folds-stratified-parquet-feather\n\nFormat: Pickle\n\"Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\"\n\nThanks to Vopani for having inpired this topic:\nhttps://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\n\nAnd all the others that are sharing their work in this competition.",
    "1514351": "Thank you for sharing !\nKaggle is also using these methods more and more when using huge data.\nI will be studying these.",
    "1514355": "Thank you for answering and the support too Crinoid."
  },
  "source": "meta"
}