{
  "id": 327195,
  "title": "Reading & Working with Large Dataset",
  "url": "/competitions/amex-default-prediction/discussion/327195",
  "author_name": "Dev Khant",
  "post_date": "2022-05-26T04:41:39.502000",
  "votes": 30,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As this competition has large dataset simply using <strong><em>pd.read_csv</em></strong> will result into in an out-of-memory error on Kaggle Notebooks. So here we will see different methods and libraries to work with huge data.</p>\n<h2><strong>Libraries</strong></h2>\n<p>There are few libraries which can be used to read Large Dataset. Here we will see 3 such libraries.</p>\n<h3>1. Dask</h3>\n<p>Dask is an open-sourced Python library that provides multi-core and distributed parallel execution of larger-than-memory datasets. It provides parallelized NumPy array and Pandas DataFrame objects. And it also provides API similar to Pandas and Numpy.</p>\n<p><code>import dask.dataframe as dd</code><br>\n<code>df_dask = dd.read_csv(\"your_dataset.csv\")</code></p>\n<p>Documentation : <a href=\"https://docs.dask.org/en/latest/\" target=\"_blank\">https://docs.dask.org/en/latest/</a></p>\n<h3>2. Modin</h3>\n<p>Modin uses Ray or Dask to provide an effortless way to speed up your pandas notebooks and utilizes all the cores available in the system, only requiring users to change a single line of code in their notebooks.</p>\n<p><code>import modin.pandas as md</code><br>\n<code>modin_df = pd.read_csv(\"your_dataset.csv\")</code></p>\n<p>Documentation : <a href=\"https://modin.readthedocs.io/en/stable/\" target=\"_blank\">https://modin.readthedocs.io/en/stable/</a></p>\n<h3>3. Rapids</h3>\n<p>The RAPIDS data science framework is a collection of libraries for executing end-to-end data science pipelines completely in the GPU. It includes libraries like CuML, CuDF, Xgboost etc…</p>\n<p><code>import cudf</code><br>\n<code>df = cudf.read_csv('your_dataset.csv')</code></p>\n<p>Documentation : <a href=\"https://rapids.ai/start.html\" target=\"_blank\">https://rapids.ai/start.html</a></p>\n<h2><strong>Methods</strong></h2>\n<p>We can convert our csv to other formats for faster loading.</p>\n<h3>1. Parquet</h3>\n<p>In the Hadoop ecosystem, parquet was popularly used as the primary file format for tabular datasets and is now extensively used with Spark. It has become more available and efficient over the years and is also supported by pandas.</p>\n<p>In order to convert csv file to parquet file simply do<br>\n<code>df.to_parquet('output.parquet')</code></p>\n<p>To read parquet file in pandas<br>\n<code>data = pd.read_parquet(\"your_dataset.parquet\")</code></p>\n<p>Documentation : <a href=\"https://parquet.apache.org/\" target=\"_blank\">https://parquet.apache.org/</a></p>\n<h3>2. Pickle</h3>\n<p>Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.</p>\n<p>In order to convert csv file to pickle file simply do<br>\n<code>df.to_parquet('output.pkl')</code></p>\n<p>To read pickle file in pandas<br>\n<code>data = pd.read_parquet(\"your_dataset.pkl\")</code></p>\n<p><strong>I hope you learnt something new. So keep Learning &amp; keep Kaggling :)</strong></p>",
  "messages": [
    {
      "id": 1801715,
      "postDate": "2022-05-26T04:41:39.503Z",
      "content": "<p>As this competition has large dataset simply using <strong><em>pd.read_csv</em></strong> will result into in an out-of-memory error on Kaggle Notebooks. So here we will see different methods and libraries to work with huge data.</p>\n<h2><strong>Libraries</strong></h2>\n<p>There are few libraries which can be used to read Large Dataset. Here we will see 3 such libraries.</p>\n<h3>1. Dask</h3>\n<p>Dask is an open-sourced Python library that provides multi-core and distributed parallel execution of larger-than-memory datasets. It provides parallelized NumPy array and Pandas DataFrame objects. And it also provides API similar to Pandas and Numpy.</p>\n<p><code>import dask.dataframe as dd</code><br>\n<code>df_dask = dd.read_csv(\"your_dataset.csv\")</code></p>\n<p>Documentation : <a href=\"https://docs.dask.org/en/latest/\" target=\"_blank\">https://docs.dask.org/en/latest/</a></p>\n<h3>2. Modin</h3>\n<p>Modin uses Ray or Dask to provide an effortless way to speed up your pandas notebooks and utilizes all the cores available in the system, only requiring users to change a single line of code in their notebooks.</p>\n<p><code>import modin.pandas as md</code><br>\n<code>modin_df = pd.read_csv(\"your_dataset.csv\")</code></p>\n<p>Documentation : <a href=\"https://modin.readthedocs.io/en/stable/\" target=\"_blank\">https://modin.readthedocs.io/en/stable/</a></p>\n<h3>3. Rapids</h3>\n<p>The RAPIDS data science framework is a collection of libraries for executing end-to-end data science pipelines completely in the GPU. It includes libraries like CuML, CuDF, Xgboost etc…</p>\n<p><code>import cudf</code><br>\n<code>df = cudf.read_csv('your_dataset.csv')</code></p>\n<p>Documentation : <a href=\"https://rapids.ai/start.html\" target=\"_blank\">https://rapids.ai/start.html</a></p>\n<h2><strong>Methods</strong></h2>\n<p>We can convert our csv to other formats for faster loading.</p>\n<h3>1. Parquet</h3>\n<p>In the Hadoop ecosystem, parquet was popularly used as the primary file format for tabular datasets and is now extensively used with Spark. It has become more available and efficient over the years and is also supported by pandas.</p>\n<p>In order to convert csv file to parquet file simply do<br>\n<code>df.to_parquet('output.parquet')</code></p>\n<p>To read parquet file in pandas<br>\n<code>data = pd.read_parquet(\"your_dataset.parquet\")</code></p>\n<p>Documentation : <a href=\"https://parquet.apache.org/\" target=\"_blank\">https://parquet.apache.org/</a></p>\n<h3>2. Pickle</h3>\n<p>Python objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.</p>\n<p>In order to convert csv file to pickle file simply do<br>\n<code>df.to_parquet('output.pkl')</code></p>\n<p>To read pickle file in pandas<br>\n<code>data = pd.read_parquet(\"your_dataset.pkl\")</code></p>\n<p><strong>I hope you learnt something new. So keep Learning &amp; keep Kaggling :)</strong></p>",
      "rawMarkdown": "As this competition has large dataset simply using ***pd.read_csv*** will result into in an out-of-memory error on Kaggle Notebooks. So here we will see different methods and libraries to work with huge data.\n\n## **Libraries**\nThere are few libraries which can be used to read Large Dataset. Here we will see 3 such libraries.\n\n### 1. Dask\nDask is an open-sourced Python library that provides multi-core and distributed parallel execution of larger-than-memory datasets. It provides parallelized NumPy array and Pandas DataFrame objects. And it also provides API similar to Pandas and Numpy.\n\n`import dask.dataframe as dd`\n`df_dask = dd.read_csv(\"your_dataset.csv\")`\n\nDocumentation : https://docs.dask.org/en/latest/\n\n### 2. Modin\nModin uses Ray or Dask to provide an effortless way to speed up your pandas notebooks and utilizes all the cores available in the system, only requiring users to change a single line of code in their notebooks.\n\n`import modin.pandas as md`\n`modin_df = pd.read_csv(\"your_dataset.csv\")`\n\nDocumentation : https://modin.readthedocs.io/en/stable/\n\n### 3. Rapids\nThe RAPIDS data science framework is a collection of libraries for executing end-to-end data science pipelines completely in the GPU. It includes libraries like CuML, CuDF, Xgboost etc...\n\n`import cudf`\n`df = cudf.read_csv('your_dataset.csv')`\n\nDocumentation : https://rapids.ai/start.html\n\n## **Methods**\nWe can convert our csv to other formats for faster loading.\n\n### 1. Parquet\nIn the Hadoop ecosystem, parquet was popularly used as the primary file format for tabular datasets and is now extensively used with Spark. It has become more available and efficient over the years and is also supported by pandas.\n\nIn order to convert csv file to parquet file simply do\n`df.to_parquet('output.parquet')`\n\nTo read parquet file in pandas\n`data = pd.read_parquet(\"your_dataset.parquet\")`\n\nDocumentation : https://parquet.apache.org/\n\n### 2. Pickle\nPython objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\n\nIn order to convert csv file to pickle file simply do\n`df.to_parquet('output.pkl')`\n\nTo read pickle file in pandas\n`data = pd.read_parquet(\"your_dataset.pkl\")`\n\n\n\n**I hope you learnt something new. So keep Learning & keep Kaggling :)**",
      "votes": 30
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1801715": "As this competition has large dataset simply using ***pd.read_csv*** will result into in an out-of-memory error on Kaggle Notebooks. So here we will see different methods and libraries to work with huge data.\n\n## **Libraries**\nThere are few libraries which can be used to read Large Dataset. Here we will see 3 such libraries.\n\n### 1. Dask\nDask is an open-sourced Python library that provides multi-core and distributed parallel execution of larger-than-memory datasets. It provides parallelized NumPy array and Pandas DataFrame objects. And it also provides API similar to Pandas and Numpy.\n\n`import dask.dataframe as dd`\n`df_dask = dd.read_csv(\"your_dataset.csv\")`\n\nDocumentation : https://docs.dask.org/en/latest/\n\n### 2. Modin\nModin uses Ray or Dask to provide an effortless way to speed up your pandas notebooks and utilizes all the cores available in the system, only requiring users to change a single line of code in their notebooks.\n\n`import modin.pandas as md`\n`modin_df = pd.read_csv(\"your_dataset.csv\")`\n\nDocumentation : https://modin.readthedocs.io/en/stable/\n\n### 3. Rapids\nThe RAPIDS data science framework is a collection of libraries for executing end-to-end data science pipelines completely in the GPU. It includes libraries like CuML, CuDF, Xgboost etc...\n\n`import cudf`\n`df = cudf.read_csv('your_dataset.csv')`\n\nDocumentation : https://rapids.ai/start.html\n\n## **Methods**\nWe can convert our csv to other formats for faster loading.\n\n### 1. Parquet\nIn the Hadoop ecosystem, parquet was popularly used as the primary file format for tabular datasets and is now extensively used with Spark. It has become more available and efficient over the years and is also supported by pandas.\n\nIn order to convert csv file to parquet file simply do\n`df.to_parquet('output.parquet')`\n\nTo read parquet file in pandas\n`data = pd.read_parquet(\"your_dataset.parquet\")`\n\nDocumentation : https://parquet.apache.org/\n\n### 2. Pickle\nPython objects can be stored in the form of pickle files and pandas has inbuilt functions to read and write dataframes as pickle objects.\n\nIn order to convert csv file to pickle file simply do\n`df.to_parquet('output.pkl')`\n\nTo read pickle file in pandas\n`data = pd.read_parquet(\"your_dataset.pkl\")`\n\n\n\n**I hope you learnt something new. So keep Learning & keep Kaggling :)**"
  }
}