{
  "id": 341633,
  "title": "how to load data and process data",
  "url": "/competitions/amex-default-prediction/discussion/341633",
  "author_name": "",
  "post_date": "2022-08-03T17:37:18.926722600Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>i copy and edit from Amex LGBM Dart CV 0.7977 baseline , the first pargraph write  pd.read_parquet, but the cpu is too small to load data, how did it work?   </p>\n<p>i try to dd.read_csv(data)  and to_parquet , but  the  for k,v in dask.dataframe .groupby()  false   and the cpu can't solve the dask.compute() to pandas  even solve  I consider the next step will go wrong,    so how to do with the data?</p>",
  "messages": [
    {
      "id": "1883273",
      "postDate": "08/03/2022 17:37:18",
      "content": "<p>i copy and edit from Amex LGBM Dart CV 0.7977 baseline , the first pargraph write  pd.read_parquet, but the cpu is too small to load data, how did it work?   </p>\n<p>i try to dd.read_csv(data)  and to_parquet , but  the  for k,v in dask.dataframe .groupby()  false   and the cpu can't solve the dask.compute() to pandas  even solve  I consider the next step will go wrong,    so how to do with the data?</p>",
      "rawMarkdown": "i copy and edit from Amex LGBM Dart CV 0.7977 baseline , the first pargraph write  pd.read_parquet, but the cpu is too small to load data, how did it work?   \n\ni try to dd.read_csv(data)  and to_parquet , but  the  for k,v in dask.dataframe .groupby()  false   and the cpu can't solve the dask.compute() to pandas  even solve  I consider the next step will go wrong,    so how to do with the data?",
      "votes": null
    },
    {
      "id": "1883301",
      "postDate": "08/03/2022 18:04:36",
      "content": "<p>You should use Raddar's parquet file. The entire train data is only 1.6GB and will easily fit in 16GB of memory. Discussion <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>, dataset <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "You should use Raddar's parquet file. The entire train data is only 1.6GB and will easily fit in 16GB of memory. Discussion [here][1], dataset [here][2]\n\n[1]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n[2]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514",
      "votes": null
    },
    {
      "id": "1883789",
      "postDate": "08/04/2022 05:26:19",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1883301,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/03/2022 18:04:36",
      "content": "<p>You should use Raddar's parquet file. The entire train data is only 1.6GB and will easily fit in 16GB of memory. Discussion <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>, dataset <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1883789,
          "author_name": "wcqglhf",
          "author_url": "",
          "post_date": "08/04/2022 05:26:19",
          "content": "<p>thanks a lot</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1883273": "i copy and edit from Amex LGBM Dart CV 0.7977 baseline , the first pargraph write  pd.read_parquet, but the cpu is too small to load data, how did it work?   \n\ni try to dd.read_csv(data)  and to_parquet , but  the  for k,v in dask.dataframe .groupby()  false   and the cpu can't solve the dask.compute() to pandas  even solve  I consider the next step will go wrong,    so how to do with the data?",
    "1883301": "You should use Raddar's parquet file. The entire train data is only 1.6GB and will easily fit in 16GB of memory. Discussion [here][1], dataset [here][2]\n\n[1]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n[2]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514",
    "1883789": "thanks a lot"
  },
  "source": "meta"
}