{
  "id": 356544,
  "title": "Read your data faster (Wall time: 12.7)",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356544",
  "author_name": "",
  "post_date": "2022-10-01T01:29:38.028706300Z",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Dear kagglers,</p>\n<p>There are many tools to read tabular data faster than by using pandas.<br>\nIf you don't want to spend the time to run it all over, just use this code snippet.</p>\n<pre><code>!pip install datatable\nimport datatable as dt\nimport pandas as pd\n\ndtypes_dict = {\n    'game_num': 'int16', 'event_id': 'int32', 'event_time': 'float16',\n    'ball_pos_x': 'float16', 'ball_pos_y': 'float16', 'ball_pos_z': 'float16',\n    'ball_vel_x': 'float16', 'ball_vel_y': 'float16', 'ball_vel_z': 'float16',\n    'p0_pos_x': 'float16', 'p0_pos_y': 'float16', 'p0_pos_z': 'float16',\n    'p0_vel_x': 'float16', 'p0_vel_y': 'float16', 'p0_vel_z': 'float16',\n    'p0_boost': 'float16', 'p1_pos_x': 'float16', 'p1_pos_y': 'float16',\n    'p1_pos_z': 'float16', 'p1_vel_x': 'float16', 'p1_vel_y': 'float16',\n    'p1_vel_z': 'float16', 'p1_boost': 'float16', 'p2_pos_x': 'float16',\n    'p2_pos_y': 'float16', 'p2_pos_z': 'float16', 'p2_vel_x': 'float16',\n    'p2_vel_y': 'float16', 'p2_vel_z': 'float16', 'p2_boost': 'float16',\n    'p3_pos_x': 'float16', 'p3_pos_y': 'float16', 'p3_pos_z': 'float16',\n    'p3_vel_x': 'float16', 'p3_vel_y': 'float16', 'p3_vel_z': 'float16',\n    'p3_boost': 'float16', 'p4_pos_x': 'float16', 'p4_pos_y': 'float16',\n    'p4_pos_z': 'float16', 'p4_vel_x': 'float16', 'p4_vel_y': 'float16',\n    'p4_vel_z': 'float16', 'p4_boost': 'float16', 'p5_pos_x': 'float16',\n    'p5_pos_y': 'float16', 'p5_pos_z': 'float16', 'p5_vel_x': 'float16',\n    'p5_vel_y': 'float16', 'p5_vel_z': 'float16', 'p5_boost': 'float16',\n    'boost0_timer': 'float16', 'boost1_timer': 'float16', 'boost2_timer': 'float16',\n    'boost3_timer': 'float16', 'boost4_timer': 'float16', 'boost5_timer': 'float16',\n    'player_scoring_next': 'O', 'team_scoring_next': 'O', 'team_A_scoring_within_10sec': 'O',\n    'team_B_scoring_within_10sec': 'O'\n}\n\n\npath_to_data = 'YOUR_FOLDER_WITH_TRAIN_FILES'\ndf = pd.DataFrame({}, columns=dtypes_dict.keys())\nfor i in range(10):\n    dt_read = dt.fread(f'{path_to_data}/train_{i}.csv').to_pandas()\n    dt_read = dt_read.astype(dtypes_dict)\n    df = pd.concat([df, dt_read])\n</code></pre>\n<p>Wall time: 12.6 s</p>\n<p>P.s. I have added <a href=\"https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet\" target=\"_blank\">the dataset with explicit dtypes csv files</a> and main parquet dataset.</p>\n<pre><code>train = dt.fread('../input/tabular-playground-series-oct-2022/train_0.csv').to_pandas()\ndtypes_dict_train = dict(pd.read_csv('../input/tps2022octparquet/dtypes_train.csv').values)\n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\n</code></pre>",
  "messages": [
    {
      "id": "1964895",
      "postDate": "10/01/2022 01:29:38",
      "content": "<p>Dear kagglers,</p>\n<p>There are many tools to read tabular data faster than by using pandas.<br>\nIf you don't want to spend the time to run it all over, just use this code snippet.</p>\n<pre><code>!pip install datatable\nimport datatable as dt\nimport pandas as pd\n\ndtypes_dict = {\n    'game_num': 'int16', 'event_id': 'int32', 'event_time': 'float16',\n    'ball_pos_x': 'float16', 'ball_pos_y': 'float16', 'ball_pos_z': 'float16',\n    'ball_vel_x': 'float16', 'ball_vel_y': 'float16', 'ball_vel_z': 'float16',\n    'p0_pos_x': 'float16', 'p0_pos_y': 'float16', 'p0_pos_z': 'float16',\n    'p0_vel_x': 'float16', 'p0_vel_y': 'float16', 'p0_vel_z': 'float16',\n    'p0_boost': 'float16', 'p1_pos_x': 'float16', 'p1_pos_y': 'float16',\n    'p1_pos_z': 'float16', 'p1_vel_x': 'float16', 'p1_vel_y': 'float16',\n    'p1_vel_z': 'float16', 'p1_boost': 'float16', 'p2_pos_x': 'float16',\n    'p2_pos_y': 'float16', 'p2_pos_z': 'float16', 'p2_vel_x': 'float16',\n    'p2_vel_y': 'float16', 'p2_vel_z': 'float16', 'p2_boost': 'float16',\n    'p3_pos_x': 'float16', 'p3_pos_y': 'float16', 'p3_pos_z': 'float16',\n    'p3_vel_x': 'float16', 'p3_vel_y': 'float16', 'p3_vel_z': 'float16',\n    'p3_boost': 'float16', 'p4_pos_x': 'float16', 'p4_pos_y': 'float16',\n    'p4_pos_z': 'float16', 'p4_vel_x': 'float16', 'p4_vel_y': 'float16',\n    'p4_vel_z': 'float16', 'p4_boost': 'float16', 'p5_pos_x': 'float16',\n    'p5_pos_y': 'float16', 'p5_pos_z': 'float16', 'p5_vel_x': 'float16',\n    'p5_vel_y': 'float16', 'p5_vel_z': 'float16', 'p5_boost': 'float16',\n    'boost0_timer': 'float16', 'boost1_timer': 'float16', 'boost2_timer': 'float16',\n    'boost3_timer': 'float16', 'boost4_timer': 'float16', 'boost5_timer': 'float16',\n    'player_scoring_next': 'O', 'team_scoring_next': 'O', 'team_A_scoring_within_10sec': 'O',\n    'team_B_scoring_within_10sec': 'O'\n}\n\n\npath_to_data = 'YOUR_FOLDER_WITH_TRAIN_FILES'\ndf = pd.DataFrame({}, columns=dtypes_dict.keys())\nfor i in range(10):\n    dt_read = dt.fread(f'{path_to_data}/train_{i}.csv').to_pandas()\n    dt_read = dt_read.astype(dtypes_dict)\n    df = pd.concat([df, dt_read])\n</code></pre>\n<p>Wall time: 12.6 s</p>\n<p>P.s. I have added <a href=\"https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet\" target=\"_blank\">the dataset with explicit dtypes csv files</a> and main parquet dataset.</p>\n<pre><code>train = dt.fread('../input/tabular-playground-series-oct-2022/train_0.csv').to_pandas()\ndtypes_dict_train = dict(pd.read_csv('../input/tps2022octparquet/dtypes_train.csv').values)\n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\n</code></pre>",
      "rawMarkdown": "Dear kagglers,\n\nThere are many tools to read tabular data faster than by using pandas.\nIf you don't want to spend the time to run it all over, just use this code snippet.\n\n```\n!pip install datatable\nimport datatable as dt\nimport pandas as pd\n\ndtypes_dict = {\n    'game_num': 'int16', 'event_id': 'int32', 'event_time': 'float16',\n    'ball_pos_x': 'float16', 'ball_pos_y': 'float16', 'ball_pos_z': 'float16',\n    'ball_vel_x': 'float16', 'ball_vel_y': 'float16', 'ball_vel_z': 'float16',\n    'p0_pos_x': 'float16', 'p0_pos_y': 'float16', 'p0_pos_z': 'float16',\n    'p0_vel_x': 'float16', 'p0_vel_y': 'float16', 'p0_vel_z': 'float16',\n    'p0_boost': 'float16', 'p1_pos_x': 'float16', 'p1_pos_y': 'float16',\n    'p1_pos_z': 'float16', 'p1_vel_x': 'float16', 'p1_vel_y': 'float16',\n    'p1_vel_z': 'float16', 'p1_boost': 'float16', 'p2_pos_x': 'float16',\n    'p2_pos_y': 'float16', 'p2_pos_z': 'float16', 'p2_vel_x': 'float16',\n    'p2_vel_y': 'float16', 'p2_vel_z': 'float16', 'p2_boost': 'float16',\n    'p3_pos_x': 'float16', 'p3_pos_y': 'float16', 'p3_pos_z': 'float16',\n    'p3_vel_x': 'float16', 'p3_vel_y': 'float16', 'p3_vel_z': 'float16',\n    'p3_boost': 'float16', 'p4_pos_x': 'float16', 'p4_pos_y': 'float16',\n    'p4_pos_z': 'float16', 'p4_vel_x': 'float16', 'p4_vel_y': 'float16',\n    'p4_vel_z': 'float16', 'p4_boost': 'float16', 'p5_pos_x': 'float16',\n    'p5_pos_y': 'float16', 'p5_pos_z': 'float16', 'p5_vel_x': 'float16',\n    'p5_vel_y': 'float16', 'p5_vel_z': 'float16', 'p5_boost': 'float16',\n    'boost0_timer': 'float16', 'boost1_timer': 'float16', 'boost2_timer': 'float16',\n    'boost3_timer': 'float16', 'boost4_timer': 'float16', 'boost5_timer': 'float16',\n    'player_scoring_next': 'O', 'team_scoring_next': 'O', 'team_A_scoring_within_10sec': 'O',\n    'team_B_scoring_within_10sec': 'O'\n}\n\n\npath_to_data = 'YOUR_FOLDER_WITH_TRAIN_FILES'\ndf = pd.DataFrame({}, columns=dtypes_dict.keys())\nfor i in range(10):\n    dt_read = dt.fread(f'{path_to_data}/train_{i}.csv').to_pandas()\n    dt_read = dt_read.astype(dtypes_dict)\n    df = pd.concat([df, dt_read])\n```\n\nWall time: 12.6 s\n\nP.s. I have added [the dataset with explicit dtypes csv files](https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet) and main parquet dataset.\n\n```\ntrain = dt.fread('../input/tabular-playground-series-oct-2022/train_0.csv').to_pandas()\ndtypes_dict_train = dict(pd.read_csv('../input/tps2022octparquet/dtypes_train.csv').values)\n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\n```",
      "votes": null
    },
    {
      "id": "1964907",
      "postDate": "10/01/2022 01:54:08",
      "content": "<p>I also used datatable <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. Thanks for sharing you code!</p>",
      "rawMarkdown": "I also used datatable @sergiosaharovskiy. Thanks for sharing you code!",
      "votes": null
    },
    {
      "id": "1964910",
      "postDate": "10/01/2022 02:00:07",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> just some fixes on the code. <code>columns</code> doesn't exist but can be replaced for <code>dtypes_dict.keys()</code>. Second it should be <code>range(10)</code></p>",
      "rawMarkdown": "Hey @sergiosaharovskiy just some fixes on the code. `columns` doesn't exist but can be replaced for `dtypes_dict.keys()`. Second it should be `range(10)`",
      "votes": null
    },
    {
      "id": "1964921",
      "postDate": "10/01/2022 02:13:39",
      "content": "<p><a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> Fixed, thank you for quickly pointing out.</p>",
      "rawMarkdown": "jcaliz Fixed, thank you for quickly pointing out.",
      "votes": null
    },
    {
      "id": "1964934",
      "postDate": "10/01/2022 02:31:46",
      "content": "<p>Well that makes sense.<br>\nAbout pandas, I'm talking about this line <code>df = pd.DataFrame({}, columns=columns)</code> the parameter exists but the variable does not.</p>",
      "rawMarkdown": "Well that makes sense.\nAbout pandas, I'm talking about this line `df = pd.DataFrame({}, columns=columns)` the parameter exists but the variable does not.",
      "votes": null
    },
    {
      "id": "1969255",
      "postDate": "10/03/2022 11:58:07",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <br>\nVery helpful! Thank you for sharing that code!</p>",
      "rawMarkdown": "sergiosaharovskiy \nVery helpful! Thank you for sharing that code!",
      "votes": null
    },
    {
      "id": "1981891",
      "postDate": "10/11/2022 06:05:47",
      "content": "<p>So saving the data (all 10 files) as parquet reduces the size of the data by 70%?</p>",
      "rawMarkdown": "So saving the data (all 10 files) as parquet reduces the size of the data by 70%?",
      "votes": null
    },
    {
      "id": "1982484",
      "postDate": "10/11/2022 12:57:35",
      "content": "<p><a href=\"https://www.kaggle.com/osopova\" target=\"_blank\">@osopova</a> casting lower int and float dtypes reduces the data size (both memory and storage). The parquet format just compresses the data and you have less storage space occupied on your disk.</p>",
      "rawMarkdown": "osopova casting lower int and float dtypes reduces the data size (both memory and storage). The parquet format just compresses the data and you have less storage space occupied on your disk.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1964907,
      "author_name": "oscarm524",
      "author_url": "",
      "post_date": "10/01/2022 01:54:08",
      "content": "<p>I also used datatable <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. Thanks for sharing you code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1964910,
      "author_name": "jcaliz",
      "author_url": "",
      "post_date": "10/01/2022 02:00:07",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> just some fixes on the code. <code>columns</code> doesn't exist but can be replaced for <code>dtypes_dict.keys()</code>. Second it should be <code>range(10)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1964921,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "10/01/2022 02:13:39",
          "content": "<p><a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> Fixed, thank you for quickly pointing out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1964934,
          "author_name": "jcaliz",
          "author_url": "",
          "post_date": "10/01/2022 02:31:46",
          "content": "<p>Well that makes sense.<br>\nAbout pandas, I'm talking about this line <code>df = pd.DataFrame({}, columns=columns)</code> the parameter exists but the variable does not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1969255,
      "author_name": "imnaho",
      "author_url": "",
      "post_date": "10/03/2022 11:58:07",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <br>\nVery helpful! Thank you for sharing that code!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1981891,
      "author_name": "osopova",
      "author_url": "",
      "post_date": "10/11/2022 06:05:47",
      "content": "<p>So saving the data (all 10 files) as parquet reduces the size of the data by 70%?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1982484,
          "author_name": "sergiosaharovskiy",
          "author_url": "",
          "post_date": "10/11/2022 12:57:35",
          "content": "<p><a href=\"https://www.kaggle.com/osopova\" target=\"_blank\">@osopova</a> casting lower int and float dtypes reduces the data size (both memory and storage). The parquet format just compresses the data and you have less storage space occupied on your disk.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1964895": "Dear kagglers,\n\nThere are many tools to read tabular data faster than by using pandas.\nIf you don't want to spend the time to run it all over, just use this code snippet.\n\n```\n!pip install datatable\nimport datatable as dt\nimport pandas as pd\n\ndtypes_dict = {\n    'game_num': 'int16', 'event_id': 'int32', 'event_time': 'float16',\n    'ball_pos_x': 'float16', 'ball_pos_y': 'float16', 'ball_pos_z': 'float16',\n    'ball_vel_x': 'float16', 'ball_vel_y': 'float16', 'ball_vel_z': 'float16',\n    'p0_pos_x': 'float16', 'p0_pos_y': 'float16', 'p0_pos_z': 'float16',\n    'p0_vel_x': 'float16', 'p0_vel_y': 'float16', 'p0_vel_z': 'float16',\n    'p0_boost': 'float16', 'p1_pos_x': 'float16', 'p1_pos_y': 'float16',\n    'p1_pos_z': 'float16', 'p1_vel_x': 'float16', 'p1_vel_y': 'float16',\n    'p1_vel_z': 'float16', 'p1_boost': 'float16', 'p2_pos_x': 'float16',\n    'p2_pos_y': 'float16', 'p2_pos_z': 'float16', 'p2_vel_x': 'float16',\n    'p2_vel_y': 'float16', 'p2_vel_z': 'float16', 'p2_boost': 'float16',\n    'p3_pos_x': 'float16', 'p3_pos_y': 'float16', 'p3_pos_z': 'float16',\n    'p3_vel_x': 'float16', 'p3_vel_y': 'float16', 'p3_vel_z': 'float16',\n    'p3_boost': 'float16', 'p4_pos_x': 'float16', 'p4_pos_y': 'float16',\n    'p4_pos_z': 'float16', 'p4_vel_x': 'float16', 'p4_vel_y': 'float16',\n    'p4_vel_z': 'float16', 'p4_boost': 'float16', 'p5_pos_x': 'float16',\n    'p5_pos_y': 'float16', 'p5_pos_z': 'float16', 'p5_vel_x': 'float16',\n    'p5_vel_y': 'float16', 'p5_vel_z': 'float16', 'p5_boost': 'float16',\n    'boost0_timer': 'float16', 'boost1_timer': 'float16', 'boost2_timer': 'float16',\n    'boost3_timer': 'float16', 'boost4_timer': 'float16', 'boost5_timer': 'float16',\n    'player_scoring_next': 'O', 'team_scoring_next': 'O', 'team_A_scoring_within_10sec': 'O',\n    'team_B_scoring_within_10sec': 'O'\n}\n\n\npath_to_data = 'YOUR_FOLDER_WITH_TRAIN_FILES'\ndf = pd.DataFrame({}, columns=dtypes_dict.keys())\nfor i in range(10):\n    dt_read = dt.fread(f'{path_to_data}/train_{i}.csv').to_pandas()\n    dt_read = dt_read.astype(dtypes_dict)\n    df = pd.concat([df, dt_read])\n```\n\nWall time: 12.6 s\n\nP.s. I have added [the dataset with explicit dtypes csv files](https://www.kaggle.com/datasets/sergiosaharovskiy/tps2022octparquet) and main parquet dataset.\n\n```\ntrain = dt.fread('../input/tabular-playground-series-oct-2022/train_0.csv').to_pandas()\ndtypes_dict_train = dict(pd.read_csv('../input/tps2022octparquet/dtypes_train.csv').values)\n# Reduce memory usage by 70%.\ntrain = train.astype(dtypes_dict_train)\n```",
    "1964907": "I also used datatable @sergiosaharovskiy. Thanks for sharing you code!",
    "1964910": "Hey @sergiosaharovskiy just some fixes on the code. `columns` doesn't exist but can be replaced for `dtypes_dict.keys()`. Second it should be `range(10)`",
    "1964921": "jcaliz Fixed, thank you for quickly pointing out.",
    "1964934": "Well that makes sense.\nAbout pandas, I'm talking about this line `df = pd.DataFrame({}, columns=columns)` the parameter exists but the variable does not.",
    "1969255": "sergiosaharovskiy \nVery helpful! Thank you for sharing that code!",
    "1981891": "So saving the data (all 10 files) as parquet reduces the size of the data by 70%?",
    "1982484": "osopova casting lower int and float dtypes reduces the data size (both memory and storage). The parquet format just compresses the data and you have less storage space occupied on your disk."
  },
  "source": "meta"
}