{
  "id": 327205,
  "title": "Tutorial on reading large datasets by Rohan",
  "url": "/competitions/amex-default-prediction/discussion/327205",
  "author_name": "",
  "post_date": "2022-05-26T06:43:32.762052800Z",
  "votes": 48,
  "comment_count": 12,
  "views": 0,
  "content": "<p>One of the biggest challenge in this competition for all of us is to read such a huge dataset. </p>\n<p>One of our KGMs Rohan has created an excellent tutorial on reading large datasets where he compared different methods and different file formats, which will be super helpful for this competition.</p>\n<p>Methods covered:</p>\n<ul>\n<li>Pandas</li>\n<li>Dask</li>\n<li>Datatable</li>\n<li>Rapids</li>\n</ul>\n<p>File Formats covered:</p>\n<ul>\n<li>csv</li>\n<li>feather</li>\n<li>hdf5</li>\n<li>jay</li>\n<li>parquet</li>\n<li>pickle</li>\n</ul>\n<p>Link to notebook - <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\" target=\"_blank\">https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook</a></p>\n<p>All the best!</p>",
  "messages": [
    {
      "id": "1801790",
      "postDate": "05/26/2022 06:43:32",
      "content": "<p>One of the biggest challenge in this competition for all of us is to read such a huge dataset. </p>\n<p>One of our KGMs Rohan has created an excellent tutorial on reading large datasets where he compared different methods and different file formats, which will be super helpful for this competition.</p>\n<p>Methods covered:</p>\n<ul>\n<li>Pandas</li>\n<li>Dask</li>\n<li>Datatable</li>\n<li>Rapids</li>\n</ul>\n<p>File Formats covered:</p>\n<ul>\n<li>csv</li>\n<li>feather</li>\n<li>hdf5</li>\n<li>jay</li>\n<li>parquet</li>\n<li>pickle</li>\n</ul>\n<p>Link to notebook - <a href=\"https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\" target=\"_blank\">https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook</a></p>\n<p>All the best!</p>",
      "rawMarkdown": "One of the biggest challenge in this competition for all of us is to read such a huge dataset. \n\nOne of our KGMs Rohan has created an excellent tutorial on reading large datasets where he compared different methods and different file formats, which will be super helpful for this competition.\n\nMethods covered:\n* Pandas\n* Dask\n* Datatable\n* Rapids\n\nFile Formats covered:\n* csv\n* feather\n* hdf5\n* jay\n* parquet\n* pickle\n\nLink to notebook - https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\n\nAll the best!",
      "votes": null
    },
    {
      "id": "1801795",
      "postDate": "05/26/2022 06:48:23",
      "content": "<p>Datatype dict for our competition which will be useful for reading data proposed in the above notebook.</p>\n<p>Converted all the <code>float64</code> to <code>float16</code>.</p>\n<pre><code>dtype_dict = {'customer_ID': \"object\",\n 'S_2': \"object\",\n 'P_2': 'float16',\n 'D_39': 'float16',\n 'B_1': 'float16',\n 'B_2': 'float16',\n 'R_1': 'float16',\n 'S_3': 'float16',\n 'D_41': 'float16',\n 'B_3': 'float16',\n 'D_42': 'float16',\n 'D_43': 'float16',\n 'D_44': 'float16',\n 'B_4': 'float16',\n 'D_45': 'float16',\n 'B_5': 'float16',\n 'R_2': 'float16',\n 'D_46': 'float16',\n 'D_47': 'float16',\n 'D_48': 'float16',\n 'D_49': 'float16',\n 'B_6': 'float16',\n 'B_7': 'float16',\n 'B_8': 'float16',\n 'D_50': 'float16',\n 'D_51': 'float16',\n 'B_9': 'float16',\n 'R_3': 'float16',\n 'D_52': 'float16',\n 'P_3': 'float16',\n 'B_10': 'float16',\n 'D_53': 'float16',\n 'S_5': 'float16',\n 'B_11': 'float16',\n 'S_6': 'float16',\n 'D_54': 'float16',\n 'R_4': 'float16',\n 'S_7': 'float16',\n 'B_12': 'float16',\n 'S_8': 'float16',\n 'D_55': 'float16',\n 'D_56': 'float16',\n 'B_13': 'float16',\n 'R_5': 'float16',\n 'D_58': 'float16',\n 'S_9': 'float16',\n 'B_14': 'float16',\n 'D_59': 'float16',\n 'D_60': 'float16',\n 'D_61': 'float16',\n 'B_15': 'float16',\n 'S_11': 'float16',\n 'D_62': 'float16',\n 'D_63': 'object',\n 'D_64': 'object',\n 'D_65': 'float16',\n 'B_16': 'float16',\n 'B_17': 'float16',\n 'B_18': 'float16',\n 'B_19': 'float16',\n 'D_66': 'float16',\n 'B_20': 'float16',\n 'D_68': 'float16',\n 'S_12': 'float16',\n 'R_6': 'float16',\n 'S_13': 'float16',\n 'B_21': 'float16',\n 'D_69': 'float16',\n 'B_22': 'float16',\n 'D_70': 'float16',\n 'D_71': 'float16',\n 'D_72': 'float16',\n 'S_15': 'float16',\n 'B_23': 'float16',\n 'D_73': 'float16',\n 'P_4': 'float16',\n 'D_74': 'float16',\n 'D_75': 'float16',\n 'D_76': 'float16',\n 'B_24': 'float16',\n 'R_7': 'float16',\n 'D_77': 'float16',\n 'B_25': 'float16',\n 'B_26': 'float16',\n 'D_78': 'float16',\n 'D_79': 'float16',\n 'R_8': 'float16',\n 'R_9': 'float16',\n 'S_16': 'float16',\n 'D_80': 'float16',\n 'R_10': 'float16',\n 'R_11': 'float16',\n 'B_27': 'float16',\n 'D_81': 'float16',\n 'D_82': 'float16',\n 'S_17': 'float16',\n 'R_12': 'float16',\n 'B_28': 'float16',\n 'R_13': 'float16',\n 'D_83': 'float16',\n 'R_14': 'float16',\n 'R_15': 'float16',\n 'D_84': 'float16',\n 'R_16': 'float16',\n 'B_29': 'float16',\n 'B_30': 'float16',\n 'S_18': 'float16',\n 'D_86': 'float16',\n 'D_87': 'float16',\n 'R_17': 'float16',\n 'R_18': 'float16',\n 'D_88': 'float16',\n 'B_31': 'int64',\n 'S_19': 'float16',\n 'R_19': 'float16',\n 'B_32': 'float16',\n 'S_20': 'float16',\n 'R_20': 'float16',\n 'R_21': 'float16',\n 'B_33': 'float16',\n 'D_89': 'float16',\n 'R_22': 'float16',\n 'R_23': 'float16',\n 'D_91': 'float16',\n 'D_92': 'float16',\n 'D_93': 'float16',\n 'D_94': 'float16',\n 'R_24': 'float16',\n 'R_25': 'float16',\n 'D_96': 'float16',\n 'S_22': 'float16',\n 'S_23': 'float16',\n 'S_24': 'float16',\n 'S_25': 'float16',\n 'S_26': 'float16',\n 'D_102': 'float16',\n 'D_103': 'float16',\n 'D_104': 'float16',\n 'D_105': 'float16',\n 'D_106': 'float16',\n 'D_107': 'float16',\n 'B_36': 'float16',\n 'B_37': 'float16',\n 'R_26': 'float16',\n 'R_27': 'float16',\n 'B_38': 'float16',\n 'D_108': 'float16',\n 'D_109': 'float16',\n 'D_110': 'float16',\n 'D_111': 'float16',\n 'B_39': 'float16',\n 'D_112': 'float16',\n 'B_40': 'float16',\n 'S_27': 'float16',\n 'D_113': 'float16',\n 'D_114': 'float16',\n 'D_115': 'float16',\n 'D_116': 'float16',\n 'D_117': 'float16',\n 'D_118': 'float16',\n 'D_119': 'float16',\n 'D_120': 'float16',\n 'D_121': 'float16',\n 'D_122': 'float16',\n 'D_123': 'float16',\n 'D_124': 'float16',\n 'D_125': 'float16',\n 'D_126': 'float16',\n 'D_127': 'float16',\n 'D_128': 'float16',\n 'D_129': 'float16',\n 'B_41': 'float16',\n 'B_42': 'float16',\n 'D_130': 'float16',\n 'D_131': 'float16',\n 'D_132': 'float16',\n 'D_133': 'float16',\n 'R_28': 'float16',\n 'D_134': 'float16',\n 'D_135': 'float16',\n 'D_136': 'float16',\n 'D_137': 'float16',\n 'D_138': 'float16',\n 'D_139': 'float16',\n 'D_140': 'float16',\n 'D_141': 'float16',\n 'D_142': 'float16',\n 'D_143': 'float16',\n 'D_144': 'float16',\n 'D_145': 'float16'}\n</code></pre>\n<p>Using this datatype dict, we can read the train data in our Kaggle notebooks environment.</p>\n<pre><code>df = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", dtype=dtype_dict)\n</code></pre>",
      "rawMarkdown": "Datatype dict for our competition which will be useful for reading data proposed in the above notebook.\n\nConverted all the `float64` to `float16`.\n\n```\ndtype_dict = {'customer_ID': \"object\",\n 'S_2': \"object\",\n 'P_2': 'float16',\n 'D_39': 'float16',\n 'B_1': 'float16',\n 'B_2': 'float16',\n 'R_1': 'float16',\n 'S_3': 'float16',\n 'D_41': 'float16',\n 'B_3': 'float16',\n 'D_42': 'float16',\n 'D_43': 'float16',\n 'D_44': 'float16',\n 'B_4': 'float16',\n 'D_45': 'float16',\n 'B_5': 'float16',\n 'R_2': 'float16',\n 'D_46': 'float16',\n 'D_47': 'float16',\n 'D_48': 'float16',\n 'D_49': 'float16',\n 'B_6': 'float16',\n 'B_7': 'float16',\n 'B_8': 'float16',\n 'D_50': 'float16',\n 'D_51': 'float16',\n 'B_9': 'float16',\n 'R_3': 'float16',\n 'D_52': 'float16',\n 'P_3': 'float16',\n 'B_10': 'float16',\n 'D_53': 'float16',\n 'S_5': 'float16',\n 'B_11': 'float16',\n 'S_6': 'float16',\n 'D_54': 'float16',\n 'R_4': 'float16',\n 'S_7': 'float16',\n 'B_12': 'float16',\n 'S_8': 'float16',\n 'D_55': 'float16',\n 'D_56': 'float16',\n 'B_13': 'float16',\n 'R_5': 'float16',\n 'D_58': 'float16',\n 'S_9': 'float16',\n 'B_14': 'float16',\n 'D_59': 'float16',\n 'D_60': 'float16',\n 'D_61': 'float16',\n 'B_15': 'float16',\n 'S_11': 'float16',\n 'D_62': 'float16',\n 'D_63': 'object',\n 'D_64': 'object',\n 'D_65': 'float16',\n 'B_16': 'float16',\n 'B_17': 'float16',\n 'B_18': 'float16',\n 'B_19': 'float16',\n 'D_66': 'float16',\n 'B_20': 'float16',\n 'D_68': 'float16',\n 'S_12': 'float16',\n 'R_6': 'float16',\n 'S_13': 'float16',\n 'B_21': 'float16',\n 'D_69': 'float16',\n 'B_22': 'float16',\n 'D_70': 'float16',\n 'D_71': 'float16',\n 'D_72': 'float16',\n 'S_15': 'float16',\n 'B_23': 'float16',\n 'D_73': 'float16',\n 'P_4': 'float16',\n 'D_74': 'float16',\n 'D_75': 'float16',\n 'D_76': 'float16',\n 'B_24': 'float16',\n 'R_7': 'float16',\n 'D_77': 'float16',\n 'B_25': 'float16',\n 'B_26': 'float16',\n 'D_78': 'float16',\n 'D_79': 'float16',\n 'R_8': 'float16',\n 'R_9': 'float16',\n 'S_16': 'float16',\n 'D_80': 'float16',\n 'R_10': 'float16',\n 'R_11': 'float16',\n 'B_27': 'float16',\n 'D_81': 'float16',\n 'D_82': 'float16',\n 'S_17': 'float16',\n 'R_12': 'float16',\n 'B_28': 'float16',\n 'R_13': 'float16',\n 'D_83': 'float16',\n 'R_14': 'float16',\n 'R_15': 'float16',\n 'D_84': 'float16',\n 'R_16': 'float16',\n 'B_29': 'float16',\n 'B_30': 'float16',\n 'S_18': 'float16',\n 'D_86': 'float16',\n 'D_87': 'float16',\n 'R_17': 'float16',\n 'R_18': 'float16',\n 'D_88': 'float16',\n 'B_31': 'int64',\n 'S_19': 'float16',\n 'R_19': 'float16',\n 'B_32': 'float16',\n 'S_20': 'float16',\n 'R_20': 'float16',\n 'R_21': 'float16',\n 'B_33': 'float16',\n 'D_89': 'float16',\n 'R_22': 'float16',\n 'R_23': 'float16',\n 'D_91': 'float16',\n 'D_92': 'float16',\n 'D_93': 'float16',\n 'D_94': 'float16',\n 'R_24': 'float16',\n 'R_25': 'float16',\n 'D_96': 'float16',\n 'S_22': 'float16',\n 'S_23': 'float16',\n 'S_24': 'float16',\n 'S_25': 'float16',\n 'S_26': 'float16',\n 'D_102': 'float16',\n 'D_103': 'float16',\n 'D_104': 'float16',\n 'D_105': 'float16',\n 'D_106': 'float16',\n 'D_107': 'float16',\n 'B_36': 'float16',\n 'B_37': 'float16',\n 'R_26': 'float16',\n 'R_27': 'float16',\n 'B_38': 'float16',\n 'D_108': 'float16',\n 'D_109': 'float16',\n 'D_110': 'float16',\n 'D_111': 'float16',\n 'B_39': 'float16',\n 'D_112': 'float16',\n 'B_40': 'float16',\n 'S_27': 'float16',\n 'D_113': 'float16',\n 'D_114': 'float16',\n 'D_115': 'float16',\n 'D_116': 'float16',\n 'D_117': 'float16',\n 'D_118': 'float16',\n 'D_119': 'float16',\n 'D_120': 'float16',\n 'D_121': 'float16',\n 'D_122': 'float16',\n 'D_123': 'float16',\n 'D_124': 'float16',\n 'D_125': 'float16',\n 'D_126': 'float16',\n 'D_127': 'float16',\n 'D_128': 'float16',\n 'D_129': 'float16',\n 'B_41': 'float16',\n 'B_42': 'float16',\n 'D_130': 'float16',\n 'D_131': 'float16',\n 'D_132': 'float16',\n 'D_133': 'float16',\n 'R_28': 'float16',\n 'D_134': 'float16',\n 'D_135': 'float16',\n 'D_136': 'float16',\n 'D_137': 'float16',\n 'D_138': 'float16',\n 'D_139': 'float16',\n 'D_140': 'float16',\n 'D_141': 'float16',\n 'D_142': 'float16',\n 'D_143': 'float16',\n 'D_144': 'float16',\n 'D_145': 'float16'}\n```\n\nUsing this datatype dict, we can read the train data in our Kaggle notebooks environment.\n```\ndf = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", dtype=dtype_dict)\n```",
      "votes": null
    },
    {
      "id": "1801957",
      "postDate": "05/26/2022 09:37:00",
      "content": "<p>Didn't know it would make such a big difference. Thanks for letting us know.</p>",
      "rawMarkdown": "Didn't know it would make such a big difference. Thanks for letting us know.",
      "votes": null
    },
    {
      "id": "1801986",
      "postDate": "05/26/2022 10:08:13",
      "content": "<p>Thank you for sharing it <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> </p>",
      "rawMarkdown": "Thank you for sharing it @sudalairajkumar",
      "votes": null
    },
    {
      "id": "1802173",
      "postDate": "05/26/2022 14:20:58",
      "content": "<p>Thanks for this.I am looking for long time something like this.</p>",
      "rawMarkdown": "Thanks for this.I am looking for long time something like this.",
      "votes": null
    },
    {
      "id": "1803715",
      "postDate": "05/28/2022 05:02:19",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> 🤝😍</p>",
      "rawMarkdown": "Thanks for sharing @sudalairajkumar 🤝😍",
      "votes": null
    },
    {
      "id": "1803741",
      "postDate": "05/28/2022 05:35:52",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> 👍</p>",
      "rawMarkdown": "Thanks for sharing @sudalairajkumar 👍",
      "votes": null
    },
    {
      "id": "1803766",
      "postDate": "05/28/2022 06:46:53",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> </p>",
      "rawMarkdown": "Thanks for sharing @sudalairajkumar",
      "votes": null
    },
    {
      "id": "1803910",
      "postDate": "05/28/2022 09:44:13",
      "content": "<p>Thanks very much for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> 😊 Was going to try and figure this out but you have saved me the time!</p>",
      "rawMarkdown": "Thanks very much for sharing @sudalairajkumar 😊 Was going to try and figure this out but you have saved me the time!",
      "votes": null
    },
    {
      "id": "1804185",
      "postDate": "05/28/2022 16:21:25",
      "content": "<p>Shouldn't <code>B_30</code>,<code>B_38</code>,etc be categorical variables?</p>",
      "rawMarkdown": "Shouldn't `B_30`,`B_38`,etc be categorical variables?",
      "votes": null
    },
    {
      "id": "1805148",
      "postDate": "05/29/2022 20:17:32",
      "content": "<p>Thanks for sharing. I have just one question: Shouldn't the list of columns mentioned in the data description (['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']) be set to \"category\" ?</p>",
      "rawMarkdown": "Thanks for sharing. I have just one question: Shouldn't the list of columns mentioned in the data description (['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']) be set to \"category\" ?",
      "votes": null
    },
    {
      "id": "1805569",
      "postDate": "05/30/2022 10:08:31",
      "content": "<p>Fair point <a href=\"https://www.kaggle.com/tawejssh\" target=\"_blank\">@tawejssh</a> and <a href=\"https://www.kaggle.com/itacdonev\" target=\"_blank\">@itacdonev</a> </p>\n<p>I read the data into pandas data frame and then used that as a baseline to do the float conversion. But since we have been given the <code>categorical features</code> list, we can very well specify the set of features as <code>categorical</code> than <code>float</code>. Thank you both. </p>",
      "rawMarkdown": "Fair point @tawejssh and @itacdonev \n\nI read the data into pandas data frame and then used that as a baseline to do the float conversion. But since we have been given the `categorical features` list, we can very well specify the set of features as `categorical` than `float`. Thank you both.",
      "votes": null
    },
    {
      "id": "1903522",
      "postDate": "08/17/2022 13:34:51",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> sir</p>",
      "rawMarkdown": "Thanks for sharing @sudalairajkumar sir",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1801795,
      "author_name": "sudalairajkumar",
      "author_url": "",
      "post_date": "05/26/2022 06:48:23",
      "content": "<p>Datatype dict for our competition which will be useful for reading data proposed in the above notebook.</p>\n<p>Converted all the <code>float64</code> to <code>float16</code>.</p>\n<pre><code>dtype_dict = {'customer_ID': \"object\",\n 'S_2': \"object\",\n 'P_2': 'float16',\n 'D_39': 'float16',\n 'B_1': 'float16',\n 'B_2': 'float16',\n 'R_1': 'float16',\n 'S_3': 'float16',\n 'D_41': 'float16',\n 'B_3': 'float16',\n 'D_42': 'float16',\n 'D_43': 'float16',\n 'D_44': 'float16',\n 'B_4': 'float16',\n 'D_45': 'float16',\n 'B_5': 'float16',\n 'R_2': 'float16',\n 'D_46': 'float16',\n 'D_47': 'float16',\n 'D_48': 'float16',\n 'D_49': 'float16',\n 'B_6': 'float16',\n 'B_7': 'float16',\n 'B_8': 'float16',\n 'D_50': 'float16',\n 'D_51': 'float16',\n 'B_9': 'float16',\n 'R_3': 'float16',\n 'D_52': 'float16',\n 'P_3': 'float16',\n 'B_10': 'float16',\n 'D_53': 'float16',\n 'S_5': 'float16',\n 'B_11': 'float16',\n 'S_6': 'float16',\n 'D_54': 'float16',\n 'R_4': 'float16',\n 'S_7': 'float16',\n 'B_12': 'float16',\n 'S_8': 'float16',\n 'D_55': 'float16',\n 'D_56': 'float16',\n 'B_13': 'float16',\n 'R_5': 'float16',\n 'D_58': 'float16',\n 'S_9': 'float16',\n 'B_14': 'float16',\n 'D_59': 'float16',\n 'D_60': 'float16',\n 'D_61': 'float16',\n 'B_15': 'float16',\n 'S_11': 'float16',\n 'D_62': 'float16',\n 'D_63': 'object',\n 'D_64': 'object',\n 'D_65': 'float16',\n 'B_16': 'float16',\n 'B_17': 'float16',\n 'B_18': 'float16',\n 'B_19': 'float16',\n 'D_66': 'float16',\n 'B_20': 'float16',\n 'D_68': 'float16',\n 'S_12': 'float16',\n 'R_6': 'float16',\n 'S_13': 'float16',\n 'B_21': 'float16',\n 'D_69': 'float16',\n 'B_22': 'float16',\n 'D_70': 'float16',\n 'D_71': 'float16',\n 'D_72': 'float16',\n 'S_15': 'float16',\n 'B_23': 'float16',\n 'D_73': 'float16',\n 'P_4': 'float16',\n 'D_74': 'float16',\n 'D_75': 'float16',\n 'D_76': 'float16',\n 'B_24': 'float16',\n 'R_7': 'float16',\n 'D_77': 'float16',\n 'B_25': 'float16',\n 'B_26': 'float16',\n 'D_78': 'float16',\n 'D_79': 'float16',\n 'R_8': 'float16',\n 'R_9': 'float16',\n 'S_16': 'float16',\n 'D_80': 'float16',\n 'R_10': 'float16',\n 'R_11': 'float16',\n 'B_27': 'float16',\n 'D_81': 'float16',\n 'D_82': 'float16',\n 'S_17': 'float16',\n 'R_12': 'float16',\n 'B_28': 'float16',\n 'R_13': 'float16',\n 'D_83': 'float16',\n 'R_14': 'float16',\n 'R_15': 'float16',\n 'D_84': 'float16',\n 'R_16': 'float16',\n 'B_29': 'float16',\n 'B_30': 'float16',\n 'S_18': 'float16',\n 'D_86': 'float16',\n 'D_87': 'float16',\n 'R_17': 'float16',\n 'R_18': 'float16',\n 'D_88': 'float16',\n 'B_31': 'int64',\n 'S_19': 'float16',\n 'R_19': 'float16',\n 'B_32': 'float16',\n 'S_20': 'float16',\n 'R_20': 'float16',\n 'R_21': 'float16',\n 'B_33': 'float16',\n 'D_89': 'float16',\n 'R_22': 'float16',\n 'R_23': 'float16',\n 'D_91': 'float16',\n 'D_92': 'float16',\n 'D_93': 'float16',\n 'D_94': 'float16',\n 'R_24': 'float16',\n 'R_25': 'float16',\n 'D_96': 'float16',\n 'S_22': 'float16',\n 'S_23': 'float16',\n 'S_24': 'float16',\n 'S_25': 'float16',\n 'S_26': 'float16',\n 'D_102': 'float16',\n 'D_103': 'float16',\n 'D_104': 'float16',\n 'D_105': 'float16',\n 'D_106': 'float16',\n 'D_107': 'float16',\n 'B_36': 'float16',\n 'B_37': 'float16',\n 'R_26': 'float16',\n 'R_27': 'float16',\n 'B_38': 'float16',\n 'D_108': 'float16',\n 'D_109': 'float16',\n 'D_110': 'float16',\n 'D_111': 'float16',\n 'B_39': 'float16',\n 'D_112': 'float16',\n 'B_40': 'float16',\n 'S_27': 'float16',\n 'D_113': 'float16',\n 'D_114': 'float16',\n 'D_115': 'float16',\n 'D_116': 'float16',\n 'D_117': 'float16',\n 'D_118': 'float16',\n 'D_119': 'float16',\n 'D_120': 'float16',\n 'D_121': 'float16',\n 'D_122': 'float16',\n 'D_123': 'float16',\n 'D_124': 'float16',\n 'D_125': 'float16',\n 'D_126': 'float16',\n 'D_127': 'float16',\n 'D_128': 'float16',\n 'D_129': 'float16',\n 'B_41': 'float16',\n 'B_42': 'float16',\n 'D_130': 'float16',\n 'D_131': 'float16',\n 'D_132': 'float16',\n 'D_133': 'float16',\n 'R_28': 'float16',\n 'D_134': 'float16',\n 'D_135': 'float16',\n 'D_136': 'float16',\n 'D_137': 'float16',\n 'D_138': 'float16',\n 'D_139': 'float16',\n 'D_140': 'float16',\n 'D_141': 'float16',\n 'D_142': 'float16',\n 'D_143': 'float16',\n 'D_144': 'float16',\n 'D_145': 'float16'}\n</code></pre>\n<p>Using this datatype dict, we can read the train data in our Kaggle notebooks environment.</p>\n<pre><code>df = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", dtype=dtype_dict)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1801957,
          "author_name": "devkhant24",
          "author_url": "",
          "post_date": "05/26/2022 09:37:00",
          "content": "<p>Didn't know it would make such a big difference. Thanks for letting us know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1803715,
          "author_name": "venkatkumar001",
          "author_url": "",
          "post_date": "05/28/2022 05:02:19",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> 🤝😍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1803910,
          "author_name": "datajmcn",
          "author_url": "",
          "post_date": "05/28/2022 09:44:13",
          "content": "<p>Thanks very much for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> 😊 Was going to try and figure this out but you have saved me the time!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1804185,
          "author_name": "itacdonev",
          "author_url": "",
          "post_date": "05/28/2022 16:21:25",
          "content": "<p>Shouldn't <code>B_30</code>,<code>B_38</code>,etc be categorical variables?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1805569,
          "author_name": "sudalairajkumar",
          "author_url": "",
          "post_date": "05/30/2022 10:08:31",
          "content": "<p>Fair point <a href=\"https://www.kaggle.com/tawejssh\" target=\"_blank\">@tawejssh</a> and <a href=\"https://www.kaggle.com/itacdonev\" target=\"_blank\">@itacdonev</a> </p>\n<p>I read the data into pandas data frame and then used that as a baseline to do the float conversion. But since we have been given the <code>categorical features</code> list, we can very well specify the set of features as <code>categorical</code> than <code>float</code>. Thank you both. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1801986,
      "author_name": "sanjaylalwani",
      "author_url": "",
      "post_date": "05/26/2022 10:08:13",
      "content": "<p>Thank you for sharing it <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1802173,
      "author_name": "gomohit",
      "author_url": "",
      "post_date": "05/26/2022 14:20:58",
      "content": "<p>Thanks for this.I am looking for long time something like this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1803741,
      "author_name": "hassanshehzadk",
      "author_url": "",
      "post_date": "05/28/2022 05:35:52",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1803766,
      "author_name": "subhajitkundu",
      "author_url": "",
      "post_date": "05/28/2022 06:46:53",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1805148,
      "author_name": "tawejssh",
      "author_url": "",
      "post_date": "05/29/2022 20:17:32",
      "content": "<p>Thanks for sharing. I have just one question: Shouldn't the list of columns mentioned in the data description (['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']) be set to \"category\" ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1903522,
      "author_name": "saiteja38",
      "author_url": "",
      "post_date": "08/17/2022 13:34:51",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/sudalairajkumar\" target=\"_blank\">@sudalairajkumar</a> sir</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1801790": "One of the biggest challenge in this competition for all of us is to read such a huge dataset. \n\nOne of our KGMs Rohan has created an excellent tutorial on reading large datasets where he compared different methods and different file formats, which will be super helpful for this competition.\n\nMethods covered:\n* Pandas\n* Dask\n* Datatable\n* Rapids\n\nFile Formats covered:\n* csv\n* feather\n* hdf5\n* jay\n* parquet\n* pickle\n\nLink to notebook - https://www.kaggle.com/code/rohanrao/tutorial-on-reading-large-datasets/notebook\n\nAll the best!",
    "1801795": "Datatype dict for our competition which will be useful for reading data proposed in the above notebook.\n\nConverted all the `float64` to `float16`.\n\n```\ndtype_dict = {'customer_ID': \"object\",\n 'S_2': \"object\",\n 'P_2': 'float16',\n 'D_39': 'float16',\n 'B_1': 'float16',\n 'B_2': 'float16',\n 'R_1': 'float16',\n 'S_3': 'float16',\n 'D_41': 'float16',\n 'B_3': 'float16',\n 'D_42': 'float16',\n 'D_43': 'float16',\n 'D_44': 'float16',\n 'B_4': 'float16',\n 'D_45': 'float16',\n 'B_5': 'float16',\n 'R_2': 'float16',\n 'D_46': 'float16',\n 'D_47': 'float16',\n 'D_48': 'float16',\n 'D_49': 'float16',\n 'B_6': 'float16',\n 'B_7': 'float16',\n 'B_8': 'float16',\n 'D_50': 'float16',\n 'D_51': 'float16',\n 'B_9': 'float16',\n 'R_3': 'float16',\n 'D_52': 'float16',\n 'P_3': 'float16',\n 'B_10': 'float16',\n 'D_53': 'float16',\n 'S_5': 'float16',\n 'B_11': 'float16',\n 'S_6': 'float16',\n 'D_54': 'float16',\n 'R_4': 'float16',\n 'S_7': 'float16',\n 'B_12': 'float16',\n 'S_8': 'float16',\n 'D_55': 'float16',\n 'D_56': 'float16',\n 'B_13': 'float16',\n 'R_5': 'float16',\n 'D_58': 'float16',\n 'S_9': 'float16',\n 'B_14': 'float16',\n 'D_59': 'float16',\n 'D_60': 'float16',\n 'D_61': 'float16',\n 'B_15': 'float16',\n 'S_11': 'float16',\n 'D_62': 'float16',\n 'D_63': 'object',\n 'D_64': 'object',\n 'D_65': 'float16',\n 'B_16': 'float16',\n 'B_17': 'float16',\n 'B_18': 'float16',\n 'B_19': 'float16',\n 'D_66': 'float16',\n 'B_20': 'float16',\n 'D_68': 'float16',\n 'S_12': 'float16',\n 'R_6': 'float16',\n 'S_13': 'float16',\n 'B_21': 'float16',\n 'D_69': 'float16',\n 'B_22': 'float16',\n 'D_70': 'float16',\n 'D_71': 'float16',\n 'D_72': 'float16',\n 'S_15': 'float16',\n 'B_23': 'float16',\n 'D_73': 'float16',\n 'P_4': 'float16',\n 'D_74': 'float16',\n 'D_75': 'float16',\n 'D_76': 'float16',\n 'B_24': 'float16',\n 'R_7': 'float16',\n 'D_77': 'float16',\n 'B_25': 'float16',\n 'B_26': 'float16',\n 'D_78': 'float16',\n 'D_79': 'float16',\n 'R_8': 'float16',\n 'R_9': 'float16',\n 'S_16': 'float16',\n 'D_80': 'float16',\n 'R_10': 'float16',\n 'R_11': 'float16',\n 'B_27': 'float16',\n 'D_81': 'float16',\n 'D_82': 'float16',\n 'S_17': 'float16',\n 'R_12': 'float16',\n 'B_28': 'float16',\n 'R_13': 'float16',\n 'D_83': 'float16',\n 'R_14': 'float16',\n 'R_15': 'float16',\n 'D_84': 'float16',\n 'R_16': 'float16',\n 'B_29': 'float16',\n 'B_30': 'float16',\n 'S_18': 'float16',\n 'D_86': 'float16',\n 'D_87': 'float16',\n 'R_17': 'float16',\n 'R_18': 'float16',\n 'D_88': 'float16',\n 'B_31': 'int64',\n 'S_19': 'float16',\n 'R_19': 'float16',\n 'B_32': 'float16',\n 'S_20': 'float16',\n 'R_20': 'float16',\n 'R_21': 'float16',\n 'B_33': 'float16',\n 'D_89': 'float16',\n 'R_22': 'float16',\n 'R_23': 'float16',\n 'D_91': 'float16',\n 'D_92': 'float16',\n 'D_93': 'float16',\n 'D_94': 'float16',\n 'R_24': 'float16',\n 'R_25': 'float16',\n 'D_96': 'float16',\n 'S_22': 'float16',\n 'S_23': 'float16',\n 'S_24': 'float16',\n 'S_25': 'float16',\n 'S_26': 'float16',\n 'D_102': 'float16',\n 'D_103': 'float16',\n 'D_104': 'float16',\n 'D_105': 'float16',\n 'D_106': 'float16',\n 'D_107': 'float16',\n 'B_36': 'float16',\n 'B_37': 'float16',\n 'R_26': 'float16',\n 'R_27': 'float16',\n 'B_38': 'float16',\n 'D_108': 'float16',\n 'D_109': 'float16',\n 'D_110': 'float16',\n 'D_111': 'float16',\n 'B_39': 'float16',\n 'D_112': 'float16',\n 'B_40': 'float16',\n 'S_27': 'float16',\n 'D_113': 'float16',\n 'D_114': 'float16',\n 'D_115': 'float16',\n 'D_116': 'float16',\n 'D_117': 'float16',\n 'D_118': 'float16',\n 'D_119': 'float16',\n 'D_120': 'float16',\n 'D_121': 'float16',\n 'D_122': 'float16',\n 'D_123': 'float16',\n 'D_124': 'float16',\n 'D_125': 'float16',\n 'D_126': 'float16',\n 'D_127': 'float16',\n 'D_128': 'float16',\n 'D_129': 'float16',\n 'B_41': 'float16',\n 'B_42': 'float16',\n 'D_130': 'float16',\n 'D_131': 'float16',\n 'D_132': 'float16',\n 'D_133': 'float16',\n 'R_28': 'float16',\n 'D_134': 'float16',\n 'D_135': 'float16',\n 'D_136': 'float16',\n 'D_137': 'float16',\n 'D_138': 'float16',\n 'D_139': 'float16',\n 'D_140': 'float16',\n 'D_141': 'float16',\n 'D_142': 'float16',\n 'D_143': 'float16',\n 'D_144': 'float16',\n 'D_145': 'float16'}\n```\n\nUsing this datatype dict, we can read the train data in our Kaggle notebooks environment.\n```\ndf = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", dtype=dtype_dict)\n```",
    "1801957": "Didn't know it would make such a big difference. Thanks for letting us know.",
    "1801986": "Thank you for sharing it @sudalairajkumar",
    "1802173": "Thanks for this.I am looking for long time something like this.",
    "1803715": "Thanks for sharing @sudalairajkumar 🤝😍",
    "1803741": "Thanks for sharing @sudalairajkumar 👍",
    "1803766": "Thanks for sharing @sudalairajkumar",
    "1803910": "Thanks very much for sharing @sudalairajkumar 😊 Was going to try and figure this out but you have saved me the time!",
    "1804185": "Shouldn't `B_30`,`B_38`,etc be categorical variables?",
    "1805148": "Thanks for sharing. I have just one question: Shouldn't the list of columns mentioned in the data description (['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']) be set to \"category\" ?",
    "1805569": "Fair point @tawejssh and @itacdonev \n\nI read the data into pandas data frame and then used that as a baseline to do the float conversion. But since we have been given the `categorical features` list, we can very well specify the set of features as `categorical` than `float`. Thank you both.",
    "1903522": "Thanks for sharing @sudalairajkumar sir"
  },
  "source": "meta"
}