{
  "id": 327160,
  "title": "Minimize memory load",
  "url": "/competitions/amex-default-prediction/discussion/327160",
  "author_name": "",
  "post_date": "2022-05-25T22:02:48.629722200Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my effort to try to reduce the memory consumption reading in the data. Im only reading in 100000 rows to begin with to experiment with. If you have any further advice please comment :)</p>\n<p>`<br>\nnrows = 100000</p>\n<p>float_cols = ['B_1', 'B_10', 'B_11', 'B_12', 'B_13', 'B_14', 'B_15', 'B_16', 'B_17', 'B_18', 'B_19', 'B_2', 'B_20', 'B_21', 'B_22', 'B_23', 'B_24', 'B_25', 'B_26', 'B_27', 'B_28', 'B_29', 'B_3', 'B_32', 'B_33', 'B_36', <br>\n              'B_37', 'B_39', 'B_4', 'B_40', 'B_41', 'B_42', 'B_5', 'B_6', 'B_7', 'B_8', 'B_9', 'D_102', 'D_103', 'D_104', 'D_105', 'D_106', 'D_107', 'D_108', 'D_109', 'D_110', 'D_111', 'D_112', 'D_113', 'D_115', 'D_118',<br>\n              'D_119', 'D_121', 'D_122', 'D_123', 'D_124', 'D_125', 'D_127', 'D_128', 'D_129', 'D_130', 'D_131', 'D_132', 'D_133', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138', 'D_139', 'D_140', 'D_141', 'D_142', 'D_143',<br>\n              'D_144', 'D_145', 'D_39', 'D_41', 'D_42', 'D_43', 'D_44', 'D_45', 'D_46', 'D_47', 'D_48', 'D_49', 'D_50', 'D_51', 'D_52', 'D_53', 'D_54', 'D_55', 'D_56', 'D_58', 'D_59', 'D_60', 'D_61', 'D_62', 'D_65', 'D_69',<br>\n              'D_70', 'D_71', 'D_72', 'D_73', 'D_74', 'D_75', 'D_76', 'D_77', 'D_78', 'D_79', 'D_80', 'D_81', 'D_82', 'D_83', 'D_84', 'D_86', 'D_87', 'D_88', 'D_89', 'D_91', 'D_92', 'D_93', 'D_94', 'D_96', 'P_2', 'P_3', <br>\n              'P_4', 'R_1', 'R_10', 'R_11', 'R_12', 'R_13', 'R_14', 'R_15', 'R_16', 'R_17', 'R_18', 'R_19', 'R_2', 'R_20', 'R_21', 'R_22', 'R_23', 'R_24', 'R_25', 'R_26', 'R_27', 'R_28', 'R_3', 'R_4', 'R_5', 'R_6', 'R_7', <br>\n              'R_8', 'R_9', 'S_11', 'S_12', 'S_13', 'S_15', 'S_16', 'S_17', 'S_18', 'S_19', 'S_20', 'S_22', 'S_23', 'S_24', 'S_25', 'S_26', 'S_27', 'S_3', 'S_5', 'S_6', 'S_7', 'S_8', 'S_9']</p>\n<p>df_t_f = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=float_cols, dtype=\"float16\")</p>\n<p>cat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']<br>\ndf_t_c = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=cat_cols, dtype=\"category\")</p>\n<p>rest_cols = ['customer_ID', 'S_2', 'B_31']<br>\nparse_dates = ['S_2']<br>\ndf_t_id_d = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, parse_dates=parse_dates, usecols=rest_cols)#, dtype={\"B31\":\"category\")</p>\n<p>df_train = pd.concat([df_t_id_d, df_t_c, df_t_f], axis=1)`</p>",
  "messages": [
    {
      "id": "1801565",
      "postDate": "05/25/2022 22:02:48",
      "content": "<p>This is my effort to try to reduce the memory consumption reading in the data. Im only reading in 100000 rows to begin with to experiment with. If you have any further advice please comment :)</p>\n<p>`<br>\nnrows = 100000</p>\n<p>float_cols = ['B_1', 'B_10', 'B_11', 'B_12', 'B_13', 'B_14', 'B_15', 'B_16', 'B_17', 'B_18', 'B_19', 'B_2', 'B_20', 'B_21', 'B_22', 'B_23', 'B_24', 'B_25', 'B_26', 'B_27', 'B_28', 'B_29', 'B_3', 'B_32', 'B_33', 'B_36', <br>\n              'B_37', 'B_39', 'B_4', 'B_40', 'B_41', 'B_42', 'B_5', 'B_6', 'B_7', 'B_8', 'B_9', 'D_102', 'D_103', 'D_104', 'D_105', 'D_106', 'D_107', 'D_108', 'D_109', 'D_110', 'D_111', 'D_112', 'D_113', 'D_115', 'D_118',<br>\n              'D_119', 'D_121', 'D_122', 'D_123', 'D_124', 'D_125', 'D_127', 'D_128', 'D_129', 'D_130', 'D_131', 'D_132', 'D_133', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138', 'D_139', 'D_140', 'D_141', 'D_142', 'D_143',<br>\n              'D_144', 'D_145', 'D_39', 'D_41', 'D_42', 'D_43', 'D_44', 'D_45', 'D_46', 'D_47', 'D_48', 'D_49', 'D_50', 'D_51', 'D_52', 'D_53', 'D_54', 'D_55', 'D_56', 'D_58', 'D_59', 'D_60', 'D_61', 'D_62', 'D_65', 'D_69',<br>\n              'D_70', 'D_71', 'D_72', 'D_73', 'D_74', 'D_75', 'D_76', 'D_77', 'D_78', 'D_79', 'D_80', 'D_81', 'D_82', 'D_83', 'D_84', 'D_86', 'D_87', 'D_88', 'D_89', 'D_91', 'D_92', 'D_93', 'D_94', 'D_96', 'P_2', 'P_3', <br>\n              'P_4', 'R_1', 'R_10', 'R_11', 'R_12', 'R_13', 'R_14', 'R_15', 'R_16', 'R_17', 'R_18', 'R_19', 'R_2', 'R_20', 'R_21', 'R_22', 'R_23', 'R_24', 'R_25', 'R_26', 'R_27', 'R_28', 'R_3', 'R_4', 'R_5', 'R_6', 'R_7', <br>\n              'R_8', 'R_9', 'S_11', 'S_12', 'S_13', 'S_15', 'S_16', 'S_17', 'S_18', 'S_19', 'S_20', 'S_22', 'S_23', 'S_24', 'S_25', 'S_26', 'S_27', 'S_3', 'S_5', 'S_6', 'S_7', 'S_8', 'S_9']</p>\n<p>df_t_f = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=float_cols, dtype=\"float16\")</p>\n<p>cat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']<br>\ndf_t_c = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=cat_cols, dtype=\"category\")</p>\n<p>rest_cols = ['customer_ID', 'S_2', 'B_31']<br>\nparse_dates = ['S_2']<br>\ndf_t_id_d = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, parse_dates=parse_dates, usecols=rest_cols)#, dtype={\"B31\":\"category\")</p>\n<p>df_train = pd.concat([df_t_id_d, df_t_c, df_t_f], axis=1)`</p>",
      "rawMarkdown": "This is my effort to try to reduce the memory consumption reading in the data. Im only reading in 100000 rows to begin with to experiment with. If you have any further advice please comment :)\n\n`\nnrows = 100000\n\nfloat_cols = ['B_1', 'B_10', 'B_11', 'B_12', 'B_13', 'B_14', 'B_15', 'B_16', 'B_17', 'B_18', 'B_19', 'B_2', 'B_20', 'B_21', 'B_22', 'B_23', 'B_24', 'B_25', 'B_26', 'B_27', 'B_28', 'B_29', 'B_3', 'B_32', 'B_33', 'B_36', \n              'B_37', 'B_39', 'B_4', 'B_40', 'B_41', 'B_42', 'B_5', 'B_6', 'B_7', 'B_8', 'B_9', 'D_102', 'D_103', 'D_104', 'D_105', 'D_106', 'D_107', 'D_108', 'D_109', 'D_110', 'D_111', 'D_112', 'D_113', 'D_115', 'D_118',\n              'D_119', 'D_121', 'D_122', 'D_123', 'D_124', 'D_125', 'D_127', 'D_128', 'D_129', 'D_130', 'D_131', 'D_132', 'D_133', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138', 'D_139', 'D_140', 'D_141', 'D_142', 'D_143',\n              'D_144', 'D_145', 'D_39', 'D_41', 'D_42', 'D_43', 'D_44', 'D_45', 'D_46', 'D_47', 'D_48', 'D_49', 'D_50', 'D_51', 'D_52', 'D_53', 'D_54', 'D_55', 'D_56', 'D_58', 'D_59', 'D_60', 'D_61', 'D_62', 'D_65', 'D_69',\n              'D_70', 'D_71', 'D_72', 'D_73', 'D_74', 'D_75', 'D_76', 'D_77', 'D_78', 'D_79', 'D_80', 'D_81', 'D_82', 'D_83', 'D_84', 'D_86', 'D_87', 'D_88', 'D_89', 'D_91', 'D_92', 'D_93', 'D_94', 'D_96', 'P_2', 'P_3', \n              'P_4', 'R_1', 'R_10', 'R_11', 'R_12', 'R_13', 'R_14', 'R_15', 'R_16', 'R_17', 'R_18', 'R_19', 'R_2', 'R_20', 'R_21', 'R_22', 'R_23', 'R_24', 'R_25', 'R_26', 'R_27', 'R_28', 'R_3', 'R_4', 'R_5', 'R_6', 'R_7', \n              'R_8', 'R_9', 'S_11', 'S_12', 'S_13', 'S_15', 'S_16', 'S_17', 'S_18', 'S_19', 'S_20', 'S_22', 'S_23', 'S_24', 'S_25', 'S_26', 'S_27', 'S_3', 'S_5', 'S_6', 'S_7', 'S_8', 'S_9']\n\ndf_t_f = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=float_cols, dtype=\"float16\")\n\ncat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ndf_t_c = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=cat_cols, dtype=\"category\")\n\nrest_cols = ['customer_ID', 'S_2', 'B_31']\nparse_dates = ['S_2']\ndf_t_id_d = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, parse_dates=parse_dates, usecols=rest_cols)#, dtype={\"B31\":\"category\")\n\ndf_train = pd.concat([df_t_id_d, df_t_c, df_t_f], axis=1)`",
      "votes": null
    },
    {
      "id": "1882769",
      "postDate": "08/03/2022 12:19:49",
      "content": "<p>How did you find out which column is categorical and which is float? </p>",
      "rawMarkdown": "How did you find out which column is categorical and which is float?",
      "votes": null
    },
    {
      "id": "1883383",
      "postDate": "08/03/2022 19:23:50",
      "content": "<p>It was given on the data page of the competition.</p>",
      "rawMarkdown": "It was given on the data page of the competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1882769,
      "author_name": "kulkarniishwarinitin",
      "author_url": "",
      "post_date": "08/03/2022 12:19:49",
      "content": "<p>How did you find out which column is categorical and which is float? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1883383,
          "author_name": "memocan64",
          "author_url": "",
          "post_date": "08/03/2022 19:23:50",
          "content": "<p>It was given on the data page of the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1801565": "This is my effort to try to reduce the memory consumption reading in the data. Im only reading in 100000 rows to begin with to experiment with. If you have any further advice please comment :)\n\n`\nnrows = 100000\n\nfloat_cols = ['B_1', 'B_10', 'B_11', 'B_12', 'B_13', 'B_14', 'B_15', 'B_16', 'B_17', 'B_18', 'B_19', 'B_2', 'B_20', 'B_21', 'B_22', 'B_23', 'B_24', 'B_25', 'B_26', 'B_27', 'B_28', 'B_29', 'B_3', 'B_32', 'B_33', 'B_36', \n              'B_37', 'B_39', 'B_4', 'B_40', 'B_41', 'B_42', 'B_5', 'B_6', 'B_7', 'B_8', 'B_9', 'D_102', 'D_103', 'D_104', 'D_105', 'D_106', 'D_107', 'D_108', 'D_109', 'D_110', 'D_111', 'D_112', 'D_113', 'D_115', 'D_118',\n              'D_119', 'D_121', 'D_122', 'D_123', 'D_124', 'D_125', 'D_127', 'D_128', 'D_129', 'D_130', 'D_131', 'D_132', 'D_133', 'D_134', 'D_135', 'D_136', 'D_137', 'D_138', 'D_139', 'D_140', 'D_141', 'D_142', 'D_143',\n              'D_144', 'D_145', 'D_39', 'D_41', 'D_42', 'D_43', 'D_44', 'D_45', 'D_46', 'D_47', 'D_48', 'D_49', 'D_50', 'D_51', 'D_52', 'D_53', 'D_54', 'D_55', 'D_56', 'D_58', 'D_59', 'D_60', 'D_61', 'D_62', 'D_65', 'D_69',\n              'D_70', 'D_71', 'D_72', 'D_73', 'D_74', 'D_75', 'D_76', 'D_77', 'D_78', 'D_79', 'D_80', 'D_81', 'D_82', 'D_83', 'D_84', 'D_86', 'D_87', 'D_88', 'D_89', 'D_91', 'D_92', 'D_93', 'D_94', 'D_96', 'P_2', 'P_3', \n              'P_4', 'R_1', 'R_10', 'R_11', 'R_12', 'R_13', 'R_14', 'R_15', 'R_16', 'R_17', 'R_18', 'R_19', 'R_2', 'R_20', 'R_21', 'R_22', 'R_23', 'R_24', 'R_25', 'R_26', 'R_27', 'R_28', 'R_3', 'R_4', 'R_5', 'R_6', 'R_7', \n              'R_8', 'R_9', 'S_11', 'S_12', 'S_13', 'S_15', 'S_16', 'S_17', 'S_18', 'S_19', 'S_20', 'S_22', 'S_23', 'S_24', 'S_25', 'S_26', 'S_27', 'S_3', 'S_5', 'S_6', 'S_7', 'S_8', 'S_9']\n\ndf_t_f = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=float_cols, dtype=\"float16\")\n\ncat_cols = ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ndf_t_c = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, usecols=cat_cols, dtype=\"category\")\n\nrest_cols = ['customer_ID', 'S_2', 'B_31']\nparse_dates = ['S_2']\ndf_t_id_d = pd.read_csv(\"/kaggle/input/amex-default-prediction/train_data.csv\", nrows=nrows, parse_dates=parse_dates, usecols=rest_cols)#, dtype={\"B31\":\"category\")\n\ndf_train = pd.concat([df_t_id_d, df_t_c, df_t_f], axis=1)`",
    "1882769": "How did you find out which column is categorical and which is float?",
    "1883383": "It was given on the data page of the competition."
  },
  "source": "meta"
}