{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Experimenting on @radek1's [Howto] Full dataset as parquet/csv file notebook","metadata":{}},{"cell_type":"markdown","source":"I converated the dataset for this competition from `jsonl` to `csv` and `parquet` so that it is easy to work with using our favorite set of tools! 🙂 You can find the converted dataset [here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).\n\nUnfotunately, it was impossible to process this data on Kaggle due to not enough RAM. I carried out the processing on my local machine and uploaded the processed data (will share the code I used for processing as well, please see [the associate thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)).\n\nThe 'type' information was represented as a string, which takes up a lot of memory. Instead, I cast it to `np.uint8`. This makes the data much smaller and easier to work with without any loss of information!\n\nLet me walk you through how everything is set up so that you can use this data in your work.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"### rd: recsys - otto - access parquet - copy and paste dataset path - !ls ../input/otto-full-optimized-memory-footprint/","metadata":{}},{"cell_type":"code","source":"# Here are the files\n!ls ../input/otto-full-optimized-memory-footprint/","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:44:15.414543Z","iopub.execute_input":"2022-11-11T05:44:15.415691Z","iopub.status.idle":"2022-11-11T05:44:16.426362Z","shell.execute_reply.started":"2022-11-11T05:44:15.415569Z","shell.execute_reply":"2022-11-11T05:44:16.425130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### rd: recsys - otto - access parquet - pd.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# There is a version of the data stored as `csv` as well\n# but I recommend you use `parquet` as I do here -- it is much faster\n\ntrain = pd.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = pd.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:44:16.625657Z","iopub.execute_input":"2022-11-11T05:44:16.626031Z","iopub.status.idle":"2022-11-11T05:44:34.319824Z","shell.execute_reply.started":"2022-11-11T05:44:16.626002Z","shell.execute_reply":"2022-11-11T05:44:34.318677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:44:34.321625Z","iopub.execute_input":"2022-11-11T05:44:34.322040Z","iopub.status.idle":"2022-11-11T05:44:34.339337Z","shell.execute_reply.started":"2022-11-11T05:44:34.322000Z","shell.execute_reply":"2022-11-11T05:44:34.338227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `type` column has been encoded as integers. To translate between the integer and original representations, please use the following.","metadata":{}},{"cell_type":"markdown","source":"### rd: recsys - otto - access parquet - load a function from a pickle file - import pickle5 as pickle - with open('../input/otto-full-optimized-memory-footprint/id2type.pkl', \"rb\") as fh: - id2type = pickle.load(fh)","metadata":{}},{"cell_type":"code","source":"!pip install pickle5\n\nimport pickle5 as pickle\n\nwith open('../input/otto-full-optimized-memory-footprint/id2type.pkl', \"rb\") as fh:\n    id2type = pickle.load(fh)\nwith open('../input/otto-full-optimized-memory-footprint/type2id.pkl', \"rb\") as fh:\n    type2id = pickle.load(fh)","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:46:26.208889Z","iopub.execute_input":"2022-11-11T05:46:26.210132Z","iopub.status.idle":"2022-11-11T05:46:38.359139Z","shell.execute_reply.started":"2022-11-11T05:46:26.210086Z","shell.execute_reply":"2022-11-11T05:46:38.357928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using `id2type` we can convert from integer to string representation (and we can use `type2id` to go in the other direction)","metadata":{}},{"cell_type":"code","source":"id2type, type2id","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:46:38.361463Z","iopub.execute_input":"2022-11-11T05:46:38.361888Z","iopub.status.idle":"2022-11-11T05:46:38.368942Z","shell.execute_reply.started":"2022-11-11T05:46:38.361856Z","shell.execute_reply":"2022-11-11T05:46:38.367929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id2type","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:46:38.370103Z","iopub.execute_input":"2022-11-11T05:46:38.370360Z","iopub.status.idle":"2022-11-11T05:46:38.380456Z","shell.execute_reply.started":"2022-11-11T05:46:38.370336Z","shell.execute_reply":"2022-11-11T05:46:38.379541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### rd: recsys - otto - access parquet - access the first 1000 rows and convert int back to string, train.iloc[:1000].type.map(lambda i: id2type[i])","metadata":{}},{"cell_type":"code","source":"type_as_string = train.iloc[:1000].type.map(lambda i: id2type[i])\ntype_as_string.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:47:05.222108Z","iopub.execute_input":"2022-11-11T05:47:05.222547Z","iopub.status.idle":"2022-11-11T05:47:05.240315Z","shell.execute_reply.started":"2022-11-11T05:47:05.222512Z","shell.execute_reply":"2022-11-11T05:47:05.239101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we can just as easily go back from strings to idxs.","metadata":{}},{"cell_type":"markdown","source":"### rd: recsys - otto - access parquet - how to use Series, map, lambda, dict together - type_as_string.map(lambda i: type2id[i])","metadata":{}},{"cell_type":"code","source":"type_as_string.map(lambda i: type2id[i]).head()","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:53:23.610784Z","iopub.execute_input":"2022-11-11T05:53:23.611183Z","iopub.status.idle":"2022-11-11T05:53:23.619455Z","shell.execute_reply.started":"2022-11-11T05:53:23.611154Z","shell.execute_reply":"2022-11-11T05:53:23.618548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_as_string.map(type2id).head()","metadata":{"execution":{"iopub.status.busy":"2022-11-11T05:48:20.592646Z","iopub.execute_input":"2022-11-11T05:48:20.593107Z","iopub.status.idle":"2022-11-11T05:48:20.603238Z","shell.execute_reply.started":"2022-11-11T05:48:20.593070Z","shell.execute_reply":"2022-11-11T05:48:20.602126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And that's it! I hope this will speed you along in your work 🙂\n\nIf you found this useful, I would be extremely grateful if you could please upvote this notebook and [the associated dataset](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).\n\nThank you so much for your support! Happy kaggling! 🥳","metadata":{}}]}