{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I converated the dataset for this competition from `jsonl` to `csv` and `parquet` so that it is easy to work with using our favorite set of tools! 🙂 You can find the converted dataset [here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).\n\nUnfotunately, it was impossible to process this data on Kaggle due to not enough RAM. I carried out the processing on my local machine and uploaded the processed data (will share the code I used for processing as well, please see [the associate thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)).\n\nThe 'type' information was represented as a string, which takes up a lot of memory. Instead, I cast it to `np.uint8`. This makes the data much smaller and easier to work with without any loss of information!\n\n## Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n\nLet me walk you through how everything is set up so that you can use this data in your work.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"# Here are the files\n!ls ../input/otto-full-optimized-memory-footprint/","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:53:40.357527Z","iopub.execute_input":"2022-11-03T14:53:40.357861Z","iopub.status.idle":"2022-11-03T14:53:40.660284Z","shell.execute_reply.started":"2022-11-03T14:53:40.357786Z","shell.execute_reply":"2022-11-03T14:53:40.658920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\n\n# There is a version of the data stored as `csv` as well\n# but I recommend you use `parquet` as I do here -- it is much faster\n\ntrain = pd.read_parquet('../input/otto-full-optimized-memory-footprint/train.parquet')\ntest = pd.read_parquet('../input/otto-full-optimized-memory-footprint/test.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:53:40.661822Z","iopub.execute_input":"2022-11-03T14:53:40.663867Z","iopub.status.idle":"2022-11-03T14:54:03.879464Z","shell.execute_reply.started":"2022-11-03T14:53:40.663788Z","shell.execute_reply":"2022-11-03T14:54:03.878203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:54:03.880440Z","iopub.execute_input":"2022-11-03T14:54:03.880721Z","iopub.status.idle":"2022-11-03T14:54:03.910742Z","shell.execute_reply.started":"2022-11-03T14:54:03.880696Z","shell.execute_reply":"2022-11-03T14:54:03.909554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `type` column has been encoded as integers. To translate between the integer and original representations, please use the following.","metadata":{}},{"cell_type":"code","source":"!pip install pickle5\n\nimport pickle5 as pickle\n\nwith open('../input/otto-full-optimized-memory-footprint/id2type.pkl', \"rb\") as fh:\n    id2type = pickle.load(fh)\nwith open('../input/otto-full-optimized-memory-footprint/type2id.pkl', \"rb\") as fh:\n    type2id = pickle.load(fh)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:54:03.915814Z","iopub.execute_input":"2022-11-03T14:54:03.916437Z","iopub.status.idle":"2022-11-03T14:54:15.083942Z","shell.execute_reply.started":"2022-11-03T14:54:03.916373Z","shell.execute_reply":"2022-11-03T14:54:15.082490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using `id2type` we can convert from integer to string representation (and we can use `type2id` to go in the other direction)","metadata":{}},{"cell_type":"code","source":"id2type, type2id","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:54:15.086370Z","iopub.execute_input":"2022-11-03T14:54:15.086764Z","iopub.status.idle":"2022-11-03T14:54:15.095172Z","shell.execute_reply.started":"2022-11-03T14:54:15.086729Z","shell.execute_reply":"2022-11-03T14:54:15.093752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id2type","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:54:15.097080Z","iopub.execute_input":"2022-11-03T14:54:15.097489Z","iopub.status.idle":"2022-11-03T14:54:15.110567Z","shell.execute_reply.started":"2022-11-03T14:54:15.097454Z","shell.execute_reply":"2022-11-03T14:54:15.109100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_as_string = train.iloc[:1000].type.map(lambda i: id2type[i])\ntype_as_string.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:54:27.669862Z","iopub.execute_input":"2022-11-03T14:54:27.670897Z","iopub.status.idle":"2022-11-03T14:54:27.681744Z","shell.execute_reply.started":"2022-11-03T14:54:27.670853Z","shell.execute_reply":"2022-11-03T14:54:27.680388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we can just as easily go back from strings to idxs.","metadata":{}},{"cell_type":"code","source":"type_as_string.map(type2id).head()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:54:41.165962Z","iopub.execute_input":"2022-11-03T14:54:41.166484Z","iopub.status.idle":"2022-11-03T14:54:41.176862Z","shell.execute_reply.started":"2022-11-03T14:54:41.166447Z","shell.execute_reply":"2022-11-03T14:54:41.175236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And that's it! I hope this will speed you along in your work 🙂\n\n**If you found this useful, I would be extremely grateful if you could please upvote this notebook and [the associated dataset](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).**\n\nThank you so much for your support! Happy kaggling! 🥳","metadata":{}}]}