{"cells":[{"metadata":{},"cell_type":"markdown","source":"There's been some discussion around dealing with PlayerTrackData.csv and it's strain on system memory. As some have pointed out, one solution is to use Kaggle kernels. Personally I don't like to develop with a Jupyter notebook, especially a remote one. I'd much rather work on my local machine and use something like VS Code with the python extension and all the other goodies. \n\nToday I finally got into the player track data. My desktop with 24<s>MB</s>GB RAM could barely handle the file straight up. So I used Pandas chunk feature along with my MemReducer utility script. There are many versions of memory reducers out there, but this one is more effective than most. It's actually pretty simple - you can find it [here](https://www.kaggle.com/jpmiller/skmem).\n\nAlso, for this data I prefer to have the Players, Games, and Plays separated vs. the hyphenated thing (makes it easier for grouping and filtering plus gives us memory-efficient integers instead of silly strings). Here's what I'm doing:"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import sys\nimport numpy as np\nimport pandas as pd\npd.options.display.float_format = '{:,.2f}'.format\nfrom tqdm.notebook import tqdm\n\nimport skmem #utility script","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"csize = 4_000_000 #set this to fit your situation\nchunker = pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv',\n                      chunksize=csize)\ntrack_list = []\nmr = skmem.MemReducer()\nfor chunk in tqdm(chunker, total = int(80_000_000/csize)):\n    chunk['PlayKey'] = chunk.PlayKey.fillna('0-0-0')\n    id_array = chunk.PlayKey.str.split('-', expand=True).to_numpy()\n    chunk['PlayerKey'] = id_array[:,0]\n    chunk['GameID'] = id_array[:,1]\n    chunk['PlayKey'] = id_array[:,2]\n    chunk['event'] = chunk.event.fillna('none')\n    floaters = chunk.select_dtypes('float').columns.tolist()\n    chunk = mr.fit_transform(chunk, float_cols=floaters) #float downcast is optional\n    track_list.append(chunk)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tracks = pd.concat(track_list)\ntracks['event'] = tracks.event.astype('category') #retype after concat\ncol_order = [9,10,0,1,2,3,4,5,7,6,8]\ntracks = tracks[[tracks.columns[idx] for idx in col_order]]\ndisplay(tracks.dtypes, tracks.head())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tracks.to_parquet('InjuryRecord.parq')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"You can see from the output that the memory footprint is reduced by a factor of 10. Saving as parquet is also nice. It preserves data types and is very efficient. Just make sure you have fastparquet and python-snappy installed in your environment.\n\nAnd/or you can download the parquet version here and get going - good luck with the challenge!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}