{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Question: How To Get 32GB RAM?\n# Answer: Add This Notebook's Output to a New Kaggle Notebook\nIn Kaggle's Competition \"Predict Student Performance from Game Play\", we must submit to the leaderboard with a 8GB RAM CPU Kaggle notebook. The train data (after Kaggle increased data on Mar 20th 2023) is 4.7GB. This makes it hard to both train and infer in the same 8GB RAM notebook. The best solution is to use two notebook. Infer in 8GB notebook but train your models in 32GB RAM notebook either: \n* 32GB CPU Notebook\n* 32GB GPU Notebook (= 16GB GPU + 16GB CPU)\n\nThere is a trick to get a 32GB RAM notebook. We cannot create a notebook from this competition's dataset because if we do, then Kaggle changes the memory to 8GB. To get 32GB RAM, we must make a new notebook outside of this competition. Then attach the output of the notebook you are reading right now which contains this competition's `train.csv` and `train_labels.csv` as parquet files. After doing this, your new notebook will have 32GB RAM!\n\nTrain your model with this 32GB notebook and save your model to that notebook's output files. Then create a new (second) 8GB CPU Notebook (from Kaggle's competition data) and attach your first notebook's output as input to your second notebook. Use your second notebook to only load your models and infer test. \n\nDiscussion about using two notebooks is [here][3] and [here][1]. If you prefer to do everything in one 8GB notebook, then you must read the train data in chunks and feature engineer in chunks to avoid memory error. An example using one notebook is [here][2].\n\n[1]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/386218\n[2]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\n[3]: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396979","metadata":{}},{"cell_type":"code","source":"%%time\nimport pandas as pd, gc\ndf = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train.csv') \ndf.to_parquet('train.parquet',index=False)","metadata":{"execution":{"iopub.status.busy":"2023-03-23T07:46:33.930113Z","iopub.execute_input":"2023-03-23T07:46:33.931024Z","iopub.status.idle":"2023-03-23T07:48:20.901492Z","shell.execute_reply.started":"2023-03-23T07:46:33.930978Z","shell.execute_reply":"2023-03-23T07:48:20.900459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndel df; gc.collect()\ndf = pd.read_csv('/kaggle/input/predict-student-performance-from-game-play/train_labels.csv') \ndf.to_parquet('train_labels.parquet',index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Six Steps To Get 32GB RAM\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2023/m1.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2023/m2.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2023/m3.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2023/m4.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2023/m5.png)","metadata":{}}]}