{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#### @[Seeker](https://www.kaggle.com/taran8727) hypothesized in [this discussion](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327602#1805287), that the last statement time is different for public and private parts of the test set. The ones with the last statement month of 2019-04 belong to the public, and those with 2019-10 belong to the private part.\n\n#### Here, I'll suggest a way to immediately utilize this information, especially for those who are struggling with memory issues, and also, I will \"prove\" that by changing the predictions that are made for the customers with the last statement month of 2019-10, of an existing submission that scores 0.791, and if this is correct, the score still be 0.791.","metadata":{"execution":{"iopub.status.busy":"2022-05-30T06:49:15.699449Z","iopub.execute_input":"2022-05-30T06:49:15.699876Z","iopub.status.idle":"2022-05-30T06:49:15.705342Z","shell.execute_reply.started":"2022-05-30T06:49:15.699842Z","shell.execute_reply":"2022-05-30T06:49:15.704103Z"}}},{"cell_type":"code","source":"import pandas as pd\n\ntest = pd.read_csv(\"../input/testing-public-private-split-amex/test_only_s2.csv\", parse_dates=[\"S_2\"])\nsub = pd.read_csv(\"../input/testing-public-private-split-amex/sub8.csv\")\ntest[\"S_2\"].dt.month.value_counts(normalize=True).round(2)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-05-30T07:02:46.693540Z","iopub.execute_input":"2022-05-30T07:02:46.694000Z","iopub.status.idle":"2022-05-30T07:02:48.545238Z","shell.execute_reply.started":"2022-05-30T07:02:46.693968Z","shell.execute_reply":"2022-05-30T07:02:48.544572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Seeing this notebook's score as 0.791, should prove the point.","metadata":{}},{"cell_type":"code","source":"sub.loc[test[\"S_2\"].dt.month==10, \"prediction\"] = 0\nsub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-05-30T07:03:01.464652Z","iopub.execute_input":"2022-05-30T07:03:01.465139Z","iopub.status.idle":"2022-05-30T07:03:06.282021Z","shell.execute_reply.started":"2022-05-30T07:03:01.465102Z","shell.execute_reply":"2022-05-30T07:03:06.280845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Now to utilize this information, here's an option. Just using the data for the public dataset and throwing out the rest, until the late stages of the competition of course. This will halve the size of the test set and make inference a lot faster.\n#### When making a submission, just map the predictions using customer_ID and fill the nan's and it should work fine.","metadata":{"execution":{"iopub.status.busy":"2022-05-30T07:03:42.324215Z","iopub.execute_input":"2022-05-30T07:03:42.324865Z","iopub.status.idle":"2022-05-30T07:03:42.329185Z","shell.execute_reply.started":"2022-05-30T07:03:42.324828Z","shell.execute_reply":"2022-05-30T07:03:42.328462Z"}}},{"cell_type":"code","source":"test = test[test[\"S_2\"].dt.month==4]","metadata":{"execution":{"iopub.status.busy":"2022-05-30T07:04:26.126463Z","iopub.execute_input":"2022-05-30T07:04:26.127458Z","iopub.status.idle":"2022-05-30T07:04:26.244794Z","shell.execute_reply.started":"2022-05-30T07:04:26.127420Z","shell.execute_reply":"2022-05-30T07:04:26.243892Z"},"trusted":true},"execution_count":null,"outputs":[]}]}