{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!pip install july\n\nimport pandas as pd\nimport numpy as np\nimport datetime\nimport july\n\ntrain = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet')[['customer_ID','S_2']]\ntest = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/test.parquet')[['customer_ID','S_2']]\n\nlabels = pd.read_csv('../input/amex-default-prediction/train_labels.csv')\ntrain = pd.merge(train, labels, on = 'customer_ID')\n\ntrain['S_2'] = train['S_2'].apply(datetime.datetime.fromisoformat)\ntrain['N'] = train.groupby('customer_ID')['customer_ID'].transform('count')\ntrain = train.loc[train.N==13] # let's deal with bias of new customers\n\ntest['S_2'] = test['S_2'].apply(datetime.datetime.fromisoformat)\ntest['N'] = test.groupby('customer_ID')['customer_ID'].transform('count')\ntest = test.loc[test.N==13] # let's deal with bias of new customers\n\ntest['max_date'] = test.groupby('customer_ID')['S_2'].transform('max')\ntest1 = test.loc[test.max_date.apply(lambda t: t.month == 4)]\ntest2 = test.loc[test.max_date.apply(lambda t: t.month == 10)]\n\ntx = train.groupby('S_2')['customer_ID'].count().reset_index(name='N')\nty = train.groupby('S_2')['target'].mean().reset_index(name='target_rate')\ntw1 = test1.groupby('S_2')['customer_ID'].count().reset_index(name='N')\ntw2 = test2.groupby('S_2')['customer_ID'].count().reset_index(name='N')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-16T16:25:03.217659Z","iopub.execute_input":"2022-07-16T16:25:03.218887Z","iopub.status.idle":"2022-07-16T16:25:14.253979Z","shell.execute_reply.started":"2022-07-16T16:25:03.218834Z","shell.execute_reply":"2022-07-16T16:25:14.252815Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Quick look at distribution by S_2\n\nLet's look how data distributes by weekday. ```july``` is a nice package doing everything we need under the hood.","metadata":{}},{"cell_type":"code","source":"july.heatmap(tx['S_2'], tx['N'], title='train count', cmap=\"golden\", colorbar=True)\njuly.heatmap(tw1['S_2'], tw1['N'], title='test (public) count', cmap=\"golden\", colorbar=True)\njuly.heatmap(tw2['S_2'], tw2['N'], title='test (private) count', cmap=\"golden\", colorbar=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T16:34:17.922565Z","iopub.status.idle":"2022-07-16T16:34:17.922993Z","shell.execute_reply.started":"2022-07-16T16:34:17.922778Z","shell.execute_reply":"2022-07-16T16:34:17.922812Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems Saturdays are most popular dates, but Sundays the least popular. Could we see a pattern for target rate?","metadata":{}},{"cell_type":"code","source":"july.heatmap(ty['S_2'], ty['target_rate'], title='target rate', cmap=\"golden\", colorbar=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T16:34:17.924598Z","iopub.status.idle":"2022-07-16T16:34:17.925105Z","shell.execute_reply.started":"2022-07-16T16:34:17.924895Z","shell.execute_reply":"2022-07-16T16:34:17.924918Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sundays although rare in ```S_2```, however consistent target rate across multiple months is observed.\n\n### Some ideas\n\n- Distribution of weekdays could be considered as features? Seems risky though as we would have to assume that it holds for test set as well.\n- On the safe side, maybe just count how many Sundays were observed?","metadata":{}}]}