{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M - Baseline: 12 most popular items\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/31254/logos/header.png)\n\n## A content-based naive recommendation system using the 12 items that were purchased the most (\"most popular\") in the full transaction history for the competition [H&M Personalized Fashion Recommendations](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations).\n\n## Total lines of code: 5\n\n\n# Please, _DO_ upvote if you find this kernel useful or interesting!","metadata":{}},{"cell_type":"markdown","source":"# Read dataframes","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv', dtype={'article_id': str})\ndf_sub = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T01:45:25.52007Z","iopub.execute_input":"2022-02-08T01:45:25.520991Z","iopub.status.idle":"2022-02-08T01:46:31.633528Z","shell.execute_reply.started":"2022-02-08T01:45:25.520863Z","shell.execute_reply":"2022-02-08T01:46:31.632669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.t_dat.max()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T01:47:11.676547Z","iopub.execute_input":"2022-02-08T01:47:11.677065Z","iopub.status.idle":"2022-02-08T01:47:16.00573Z","shell.execute_reply.started":"2022-02-08T01:47:11.677025Z","shell.execute_reply":"2022-02-08T01:47:16.004931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get the 12 most popular items in the last month\n\nhttps://www.kaggle.com/hengzheng/time-is-our-best-friend","metadata":{}},{"cell_type":"code","source":"top_12_items = df[df['t_dat'] > '2020-08-21'].groupby('article_id')['customer_id'].nunique().sort_values(ascending=False).head(12).index.tolist()\ntop_12_items","metadata":{"execution":{"iopub.status.busy":"2022-02-08T01:49:41.386094Z","iopub.execute_input":"2022-02-08T01:49:41.386939Z","iopub.status.idle":"2022-02-08T01:49:49.931013Z","shell.execute_reply.started":"2022-02-08T01:49:41.386881Z","shell.execute_reply":"2022-02-08T01:49:49.930502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submit","metadata":{}},{"cell_type":"code","source":"# Predict the same for everyone\ndf_sub['prediction'] =  ' '.join(top_12_items)\ndf_sub.to_csv('submission.csv', index=False)\ndf_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T01:28:30.566511Z","iopub.execute_input":"2022-02-08T01:28:30.567312Z","iopub.status.idle":"2022-02-08T01:28:43.676827Z","shell.execute_reply.started":"2022-02-08T01:28:30.567268Z","shell.execute_reply":"2022-02-08T01:28:43.676199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Condensed in 5 lines of code:\n\n```python\nimport pandas as pd\ndf = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv', dtype={'article_id': str})\ndf_sub = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\ndf_sub['prediction'] =  ' '.join(df[df['t_dat'] > '2020-08-21'].groupby('article_id')['customer_id'].count().sort_values(ascending=False).head(12).index.tolist())\ndf_sub.to_csv('submission.csv', index=False)\n```","metadata":{}},{"cell_type":"markdown","source":"# Please, _DO_ upvote if you find this kernel useful or interesting!","metadata":{}}]}