{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Exponential weighted ensemble**","metadata":{}},{"cell_type":"markdown","source":"inspired on this code https://www.kaggle.com/code/heyspaceturtle/getting-my-life-together","metadata":{}},{"cell_type":"markdown","source":"The forumla used is: exp(b*(x-S))\n\nx is the LB score of each model. Larger scores are better, hence the weights get larger as x increases and get smaller as x decreases (The better the model, the bigger the weight).\n\nb is the one meaningfully adjustable parameter of the model, the larger b is then the faster the weights decay as the score gets worse. S is a calibration parameter defined such that if S is set to the best single model score then the highest unnormalised weight exp(b*(x-S)) is 1.0, which is convenient, but not essential.\n\nq is the sum of the unnormalised weights. Once all weights have been calculated, these weights are normalised by dividing them all by q.","metadata":{}},{"cell_type":"markdown","source":"**Notebooks used**","metadata":{}},{"cell_type":"markdown","source":"https://www.kaggle.com/code/swimmy/tuffline-amex-anotherfeaturelgbm\n\nhttps://www.kaggle.com/code/adhithyasrinivasan/amex-feature-ensemble\n\nhttps://www.kaggle.com/code/zb1373/blend-boosting-study\n\nhttps://www.kaggle.com/code/manavtrivedi/tuffline-plotly-amex","metadata":{}},{"cell_type":"markdown","source":"**ensemble**","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport glob","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:36.266475Z","iopub.execute_input":"2022-08-15T12:32:36.266787Z","iopub.status.idle":"2022-08-15T12:32:36.271294Z","shell.execute_reply.started":"2022-08-15T12:32:36.266764Z","shell.execute_reply":"2022-08-15T12:32:36.270401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"b = 4000.0\nS = 0.799\nq = 0.0","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:36.297679Z","iopub.execute_input":"2022-08-15T12:32:36.298284Z","iopub.status.idle":"2022-08-15T12:32:36.302696Z","shell.execute_reply.started":"2022-08-15T12:32:36.298247Z","shell.execute_reply":"2022-08-15T12:32:36.301881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = pd.read_csv(\"../input/amex-default-prediction/sample_submission.csv\")\ndisplay(sub)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:36.332474Z","iopub.execute_input":"2022-08-15T12:32:36.332802Z","iopub.status.idle":"2022-08-15T12:32:37.352897Z","shell.execute_reply.started":"2022-08-15T12:32:36.332777Z","shell.execute_reply":"2022-08-15T12:32:37.352011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scores = [0.799,0.799,0.799,0.799]","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:37.354235Z","iopub.execute_input":"2022-08-15T12:32:37.354607Z","iopub.status.idle":"2022-08-15T12:32:37.357976Z","shell.execute_reply.started":"2022-08-15T12:32:37.354583Z","shell.execute_reply":"2022-08-15T12:32:37.357265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfs = [\n    \"../input/tuffline-plotly-amex/submission.csv\",\n    \"../input/blend-boosting-study/submission.csv\",\n    \"../input/amex-feature-ensemble/submission.csv\",\n    \"../input/tuffline-amex-anotherfeaturelgbm/submission.csv\"\n]","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:37.358873Z","iopub.execute_input":"2022-08-15T12:32:37.359235Z","iopub.status.idle":"2022-08-15T12:32:37.370321Z","shell.execute_reply.started":"2022-08-15T12:32:37.359214Z","shell.execute_reply":"2022-08-15T12:32:37.369140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for df,score in zip(dfs,scores):\n    df = pd.read_csv(df)\n    sub[\"prediction\"] = sub[\"prediction\"]+df[\"prediction\"]*np.exp(b*(score-S))\n    q = q+np.exp(b*(score-S))\n    \nsub[\"prediction\"] = sub[\"prediction\"]/q\nprint(q)\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:37.372224Z","iopub.execute_input":"2022-08-15T12:32:37.373255Z","iopub.status.idle":"2022-08-15T12:32:41.420004Z","shell.execute_reply.started":"2022-08-15T12:32:37.373230Z","shell.execute_reply":"2022-08-15T12:32:41.418871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv(\"submission.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-15T12:32:41.421416Z","iopub.execute_input":"2022-08-15T12:32:41.422243Z","iopub.status.idle":"2022-08-15T12:32:44.084157Z","shell.execute_reply.started":"2022-08-15T12:32:41.422210Z","shell.execute_reply":"2022-08-15T12:32:44.083177Z"},"trusted":true},"execution_count":null,"outputs":[]}]}