{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\">Smart Ensembling</h1>    \n    <h1 align=\"center\">G2Net Gravitational Wave Detection</h1> \n    <h4 align=\"center\">By: Somayyeh Gholami & Mehran Kazeminia</h4>\n    <h5 align=\"center\">Modified by: kuroyuli</h4>\n</div>","metadata":{"_uuid":"ddb5388e-e8a7-4738-9e59-02d3f59acd26","_cell_guid":"5105fb34-d535-4af0-b929-ce9235839f80","trusted":true}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">\n    <h1 align=\"center\">If you find this work useful, please don't forget upvoting :)</h1>\n</div>","metadata":{"_uuid":"6b47d287-bfc1-4e44-8cea-b1592a601c13","_cell_guid":"aa8a9c10-1e22-46f3-a2e2-10f083207680","trusted":true}},{"cell_type":"markdown","source":"## Below code is a modification of great notebook by [Somayyeh Gholami](https://www.kaggle.com/somayyehgholami) & [Mehran Kazeminia](https://www.kaggle.com/mehrankazeminia).<br> Check their [original work](https://www.kaggle.com/somayyehgholami/1-g2net-smart-ensembling).\n\nI want more control over how to ensemble models.  Therefore, I modified the code so that I can change the coefficient for each area in the scatter plot of the main and support model prediction value.\n\nSomayyeh Gholami & Mehran Kazeminiaさんが作ったアンサンブルのためのコードを一部変更しました。<br>\n元コードはidの先頭文字によって係数を変えているのですが、ランダム性を加える以上の意味が分かりませんでした。<br>\nそこでアンサンブルに使うメインモデルとサポートモデルの予測値分布を表す散布図のエリアごとに係数を変えられるようにしました。<br>\n（メインモデルはそれまでのアンサンブル・モデル、サポートモデルは新しくアンサンブルに入れるモデルを表します。）","metadata":{"_uuid":"0a8738bf-edde-4fba-ae2a-d79af9167d04","_cell_guid":"064fd12d-2bce-4b99-8585-f08230d09906","trusted":true}},{"cell_type":"markdown","source":"## Areas in the scatter plot","metadata":{"_uuid":"75bd2f74-4bad-4ebc-9560-5fc202604529","_cell_guid":"cc16b637-da17-461b-ae01-9a319ee218ce","trusted":true}},{"cell_type":"code","source":"import numpy as np \nimport matplotlib.pyplot as plt\nfrom PIL import Image\n\n%matplotlib inline\nim = Image.open(\"../input/coef-pic/coef.png\")\n\nplt.figure(figsize=(10, 10), dpi=50)\nim_list = np.asarray(im)\nplt.imshow(im_list)\nplt.show()","metadata":{"_uuid":"4a5491fb-c028-49bf-975b-5ff3c6b9ec06","_cell_guid":"7f1565f5-9f7f-4b43-81d0-fa7056defea1","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:38.829881Z","iopub.execute_input":"2021-08-22T00:00:38.830338Z","iopub.status.idle":"2021-08-22T00:00:39.223799Z","shell.execute_reply.started":"2021-08-22T00:00:38.830247Z","shell.execute_reply":"2021-08-22T00:00:39.222739Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import","metadata":{"_uuid":"79aaed43-020a-4a71-9499-8dad97a249a0","_cell_guid":"c1bdd216-3bbc-4fdc-8832-1d60b9f25d6b","trusted":true}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\n\nimport plotly.figure_factory as ff\nimport plotly.express as px\n\nfrom sklearn.metrics import roc_auc_score, roc_curve, auc","metadata":{"_uuid":"8ff0871f-0ae6-4c8b-94a0-35e9579f59a4","_cell_guid":"c1f48bb8-f644-4ac9-a651-f6b6e393d967","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:39.225666Z","iopub.execute_input":"2021-08-22T00:00:39.226094Z","iopub.status.idle":"2021-08-22T00:00:43.10777Z","shell.execute_reply.started":"2021-08-22T00:00:39.226056Z","shell.execute_reply":"2021-08-22T00:00:43.106887Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{"_uuid":"b5dfdfee-c22c-4c1e-a74b-aa32bd3ae2c6","_cell_guid":"e8418a94-6146-40a4-b92d-6f7d03bcf0da","trusted":true}},{"cell_type":"markdown","source":"## Functions","metadata":{"_uuid":"3911d2a3-468f-4358-90f4-b93b4253e76f","_cell_guid":"493b8c76-a946-44c2-9963-46525dfeb48a","trusted":true}},{"cell_type":"code","source":"def get_var_name(var):\n    for k,v in globals().items():\n        if id(v) == id(var):\n            name=k\n    return name","metadata":{"_uuid":"aac313a5-cdc2-4fa7-bd7d-e17af5c0ba02","_cell_guid":"a5d1ae1e-7f19-4dc8-a284-1634e3a6f944","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:43.109312Z","iopub.execute_input":"2021-08-22T00:00:43.109732Z","iopub.status.idle":"2021-08-22T00:00:43.113892Z","shell.execute_reply.started":"2021-08-22T00:00:43.109702Z","shell.execute_reply":"2021-08-22T00:00:43.112887Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def ensembling(main, support, coef1, coef2, coef3, coef4, coef5, coef6, coef7, coef8):\n    \n    main_name = get_var_name(locals().get('main'))\n    supp_name = get_var_name(locals().get('support'))\n    \n    mod_a  = main.copy()\n    mod_av = mod_a.values\n    \n    mod_b  = support.copy()\n    mod_bv = mod_b.values    \n              \n    ense  = main.copy()\n    ense_v = ense.values\n \n    for i in range (len(main)):       \n        diff = mod_bv[i, 1] - mod_av[i, 1]\n        pred_a = mod_av[i, 1]       \n        pred_b = mod_bv[i, 1] \n        \n        if ((pred_a < 0.25) & (diff >= 0)):        \n            pred = (pred_a * coef1) + (pred_b * (1.0 - coef1))\n\n        elif ((pred_a < 0.25) & (diff < 0)):        \n            pred = (pred_a * coef2) + (pred_b * (1.0 - coef2))\n            \n        elif ((pred_a >= 0.25) & (pred_a < 0.5) & (diff >= 0)):        \n            pred = (pred_a * coef3) + (pred_b * (1.0 - coef3))\n\n        elif ((pred_a >= 0.25) & (pred_a < 0.5) & (diff < 0)):        \n            pred = (pred_a * coef4) + (pred_b * (1.0 - coef4))   \n                      \n        elif ((pred_a >= 0.5) & (pred_a < 0.75) & (diff >= 0)):        \n            pred = (pred_a * coef5) + (pred_b * (1.0 - coef5))\n\n        elif ((pred_a >= 0.5) & (pred_a < 0.75) & (diff < 0)):        \n            pred = (pred_a * coef6) + (pred_b * (1.0 - coef6))\n            \n        elif ((pred_a >= 0.75) & (diff >= 0)):        \n            pred = (pred_a * coef7) + (pred_b * (1.0 - coef7))\n\n        elif ((pred_a >= 0.75) & (diff < 0)):        \n            pred = (pred_a * coef8) + (pred_b * (1.0 - coef8))\n        \n        else:\n            raise ValueError(\"if sentence error\")\n           \n        ense_v[i, 1] = pred\n        \n    ense.iloc[:, 1] = ense_v[:, 1]\n\n    ###############################    \n    X  = mod_a.iloc[:, 1]\n    Y1 = mod_b.iloc[:, 1]\n    Y2 = ense.iloc[:, 1]\n    \n    plt.style.use('seaborn-whitegrid') \n    plt.figure(figsize=(9, 9), facecolor='lightgray')\n    plt.title(f'\\nE N S E M B L I N G\\n')   \n      \n    plt.scatter(X, Y1, s=1.5, label=f'Support: {supp_name} {support.score}')    \n    plt.scatter(X, Y2, s=1.5, label='Generated')\n    plt.scatter(X, X, s=0.1, label=f'Main: {main_name}')\n    \n    plt.legend(fontsize=12, loc=2)\n    #plt.savefig('Ensembling_1.png')\n    plt.show()     \n    ###############################   \n    ense.iloc[:, 1] = ense.iloc[:, 1].astype(float)\n    hist_data = [mod_b.iloc[:, 1], ense.iloc[:, 1], mod_a.iloc[:, 1]] \n    group_labels = [f'Support: {supp_name} {support.score}', 'Ensembling', f'Main: {main_name}']\n    \n    fig = ff.create_distplot(hist_data, group_labels, bin_size=.2, show_hist=False, show_rug=False)\n    fig.show()   \n    ###############################   \n    \n    return ense","metadata":{"_uuid":"cae9a85d-471b-48a6-98fb-9c47c5ca8371","_cell_guid":"bd1c7a4a-9639-4f06-ab10-259a0e01fb1e","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:43.115553Z","iopub.execute_input":"2021-08-22T00:00:43.116067Z","iopub.status.idle":"2021-08-22T00:00:43.137738Z","shell.execute_reply.started":"2021-08-22T00:00:43.115923Z","shell.execute_reply":"2021-08-22T00:00:43.136866Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{"_uuid":"02281e72-950d-4411-9685-72410a8e1e94","_cell_guid":"695cc1b8-2212-4bf6-be21-b84719ae8410","trusted":true}},{"cell_type":"markdown","source":"## Data Set","metadata":{"_uuid":"d858ee55-fb3e-4997-9cd0-fd02291b482d","_cell_guid":"c4b060c0-627a-4f66-aab9-f9e604e2df04","trusted":true}},{"cell_type":"markdown","source":"Revise the path and its Public Score when you replace the model.  Put '860b' and so forth, if you have the same score.<br>\nモデルを入れ替える場合、パスとスコアを修正して下さい。同じスコアのモデルがある場合、'860b'などとして区別して下さい。","metadata":{"_uuid":"dfff95c5-fd90-4a40-9d7a-517dade27f59","_cell_guid":"bfc8c4d0-9df7-4852-b3ea-ed340b77c10d","execution":{"iopub.status.busy":"2021-08-19T23:33:54.524989Z","iopub.execute_input":"2021-08-19T23:33:54.525465Z","iopub.status.idle":"2021-08-19T23:33:54.538746Z","shell.execute_reply.started":"2021-08-19T23:33:54.525433Z","shell.execute_reply":"2021-08-19T23:33:54.536735Z"},"trusted":true}},{"cell_type":"markdown","source":"Thanks to: @miklgr500 https://www.kaggle.com/miklgr500/g2net-efficientnetb1-tpu-evaluate/output","metadata":{"_uuid":"05a96b26-3273-46b4-b002-6e5685031088","_cell_guid":"75ec0c81-939a-48b6-8e1d-c7db1894c863","trusted":true}},{"cell_type":"code","source":"path0 = '../input/g2net-834/submission.csv'\n\nmodel0 = pd.read_csv(path0).sort_values('id')\nmodel0.score = '834a'","metadata":{"_uuid":"f2dc2ab1-ef55-4d26-8274-23e6f48162ea","_cell_guid":"0aacc3ec-c89c-48c0-b4c0-e29a68d47a0d","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:43.138768Z","iopub.execute_input":"2021-08-22T00:00:43.139861Z","iopub.status.idle":"2021-08-22T00:00:43.714701Z","shell.execute_reply.started":"2021-08-22T00:00:43.139828Z","shell.execute_reply":"2021-08-22T00:00:43.713845Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: @mrigendraagrawal https://www.kaggle.com/mrigendraagrawal/tf-g2net-eda-and-starter","metadata":{"_uuid":"9b2fcf2e-cbfb-4d76-858b-fd5838dbb8d5","_cell_guid":"99458784-2cd1-4a57-b995-57d2fdfa5094","trusted":true}},{"cell_type":"code","source":"path1 = '../input/g2net-855a/submission.csv'\n\nmodel1 = pd.read_csv(path1).sort_values('id')\nmodel1.score = '855a'","metadata":{"_uuid":"54fcf6a2-ed9a-47a2-ad29-c6163f32db5a","_cell_guid":"13d38e7e-06b1-4592-88c8-0bcd5123183c","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:43.715705Z","iopub.execute_input":"2021-08-22T00:00:43.716099Z","iopub.status.idle":"2021-08-22T00:00:44.220028Z","shell.execute_reply.started":"2021-08-22T00:00:43.71607Z","shell.execute_reply":"2021-08-22T00:00:44.219035Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: @yasufuminakama https://www.kaggle.com/yasufuminakama/g2net-efficientnet-b7-baseline-inference","metadata":{"_uuid":"4a993009-fad7-4603-b542-42b31478c2c8","_cell_guid":"a923b033-6393-482f-84af-4f9ce68e5860","trusted":true}},{"cell_type":"code","source":"path2 = '../input/g2net-860/submission.csv'\n\nmodel2 = pd.read_csv(path2).sort_values('id')\nmodel2.score = '860a'","metadata":{"_uuid":"79643152-d083-4921-905e-a9ae2eabff37","_cell_guid":"efa2506e-9cf2-4824-a389-d4399938e130","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:44.221559Z","iopub.execute_input":"2021-08-22T00:00:44.222013Z","iopub.status.idle":"2021-08-22T00:00:44.741703Z","shell.execute_reply.started":"2021-08-22T00:00:44.221969Z","shell.execute_reply":"2021-08-22T00:00:44.740833Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: @wabinab https://www.kaggle.com/wabinab/submission-baseline","metadata":{"_uuid":"52e995fd-2cc9-486a-8084-9a3799ad7371","_cell_guid":"2cb11031-248f-4313-be27-e92eb9c34c68","trusted":true}},{"cell_type":"code","source":"path3 = '../input/g2net-861/submission.csv'\n\nmodel3 = pd.read_csv(path3).sort_values('id')\nmodel3.score = '861a'","metadata":{"_uuid":"7ada6a45-5857-41da-b817-870a87ce376c","_cell_guid":"652a12b3-5e6a-4454-81ad-4ae0166413bd","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:44.743696Z","iopub.execute_input":"2021-08-22T00:00:44.744117Z","iopub.status.idle":"2021-08-22T00:00:45.241702Z","shell.execute_reply.started":"2021-08-22T00:00:44.744087Z","shell.execute_reply":"2021-08-22T00:00:45.240819Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: @ihelon https://www.kaggle.com/ihelon/g2net-eda-and-modeling/output?select=model_submission.csv","metadata":{"_uuid":"1ab5ab25-b383-4685-bda1-1352b77eecd9","_cell_guid":"b321fcf5-8b05-4a43-83de-d7a61fdcafc4","trusted":true}},{"cell_type":"code","source":"path4 = '../input/g2net-864/model_submission.csv'\n\nmodel4 = pd.read_csv(path4).sort_values('id')\nmodel4.score = '864a'","metadata":{"_uuid":"c03a8360-1875-4291-a618-a8333341d947","_cell_guid":"d1a5e7b9-9e1b-4ef6-bf63-020f8b64ac72","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:45.242981Z","iopub.execute_input":"2021-08-22T00:00:45.243409Z","iopub.status.idle":"2021-08-22T00:00:45.805364Z","shell.execute_reply.started":"2021-08-22T00:00:45.243379Z","shell.execute_reply":"2021-08-22T00:00:45.804315Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: @miklgr500 https://www.kaggle.com/miklgr500/cqt-g2net-efficientnetb1-tpu-inference/output?select=submission.csv\n\nThanks to: @xuxu1234 https://www.kaggle.com/xuxu1234/lb-0-866-g2net-efficientnetb7-tpu-inference","metadata":{"_uuid":"bdc48b24-ef05-407b-88e6-7d17e20edfe4","_cell_guid":"d19b9dfe-13b9-447d-8875-94f9122d0a6e","trusted":true}},{"cell_type":"code","source":"path5 = '../input/g2net-866/submission.csv'\n\nmodel5 = pd.read_csv(path5).sort_values('id')\nmodel5.score = '866a'","metadata":{"_uuid":"00cdb8a1-17f2-4f8a-8c98-5df58aa7025c","_cell_guid":"8f821672-808f-4876-b3c5-9e9484acc53a","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:45.806764Z","iopub.execute_input":"2021-08-22T00:00:45.807088Z","iopub.status.idle":"2021-08-22T00:00:46.321839Z","shell.execute_reply.started":"2021-08-22T00:00:45.807056Z","shell.execute_reply":"2021-08-22T00:00:46.320651Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks to: @hidehisaarai1213 https://www.kaggle.com/hidehisaarai1213/g2net-tf-on-the-fly-cqt-tpu-inference","metadata":{"_uuid":"02246326-e373-4237-ac92-e7917757f4bf","_cell_guid":"8773c875-63c9-443d-b509-e26f30d9795c","trusted":true}},{"cell_type":"code","source":"path6 = '../input/g2net-869/submission.csv'\n\nmodel6 = pd.read_csv(path6).sort_values('id')\nmodel6.score = '869a'","metadata":{"_uuid":"69949651-d254-40ac-8e83-4f131c2d5b87","_cell_guid":"5caaeba5-e58c-44ab-b449-bb84db60f056","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:46.323482Z","iopub.execute_input":"2021-08-22T00:00:46.323914Z","iopub.status.idle":"2021-08-22T00:00:47.076369Z","shell.execute_reply.started":"2021-08-22T00:00:46.32387Z","shell.execute_reply":"2021-08-22T00:00:47.075293Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_names = [f'm_{model0.score}', f'm_{model1.score}', f'm_{model2.score}', f'm_{model3.score}', f'm_{model4.score}', f'm_{model5.score}', f'm_{model6.score}']\n\nhist_data = [model0.target, model1.target, model2.target, model3.target, model4.target, model5.target, model6.target]  \n   \nfig = ff.create_distplot(hist_data, model_names, bin_size=.2, show_hist=False, show_rug=False) \n\nfig.show()","metadata":{"_uuid":"bf9e6ef1-93ba-4d42-bd53-b5e6019a03a2","_cell_guid":"771f5d27-6e82-4410-95ab-cc7ace85b2f8","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:00:47.077719Z","iopub.execute_input":"2021-08-22T00:00:47.078153Z","iopub.status.idle":"2021-08-22T00:01:05.220656Z","shell.execute_reply.started":"2021-08-22T00:00:47.078114Z","shell.execute_reply":"2021-08-22T00:01:05.219933Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{"_uuid":"46a18e5a-57cf-4040-adf5-eac4af2e8c66","_cell_guid":"a60b92b6-5853-402d-92e7-9df1d7a0250f","trusted":true}},{"cell_type":"markdown","source":"## Model prediction average and correlation table","metadata":{"_uuid":"155e7216-e200-4ce3-a08a-5f65fc4c2994","_cell_guid":"c00284bd-99ab-47e1-a323-09ffe07b4e8b","trusted":true}},{"cell_type":"code","source":"mean_data=np.mean(hist_data, axis=1)\ncorr_data=np.corrcoef([model0.target, model1.target, model2.target, model3.target, model4.target, model5.target, model6.target])\ncorr_data = pd.DataFrame(corr_data)\ncorr = corr_data.applymap(lambda x : 0 if x >= 0.9999 else x)\ncorr_ = corr.copy()\ncorr.columns = model_names\ncorr['model'] = model_names\ncorr['pred_avg'] = mean_data\ncorr['cor_avg'] = (corr_.mean(axis='columns'))*len(corr_)/(len(corr_)-1)\ncorr['cor_max'] = corr_.max(axis='columns')\ncorr","metadata":{"_uuid":"941ad647-3ca4-4993-be5e-7d1d72e6d0ff","_cell_guid":"0e390d20-5086-4eaf-bce8-6375772b3aee","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:01:05.221777Z","iopub.execute_input":"2021-08-22T00:01:05.22226Z","iopub.status.idle":"2021-08-22T00:01:05.304836Z","shell.execute_reply.started":"2021-08-22T00:01:05.222227Z","shell.execute_reply":"2021-08-22T00:01:05.304097Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I prefer this table rather than a heatmap.<br>\nMy strategy is to ensemble the most correlated models first and leave the less correlated and low scoring models for the latter ensemble.<br>In this way, I can judge the model is good for ensemble or not, and remove unnecessary models easily.<br>\nIn this version, I didn't use model0(834a) and model1 (855a).\n\n各モデル同士の相関表です。Heatmapでもいいんですが、細かい違いを知るため表を使いました。<br>\n私の戦略は「最初に相関が高いモデル同士でアンサンブルして、他との相関が低く、スコアも低いモデルは後回しにする。」というものです。<br>他との相関が低く、精度の低いモデルは「精度は低いものの、他のモデルの弱い部分を補ってアンサンブルとして良いモデルを作れるもの」か「精度が低くアンサンブルに悪影響を与えるもの」の可能性があり、それを見極めるのは後の方が良いのではという考えからです。（最初から入れてしまうと、外しにくくなる。）<br>このバージョンでは、model0(834a)とmodel1 (855a)をアンサンブルに加えていません。<br>この戦略が正しいかどうかは分かりません。みなさん試行錯誤してみて下さい。","metadata":{"_uuid":"6c05fc66-d40e-4a39-81d9-df6ca525c570","_cell_guid":"f2b04236-8978-46c2-bd4f-7b9fba8dacd3","trusted":true}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">\n    <h1 align=\"center\">Ensembling</h1>\n</div>","metadata":{"_uuid":"46a8a108-0b5f-4e7c-9e0e-d4006fa5e177","_cell_guid":"a8e8e01a-55e6-422c-b1d8-a6582075e84a","trusted":true}},{"cell_type":"markdown","source":"Large blue area and low correlation reflect big differences between main and support models. They may be an excellent combo for ensemble or one of which is an inferior model.<br>A smaller coefficient widens the orange area, thus dragging the ensemble close to the support model.\n\n青いエリアの広さはmainとsupportモデルの違いの大きさを示す傾向があります。correlationの方がより正確に違いが分かると思います。<br>小さなCoefficientはオレンジのエリアを広くし、アンサンブル・モデルをsupportモデルに近づけます。","metadata":{"_uuid":"97a52b61-cec4-4cfc-bcd0-e233f68b58f3","_cell_guid":"9d1414bf-2180-4e12-9ba6-ac6b302da57d","trusted":true}},{"cell_type":"code","source":"sub1 = ensembling(model2, model3, 0.3, 0.7, 0.3, 0.7, 0.3, 0.7, 0.3, 0.7)\nprint('sub1_average', np.mean(sub1.target))","metadata":{"execution":{"iopub.status.busy":"2021-08-22T01:45:48.783776Z","iopub.execute_input":"2021-08-22T01:45:48.784172Z","iopub.status.idle":"2021-08-22T01:45:57.749242Z","shell.execute_reply.started":"2021-08-22T01:45:48.784128Z","shell.execute_reply":"2021-08-22T01:45:57.748149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub2 = ensembling(sub1, model4, 0.3, 0.8, 0.3, 0.8, 0.2, 0.6, 0.2, 0.6)\nprint('sub2_average', np.mean(sub2.target))","metadata":{"execution":{"iopub.status.busy":"2021-08-22T02:08:50.307623Z","iopub.execute_input":"2021-08-22T02:08:50.308129Z","iopub.status.idle":"2021-08-22T02:08:59.11424Z","shell.execute_reply.started":"2021-08-22T02:08:50.308098Z","shell.execute_reply":"2021-08-22T02:08:59.113405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub3 = ensembling(sub2, model6, 0.3, 0.5, 0.3, 0.5, 0.3, 0.7, 0.6, 0.4)\nprint('sub3_average', np.mean(sub3.target))","metadata":{"execution":{"iopub.status.busy":"2021-08-22T02:17:09.845477Z","iopub.execute_input":"2021-08-22T02:17:09.846093Z","iopub.status.idle":"2021-08-22T02:17:18.702989Z","shell.execute_reply.started":"2021-08-22T02:17:09.846056Z","shell.execute_reply":"2021-08-22T02:17:18.702176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub4 = ensembling(sub3, model5, 0.4, 0.5, 0.4, 0.5, 0.5, 0.7, 0.65, 0.55)\nprint('sub4_average', np.mean(sub4.target))","metadata":{"execution":{"iopub.status.busy":"2021-08-22T02:34:25.127073Z","iopub.execute_input":"2021-08-22T02:34:25.127472Z","iopub.status.idle":"2021-08-22T02:34:34.017277Z","shell.execute_reply.started":"2021-08-22T02:34:25.127439Z","shell.execute_reply":"2021-08-22T02:34:34.016102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub5 = ensembling(sub4, model0, 1, 1, 1, 1, 1, 1, 1, 1)\nprint('sub5_average', np.mean(sub5.target))","metadata":{"execution":{"iopub.status.busy":"2021-08-22T01:52:43.642565Z","iopub.execute_input":"2021-08-22T01:52:43.642937Z","iopub.status.idle":"2021-08-22T01:52:52.537345Z","shell.execute_reply.started":"2021-08-22T01:52:43.642906Z","shell.execute_reply":"2021-08-22T01:52:52.536264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub6 = ensembling(sub5, model1, 1, 1, 1, 1, 1, 1, 1, 1)\nprint('sub6_average', np.mean(sub6.target))","metadata":{"_uuid":"d31d4d82-61ec-43ed-89b0-98d393f4e672","_cell_guid":"cc494097-18bb-474b-976b-d6525c6ab895","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2021-08-22T01:41:54.893263Z","iopub.execute_input":"2021-08-22T01:41:54.893674Z","iopub.status.idle":"2021-08-22T01:42:47.952217Z","shell.execute_reply.started":"2021-08-22T01:41:54.893642Z","shell.execute_reply":"2021-08-22T01:42:47.951077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{"_uuid":"cd6f2681-c390-4afe-9e6a-58572675d898","_cell_guid":"602461c2-e562-4010-89a6-68b312030f6c","trusted":true}},{"cell_type":"markdown","source":"## Submission","metadata":{"_uuid":"1db3701f-b95a-424d-abfd-514723efa586","_cell_guid":"e0bff238-6947-4dab-8534-021a82e5a1c9","trusted":true}},{"cell_type":"markdown","source":"**It's important to make sure the Public Score is getting better for every step (sub1→sub2→...).**<br>\nSubmitting all csv files leverage the value of this ensembling approach, even though it consumes daily submission allowance. If you find the step which worsen the score, you should change the coefficients or remove the model from ensemble.<br>\nI advise you to add \"-sub1\" and so forth to the Submission Description, for your record.<br>\nOut of Fold data may help you not to waste daily submission allowance, but that may be time consuming.\n\n大事なのは、sub1→sub2→...と進んでいくにつれて常にスコアが良くなること。なので、日毎のsubmit制限数を食ってしまうものの、sub1から順番にsubmitしていくことで、このアンサンブル法の真価が発揮できると思われます。スコアが悪くなるステップがあれば、そこで入れたsupportモデルの質・相性が悪いか Coefficientが不適切なので、Coefficientを調整するか、そのモデルをアンサンブルから抜く必要があります。<br>\nsubmitした後、どのcsvを出したのか分からなくなるので、Submission Description欄に\"-sub1\"などと加えておくと便利です。<br>\nOut of Foldデータがあれば、submitせずに確認できますが、全てのモデルに対して用意するのは大変そうですね。","metadata":{"_uuid":"28636755-1773-46ef-837d-709ad67b95aa","_cell_guid":"c336cff4-4aab-4580-8943-5c3938162df4","trusted":true}},{"cell_type":"code","source":"sub1.to_csv(\"submission1.csv\",index=False)\nsub2.to_csv(\"submission2.csv\",index=False)\nsub3.to_csv(\"submission3.csv\",index=False)\nsub4.to_csv(\"submission4.csv\",index=False)\nsub5.to_csv(\"submission5.csv\",index=False)\n\nsub6.to_csv(\"submission_final.csv\",index=False)\n!ls","metadata":{"_uuid":"14c1417a-170a-4193-810b-142e8995311f","_cell_guid":"f65ddd22-1e22-4186-859b-6396aab3aaba","collapsed":false,"execution":{"iopub.status.busy":"2021-08-22T00:01:58.431764Z","iopub.execute_input":"2021-08-22T00:01:58.432223Z","iopub.status.idle":"2021-08-22T00:02:04.272435Z","shell.execute_reply.started":"2021-08-22T00:01:58.432176Z","shell.execute_reply":"2021-08-22T00:02:04.271379Z"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{"_uuid":"60515760-45a0-41db-91b1-d850a2388ef4","_cell_guid":"e66161eb-5ad2-4cf4-bebf-c690985bbdb9","trusted":true}},{"cell_type":"markdown","source":"## Thank you very much for reading.","metadata":{"_uuid":"407bc4be-4980-4028-8288-263a9c92896e","_cell_guid":"d24ac7c8-de5a-4b06-8dba-9525a50f0105","trusted":true}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{"_uuid":"4c2a4ff4-379b-4b19-aacd-9811c6ef8e92","_cell_guid":"32394174-8da6-4d86-a785-5c08cde0cc6f","trusted":true}}]}