{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### The main consept and code comes from [the great work by Mehran Kazeminia](https://www.kaggle.com/code/mehrankazeminia/gsdc22-coordinate-with-nearestneighbors) in GSDC2022 competition. Thank you very much.\n> - Usually when you \"Ensemble\" between the results of two or more notebooks, you specify coefficients that are multiplied in all rows. In this way, the results will probably get better in some rows, and at the same time, in some other rows, the results may get worse. However, if the results are generally better, we consider this \"ensemble\" successful.\n> - In any case, we should not forget that finding the coefficients for \"Ensembling\" with our eyes closed will probably make the results of some rows worse, but sometimes we do not realize this, because our overall score has improved anyway.\n> - In this notebook, we will share our innovative method for \"Coordinate [One by One]\" the results and you will see that for each row, we perform separate calculations and And we determine the order of proximity of all points in a row. Then we can use the point that has the highest score in this row as the main basis and, for example, combine the value of this point with the point closest to itself (Blend or Snap).\n> <img src=\"https://raw.githubusercontent.com/MehranKazeminia/fifa-worldcup-2018/master/dart101.png\">\n\n> - Suppose in a match, shooters throw their arrows at a target. Then, for some reason, the point that marks the center of the target disappears. If the number of arrows is large enough, the actual position of the target can be regained.\n\n### There are a lot of public notebooks which employ ensemble. I suppose that modification of the ensemble method itself is much more meaningful and exciting than just varying coefficients.","metadata":{}},{"cell_type":"markdown","source":"### **Note**: The base prediction files are dummy. Please replace them by other files as you like.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport glob\nimport pprint\nfrom tqdm import tqdm\nfrom sklearn.neighbors import NearestNeighbors\n\ndebug = False","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:38:20.412512Z","iopub.execute_input":"2022-08-06T00:38:20.413812Z","iopub.status.idle":"2022-08-06T00:38:20.420436Z","shell.execute_reply.started":"2022-08-06T00:38:20.413756Z","shell.execute_reply":"2022-08-06T00:38:20.418945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"paths = [\"../input/amex-dummy-predictions/dummy_submit_0.csv\", # Please replace them with the predictions you want to use.\n         \"../input/amex-dummy-predictions/dummy_submit_1.csv\",\n         \"../input/amex-dummy-predictions/dummy_submit_2.csv\",\n         \"../input/amex-dummy-predictions/dummy_submit_3.csv\", \n         \"../input/amex-dummy-predictions/dummy_submit_4.csv\"  # The last file should be the best one (: best score before ensembling).\n        ]\nn_files=len(paths)\nfor i in range(n_files):\n    print(i, \" : \",paths[i])\n    \nsubmit = pd.read_csv('../input/amex-default-prediction/sample_submission.csv')\nlen_sub=len(submit)\nif debug:\n    len_sub=1000\n    submit=submit[0:len_sub]","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:38:20.457889Z","iopub.execute_input":"2022-08-06T00:38:20.458796Z","iopub.status.idle":"2022-08-06T00:38:21.870076Z","shell.execute_reply.started":"2022-08-06T00:38:20.458749Z","shell.execute_reply":"2022-08-06T00:38:21.868585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"QT = [[] for _ in range(n_files)]\n\nfor k in range(len(paths)):    \n    sub_k = pd.read_csv(paths[k],nrows=len_sub).values  \n    PT = []\n    for j in range(len_sub):\n        PT.append([sub_k[j][1]])      \n    QT[k] = PT  \n\ncounts_1 = {i: 0 for i in list(range(n_files))}\ncounts_2 = {i: 0 for i in list(range(n_files))}\nT = []\nfor i in tqdm(range(len_sub)): \n    XT = [QT[j][i] for j in range(n_files)]\n    nbrs = NearestNeighbors(n_neighbors=len(XT), algorithm='ball_tree').fit(XT)    \n    _ , indices_T = nbrs.kneighbors(XT)\n    tt = (0.85 * XT[indices_T[-1][0]][0]) + (0.10 * XT[indices_T[-1][1]][0]) + (0.05 * XT[indices_T[-1][2]][0]) #Weights among best/nearest/2nd-nearest predictions. Should be tuned.\n    T.append(tt)\n    counts_1[indices_T[-1][1]] += 1\n    counts_2[indices_T[-1][2]] += 1","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:38:21.872603Z","iopub.execute_input":"2022-08-06T00:38:21.873081Z","iopub.status.idle":"2022-08-06T00:38:51.711098Z","shell.execute_reply.started":"2022-08-06T00:38:21.873040Z","shell.execute_reply":"2022-08-06T00:38:51.708975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Number of the predictions nearest to the best model:\") \npprint.pprint(counts_1)\nprint(\"Number of the predictions 2nd nearest to the best model:\") \npprint.pprint(counts_2)\nsubmit['prediction']  = T\nsubmit.to_csv('submission.csv', index=None)\nprint(submit.describe())\nsubmit","metadata":{"execution":{"iopub.status.busy":"2022-08-06T00:38:51.712442Z","iopub.status.idle":"2022-08-06T00:38:51.713307Z","shell.execute_reply.started":"2022-08-06T00:38:51.712924Z","shell.execute_reply":"2022-08-06T00:38:51.712959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}}]}