{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Thanks @mehrankazeminia for your great [work](https://www.kaggle.com/code/mehrankazeminia/gsdc22-coordinate-with-nearestneighbors)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkcyan;\">Google Smartphone Decimeter Challenge 2022</h1>  \n</div>\n\n<div>\n    <h1 align=\"center\" style=\"color:darkcyan;\"><< One by One >></h1>\n    <h1 align=\"center\" style=\"color:darkcyan;\">Coordinate with NearestNeighbors</h1>    \n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"#### - Usually when you \"Ensemble\" between the results of two or more notebooks, you specify coefficients that are multiplied in all rows. In this way, the results will probably get better in some rows, and at the same time, in some other rows, the results may get worse. However, if the results are generally better, we consider this \"ensemble\" successful.\n\n#### - In any case, we should not forget that finding the coefficients for \"Ensembling\" with our eyes closed will probably make the results of some rows worse, but sometimes we do not realize this, because our overall score has improved anyway. \n\n#### - By the way, if the results of the notebooks are two or more dimensions, it will be really harder to choose the coefficients for \"Ensembling\", and only by guessing or a lot of trial and error, maybe it can be successful.\n\n#### - In this notebook, we will share our innovative method for \"Coordinate [One by One]\" the results and you will see that for each row, we perform separate calculations and And we determine the order of proximity of all points in a row. Then we can use the point that has the highest score in this row as the main basis and, for example, combine the value of this point with the point closest to itself (Blend or Snap).\n\n#### - We use \"NearestNeighbors\" for each row, which makes the calculations a bit slow. Of course you can use other methods.","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://raw.githubusercontent.com/MehranKazeminia/fifa-worldcup-2018/master/dart101.png\">","metadata":{}},{"cell_type":"markdown","source":"## **<span style=\"color:darkred;\">Adolphe Quetelet (1796-1874):</span>**\n\n#### Suppose in a match, shooters throw their arrows at a target. Then, for some reason, the point that marks the center of the target disappears. If the number of arrows is large enough, the actual position of the target can be regained.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport seaborn as sns\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\n#!ls ../input/*","metadata":{"execution":{"iopub.status.busy":"2022-07-22T22:15:49.949352Z","iopub.execute_input":"2022-07-22T22:15:49.950308Z","iopub.status.idle":"2022-07-22T22:15:51.164149Z","shell.execute_reply.started":"2022-07-22T22:15:49.950203Z","shell.execute_reply":"2022-07-22T22:15:51.163050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:darkcyan;\">Data-Sets</span>\n\n#### To illustrate this, I used the results of my notebook and the results of five public notebooks. Thanks to everyone who published their great notebooks, especially:\n\n##### @**saitodevel01**, @**robikscube**, @**dienhoa**, @**ravishah1**, @**taroz1461**, @**saurabhbagchi**","metadata":{}},{"cell_type":"code","source":"SAMPLE = pd.read_csv('../input/smartphone-decimeter-2022/sample_submission.csv')\ndisplay(SAMPLE)","metadata":{"execution":{"iopub.status.busy":"2022-07-22T22:15:51.165934Z","iopub.execute_input":"2022-07-22T22:15:51.167957Z","iopub.status.idle":"2022-07-22T22:15:51.314848Z","shell.execute_reply.started":"2022-07-22T22:15:51.167912Z","shell.execute_reply":"2022-07-22T22:15:51.313949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path1 = '../input/gsdc224581/submission.csv' \npath2 = '../input/gsdc224376/submission.csv' \npath3 = '../input/gsdc223355/submission.csv'\npath4 = '../input/carriersmoothingrobust-submission-score-3013/submission.csv'\n\npath  = [path1, path2, path3, path4]","metadata":{"execution":{"iopub.status.busy":"2022-07-22T22:15:51.316377Z","iopub.execute_input":"2022-07-22T22:15:51.316984Z","iopub.status.idle":"2022-07-22T22:15:51.322392Z","shell.execute_reply.started":"2022-07-22T22:15:51.316938Z","shell.execute_reply":"2022-07-22T22:15:51.321442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n# <span style=\"color:darkcyan;\">Coordinate - One by One</span>","metadata":{}},{"cell_type":"code","source":"QT = [[], [], [], []]\nQN = [[], [], [], []]\n\nfor k in range(len(path)):    \n    sub_k = pd.read_csv(path[k]).values  \n    PT = []\n    PN = []    \n    for j in range(len(SAMPLE)):\n        PT.append([sub_k[j][2]])     \n        PN.append([sub_k[j][3]])   \n    QT[k] = PT  \n    QN[k] = PN  ","metadata":{"execution":{"iopub.status.busy":"2022-07-22T22:15:51.324910Z","iopub.execute_input":"2022-07-22T22:15:51.325422Z","iopub.status.idle":"2022-07-22T22:15:52.818390Z","shell.execute_reply.started":"2022-07-22T22:15:51.325377Z","shell.execute_reply":"2022-07-22T22:15:52.817419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def near_plt(points, best_score, support, best_1, generated):\n    \n    plt.style.use('seaborn-whitegrid') \n    plt.figure(figsize=(10, 10), facecolor='lightblue')\n    plt.title(f'\\nC O O R D I N A T E\\n\\n{SAMPLE.iloc[i][:2]}')   \n    \n    plt.scatter(points[0], points[1], s=200, facecolor='lightblue', edgecolor='black', label='All Points')\n    plt.scatter(best_score[0], best_score[1], s=200, facecolor='violet', edgecolor='black', label='Best Score')\n    plt.scatter(support[0], support[1], s=200, facecolor='yellow', edgecolor='black', label='Support')    \n    plt.scatter(generated[0], generated[1], s=150, marker='x', label='Generated')\n    plt.scatter(best_1[0], best_1[1], s=150, marker='x', label='Best-1 (To Check)')\n   \n    plt.legend(fontsize=12)\n    plt.xlabel('LatitudeDegrees', fontsize=12)\n    plt.ylabel('LongitudeDegrees', fontsize=12)\n    plt.savefig(f'Coordinate_{i}.png')\n    plt.show()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-07-22T22:15:52.819708Z","iopub.execute_input":"2022-07-22T22:15:52.820179Z","iopub.status.idle":"2022-07-22T22:15:52.832836Z","shell.execute_reply.started":"2022-07-22T22:15:52.820133Z","shell.execute_reply":"2022-07-22T22:15:52.832115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **<span style=\"color:darkred;\">Note:</span>**\n\n#### - You can compare the results of your notebook with many other notebooks. You can even use one or two or ... the nearest points. Then \"Coordinate\" each row of your results. Also, easily try different coefficients.\n\n#### - If the red cross (Best-1) is on the yellow circle, it means that the point closest to Best_Score is the point that scores exactly after Best_Score. But if the red cross (Best-1) is on the blue circle, that is, in this row, a point with a lower score is closer to the Best_Score point. Of course, we did not know about it before, but now we can use the same point to combine the results of this.\n\n#### - In order to be more accurate, in version 9, I made some changes in the calculations. I have done calculations for LatitudeDegrees and LongitudeDegrees columns separately. In addition, (as I explained earlier) individual rows are also calculated separately. For this reason, the execution time has been almost doubled.","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import NearestNeighbors\nrandom_examples = np.random.randint(len(SAMPLE), size=10) # Number of examples: 10\n\nT = [] # LatitudeDegrees\nN = [] # LongitudeDegrees\n\nfor i in range(len(SAMPLE)): \n    XT = [QT[0][i], QT[1][i], QT[2][i], QT[3][i]]\n    XN = [QN[0][i], QN[1][i], QN[2][i], QN[3][i]]\n    \n    nbrs = NearestNeighbors(n_neighbors=len(XT), algorithm='ball_tree').fit(XT)    \n    _ , indices_T = nbrs.kneighbors(XT)\n    \n    nbrs = NearestNeighbors(n_neighbors=len(XN), algorithm='ball_tree').fit(XN)    \n    _ , indices_N = nbrs.kneighbors(XN)\n \n    tt = (1.13 * XT[indices_T[-1][0]][0]) + (-0.16 * XT[indices_T[-1][1]][0]) + (0.03 * XT[indices_T[-1][2]][0])\n    T.append(tt) \n    \n    nn = (1.29 * XN[indices_N[-1][0]][0]) + (-0.29 * XN[indices_N[-1][1]][0]) + (0.00 * XN[indices_N[-1][2]][0])    \n    N.append(nn) \n    \n    \"\"\"\n    if i in random_examples:\n        print(f'\\n\\n\\nRow >>>>> {i}\\n{SAMPLE.iloc[i][:2]}')\n        near_plt([XT, XN], [XT[indices_T[-1][0]], XN[indices_N[-1][0]]],\n                 [XT[indices_T[-1][1]], XN[indices_N[-1][1]]], [XT[-2], XN[-2]], [tt, nn])\n    \"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-07-22T22:15:52.834102Z","iopub.execute_input":"2022-07-22T22:15:52.834610Z","iopub.status.idle":"2022-07-22T22:17:19.755299Z","shell.execute_reply.started":"2022-07-22T22:15:52.834577Z","shell.execute_reply":"2022-07-22T22:17:19.754426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = SAMPLE.copy()\nsub['LatitudeDegrees']  = T\nsub['LongitudeDegrees'] = N\nsub","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-07-22T22:17:19.756674Z","iopub.execute_input":"2022-07-22T22:17:19.757068Z","iopub.status.idle":"2022-07-22T22:17:19.801684Z","shell.execute_reply.started":"2022-07-22T22:17:19.757009Z","shell.execute_reply":"2022-07-22T22:17:19.800761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv(\"submission.csv\", index=False)\n!ls","metadata":{"execution":{"iopub.status.busy":"2022-07-22T22:17:19.803116Z","iopub.execute_input":"2022-07-22T22:17:19.803557Z","iopub.status.idle":"2022-07-22T22:17:21.070297Z","shell.execute_reply.started":"2022-07-22T22:17:19.803514Z","shell.execute_reply":"2022-07-22T22:17:21.069106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>\n\n<div class=\"alert alert-success\">\n    <h1 align=\"center\" style=\"color:darkcyan;\">Good Luck.</h1> \n</div>","metadata":{"_kg_hide-output":true}}]}