{"cells":[{"metadata":{},"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\"> Comparative Method - Part(A)</h1></h1>\n    <h2 align=\"center\">Rainforest Connection Species Audio Detection</h2>\n    <h3 align=\"center\">By: Somayyeh Gholami & Mehran Kazeminia</h3>\n</div>"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{},"cell_type":"markdown","source":"# Description:"},{"metadata":{},"cell_type":"markdown","source":"### - At the end of the challenge, Mr. [@meaninglesslives](https://www.kaggle.com/meaninglesslives) shared his notebook. He won third place in the challenge. The score of the notebook published in the first version is \"public score 0.96171 and private score 0.96460\". Thanks for sharing the results, we congratulate him too.\n\nhttps://www.kaggle.com/meaninglesslives/rfcx-minimal?scriptVersionId=54514070\n\n### - Then Mr. [@cdeotte](https://www.kaggle.com/cdeotte) released another notebook and with a great trick, raised the previous notebook's private score above 0.970. We also thank him for sharing this trick.\n\nhttps://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\n\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\n# An important question:\n\n### Does this trick improve all the results (all the columns)?\n\n### No, of course the results of some columns are getting worse.\n\n### This means that the results of some columns get very good and the results of some columns get worse, but in this challenge (and usually) the overall results improve.\n\n### To prove this, we wrote this notebook and share it with you. Our method is very simple. We first identified nine columns, the results of which will be reduced by performing this trick.\n\n### S0, S1, S4, S8, S11, S14, S18, S20, S21\n\n### Then we transferred the results of these columns from the original notebook (the first notebook) and replaced these results exactly with the results of the second notebook, and finally saved the entire result in the \"d\" file. That is, in file \"d\", a trick is applied for the results of 15 columns and no trick is applied for the results of 9 columns. The scores of the \"d\" file are as follows:\n\n### \"d\" : [(Private Score: 0.97915) , (Public Score: 0.97373)]\n\n### As you can see, this trick is not good for the results of these nine columns, and we got a much better score with the results of the original notebook (first version). Please note that in order to be able to compare, we used exactly the results of the first version of the original notebook.\n\n### In the end, we were able to easily improve the results once again with our own method. We saved the final results in the \"e\" file. The scores of the \"e\" file are as follows:\n\n### \"e\" : [(Private Score: 0.98022) , (Public Score: 0.97490)]\n\n"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{},"cell_type":"markdown","source":"## If you find this work useful, please don't forget upvoting :)"},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{},"cell_type":"markdown","source":"# Import & Data Set"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd \nimport matplotlib.pyplot as plt\n\n%matplotlib inline\n\n# _______________________________\n\nsub961 = pd.read_csv(\"../input/rain961/RAIN961.csv\") \n\nsub968 = pd.read_csv(\"../input/rainforest-post-process-lb-0-970/submission_with_pp.csv\") \n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{},"cell_type":"markdown","source":"# Functions"},{"metadata":{"trusted":true},"cell_type":"code","source":"def generate(main, support, coeff):\n    g1 = main.copy()\n    g2 = main.copy()\n    g3 = main.copy()\n    g4 = main.copy()\n    \n    for i in main.columns[1:]:\n        lm, Is = [], []                \n        lm = main[i].tolist()\n        ls = support[i].tolist() \n        \n        res1, res2, res3, res4 = [], [], [], []          \n        for j in range(len(main)):\n            res1.append(max(lm[j] , ls[j]))\n            res2.append(min(lm[j] , ls[j]))\n            res3.append((lm[j] + ls[j]) / 2)\n            res4.append((lm[j] * coeff) + (ls[j] * (1.- coeff)))\n            \n        g1[i] = res1\n        g2[i] = res2\n        g3[i] = res3\n        g4[i] = res4\n        \n    return g1,g2,g3,g4","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def generate1(main, support, coeff):\n    \n    g = main.copy()    \n    for i in main.columns[1:]:\n        \n        res = []\n        lm, Is = [], []        \n        lm = main[i].tolist()\n        ls = support[i].tolist()  \n        \n        for j in range(len(main)):\n            res.append((lm[j] * coeff[i]) + (ls[j] * (1.- coeff[i])))            \n        g[i] = res\n        \n    return g","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def drawing(main, support, generated, column_number):\n    \n    X  = main.iloc[:, column_number]\n    Y1 = support.iloc[:, column_number]\n    Y2 = generated.iloc[:, column_number]\n    \n    plt.style.use('seaborn-whitegrid') \n    plt.figure(figsize=(8, 8), facecolor='lightgray')\n    plt.title(f'\\nOn the X axis >>> main\\n\\nOn the Y axis >>> support\\n')           \n    plt.scatter(X, Y1, s=3)\n    plt.show() \n    \n    plt.style.use('seaborn-whitegrid') \n    plt.figure(figsize=(8, 8), facecolor='lightgray')\n    plt.title(f'\\nOn the X axis >>> main\\n\\nOn the Y axis >>> generated\\n')           \n    plt.scatter(X, Y2, s=3)\n    plt.show()     ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def drawing1(main, support, generated, column_number):\n    \n    X  = main.iloc[:, column_number]\n    Y1 = support.iloc[:, column_number]\n    Y2 = generated.iloc[:, column_number]\n    \n    plt.style.use('seaborn-whitegrid') \n    plt.figure(figsize=(8, 8), facecolor='lightgray')\n    plt.title(f'\\nBlue | X axis >> main | Y axis >> support\\n\\nOrange | X axis >> main | Y axis >> generated\\n') \n    \n    plt.scatter(X, Y1, s=3)    \n    plt.scatter(X, Y2, s=3)\n    \n    plt.show()     ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{},"cell_type":"markdown","source":"# Comparative Method\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"# print(sub968.mean() , sub961.mean())\n\nm1 = sub968.mean() + sub961.mean()\n\nm1mean = m1.mean()\n\nm2 = m1 / m1mean\n\nm2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"m3 = m2.copy()\nfor k in range(24):\n    m3[k] = 1.00\n\n# m3    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"m4 = m3.copy()\n\nm4[0]   = 0.00\n\nm4[1]   = 0.00\n\nm4[4]   = 0.00\n\nm4[8]   = 0.00\n\nm4[11]  = 0.00\n\nm4[14]  = 0.00\n\nm4[18]  = 0.00\n\nm4[20]  = 0.00\n\nm4[21]  = 0.00\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Result  \n\n## [(Private Score: 0.97892) , (Public Score: 0.97309)]\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"m4[15]   = 1.30\n\nm4","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Result \n\n## [(Private Score: 0.97915) , (Public Score: 0.97373)]"},{"metadata":{"trusted":true},"cell_type":"code","source":"d = generate1(sub968, sub961, m4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub968.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub961.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 2)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 6)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 7)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 8)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 9)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 11)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 12)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 13)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 14)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 15)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 16)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 17)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 18)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 19)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 20)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 21)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 22)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 23)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"drawing1(sub968, sub961, d, 24)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Result\n\n## [(Score: 0.968) , (Score: 0.961)] >>> d\n\n## d : [(Private Score: 0.97915) , (Public Score: 0.97373)]"},{"metadata":{"trusted":true},"cell_type":"code","source":"d.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{"trusted":true},"cell_type":"code","source":"# print(d.mean() , sub961.mean())\n\nn1 = d.mean() + sub961.mean()\n\nn1mean = n1.mean()\n\nn2 = n1 / n1mean\n\nn2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n3 = n2.copy()\nfor k in range(24):\n    n3[k] = 0.95\n\n# n3   ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Result \n\n## [(Private Score: 0.97975) , (Public Score: 0.97444)]"},{"metadata":{"trusted":true},"cell_type":"code","source":"n4 = n3.copy()\n\nn4[2]   = 0.80\n\nn4[6]   = 0.80\n\nn4[19]  = 0.70\n\nn4[22]  = 0.80\n\nn4[23]  = 0.80\n\nn4","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"e = generate1(d, sub961, n4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Result\n\n## [(Score: 0.973) , (Score: 0.961)] >>> e\n\n## e : [(Private Score: 0.98022) , (Public Score: 0.97490)]"},{"metadata":{"trusted":true},"cell_type":"code","source":"e.describe()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"},{"metadata":{},"cell_type":"markdown","source":"# Submission\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = e\nsub.to_csv(\"submission.csv\", index=False)\n\nd.to_csv(\"submission1.csv\", index=False)\ne.to_csv(\"submission2.csv\", index=False)\n\n!ls","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}