{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Introduction\n\nIn this notebook, one of the last techniques that we applied is shown. A simple post-processing technique, as described in the associated [post](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160986). The score gain is relatively low compared to other techniques we applied. However, it gave a steady increase (~0.0001) for each of languages es/tr/fr/ru both in public LB as private LB. This also secured our first place.\n\n\nHere, I present an example of how to use our earlier Russian subs to achieve the gain in score: going from public LB 9549 to 9550, and private LB 9532 to 9533. This be done in similar fashion with the other languages.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Imports","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# General imports\nimport numpy as np\nimport pandas as pd\nimport os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Specify lang\nLANG = \"ru\"\nDIR = f\"../input/{LANG}-changed-subs/\"\nWEIGHT = 1 # we kept WEIGHT between 1-2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission = pd.read_csv(\"../input/jigsaw-multilingual-toxic-comment-classification/sample_submission.csv\")\ntest = pd.read_csv(\"../input/jigsaw-multilingual-toxic-comment-classification/test.csv\")\nsub_best = pd.read_csv(os.path.join(DIR, \"sub-LB-9549.csv\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"files_sub = os.listdir(DIR)\nfiles_sub = sorted(files_sub)\nprint(len(files_sub))\nfiles_sub","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for file in files_sub:\n    test[file.replace(\".csv\", \"\")] = pd.read_csv(os.path.join(DIR, file))[\"toxic\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test = test.loc[test[\"lang\"]==LANG].reset_index(drop=True)\ntest.head(1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Derive the given sub increases or decreases in score\ntest[\"diff_good1\"] = test[f\"{LANG}-9397\"] - test[f\"{LANG}-9373\"]\ntest[\"diff_good2\"] = test[f\"{LANG}-9476\"] - test[f\"{LANG}-9475\"]\ntest[\"diff_good3\"] = test[f\"{LANG}-9529\"] - test[f\"{LANG}-9510\"]\ntest[\"diff_good4\"] = test[f\"{LANG}-9544\"] - test[f\"{LANG}-9543\"]\n\ntest[\"diff_bad1\"] = test[f\"{LANG}-9545\"] - test[f\"{LANG}-9543-from-9545\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test[\"sub_best\"] = test[\"sub-LB-9549\"]\ncol_comment = [\"id\", \"content\", \"sub_best\"]\ncol_diff = [column for column in test.columns if \"diff\" in column]\ntest_diff = test[col_comment + col_diff].reset_index(drop=True)\n\ntest_diff[\"diff_avg\"] = test_diff[col_diff].mean(axis=1) # the mean trend","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Apply the post-processing technique in one line (as explained in the pseudo-code of my post.\ntest_diff[\"sub_new\"] = test_diff.apply(lambda x: (1+WEIGHT*x[\"diff_avg\"])*x[\"sub_best\"] if x[\"diff_avg\"]<0 else (1-WEIGHT*x[\"diff_avg\"])*x[\"sub_best\"] + WEIGHT*x[\"diff_avg\"] , axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission[\"toxic\"] = sub_best[\"toxic\"]\nsubmission.loc[test[\"id\"], \"toxic\"] = test_diff[\"sub_new\"].values\nsubmission.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}