{"cells":[{"metadata":{},"cell_type":"markdown","source":"# About this notebook\n\nThis notebook ensembles result of 5 starter notebook. Its blending based on public score is a really simple, while its result can match the performance of other complex blending attempts with all availbale public kernels before competition end.\n\n<br/><br/>\n**The five starter kernels are:**\n1. Using English translation and Bert-base: https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-with-huggingface-and-keras (.9158)\n2. XLM-R with augmentation: https://www.kaggle.com/yeayates21/xlm-roberta-augmentation-ssl-0-9417-pub-lb (.9399)\n3. Finetune XLM-R by layer: https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large (.9422)\n4. Pytorch XLM-R with augmentation: https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta (.9416)\n5. validset adjusted plus simple blending: https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta (.9459)\n\n(The fifth blends the fourth with a early version of hamditarek's ensemble about .9430)\n\n\n<br/><br/>\n**Example for complex blending before competition end:**\n1. https://www.kaggle.com/hamditarek/ensemble?scriptVersionId=37216149 (version-108 .9471)\n2. https://www.kaggle.com/jazivxt/howling-with-wolf-on-l-genpresse (.9472)\n\n\n*In contrast, these simple blending solution gives .9473 before and .9482 after adjusting mean of each language follwing advice of [4th solution](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/160980).*\n\n\n# Main idea\n\n* This notebook shows the baseline score can be achieved by emsemble public kernels before competition end.\n* Because of the huge similarity in public kernels avaible before competition end, starter kernels plus simple blending can achieve decent result among all pure blending solutions. \n* The result of this notebook is used to blended with [my monolingual models kernel](https://www.kaggle.com/mint101/jmtc-20-lb-9508-mono-lingual-models) to achieve lb.9514","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nfrom scipy.special import softmax\n\nimport os","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"in_path = '/kaggle/input/'\npath = '/kaggle/input/jigsaw-multilingual-toxic-comment-classification/'\nfiles = ['tpu-inference-super-fast-xlmroberta/submission.csv',\n         'tpu-training-super-fast-xlmroberta/submission.csv',\n         'train-from-mlm-finetuned-xlm-roberta-large/submission.csv',\n         'xlm-roberta-augmentation-ssl-0-9417-pub-lb/submission.csv',\n         'jigsaw-tpu-bert-with-huggingface-and-keras/submission.csv']   \nsumbs = np.array([pd.read_csv(in_path+f).toxic for f in files])\n\npb = np.array([0.9459,0.9416,0.9422,0.9399,0.9158])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Simple blending schedule","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"\"\"\"Weighted by sofmax of 1 over 1 minus pb score.\"\"\"\nweight = lambda x: softmax(1/(1-x))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Blending and Adjusting","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"w = weight(pb)\nprint(w)\nsub = pd.read_csv(path+'sample_submission.csv')\nsub['toxic'] = sumbs.T@w\nsub.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**(The reault here obtains lb.9473)**","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test = pd.read_csv(path+'test.csv')\ndic_ids = {k:v.id for k,v in test.groupby([\"lang\"])}\nadj = {\n    \"fr\":1.04,\n    \"es\":1.06,\n    \"pt\":.96,\n    \"it\":.97,\n    \"tr\":.98,\n}\n\nfor l,v in adj.items():\n    ids = dic_ids[l]\n    sub.loc[ids,\"toxic\"] *= v\n\nsub.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Submit","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}