{"cells":[{"metadata":{},"cell_type":"markdown","source":"# How To Ensemble OOF\nIn this notebook, we learn how to use `forward selection` to ensemble OOF. First build lots of models using the same KFolds (i.e. use same `seed`). Next save all the oof files as `oof_XX.csv` and submission files as `sub_XX.csv` where the oof and submission share the same `XX` number. Then save them in a Kaggle dataset and run the code below.\n\nThe ensemble begins with the model of highest oof AUC. Next each other model is added one by one to see which additional model increases ensemble AUC the most. The best additional model is kept and the process is repeated until the ensemble AUC doesn't increase.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Read OOF Files\nWhen i get more time, I will compete this table to describe all 39 models in this notebook. For now here are the ones that get selected:\n\n| k | CV | LB | read size | crop size | effNet | ext data | upsample | misc | name |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| 1 | 0.910 | 0.950 | 384 | 384 | B6 | 2018 | no |  | oof_100 |\n| 3 | 0.916 | 0.946 | 384 | 384 | B345 | no | no |  | oof_108 |\n| 8 | 0.935 | 0.949 | 768 | 512 | B7 | 2018 | 1,1,1,1 |  | oof_113 |\n| 10 | 0.920 | 0.941 | 512 | 384 | B5 | 2019 2018 | 10,0,0,0 |  | oof_117 |\n| 12 | 0.935 | 0.937 | 768 | 512 | B6 | 2019 2018 | 3,3,0,0 |  | oof_120 |\n| 21 | 0.933 | 0.950 | 1024 | 512 | B6 | 2018 | 2,2,2,2 |  | oof_30 |\n| 26 | 0.927 | 0.942 | 768 | 384 | B4 | 2018 | no |  | oof_385 |\n| 37 | 0.936 | 0.956 | 512 | 384 | B5 | 2018 | 1,1,1,1 |  | oof_67 |\n","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd, numpy as np, os\nfrom sklearn.metrics import roc_auc_score\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"PATH = '../input/melanoma-oof-and-sub/'\nFILES = os.listdir(PATH)\n\nOOF = np.sort( [f for f in FILES if 'oof' in f] )\nOOF_CSV = [pd.read_csv(PATH+k) for k in OOF]\n\nprint('We have %i oof files...'%len(OOF))\nprint(); print(OOF)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"x = np.zeros(( len(OOF_CSV[0]),len(OOF) ))\nfor k in range(len(OOF)):\n    x[:,k] = OOF_CSV[k].pred.values\n    \nTRUE = OOF_CSV[0].target.values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"all = []\nfor k in range(x.shape[1]):\n    auc = roc_auc_score(OOF_CSV[0].target,x[:,k])\n    all.append(auc)\n    print('Model %i has OOF AUC = %.4f'%(k,auc))\n    \nm = [np.argmax(all)]; w = []","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Build OOF Ensemble. Maximize CV Score","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"old = np.max(all); \n\nRES = 200; \nPATIENCE = 10; \nTOL = 0.0003\nDUPLICATES = False\n\nprint('Ensemble AUC = %.4f by beginning with model %i'%(old,m[0]))\nprint()\n\nfor kk in range(len(OOF)):\n    \n    # BUILD CURRENT ENSEMBLE\n    md = x[:,m[0]]\n    for i,k in enumerate(m[1:]):\n        md = w[i]*x[:,k] + (1-w[i])*md\n        \n    # FIND MODEL TO ADD\n    mx = 0; mx_k = 0; mx_w = 0\n    print('Searching for best model to add... ')\n    \n    # TRY ADDING EACH MODEL\n    for k in range(x.shape[1]):\n        print(k,', ',end='')\n        if not DUPLICATES and (k in m): continue\n            \n        # EVALUATE ADDING MODEL K WITH WEIGHTS W\n        bst_j = 0; bst = 0; ct = 0\n        for j in range(RES):\n            tmp = j/RES*x[:,k] + (1-j/RES)*md\n            auc = roc_auc_score(TRUE,tmp)\n            if auc>bst:\n                bst = auc\n                bst_j = j/RES\n            else: ct += 1\n            if ct>PATIENCE: break\n        if bst>mx:\n            mx = bst\n            mx_k = k\n            mx_w = bst_j\n            \n    # STOP IF INCREASE IS LESS THAN TOL\n    inc = mx-old\n    if inc<=TOL: \n        print(); print('No increase. Stopping.')\n        break\n        \n    # DISPLAY RESULTS\n    print(); #print(kk,mx,mx_k,mx_w,'%.5f'%inc)\n    print('Ensemble AUC = %.4f after adding model %i with weight %.3f. Increase of %.4f'%(mx,mx_k,mx_w,inc))\n    print()\n    \n    old = mx; m.append(mx_k); w.append(mx_w)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print('We are using models',m)\nprint('with weights',w)\nprint('and achieve ensemble AUC = %.4f'%old)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"md = x[:,m[0]]\nfor i,k in enumerate(m[1:]):\n    md = w[i]*x[:,k] + (1-w[i])*md\nplt.hist(md,bins=100)\nplt.title('Ensemble OOF predictions')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = OOF_CSV[0].copy()\ndf.pred = md\ndf.to_csv('ensemble_oof.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Load SUB Files","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"SUB = np.sort( [f for f in FILES if 'sub' in f] )\nSUB_CSV = [pd.read_csv(PATH+k) for k in SUB]\n\nprint('We have %i submission files...'%len(SUB))\nprint(); print(SUB)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# VERFIY THAT SUBMISSION FILES MATCH OOF FILES\na = np.array( [ int( x.split('_')[1].split('.')[0]) for x in SUB ] )\nb = np.array( [ int( x.split('_')[1].split('.')[0]) for x in OOF ] )\nif len(a)!=len(b):\n    print('ERROR submission files dont match oof files')\nelse:\n    for k in range(len(a)):\n        if a[k]!=b[k]: print('ERROR submission files dont match oof files')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"y = np.zeros(( len(SUB_CSV[0]),len(SUB) ))\nfor k in range(len(SUB)):\n    y[:,k] = SUB_CSV[k].target.values","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Build SUB Ensemble","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"md2 = y[:,m[0]]\nfor i,k in enumerate(m[1:]):\n    md2 = w[i]*y[:,k] + (1-w[i])*md2\nplt.hist(md2,bins=100)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df = SUB_CSV[0].copy()\ndf.target = md2\ndf.to_csv('ensemble_sub.csv',index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}