{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on December 06, 2022 (17:52 GMT) by Marília Prata, mpwolke","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=Warning)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-06T18:44:10.569216Z","iopub.execute_input":"2022-12-06T18:44:10.569770Z","iopub.status.idle":"2022-12-06T18:44:10.592696Z","shell.execute_reply.started":"2022-12-06T18:44:10.569720Z","shell.execute_reply":"2022-12-06T18:44:10.591400Z"},"_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Optimizing probabilities for best MCC  By CPMP (6 years ago)\n\nThis precious script was written by CPMP. Vote his work, please.\n\nhttps://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\n\nfrom sklearn.metrics import matthews_corrcoef\n\nimport matplotlib.pyplot as plt\nimport numpy as np","metadata":{"execution":{"iopub.status.busy":"2022-12-06T18:40:19.780368Z","iopub.execute_input":"2022-12-06T18:40:19.780820Z","iopub.status.idle":"2022-12-06T18:40:20.370953Z","shell.execute_reply.started":"2022-12-06T18:40:19.780782Z","shell.execute_reply":"2022-12-06T18:40:20.369837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from PIL import Image, ImageDraw\nfrom pathlib import Path\nfrom tqdm.notebook import tqdm","metadata":{"execution":{"iopub.status.busy":"2022-12-06T14:46:30.475358Z","iopub.execute_input":"2022-12-06T14:46:30.476628Z","iopub.status.idle":"2022-12-06T14:46:30.621382Z","shell.execute_reply.started":"2022-12-06T14:46:30.476525Z","shell.execute_reply":"2022-12-06T14:46:30.620250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Numba a just-in-time compiler for Python\n\n\"Numba is a just-in-time compiler for Python that works best on code that uses NumPy arrays and functions, and loops. The most common way to use Numba is through its collection of decorators that can be applied to your functions to instruct Numba to compile them. When a call is made to a Numba-decorated function it is compiled to machine code “just-in-time” for execution and all or part of your code can subsequently run at native machine code speed!\"\n\nhttps://numba.readthedocs.io/en/stable/user/5minguide.html","metadata":{}},{"cell_type":"markdown","source":"![](https://www.askpython.com/wp-content/uploads/2021/03/100x-Faster-Python-using-Numba.png)","metadata":{}},{"cell_type":"markdown","source":"#Compile the code with Numba","metadata":{}},{"cell_type":"code","source":"#Snippet by CPMP  https://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\nfrom numba import jit\n\n@jit\ndef mcc(tp, tn, fp, fn):\n    sup = tp * tn - fp * fn\n    inf = (tp + fp) * (tp + fn) * (tn + fp) * (tn + fn)\n    if inf==0:\n        return 0\n    else:\n        return sup / np.sqrt(inf)\n\n@jit\ndef eval_mcc(y_true, y_prob, show=False):\n    idx = np.argsort(y_prob)\n    y_true_sort = y_true[idx]\n    n = y_true.shape[0]\n    nump = 1.0 * np.sum(y_true) # number of positive\n    numn = n - nump # number of negative\n    tp = nump\n    tn = 0.0\n    fp = numn\n    fn = 0.0\n    best_mcc = 0.0\n    best_id = -1\n    prev_proba = -1\n    best_proba = -1\n    mccs = np.zeros(n)\n    for i in range(n):\n        # all items with idx < i are predicted negative while others are predicted positive\n        # only evaluate mcc when probability changes\n        proba = y_prob[idx[i]]\n        if proba != prev_proba:\n            prev_proba = proba\n            new_mcc = mcc(tp, tn, fp, fn)\n            if new_mcc >= best_mcc:\n                best_mcc = new_mcc\n                best_id = i\n                best_proba = proba\n        mccs[i] = new_mcc\n        if y_true_sort[i] == 1:\n            tp -= 1.0\n            fn += 1.0\n        else:\n            fp -= 1.0\n            tn += 1.0\n    if show:\n        y_pred = (y_prob >= best_proba).astype(int)\n        score = matthews_corrcoef(y_true, y_pred)\n        print(score, best_mcc)\n        plt.plot(mccs)\n        return best_proba, best_mcc, y_pred\n    else:\n        return best_mcc","metadata":{"execution":{"iopub.status.busy":"2022-12-06T18:41:33.813412Z","iopub.execute_input":"2022-12-06T18:41:33.813846Z","iopub.status.idle":"2022-12-06T18:41:34.864528Z","shell.execute_reply.started":"2022-12-06T18:41:33.813811Z","shell.execute_reply":"2022-12-06T18:41:34.863144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Snippet by CPMP  https://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\ny_prob0 = np.random.rand(1000000)\ny_prob = y_prob0 + 0.4 * np.random.rand(1000000) - 0.02\n\n\ny_true = (y_prob0 > 0.6).astype(int)\nbest_proba, best_mcc, y_pred = eval_mcc(y_true, y_prob, True)","metadata":{"execution":{"iopub.status.busy":"2022-12-06T18:44:42.022712Z","iopub.execute_input":"2022-12-06T18:44:42.023127Z","iopub.status.idle":"2022-12-06T18:44:43.164726Z","shell.execute_reply.started":"2022-12-06T18:44:42.023093Z","shell.execute_reply":"2022-12-06T18:44:43.163462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%timeit eval_mcc(y_true, y_prob)","metadata":{"execution":{"iopub.status.busy":"2022-12-06T18:45:03.235141Z","iopub.execute_input":"2022-12-06T18:45:03.235624Z","iopub.status.idle":"2022-12-06T18:45:05.507281Z","shell.execute_reply.started":"2022-12-06T18:45:03.235583Z","shell.execute_reply":"2022-12-06T18:45:05.505939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Snippet by CPMP  https://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\ndef roundn(yprob, scale):\n    return np.around(y_prob * scale) / scale\n\nbest_proba, best_mcc, y_pred = eval_mcc(y_true, roundn(y_prob, 100), True)","metadata":{"execution":{"iopub.status.busy":"2022-12-06T18:46:55.375860Z","iopub.execute_input":"2022-12-06T18:46:55.376311Z","iopub.status.idle":"2022-12-06T18:46:56.409174Z","shell.execute_reply.started":"2022-12-06T18:46:55.376274Z","shell.execute_reply":"2022-12-06T18:46:56.407829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Above, \"the best mcc is lower than when probabilities were all different.\"","metadata":{}},{"cell_type":"markdown","source":"#To use with XGBOOST, wrap it to get the right input and output.","metadata":{}},{"cell_type":"code","source":"#Snippet by CPMP  https://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\ndef mcc_eval(y_prob, dtrain):\n    y_true = dtrain.get_label()\n    best_mcc = eval_mcc(y_true, y_prob)\n    return 'MCC', best_mcc","metadata":{"execution":{"iopub.status.busy":"2022-12-06T18:48:02.128685Z","iopub.execute_input":"2022-12-06T18:48:02.129793Z","iopub.status.idle":"2022-12-06T18:48:02.136113Z","shell.execute_reply.started":"2022-12-06T18:48:02.129744Z","shell.execute_reply":"2022-12-06T18:48:02.134850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Then, use it with xgboost, for instance passing it as feval parameter as follow:","metadata":{}},{"cell_type":"code","source":"#Snippet by CPMP  https://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\nimport xgboost as xgb\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2022-12-06T19:13:58.168876Z","iopub.execute_input":"2022-12-06T19:13:58.169309Z","iopub.status.idle":"2022-12-06T19:13:58.175026Z","shell.execute_reply.started":"2022-12-06T19:13:58.169276Z","shell.execute_reply":"2022-12-06T19:13:58.174031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Snippet by Chris D. https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n\n# TRAIN RANDOM SEED\nSEED = 42\n\n# XGB MODEL PARAMETERS\nxgb_parms = { \n    'max_depth':4, \n    'learning_rate':0.05, \n    'subsample':0.8,\n    'colsample_bytree':0.6, \n    'eval_metric':'logloss',\n    'objective':'binary:logistic',\n    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n    'random_state':SEED\n}","metadata":{"execution":{"iopub.status.busy":"2022-12-06T19:05:17.999668Z","iopub.execute_input":"2022-12-06T19:05:18.000181Z","iopub.status.idle":"2022-12-06T19:05:18.006852Z","shell.execute_reply.started":"2022-12-06T19:05:18.000145Z","shell.execute_reply":"2022-12-06T19:05:18.005539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Snippet by Alexis Cook https://www.kaggle.com/code/mpwolke/exercise-xgboost\n\nX = pd.read_csv('../input/nfl-player-contact-detection/train_baseline_helmets.csv')\ny = X.game_play \n\n# Break off validation set from training data\nX_train_full, X_valid_full, y_train, y_valid = train_test_split(X, y, train_size=0.8, test_size=0.2,\n\nrandom_state=0)                                                                ","metadata":{"execution":{"iopub.status.busy":"2022-12-06T19:28:15.734520Z","iopub.execute_input":"2022-12-06T19:28:15.734987Z","iopub.status.idle":"2022-12-06T19:28:23.447970Z","shell.execute_reply.started":"2022-12-06T19:28:15.734941Z","shell.execute_reply":"2022-12-06T19:28:23.446958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Snippet by Chris D. https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n\n# NEEDED WITH DeviceQuantileDMatrix BELOW\nclass IterLoadForDMatrix(xgb.core.DataIter):\n    def __init__(self, df=None, features=None, target=None, batch_size=256*1024):\n        self.features = features\n        self.target = target\n        self.df = df\n        self.it = 0 # set iterator to 0\n        self.batch_size = batch_size\n        self.batches = int( np.ceil( len(df) / self.batch_size ) )\n        super().__init__()","metadata":{"execution":{"iopub.status.busy":"2022-12-06T19:36:36.295081Z","iopub.execute_input":"2022-12-06T19:36:36.295565Z","iopub.status.idle":"2022-12-06T19:36:36.303991Z","shell.execute_reply.started":"2022-12-06T19:36:36.295521Z","shell.execute_reply":"2022-12-06T19:36:36.302573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Feval\n\nWith XGBoost, pass it as feval parameter as below:\n\nWrite your own Fevals, since I've No clue what to write in this NFL (1st and Future - Player Contact Detection) Competition","metadata":{}},{"cell_type":"markdown","source":"![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcTMoulD7MSVX-BjSDa8anFbtm-sH9IpfJUQLwi7rk9xNAGyIFqZ80kAFMD0bJqu3kYF_tw&usqp=CAU)https://slideplayer.com/slide/4351551/","metadata":{}},{"cell_type":"markdown","source":"#Feval Function:\n\n\"Feval function is used to evaluate the function output of a function by using the arguments passed inside a single parenthesis input. It accepts the function name passed as a “string”, as its first argument. Feval function is used to make the code easy to understand or read, as we use parentheses in it for both function invocation and indexing.\"\n\nhttps://www.educba.com/feval-matlab/\n\nFEVAL Description\n\n\"Multiple evaluation of a function for one or two arguments of vector type :\n\nz=feval(x,f)\nreturns the vector z defined by z(i)=f(x(i))\n\nz=feval(x,y,f)\nreturns the matrix z such as z(i,j)=f(x(i),y(j))\n\n\"f is an external (function or routine) accepting on one or two arguments which are supposed to be real. The result returned by f can be real or complex. In case of a Fortran call, the function f must be defined in the subroutine fevaltable.c (in directory SCI/modules/differential_equations/src/c).\"\n\nhttps://help.scilab.org/doc/6.0.0/en_US/feval.html","metadata":{}},{"cell_type":"code","source":"#Snippet by CPMP  https://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\nbst = xgb.train(xgb_parms, X_train_full, \n                    num_boost_round=9999,\n                    dtrain=dtrain,\n                    dvalid=dvalid, \n                    evals=[WRITE yours HERE example(dtrain,'train'),(dvalid,'valid')],#I got No Clue what eval to write in this comp\n                    early_stopping_rounds=early_stopping_rounds, \n                    evals_result=evals_result, \n                    verbose_eval=verbose_eval,\n                    feval=mcc_eval, \n                    maximize=True,)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Time to finish this Kaggle Notebook\n\nThough I copied the Matthews Correlation Coeficient script, I took 2:09h to finish it since I tried to complete the last snippet, and failed over and over. ","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nCPMP Optimizing probabilities for best MCC\nhttps://www.kaggle.com/code/cpmpml/optimizing-probabilities-for-best-mcc/notebook\n\nChris Deotte https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n\nAlexis Cook https://www.kaggle.com/code/mpwolke/exercise-xgboost","metadata":{}}]}