{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook is a quick demonstration, who to use the Fastai v2 library for a Kaggle tabular competition. Fastai v2 is based on pytorch and allows you, to build a decent machine learning application. For more information please visit the Fastai documentation: https://docs.fast.ai/. I will link to \"Chapter 9, Tabular Modelling Deep Dive\" and the notebook \"09_tabular.ipynb\".\n\nThis monthly competition is a regression problem: predict the probability of each team scoring within the next 10 second. The big challenge is the large amount of training data. There are so many records that using them fully will exceed available memory.\n\nIn this notebook i implement a iterattive training loop.  For each training data set i create a new learner instance and i will load the stored model from the previous run. After that, i will use the last created learner instance to predict the target values. This idea some published in this forum also and my was inspired by that fact, that Fastai V2 offers a lot of pre-trained model image processing. \n\nThis notebook is a copy of my initial notebook, where i use only a prat of the offered traing data sets: https://www.kaggle.com/code/casati8/kaggle-tps-2022-oct-fastai\n\nThis notebook is also inspired by Patricks notebook, which uses an older version of Fastai: https://www.kaggle.com/code/paddykb/tps-2022-10-fastai\n\nLet's start and import the needed stuff ..","metadata":{}},{"cell_type":"code","source":"from fastai.tabular.all import * \nfrom fastai.test_utils import show_install\nimport gc\n\nfrom IPython.display import display, clear_output","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-28T10:09:55.062675Z","iopub.execute_input":"2022-10-28T10:09:55.063397Z","iopub.status.idle":"2022-10-28T10:09:56.886138Z","shell.execute_reply.started":"2022-10-28T10:09:55.063304Z","shell.execute_reply":"2022-10-28T10:09:56.885003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nshow_install()","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:09:56.891573Z","iopub.execute_input":"2022-10-28T10:09:56.894336Z","iopub.status.idle":"2022-10-28T10:09:57.064759Z","shell.execute_reply.started":"2022-10-28T10:09:56.894289Z","shell.execute_reply":"2022-10-28T10:09:57.063451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\ndevice","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:09:57.069886Z","iopub.execute_input":"2022-10-28T10:09:57.072428Z","iopub.status.idle":"2022-10-28T10:09:57.089176Z","shell.execute_reply.started":"2022-10-28T10:09:57.072380Z","shell.execute_reply":"2022-10-28T10:09:57.088197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed_value(seed=718):\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n\nset_seed_value()","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:09:57.095307Z","iopub.execute_input":"2022-10-28T10:09:57.097674Z","iopub.status.idle":"2022-10-28T10:09:57.106047Z","shell.execute_reply.started":"2022-10-28T10:09:57.097634Z","shell.execute_reply":"2022-10-28T10:09:57.104893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = Path('../input/tabular-playground-series-oct-2022/')\n# feather_path = Path('../input/fast-loading-high-compression-with-feather/feather_data')\nPath.BASE_PATH = path\npath.ls()","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:09:57.111346Z","iopub.execute_input":"2022-10-28T10:09:57.113964Z","iopub.status.idle":"2022-10-28T10:09:57.126510Z","shell.execute_reply.started":"2022-10-28T10:09:57.113922Z","shell.execute_reply":"2022-10-28T10:09:57.125398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dtypes_df = pd.read_csv(os.path.join(path, 'train_dtypes.csv'))\ntrain_types = {k: v for (k, v) in zip(dtypes_df.column, dtypes_df.dtype)}\n\ntrain_df = pd.DataFrame(columns=train_types.keys())\n\nsample_submission = pd.read_csv(os.path.join(path, 'sample_submission.csv')).set_index('id')\n\ntypes_df = pd.read_csv(os.path.join(path, 'test_dtypes.csv'))\ndtypes = {k: v for (k, v) in zip(dtypes_df.column, dtypes_df.dtype)}\ntest_df= pd.read_csv(os.path.join(path, 'test.csv'))","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:09:57.131822Z","iopub.execute_input":"2022-10-28T10:09:57.134258Z","iopub.status.idle":"2022-10-28T10:10:08.554550Z","shell.execute_reply.started":"2022-10-28T10:09:57.134217Z","shell.execute_reply":"2022-10-28T10:10:08.553394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's define some of needed prepocessing steps and the depended or target values.","metadata":{}},{"cell_type":"code","source":"dep_var_A = 'team_A_scoring_within_10sec'\ndep_var_B = 'team_B_scoring_within_10sec'\ndep_vars = [dep_var_A, dep_var_B]","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.559571Z","iopub.execute_input":"2022-10-28T10:10:08.562001Z","iopub.status.idle":"2022-10-28T10:10:08.568786Z","shell.execute_reply.started":"2022-10-28T10:10:08.561959Z","shell.execute_reply":"2022-10-28T10:10:08.567747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_with_missing_values =  train_df.columns[train_df.isnull().any()]\n\nfor col in cols_with_missing_values:\n    print(col, train_df[col].isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.570550Z","iopub.execute_input":"2022-10-28T10:10:08.571706Z","iopub.status.idle":"2022-10-28T10:10:08.581840Z","shell.execute_reply.started":"2022-10-28T10:10:08.571668Z","shell.execute_reply":"2022-10-28T10:10:08.580652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I need a list of the used or interesting columns. The game number and the event datas don't belong to this list.","metadata":{}},{"cell_type":"code","source":"used_cols = list()\nfor i in range(0,6):\n    used_cols +=  [col for col in train_types.keys() if col.startswith(f'p{i}') ]\n    \nused_cols += [col for col in train_types.keys() if col.startswith(f'ball') or col.startswith(f'boost')]\nprint(used_cols)","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.583619Z","iopub.execute_input":"2022-10-28T10:10:08.584294Z","iopub.status.idle":"2022-10-28T10:10:08.594314Z","shell.execute_reply.started":"2022-10-28T10:10:08.584256Z","shell.execute_reply":"2022-10-28T10:10:08.593153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I know, that the offered traing data sets contian some NaN values. Therefore we will impute them and we will add the value 0 for them. In a later version of this notebook we can implenent a more enhanced verion of this function.","metadata":{}},{"cell_type":"code","source":"def do_impute(df, used_cols:dict):\n    print(f\"Impute missing values with fillna\")\n    for col in used_cols:\n        df[col] = df[col].fillna(0).astype('float16')\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.596049Z","iopub.execute_input":"2022-10-28T10:10:08.596935Z","iopub.status.idle":"2022-10-28T10:10:08.603455Z","shell.execute_reply.started":"2022-10-28T10:10:08.596897Z","shell.execute_reply":"2022-10-28T10:10:08.602384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def do_scale_position(df):\n    print(f\"Scale the position values\")\n    \n    for col in df.columns:\n        if col.endswith(f'pos_x'):\n            df[col] = df[col]/120.0\n\n        if col.endswith(f'pos_y'):\n            df[col] = df[col]/82.0\n        \n        if col.endswith(f'pos_z'):\n            df[col] = df[col]/40.0\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.605278Z","iopub.execute_input":"2022-10-28T10:10:08.606142Z","iopub.status.idle":"2022-10-28T10:10:08.614333Z","shell.execute_reply.started":"2022-10-28T10:10:08.606104Z","shell.execute_reply":"2022-10-28T10:10:08.612614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the iteration loop we need some definitions like number of used epochs and the number of used trainig data sets.","metadata":{}},{"cell_type":"code","source":"num_of_epochs = 25\nnum_traing_data_sets = 10\nlearning_rate = 5e-3","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.616921Z","iopub.execute_input":"2022-10-28T10:10:08.618702Z","iopub.status.idle":"2022-10-28T10:10:08.627617Z","shell.execute_reply.started":"2022-10-28T10:10:08.618664Z","shell.execute_reply":"2022-10-28T10:10:08.626595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some more definitions for Fastai and the usage in the iteration loop","metadata":{}},{"cell_type":"code","source":"results = pd.DataFrame(columns=['name',  'min. valid loss', 'min. mae', 'time (sec)'])\nmy_config = tabular_config(ps=0.1, bn_cont=True, use_bn=True, y_range=(0,1),  act_cls = Mish(),)\n\ntrain_df = None\nlearn = None","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.640739Z","iopub.execute_input":"2022-10-28T10:10:08.641481Z","iopub.status.idle":"2022-10-28T10:10:08.653083Z","shell.execute_reply.started":"2022-10-28T10:10:08.641442Z","shell.execute_reply":"2022-10-28T10:10:08.651677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I need a list of the column names, which are candidates for category variables and which are no candidates, also called continous variables. The Fastai library offers the function 'cont_cat_split' to do this for us. Our training data set contains only floating values for the independed variables, therefore we expect that no category variables are available.","metadata":{}},{"cell_type":"markdown","source":"The next step is to create a data loader. The Fastai library offers a powerful helper called 'TabularPandas'. It needs the data frame, list of the category and continous variables, the depened variable and a splitter. The splitter divides the data set into two parts: one for the training and one for the validation and for internal optimization step in each epoch. The batch size is set to 8196, because we have a large data set. We can use a random split because the rows in the data set are independed.","metadata":{}},{"cell_type":"code","source":"def getData(df, batchSize=8196):\n    \n    cont_vars, cat_vars = cont_cat_split(df, dep_var=dep_vars)\n    to_train = TabularPandas(df, \n                           [Normalize],\n                           cat_names=cat_vars,\n                           cont_names=cont_vars, \n                           splits=RandomSplitter(valid_pct=0.20)(df),  \n                           device = device,\n                           y_block=RegressionBlock(),\n                           y_names=dep_vars) \n\n    return to_train.dataloaders(bs=batchSize)","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.655948Z","iopub.execute_input":"2022-10-28T10:10:08.657374Z","iopub.status.idle":"2022-10-28T10:10:08.670416Z","shell.execute_reply.started":"2022-10-28T10:10:08.657336Z","shell.execute_reply":"2022-10-28T10:10:08.669175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The iteration loop works as follows:\nFor each traing data set train_i do:\n1. Load the data\n2. Impute the missing values\n3. Create the dataloader with the training and validation batches\n4. Create a new tabular learner and the model. The default settings are two hidden layers with 200 and 100 elements. I will use this structure for the baseline. Increasing the number of parameters in the neural network will improve the accuarcy and score, hopefully.\n5. Load the stored model parameters if they are available. This is the case of all data sets, but not for the first one.\n6. Train the new created model for the specified number of epochs and the learning rate. The Fastai library offers the SaveModelCallback callback to store the 'best' model. I use this option too transfer to used model between the iteration steps.\n7. Store the minimum validation loss and some other metrics in a global dataframe, to keep track with the learning success for each iteration step.","metadata":{}},{"cell_type":"code","source":"for i in range(num_traing_data_sets):\n    \n    name = f\"train_{i}\"\n    print(f\"Load file {name}\")\n    \n    train_df = pd.read_csv(os.path.join(path, f\"train_{i}.csv\"), dtype=train_types)\n    \n    train_df = train_df[used_cols + dep_vars].copy()\n    train_df[dep_vars] = train_df[dep_vars].astype('int8')\n    \n    gc.collect()\n    \n    train_df = do_impute(train_df, used_cols)\n    train_df = do_scale_position(train_df)\n    \n    dls = getData(train_df)\n    print(f\"Number of training batches {len(dls.train)} and validation batchs {len(dls.valid)}\")\n    \n    if i == 0:\n        dls.show_batch()\n        \n    print(f\"Create model file {name}\")\n    learn = tabular_learner(dls,\n                            config = my_config,\n                          #  loss_func=BCELossFlat(),\n                            metrics=[mae, rmse],\n                           # layers=[512, 256, 128,64],\n                            # layers = [2048, 2048, 1024, 1024]\n                           )\n    \n    if i == 0:\n        learn.model\n        \n    if i > 0 and os.path.exists(\"models/kaggle_2022_oct_iter.pth\"):\n        learn.load('kaggle_2022_oct_iter')\n        learn.unfreeze()\n        \n    print(f\"Fit model  file {name} for {num_of_epochs} epochs \")\n    start = time.time()\n    learn.fit_one_cycle(num_of_epochs, learning_rate, wd=0.1, cbs=SaveModelCallback(fname='kaggle_2022_oct_iter', with_opt=True))\n    elapsed = time.time() - start\n    \n    vals_df = pd.DataFrame(learn.recorder.values)\n    \n    results.loc[i] = [name, vals_df.iloc[:,1].min(), vals_df.iloc[:,2].min(), int(elapsed)]\n    \n    clear_output()\n    display(results)\n    del train_df","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:10:08.672439Z","iopub.execute_input":"2022-10-28T10:10:08.673474Z","iopub.status.idle":"2022-10-28T10:52:25.678091Z","shell.execute_reply.started":"2022-10-28T10:10:08.673437Z","shell.execute_reply":"2022-10-28T10:52:25.676811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To calculate the predictions for this competition, i will load the best model from the training process. Best model means the model with the lowest validation loss value.","metadata":{}},{"cell_type":"code","source":"learn.load('kaggle_2022_oct_iter')","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:25.679784Z","iopub.execute_input":"2022-10-28T10:52:25.681481Z","iopub.status.idle":"2022-10-28T10:52:25.701516Z","shell.execute_reply.started":"2022-10-28T10:52:25.681437Z","shell.execute_reply":"2022-10-28T10:52:25.700505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now it's time to calculate the predictions for the test data set.","metadata":{}},{"cell_type":"code","source":"test_df = do_impute(test_df, used_cols)\n\ndlt = learn.dls.test_dl(test_df) \npreds = learn.get_preds(dl=dlt ) \n\npreds[0]","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:25.702779Z","iopub.execute_input":"2022-10-28T10:52:25.703083Z","iopub.status.idle":"2022-10-28T10:52:30.647129Z","shell.execute_reply.started":"2022-10-28T10:52:25.703055Z","shell.execute_reply":"2022-10-28T10:52:30.645993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission = pd.read_csv(os.path.join(path, 'sample_submission.csv')).set_index('id')\npx = pd.DataFrame(preds[0])\nsample_submission[dep_var_A] = px[0]\nsample_submission[dep_var_B] = px[1]\n\nsample_submission.to_csv(\"submission.csv\")\nsample_submission.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:30.648788Z","iopub.execute_input":"2022-10-28T10:52:30.649794Z","iopub.status.idle":"2022-10-28T10:52:32.616566Z","shell.execute_reply.started":"2022-10-28T10:52:30.649750Z","shell.execute_reply":"2022-10-28T10:52:32.615317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"prediction {dep_var_A} mean: {sample_submission[dep_var_A].mean():.6}\")\nprint(f\"prediction {dep_var_B} mean: {sample_submission[dep_var_B].mean():.6}\")","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:32.618322Z","iopub.execute_input":"2022-10-28T10:52:32.619015Z","iopub.status.idle":"2022-10-28T10:52:32.631155Z","shell.execute_reply.started":"2022-10-28T10:52:32.618973Z","shell.execute_reply":"2022-10-28T10:52:32.629825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -la ","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:32.632591Z","iopub.execute_input":"2022-10-28T10:52:32.633195Z","iopub.status.idle":"2022-10-28T10:52:33.845202Z","shell.execute_reply.started":"2022-10-28T10:52:32.633133Z","shell.execute_reply":"2022-10-28T10:52:33.843900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next class is a direct copy from Zachary website: https://walkwithfastai.com/Regression_and_Permutation_Importance#Permutation-Importance\n\nThe implementation shuffles each column/feature in my training data and calculates how the prediction is changed. The more the prediction is changed, the more important a column/feature is. ","metadata":{}},{"cell_type":"code","source":"class PermutationImportance():\n      \"Calculate and plot the permutation importance\"\n      def __init__(self, learn:Learner, df=None, bs=None):\n        \"Initialize with a test dataframe, a learner, and a metric\"\n        self.learn = learn\n        self.df = df\n        bs = bs if bs is not None else learn.dls.bs\n        if self.df is not None:\n          self.dl = learn.dls.test_dl(self.df, bs=bs)\n        else:\n          self.dl = learn.dls[1]\n        self.x_names = learn.dls.x_names.filter(lambda x: '_na' not in x)\n        self.na = learn.dls.x_names.filter(lambda x: '_na' in x)\n        self.y = dls.y_names\n        self.results = self.calc_feat_importance()\n        self.plot_importance(self.ord_dic_to_df(self.results))\n    \n        \n      def get_results(self): return self.results\n        \n      def measure_col(self, name:str):\n          \"Measures change after column shuffle\"\n          col = [name]\n          if f'{name}_na' in self.na: col.append(name)\n          orig = self.dl.items[col].values\n          perm = np.random.permutation(len(orig))\n          self.dl.items[col] = self.dl.items[col].values[perm]\n          metric = learn.validate(dl=self.dl)[1]\n          self.dl.items[col] = orig\n          clear_output()\n          return metric\n\n      def calc_feat_importance(self):\n          \"Calculates permutation importance by shuffling a column on a percentage scale\"\n          print('Getting base error')\n          base_error = self.learn.validate(dl=self.dl)[1]\n          self.importance = {}\n          pbar = progress_bar(self.x_names)\n          print('Calculating Permutation Importance')\n          for col in pbar:\n            self.importance[col] = self.measure_col(col)\n          for key, value in self.importance.items():\n            self.importance[key] = np.abs(base_error-value)/base_error #this can be adjusted\n          return OrderedDict(sorted(self.importance.items(), key=lambda kv: kv[1], reverse=True))\n\n      def ord_dic_to_df(self, dict:OrderedDict):\n          return pd.DataFrame([[k, v] for k, v in dict.items()], columns=['feature', 'importance'])\n\n      def plot_importance(self, df:pd.DataFrame, limit=20, asc=False, **kwargs):\n          \"Plot importance with an optional limit to how many variables shown\"\n          df_copy = df.copy()\n          df_copy['feature'] = df_copy['feature'].str.slice(0,40)\n          df_copy = df_copy.sort_values(by='importance', ascending=asc)[:limit].sort_values(by='importance', ascending=not(asc))\n          ax = df_copy.plot.barh(x='feature', y='importance', sort_columns=True, **kwargs)\n          for p in ax.patches:\n            ax.annotate(f'{p.get_width():.4f}', ((p.get_width() * 1.005), p.get_y()  * 1.005))","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:33.847471Z","iopub.execute_input":"2022-10-28T10:52:33.847789Z","iopub.status.idle":"2022-10-28T10:52:33.867516Z","shell.execute_reply.started":"2022-10-28T10:52:33.847759Z","shell.execute_reply":"2022-10-28T10:52:33.866452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I need to reload one of the training data sets and prepare the data to calculate the permutation importance","metadata":{}},{"cell_type":"code","source":"importance_df = pd.read_csv(os.path.join(path, f\"train_1.csv\"), dtype=train_types)\n    \nimportance_df = importance_df[used_cols + dep_vars].copy()\nimportance_df[dep_vars] = importance_df[dep_vars].astype('int8')\n       \nimportance_df = do_impute(importance_df, used_cols)","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:52:33.869243Z","iopub.execute_input":"2022-10-28T10:52:33.869644Z","iopub.status.idle":"2022-10-28T10:53:00.720550Z","shell.execute_reply.started":"2022-10-28T10:52:33.869606Z","shell.execute_reply":"2022-10-28T10:53:00.719525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res = PermutationImportance(learn, importance_df.iloc[:250000], bs=8196)\nres.get_results()","metadata":{"execution":{"iopub.status.busy":"2022-10-28T10:53:00.722227Z","iopub.execute_input":"2022-10-28T10:53:00.722606Z","iopub.status.idle":"2022-10-28T10:53:30.232208Z","shell.execute_reply.started":"2022-10-28T10:53:00.722569Z","shell.execute_reply":"2022-10-28T10:53:30.231101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The happy end!","metadata":{}}]}