{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook is a quick demonstration, who to use the Fastai v2 library for a Kaggle tabular competition. Fastai v2 is based on pytorch and allows you, to build a decent machine learning application. For more information please visit the Fastai documentation: https://docs.fast.ai/. I will link to \"Chapter 9, Tabular Modelling Deep Dive\" and the notebook \"09_tabular.ipynb\".\n\nThis monthly competition is a binary classification problem: find the failure state of measurement values for different products. \nFor eache measurement an tabular entries with 24 different values exists.  In this notebook i will use a neural network approach and i will train this network with the offered traing data set.\n\nLet's start and import the needed stuff ..","metadata":{}},{"cell_type":"code","source":"from fastai.tabular.all import * \nfrom fastai.test_utils import show_install\nfrom IPython.display import display, clear_output\n\nimport xgboost\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import SimpleImputer, IterativeImputer \n\nimport seaborn as sns\n\nshow_install()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-06T14:44:00.955828Z","iopub.execute_input":"2022-08-06T14:44:00.956322Z","iopub.status.idle":"2022-08-06T14:44:02.217885Z","shell.execute_reply.started":"2022-08-06T14:44:00.956224Z","shell.execute_reply":"2022-08-06T14:44:02.216451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\ndevice","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.220455Z","iopub.execute_input":"2022-08-06T14:44:02.221562Z","iopub.status.idle":"2022-08-06T14:44:02.232438Z","shell.execute_reply.started":"2022-08-06T14:44:02.221520Z","shell.execute_reply":"2022-08-06T14:44:02.231495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed_value(seed=718):\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n\nset_seed_value()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.233862Z","iopub.execute_input":"2022-08-06T14:44:02.234442Z","iopub.status.idle":"2022-08-06T14:44:02.253282Z","shell.execute_reply.started":"2022-08-06T14:44:02.234406Z","shell.execute_reply":"2022-08-06T14:44:02.252105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = Path('../input/tabular-playground-series-aug-2022/')\nPath.BASE_PATH = path\npath.ls()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.255728Z","iopub.execute_input":"2022-08-06T14:44:02.256320Z","iopub.status.idle":"2022-08-06T14:44:02.268739Z","shell.execute_reply.started":"2022-08-06T14:44:02.256284Z","shell.execute_reply":"2022-08-06T14:44:02.267631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(os.path.join(path, 'train.csv')).set_index('id')\ntest_df = pd.read_csv(os.path.join(path, 'test.csv')).set_index('id')\nsample_submission = pd.read_csv(os.path.join(path, 'sample_submission.csv'))\n\ndep_var = 'failure'\n\nadd_missing_values = True","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.270279Z","iopub.execute_input":"2022-08-06T14:44:02.270724Z","iopub.status.idle":"2022-08-06T14:44:02.516870Z","shell.execute_reply.started":"2022-08-06T14:44:02.270687Z","shell.execute_reply":"2022-08-06T14:44:02.515651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next step is to look at the info of loaded data. ","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.518680Z","iopub.execute_input":"2022-08-06T14:44:02.519331Z","iopub.status.idle":"2022-08-06T14:44:02.547930Z","shell.execute_reply.started":"2022-08-06T14:44:02.519283Z","shell.execute_reply":"2022-08-06T14:44:02.546546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I use Pandas to import them and to verify, where null values are there or some values are missing. The result shows, that the some fatures have missing values. Therefore i need a strategy to fill these missing values. The Fastai offers feature to fill these values when the dataloader is created. This is the Fastai default behaviour. \n\nI will offer a function to impute the missing values before. Therefore i will use the impution techniques from the Kaggle competition in June 2022.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.549557Z","iopub.execute_input":"2022-08-06T14:44:02.550991Z","iopub.status.idle":"2022-08-06T14:44:02.567458Z","shell.execute_reply.started":"2022-08-06T14:44:02.550937Z","shell.execute_reply":"2022-08-06T14:44:02.566134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The info of the dataframe shows me that some object features exists. These features are candidates for categorized features later on. \nThe fastai library converts the categorized features into embeddings per deafult. Therefore don't implement any special handling for these columns in this baseline notebook. Show me how many unique values these features have:\n\nI sees that a handfull differnt values exists. 'attribute_0' and 'attribute_1'have the same unique values in tge training and in the test dataset. But the product_code values are different: There is no overlap of the products in the training and in the test data sets!","metadata":{}},{"cell_type":"code","source":"print(f\"train_df['attribute_0']  : {train_df['attribute_0'].unique()}\")\nprint(f\"train_df['attribute_1']  : {train_df['attribute_1'].unique()}\")\nprint(f\"train_df['product_code'] : {train_df['product_code'].unique()}\")\n\nprint(f\"train_df['failure'] : {train_df[dep_var].unique()}\")\n\nprint(f\"test_df['attribute_0'] : {test_df['attribute_0'].unique()}\")\nprint(f\"test_df['attribute_1'] : {test_df['attribute_1'].unique()}\")\nprint(f\"test_df['product_code'] : {test_df['product_code'].unique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.569032Z","iopub.execute_input":"2022-08-06T14:44:02.569499Z","iopub.status.idle":"2022-08-06T14:44:02.592529Z","shell.execute_reply.started":"2022-08-06T14:44:02.569457Z","shell.execute_reply":"2022-08-06T14:44:02.591138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In an optimized version of this notebook you can implement different value mapping function or strategies for these featues: \n* use a 'one got encoded' strategy with DataFrame.get_dummies()\n* map the string values 'material_x' into integer values \n* implement some of the ideas which are discussed here in the forum","metadata":{}},{"cell_type":"markdown","source":"Let's see how the values for the depended variable, the taget, are distributed: There are to unique values and these quite balanced, there much more entries without a failure ('failure' == 0) than with failures. This can be a challenge for the neural network later on.","metadata":{}},{"cell_type":"code","source":"train_df.hist(column=dep_var)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:02.594597Z","iopub.execute_input":"2022-08-06T14:44:02.595115Z","iopub.status.idle":"2022-08-06T14:44:03.018784Z","shell.execute_reply.started":"2022-08-06T14:44:02.595070Z","shell.execute_reply":"2022-08-06T14:44:03.017778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr = train_df.corr()\n\nfig, axes = plt.subplots(figsize=(25, 25))\nmask = np.zeros_like(corr)\nmask[np.triu_indices_from(mask)] = True\nsns.heatmap(corr, mask=mask, linewidths=.5, annot=True, cmap='rainbow')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:03.024019Z","iopub.execute_input":"2022-08-06T14:44:03.025196Z","iopub.status.idle":"2022-08-06T14:44:04.740336Z","shell.execute_reply.started":"2022-08-06T14:44:03.025143Z","shell.execute_reply":"2022-08-06T14:44:04.738949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I offer a function to fill in the missing values. Based on my experience from the June 2022 competition, the function uses a SimpleImputer or a IterativeImputer. The function parameter 'use_simple_imputer' specifies, which one is used. \nOnly the float columns/features have missing values, therefore can limit the dataframe on these columns.","metadata":{}},{"cell_type":"code","source":"def do_add_missing_values(df, use_simple_imputer:bool=True, strategy:str = 'mean'): #'median' # mean' #most_frequent):\n    \n    f_columns = list(df.columns[(df.dtypes.values == np.dtype('float64'))])\n    \n    imputer = None\n    if use_simple_imputer:\n        print(f\"use SimpleImputer with strategy {strategy} to fill missing float values\")\n        imputer = SimpleImputer(missing_values=np.nan, strategy=strategy)\n    else:\n        print(f\"use IterativeImputer with strategy {strategy} to fill missing float values\")\n        imp_order = 'arabic' \n        xgb = xgboost.XGBRegressor(n_estimators=500, random_state=718, tree_method='gpu_hist')\n        imputer = IterativeImputer(missing_values = np.nan,\n                                   estimator = xgb,\n                                   initial_strategy=strategy,\n                                   max_iter=10,\n                                   sample_posterior=False,\n                                   verbose=2,\n                                   imputation_order=imp_order,\n                                   random_state=718)\n\n    df[f_columns] = imputer.fit_transform(df[f_columns])\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.741751Z","iopub.execute_input":"2022-08-06T14:44:04.742182Z","iopub.status.idle":"2022-08-06T14:44:04.754507Z","shell.execute_reply.started":"2022-08-06T14:44:04.742144Z","shell.execute_reply":"2022-08-06T14:44:04.752684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if add_missing_values:\n    train_df = do_add_missing_values(train_df)\n    test_df = do_add_missing_values(test_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.756846Z","iopub.execute_input":"2022-08-06T14:44:04.757903Z","iopub.status.idle":"2022-08-06T14:44:04.809316Z","shell.execute_reply.started":"2022-08-06T14:44:04.757850Z","shell.execute_reply":"2022-08-06T14:44:04.807938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check if there are any missing values yet.","metadata":{}},{"cell_type":"code","source":"print(f\"Number of missing values in train_df : {train_df.isna().sum().sum()}\")\nprint(f\"Number of missing values in test_df  : {test_df.isna().sum().sum()}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.811056Z","iopub.execute_input":"2022-08-06T14:44:04.812225Z","iopub.status.idle":"2022-08-06T14:44:04.832677Z","shell.execute_reply.started":"2022-08-06T14:44:04.812184Z","shell.execute_reply":"2022-08-06T14:44:04.831303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I need a list of the column names, which are candidates for category variables and which are no candidates, also called continous variables. The Fastai library offers the function 'cont_cat_split' to do this for us. Our training data set contains only floating values for the independed variables, therefore we expect that no category variables are available.","metadata":{}},{"cell_type":"code","source":"cont_vars, cat_vars = cont_cat_split(train_df, dep_var=dep_var)\nlen(cont_vars), len(cat_vars),cont_vars, cat_vars","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.834508Z","iopub.execute_input":"2022-08-06T14:44:04.835200Z","iopub.status.idle":"2022-08-06T14:44:04.847133Z","shell.execute_reply.started":"2022-08-06T14:44:04.835159Z","shell.execute_reply":"2022-08-06T14:44:04.845715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next step is to create a data loader. The Fastai library offers a powerful helper called 'TabularPandas'. It needs the data frame, list of the category and continous variables, the depened variable and a splitter. The splitter divides the data set into two parts: one for the training and one for the validation and for internal optimization step in each epoch. The batch size is set to 128. I can use a random split because the rows in the data set are independed.","metadata":{}},{"cell_type":"code","source":"def getData(df, batchSize=128):\n    \n    to_train = TabularPandas(df, \n                           [Normalize, Categorify, FillMissing],\n                           cat_vars,\n                           cont_vars, \n                           splits=RandomSplitter(valid_pct=0.2)(df),  \n                           device = device,\n                           y_block=CategoryBlock(),\n                           y_names=dep_var) \n\n    return to_train.dataloaders(bs=batchSize)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.848477Z","iopub.execute_input":"2022-08-06T14:44:04.848906Z","iopub.status.idle":"2022-08-06T14:44:04.857313Z","shell.execute_reply.started":"2022-08-06T14:44:04.848860Z","shell.execute_reply":"2022-08-06T14:44:04.855949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = getData(train_df)\nlen(dls.train), len(dls.valid)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.859237Z","iopub.execute_input":"2022-08-06T14:44:04.860050Z","iopub.status.idle":"2022-08-06T14:44:04.977003Z","shell.execute_reply.started":"2022-08-06T14:44:04.859996Z","shell.execute_reply":"2022-08-06T14:44:04.975673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Show me the transformed data, which will be used in the network later.","metadata":{}},{"cell_type":"code","source":"dls.show_batch()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:04.978857Z","iopub.execute_input":"2022-08-06T14:44:04.979674Z","iopub.status.idle":"2022-08-06T14:44:05.046701Z","shell.execute_reply.started":"2022-08-06T14:44:04.979622Z","shell.execute_reply":"2022-08-06T14:44:05.045687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At least i create a learner pasing the dataloader into it. The default settings are two hidden layers with 200 and 100 elements. Increasing the number of parameters in the neural network will improve the accuarcy and score, hopefully.","metadata":{}},{"cell_type":"code","source":"my_config = tabular_config(ps=0.1, use_bn=True)\nlearn = tabular_learner(dls,\n                        config = my_config,\n                        metrics=[accuracy])\n\nlearn.summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:05.048136Z","iopub.execute_input":"2022-08-06T14:44:05.048739Z","iopub.status.idle":"2022-08-06T14:44:05.108435Z","shell.execute_reply.started":"2022-08-06T14:44:05.048702Z","shell.execute_reply":"2022-08-06T14:44:05.107501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr_min,lr_steep = learn.lr_find(suggest_funcs=(minimum, steep))\nprint(f\"Minimum: {lr_min:.2e}, steepest point: {lr_steep:.2e}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:05.109860Z","iopub.execute_input":"2022-08-06T14:44:05.110786Z","iopub.status.idle":"2022-08-06T14:44:06.940348Z","shell.execute_reply.started":"2022-08-06T14:44:05.110749Z","shell.execute_reply":"2022-08-06T14:44:06.939053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I will use a maximum learning rate of 5e-3. Starting the learning process is quite easy, i will run for 20 epochs. I will save the model with the best, with the lowest validation lost value. The Fastai library offers the SaveModelCallback callback. You must specify the file name only. The option with_opt=True stores the values of the optimizer also. You will find the new file in the subdirectory 'models'.","metadata":{}},{"cell_type":"code","source":"learn.fit_one_cycle(20, 5e-3, cbs=SaveModelCallback(fname='kaggle_tps_2022_aug', with_opt=True))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:06.942059Z","iopub.execute_input":"2022-08-06T14:44:06.943213Z","iopub.status.idle":"2022-08-06T14:44:51.752637Z","shell.execute_reply.started":"2022-08-06T14:44:06.943161Z","shell.execute_reply":"2022-08-06T14:44:51.751460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.show_results()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:51.754198Z","iopub.execute_input":"2022-08-06T14:44:51.756061Z","iopub.status.idle":"2022-08-06T14:44:51.855457Z","shell.execute_reply.started":"2022-08-06T14:44:51.756004Z","shell.execute_reply":"2022-08-06T14:44:51.853723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The confusion matrix below shows us the quality of test data predictions.\nI see that the neural netowrk is not able predict the failure value '1' correct, the network couldn't detect any failure, is is a bad news!","metadata":{}},{"cell_type":"code","source":"interp = ClassificationInterpretation.from_learner(learn)\ninterp.plot_confusion_matrix(normalize=True, norm_dec=4)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:51.857924Z","iopub.execute_input":"2022-08-06T14:44:51.858432Z","iopub.status.idle":"2022-08-06T14:44:52.682661Z","shell.execute_reply.started":"2022-08-06T14:44:51.858384Z","shell.execute_reply":"2022-08-06T14:44:52.681218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Low it's time to calculate the predictions for the test data set. Therefore i load the best model from the learining step above.","metadata":{}},{"cell_type":"code","source":"learn.load('kaggle_tps_2022_aug')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:52.684469Z","iopub.execute_input":"2022-08-06T14:44:52.688386Z","iopub.status.idle":"2022-08-06T14:44:52.717258Z","shell.execute_reply.started":"2022-08-06T14:44:52.688315Z","shell.execute_reply":"2022-08-06T14:44:52.715852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dlt = learn.dls.test_dl(test_df, bs=1024) \npreds = learn.get_preds(dl=dlt)[0].numpy()[:, 1]\npreds","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:52.724190Z","iopub.execute_input":"2022-08-06T14:44:52.727773Z","iopub.status.idle":"2022-08-06T14:44:52.972268Z","shell.execute_reply.started":"2022-08-06T14:44:52.727705Z","shell.execute_reply":"2022-08-06T14:44:52.970975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission[dep_var] = preds\nsample_submission.to_csv(\"submission.csv\", index=False)\nsample_submission.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:52.974297Z","iopub.execute_input":"2022-08-06T14:44:52.974882Z","iopub.status.idle":"2022-08-06T14:44:53.031058Z","shell.execute_reply.started":"2022-08-06T14:44:52.974834Z","shell.execute_reply":"2022-08-06T14:44:53.028808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class PermutationImportance():\n      \"Calculate and plot the permutation importance\"\n      def __init__(self, learn:Learner, df=None, bs=None):\n        \"Initialize with a test dataframe, a learner, and a metric\"\n        self.learn = learn\n        self.df = df\n        bs = bs if bs is not None else learn.dls.bs\n        if self.df is not None:\n          self.dl = learn.dls.test_dl(self.df, bs=bs)\n        else:\n          self.dl = learn.dls[1]\n        self.x_names = learn.dls.x_names.filter(lambda x: '_na' not in x)\n        self.na = learn.dls.x_names.filter(lambda x: '_na' in x)\n        self.y = dls.y_names\n        self.results = self.calc_feat_importance()\n        self.plot_importance(self.ord_dic_to_df(self.results))\n    \n        \n      def get_results(self): return self.results\n        \n      def measure_col(self, name:str):\n          \"Measures change after column shuffle\"\n          col = [name]\n          if f'{name}_na' in self.na: col.append(name)\n          orig = self.dl.items[col].values\n          perm = np.random.permutation(len(orig))\n          self.dl.items[col] = self.dl.items[col].values[perm]\n          metric = learn.validate(dl=self.dl)[1]\n          self.dl.items[col] = orig\n          clear_output()\n          return metric\n\n      def calc_feat_importance(self):\n          \"Calculates permutation importance by shuffling a column on a percentage scale\"\n          print('Getting base error')\n          base_error = self.learn.validate(dl=self.dl)[1]\n          self.importance = {}\n          pbar = progress_bar(self.x_names)\n          print('Calculating Permutation Importance')\n          for col in pbar:\n            self.importance[col] = self.measure_col(col)\n          for key, value in self.importance.items():\n            self.importance[key] = np.abs(base_error-value)/base_error #this can be adjusted\n          return OrderedDict(sorted(self.importance.items(), key=lambda kv: kv[1], reverse=True))\n\n      def ord_dic_to_df(self, dict:OrderedDict):\n          return pd.DataFrame([[k, v] for k, v in dict.items()], columns=['feature', 'importance'])\n\n      def plot_importance(self, df:pd.DataFrame, limit=20, asc=False, **kwargs):\n          \"Plot importance with an optional limit to how many variables shown\"\n          df_copy = df.copy()\n          df_copy['feature'] = df_copy['feature'].str.slice(0,40)\n          df_copy = df_copy.sort_values(by='importance', ascending=asc)[:limit].sort_values(by='importance', ascending=not(asc))\n          ax = df_copy.plot.barh(x='feature', y='importance', sort_columns=True, **kwargs)\n          for p in ax.patches:\n            ax.annotate(f'{p.get_width():.4f}', ((p.get_width() * 1.005), p.get_y()  * 1.005))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:53.032631Z","iopub.execute_input":"2022-08-06T14:44:53.033380Z","iopub.status.idle":"2022-08-06T14:44:53.054029Z","shell.execute_reply.started":"2022-08-06T14:44:53.033331Z","shell.execute_reply":"2022-08-06T14:44:53.052982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res = PermutationImportance(learn, train_df, bs=256)\nres.get_results()","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:44:53.055760Z","iopub.execute_input":"2022-08-06T14:44:53.056945Z","iopub.status.idle":"2022-08-06T14:45:13.619041Z","shell.execute_reply.started":"2022-08-06T14:44:53.056895Z","shell.execute_reply":"2022-08-06T14:45:13.617672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -la","metadata":{"execution":{"iopub.status.busy":"2022-08-06T14:45:13.624209Z","iopub.execute_input":"2022-08-06T14:45:13.624658Z","iopub.status.idle":"2022-08-06T14:45:14.732017Z","shell.execute_reply.started":"2022-08-06T14:45:13.624619Z","shell.execute_reply":"2022-08-06T14:45:14.730069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see we achieve an accuracy value of 0.577 with the default fastai settings and with a minimal features engineering for the missing values. That is a greate result and is the baseline to investigate in more feature engineering and/or modeling to get a better final result. At this point you can start your own experience. Feel free and use my notebook if you like. \nShare your results.","metadata":{}}]}