{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook is a quick demonstration, who to use the Fastai v2 library for a Kaggle tabular competition. Fastai v2 is based on pytorch and allows you, to build a decent machine learning application. For more information please visit the Fastai documentation: https://docs.fast.ai/. I will link to \"Chapter 9, Tabular Modelling Deep Dive\" and the notebook \"09_tabular.ipynb\".\n\nThis competition is a binary classification problem: find the correct state, wheter a passenger is transported. The offered dataset has 14 differend features and for many rows, some values are missing.\nIn this notebook i will use a neural network approach and i will train this network with the traing data set.\n\nLet's start and import the needed stuff ..","metadata":{}},{"cell_type":"code","source":"from fastai.tabular.all import * \nfrom fastai.test_utils import show_install\nfrom IPython.display import display, clear_output\n\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import SimpleImputer, IterativeImputer \n\nimport seaborn as sns\nshow_install()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-16T10:59:33.930714Z","iopub.execute_input":"2022-07-16T10:59:33.931325Z","iopub.status.idle":"2022-07-16T10:59:35.667849Z","shell.execute_reply.started":"2022-07-16T10:59:33.931228Z","shell.execute_reply":"2022-07-16T10:59:35.666661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\ndevice","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.672499Z","iopub.execute_input":"2022-07-16T10:59:35.675392Z","iopub.status.idle":"2022-07-16T10:59:35.689533Z","shell.execute_reply.started":"2022-07-16T10:59:35.675344Z","shell.execute_reply":"2022-07-16T10:59:35.688579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed_value(seed=718):\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n\nset_seed_value()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.693893Z","iopub.execute_input":"2022-07-16T10:59:35.696400Z","iopub.status.idle":"2022-07-16T10:59:35.703549Z","shell.execute_reply.started":"2022-07-16T10:59:35.696361Z","shell.execute_reply":"2022-07-16T10:59:35.702513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = Path('../input/spaceship-titanic/')\nPath.BASE_PATH = path\npath.ls()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.709352Z","iopub.execute_input":"2022-07-16T10:59:35.712036Z","iopub.status.idle":"2022-07-16T10:59:35.723228Z","shell.execute_reply.started":"2022-07-16T10:59:35.711993Z","shell.execute_reply":"2022-07-16T10:59:35.722489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Load the datasets and define the depending variable: Transported","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(os.path.join(path, 'train.csv'))\ntest_df = pd.read_csv(os.path.join(path, 'test.csv'))\nsample_submission = pd.read_csv(os.path.join(path, 'sample_submission.csv'))\n\ndep_var = 'Transported'","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.725755Z","iopub.execute_input":"2022-07-16T10:59:35.727380Z","iopub.status.idle":"2022-07-16T10:59:35.801491Z","shell.execute_reply.started":"2022-07-16T10:59:35.727351Z","shell.execute_reply":"2022-07-16T10:59:35.800696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see the first rows of our training data set to get an overview:","metadata":{}},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.805230Z","iopub.execute_input":"2022-07-16T10:59:35.805913Z","iopub.status.idle":"2022-07-16T10:59:35.841864Z","shell.execute_reply.started":"2022-07-16T10:59:35.805873Z","shell.execute_reply":"2022-07-16T10:59:35.841015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see the columns and their types. The function info() shows the number of rows with values for each column. If these numbers differ from row to row and the total amount of rows, we have a dataset with missing values, mostly NaN named.","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.845571Z","iopub.execute_input":"2022-07-16T10:59:35.845925Z","iopub.status.idle":"2022-07-16T10:59:35.876582Z","shell.execute_reply.started":"2022-07-16T10:59:35.845890Z","shell.execute_reply":"2022-07-16T10:59:35.875861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.880587Z","iopub.execute_input":"2022-07-16T10:59:35.882772Z","iopub.status.idle":"2022-07-16T10:59:35.927007Z","shell.execute_reply.started":"2022-07-16T10:59:35.882735Z","shell.execute_reply":"2022-07-16T10:59:35.926104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's print the number of NaN rows for each column.","metadata":{}},{"cell_type":"code","source":"print(train_df.isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.931022Z","iopub.execute_input":"2022-07-16T10:59:35.933120Z","iopub.status.idle":"2022-07-16T10:59:35.952410Z","shell.execute_reply.started":"2022-07-16T10:59:35.933084Z","shell.execute_reply":"2022-07-16T10:59:35.951469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I will offer two different solutions for handling the missing values.\nThe first solution bases on the passgener can't consume any amenities when he is cyrosleep state: If the value 'CryoSleep' is True, all missign values for 'ShoppingMall', 'Spa', 'FoodCourt' and 'RoomService' must be 0. An if one of these values is greater 0, the passgener is using one of the amenities, the value for 'CyroSleep' must be False. \nThe second solution is inspired by the tabular playground series for July 2022: https://www.kaggle.com/competitions/tabular-playground-series-jun-2022.","metadata":{}},{"cell_type":"code","source":"def do_correct_cryoSleep(df):\n    print(\"Correct NaN value for cyroSleep\")\n    df['CryoSleep'] = np.where((df['CryoSleep'].isnull()) & \n                               ((df['RoomService'] == 0.0) & (df['FoodCourt'] == 0.0) & \n                                (df['ShoppingMall'] == 0.0) & (df['Spa'] == 0.0) & \n                                (df['VRDeck'] == 0.0)), True, df['CryoSleep'])\n    \n    df['CryoSleep'] = np.where((df['CryoSleep'].isnull()) & \n                               ((df['RoomService'] > 0.0) | (df['FoodCourt'] > 0.0) | \n                                (df['ShoppingMall'] > 0.0) | (df['Spa'] > 0.0) | \n                                (df['VRDeck'] > 0.0)), False, df['CryoSleep'])\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.956636Z","iopub.execute_input":"2022-07-16T10:59:35.958897Z","iopub.status.idle":"2022-07-16T10:59:35.969563Z","shell.execute_reply.started":"2022-07-16T10:59:35.958861Z","shell.execute_reply":"2022-07-16T10:59:35.968673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def do_correct_amenity_values(df, add_bool_flag_only:bool=False):\n    print(\"Correct NaN value for the amenities\")\n    amenity_cols = ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']\n    \n    if add_bool_flag_only:\n        df['consumes_amenities'] = (df[amenity_cols].sum(axis=1)>0.0)\n        df.drop(amenity_cols, axis=1, inplace=True)        \n    else:\n        for service in amenity_cols:\n\n            df[service] = np.where((df[service].isnull()) & (df['CryoSleep'] == True), 0.0, df[service])\n         #   df[service] = df[service]/1024.0\n            df[service] = df[service].fillna(0)\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.973952Z","iopub.execute_input":"2022-07-16T10:59:35.976675Z","iopub.status.idle":"2022-07-16T10:59:35.986599Z","shell.execute_reply.started":"2022-07-16T10:59:35.976628Z","shell.execute_reply":"2022-07-16T10:59:35.985591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = do_correct_cryoSleep(train_df)\ntest_df = do_correct_cryoSleep(test_df)\ntrain_df = do_correct_amenity_values(train_df)\ntest_df = do_correct_amenity_values(test_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:35.991615Z","iopub.execute_input":"2022-07-16T10:59:35.992042Z","iopub.status.idle":"2022-07-16T10:59:36.042871Z","shell.execute_reply.started":"2022-07-16T10:59:35.992004Z","shell.execute_reply":"2022-07-16T10:59:36.041685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The name values are mostly unique and should'nt have any influence on the 'Transported' value, therefore i will remove them.","metadata":{}},{"cell_type":"code","source":"train_df.drop(['Name'], axis=1, inplace=True)\ntest_df.drop(['Name'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.049692Z","iopub.execute_input":"2022-07-16T10:59:36.050091Z","iopub.status.idle":"2022-07-16T10:59:36.063515Z","shell.execute_reply.started":"2022-07-16T10:59:36.050054Z","shell.execute_reply":"2022-07-16T10:59:36.061770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As mentioned the description of the data, the values inn the columns 'Cabin' and 'PassengerId' are combined string values, The values for 'Cabin' constist of the value for the deck, a number and s side value. The values for 'PassengerId' are are the string concatenation of a group value and unique number inside this group. \nI will define a function to replace the values in the columns 'Cabin' and 'PassengerId' with these sub values. The original columns can be droped.","metadata":{}},{"cell_type":"code","source":"def split_columns_with_combinded_data(df, drop_orgin:bool=False):\n    df[['Deck','Num', 'Side']] = df['Cabin'].str.split('/', expand=True)\n    df[['PGroup','PNr']] = df['PassengerId'].str.split('_', expand=True)\n    if drop_orgin:\n        df.drop(['Cabin', 'PassengerId'], axis=1, inplace=True)\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.069691Z","iopub.execute_input":"2022-07-16T10:59:36.070164Z","iopub.status.idle":"2022-07-16T10:59:36.080217Z","shell.execute_reply.started":"2022-07-16T10:59:36.070127Z","shell.execute_reply":"2022-07-16T10:59:36.079414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = split_columns_with_combinded_data(train_df, drop_orgin=True)\ntest_df = split_columns_with_combinded_data(test_df, drop_orgin=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.084613Z","iopub.execute_input":"2022-07-16T10:59:36.087581Z","iopub.status.idle":"2022-07-16T10:59:36.346212Z","shell.execute_reply.started":"2022-07-16T10:59:36.087530Z","shell.execute_reply":"2022-07-16T10:59:36.344760Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that the columns 'Destination' and 'HomePlanet' are enumeration types. I will print the number of thier unique values to check my assumption.","metadata":{}},{"cell_type":"code","source":"train_df['Destination'].nunique(), train_df['HomePlanet'].nunique(),","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.350875Z","iopub.execute_input":"2022-07-16T10:59:36.353685Z","iopub.status.idle":"2022-07-16T10:59:36.369947Z","shell.execute_reply.started":"2022-07-16T10:59:36.353590Z","shell.execute_reply":"2022-07-16T10:59:36.367348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.376194Z","iopub.execute_input":"2022-07-16T10:59:36.378376Z","iopub.status.idle":"2022-07-16T10:59:36.409095Z","shell.execute_reply.started":"2022-07-16T10:59:36.378338Z","shell.execute_reply":"2022-07-16T10:59:36.407385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Okay there a a handful unique values in both columns. That's lighten handling of the missing values in these columns. We can convert them into 'one-hot-encoded' values. If a row contains a Nan value in the orginal column, none of the derived rows contain the value 1, all rows have the value 0.","metadata":{}},{"cell_type":"code","source":"def convert_to_dummies(df):\n    df = pd.get_dummies(df, columns=['Destination'], prefix=\"D\")\n    df = pd.get_dummies(df, columns=['HomePlanet'], prefix=\"H\")\n    df = pd.get_dummies(df, columns=['Side'])\n    df = pd.get_dummies(df, columns=['VIP'])\n    df = pd.get_dummies(df, columns=['CryoSleep'])\n    df = pd.get_dummies(df, columns=['Deck'])\n  \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.413903Z","iopub.execute_input":"2022-07-16T10:59:36.416020Z","iopub.status.idle":"2022-07-16T10:59:36.425294Z","shell.execute_reply.started":"2022-07-16T10:59:36.415982Z","shell.execute_reply":"2022-07-16T10:59:36.424323Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = convert_to_dummies(train_df)\ntest_df = convert_to_dummies(test_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.430182Z","iopub.execute_input":"2022-07-16T10:59:36.432933Z","iopub.status.idle":"2022-07-16T10:59:36.518896Z","shell.execute_reply.started":"2022-07-16T10:59:36.432894Z","shell.execute_reply":"2022-07-16T10:59:36.517975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I will change the type for some columns, that needed for the neural network later ..","metadata":{}},{"cell_type":"code","source":"def change_column_type(df):\n    df['Num'] = df['Num'].astype('float')\n    df['PNr'] = df['PNr'].astype('int')\n    df['PGroup'] = df['PGroup'].astype('int')\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.522707Z","iopub.execute_input":"2022-07-16T10:59:36.524808Z","iopub.status.idle":"2022-07-16T10:59:36.532518Z","shell.execute_reply.started":"2022-07-16T10:59:36.524770Z","shell.execute_reply":"2022-07-16T10:59:36.531519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = change_column_type(train_df)\ntestn_df = change_column_type(test_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.536897Z","iopub.execute_input":"2022-07-16T10:59:36.539672Z","iopub.status.idle":"2022-07-16T10:59:36.563307Z","shell.execute_reply.started":"2022-07-16T10:59:36.539632Z","shell.execute_reply":"2022-07-16T10:59:36.561939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's recheck the number of NaN rows for each column after the preprocessing:","metadata":{}},{"cell_type":"code","source":"print(train_df.isna().sum())","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.570749Z","iopub.execute_input":"2022-07-16T10:59:36.571508Z","iopub.status.idle":"2022-07-16T10:59:36.589993Z","shell.execute_reply.started":"2022-07-16T10:59:36.571462Z","shell.execute_reply":"2022-07-16T10:59:36.589222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First i will look at the correlation matrix to verify how important a feature is.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15,15))\n\ncorr=train_df.corr()\nmask = np.triu(np.ones_like(corr, dtype=bool))\nsns.heatmap(corr, mask=mask, robust=True, center=0,square=True, linewidths=.6,cmap='rainbow')\nplt.title('Correlation')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:36.592996Z","iopub.execute_input":"2022-07-16T10:59:36.596380Z","iopub.status.idle":"2022-07-16T10:59:37.868402Z","shell.execute_reply.started":"2022-07-16T10:59:36.596339Z","shell.execute_reply":"2022-07-16T10:59:37.867260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I need a list of the column names, which are candidates for category variables and which are no candidates, also called continous variables. The Fastai library offers the function 'cont_cat_split' to do this for us. Our training data set contains only floating values for the independed variables, therefore we expect that no category variables are available.","metadata":{}},{"cell_type":"code","source":"cont_vars, cat_vars = cont_cat_split(train_df, dep_var=dep_var)\ncont_vars, cat_vars","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:37.872573Z","iopub.execute_input":"2022-07-16T10:59:37.872962Z","iopub.status.idle":"2022-07-16T10:59:37.894891Z","shell.execute_reply.started":"2022-07-16T10:59:37.872922Z","shell.execute_reply":"2022-07-16T10:59:37.894053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The next step is to create a data loader. The Fastai library offers a powerful helper called 'TabularPandas'. It needs the data frame, list of the category and continous variables, the depened variable and a splitter. The splitter divides the data set into two parts: one for the training and one for the validation and for internal optimization step in each epoch. The batch size is set to 1024, because we have a large data set. We can use a random split because the rows in the data set are independed.","metadata":{}},{"cell_type":"code","source":"def getData(df, batchSize=128):\n    \n    to_train = TabularPandas(df, \n                           [Normalize, Categorify, FillMissing],\n                           cat_vars,\n                           cont_vars, \n                           splits=RandomSplitter(valid_pct=0.2)(df),  \n                           device = device,\n                           y_block=CategoryBlock(),\n                           y_names=dep_var) \n\n    return to_train.dataloaders(bs=batchSize)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:37.895995Z","iopub.execute_input":"2022-07-16T10:59:37.900414Z","iopub.status.idle":"2022-07-16T10:59:37.908535Z","shell.execute_reply.started":"2022-07-16T10:59:37.900373Z","shell.execute_reply":"2022-07-16T10:59:37.907676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls = getData(train_df)\nlen(dls.train), len(dls.valid)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:37.912938Z","iopub.execute_input":"2022-07-16T10:59:37.915599Z","iopub.status.idle":"2022-07-16T10:59:40.245093Z","shell.execute_reply.started":"2022-07-16T10:59:37.915560Z","shell.execute_reply":"2022-07-16T10:59:40.244151Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Show me the transformed data, which will be used in the network later.","metadata":{}},{"cell_type":"code","source":"dls.show_batch()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:40.249910Z","iopub.execute_input":"2022-07-16T10:59:40.252142Z","iopub.status.idle":"2022-07-16T10:59:40.342775Z","shell.execute_reply.started":"2022-07-16T10:59:40.252100Z","shell.execute_reply":"2022-07-16T10:59:40.342005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"At least i create a learner pasing the dataloader into it. The default settings are two hidden layers with 200 and 100 elements. Increasing the number of parameters in the neural network will improve the accuarcy and score, hopefully: Change number and the depth of the hidden layers, use a a batch normalization and/or a dropout layer, etc.","metadata":{}},{"cell_type":"code","source":"my_config = tabular_config(y_range=(0,1), use_bn=True, ps=0.1, embed_p=0.1)\n\nlearn = tabular_learner(dls,\n                        config = my_config,\n                        layers=[200,100],\n                        metrics=[accuracy])\n\nlearn.summary()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:40.346789Z","iopub.execute_input":"2022-07-16T10:59:40.349177Z","iopub.status.idle":"2022-07-16T10:59:40.787475Z","shell.execute_reply.started":"2022-07-16T10:59:40.349138Z","shell.execute_reply":"2022-07-16T10:59:40.786570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We needd a proper leraning rate. The The Fastai library offers the funtcion lr_find() for this job.","metadata":{}},{"cell_type":"code","source":"lr_min,lr_steep = learn.lr_find(suggest_funcs=(minimum, steep))","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:40.791374Z","iopub.execute_input":"2022-07-16T10:59:40.793786Z","iopub.status.idle":"2022-07-16T10:59:44.567541Z","shell.execute_reply.started":"2022-07-16T10:59:40.793739Z","shell.execute_reply":"2022-07-16T10:59:44.566529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Minimum/10: {lr_min:.2e}, steepest point: {lr_steep:.2e}\")","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:44.571841Z","iopub.execute_input":"2022-07-16T10:59:44.573933Z","iopub.status.idle":"2022-07-16T10:59:44.582738Z","shell.execute_reply.started":"2022-07-16T10:59:44.573893Z","shell.execute_reply":"2022-07-16T10:59:44.581939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I will use a maximum learning rate of 5e-3. Starting the learning process is quite easy, i will run for 30 epochs. I will save the model with the best, with the lowest validation lost value. The Fastai library offers the SaveModelCallback callback. You must specify the file name only. The option with_opt=True stores the values of the optimizer also. You will find the new file in the subdirectory 'models'.","metadata":{}},{"cell_type":"code","source":"learn.fit_one_cycle(30, 5e-3, wd=0.01, cbs=SaveModelCallback(fname='kaggle_spaceship_titanic', with_opt=True))","metadata":{"execution":{"iopub.status.busy":"2022-07-16T10:59:44.587779Z","iopub.execute_input":"2022-07-16T10:59:44.588609Z","iopub.status.idle":"2022-07-16T11:00:21.669539Z","shell.execute_reply.started":"2022-07-16T10:59:44.588566Z","shell.execute_reply":"2022-07-16T11:00:21.668733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The confusion matrix below shows us the quality of data prediction during the learning phase.","metadata":{}},{"cell_type":"code","source":"interp = ClassificationInterpretation.from_learner(learn)\ninterp.plot_confusion_matrix(normalize=True, norm_dec=3)","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:21.673640Z","iopub.execute_input":"2022-07-16T11:00:21.675749Z","iopub.status.idle":"2022-07-16T11:00:22.157937Z","shell.execute_reply.started":"2022-07-16T11:00:21.675708Z","shell.execute_reply":"2022-07-16T11:00:22.157195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now it's time to calculate the predictions for the test data set. Therefore i load the 'best' model with the lowest validation loss value","metadata":{}},{"cell_type":"code","source":"learn.load('kaggle_spaceship_titanic')","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:22.162005Z","iopub.execute_input":"2022-07-16T11:00:22.164434Z","iopub.status.idle":"2022-07-16T11:00:22.230137Z","shell.execute_reply.started":"2022-07-16T11:00:22.164391Z","shell.execute_reply":"2022-07-16T11:00:22.229437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.show_results()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:22.234037Z","iopub.execute_input":"2022-07-16T11:00:22.236115Z","iopub.status.idle":"2022-07-16T11:00:22.319943Z","shell.execute_reply.started":"2022-07-16T11:00:22.236077Z","shell.execute_reply":"2022-07-16T11:00:22.319029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I got the 'one hot encoded' prediction values, which are probabilities for the different target values. np.argmax returns the index with the maximum probability value, like 0 or 1.","metadata":{}},{"cell_type":"code","source":"dlt = learn.dls.test_dl(test_df) \nnn_preds,_ ,preds = learn.get_preds(dl=dlt , with_decoded=True) \n\nnn_preds","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:22.323959Z","iopub.execute_input":"2022-07-16T11:00:22.326065Z","iopub.status.idle":"2022-07-16T11:00:22.635121Z","shell.execute_reply.started":"2022-07-16T11:00:22.326028Z","shell.execute_reply":"2022-07-16T11:00:22.634404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission[dep_var] = np.argmax(nn_preds, axis=1) == 1\nsample_submission.to_csv(\"submission.csv\", index=False)\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:22.638903Z","iopub.execute_input":"2022-07-16T11:00:22.640997Z","iopub.status.idle":"2022-07-16T11:00:22.668902Z","shell.execute_reply.started":"2022-07-16T11:00:22.640958Z","shell.execute_reply":"2022-07-16T11:00:22.668225Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class PermutationImportance():\n      \"Calculate and plot the permutation importance\"\n      def __init__(self, learn:Learner, df=None, bs=None):\n        \"Initialize with a test dataframe, a learner, and a metric\"\n        self.learn = learn\n        self.df = df\n        bs = bs if bs is not None else learn.dls.bs\n        if self.df is not None:\n          self.dl = learn.dls.test_dl(self.df, bs=bs)\n        else:\n          self.dl = learn.dls[1]\n        self.x_names = learn.dls.x_names.filter(lambda x: '_na' not in x)\n        self.na = learn.dls.x_names.filter(lambda x: '_na' in x)\n        self.y = dls.y_names\n        self.results = self.calc_feat_importance()\n        self.plot_importance(self.ord_dic_to_df(self.results))\n    \n        \n      def get_results(self): return self.results\n        \n      def measure_col(self, name:str):\n          \"Measures change after column shuffle\"\n          col = [name]\n          if f'{name}_na' in self.na: col.append(name)\n          orig = self.dl.items[col].values\n          perm = np.random.permutation(len(orig))\n          self.dl.items[col] = self.dl.items[col].values[perm]\n          metric = learn.validate(dl=self.dl)[1]\n          self.dl.items[col] = orig\n          clear_output()\n          return metric\n\n      def calc_feat_importance(self):\n          \"Calculates permutation importance by shuffling a column on a percentage scale\"\n          print('Getting base error')\n          base_error = self.learn.validate(dl=self.dl)[1]\n          self.importance = {}\n          pbar = progress_bar(self.x_names)\n          print('Calculating Permutation Importance')\n          for col in pbar:\n            self.importance[col] = self.measure_col(col)\n          for key, value in self.importance.items():\n            self.importance[key] = np.abs(base_error-value)/base_error #this can be adjusted\n          return OrderedDict(sorted(self.importance.items(), key=lambda kv: kv[1], reverse=True))\n\n      def ord_dic_to_df(self, dict:OrderedDict):\n          return pd.DataFrame([[k, v] for k, v in dict.items()], columns=['feature', 'importance'])\n\n      def plot_importance(self, df:pd.DataFrame, limit=20, asc=False, **kwargs):\n          \"Plot importance with an optional limit to how many variables shown\"\n          df_copy = df.copy()\n          df_copy['feature'] = df_copy['feature'].str.slice(0,25)\n          df_copy = df_copy.sort_values(by='importance', ascending=asc)[:limit].sort_values(by='importance', ascending=not(asc))\n          ax = df_copy.plot.barh(x='feature', y='importance', sort_columns=True, **kwargs)\n          for p in ax.patches:\n            ax.annotate(f'{p.get_width():.4f}', ((p.get_width() * 1.005), p.get_y()  * 1.005))","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:22.672971Z","iopub.execute_input":"2022-07-16T11:00:22.675150Z","iopub.status.idle":"2022-07-16T11:00:22.701756Z","shell.execute_reply.started":"2022-07-16T11:00:22.675113Z","shell.execute_reply":"2022-07-16T11:00:22.700697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"res = PermutationImportance(learn, train_df, bs=128)\nres.get_results()","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:22.705617Z","iopub.execute_input":"2022-07-16T11:00:22.707992Z","iopub.status.idle":"2022-07-16T11:00:41.368533Z","shell.execute_reply.started":"2022-07-16T11:00:22.707924Z","shell.execute_reply":"2022-07-16T11:00:41.367743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls -la","metadata":{"execution":{"iopub.status.busy":"2022-07-16T11:00:41.372514Z","iopub.execute_input":"2022-07-16T11:00:41.374883Z","iopub.status.idle":"2022-07-16T11:00:42.233291Z","shell.execute_reply.started":"2022-07-16T11:00:41.374842Z","shell.execute_reply":"2022-07-16T11:00:42.232103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see i achieve an accuracy value roughly 0.802 - 0.805 with the default Fastai settings and with a minimal features engineering. That is a great result and is the baseline to investigate in more feature engineering and/or modeling to get a better final result. At this point you can start your own experience. Fell free and use my notebokk if you like, or tell me your concerns. Feedback is wellcome!","metadata":{}}]}