{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from IPython.core.display import display, HTML, Javascript\n\n# ----- Notebook Theme -----\ncolor_map = ['#16a085', '#e8f6f3', '#d0ece7', '#a2d9ce', '#73c6b6', '#45b39d', \n                        '#16a085', '#138d75', '#117a65', '#0e6655', '#0b5345']\n\nprompt = color_map[-1]\nmain_color = color_map[0]\nstrong_main_color = color_map[1]\ncustom_colors = [strong_main_color, main_color]\n\ncss_file = ''' \n\ndiv #notebook {\nbackground-color: white;\nline-height: 20px;\n}\n\n#notebook-container {\n%s\nmargin-top: 2em;\npadding-top: 2em;\nborder-top: 4px solid %s; /* light orange */\n-webkit-box-shadow: 0px 0px 8px 2px rgba(224, 212, 226, 0.5); /* pink */\n    box-shadow: 0px 0px 8px 2px rgba(224, 212, 226, 0.5); /* pink */\n}\n\ndiv .input {\nmargin-bottom: 1em;\n}\n\n.rendered_html h1, .rendered_html h2, .rendered_html h3, .rendered_html h4, .rendered_html h5, .rendered_html h6 {\ncolor: %s; /* light orange */\nfont-weight: 600;\n}\n\ndiv.input_area {\nborder: none;\n    background-color: %s; /* rgba(229, 143, 101, 0.1); light orange [exactly #E58F65] */\n    border-top: 2px solid %s; /* light orange */\n}\n\ndiv.input_prompt {\ncolor: %s; /* light blue */\n}\n\ndiv.output_prompt {\ncolor: %s; /* strong orange */\n}\n\ndiv.cell.selected:before, div.cell.selected.jupyter-soft-selected:before {\nbackground: %s; /* light orange */\n}\n\ndiv.cell.selected, div.cell.selected.jupyter-soft-selected {\n    border-color: %s; /* light orange */\n}\n\n.edit_mode div.cell.selected:before {\nbackground: %s; /* light orange */\n}\n\n.edit_mode div.cell.selected {\nborder-color: %s; /* light orange */\n\n}\n'''\ndef to_rgb(h): \n    return tuple(int(h[i:i+2], 16) for i in [0, 2, 4])\n\nmain_color_rgba = 'rgba(%s, %s, %s, 0.1)' % (to_rgb(main_color[1:]))\nopen('notebook.css', 'w').write(css_file % ('width: 95%;', main_color, main_color, main_color_rgba, main_color,  main_color, prompt, main_color, main_color, main_color, main_color))\n\ndef nb(): \n    return HTML(\"<style>\" + open(\"notebook.css\", \"r\").read() + \"</style>\")\nnb()","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-01T13:27:30.918089Z","iopub.execute_input":"2022-08-01T13:27:30.918537Z","iopub.status.idle":"2022-08-01T13:27:30.964135Z","shell.execute_reply.started":"2022-08-01T13:27:30.918452Z","shell.execute_reply":"2022-08-01T13:27:30.962679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<img src=\"https://github.com/AILab-MLTools/LightAutoML/raw/master/imgs/LightAutoML_logo_big.png\" alt=\"LightAutoML logo\" style=\"width:70%;\"/>","metadata":{"papermill":{"duration":0.032379,"end_time":"2021-06-22T20:10:29.835505","exception":false,"start_time":"2021-06-22T20:10:29.803126","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# LightAutoML baseline\n\n### Welcome back from 2x Kaggle Grandmaster 😎\n\nOfficial LightAutoML github repository is [here](https://github.com/sb-ai-lab/LightAutoML). \n\n### Do not forget to put upvote for the notebook, follow me on Kaggle and the ⭐️ for github repo if you like it - one click for you, great pleasure for us ☺️ ","metadata":{}},{"cell_type":"code","source":"s = '<iframe src=\"https://ghbtns.com/github-btn.html?user=sb-ai-lab&repo=LightAutoML&type=star&count=true&size=large\" frameborder=\"0\" scrolling=\"0\" width=\"170\" height=\"30\" title=\"LightAutoML GitHub\"></iframe>'\nHTML(s)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-01T13:27:32.392031Z","iopub.execute_input":"2022-08-01T13:27:32.392447Z","iopub.status.idle":"2022-08-01T13:27:32.405310Z","shell.execute_reply.started":"2022-08-01T13:27:32.392411Z","shell.execute_reply":"2022-08-01T13:27:32.404364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## This notebook is the updated copy of our [Tutorial_1 from the GIT repository](https://github.com/sb-ai-lab/LightAutoML/blob/master/examples/tutorials/Tutorial_1_basics.ipynb). Please check our [tutorials folder](https://github.com/sb-ai-lab/LightAutoML/blob/master/examples/tutorials) if you are interested in other examples of LightAutoML functionality.","metadata":{}},{"cell_type":"markdown","source":"## 0. Prerequisites","metadata":{}},{"cell_type":"markdown","source":"### 0.0. install LightAutoML","metadata":{}},{"cell_type":"code","source":"%%capture\n!pip install lightautoml\n\n# QUICK WORKAROUND FOR PROBLEM WITH PANDAS\n!pip install -U pandas","metadata":{"_kg_hide-output":true,"papermill":{"duration":23.023261,"end_time":"2021-06-22T20:10:52.955691","exception":false,"start_time":"2021-06-22T20:10:29.93243","status":"completed"},"scrolled":true,"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:27:34.564444Z","iopub.execute_input":"2022-08-01T13:27:34.565170Z","iopub.status.idle":"2022-08-01T13:30:12.522797Z","shell.execute_reply.started":"2022-08-01T13:27:34.565129Z","shell.execute_reply":"2022-08-01T13:30:12.519957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 0.1. Import libraries\n\nHere we will import the libraries we use in this kernel:\n- Standard python libraries for timing, working with OS etc.\n- Essential python DS libraries like numpy, pandas, scikit-learn and torch (the last we will use in the next cell)\n- LightAutoML modules: `TabularAutoML` preset for AutoML model creation and `Task` class to setup what kind of ML problem we solve (binary/multiclass classification or regression)","metadata":{"papermill":{"duration":0.066681,"end_time":"2021-06-22T20:10:53.090975","exception":false,"start_time":"2021-06-22T20:10:53.024294","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Standard python libraries\nimport os\nimport time\n\n# Essential DS libraries\nimport numpy as np\nimport pandas as pd\nfrom sklearn.metrics import roc_auc_score\nimport torch\n\n# LightAutoML presets, task and report generation\nfrom lightautoml.automl.presets.tabular_presets import TabularAutoML, TabularUtilizedAutoML\nfrom lightautoml.tasks import Task","metadata":{"papermill":{"duration":8.32949,"end_time":"2021-06-22T20:11:01.487788","exception":false,"start_time":"2021-06-22T20:10:53.158298","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:12.528062Z","iopub.execute_input":"2022-08-01T13:30:12.529013Z","iopub.status.idle":"2022-08-01T13:30:15.665730Z","shell.execute_reply.started":"2022-08-01T13:30:12.528943Z","shell.execute_reply":"2022-08-01T13:30:15.663939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 0.2. Constants\n\nHere we setup the constants to use in the kernel:\n- `N_THREADS` - number of vCPUs for LightAutoML model creation\n- `N_FOLDS` - number of folds in LightAutoML inner CV\n- `RANDOM_STATE` - random seed for better reproducibility\n- `TIMEOUT` - limit in seconds for model to train\n- `TARGET_NAME` - target column name in dataset","metadata":{"papermill":{"duration":0.064234,"end_time":"2021-06-22T20:11:01.61901","exception":false,"start_time":"2021-06-22T20:11:01.554776","status":"completed"},"tags":[]}},{"cell_type":"code","source":"N_THREADS = 4\nN_FOLDS = 5\nRANDOM_STATE = 42\nTIMEOUT = 1800 # equal to 30 minutes\nTARGET_NAME = 'failure'","metadata":{"papermill":{"duration":0.077787,"end_time":"2021-06-22T20:11:01.76103","exception":false,"start_time":"2021-06-22T20:11:01.683243","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:15.668019Z","iopub.execute_input":"2022-08-01T13:30:15.668541Z","iopub.status.idle":"2022-08-01T13:30:15.675924Z","shell.execute_reply.started":"2022-08-01T13:30:15.668487Z","shell.execute_reply":"2022-08-01T13:30:15.674578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 0.3. Imported models setup\n\nFor better reproducibility fix numpy random seed with max number of threads for Torch (which usually try to use all the threads on server):","metadata":{"papermill":{"duration":0.086481,"end_time":"2021-06-22T20:11:01.927314","exception":false,"start_time":"2021-06-22T20:11:01.840833","status":"completed"},"tags":[]}},{"cell_type":"code","source":"np.random.seed(RANDOM_STATE)\ntorch.set_num_threads(N_THREADS)","metadata":{"papermill":{"duration":0.087268,"end_time":"2021-06-22T20:11:02.092497","exception":false,"start_time":"2021-06-22T20:11:02.005229","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:15.679539Z","iopub.execute_input":"2022-08-01T13:30:15.680158Z","iopub.status.idle":"2022-08-01T13:30:15.729134Z","shell.execute_reply.started":"2022-08-01T13:30:15.680118Z","shell.execute_reply":"2022-08-01T13:30:15.727866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 0.4. Data loading\nLet's check the data we have:","metadata":{"papermill":{"duration":0.072033,"end_time":"2021-06-22T20:11:02.238196","exception":false,"start_time":"2021-06-22T20:11:02.166163","status":"completed"},"tags":[]}},{"cell_type":"code","source":"INPUT_DIR = '../input/tabular-playground-series-aug-2022/'","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:30:15.735489Z","iopub.execute_input":"2022-08-01T13:30:15.738482Z","iopub.status.idle":"2022-08-01T13:30:15.746599Z","shell.execute_reply.started":"2022-08-01T13:30:15.738422Z","shell.execute_reply":"2022-08-01T13:30:15.745439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = pd.read_csv(INPUT_DIR + 'train.csv')\nprint(train_data.shape)\ntrain_data.head()","metadata":{"papermill":{"duration":12.710747,"end_time":"2021-06-22T20:11:15.01836","exception":false,"start_time":"2021-06-22T20:11:02.307613","status":"completed"},"scrolled":true,"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:15.751891Z","iopub.execute_input":"2022-08-01T13:30:15.755509Z","iopub.status.idle":"2022-08-01T13:30:16.162997Z","shell.execute_reply.started":"2022-08-01T13:30:15.755196Z","shell.execute_reply":"2022-08-01T13:30:16.162023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data = pd.read_csv(INPUT_DIR + 'test.csv')\nprint(test_data.shape)\ntest_data.head()","metadata":{"papermill":{"duration":0.077509,"end_time":"2021-06-22T20:11:15.161419","exception":false,"start_time":"2021-06-22T20:11:15.08391","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:16.164403Z","iopub.execute_input":"2022-08-01T13:30:16.165535Z","iopub.status.idle":"2022-08-01T13:30:16.361071Z","shell.execute_reply.started":"2022-08-01T13:30:16.165492Z","shell.execute_reply":"2022-08-01T13:30:16.359495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv(INPUT_DIR + 'sample_submission.csv')\nprint(submission.shape)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:30:16.362636Z","iopub.execute_input":"2022-08-01T13:30:16.362947Z","iopub.status.idle":"2022-08-01T13:30:16.389507Z","shell.execute_reply.started":"2022-08-01T13:30:16.362920Z","shell.execute_reply":"2022-08-01T13:30:16.388541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Task definition","metadata":{"papermill":{"duration":0.071526,"end_time":"2021-06-22T20:11:22.853156","exception":false,"start_time":"2021-06-22T20:11:22.78163","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### 1.1. Task type\n\nOn the cell below we create Task object - the class to setup what task LightAutoML model should solve with specific loss and metric if necessary (more info can be found [here](https://lightautoml.readthedocs.io/en/latest/pages/modules/generated/lightautoml.tasks.base.Task.html#lightautoml.tasks.base.Task) in our documentation):","metadata":{}},{"cell_type":"code","source":"task = Task('binary', )","metadata":{"papermill":{"duration":0.086442,"end_time":"2021-06-22T20:11:23.010643","exception":false,"start_time":"2021-06-22T20:11:22.924201","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:16.390959Z","iopub.execute_input":"2022-08-01T13:30:16.391813Z","iopub.status.idle":"2022-08-01T13:30:16.403493Z","shell.execute_reply.started":"2022-08-01T13:30:16.391769Z","shell.execute_reply":"2022-08-01T13:30:16.401672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1.2. Feature roles setup","metadata":{"papermill":{"duration":0.070103,"end_time":"2021-06-22T20:11:23.150929","exception":false,"start_time":"2021-06-22T20:11:23.080826","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"To solve the task, we need to setup columns roles. The **only role you must setup is target role**, everything else (drop, numeric, categorical, group, weights etc.) is up to user - LightAutoML models have automatic columns typization inside:","metadata":{"papermill":{"duration":0.069372,"end_time":"2021-06-22T20:11:23.290153","exception":false,"start_time":"2021-06-22T20:11:23.220781","status":"completed"},"tags":[]}},{"cell_type":"code","source":"roles = {\n    'target': TARGET_NAME,\n    'drop': ['id'],\n    'group': 'product_code'\n}","metadata":{"papermill":{"duration":0.07715,"end_time":"2021-06-22T20:11:23.43883","exception":false,"start_time":"2021-06-22T20:11:23.36168","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2022-08-01T13:30:16.407511Z","iopub.execute_input":"2022-08-01T13:30:16.407955Z","iopub.status.idle":"2022-08-01T13:30:16.413957Z","shell.execute_reply.started":"2022-08-01T13:30:16.407917Z","shell.execute_reply":"2022-08-01T13:30:16.412710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1.3. LightAutoML model creation - TabularAutoML preset","metadata":{"papermill":{"duration":0.074284,"end_time":"2021-06-22T20:11:23.582462","exception":false,"start_time":"2021-06-22T20:11:23.508178","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"In next the cell we are going to create LightAutoML model with `TabularAutoML` class - preset with default model structure like in the image below:\n\n<img src=\"https://github.com/AILab-MLTools/LightAutoML/raw/master/imgs/tutorial_blackbox_pipeline.png\" alt=\"TabularAutoML preset pipeline\" style=\"width:85%;\"/>\n\nin just several lines. Let's discuss the params we can setup:\n- `task` - the type of the ML task (the only **must have** parameter)\n- `timeout` - time limit in seconds for model to train\n- `cpu_limit` - vCPU count for model to use\n- `reader_params` - parameter change for Reader object inside preset, which works on the first step of data preparation: automatic feature typization, preliminary almost-constant features, correct CV setup etc. For example, we setup `n_jobs` threads for typization algo, `cv` folds and `random_state` as inside CV seed.\n\n**Important note**: `reader_params` key is one of the YAML config keys, which is used inside `TabularAutoML` preset. [More details](https://github.com/AILab-MLTools/LightAutoML/blob/master/lightautoml/automl/presets/tabular_config.yml) on its structure with explanation comments can be found on the link attached. Each key from this config can be modified with user settings during preset object initialization. To get more info about different parameters setting (for example, ML algos which can be used in `general_params->use_algos`) please take a look at our [article on TowardsDataScience](https://towardsdatascience.com/lightautoml-preset-usage-tutorial-2cce7da6f936).\n\nMoreover, to receive the automatic report for our model we can use `ReportDeco` decorator and work with the decorated version in the same way as we do with usual one (more details in [this tutorial](https://github.com/AILab-MLTools/LightAutoML/blob/master/examples/tutorials/Tutorial_1_basics.ipynb))","metadata":{"papermill":{"duration":0.072649,"end_time":"2021-06-22T20:11:23.726154","exception":false,"start_time":"2021-06-22T20:11:23.653505","status":"completed"},"tags":[]}},{"cell_type":"code","source":"automl = TabularAutoML(\n    task = task, \n    timeout = TIMEOUT,\n    cpu_limit = N_THREADS,\n    reader_params = {'n_jobs': N_THREADS, 'cv': N_FOLDS, 'random_state': RANDOM_STATE}\n)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:30:16.416006Z","iopub.execute_input":"2022-08-01T13:30:16.416647Z","iopub.status.idle":"2022-08-01T13:30:16.452627Z","shell.execute_reply.started":"2022-08-01T13:30:16.416608Z","shell.execute_reply":"2022-08-01T13:30:16.451189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. AutoML training","metadata":{}},{"cell_type":"markdown","source":"To run autoML training use fit_predict method:\n- `train_data` - Dataset to train.\n- `roles` - Roles dict.\n- `verbose` - Controls the verbosity: the higher, the more messages.\n        <1  : messages are not displayed;\n        >=1 : the computation process for layers is displayed;\n        >=2 : the information about folds processing is also displayed;\n        >=3 : the hyperparameters optimization process is also displayed;\n        >=4 : the training process for every algorithm is displayed;\n\nNote: out-of-fold prediction is calculated during training and returned from the fit_predict method","metadata":{}},{"cell_type":"code","source":"%%time \noof_pred = automl.fit_predict(train_data, roles = roles, verbose = 3)","metadata":{"scrolled":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-01T13:30:16.454633Z","iopub.execute_input":"2022-08-01T13:30:16.455356Z","iopub.status.idle":"2022-08-01T13:37:01.919334Z","shell.execute_reply.started":"2022-08-01T13:30:16.455283Z","shell.execute_reply":"2022-08-01T13:37:01.917397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's now check how the final model from `TabularAutoML` looks like:","metadata":{}},{"cell_type":"code","source":"print(automl.create_model_str_desc())","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:37:01.922006Z","iopub.execute_input":"2022-08-01T13:37:01.922455Z","iopub.status.idle":"2022-08-01T13:37:01.928670Z","shell.execute_reply.started":"2022-08-01T13:37:01.922421Z","shell.execute_reply":"2022-08-01T13:37:01.927625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ... and check the AUC score for the model based on train target and Out-Of-Fold predictions:","metadata":{}},{"cell_type":"code","source":"print(f'TRAIN out-of-fold score: {roc_auc_score(train_data[TARGET_NAME].values, oof_pred.data[:, 0])}')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:37:01.929748Z","iopub.execute_input":"2022-08-01T13:37:01.930099Z","iopub.status.idle":"2022-08-01T13:37:01.955987Z","shell.execute_reply.started":"2022-08-01T13:37:01.930062Z","shell.execute_reply":"2022-08-01T13:37:01.954300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Prediction for test dataset from `TabularAutoML`\n\nWe now have the trained model called `automl` and it's time to create predictions for the test file:","metadata":{"papermill":{"duration":0.145098,"end_time":"2021-06-22T20:34:32.530768","exception":false,"start_time":"2021-06-22T20:34:32.38567","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n\ntest_pred = automl.predict(test_data)\nprint(f'Prediction for te_data:\\n{test_pred}\\nShape = {test_pred.shape}')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:37:01.958679Z","iopub.execute_input":"2022-08-01T13:37:01.959456Z","iopub.status.idle":"2022-08-01T13:37:02.824783Z","shell.execute_reply.started":"2022-08-01T13:37:01.959393Z","shell.execute_reply":"2022-08-01T13:37:02.822856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[TARGET_NAME] = test_pred.data[:, 0]\nsubmission.to_csv('LightAutoML_group_TabularAutoML.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:37:02.829854Z","iopub.execute_input":"2022-08-01T13:37:02.833110Z","iopub.status.idle":"2022-08-01T13:37:02.909835Z","shell.execute_reply.started":"2022-08-01T13:37:02.833056Z","shell.execute_reply":"2022-08-01T13:37:02.908461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Feature importances calculation \n\nFor feature importances calculation we have 2 different methods in LightAutoML:\n- Fast (`fast`) - this method uses feature importances from feature selector LGBM model inside LightAutoML. It works extremely fast and almost always (almost because of situations, when feature selection is turned off or selector was removed from the final models with all GBM models). no need to use new labelled data.\n- Accurate (`accurate`) - this method calculate *features permutation importances* for the whole LightAutoML model based on the **new labelled data**. It always works but can take a lot of time to finish (depending on the model structure, new labelled dataset size etc.).","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Fast feature importances calculation\nfast_fi = automl.get_feature_scores('fast')\nfast_fi.set_index('Feature')['Importance'].plot.bar(figsize = (30, 10), grid = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T13:37:02.911530Z","iopub.execute_input":"2022-08-01T13:37:02.911943Z","iopub.status.idle":"2022-08-01T13:37:03.362640Z","shell.execute_reply.started":"2022-08-01T13:37:02.911907Z","shell.execute_reply":"2022-08-01T13:37:03.361132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%time\n\n# # Accurate feature importances calculation (Permutation importances) -  can take long time to calculate on bigger datasets\n# accurate_fi = automl.get_feature_scores('accurate', te_data, silent = False)","metadata":{"execution":{"iopub.status.busy":"2022-06-09T09:09:33.425375Z","iopub.execute_input":"2022-06-09T09:09:33.425735Z","iopub.status.idle":"2022-06-09T09:09:40.154496Z","shell.execute_reply.started":"2022-06-09T09:09:33.425688Z","shell.execute_reply":"2022-06-09T09:09:40.153791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# accurate_fi.set_index('Feature')['Importance'].plot.bar(figsize = (30, 10), grid = True)","metadata":{"execution":{"iopub.status.busy":"2022-06-09T09:09:40.158256Z","iopub.execute_input":"2022-06-09T09:09:40.160045Z","iopub.status.idle":"2022-06-09T09:09:40.522398Z","shell.execute_reply.started":"2022-06-09T09:09:40.160003Z","shell.execute_reply":"2022-06-09T09:09:40.521228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Spending more from `TIMEOUT` - `TabularUtilizedAutoML` usage\n\nUsing `TabularAutoML` we spent only ~5 minutes to build the model with setup `TIMEOUT` equal to 30 minutes. To spend (almost) all the `TIMEOUT` we can use `TabularUtilizedAutoML` preset instead of `TabularAutoML`, which has the same API:","metadata":{}},{"cell_type":"code","source":"# utilized_automl = TabularUtilizedAutoML(\n#     task = task, \n#     timeout = TIMEOUT,\n#     cpu_limit = N_THREADS,\n#     reader_params = {'n_jobs': N_THREADS, 'cv': N_FOLDS, 'random_state': RANDOM_STATE},\n# )","metadata":{"execution":{"iopub.status.busy":"2022-08-01T12:50:19.985719Z","iopub.execute_input":"2022-08-01T12:50:19.986164Z","iopub.status.idle":"2022-08-01T12:50:19.992169Z","shell.execute_reply.started":"2022-08-01T12:50:19.986131Z","shell.execute_reply":"2022-08-01T12:50:19.990722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training the model on the full dataset:","metadata":{}},{"cell_type":"code","source":"# %%time \n\n# oof_pred = utilized_automl.fit_predict(train_data, roles = roles, verbose = 1)","metadata":{"scrolled":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-01T12:50:22.973529Z","iopub.execute_input":"2022-08-01T12:50:22.973952Z","iopub.status.idle":"2022-08-01T12:50:22.979039Z","shell.execute_reply.started":"2022-08-01T12:50:22.973921Z","shell.execute_reply":"2022-08-01T12:50:22.977505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Let's now check how the final model from `TabularUtilizedAutoML` looks like:","metadata":{}},{"cell_type":"code","source":"# print(utilized_automl.create_model_str_desc())","metadata":{"execution":{"iopub.status.busy":"2022-08-01T12:50:26.273182Z","iopub.execute_input":"2022-08-01T12:50:26.273576Z","iopub.status.idle":"2022-08-01T12:50:26.278583Z","shell.execute_reply.started":"2022-08-01T12:50:26.273547Z","shell.execute_reply":"2022-08-01T12:50:26.277476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ... and check its AUC score:","metadata":{}},{"cell_type":"code","source":"# print(f'TRAIN out-of-fold utilized score: {roc_auc_score(train_data[TARGET_NAME].values, oof_pred.data[:, 0])}')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T12:50:34.579031Z","iopub.execute_input":"2022-08-01T12:50:34.579624Z","iopub.status.idle":"2022-08-01T12:50:34.586350Z","shell.execute_reply.started":"2022-08-01T12:50:34.579578Z","shell.execute_reply":"2022-08-01T12:50:34.584862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ... and the feature importances:","metadata":{}},{"cell_type":"code","source":"# %%time\n\n# # Fast feature importances calculation\n# fast_fi = utilized_automl.get_feature_scores('fast')\n# fast_fi.set_index('Feature')['Importance'].plot.bar(figsize = (30, 10), grid = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T12:50:41.706919Z","iopub.execute_input":"2022-08-01T12:50:41.707303Z","iopub.status.idle":"2022-08-01T12:50:41.712121Z","shell.execute_reply.started":"2022-08-01T12:50:41.707275Z","shell.execute_reply":"2022-08-01T12:50:41.711133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Predict for test dataset using `TabularUtilizedAutoML` model\n\nWe are also ready to predict for our test competition dataset using `utilized_automl` model and submission file creation:","metadata":{}},{"cell_type":"code","source":"# test_pred = utilized_automl.predict(test_data)\n# print(f'Prediction for te_data:\\n{test_pred}\\nShape = {test_pred.shape}')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T12:50:45.798337Z","iopub.execute_input":"2022-08-01T12:50:45.798808Z","iopub.status.idle":"2022-08-01T12:50:45.804454Z","shell.execute_reply.started":"2022-08-01T12:50:45.798775Z","shell.execute_reply":"2022-08-01T12:50:45.803039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# submission[TARGET_NAME] = test_pred.data[:, 0]\n# submission.to_csv('LightAutoML_TabularUtilizedAutoML.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T12:50:48.709561Z","iopub.execute_input":"2022-08-01T12:50:48.710201Z","iopub.status.idle":"2022-08-01T12:50:48.715639Z","shell.execute_reply.started":"2022-08-01T12:50:48.710155Z","shell.execute_reply":"2022-08-01T12:50:48.714372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Additional materials","metadata":{"papermill":{"duration":0.14221,"end_time":"2021-06-22T20:35:48.782561","exception":false,"start_time":"2021-06-22T20:35:48.640351","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"- [Official LightAutoML github repo](https://github.com/AILab-MLTools/LightAutoML)\n- [LightAutoML documentation](https://lightautoml.readthedocs.io/en/latest)\n- [LightAutoML tutorials](https://github.com/AILab-MLTools/LightAutoML/tree/master/examples/tutorials)\n- LightAutoML course:\n    - [Part 1 - general overview](https://ods.ai/tracks/automl-course-part1) \n    - [Part 2 - LightAutoML specific applications](https://ods.ai/tracks/automl-course-part2)\n    - [Part 3 - LightAutoML customization](https://ods.ai/tracks/automl-course-part3)\n- [OpenDataScience AutoML benchmark leaderboard](https://ods.ai/competitions/automl-benchmark/leaderboard)","metadata":{"papermill":{"duration":0.147943,"end_time":"2021-06-22T20:35:49.074531","exception":false,"start_time":"2021-06-22T20:35:48.926588","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"### If you still like the notebook, do not forget to put upvote for the notebook and the ⭐️ for github repo if you like it using the button below - one click for you, great pleasure for us ☺️","metadata":{}},{"cell_type":"code","source":"s = '<iframe src=\"https://ghbtns.com/github-btn.html?user=sb-ai-lab&repo=LightAutoML&type=star&count=true&size=large\" frameborder=\"0\" scrolling=\"0\" width=\"170\" height=\"30\" title=\"LightAutoML GitHub\"></iframe>'\nHTML(s)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-01T13:37:17.047143Z","iopub.execute_input":"2022-08-01T13:37:17.047763Z","iopub.status.idle":"2022-08-01T13:37:17.057775Z","shell.execute_reply.started":"2022-08-01T13:37:17.047721Z","shell.execute_reply":"2022-08-01T13:37:17.056718Z"},"trusted":true},"execution_count":null,"outputs":[]}]}