{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Predict Bike Sharing Demand with AutoGluon Template","metadata":{}},{"cell_type":"markdown","source":"## Project: Predict Bike Sharing Demand with AutoGluon\nThis notebook is a template with each step that you need to complete for the project.\n\nPlease fill in your code where there are explicit `?` markers in the notebook. You are welcome to add more cells and code as you see fit.\n\nOnce you have completed all the code implementations, please export your notebook as a HTML file so the reviews can view your code. Make sure you have all outputs correctly outputted.\n\n`File-> Export Notebook As... -> Export Notebook as HTML`\n\nThere is a writeup to complete as well after all code implememtation is done. Please answer all questions and attach the necessary tables and charts. You can complete the writeup in either markdown or PDF.\n\nCompleting the code template and writeup template will cover all of the rubric points for this project.\n\nThe rubric contains \"Stand Out Suggestions\" for enhancing the project beyond the minimum requirements. The stand out suggestions are optional. If you decide to pursue the \"stand out suggestions\", you can include the code in this notebook and also discuss the results in the writeup file.","metadata":{}},{"cell_type":"markdown","source":"## Step 1: Create an account with Kaggle","metadata":{}},{"cell_type":"markdown","source":"### Create Kaggle Account and download API key\nBelow is example of steps to get the API username and key. Each student will have their own username and key.","metadata":{}},{"cell_type":"markdown","source":"## Step 2: Download the Kaggle dataset using the kaggle python library","metadata":{}},{"cell_type":"markdown","source":"### Open up Sagemaker Studio and use starter template","metadata":{}},{"cell_type":"markdown","source":"1. Notebook should be using a `ml.t3.medium` instance (2 vCPU + 4 GiB)\n2. Notebook should be using kernal: `Python 3 (MXNet 1.8 Python 3.7 CPU Optimized)`","metadata":{}},{"cell_type":"markdown","source":"### Install packages","metadata":{}},{"cell_type":"code","source":"# !pip install -U pip\n# !pip install -U setuptools wheel\n!pip install -U \"mxnet<2.0.0\" bokeh==2.0.1\n!pip install autogluon --no-cache-dir\n# Without --no-cache-dir, smaller aws instances may have trouble installing","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:22:57.980243Z","iopub.execute_input":"2022-08-08T11:22:57.980682Z","iopub.status.idle":"2022-08-08T11:24:14.042218Z","shell.execute_reply.started":"2022-08-08T11:22:57.980647Z","shell.execute_reply":"2022-08-08T11:24:14.040995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Setup Kaggle API Key","metadata":{}},{"cell_type":"code","source":"# # create the .kaggle directory and an empty kaggle.json file\n# !mkdir -p /root/.kaggle\n# !touch /root/.kaggle/kaggle.jsona\n# !chmod 600 /root/.kaggle/kaggle.json","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Fill in your user name and key from creating the kaggle account and API token file\n# import json\n# kaggle_username = \"FILL_IN_USERNAME\"\n# kaggle_key = \"FILL_IN_KEY\"\n\n# # Save API token the kaggle.json file\n# with open(\"/root/.kaggle/kaggle.json\", \"w\") as f:\n#     f.write(json.dumps({\"username\": kaggle_username, \"key\": kaggle_key}))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Download and explore dataset","metadata":{}},{"cell_type":"code","source":"# !pip install kaggle","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Download the dataset, it will be in a .zip file so you'll need to unzip it as well.\n# !kaggle competitions download -c bike-sharing-demand\n# If you already downloaded it you can use the -o command to overwrite the file\n# !unzip -o bike-sharing-demand.zip","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom autogluon.tabular import TabularPredictor","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:14.044914Z","iopub.execute_input":"2022-08-08T11:24:14.045249Z","iopub.status.idle":"2022-08-08T11:24:15.091789Z","shell.execute_reply.started":"2022-08-08T11:24:14.045217Z","shell.execute_reply":"2022-08-08T11:24:15.090676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the train dataset in pandas by reading the csv\n# Set the parsing of the datetime column so you can use some of the `dt` features in pandas later\ntrain = pd.read_csv('../input/bike-sharing-demand/train.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.093092Z","iopub.execute_input":"2022-08-08T11:24:15.093428Z","iopub.status.idle":"2022-08-08T11:24:15.153080Z","shell.execute_reply.started":"2022-08-08T11:24:15.093398Z","shell.execute_reply":"2022-08-08T11:24:15.152280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['datetime'] = pd.to_datetime(train['datetime'])","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.154983Z","iopub.execute_input":"2022-08-08T11:24:15.155692Z","iopub.status.idle":"2022-08-08T11:24:15.166622Z","shell.execute_reply.started":"2022-08-08T11:24:15.155661Z","shell.execute_reply":"2022-08-08T11:24:15.165768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Simple output of the train dataset to view some of the min/max/varition of the dataset features.","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.168067Z","iopub.execute_input":"2022-08-08T11:24:15.168605Z","iopub.status.idle":"2022-08-08T11:24:15.174822Z","shell.execute_reply.started":"2022-08-08T11:24:15.168573Z","shell.execute_reply":"2022-08-08T11:24:15.173926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.177768Z","iopub.execute_input":"2022-08-08T11:24:15.178405Z","iopub.status.idle":"2022-08-08T11:24:15.232145Z","shell.execute_reply.started":"2022-08-08T11:24:15.178365Z","shell.execute_reply":"2022-08-08T11:24:15.230954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the test pandas dataframe in pandas by reading the csv, remember to parse the datetime!\ntest = pd.read_csv('../input/bike-sharing-demand/test.csv')\ntest['datetime'] = pd.to_datetime(test['datetime'])\ntest.head()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.234599Z","iopub.execute_input":"2022-08-08T11:24:15.234944Z","iopub.status.idle":"2022-08-08T11:24:15.269646Z","shell.execute_reply.started":"2022-08-08T11:24:15.234914Z","shell.execute_reply":"2022-08-08T11:24:15.268826Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same thing as train and test dataset\nsubmission = pd.read_csv('../input/bike-sharing-demand/sampleSubmission.csv')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.271170Z","iopub.execute_input":"2022-08-08T11:24:15.271846Z","iopub.status.idle":"2022-08-08T11:24:15.300814Z","shell.execute_reply.started":"2022-08-08T11:24:15.271805Z","shell.execute_reply":"2022-08-08T11:24:15.299845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 3: Train a model using AutoGluon’s Tabular Prediction","metadata":{}},{"cell_type":"markdown","source":"Requirements:\n* We are prediting `count`, so it is the label we are setting.\n* Ignore `casual` and `registered` columns as they are also not present in the test dataset. \n* Use the `root_mean_squared_error` as the metric to use for evaluation.\n* Set a time limit of 10 minutes (600 seconds).\n* Use the preset `best_quality` to focus on creating the best model.","metadata":{}},{"cell_type":"code","source":"train.drop(columns = ['casual', 'registered'], inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.302282Z","iopub.execute_input":"2022-08-08T11:24:15.304123Z","iopub.status.idle":"2022-08-08T11:24:15.309619Z","shell.execute_reply.started":"2022-08-08T11:24:15.304089Z","shell.execute_reply":"2022-08-08T11:24:15.308762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.313464Z","iopub.execute_input":"2022-08-08T11:24:15.313782Z","iopub.status.idle":"2022-08-08T11:24:15.331147Z","shell.execute_reply.started":"2022-08-08T11:24:15.313752Z","shell.execute_reply":"2022-08-08T11:24:15.329942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor = TabularPredictor(label = 'count', eval_metric = 'root_mean_squared_error').fit(train_data = train, presets = 'best_quality', time_limit = 600)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:24:15.332780Z","iopub.execute_input":"2022-08-08T11:24:15.333260Z","iopub.status.idle":"2022-08-08T11:34:22.224573Z","shell.execute_reply.started":"2022-08-08T11:24:15.333208Z","shell.execute_reply":"2022-08-08T11:34:22.223289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Review AutoGluon's training run with ranking of models that did the best.","metadata":{}},{"cell_type":"code","source":"predictor.fit_summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:34:22.226652Z","iopub.execute_input":"2022-08-08T11:34:22.227431Z","iopub.status.idle":"2022-08-08T11:34:22.598273Z","shell.execute_reply.started":"2022-08-08T11:34:22.227384Z","shell.execute_reply":"2022-08-08T11:34:22.597224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create predictions from test dataset","metadata":{}},{"cell_type":"code","source":"predictions = predictor.predict(test)\npredictions.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:34:22.600047Z","iopub.execute_input":"2022-08-08T11:34:22.601177Z","iopub.status.idle":"2022-08-08T11:34:42.807938Z","shell.execute_reply.started":"2022-08-08T11:34:22.601135Z","shell.execute_reply":"2022-08-08T11:34:42.806673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### NOTE: Kaggle will reject the submission if we don't set everything to be > 0.","metadata":{}},{"cell_type":"code","source":"# Describe the `predictions` series to see if there are any negative values\npredictions.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:34:42.809428Z","iopub.execute_input":"2022-08-08T11:34:42.810632Z","iopub.status.idle":"2022-08-08T11:34:42.823695Z","shell.execute_reply.started":"2022-08-08T11:34:42.810582Z","shell.execute_reply":"2022-08-08T11:34:42.822280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# How many negative values do we have?\n(predictions < 0).sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:35:42.508785Z","iopub.execute_input":"2022-08-08T11:35:42.509186Z","iopub.status.idle":"2022-08-08T11:35:42.516536Z","shell.execute_reply.started":"2022-08-08T11:35:42.509155Z","shell.execute_reply":"2022-08-08T11:35:42.515430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set them to zero\npredictions[predictions < 0] = 0","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:35:39.229332Z","iopub.execute_input":"2022-08-08T11:35:39.230570Z","iopub.status.idle":"2022-08-08T11:35:39.239311Z","shell.execute_reply.started":"2022-08-08T11:35:39.230519Z","shell.execute_reply":"2022-08-08T11:35:39.238224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Set predictions to submission dataframe, save, and submit","metadata":{}},{"cell_type":"code","source":"submission[\"count\"] = predictions\nsubmission.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:35:59.067324Z","iopub.execute_input":"2022-08-08T11:35:59.067767Z","iopub.status.idle":"2022-08-08T11:35:59.093638Z","shell.execute_reply.started":"2022-08-08T11:35:59.067730Z","shell.execute_reply":"2022-08-08T11:35:59.092669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install kaggle","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:36:14.804954Z","iopub.execute_input":"2022-08-08T11:36:14.806007Z","iopub.status.idle":"2022-08-08T11:36:25.941809Z","shell.execute_reply.started":"2022-08-08T11:36:14.805963Z","shell.execute_reply":"2022-08-08T11:36:25.940670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!kaggle competitions submit -c bike-sharing-demand -f submission.csv -m \"first raw submission\"","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:36:33.876939Z","iopub.execute_input":"2022-08-08T11:36:33.877379Z","iopub.status.idle":"2022-08-08T11:36:35.347406Z","shell.execute_reply.started":"2022-08-08T11:36:33.877335Z","shell.execute_reply":"2022-08-08T11:36:35.346277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### View submission via the command line or in the web browser under the competition's page - `My Submissions`","metadata":{}},{"cell_type":"code","source":"!kaggle competitions submissions -c bike-sharing-demand | tail -n +1 | head -n 6","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Initial score of `?`","metadata":{}},{"cell_type":"markdown","source":"## Step 4: Exploratory Data Analysis and Creating an additional feature\n* Any additional feature will do, but a great suggestion would be to separate out the datetime into hour, day, or month parts.","metadata":{}},{"cell_type":"code","source":"# Create a histogram of all features to show the distribution of each one relative to the data. This is part of the exploritory data analysis\ntrain.plot(kind = 'hist');","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:45:59.735563Z","iopub.execute_input":"2022-08-08T11:45:59.736017Z","iopub.status.idle":"2022-08-08T11:46:00.226991Z","shell.execute_reply.started":"2022-08-08T11:45:59.735980Z","shell.execute_reply":"2022-08-08T11:46:00.225779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create a new feature\ntrain['year'] = train.datetime.dt.year\ntest['year'] = test.datetime.dt.year\n\ntrain['month'] = train.datetime.dt.month\ntest['month'] = test.datetime.dt.month\n\ntrain['hour'] = train.datetime.dt.hour\ntest['hour'] = test.datetime.dt.hour","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:51:16.275383Z","iopub.execute_input":"2022-08-08T11:51:16.275788Z","iopub.status.idle":"2022-08-08T11:51:16.292572Z","shell.execute_reply.started":"2022-08-08T11:51:16.275750Z","shell.execute_reply":"2022-08-08T11:51:16.291735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Make category types for these so models know they are not just numbers\n* AutoGluon originally sees these as ints, but in reality they are int representations of a category.\n* Setting the dtype to category will classify these as categories in AutoGluon.","metadata":{}},{"cell_type":"code","source":"train[\"season\"] = train[\"season\"].astype('category')\ntrain[\"weather\"] = train[\"weather\"].astype('category')\ntest[\"season\"] = test[\"season\"].astype('category')\ntest[\"weather\"] = test[\"weather\"].astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:49:41.037649Z","iopub.execute_input":"2022-08-08T11:49:41.038084Z","iopub.status.idle":"2022-08-08T11:49:41.050670Z","shell.execute_reply.started":"2022-08-08T11:49:41.038047Z","shell.execute_reply":"2022-08-08T11:49:41.049591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# View are new feature\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:51:25.159116Z","iopub.execute_input":"2022-08-08T11:51:25.159835Z","iopub.status.idle":"2022-08-08T11:51:25.179178Z","shell.execute_reply.started":"2022-08-08T11:51:25.159789Z","shell.execute_reply":"2022-08-08T11:51:25.178136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# View histogram of all features again now with the hour feature\ntrain.plot(kind = 'hist', bins=30, figsize=(15, 10))","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:54:29.270350Z","iopub.execute_input":"2022-08-08T11:54:29.270897Z","iopub.status.idle":"2022-08-08T11:54:30.360214Z","shell.execute_reply.started":"2022-08-08T11:54:29.270835Z","shell.execute_reply":"2022-08-08T11:54:30.359083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 5: Rerun the model with the same settings as before, just with more features","metadata":{}},{"cell_type":"code","source":"predictor_new_features = TabularPredictor(label = 'count', eval_metric = 'root_mean_squared_error').fit(train_data = train, presets = 'best_quality', time_limit = 600)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T11:55:48.917597Z","iopub.execute_input":"2022-08-08T11:55:48.918037Z","iopub.status.idle":"2022-08-08T12:05:56.040384Z","shell.execute_reply.started":"2022-08-08T11:55:48.918001Z","shell.execute_reply":"2022-08-08T12:05:56.035293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor_new_features.fit_summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:05:56.043839Z","iopub.execute_input":"2022-08-08T12:05:56.045104Z","iopub.status.idle":"2022-08-08T12:05:56.175112Z","shell.execute_reply.started":"2022-08-08T12:05:56.045049Z","shell.execute_reply":"2022-08-08T12:05:56.173695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Remember to set all negative values to zero\npredictions2 = predictor_new_features.predict(test)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:10:18.106993Z","iopub.execute_input":"2022-08-08T12:10:18.108315Z","iopub.status.idle":"2022-08-08T12:10:44.047175Z","shell.execute_reply.started":"2022-08-08T12:10:18.108259Z","shell.execute_reply":"2022-08-08T12:10:44.045804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(predictions2 < 0).sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:10:47.854244Z","iopub.execute_input":"2022-08-08T12:10:47.854665Z","iopub.status.idle":"2022-08-08T12:10:47.863361Z","shell.execute_reply.started":"2022-08-08T12:10:47.854630Z","shell.execute_reply":"2022-08-08T12:10:47.861869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same thing as train and test dataset\nsubmission_new_features = pd.read_csv('../input/bike-sharing-demand/sampleSubmission.csv')\nsubmission_new_features.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:13:15.940343Z","iopub.execute_input":"2022-08-08T12:13:15.941401Z","iopub.status.idle":"2022-08-08T12:13:15.968377Z","shell.execute_reply.started":"2022-08-08T12:13:15.941354Z","shell.execute_reply":"2022-08-08T12:13:15.967124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same submitting predictions\nsubmission_new_features[\"count\"] = predictions2\nsubmission_new_features.to_csv(\"submission_new_features.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:14:11.953248Z","iopub.execute_input":"2022-08-08T12:14:11.953700Z","iopub.status.idle":"2022-08-08T12:14:11.975779Z","shell.execute_reply.started":"2022-08-08T12:14:11.953661Z","shell.execute_reply":"2022-08-08T12:14:11.974637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!kaggle competitions submit -c bike-sharing-demand -f submission_new_features.csv -m \"new features\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!kaggle competitions submissions -c bike-sharing-demand | tail -n +1 | head -n 6","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### New Score of 0.449","metadata":{}},{"cell_type":"markdown","source":"## Step 6: Hyper parameter optimization\n* There are many options for hyper parameter optimization.\n* Options are to change the AutoGluon higher level parameters or the individual model hyperparameters.\n* The hyperparameters of the models themselves that are in AutoGluon. Those need the `hyperparameter` and `hyperparameter_tune_kwargs` arguments.","metadata":{}},{"cell_type":"code","source":"predictor_new_hpo = TabularPredictor(label = 'count', eval_metric = 'root_mean_squared_error').fit(train_data = train, presets = 'best_quality', time_limit = 600, num_bag_folds = 5, num_stack_levels = 3, num_bag_sets = 10)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:56:30.747925Z","iopub.execute_input":"2022-08-08T12:56:30.748351Z","iopub.status.idle":"2022-08-08T12:56:30.784768Z","shell.execute_reply.started":"2022-08-08T12:56:30.748317Z","shell.execute_reply":"2022-08-08T12:56:30.783193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor_new_hpo.fit_summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:42:15.161353Z","iopub.execute_input":"2022-08-08T12:42:15.161942Z","iopub.status.idle":"2022-08-08T12:42:15.238679Z","shell.execute_reply.started":"2022-08-08T12:42:15.161894Z","shell.execute_reply":"2022-08-08T12:42:15.237577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions_new_hpo = predictor_new_hpo.predict(test)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:46:19.340675Z","iopub.execute_input":"2022-08-08T12:46:19.342026Z","iopub.status.idle":"2022-08-08T12:46:35.215008Z","shell.execute_reply.started":"2022-08-08T12:46:19.341969Z","shell.execute_reply":"2022-08-08T12:46:35.214018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Remember to set all negative values to zero\n(predictions_new_hpo < 0).sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:46:49.919949Z","iopub.execute_input":"2022-08-08T12:46:49.920416Z","iopub.status.idle":"2022-08-08T12:46:49.928733Z","shell.execute_reply.started":"2022-08-08T12:46:49.920377Z","shell.execute_reply":"2022-08-08T12:46:49.927474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same thing as train and test dataset\nsubmission_new_hpo = pd.read_csv('../input/bike-sharing-demand/sampleSubmission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:47:00.724700Z","iopub.execute_input":"2022-08-08T12:47:00.725696Z","iopub.status.idle":"2022-08-08T12:47:00.743754Z","shell.execute_reply.started":"2022-08-08T12:47:00.725657Z","shell.execute_reply":"2022-08-08T12:47:00.742667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Same submitting predictions\nsubmission_new_hpo[\"count\"] = predictions_new_hpo\nsubmission_new_hpo.to_csv(\"submission_new_hpo.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-08T12:47:04.971710Z","iopub.execute_input":"2022-08-08T12:47:04.972166Z","iopub.status.idle":"2022-08-08T12:47:04.994458Z","shell.execute_reply.started":"2022-08-08T12:47:04.972128Z","shell.execute_reply":"2022-08-08T12:47:04.993596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !kaggle competitions submit -c bike-sharing-demand -f submission_new_hpo.csv -m \"new features with hyperparameters\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !kaggle competitions submissions -c bike-sharing-demand | tail -n +1 | head -n 6","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### New Score of `?`","metadata":{}},{"cell_type":"markdown","source":"## Step 7: Write a Report\n### Refer to the markdown file for the full report\n### Creating plots and table for report","metadata":{}},{"cell_type":"code","source":"# Taking the top model score from each training run and creating a line plot to show improvement\n# You can create these in the notebook and save them to PNG or use some other tool (e.g. google sheets, excel)\nfig = pd.DataFrame(\n    {\n        \"model\": [\"initial\", \"add_features\", \"hpo\"],\n        \"score\": [-50.21, -31.79, -32.82]\n    }\n).plot(x=\"model\", y=\"score\", figsize=(8, 6)).get_figure()\nfig.savefig('model_train_score.png')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Take the 3 kaggle scores and creating a line plot to show improvement\nfig = pd.DataFrame(\n    {\n        \"test_eval\": [\"initial\", \"add_features\", \"hpo\"],\n        \"score\": [1.845, 0.449, 0.455]\n    }\n).plot(x=\"test_eval\", y=\"score\", figsize=(8, 6)).get_figure()\nfig.savefig('model_test_score.png')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Hyperparameter table","metadata":{}},{"cell_type":"code","source":"# The 3 hyperparameters we tuned with the kaggle score as the result\npd.DataFrame({\n    \"model\": [\"initial\", \"add_features\", \"hpo\"],\n    \"hpo1\": [8, 8, 5],\n    \"hpo2\": [3, 3, 3],\n    \"hpo3\": [20, 20, 10],\n    \"score\": [1.84, 0.44, 0.52]\n})","metadata":{},"execution_count":null,"outputs":[]}]}