{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Bike Sharing Demand","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\npath = '/kaggle/input/bike-sharing-demand/'\n\ntrain = pd.read_csv(path + 'train.csv')\ntest = pd.read_csv(path + 'test.csv')\nsubmission = pd.read_csv(path + 'sampleSubmission.csv')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-26T01:02:52.569161Z","iopub.execute_input":"2022-07-26T01:02:52.569561Z","iopub.status.idle":"2022-07-26T01:02:52.655715Z","shell.execute_reply.started":"2022-07-26T01:02:52.569531Z","shell.execute_reply":"2022-07-26T01:02:52.654770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Fields\n\n- datetime - hourly date + timestamp  \n- season -  1 = spring, 2 = summer, 3 = fall, 4 = winter \n- holiday - whether the day is considered a holiday\n- workingday - whether the day is neither a weekend nor holiday\n- weather\n  - 1: Clear, Few clouds, Partly cloudy, Partly cloudy \n  - 2: Mist + Cloudy, Mist + Broken clouds, Mist + Few clouds, Mist \n  - 3: Light Snow, Light Rain + Thunderstorm + Scattered clouds, Light Rain + Scattered clouds \n  - 4: Heavy Rain + Ice Pallets + Thunderstorm + Mist, Snow + Fog \n- temp - temperature in Celsius\n- atemp - \"feels like\" temperature in Celsius\n- humidity - relative humidity\n- windspeed - wind speed\n- casual - number of non-registered user rentals initiated\n- registered - number of registered user rentals initiated\n- count - number of total rentals","metadata":{}},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:52.679099Z","iopub.execute_input":"2022-07-26T01:02:52.679572Z","iopub.status.idle":"2022-07-26T01:02:52.718687Z","shell.execute_reply.started":"2022-07-26T01:02:52.679536Z","shell.execute_reply":"2022-07-26T01:02:52.717692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:52.769489Z","iopub.execute_input":"2022-07-26T01:02:52.769867Z","iopub.status.idle":"2022-07-26T01:02:52.788107Z","shell.execute_reply.started":"2022-07-26T01:02:52.769836Z","shell.execute_reply":"2022-07-26T01:02:52.787056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- I should remove `casual`, `registered` features because there are no column names in test file.","metadata":{}},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:52.870943Z","iopub.execute_input":"2022-07-26T01:02:52.871371Z","iopub.status.idle":"2022-07-26T01:02:52.882688Z","shell.execute_reply.started":"2022-07-26T01:02:52.871339Z","shell.execute_reply":"2022-07-26T01:02:52.881262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- I should remove datetime feature because it's not helpful for prediction.","metadata":{}},{"cell_type":"markdown","source":"## Getting Information of DateFrame","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:52.885503Z","iopub.execute_input":"2022-07-26T01:02:52.886692Z","iopub.status.idle":"2022-07-26T01:02:52.916438Z","shell.execute_reply.started":"2022-07-26T01:02:52.886640Z","shell.execute_reply":"2022-07-26T01:02:52.915519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There are no missing values because the number of data matches Non-Null Count.","metadata":{}},{"cell_type":"markdown","source":"## Feature Engineering","metadata":{}},{"cell_type":"code","source":"train['datetime'] = pd.to_datetime(train['datetime']) #type casting\ntrain.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:52.917587Z","iopub.execute_input":"2022-07-26T01:02:52.918067Z","iopub.status.idle":"2022-07-26T01:02:52.941449Z","shell.execute_reply.started":"2022-07-26T01:02:52.918037Z","shell.execute_reply":"2022-07-26T01:02:52.940624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['year'] = train['datetime'].dt.year\ntrain['month'] = train['datetime'].dt.month\ntrain['day'] = train['datetime'].dt.day\ntrain['hour'] = train['datetime'].dt.hour\ntrain['minute'] = train['datetime'].dt.minute\ntrain['second'] = train['datetime'].dt.second\ntrain['weekday'] = train['datetime'].dt.day_name()\ntrain['season'] = train['season'].map({1: 'spring', 2: 'summer', 3: 'fall', 4: 'winter'})\ntrain['weather'] = train['weather'].map({1: 'Clear', 2: 'Mist, Few clouds', 3: 'Light Snow, Rain, Thunderstorm', 4: 'Heavy Rain, Thunderstorm, Snow, Fog'})\n\ntrain","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:52.975134Z","iopub.execute_input":"2022-07-26T01:02:52.975837Z","iopub.status.idle":"2022-07-26T01:02:53.045681Z","shell.execute_reply.started":"2022-07-26T01:02:52.975799Z","shell.execute_reply":"2022-07-26T01:02:53.044806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You should not use string value for training model. The only reason for converting weekday from integer to string would be a readibility. \n","metadata":{}},{"cell_type":"markdown","source":"## Data Visualization","metadata":{"execution":{"iopub.status.busy":"2022-07-24T13:30:51.237028Z","iopub.execute_input":"2022-07-24T13:30:51.238050Z","iopub.status.idle":"2022-07-24T13:30:51.243449Z","shell.execute_reply.started":"2022-07-24T13:30:51.237955Z","shell.execute_reply":"2022-07-24T13:30:51.242057Z"}}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:53.047896Z","iopub.execute_input":"2022-07-26T01:02:53.049216Z","iopub.status.idle":"2022-07-26T01:02:53.723776Z","shell.execute_reply.started":"2022-07-26T01:02:53.049155Z","shell.execute_reply":"2022-07-26T01:02:53.722964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Distribution plot","metadata":{}},{"cell_type":"code","source":"sns.displot(train['count'])","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:53.725412Z","iopub.execute_input":"2022-07-26T01:02:53.726233Z","iopub.status.idle":"2022-07-26T01:02:54.151170Z","shell.execute_reply.started":"2022-07-26T01:02:53.726183Z","shell.execute_reply":"2022-07-26T01:02:54.150083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.displot(np.log(train['count']))","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:54.152597Z","iopub.execute_input":"2022-07-26T01:02:54.153077Z","iopub.status.idle":"2022-07-26T01:02:54.525820Z","shell.execute_reply.started":"2022-07-26T01:02:54.153045Z","shell.execute_reply":"2022-07-26T01:02:54.525006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'm going to use `log(count)` as a target value because this is close to normal distribution.","metadata":{}},{"cell_type":"markdown","source":"## Bar graphs","metadata":{}},{"cell_type":"code","source":"mpl.rc('font', size = 14)\nmpl.rc('axes', titlesize = 15)\nfigure, axes = plt.subplots(nrows = 3, ncols = 2)\nplt.tight_layout()\nfigure.set_size_inches(10, 9)\n\naxes[0, 0].set(title = 'Rental amounts by year')\nsns.barplot(x = 'year', y = 'count', data = train, ax = axes[0, 0])\n\naxes[0, 1].set(title = 'Rental amounts by month')\nsns.barplot(x = 'month', y = 'count', data = train, ax = axes[0, 1])\n\naxes[1, 0].set(title = 'Rental amounts by day')\naxes[1, 0].tick_params(axis = 'x', labelrotation = 90)\nsns.barplot(x = 'day', y = 'count', data = train, ax = axes[1, 0])\n\naxes[1, 1].set(title = 'Rental amounts by hour')\naxes[1, 1].tick_params(axis = 'x', labelrotation = 90)\nsns.barplot(x = 'hour', y = 'count', data = train, ax = axes[1, 1])\n\naxes[2, 0].set(title = 'Rental amounts by minute')\nsns.barplot(x = 'minute', y = 'count', data = train, ax = axes[2, 0])\n\naxes[2, 1].set(title = 'Rental amounts by second')\nsns.barplot(x = 'second', y = 'count', data = train, ax = axes[2, 1])","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:54.529925Z","iopub.execute_input":"2022-07-26T01:02:54.530772Z","iopub.status.idle":"2022-07-26T01:02:58.241011Z","shell.execute_reply.started":"2022-07-26T01:02:54.530735Z","shell.execute_reply":"2022-07-26T01:02:58.239833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I'm goint to drop day, minute and second features which are useless.","metadata":{}},{"cell_type":"markdown","source":"## Box plot","metadata":{}},{"cell_type":"code","source":"figure, axes = plt.subplots(nrows = 2, ncols = 2)\nplt.tight_layout()\nfigure.set_size_inches(10, 10)\n\nsns.boxplot(x = 'season', y = 'count', data = train, ax = axes[0, 0])\nsns.boxplot(x = 'weather', y = 'count', data = train, ax = axes[0, 1])\nsns.boxplot(x = 'holiday', y = 'count', data = train, ax = axes[1, 0])\nsns.boxplot(x = 'workingday', y = 'count', data = train, ax = axes[1, 1])\n\naxes[0, 0].set(title = 'Box Plot On Count Across Season')\naxes[0, 1].set(title = 'Box Plot On Count Across Weather')\naxes[1, 0].set(title = 'Box Plot On Count Across Holiday')\naxes[1, 1].set(title = 'Box Plot On Count Across Working Day')\n\naxes[0, 1].tick_params(axis = 'x', labelrotation=10)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:58.242719Z","iopub.execute_input":"2022-07-26T01:02:58.243448Z","iopub.status.idle":"2022-07-26T01:02:58.927348Z","shell.execute_reply.started":"2022-07-26T01:02:58.243403Z","shell.execute_reply":"2022-07-26T01:02:58.926137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Point Plot","metadata":{}},{"cell_type":"code","source":"mpl.rc('font', size = 11)\nfigure, axes = plt.subplots(nrows = 5)\nfigure.set_size_inches(12, 18)\n\nsns.pointplot(x = 'hour', y = 'count', data = train, hue = 'workingday', ax=axes[0])\nsns.pointplot(x = 'hour', y = 'count', data = train, hue = 'holiday', ax=axes[1])\nsns.pointplot(x = 'hour', y = 'count', data = train, hue = 'weekday', ax=axes[2])\nsns.pointplot(x = 'hour', y = 'count', data = train, hue = 'season', ax=axes[3])\nsns.pointplot(x = 'hour', y = 'count', data = train, hue = 'weather', ax=axes[4])","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:02:58.928769Z","iopub.execute_input":"2022-07-26T01:02:58.929893Z","iopub.status.idle":"2022-07-26T01:03:16.056931Z","shell.execute_reply.started":"2022-07-26T01:02:58.929842Z","shell.execute_reply":"2022-07-26T01:03:16.055702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is a outlier in the last graph(Heavy Rain, Thunderstorm, Snow, Fog). You might consider to remove it.","metadata":{}},{"cell_type":"markdown","source":"## Scatter plot","metadata":{}},{"cell_type":"code","source":"mpl.rc('font', size = 15)\nfigure, axes = plt.subplots(nrows = 2, ncols = 2)\nplt.tight_layout()\nfigure.set_size_inches(7, 6)\n\nsns.regplot(x = 'temp', y = 'count', data = train, ax = axes[0, 0], scatter_kws = {'alpha':0.2}, line_kws = {'color':'blue'})\nsns.regplot(x = 'atemp', y = 'count', data = train, ax = axes[0, 1], scatter_kws = {'alpha':0.2}, line_kws = {'color':'blue'})\nsns.regplot(x = 'windspeed', y = 'count', data = train, ax = axes[1, 0], scatter_kws = {'alpha':0.2}, line_kws = {'color':'blue'})\nsns.regplot(x = 'humidity', y = 'count', data = train, ax = axes[1, 1], scatter_kws = {'alpha':0.2}, line_kws = {'color':'blue'})","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:16.058762Z","iopub.execute_input":"2022-07-26T01:03:16.059887Z","iopub.status.idle":"2022-07-26T01:03:19.455092Z","shell.execute_reply.started":"2022-07-26T01:03:16.059842Z","shell.execute_reply":"2022-07-26T01:03:19.453961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are many objects which windspeed is zero. It might not be actual data so I'm going to remove these missing values.","metadata":{}},{"cell_type":"markdown","source":"## Heatmap","metadata":{}},{"cell_type":"code","source":"corrMat = train[['temp', 'atemp', 'humidity', 'windspeed', 'count']].corr()\ncorrMat","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.457099Z","iopub.execute_input":"2022-07-26T01:03:19.457862Z","iopub.status.idle":"2022-07-26T01:03:19.479719Z","shell.execute_reply.started":"2022-07-26T01:03:19.457818Z","shell.execute_reply":"2022-07-26T01:03:19.478450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`corr()` computes pairwise correlation of columns, excluding NA/null values.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(10, 10)\nsns.heatmap(corrMat, annot = True)\nax.set(title = 'Heatmap of Numerical Data')","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.481716Z","iopub.execute_input":"2022-07-26T01:03:19.482180Z","iopub.status.idle":"2022-07-26T01:03:19.821462Z","shell.execute_reply.started":"2022-07-26T01:03:19.482136Z","shell.execute_reply":"2022-07-26T01:03:19.818150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline Model\n\nA baseline model is essentially a simple model that acts as a reference in a machine learning project. Its main function is to contextualize the results of trained models.\nhttps://towardsdatascience.com/baseline-models-your-guide-for-model-building-1ec3aa244b8d","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(path + 'train.csv')\ntest = pd.read_csv(path + 'test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.823598Z","iopub.execute_input":"2022-07-26T01:03:19.824452Z","iopub.status.idle":"2022-07-26T01:03:19.862171Z","shell.execute_reply.started":"2022-07-26T01:03:19.824396Z","shell.execute_reply":"2022-07-26T01:03:19.861259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Engineering","metadata":{}},{"cell_type":"code","source":"train = train[train['weather'] != 4]","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.863842Z","iopub.execute_input":"2022-07-26T01:03:19.864774Z","iopub.status.idle":"2022-07-26T01:03:19.873486Z","shell.execute_reply.started":"2022-07-26T01:03:19.864723Z","shell.execute_reply":"2022-07-26T01:03:19.872281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Merging training and test data","metadata":{}},{"cell_type":"code","source":"all_data = pd.concat([train, test], ignore_index = True)\nall_data","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.875000Z","iopub.execute_input":"2022-07-26T01:03:19.875381Z","iopub.status.idle":"2022-07-26T01:03:19.906887Z","shell.execute_reply.started":"2022-07-26T01:03:19.875349Z","shell.execute_reply":"2022-07-26T01:03:19.905717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Adding Features","metadata":{}},{"cell_type":"code","source":"all_data['datetime'] = pd.to_datetime(all_data['datetime']) #type casting\n\nall_data['year'] = all_data['datetime'].dt.year\nall_data['month'] = all_data['datetime'].dt.month\nall_data['hour'] = all_data['datetime'].dt.hour\nall_data['weekday'] = all_data['datetime'].dt.weekday\n\nall_data","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.912119Z","iopub.execute_input":"2022-07-26T01:03:19.912861Z","iopub.status.idle":"2022-07-26T01:03:19.959631Z","shell.execute_reply.started":"2022-07-26T01:03:19.912811Z","shell.execute_reply":"2022-07-26T01:03:19.958473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Dropping Features","metadata":{}},{"cell_type":"code","source":"all_data = all_data.drop(['datetime', 'casual', 'registered', 'windspeed', 'month'], axis = 1)\nall_data","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.961216Z","iopub.execute_input":"2022-07-26T01:03:19.962246Z","iopub.status.idle":"2022-07-26T01:03:19.987791Z","shell.execute_reply.started":"2022-07-26T01:03:19.962210Z","shell.execute_reply":"2022-07-26T01:03:19.986734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Seperating training and test data","metadata":{}},{"cell_type":"code","source":"X_train = all_data[~pd.isnull(all_data['count'])]\nX_test = all_data[pd.isnull(all_data['count'])]\n\nX_train = X_train.drop(['count'], axis = 1)\nX_test = X_test.drop(['count'], axis = 1)\n\ny = train['count']","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:19.989429Z","iopub.execute_input":"2022-07-26T01:03:19.989757Z","iopub.status.idle":"2022-07-26T01:03:20.001738Z","shell.execute_reply.started":"2022-07-26T01:03:19.989727Z","shell.execute_reply":"2022-07-26T01:03:20.000396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Defining RMSLE function","metadata":{}},{"cell_type":"code","source":"def rmsle(y_true, y_prediction, isExponentialConversion = True):\n    if isExponentialConversion:\n        y_true = np.exp(y_true)\n        y_prediction = np.exp(y_prediction)\n        \n    log_true = np.nan_to_num(np.log(y_true + 1))\n    log_prediction = np.nan_to_num(np.log(y_prediction + 1))\n    \n    return np.sqrt(np.mean((log_true - log_prediction) ** 2))","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:20.003217Z","iopub.execute_input":"2022-07-26T01:03:20.004110Z","iopub.status.idle":"2022-07-26T01:03:20.012486Z","shell.execute_reply.started":"2022-07-26T01:03:20.004075Z","shell.execute_reply":"2022-07-26T01:03:20.011707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training Model","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression\n\nlinear_regression_model = LinearRegression()\nlog_y = np.log(y)\nlinear_regression_model.fit(X_train, log_y)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:20.014268Z","iopub.execute_input":"2022-07-26T01:03:20.015055Z","iopub.status.idle":"2022-07-26T01:03:20.242449Z","shell.execute_reply.started":"2022-07-26T01:03:20.015010Z","shell.execute_reply":"2022-07-26T01:03:20.240818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Evaluating Model","metadata":{}},{"cell_type":"code","source":"predictions = linear_regression_model.predict(X_train)\n\nprint(f'RMSLE: {rmsle(log_y, predictions):.4f}')","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:20.244810Z","iopub.execute_input":"2022-07-26T01:03:20.245421Z","iopub.status.idle":"2022-07-26T01:03:20.260587Z","shell.execute_reply.started":"2022-07-26T01:03:20.245372Z","shell.execute_reply":"2022-07-26T01:03:20.259230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission","metadata":{}},{"cell_type":"code","source":"linear_regression_predictions = linear_regression_model.predict(X_test)\n\nsubmission['count'] = np.exp(linear_regression_predictions)\nsubmission.to_csv('submission.csv', index = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-26T01:03:20.262541Z","iopub.execute_input":"2022-07-26T01:03:20.263749Z","iopub.status.idle":"2022-07-26T01:03:20.363382Z","shell.execute_reply.started":"2022-07-26T01:03:20.263695Z","shell.execute_reply":"2022-07-26T01:03:20.361968Z"},"trusted":true},"execution_count":null,"outputs":[]}]}