{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **About Dataset**\n","metadata":{}},{"cell_type":"markdown","source":"**Overview**","metadata":{}},{"cell_type":"markdown","source":"\n\nBike sharing systems are a means of renting bicycles where the process of obtaining membership, rental, and bike return is automated via a network of kiosk locations throughout a city. Using these systems, people are able rent a bike from a one location and return it to a different place on an as-needed basis. Currently, there are over 500 bike-sharing programs around the world.\n\nData Fields\ndatetime - hourly date + timestamp\nseason - 1 = spring, 2 = summer, 3 = fall, 4 = winter\nholiday - whether the day is considered a holiday\nworkingday - whether the day is neither a weekend nor holiday\nweather -\n1: Clear, Few clouds, Partly cloudy, Partly cloudy\n2: Mist + Cloudy, Mist + Broken clouds, Mist + Few clouds, Mist\n3: Light Snow, Light Rain + Thunderstorm + Scattered clouds, Light Rain + Scattered clouds\n4: Heavy Rain + Ice Pallets + Thunderstorm + Mist, Snow + Fog\ntemp - temperature in Celsius\natemp - \"feels like\" temperature in Celsius\nhumidity - relative humidity\nwindspeed - wind speed\ncasual - number of non-registered user rentals initiated\nregistered - number of registered user rentals initiated\ncount - number of total rentals (Dependent Variable)","metadata":{}},{"cell_type":"markdown","source":"About Dataset\nOverview\nBike sharing systems are a means of renting bicycles where the process of obtaining membership, rental, and bike return is automated via a network of kiosk locations throughout a city. Using these systems, people are able rent a bike from a one location and return it to a different place on an as-needed basis. Currently, there are over 500 bike-sharing programs around the world.\n\nData Fields\ndatetime - hourly date + timestamp\nseason - 1 = spring, 2 = summer, 3 = fall, 4 = winter\nholiday - whether the day is considered a holiday\nworkingday - whether the day is neither a weekend nor holiday\nweather -\n1: Clear, Few clouds, Partly cloudy, Partly cloudy\n2: Mist + Cloudy, Mist + Broken clouds, Mist + Few clouds, Mist\n3: Light Snow, Light Rain + Thunderstorm + Scattered clouds, Light Rain + Scattered clouds\n4: Heavy Rain + Ice Pallets + Thunderstorm + Mist, Snow + Fog\ntemp - temperature in Celsius\natemp - \"feels like\" temperature in Celsius\nhumidity - relative humidity\nwindspeed - wind speed\ncasual - number of non-registered user rentals initiated\nregistered - number of registered user rentals initiated\ncount - number of total rentals (Dependent Variable)","metadata":{}},{"cell_type":"markdown","source":"This work refers to https://github.com/qinhanmin2014/kaggle-bike-sharing-demand","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport calendar\n\nfrom matplotlib.gridspec import GridSpec\nfrom sklearn.ensemble import GradientBoostingRegressor\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import make_scorer\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import GridSearchCV\n\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-14T16:43:22.940742Z","iopub.execute_input":"2022-07-14T16:43:22.941188Z","iopub.status.idle":"2022-07-14T16:43:23.877831Z","shell.execute_reply.started":"2022-07-14T16:43:22.941078Z","shell.execute_reply":"2022-07-14T16:43:23.877097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train=pd.read_csv(\"/kaggle/input/bike-sharing-demand/train.csv\")\ntest=pd.read_csv(\"/kaggle/input/bike-sharing-demand/test.csv\")\n# submission=pd.read_csv(\"/kaggle/input/bike-sharing-demand/sampleSubmission.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:23.882048Z","iopub.execute_input":"2022-07-14T16:43:23.884003Z","iopub.status.idle":"2022-07-14T16:43:23.956641Z","shell.execute_reply.started":"2022-07-14T16:43:23.883965Z","shell.execute_reply":"2022-07-14T16:43:23.955909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# transfer the target from x to log(x + 1)\nfor col in ['casual', 'registered', 'count']:\n    train['%s_log' % col] = np.log(train[col] + 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:23.960832Z","iopub.execute_input":"2022-07-14T16:43:23.962834Z","iopub.status.idle":"2022-07-14T16:43:23.979899Z","shell.execute_reply.started":"2022-07-14T16:43:23.962796Z","shell.execute_reply":"2022-07-14T16:43:23.979032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 10))\nplt.subplot(221)\nsns.distplot(train['casual'])\nplt.xlabel(\"casual (before transformation)\")\nplt.subplot(222)\nsns.distplot(np.log(train['casual'] + 1))\nplt.xlabel(\"casual (after transformation)\")\nplt.subplot(223)\nsns.distplot(train['registered'])\nplt.xlabel(\"registered (before transformation)\")\nplt.subplot(224)\nsns.distplot(np.log(train['registered'] + 1))\nplt.xlabel(\"registered (after transformation)\")\nplt.show()\n   ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:23.984928Z","iopub.execute_input":"2022-07-14T16:43:23.986969Z","iopub.status.idle":"2022-07-14T16:43:25.088887Z","shell.execute_reply.started":"2022-07-14T16:43:23.986929Z","shell.execute_reply":"2022-07-14T16:43:25.088109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# extract information from the timestamp\ntrain_date = pd.DatetimeIndex(train['datetime'])\ntrain['year'] = train_date.year\ntrain['month'] = train_date.month\ntrain['hour'] = train_date.hour\ntrain['dayofweek'] = train_date.dayofweek\ntest_date = pd.DatetimeIndex(test['datetime'])\ntest['year'] = test_date.year\ntest['month'] = test_date.month\ntest['hour'] = test_date.hour\ntest['dayofweek'] = test_date.dayofweek\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:25.090147Z","iopub.execute_input":"2022-07-14T16:43:25.090525Z","iopub.status.idle":"2022-07-14T16:43:25.115746Z","shell.execute_reply.started":"2022-07-14T16:43:25.090488Z","shell.execute_reply":"2022-07-14T16:43:25.115125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# new feature\n# non-registered user: more rentals during daytime\n# registered user: more rentals when going to work / going off work\nfig = plt.figure(figsize=(15, 10))\ngs1 = GridSpec(4, 4, fig, wspace=0.5, hspace=0.5)\nplt.subplot(gs1[:2, 1:3])\nsns.boxplot(x='hour', y='count', hue='workingday', data=train)\nplt.subplot(gs1[2:, :2])\nsns.boxplot(x='hour', y='casual', hue='workingday', data=train)\nplt.subplot(gs1[2:, 2:])\nsns.boxplot(x='hour', y='registered', hue='workingday', data=train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:25.116862Z","iopub.execute_input":"2022-07-14T16:43:25.117093Z","iopub.status.idle":"2022-07-14T16:43:27.734988Z","shell.execute_reply.started":"2022-07-14T16:43:25.117061Z","shell.execute_reply":"2022-07-14T16:43:27.734308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# new feature\n# combine year and season\ntrain['year_season'] = train_date.year + train.season / 10\nfig = plt.figure(figsize=(12, 10))\ngs1 = GridSpec(4, 4, fig, wspace=0.5, hspace=0.5)\nplt.subplot(gs1[:2, 1:3])\nsns.boxplot(x='year_season', y='count', data=train)\nplt.subplot(gs1[2:, :2])\nsns.boxplot(x='year_season', y='casual', data=train)\nplt.subplot(gs1[2:, 2:])\nsns.boxplot(x='year_season', y='registered', data=train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:27.736107Z","iopub.execute_input":"2022-07-14T16:43:27.736962Z","iopub.status.idle":"2022-07-14T16:43:28.320444Z","shell.execute_reply.started":"2022-07-14T16:43:27.736930Z","shell.execute_reply":"2022-07-14T16:43:28.319756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# new feature\nfor df in [train, test]:\n    df['year_season'] = df['year'] + df['season'] / 10\n    df['hour_workingday_casual'] = df[['hour', 'workingday']].apply(\n        lambda x: int(10 <= x['hour'] <= 19), axis=1)\n    df['hour_workingday_registered'] = df[['hour', 'workingday']].apply(\n      lambda x: int(\n        (x['workingday'] == 1 and (x['hour'] == 8 or 17 <= x['hour'] <= 18))\n        or (x['workingday'] == 0 and 10 <= x['hour'] <= 19)), axis=1)\n\nby_season = train.groupby('year_season')[['count']].median()\nby_season.columns = ['count_season']\ntrain = train.join(by_season, on='year_season')\ntest = test.join(by_season, on='year_season')","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:28.321995Z","iopub.execute_input":"2022-07-14T16:43:28.322321Z","iopub.status.idle":"2022-07-14T16:43:29.005853Z","shell.execute_reply.started":"2022-07-14T16:43:28.322271Z","shell.execute_reply":"2022-07-14T16:43:29.005072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GradientBoostingRegressor\n# features used to train the model\n# removing month improves the performance\nfeatures = ['season', 'holiday', 'workingday', 'weather',\n            'temp', 'atemp', 'humidity', 'windspeed',\n            'year', 'hour', 'dayofweek', 'hour_workingday_casual', 'count_season']\nreg = GradientBoostingRegressor(n_estimators=1000, min_samples_leaf=6, random_state=0)\nreg.fit(train[features], train['casual_log'])\npred_casual = reg.predict(test[features])\npred_casual = np.exp(pred_casual) - 1\npred_casual[pred_casual < 0] = 0\nfeatures = ['season', 'holiday', 'workingday', 'weather',\n            'temp', 'atemp', 'humidity', 'windspeed',\n            'year', 'hour', 'dayofweek', 'hour_workingday_registered', 'count_season']\nreg = GradientBoostingRegressor(n_estimators=1000, min_samples_leaf=6, random_state=0)\nreg.fit(train[features], train['registered_log'])\npred_registered = reg.predict(test[features])\npred_registered = np.exp(pred_registered) - 1\npred_registered[pred_registered < 0] = 0\npred1 = pred_casual + pred_registered","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:29.007360Z","iopub.execute_input":"2022-07-14T16:43:29.007608Z","iopub.status.idle":"2022-07-14T16:43:45.697025Z","shell.execute_reply.started":"2022-07-14T16:43:29.007574Z","shell.execute_reply":"2022-07-14T16:43:45.695692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# RandomForestRegressor\nfeatures = ['season', 'holiday', 'workingday', 'weather',\n            'temp', 'atemp', 'humidity', 'windspeed',\n            'year', 'hour', 'dayofweek', 'hour_workingday_casual', 'count_season']\nreg = RandomForestRegressor(n_estimators=1000, min_samples_leaf=2, random_state=0, n_jobs=-1)\nreg.fit(train[features], train['casual_log'])\npred_casual = reg.predict(test[features])\npred_casual = np.exp(pred_casual) - 1\npred_casual[pred_casual < 0] = 0\nfeatures = ['season', 'holiday', 'workingday', 'weather',\n            'temp', 'atemp', 'humidity', 'windspeed',\n            'year', 'hour', 'dayofweek', 'hour_workingday_registered', 'count_season']\nreg = RandomForestRegressor(n_estimators=1000, min_samples_leaf=2, random_state=0, n_jobs=-1)\nreg.fit(train[features], train['registered_log'])\npred_registered = reg.predict(test[features])\npred_registered = np.exp(pred_registered) - 1\npred_registered[pred_registered < 0] = 0\npred2 = pred_casual + pred_registered","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:43:45.700056Z","iopub.execute_input":"2022-07-14T16:43:45.701835Z","iopub.status.idle":"2022-07-14T16:44:27.821747Z","shell.execute_reply.started":"2022-07-14T16:43:45.701804Z","shell.execute_reply":"2022-07-14T16:44:27.820893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# higher weight for better model\n# rank 24/3251 public score 0.36535\npred = 0.7 * pred1 + 0.3 * pred2\nsubmission = pd.DataFrame({'datetime':test.datetime, 'count':pred},\n                          columns = ['datetime', 'count'])\nsubmission.to_csv(\"./submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T16:44:59.254626Z","iopub.execute_input":"2022-07-14T16:44:59.255093Z","iopub.status.idle":"2022-07-14T16:44:59.287665Z","shell.execute_reply.started":"2022-07-14T16:44:59.255057Z","shell.execute_reply":"2022-07-14T16:44:59.286779Z"},"trusted":true},"execution_count":null,"outputs":[]}]}