{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-14T01:53:10.596939Z","iopub.execute_input":"2022-07-14T01:53:10.597897Z","iopub.status.idle":"2022-07-14T01:53:10.607934Z","shell.execute_reply.started":"2022-07-14T01:53:10.597854Z","shell.execute_reply":"2022-07-14T01:53:10.606835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Bike Sharing Demand\n\n\n# 定义问题\n你有横跨两年的每小时租金数据。在这次比赛中，训练集由每个月的前19天组成，而测试集是每月20日到月底的数据。你必须预测测试集所涵盖的每小时内的自行车租赁总数，只使用租赁期之前的信息。\n\n### 导包\n\n在本环节中，我们导入了在数据挖掘过程中需要的各种python包。\n","metadata":{}},{"cell_type":"code","source":"# 数据整理和分析\nimport pandas as pd\nimport numpy as np\nimport random as rnd\n\n# 可视化\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport warnings\nsns.set_style('whitegrid')\nplt.rcParams['font.sans-serif'] = ['SimHei']\nsns.set(font='SimHei')  #seaborn作图默认不支持中文，需要额外设置\nwarnings.filterwarnings('ignore')\n%matplotlib inline\n\n# 机器学习\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC,LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import Perceptron\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import cross_val_score","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:54:28.832669Z","iopub.execute_input":"2022-07-14T01:54:28.833068Z","iopub.status.idle":"2022-07-14T01:54:29.764860Z","shell.execute_reply.started":"2022-07-14T01:54:28.833032Z","shell.execute_reply":"2022-07-14T01:54:29.764039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 获取数据\n\n使用pandas方法read_csv获取train.csv数据和test.csv数据。","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/bike-sharing-demand/train.csv')\ntest_df = pd.read_csv('/kaggle/input/bike-sharing-demand/test.csv')\ncombine = [train_df,test_df]\nsampleSubmission_df = pd.read_csv('/kaggle/input/bike-sharing-demand/sampleSubmission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:56:56.192907Z","iopub.execute_input":"2022-07-14T01:56:56.193317Z","iopub.status.idle":"2022-07-14T01:56:56.245557Z","shell.execute_reply.started":"2022-07-14T01:56:56.193287Z","shell.execute_reply":"2022-07-14T01:56:56.244369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 探索性数据分析（EDA）\n## 描述性数据分析\n\n### 数据集中包含哪些特征？\n\n使用train_df.columns.values获取训练集的列名称。","metadata":{}},{"cell_type":"markdown","source":"我们得到的列名称如下所示。","metadata":{}},{"cell_type":"code","source":"train_df.columns.values","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:57:13.186660Z","iopub.execute_input":"2022-07-14T01:57:13.187088Z","iopub.status.idle":"2022-07-14T01:57:13.197489Z","shell.execute_reply.started":"2022-07-14T01:57:13.187039Z","shell.execute_reply":"2022-07-14T01:57:13.196362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"对以上字段进行解释，他们的意义如下\n\n| 字段名        | 定义         | 值                              |\n|------------|------------|--------------------------------|\n| datetime   | 时间日期       |                                |\n| season     | 季节         | 1 春季 2 夏季 3 秋季 4 冬季            |\n| holiday    | 是否为节假日     |                                |\n| workingday | 是否为工作日     |                                |\n| weather    | 天气         | 1 晴朗少云 2 雾天 3 轻微雨雪天气 4 重度雨雪天气  |\n| temp       | 气温（摄氏度）    |                                |\n| atemp      |  体感温度（摄氏度） |                                |\n| humidity   |  相对湿度      |                                |\n| windspeed  |  风速        |                                |\n| casual     |  未注册用户     |                                |\n| registered |  注册用户      |                                |\n| count      |  总体租借数量    |                                |\n\n\n#### 哪些数据是分类型(定性)的？\n* 标称型数据有season，holiday，workingday\n* 序数性数据有weather\n\n#### 哪些数据是数值型（定量）的？\n* 连续型数据有temp，atemp，humidity，windspeed\n* 离散型数据有datetime，casual，registered，count\n\n#### 哪些数据有缺失值？\n\n检查训练集数据是否具有缺失值。","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:58:18.904889Z","iopub.execute_input":"2022-07-14T01:58:18.906062Z","iopub.status.idle":"2022-07-14T01:58:18.920010Z","shell.execute_reply.started":"2022-07-14T01:58:18.905995Z","shell.execute_reply":"2022-07-14T01:58:18.918752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"检查测试集数据是否含有缺失值。","metadata":{}},{"cell_type":"code","source":"test_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:58:21.157826Z","iopub.execute_input":"2022-07-14T01:58:21.158501Z","iopub.status.idle":"2022-07-14T01:58:21.167399Z","shell.execute_reply.started":"2022-07-14T01:58:21.158466Z","shell.execute_reply":"2022-07-14T01:58:21.166455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"经检查可知train_df和test_df两部分数据都没有缺失值。\n\n#### 哪些数据具有异常值？\n我们去除掉训练集中count数据落在三倍标准差之外的异常值。\n\n","metadata":{}},{"cell_type":"code","source":"# 把超出3倍标准差的数据，共147个剔除\ntrain_df = train_df.loc[np.abs(train_df['count']-train_df['count'].mean()) < (3*train_df['count'].std())]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:58:52.797879Z","iopub.execute_input":"2022-07-14T01:58:52.798288Z","iopub.status.idle":"2022-07-14T01:58:52.813549Z","shell.execute_reply.started":"2022-07-14T01:58:52.798257Z","shell.execute_reply":"2022-07-14T01:58:52.812502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 各个特征的数据类型是什么？\n","metadata":{}},{"cell_type":"code","source":"print(\"*\" * 14 +\"训练集信息\"+\"*\" * 14)\nprint(train_df.info())\nprint(\"*\" * 14 +\"测试集信息\"+\"*\" * 14)\nprint(test_df.info())","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:59:07.419832Z","iopub.execute_input":"2022-07-14T01:59:07.420267Z","iopub.status.idle":"2022-07-14T01:59:07.451629Z","shell.execute_reply.started":"2022-07-14T01:59:07.420229Z","shell.execute_reply":"2022-07-14T01:59:07.450471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"我们可以看到，我们处理的数据除日期为datetime数据类型之外，其他均为整形或者浮点型数据，无字符串类型数据。\n\n\n#### 数值型特征的经验分布如何？\n数据集的经验分布有助于我们对数据集进行初步观察，判断经验分布能否代表真实分布，在这里我们对\n\n我们对训练集数据中的连续型数据进行分析。","metadata":{}},{"cell_type":"code","source":"train_df.drop(labels=['season','holiday','workingday','weather'],axis=1).describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:59:24.220725Z","iopub.execute_input":"2022-07-14T01:59:24.221851Z","iopub.status.idle":"2022-07-14T01:59:24.264484Z","shell.execute_reply.started":"2022-07-14T01:59:24.221804Z","shell.execute_reply":"2022-07-14T01:59:24.263633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"从百分位数上看，我们可以看到casual、registered和count的分布非常不平衡，绘制出他们的概率密度曲线。","metadata":{}},{"cell_type":"code","source":"f, [ax1, ax2, ax3] = plt.subplots(1,3, figsize=(15,6))\nsns.distplot(train_df['casual'], ax=ax1)\nax1.set_title('Distribution of casual')\nsns.distplot(train_df['registered'], ax=ax2)\nax2.set_title('Distribution of registered')\nsns.distplot(train_df['count'], ax=ax3)\nax3.set_title('Distribution of count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:59:41.395442Z","iopub.execute_input":"2022-07-14T01:59:41.395836Z","iopub.status.idle":"2022-07-14T01:59:42.454034Z","shell.execute_reply.started":"2022-07-14T01:59:41.395802Z","shell.execute_reply":"2022-07-14T01:59:42.453183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"从上图中，我们可看到casual、registered和count具有相似的概率密度曲线。\n\n使用对数变换实现纠偏","metadata":{}},{"cell_type":"code","source":"count_log = np.log(train_df['count']+1)\ncasual_log = np.log(train_df['casual'] +1)\nregistered_log = np.log(train_df['registered'] + 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:59:57.949844Z","iopub.execute_input":"2022-07-14T01:59:57.950261Z","iopub.status.idle":"2022-07-14T01:59:57.957737Z","shell.execute_reply.started":"2022-07-14T01:59:57.950227Z","shell.execute_reply":"2022-07-14T01:59:57.956978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f, [ax1, ax2, ax3] = plt.subplots(1,3, figsize=(15,6))\nsns.distplot(casual_log, ax=ax1)\nax1.set_title('Distribution of casual')\nsns.distplot(registered_log, ax=ax2)\nax2.set_title('Distribution of registered')\nsns.distplot(count_log, ax=ax3)\nax3.set_title('Distribution of count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:00:06.322405Z","iopub.execute_input":"2022-07-14T02:00:06.322780Z","iopub.status.idle":"2022-07-14T02:00:07.288485Z","shell.execute_reply.started":"2022-07-14T02:00:06.322749Z","shell.execute_reply":"2022-07-14T02:00:07.287012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['casual','registered','count']].corr()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:00:14.512864Z","iopub.execute_input":"2022-07-14T02:00:14.513283Z","iopub.status.idle":"2022-07-14T02:00:14.530867Z","shell.execute_reply.started":"2022-07-14T02:00:14.513251Z","shell.execute_reply":"2022-07-14T02:00:14.529747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"我们可以看到causal和count之间呈现中等相关，registered和count之间呈现强相关。考虑到causal和count之间呈现中等相关，在进行预测时，我们应该分别对causal和registered进行估计，并将其相加以获取最终的预测结果。","metadata":{}},{"cell_type":"code","source":"f, [[ax1, ax2],[ ax3,ax4]] = plt.subplots(2,2, figsize=(15,12))\nsns.distplot(train_df['atemp'], ax=ax1)\nax1.set_title('Distribution of atemp')\nsns.distplot(train_df['temp'], ax=ax2)\nax2.set_title('Distribution of temp')\nsns.distplot(train_df['humidity'], ax=ax3)\nax3.set_title('Distribution of humidity')\nsns.distplot(train_df['windspeed'], ax=ax4)\nax4.set_title('Distribution of windspeed')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:00:38.337778Z","iopub.execute_input":"2022-07-14T02:00:38.338614Z","iopub.status.idle":"2022-07-14T02:00:39.835353Z","shell.execute_reply.started":"2022-07-14T02:00:38.338569Z","shell.execute_reply":"2022-07-14T02:00:39.834086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"我们可以看到风速为0的数据偏多，因此我们推测该部分原始数据已经使用零来替代了风速方面的缺失值，因此我们采用随机森林的方法填充风速为零的部分。首先，我们可以将train_df和test_df合并到一起以方便特征工程。\n","metadata":{}},{"cell_type":"code","source":"#首先合并数据\ndf = train_df.append(test_df, ignore_index=True)\n#整理列顺序\ndf = pd.DataFrame(df, columns=train_df.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:00:50.188788Z","iopub.execute_input":"2022-07-14T02:00:50.189206Z","iopub.status.idle":"2022-07-14T02:00:50.198886Z","shell.execute_reply.started":"2022-07-14T02:00:50.189171Z","shell.execute_reply":"2022-07-14T02:00:50.197894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"将df中的datetime拆分为多种时间格式，方便后续分析。","metadata":{}},{"cell_type":"code","source":"#转化格式\ndf['datetime'] = pd.to_datetime(df['datetime'], format='%Y-%m-%d %H:%M:%S')\n#细化，这里去了weekday，方便看每周的变化情况\ndf['year'] = df['datetime'].dt.year\ndf['month'] = df['datetime'].dt.month\ndf['hour'] = df['datetime'].dt.hour\ndf['weekday'] = df['datetime'].dt.weekday","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:01:24.985359Z","iopub.execute_input":"2022-07-14T02:01:24.985786Z","iopub.status.idle":"2022-07-14T02:01:25.012867Z","shell.execute_reply.started":"2022-07-14T02:01:24.985747Z","shell.execute_reply":"2022-07-14T02:01:25.011433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"我们猜测风速和天气以及时间都有关，因此我们使用随机森林的方法填充风速为零的异常值。","metadata":{}},{"cell_type":"code","source":"wind_0 = df[df['windspeed']==0]\nwind_not0 = df[df['windspeed']!=0]\ny_label = wind_not0['windspeed']\n#猜测风速和天气以及时间都有关\nfrom sklearn.ensemble import RandomForestClassifier\nmodel = RandomForestClassifier()\nwindcolunms = ['season', 'weather', 'temp', 'atemp', 'humidity', 'hour', 'month']\nmodel.fit(wind_not0[windcolunms], y_label.astype('int'))\npred_y = model.predict(wind_0[windcolunms])\n#预测结果填充\nwind_0['windspeed'] = pred_y\ndf_rfw = wind_not0.append(wind_0)\ndf_rfw.reset_index(inplace=True)\ndf_rfw.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:01:27.127678Z","iopub.execute_input":"2022-07-14T02:01:27.128110Z","iopub.status.idle":"2022-07-14T02:01:29.582281Z","shell.execute_reply.started":"2022-07-14T02:01:27.128073Z","shell.execute_reply":"2022-07-14T02:01:29.581216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_rfw = df_rfw.drop('index', axis=1)\n#查看处理后的风速情况\nf, ax = plt.subplots(figsize=(8,5))\nsns.distplot(df_rfw['windspeed'], ax=ax)\nax.set_title('Distribution of handled windspeed')","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:01:43.684618Z","iopub.execute_input":"2022-07-14T02:01:43.685007Z","iopub.status.idle":"2022-07-14T02:01:44.136081Z","shell.execute_reply.started":"2022-07-14T02:01:43.684976Z","shell.execute_reply":"2022-07-14T02:01:44.134924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 分析数据\n\n首先，我们绘制各种类型数据之间的相关系数矩阵，如下图示。","metadata":{}},{"cell_type":"code","source":"#查看各组数据和count的相关性\nf, ax = plt.subplots(figsize=(20,20))\ncmap = sns.diverging_palette(220, 10, as_cmap=True)\nsns.heatmap(df_rfw[df_rfw['count'].notnull()].corr(), cmap=cmap, ax=ax, annot=True, lw=.1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:01:58.416758Z","iopub.execute_input":"2022-07-14T02:01:58.417844Z","iopub.status.idle":"2022-07-14T02:01:59.667122Z","shell.execute_reply.started":"2022-07-14T02:01:58.417798Z","shell.execute_reply":"2022-07-14T02:01:59.665975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"我们获取\"count\"、\"registered\"和\"casual\"和其他字段的相关性排序。\n","metadata":{}},{"cell_type":"code","source":"#相关性排序\ndf_rfw[df_rfw['count'].notnull()].corr()['count'].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:02:14.226828Z","iopub.execute_input":"2022-07-14T02:02:14.227263Z","iopub.status.idle":"2022-07-14T02:02:14.249360Z","shell.execute_reply.started":"2022-07-14T02:02:14.227225Z","shell.execute_reply":"2022-07-14T02:02:14.248122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#相关性排序\ndf_rfw[df_rfw['count'].notnull()].corr()['registered'].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:02:25.996844Z","iopub.execute_input":"2022-07-14T02:02:25.997774Z","iopub.status.idle":"2022-07-14T02:02:26.017068Z","shell.execute_reply.started":"2022-07-14T02:02:25.997729Z","shell.execute_reply":"2022-07-14T02:02:26.015976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_rfw[df_rfw['count'].notnull()].corr()['casual'].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:02:33.220417Z","iopub.execute_input":"2022-07-14T02:02:33.220804Z","iopub.status.idle":"2022-07-14T02:02:33.245636Z","shell.execute_reply.started":"2022-07-14T02:02:33.220771Z","shell.execute_reply":"2022-07-14T02:02:33.244730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"排除掉自身之后，我们有registered和casual的相关性排序如下所示。\n\n| registered |           | casual     |            |\n|------------|-----------|------------|------------|\n| hour       | 0.388358  | temp       | 0.460774   |\n| temp       | 0.304328  | atemp      | 0.456319   |\n| atemp      | 0.301074  | hour       | 0.295037   |\n| year       | 0.239181  | weekday    | 0.257293   |\n| month      | 0.168459  | year       | 0.13212    |\n| season     | 0.160967  | windspeed  | 0.112242   |\n| windspeed  | 0.096816  | season     | 0.09455    |\n| workingday | 0.095554  | month      | 0.09032    |\n| holiday    | -0.013535 | holiday    | 0.04727    |\n| weekday    | -0.065856 | weather    | -0.133326  |\n| weather    | -0.107421 | workingday | -0.332853  |\n| humidity   | -0.263525 | humidity   | -0.341204  |\n\n\n由上表可知，对于注册用户，下列数据有较为明显的正相关。\n* hour\n* temp\n* atemp\n* year\n* month\n* season\n\n下列数据有较为明显的负相关。\n* weather\n* humidity\n\n对于非注册用户，下列数据有较为明显的正相关。\n* temp\n* atemp\n* hour\n* weekday\n* year\n* windspeed\n\n下列数据有较为明显的负相关\n* weather\n* workingday\n* humidity\n\n综上所述，我们选取hour、year、temp、atemp、month、season、weather、humidity数据字段以实现对注册用户的预测，选取temp、atemp、hour、weekday、year、windspeed、weather、workingday、humidity字段时间对非注册用户的预测。\n\n对数据进行独热编码。","metadata":{}},{"cell_type":"code","source":"season_dummy=pd.get_dummies(df_rfw['season'],prefix='season')\nweather_dummy=pd.get_dummies(df_rfw['weather'],prefix='weather')\nmonth_dummy=pd.get_dummies(df_rfw['month'],prefix='month')\ndf_rfw1 = pd.concat([df_rfw,season_dummy,weather_dummy,month_dummy],axis=1)\ndf_rfw1.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:02:48.978765Z","iopub.execute_input":"2022-07-14T02:02:48.979135Z","iopub.status.idle":"2022-07-14T02:02:49.010687Z","shell.execute_reply.started":"2022-07-14T02:02:48.979106Z","shell.execute_reply":"2022-07-14T02:02:49.009439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 预测共享单车出租量\n\n#### 对注册用户数据进行预测\n\n首先，我们去除掉对于影响注册用户决策关系不大的字段。","metadata":{}},{"cell_type":"code","source":"# 对于注册用户，我们需要的字段为hour、year、temp、atemp、month、season、weather、humidity\ndrop_column_reg= ['datetime','season','holiday','workingday','weather','windspeed','casual', 'registered', 'count', 'month',]\n# 分离训练集和测试集\ndf_train=df_rfw1[df_rfw1['count'].notnull()].sort_values('datetime',ascending=True)\ndf_test=df_rfw1[df_rfw['count'].isnull()].sort_values('datetime',ascending=True)\n\n# 注册用户\nreg_rental = df_train['registered']\ndf_train_reg = df_train.drop(columns=drop_column_reg,axis=1)\ndf_test_reg = df_test.drop(columns=drop_column_reg,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:03:11.014256Z","iopub.execute_input":"2022-07-14T02:03:11.014639Z","iopub.status.idle":"2022-07-14T02:03:11.039431Z","shell.execute_reply.started":"2022-07-14T02:03:11.014607Z","shell.execute_reply":"2022-07-14T02:03:11.038314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"使用随机森林方法进行预测。","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nmodel_reg =RandomForestRegressor(n_estimators=1000,random_state=42)\n# model_reg.fit(df_train_reg,reg_rental)\nmodel_reg.fit(df_train_reg,registered_log)\npred_reg=np.exp(model_reg.predict(df_test_reg))-1","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:03:22.290706Z","iopub.execute_input":"2022-07-14T02:03:22.291474Z","iopub.status.idle":"2022-07-14T02:04:06.334994Z","shell.execute_reply.started":"2022-07-14T02:03:22.291433Z","shell.execute_reply":"2022-07-14T02:04:06.333895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 对非注册用户进行预测\n\n首先，我们去除掉对非注册用户决策影响不大的字段，以此作为训练集数据。","metadata":{}},{"cell_type":"code","source":"# 对于非注册用户，我们需要的字段为temp、atemp、hour、weekday、year、windspeed、weather、workingday、humidity\ndrop_column_cas= ['datetime','season','holiday','weather','casual', 'registered', 'count', 'month',]\n# 分离训练集和测试集\ndf_train=df_rfw1[df_rfw1['count'].notnull()].sort_values('datetime',ascending=True)\ndf_test=df_rfw1[df_rfw['count'].isnull()].sort_values('datetime',ascending=True)\n\n# 非注册用户\ncas_rental = df_train['casual']\ndf_train_cas = df_train.drop(columns=drop_column_cas,axis=1)\ndf_test_cas = df_test.drop(columns=drop_column_cas,axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:04:06.361823Z","iopub.execute_input":"2022-07-14T02:04:06.362617Z","iopub.status.idle":"2022-07-14T02:04:06.385800Z","shell.execute_reply.started":"2022-07-14T02:04:06.362569Z","shell.execute_reply":"2022-07-14T02:04:06.384265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"使用随机森林方法对非注册用户进行预测。","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\nmodel_cas =RandomForestRegressor(n_estimators=1000,random_state=42)\n# model_cas.fit(df_train_cas,cas_rental)\nmodel_cas.fit(df_train_cas,casual_log)\n\npred_cas=np.exp(model_cas.predict(df_test_cas))-1","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:04:44.771984Z","iopub.execute_input":"2022-07-14T02:04:44.772375Z","iopub.status.idle":"2022-07-14T02:05:30.854240Z","shell.execute_reply.started":"2022-07-14T02:04:44.772341Z","shell.execute_reply":"2022-07-14T02:05:30.853111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"将预测结果保存为bike_submission.csv。","metadata":{}},{"cell_type":"code","source":"pred=pd.Series(pred_cas + pred_reg ,name='count')\n\npred_concat=pd.concat([sampleSubmission_df['datetime'],pred],axis=1)\npred_concat.to_csv('output\\\\bike_submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T02:05:47.642404Z","iopub.execute_input":"2022-07-14T02:05:47.642791Z","iopub.status.idle":"2022-07-14T02:05:47.667295Z","shell.execute_reply.started":"2022-07-14T02:05:47.642761Z","shell.execute_reply":"2022-07-14T02:05:47.666198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 提交测试结果\n\n我们的kaggle分数为![image-20220714094916337](https://img1.imgtp.com/2022/07/14/gkvMNj0L.png)\n\n因为一些奇怪的原因，我们没有在kaggle排行榜上找到自己的名字。但是根据已经存在的排行榜可知，我们的kaggle排名为485名左右。\n\n![image-20220714093540025](https://img1.imgtp.com/2022/07/14/RaJSXz79.png)\n\n# 问题回答\n\n在本次实验中我们采用使用$X = log(X + 1)$函数对count、causal、registered的数据在分数提高上具有显著作用。\n\n其原因在于在进行纠偏之前，我们使用的训练数据具有明显的长尾现象。\n\n![png](https://img1.imgtp.com/2022/07/14/WcSOveeK.png)\n\n使得随机森林预测方法不能很好适应count、casual、registered取值较大的情况。\n\n在进行纠偏之前，我们获得的最大的预测分数在0.44左右，纠偏操作使得我们的分数提高到了0.42。\n\n在本次实验中，我们采用的主要改进为根据causal和registered的分布情况显著不同的情况，使用随机森林方法对两者进行分别预测，以提高预测的准确率。\n\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}