{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Bike Sharing Demand\n## 경진대회 이해\n- 워싱턴 D.C의 자전거 무인 대여 시스템 과거 기록을 기반으로 향후 자전거 대여 수요를 예측하는 대회\n- 2011년부터 2012년까지 2년간의 자전거 대여 데이터\n- 대여 데이터는 한 시간 간격으로 기록\n- 훈련 데이터는 매달 1일 부터 19일까지의 기록\n- 테스트 데이터는 매달 20일부터 월말까지의 기록\n- 피처는 대여 날짜, 시간, 요일, 계절, 날씨, 실제 온도, 체감 온도, 습도, 풍속, 회원 여부\n- 위 데이터를 활용해 시간별 자전거 대여 수량을 예측하는 회귀 문제\n\n## 평가 지표\n- RMSLE(Root Mean Squared Logarithmic Error)\n\n\n$$\\sqrt{{1\\over{n}} \\sum_{i=1}^n (log((p_i + 1) - log(a_i + 1))^2}$$","metadata":{"papermill":{"duration":0.026639,"end_time":"2021-08-16T04:01:01.662249","exception":false,"start_time":"2021-08-16T04:01:01.63561","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"## EDA\n### 데이터 둘러보기","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\ndata_path = '/kaggle/input/bike-sharing-demand/'\n\ntrain = pd.read_csv(data_path + 'train.csv')\ntest = pd.read_csv(data_path + 'test.csv')\nsubmission = pd.read_csv(data_path + 'sampleSubmission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.608446Z","iopub.execute_input":"2022-08-05T05:32:16.609204Z","iopub.status.idle":"2022-08-05T05:32:16.673611Z","shell.execute_reply.started":"2022-08-05T05:32:16.609070Z","shell.execute_reply":"2022-08-05T05:32:16.672376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.shape, test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.675591Z","iopub.execute_input":"2022-08-05T05:32:16.675926Z","iopub.status.idle":"2022-08-05T05:32:16.686519Z","shell.execute_reply.started":"2022-08-05T05:32:16.675885Z","shell.execute_reply":"2022-08-05T05:32:16.685108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.688660Z","iopub.execute_input":"2022-08-05T05:32:16.689326Z","iopub.status.idle":"2022-08-05T05:32:16.710519Z","shell.execute_reply.started":"2022-08-05T05:32:16.689275Z","shell.execute_reply":"2022-08-05T05:32:16.709586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"|피처명|설명|\n|----|-----|\n|datetime|기록 일시(1시간 간격)|\n|seadon|계절(1:봄, 2:여름, 3:가을, 4:겨울)|\n|holiday|공휴일 여부(0:공휴일 아님, 1:공휴일)|\n|workingday|근무일 여부(0:근무일 아님, 1:근무일)|\n|weather|날씨(1:맑음, 2:옅은 안개, 약간흐림, 3:약간의 눈, 약간의 비와 천둥 번개, 흐림, 4:폭우와 천둥 번개, 눈과 짙은 안개)|\n|temp|실제 온도|\n|atemp|체감 온도|\n|humidity|상대 습도|\n|windspeed|풍속|\n|casual|등록되지 않은 사용자(비회원) 수|\n|registered|등록된 사용자(회원) 수|\n|count|자전거 대여 수량|","metadata":{}},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.712310Z","iopub.execute_input":"2022-08-05T05:32:16.712923Z","iopub.status.idle":"2022-08-05T05:32:16.727856Z","shell.execute_reply.started":"2022-08-05T05:32:16.712885Z","shell.execute_reply":"2022-08-05T05:32:16.726934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.729564Z","iopub.execute_input":"2022-08-05T05:32:16.730113Z","iopub.status.idle":"2022-08-05T05:32:16.748322Z","shell.execute_reply.started":"2022-08-05T05:32:16.730067Z","shell.execute_reply":"2022-08-05T05:32:16.747148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.750334Z","iopub.execute_input":"2022-08-05T05:32:16.750720Z","iopub.status.idle":"2022-08-05T05:32:16.772864Z","shell.execute_reply.started":"2022-08-05T05:32:16.750612Z","shell.execute_reply":"2022-08-05T05:32:16.772187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.773981Z","iopub.execute_input":"2022-08-05T05:32:16.775059Z","iopub.status.idle":"2022-08-05T05:32:16.790318Z","shell.execute_reply.started":"2022-08-05T05:32:16.774999Z","shell.execute_reply":"2022-08-05T05:32:16.789545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 피처 엔지니어링\n- `datatime` 피처가 시각화하기에 적합하지 않음","metadata":{}},{"cell_type":"code","source":"train['datetime']","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.792133Z","iopub.execute_input":"2022-08-05T05:32:16.792576Z","iopub.status.idle":"2022-08-05T05:32:16.800709Z","shell.execute_reply.started":"2022-08-05T05:32:16.792539Z","shell.execute_reply":"2022-08-05T05:32:16.800077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train['datetime'][1]) # datetime 1번째 원소\nprint(train['datetime'][1].split()) # 공백 기준 문자열 나누기\nprint(train['datetime'][1].split()[0]) # 날짜\nprint(train['datetime'][1].split()[1]) # 시간","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.801801Z","iopub.execute_input":"2022-08-05T05:32:16.802753Z","iopub.status.idle":"2022-08-05T05:32:16.814061Z","shell.execute_reply.started":"2022-08-05T05:32:16.802713Z","shell.execute_reply":"2022-08-05T05:32:16.812720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train['datetime'][1].split()[0]) # 날짜\nprint(train['datetime'][1].split()[0].split(\"-\")) # \"-\" 기준 문자열 나누기\nprint(train['datetime'][1].split()[0].split(\"-\")[0]) # 연도\nprint(train['datetime'][1].split()[0].split(\"-\")[1]) # 월\nprint(train['datetime'][1].split()[0].split(\"-\")[2]) # 일","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.816789Z","iopub.execute_input":"2022-08-05T05:32:16.817673Z","iopub.status.idle":"2022-08-05T05:32:16.828309Z","shell.execute_reply.started":"2022-08-05T05:32:16.817639Z","shell.execute_reply":"2022-08-05T05:32:16.826873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train['datetime'][1].split()[1]) # 시간\nprint(train['datetime'][1].split()[1].split(\":\")) # \":\" 기준 문자열 나누기\nprint(train['datetime'][1].split()[1].split(\":\")[0]) # 시간\nprint(train['datetime'][1].split()[1].split(\":\")[1]) # 분\nprint(train['datetime'][1].split()[1].split(\":\")[2]) # 초","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.830022Z","iopub.execute_input":"2022-08-05T05:32:16.830460Z","iopub.status.idle":"2022-08-05T05:32:16.839031Z","shell.execute_reply.started":"2022-08-05T05:32:16.830414Z","shell.execute_reply":"2022-08-05T05:32:16.837778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\npandas apply()함수로 \n날짜(date), 연도(year), 월(month), 일(day), 시(hour), 분(minute), 초(second) 피처 생성\n\"\"\"\ntrain['date'] = train['datetime'].apply(lambda x: x.split()[0])\ntrain['year'] = train['datetime'].apply(lambda x: x.split()[0].split('-')[0])\ntrain['month'] = train['datetime'].apply(lambda x: x.split()[0].split('-')[1])\ntrain['day'] = train['datetime'].apply(lambda x: x.split()[0].split('-')[2])\ntrain['hour'] = train['datetime'].apply(lambda x: x.split()[1].split(':')[0])\ntrain['minute'] = train['datetime'].apply(lambda x: x.split()[1].split(':')[1])\ntrain['second'] = train['datetime'].apply(lambda x: x.split()[1].split(':')[2])","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.840489Z","iopub.execute_input":"2022-08-05T05:32:16.841235Z","iopub.status.idle":"2022-08-05T05:32:16.908886Z","shell.execute_reply.started":"2022-08-05T05:32:16.841193Z","shell.execute_reply":"2022-08-05T05:32:16.907837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 요일 피처 생성\n\nfrom datetime import datetime\nimport calendar\n\nprint(train['date'][1]) # 날짜\nprint(datetime.strptime(train['date'][1], '%Y-%m-%d')) # datetime 타입으로 변경\nprint(datetime.strptime(train['date'][1], '%Y-%m-%d').weekday()) # 정수로 요일 반환(0~6)\nprint(calendar.day_name[datetime.strptime(train['date'][1], '%Y-%m-%d').weekday()]) # 문자열로 요일 반환","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.910980Z","iopub.execute_input":"2022-08-05T05:32:16.911513Z","iopub.status.idle":"2022-08-05T05:32:16.921712Z","shell.execute_reply.started":"2022-08-05T05:32:16.911467Z","shell.execute_reply":"2022-08-05T05:32:16.920302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['weekday'] = train['date'].apply(\n    lambda date_string: \n    calendar.day_name[datetime.strptime(date_string, \"%Y-%m-%d\").weekday()]\n)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:16.924778Z","iopub.execute_input":"2022-08-05T05:32:16.925542Z","iopub.status.idle":"2022-08-05T05:32:17.074010Z","shell.execute_reply.started":"2022-08-05T05:32:16.925505Z","shell.execute_reply":"2022-08-05T05:32:17.073133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# season과 weater 시각화 시 의미가 잘 드러나도록 문자열로 변환\n\ntrain['season'] = train['season'].map({1: 'Spring', \n                                       2: 'Summer', \n                                       3: 'Fall', \n                                       4: 'Winter'})\n\ntrain['weather'] = train['weather'].map({1: 'Clear', \n                                         2: 'Mist, Few clouds', \n                                         3: 'Light snow, Rain, Thunderstorm', \n                                         4: 'Heavy Rain, Thunderstorm, Snow, Fog'})","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:17.075535Z","iopub.execute_input":"2022-08-05T05:32:17.075876Z","iopub.status.idle":"2022-08-05T05:32:17.087959Z","shell.execute_reply.started":"2022-08-05T05:32:17.075832Z","shell.execute_reply":"2022-08-05T05:32:17.086961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:17.089516Z","iopub.execute_input":"2022-08-05T05:32:17.089956Z","iopub.status.idle":"2022-08-05T05:32:17.122103Z","shell.execute_reply.started":"2022-08-05T05:32:17.089916Z","shell.execute_reply":"2022-08-05T05:32:17.121297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`date` 피처가 제공하는 정보는 `year`, `month`, `day` 피처에도 있어 추후 `date` 피처 제거\n\n\n세 달씩 '월'을 묶으면 '계절'이 되므로 `season` 피처만 남기고 `month` 피처는 추후 제거","metadata":{}},{"cell_type":"markdown","source":"### 데이터 시각화","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\n\nplt.style.use('seaborn')\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:17.123227Z","iopub.execute_input":"2022-08-05T05:32:17.124455Z","iopub.status.idle":"2022-08-05T05:32:17.512894Z","shell.execute_reply.started":"2022-08-05T05:32:17.124412Z","shell.execute_reply":"2022-08-05T05:32:17.511748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 분포도\n- 수치형 데이터의 집계 값을 나타내는 그래프\n- 집계 값은 총 개수나 비율 등을 의미","metadata":{}},{"cell_type":"code","source":"# 타깃값인 count의 분포도\n\nmpl.rc('font', size=15) # 폰트 크기를 15로 설정\nsns.displot(train['count']) # 분포도 출력","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:17.514835Z","iopub.execute_input":"2022-08-05T05:32:17.515149Z","iopub.status.idle":"2022-08-05T05:32:18.002950Z","shell.execute_reply.started":"2022-08-05T05:32:17.515109Z","shell.execute_reply":"2022-08-05T05:32:18.002156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 회귀 모델이 좋은 성능을 내려면 데이터가 정규분포를 따라야 하는데, 현재 타깃값은 정규분포를 따르지 않는다.\n- 데이터 분포를 정규분포에 가깝게 만들기 위해 로그변환을 많이 사용한다(데이터가 왼쪽으로 편향되어 있을때).\n- 즉, count를 예측하는 것보다 log(count)를 예측하는 편이 더 정확하다.\n- 다만, 마지막에 예측값을 실제 타깃값인 count로 지수변환해야 한다.","metadata":{}},{"cell_type":"code","source":"sns.displot(np.log(train['count']))","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:18.004556Z","iopub.execute_input":"2022-08-05T05:32:18.005113Z","iopub.status.idle":"2022-08-05T05:32:18.490636Z","shell.execute_reply.started":"2022-08-05T05:32:18.005054Z","shell.execute_reply":"2022-08-05T05:32:18.489593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"$$y = e^{log(y)}$$","metadata":{}},{"cell_type":"markdown","source":"#### 막대 그래프\n- 연도, 월, 일, 시, 분, 초별로 총 여섯 가지의 평균 대여 수량을 막대 그래프로 그려보기\n- 각 범주형 데이터에 따라 평균 대여 수량이 어떻게 다른지 파악하기 위함","metadata":{}},{"cell_type":"code","source":"mpl.rc('font', size=14) # 폰트 크기 설정\nmpl.rc('axes', titlesize=15) # 각 축의 제목 크기 설정\nfigure, axes = plt.subplots(nrows=3, ncols=2) # 3행 2열 Figure 생성\nplt.tight_layout() # 그래프 사이에 여백 확보\nfigure.set_size_inches(10, 9) # 전체 Figure 크기를 10x9인치로 설정\n\nx = ['year', 'month', 'day', 'hour', 'minute', 'second']\nfor i in range(6):\n    row = i // 2\n    col = i % 2\n    \n    # 각 축에 서브플롯 할당\n    sns.barplot(x=x[i], y='count', data=train, ax=axes[row, col])\n    \n    # 각 서브플롯에 제목 추가\n    axes[row, col].set(title='Rental amounts by ' + x[i])\n\n# 1행에 위치한 서브플롯들의 x축 라벨 90도 회전\naxes[1, 0].tick_params(axis='x', labelrotation=90)\naxes[1, 1].tick_params(axis='x', labelrotation=90)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:18.492606Z","iopub.execute_input":"2022-08-05T05:32:18.493202Z","iopub.status.idle":"2022-08-05T05:32:22.084858Z","shell.execute_reply.started":"2022-08-05T05:32:18.493145Z","shell.execute_reply":"2022-08-05T05:32:22.084173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. 연도별 평균 대여 수량은 2011년보다 2012년에 더 많았다.\n2. 평균 대여 수량은 6월에 가장 많고 1월에 가장 적었다.\n    - 날씨가 따뜻할수록 대여 수량이 많다고 짐작가능\n3. 일병 대여 수량은 뚜렷한 차이가 없다.\n    - 훈련 데이터에는 1일부터 19일까지의 데이터만 있고, 테스트 데이터에 나머지 데이터가 있어 훈련에 피처로 사용불가\n4. 시간별 평균 대여 수량은 새벽 4시가 가장 적고, 아침 8시와 저녁 5~6시에 가장 많다.\n    - 등하교 혹은 출되근 길에 자전거를 많이 이용한다고 짐작가능\n5. 분과 초는 모두 0으로 기록되어 있기 때문에 별다른 정보를 담고 있지 않다.\n    - 모델을 훈련할 때 분과 초 피처는 사용X","metadata":{}},{"cell_type":"markdown","source":"#### 박스플롯\n- 범주형 데이터에 따른 수치형 데이터 정보를 나타내는 그래프\n- 계절, 날씨, 공휴일, 근무일(범주형 데이터)별 대여 수량(수치형 데이터)을 그리기","metadata":{}},{"cell_type":"code","source":"# 2행 2열 Figure 준비\nfigure, axes = plt.subplots(nrows=2, ncols=2) # 2행 2열\nplt.tight_layout()\nfigure.set_size_inches(10, 10)\n\n# 서브플롯 할당\nx = ['season', 'weather', 'holiday', 'workingday'] \nfor i in range(4):\n    row = i // 2\n    col = i % 2\n    \n    sns.boxplot(x=x[i], y='count', data=train, ax=axes[row, col])\n    # 제목 달기\n    axes[row, col].set(title='Box Plot On Count Across ' + x[i])\n    \n# x축 라벨 겹침 해결\naxes[0, 1].tick_params(axis='x', labelrotation=10)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:22.086371Z","iopub.execute_input":"2022-08-05T05:32:22.086876Z","iopub.status.idle":"2022-08-05T05:32:23.005600Z","shell.execute_reply.started":"2022-08-05T05:32:22.086791Z","shell.execute_reply":"2022-08-05T05:32:23.004005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. 자전거 대여 수량은 봄에 가장 적고, 가을에 가장 많다.\n2. 날씨가 좋을 때 대여 수량이 가장 많고, 안 좋을수록 수량이 적다.\n3. 공휴일일 때와 아닐 때 자전거 대여 수량의 중앙값은 거의 비슷하다.\n    - 다만, 공휴일이 아닐 때(0)는 이상치가 많다.\n4. 근무일일 때 이상치가 많다.","metadata":{}},{"cell_type":"markdown","source":"#### 포인트플롯\n- 범주형 데이터에 따른 수치형 데이터의 평균과 신뢰구간을 점과 선으로 표시\n- 막대 그래프와 동일한 정보를 제공하지만, 한 화면에 여러 그래프를 그려 서로 비교해보기 더 적합하다.\n- 근무일, 공휴일, 요일, 계절, 날씨에 따른 시간대별 평균 대여 수량을 포인트플롯으로 그리기","metadata":{}},{"cell_type":"code","source":"mpl.rc('font', size=11) # 폰트 크기 설정\nfigure, axes = plt.subplots(nrows=5) # 5행 1열\nfigure.set_size_inches(12, 18)\n\n# 서브플롯 할당\nhue = ['workingday', 'holiday', 'weekday', 'season', 'weather']\nfor i in range(5):\n    sns.pointplot(x='hour', y='count', data=train, hue=hue[i], ax=axes[i])","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:23.007383Z","iopub.execute_input":"2022-08-05T05:32:23.008146Z","iopub.status.idle":"2022-08-05T05:32:37.528495Z","shell.execute_reply.started":"2022-08-05T05:32:23.008093Z","shell.execute_reply":"2022-08-05T05:32:37.527124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. 근무일에는 출퇴근 시간에 대여 수량이 많고 쉬는 날에는 오후 12~2시에 가장 많다.\n2. 공휴일 여부, 요일에 따른 포인트플롯도 근무일 여부에 따른 그래프와 비슷한 양상을 보인다.\n3. 계절에 따른 시간대별 포인트플롯에서는, 대여 수량은 가을에 가장 많고 봄에 가장 적다.\n4. 날씨에 따른 그래프에서는, 날씨가 좋을 때 대여량이 가장 많았다.\n    - 폭우, 폭설이 내릴 때 대여 건수가 있었다. 이런 이상치는 제거를 고려해보는 것이 좋다.","metadata":{}},{"cell_type":"markdown","source":"#### 회귀선을 포함한 산점도 그래프\n- 수치형 데이터 간 상관관계를 파악하는 데 사용한다.\n- 수치형 데이터인 온도, 체감 온도, 풍속, 습도별 대여 수량을 '회귀선을 포함한 산점도 그래프'로 그리기","metadata":{}},{"cell_type":"code","source":"mpl.rc('font', size=15)\nfigure, axes = plt.subplots(nrows=2, ncols=2)\nplt.tight_layout()\nfigure.set_size_inches(7, 6)\n\nx = ['temp', 'atemp', 'windspeed', 'humidity']\nfor i in range(4):\n    row = i // 2\n    col = i % 2\n    \n    # alpha: 산점도의 투명도 조절 (20% 수준으로 설정)\n    sns.regplot(x=x[i], y='count', data=train, ax=axes[row, col], \n                scatter_kws={'alpha': 0.2}, line_kws={'color': 'blue'})","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:37.530672Z","iopub.execute_input":"2022-08-05T05:32:37.531565Z","iopub.status.idle":"2022-08-05T05:32:40.851197Z","shell.execute_reply.started":"2022-08-05T05:32:37.531473Z","shell.execute_reply":"2022-08-05T05:32:40.849985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. 온도와 체감 온도가 높을수록 대여 수량이 많다.\n2. 습도는 낮을수록 대여를 많이 한다.\n3. `windspeed` 피처에 0으로 기록된 값이 많다.\n    - 실제 풍속이 0이 아니라 관측치가 없거나 오류로 인해 0으로 기록됐을 가능성이 높다.\n    - 결측값이 많아서 풍속과 대여 수량의 상관관계를 파악하기 힘들어 추후 피처를 삭제 ","metadata":{}},{"cell_type":"markdown","source":"#### 히트맵\n- 데이터 간 관계를 색상으로 표현하여, 여러 데이터를 한눈에 비교하기 좋다.\n- 수치형 데이터(temp, atemp, humidity, windspeed, count)끼리 어떤 상관관계가 있는지 히트맵으로 표현","metadata":{}},{"cell_type":"code","source":"# 피처 간 상관관계 매트릭스\ncorr_matrix = train[['temp', 'atemp', 'humidity', 'windspeed', 'count']].corr()\n\nfig, ax = plt.subplots()\nfig.set_size_inches(10, 10)\nsns.heatmap(corr_matrix, annot=True) # 상관관계 히트맵 그리기\nax.set(title='Heatmap of Numerical Data')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:40.853095Z","iopub.execute_input":"2022-08-05T05:32:40.853516Z","iopub.status.idle":"2022-08-05T05:32:41.282992Z","shell.execute_reply.started":"2022-08-05T05:32:40.853481Z","shell.execute_reply":"2022-08-05T05:32:41.281931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. 온도와 대여 수량 간 상관계수는 0.39로 양의 상관관계를 보인다.\n2. 습도와 대여 수량은 음의 상관관계를 보인다.\n3. 풍속은 상관관계가 매우 약해 별 도움을 주지 못할 것이라 판단\n    - 산점도와 마찬가지로 피처 제거 결정","metadata":{}},{"cell_type":"markdown","source":"## 분석 정리\n1. 타깃값 변환\n    - 분포도 확인 결과 타깃값인 count가 0 근처로 치우쳐 있으므로 로그변환하여 정규분포에 가깝게 만들기\n    - 예측값을 마지막에 다시 지수변환해 count로 복원\n2. 파생 피처 추가\n    - datetime 피처에서 year, month, day, hour, minute, second 피처를 생성\n    - 요일(weekday)피처 생성\n3. 피처 제거\n    - 테스트 데이터에 없고 훈련 데이터에만 잇는 casual과 registered 피처 제거\n    - datetime 피처는 인덱스 역할만 하므로 예측에 도움이 되지 않아 제거\n    - date 피처가 제공하는 정보를 year, month, day 피처에 있으므로 제거\n    - month는 season 피처의 세부 분류로 불 수 있고, 데이터가 지나치게 세분화 되어 있으면 오히려 학습에 방해가 될 수 있음\n    - 막대 그래프 확인 결과 파생 피처인 day는 분별력이 없어 제거\n    - minute과 second에는 아무런 정보가 담겨 있지 않아 제거\n    - 산점도 그래프와 히트맥 확인 결과 windspeed 피처에는 결측값이 많고 대여 수량과의 상관관계가 매우 약해 제거\n4. 이상치 제거\n    - 포인트 플롯 확인 결과 weather가 4인 데이터는 이상치","metadata":{}},{"cell_type":"markdown","source":"## 모델링\n## 기본 모델(LinearRegression)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ndata_path = '/kaggle/input/bike-sharing-demand/'\n\ntrain = pd.read_csv(data_path + 'train.csv')\ntest = pd.read_csv(data_path + 'test.csv')\nsubmission = pd.read_csv(data_path + 'sampleSubmission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.284464Z","iopub.execute_input":"2022-08-05T05:32:41.284744Z","iopub.status.idle":"2022-08-05T05:32:41.334399Z","shell.execute_reply.started":"2022-08-05T05:32:41.284712Z","shell.execute_reply":"2022-08-05T05:32:41.333191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 피처 엔지니어링\n- EDA를 바탕으로 데이터 변환\n- 훈련 데이터와 테스트 데이터에 공통으로 반영해야 하기 때문에 두 데이터를 합치고 피처 엔지니어링 끝나면 나눠주기\n\n\n#### 이상치 제거\n- 포인트 플롯에서 확인한 결과 훈련 데이터에서 weather가 4인 데이터(폭우, 폭설이 내리는 날 저녁 6시에 대여)는 이상치","metadata":{}},{"cell_type":"code","source":"# 훈련 데이터에서 weather가 4가 아닌 데이터만 추출\ntrain = train[train['weather'] != 4]","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.336446Z","iopub.execute_input":"2022-08-05T05:32:41.336717Z","iopub.status.idle":"2022-08-05T05:32:41.344384Z","shell.execute_reply.started":"2022-08-05T05:32:41.336685Z","shell.execute_reply":"2022-08-05T05:32:41.343277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 데이터 합치기\n- 훈련 데이터와 테스트 데이터에 같은 피처 엔지니어링을 적용하기 위해 두 데이터를 하나로 합친 후 분리","metadata":{}},{"cell_type":"code","source":"train.shape, test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.345837Z","iopub.execute_input":"2022-08-05T05:32:41.346311Z","iopub.status.idle":"2022-08-05T05:32:41.357217Z","shell.execute_reply.started":"2022-08-05T05:32:41.346276Z","shell.execute_reply":"2022-08-05T05:32:41.355719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data = pd.concat([train, test], ignore_index = True)\nall_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.361398Z","iopub.execute_input":"2022-08-05T05:32:41.361649Z","iopub.status.idle":"2022-08-05T05:32:41.372740Z","shell.execute_reply.started":"2022-08-05T05:32:41.361622Z","shell.execute_reply":"2022-08-05T05:32:41.371862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.374340Z","iopub.execute_input":"2022-08-05T05:32:41.374854Z","iopub.status.idle":"2022-08-05T05:32:41.400125Z","shell.execute_reply.started":"2022-08-05T05:32:41.374788Z","shell.execute_reply":"2022-08-05T05:32:41.398985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 파생 피처 추가","metadata":{}},{"cell_type":"code","source":"from datetime import datetime\n\n# 날짜 피처 생성\nall_data['date'] = all_data['datetime'].apply(lambda x: x.split()[0])\n\n# 연도 피처 생성\nall_data['year'] = all_data['datetime'].apply(lambda x: x.split()[0].split('-')[0])\n\n# 월 피처 생성\nall_data['month'] = all_data['datetime'].apply(lambda x: x.split()[0].split('-')[1])\n\n# 시 피처 생성\nall_data['hour'] = all_data['datetime'].apply(lambda x: x.split()[1].split(':')[0])\n\n# 요일 피처 생성\nall_data['weekday'] = all_data['date'].apply(lambda date_string: \n                                             datetime.strptime(date_string, \"%Y-%m-%d\").weekday())","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.401769Z","iopub.execute_input":"2022-08-05T05:32:41.402144Z","iopub.status.idle":"2022-08-05T05:32:41.623966Z","shell.execute_reply.started":"2022-08-05T05:32:41.402103Z","shell.execute_reply":"2022-08-05T05:32:41.623264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 필요 없는 피처 제거","metadata":{}},{"cell_type":"code","source":"drop_features = ['casual', 'registered', 'datetime', 'date', 'windspeed', 'month']\n\nall_data = all_data.drop(drop_features, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.625081Z","iopub.execute_input":"2022-08-05T05:32:41.626061Z","iopub.status.idle":"2022-08-05T05:32:41.640673Z","shell.execute_reply.started":"2022-08-05T05:32:41.626013Z","shell.execute_reply":"2022-08-05T05:32:41.639483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 데이터 나누기","metadata":{}},{"cell_type":"code","source":"# 훈련 데이터와 테스트 데이터 나누기\nX_train = all_data[~pd.isnull(all_data['count'])]\nX_test = all_data[pd.isnull(all_data['count'])]","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.642579Z","iopub.execute_input":"2022-08-05T05:32:41.643090Z","iopub.status.idle":"2022-08-05T05:32:41.651928Z","shell.execute_reply.started":"2022-08-05T05:32:41.643041Z","shell.execute_reply":"2022-08-05T05:32:41.650944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 타깃값 count 분리\n\nX_train = X_train.drop(['count'], axis=1)\nX_test = X_test.drop(['count'], axis=1)\n\ny = train['count']","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.653640Z","iopub.execute_input":"2022-08-05T05:32:41.654082Z","iopub.status.idle":"2022-08-05T05:32:41.666004Z","shell.execute_reply.started":"2022-08-05T05:32:41.654047Z","shell.execute_reply":"2022-08-05T05:32:41.664990Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.667621Z","iopub.execute_input":"2022-08-05T05:32:41.668002Z","iopub.status.idle":"2022-08-05T05:32:41.689393Z","shell.execute_reply.started":"2022-08-05T05:32:41.667963Z","shell.execute_reply":"2022-08-05T05:32:41.688217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 평가지표 계산 함수 작성","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef rmsle(y_true, y_pred, convert_exp=True):\n    # 지수변환\n    if convert_exp:\n        y_true = np.exp(y_true)\n        y_pred = np.exp(y_pred)\n    \n    # 로그변환 후 결측값을 0으로 변환\n    log_true = np.nan_to_num(np.log(y_true+1))\n    log_pred = np.nan_to_num(np.log(y_pred+1))\n    \n    # RMSLE 계산\n    output = np.sqrt(np.mean((log_true - log_pred)**2))\n    return output","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.691177Z","iopub.execute_input":"2022-08-05T05:32:41.691751Z","iopub.status.idle":"2022-08-05T05:32:41.701324Z","shell.execute_reply.started":"2022-08-05T05:32:41.691702Z","shell.execute_reply":"2022-08-05T05:32:41.700193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"$$\\sqrt{{1\\over{n}} \\sum_{i=1}^n (log((y_i + 1) - log(\\hat{y}_i + 1))^2}$$","metadata":{}},{"cell_type":"markdown","source":"### 모델 훈련","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LinearRegression\n\nlinear_reg_model = LinearRegression()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.702728Z","iopub.execute_input":"2022-08-05T05:32:41.703032Z","iopub.status.idle":"2022-08-05T05:32:41.786845Z","shell.execute_reply.started":"2022-08-05T05:32:41.703001Z","shell.execute_reply":"2022-08-05T05:32:41.786144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"log_y = np.log(y) # 타깃값 로그변환\nlinear_reg_model.fit(X_train, log_y) # 모델 훈련","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.788224Z","iopub.execute_input":"2022-08-05T05:32:41.788647Z","iopub.status.idle":"2022-08-05T05:32:41.814267Z","shell.execute_reply.started":"2022-08-05T05:32:41.788600Z","shell.execute_reply":"2022-08-05T05:32:41.813108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 선형 회귀 모델을 훈련한다는 것은 독립변수(피처)인 X_train과 종속변수(타깃값)인 log_y에 대응하는 최적의 선형 회귀 계수를 구한다는 의미\n\n\n$$Y = \\theta_0 + \\theta_1x_1 + \\theta_2x_2 + \\theta_3x_3$$\n\n- 독립변수 $x_1$, $x_2$, $x_3$와 종속변수 $Y$를 활용하여 최적의 선형 회귀계수 $\\theta_0$, $\\theta_1$, $\\theta_2$, $\\theta_3$을 구하는 과정","metadata":{}},{"cell_type":"markdown","source":"### 모델 성능 검증","metadata":{}},{"cell_type":"code","source":"preds = linear_reg_model.predict(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.816360Z","iopub.execute_input":"2022-08-05T05:32:41.817022Z","iopub.status.idle":"2022-08-05T05:32:41.854474Z","shell.execute_reply.started":"2022-08-05T05:32:41.816960Z","shell.execute_reply":"2022-08-05T05:32:41.853273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'선형 회귀의 RMSLE 값 : {rmsle(log_y, preds, True):.4f}')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.856550Z","iopub.execute_input":"2022-08-05T05:32:41.857205Z","iopub.status.idle":"2022-08-05T05:32:41.870538Z","shell.execute_reply.started":"2022-08-05T05:32:41.857144Z","shell.execute_reply":"2022-08-05T05:32:41.869018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 예측 및 결과 제출","metadata":{}},{"cell_type":"code","source":"linear_reg_preds = linear_reg_model.predict(X_test)\n\nsubmission['count'] = np.exp(linear_reg_preds)\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.872546Z","iopub.execute_input":"2022-08-05T05:32:41.873221Z","iopub.status.idle":"2022-08-05T05:32:41.951139Z","shell.execute_reply.started":"2022-08-05T05:32:41.873145Z","shell.execute_reply":"2022-08-05T05:32:41.949900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Score: 1.02142","metadata":{}},{"cell_type":"markdown","source":"## 성능 개선 1 : 릿지 회귀 모델\n### 하이퍼파라미터 최적화(그리드 서치)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:44:58.433057Z","iopub.execute_input":"2022-08-05T04:44:58.433405Z","iopub.status.idle":"2022-08-05T04:44:58.438331Z","shell.execute_reply.started":"2022-08-05T04:44:58.433371Z","shell.execute_reply":"2022-08-05T04:44:58.436938Z"}}},{"cell_type":"code","source":"from sklearn.linear_model import Ridge\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn import metrics\n\nridge_model = Ridge()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.953350Z","iopub.execute_input":"2022-08-05T05:32:41.954092Z","iopub.status.idle":"2022-08-05T05:32:41.961203Z","shell.execute_reply.started":"2022-08-05T05:32:41.954026Z","shell.execute_reply":"2022-08-05T05:32:41.960063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 하이퍼파라미터 값 목록\nridge_params = {\n    'max_iter': [3000],\n    'alpha': [0.1, 1, 2, 3, 4, 10, 30],\n}\n\n# 교차 검증용 평가 함수(RMSLE 점수 계산)\nrmsle_scorer = metrics.make_scorer(rmsle, greater_is_better=False)\n\n# 그리드 서치 객체 생성\ngridsearch_ridge_model = GridSearchCV(\n    estimator=ridge_model,\n    param_grid=ridge_params,\n    scoring=rmsle_scorer,\n    cv=5\n)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.963262Z","iopub.execute_input":"2022-08-05T05:32:41.964110Z","iopub.status.idle":"2022-08-05T05:32:41.977335Z","shell.execute_reply.started":"2022-08-05T05:32:41.964045Z","shell.execute_reply":"2022-08-05T05:32:41.975963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"log_y = np.log(y)\ngridsearch_ridge_model.fit(X_train, log_y)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:41.985727Z","iopub.execute_input":"2022-08-05T05:32:41.986925Z","iopub.status.idle":"2022-08-05T05:32:43.337272Z","shell.execute_reply.started":"2022-08-05T05:32:41.986852Z","shell.execute_reply":"2022-08-05T05:32:43.336145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('최적 하이퍼파라미터 :', gridsearch_ridge_model.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:43.339462Z","iopub.execute_input":"2022-08-05T05:32:43.340166Z","iopub.status.idle":"2022-08-05T05:32:43.350468Z","shell.execute_reply.started":"2022-08-05T05:32:43.340107Z","shell.execute_reply":"2022-08-05T05:32:43.349098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 성능 검증","metadata":{}},{"cell_type":"code","source":"# 예측\npreds = gridsearch_ridge_model.best_estimator_.predict(X_train)\n\n# 평가\nprint(f'릿지 회귀 RMSLE 값 : {rmsle(log_y, preds, True):.4f}')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:43.353719Z","iopub.execute_input":"2022-08-05T05:32:43.356119Z","iopub.status.idle":"2022-08-05T05:32:43.427650Z","shell.execute_reply.started":"2022-08-05T05:32:43.356055Z","shell.execute_reply":"2022-08-05T05:32:43.426331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 이전 모델과 큰 차이가 없음","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:58:17.853554Z","iopub.execute_input":"2022-08-05T04:58:17.853909Z","iopub.status.idle":"2022-08-05T04:58:17.859973Z","shell.execute_reply.started":"2022-08-05T04:58:17.853858Z","shell.execute_reply":"2022-08-05T04:58:17.859066Z"}}},{"cell_type":"markdown","source":"## 성능 개선 2 : 라쏘 회귀 모델","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:58:33.201702Z","iopub.execute_input":"2022-08-05T04:58:33.202195Z","iopub.status.idle":"2022-08-05T04:58:33.206929Z","shell.execute_reply.started":"2022-08-05T04:58:33.202165Z","shell.execute_reply":"2022-08-05T04:58:33.205739Z"}}},{"cell_type":"code","source":"from sklearn.linear_model import Lasso\n\nlasso_model = Lasso()\n\nlasso_alpha = 1 / np.array([0.1, 1, 2, 3, 4, 10, 30, 100, 200, \n                            300, 400, 800, 900, 1000])\nlasso_params = {\n    'max_iter': [3000],\n    'alpha': lasso_alpha\n}\n\ngridsearch_lasso_model = GridSearchCV(\n    estimator=lasso_model,\n    param_grid=lasso_params,\n    scoring=rmsle_scorer,\n    cv=5\n)\n\nlog_y = np.log(y)\ngridsearch_lasso_model.fit(X_train, log_y)\n\nprint('최적 하이퍼파라미터 : ', gridsearch_lasso_model.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:43.430757Z","iopub.execute_input":"2022-08-05T05:32:43.432804Z","iopub.status.idle":"2022-08-05T05:32:47.318290Z","shell.execute_reply.started":"2022-08-05T05:32:43.432741Z","shell.execute_reply":"2022-08-05T05:32:47.317118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 성능 검증","metadata":{}},{"cell_type":"code","source":"# 예측\npreds = gridsearch_lasso_model.best_estimator_.predict(X_train)\n\n# 평가\nprint(f'라쏘 회귀 RMSLE 값 : {rmsle(log_y, preds, True):.4f}')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:47.320092Z","iopub.execute_input":"2022-08-05T05:32:47.327914Z","iopub.status.idle":"2022-08-05T05:32:47.374173Z","shell.execute_reply.started":"2022-08-05T05:32:47.327784Z","shell.execute_reply":"2022-08-05T05:32:47.372855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 별 다른 개선 없음","metadata":{}},{"cell_type":"markdown","source":"## 성능 개선 3 : 랜덤 포레스트 회귀 모델","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestRegressor\n\nrf_model = RandomForestRegressor()\n\nrf_params = {\n    'random_state': [42],\n    'n_estimators': [100, 120, 140, 200],\n}\n\ngridsearch_rf = GridSearchCV(\n    estimator=rf_model,\n    param_grid=rf_params,\n    scoring=rmsle_scorer,\n    cv=5\n)\n\nlog_y = np.log(y)\ngridsearch_rf.fit(X_train, log_y)\n\nprint('최적 하이퍼파라미터 : ', gridsearch_rf.best_params_)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:32:47.377969Z","iopub.execute_input":"2022-08-05T05:32:47.379103Z","iopub.status.idle":"2022-08-05T05:34:13.450013Z","shell.execute_reply.started":"2022-08-05T05:32:47.379048Z","shell.execute_reply":"2022-08-05T05:34:13.448984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 성능 검증","metadata":{}},{"cell_type":"code","source":"# 예측\npreds = gridsearch_rf.best_estimator_.predict(X_train)\n\n# 평가\nprint(f'랜덤 포레스트 회귀 RMSLE 값 : {rmsle(log_y, preds, True):.4f}')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:34:13.451483Z","iopub.execute_input":"2022-08-05T05:34:13.451745Z","iopub.status.idle":"2022-08-05T05:34:13.760570Z","shell.execute_reply.started":"2022-08-05T05:34:13.451712Z","shell.execute_reply":"2022-08-05T05:34:13.759413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 성능 향상 성공","metadata":{}},{"cell_type":"markdown","source":"### 예측 및 결과 제출","metadata":{}},{"cell_type":"code","source":"rf_reg_preds = gridsearch_rf.best_estimator_.predict(X_test)\n\nsubmission['count'] = np.exp(rf_reg_preds)\nsubmission.to_csv('submission2.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:34:13.762303Z","iopub.execute_input":"2022-08-05T05:34:13.763229Z","iopub.status.idle":"2022-08-05T05:34:13.998228Z","shell.execute_reply.started":"2022-08-05T05:34:13.763179Z","shell.execute_reply":"2022-08-05T05:34:13.996971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Score: 0.39567","metadata":{}}]}