{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 프로젝트 개요\n\n- 강의명 : 2022년 K-디지털 직업훈련(Training) 사업 - AI데이터플랫폼을 활용한 - 빅데이터 분석전문가 과정\n- 교과목명 : 빅데이터 분석 및 시각화, AI개발 기초, 인공지능 프로그래밍\n- 프로젝트 주제 : 캐글 대회 Bike Sharing Demand 데이터를 활용한 수요 예측 대회\n- 프로젝트 마감일 : 2022년 7월 19일 화요일\n- 수강생명 : 오세영","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-14T01:52:11.568960Z","iopub.execute_input":"2022-07-14T01:52:11.569480Z","iopub.status.idle":"2022-07-14T01:52:11.580966Z","shell.execute_reply.started":"2022-07-14T01:52:11.569444Z","shell.execute_reply":"2022-07-14T01:52:11.579089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## step 01. 필수 라이브러리 불러오기","metadata":{}},{"cell_type":"code","source":"import pandas as pd # 데이터 가공\nimport numpy as np # 수치 연산\nimport matplotlib as mpl # 시각화 \nimport matplotlib.pyplot as plt \nimport seaborn as sns # 시각화 \nimport sklearn # 머신러닝\n\n# 버전 확인\nprint(\"pandas version :\", pd.__version__)\nprint(\"numpy version :\", np.__version__)\nprint(\"matplotlib :\", mpl.__version__)\nprint(\"seaborn :\", sns.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:11.657294Z","iopub.execute_input":"2022-07-14T01:52:11.657671Z","iopub.status.idle":"2022-07-14T01:52:11.666736Z","shell.execute_reply.started":"2022-07-14T01:52:11.657643Z","shell.execute_reply":"2022-07-14T01:52:11.665202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## step 02. 데이터 불러오기","metadata":{}},{"cell_type":"code","source":"DATA_PATH = '/kaggle/input/bike-sharing-demand/'\n\ntrain = pd.read_csv(DATA_PATH + 'train.csv') # 훈련 데이터\ntest = pd.read_csv(DATA_PATH + 'test.csv')   # 테스트 데이터\nsubmission = pd.read_csv(DATA_PATH + 'sampleSubmission.csv') # 제출 샘플 데이터\n\ntrain.shape, test.shape, submission.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:11.743871Z","iopub.execute_input":"2022-07-14T01:52:11.744337Z","iopub.status.idle":"2022-07-14T01:52:11.813503Z","shell.execute_reply.started":"2022-07-14T01:52:11.744300Z","shell.execute_reply":"2022-07-14T01:52:11.812319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## step 03. 데이터 확인하기","metadata":{}},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:11.819447Z","iopub.execute_input":"2022-07-14T01:52:11.819853Z","iopub.status.idle":"2022-07-14T01:52:11.837810Z","shell.execute_reply.started":"2022-07-14T01:52:11.819824Z","shell.execute_reply":"2022-07-14T01:52:11.836431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:11.892976Z","iopub.execute_input":"2022-07-14T01:52:11.893471Z","iopub.status.idle":"2022-07-14T01:52:11.915199Z","shell.execute_reply.started":"2022-07-14T01:52:11.893432Z","shell.execute_reply":"2022-07-14T01:52:11.913946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 결측치 확인하기","metadata":{}},{"cell_type":"code","source":"train.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:11.966461Z","iopub.execute_input":"2022-07-14T01:52:11.966949Z","iopub.status.idle":"2022-07-14T01:52:11.982263Z","shell.execute_reply.started":"2022-07-14T01:52:11.966886Z","shell.execute_reply":"2022-07-14T01:52:11.980947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.037852Z","iopub.execute_input":"2022-07-14T01:52:12.038321Z","iopub.status.idle":"2022-07-14T01:52:12.056190Z","shell.execute_reply.started":"2022-07-14T01:52:12.038287Z","shell.execute_reply":"2022-07-14T01:52:12.052141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## step 03. 탐색적 자료 분석 및 데이터 전처리\n- 시각화\n- 날짜 기반\n- train 데이터에 바로 변화를 주면 전처리 시 헷갈릴 수 있다.\n    + train 데이터 복제본을 만들어서 전처리","metadata":{}},{"cell_type":"code","source":"temp_df = train.copy()\ntemp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.107158Z","iopub.execute_input":"2022-07-14T01:52:12.108040Z","iopub.status.idle":"2022-07-14T01:52:12.125859Z","shell.execute_reply.started":"2022-07-14T01:52:12.108002Z","shell.execute_reply":"2022-07-14T01:52:12.124585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- train 데이터와 temp 데이터를 비교했을 때 train 데이터에서 'casual', 'registered' 컬럼을 제거해도 괜찮을 거 같다.","metadata":{}},{"cell_type":"code","source":"temp_df = temp_df.drop(labels = ['casual', 'registered'], axis = 1)\ntemp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.179264Z","iopub.execute_input":"2022-07-14T01:52:12.179723Z","iopub.status.idle":"2022-07-14T01:52:12.200061Z","shell.execute_reply.started":"2022-07-14T01:52:12.179687Z","shell.execute_reply":"2022-07-14T01:52:12.199095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 시각화를 위한 날짜 데이터 처리\n- split() 함수를 이용해서 연도, 월, 일자, 시간, 분, 초 를 분리하기\n","metadata":{}},{"cell_type":"code","source":"temp_df['datetime']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.326350Z","iopub.execute_input":"2022-07-14T01:52:12.326699Z","iopub.status.idle":"2022-07-14T01:52:12.334671Z","shell.execute_reply.started":"2022-07-14T01:52:12.326672Z","shell.execute_reply":"2022-07-14T01:52:12.333959Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 방법 1\n- 깡으로 악으로 분리하기","metadata":{}},{"cell_type":"code","source":"import time \nimport datetime \n\n# 시간 테스트 \nstart_time = time.time()\n\ntemp_df['date'] = temp_df['datetime'].apply(lambda x : x.split()[0])\ntemp_df['year'] = temp_df['datetime'].apply(lambda x : x.split()[0].split('-')[0])\ntemp_df['month'] = temp_df['datetime'].apply(lambda x : x.split()[0].split('-')[1])\ntemp_df['day'] = temp_df['datetime'].apply(lambda x : x.split()[0].split('-')[2])\ntemp_df['hour'] = temp_df['datetime'].apply(lambda x : x.split()[1].split(':')[0])\ntemp_df['minute'] = temp_df['datetime'].apply(lambda x : x.split()[1].split(':')[1])\ntemp_df['second'] = temp_df['datetime'].apply(lambda x : x.split()[1].split(':')[2])\n\nend_time = time.time() \nlambda_ctime = end_time - start_time\n\nprint(\"실행시간 (second) -> \", np.round(lambda_ctime, 3))\ntemp_df[['datetime', 'year', 'month', 'day', 'hour','minute', 'second']]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.405690Z","iopub.execute_input":"2022-07-14T01:52:12.406125Z","iopub.status.idle":"2022-07-14T01:52:12.503526Z","shell.execute_reply.started":"2022-07-14T01:52:12.406095Z","shell.execute_reply":"2022-07-14T01:52:12.502541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 방법 2\n- 데이터프레임 형식으로 변환해서 분리하기\n- 시간이 3배 이상 단축된다...\n    + 만약 1시간짜리 작업이 3배이상 단축된다면? ㄷㄷ","metadata":{}},{"cell_type":"code","source":"import time \nimport datetime \n\n# 시간 테스트 \nstart_time = time.time()\n\ntemp_df['date'] = pd.to_datetime(temp_df['datetime'])\ntemp_df['year'] = temp_df['date'].dt.year\ntemp_df['month'] = temp_df['date'].dt.month\ntemp_df['day'] = temp_df['date'].dt.day\ntemp_df['hour'] = temp_df['date'].dt.hour\ntemp_df['minute'] = temp_df['date'].dt.minute\ntemp_df['second'] = temp_df['date'].dt.second\n\nend_time = time.time() \ndt_ctime = end_time - start_time\n\nprint(\"실행시간 (second) -> \", np.round(dt_ctime, 3))\n\ntemp_df[['datetime', 'year', 'month', 'day', 'hour','minute', 'second']]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.505804Z","iopub.execute_input":"2022-07-14T01:52:12.506843Z","iopub.status.idle":"2022-07-14T01:52:12.574245Z","shell.execute_reply.started":"2022-07-14T01:52:12.506772Z","shell.execute_reply":"2022-07-14T01:52:12.572757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 'minute' 와 'second' 는 '0' 으로 동일한 패턴을 보이므로 삭제하는게 좋을 거 같다.","metadata":{}},{"cell_type":"code","source":"temp_df = temp_df.drop(labels = ['minute','second'], axis = 1)\ntemp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.577111Z","iopub.execute_input":"2022-07-14T01:52:12.578264Z","iopub.status.idle":"2022-07-14T01:52:12.606210Z","shell.execute_reply.started":"2022-07-14T01:52:12.578209Z","shell.execute_reply":"2022-07-14T01:52:12.604980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 날짜 데이터 전처리 완료","metadata":{}},{"cell_type":"code","source":"temp_df.info","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.618860Z","iopub.execute_input":"2022-07-14T01:52:12.620446Z","iopub.status.idle":"2022-07-14T01:52:12.646978Z","shell.execute_reply.started":"2022-07-14T01:52:12.620370Z","shell.execute_reply":"2022-07-14T01:52:12.645678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 12, 1, 2월은 계절상 겨울인데 데이터상으로는 1월이 봄으로 찍혀있다.","metadata":{}},{"cell_type":"markdown","source":"- 계절 데이터 전처리","metadata":{}},{"cell_type":"code","source":"month = temp_df.month\nmonth","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.685544Z","iopub.execute_input":"2022-07-14T01:52:12.688207Z","iopub.status.idle":"2022-07-14T01:52:12.702411Z","shell.execute_reply.started":"2022-07-14T01:52:12.688159Z","shell.execute_reply":"2022-07-14T01:52:12.699685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def badToRight(month):\n    if month in [12,1,2]:\n        return 4\n    elif month in [3,4,5]:\n        return 1\n    elif month in [6,7,8]:\n        return 2\n    elif month in [9,10,11]:\n        return 3\n\ntemp_df['season'] = temp_df.month.apply(badToRight)\ntemp_df['season']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.754826Z","iopub.execute_input":"2022-07-14T01:52:12.755642Z","iopub.status.idle":"2022-07-14T01:52:12.785036Z","shell.execute_reply.started":"2022-07-14T01:52:12.755603Z","shell.execute_reply":"2022-07-14T01:52:12.783613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_df['season'] = temp_df['season'].map({\n    1: 'Spring',\n    2: 'Summer',\n    3: 'Fall',\n    4: 'Winter'\n})\ntemp_df['season']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.823793Z","iopub.execute_input":"2022-07-14T01:52:12.824640Z","iopub.status.idle":"2022-07-14T01:52:12.843670Z","shell.execute_reply.started":"2022-07-14T01:52:12.824591Z","shell.execute_reply":"2022-07-14T01:52:12.842301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 날씨 데이터 문자열로 보기 쉽게 변환\ntemp_df['weather'] = temp_df['weather'].map({\n    1 : 'Clear', \n    2 : 'Few clouds', \n    3 : 'Light Snow, Rain', \n    4 : 'Heavy Snow, Rain'\n})\n\ntemp_df['weather']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.892963Z","iopub.execute_input":"2022-07-14T01:52:12.893605Z","iopub.status.idle":"2022-07-14T01:52:12.906414Z","shell.execute_reply.started":"2022-07-14T01:52:12.893575Z","shell.execute_reply":"2022-07-14T01:52:12.905019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- season 데이터 전처리 완료","metadata":{}},{"cell_type":"markdown","source":"- 요일 데이터 추출","metadata":{}},{"cell_type":"code","source":"temp_df['weekday'] = temp_df['date'].dt.day_name()\ntemp_df['weekday']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:12.962151Z","iopub.execute_input":"2022-07-14T01:52:12.962752Z","iopub.status.idle":"2022-07-14T01:52:12.977407Z","shell.execute_reply.started":"2022-07-14T01:52:12.962710Z","shell.execute_reply":"2022-07-14T01:52:12.975818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 평일, 주말 분리해보기","metadata":{}},{"cell_type":"code","source":"week = temp_df['weekday']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.030399Z","iopub.execute_input":"2022-07-14T01:52:13.030777Z","iopub.status.idle":"2022-07-14T01:52:13.036776Z","shell.execute_reply.started":"2022-07-14T01:52:13.030747Z","shell.execute_reply":"2022-07-14T01:52:13.035271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fx(week):\n    if week in ['Saturday', 'Sunday']:\n        return 1\n    else:\n        return 0\n\ntemp_df['weekend'] = temp_df['weekday'].apply(fx)\nprint(temp_df['weekend'])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.099477Z","iopub.execute_input":"2022-07-14T01:52:13.102291Z","iopub.status.idle":"2022-07-14T01:52:13.116729Z","shell.execute_reply.started":"2022-07-14T01:52:13.102243Z","shell.execute_reply":"2022-07-14T01:52:13.115960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.168360Z","iopub.execute_input":"2022-07-14T01:52:13.168987Z","iopub.status.idle":"2022-07-14T01:52:13.185668Z","shell.execute_reply.started":"2022-07-14T01:52:13.168955Z","shell.execute_reply":"2022-07-14T01:52:13.184889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 평일, 주말 분리 완료","metadata":{}},{"cell_type":"markdown","source":"## step 04. 데이터 시각화\n- 종속변수(여기서는 'count')를 그래프로 시각화","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\n\nfig.set_size_inches(10, 6)\n\nsns.histplot(train['count'],stat='density')\n\n# 옵션 제목\nax.set_title('Normal Graph')\n\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.236972Z","iopub.execute_input":"2022-07-14T01:52:13.237507Z","iopub.status.idle":"2022-07-14T01:52:13.525468Z","shell.execute_reply.started":"2022-07-14T01:52:13.237480Z","shell.execute_reply":"2022-07-14T01:52:13.524036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 뭔가 정규분포 그래프 같지는 않음. 로그변환을 해보자.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nfig.set_size_inches(10, 6)\nsns.histplot(np.log(train['count']), stat='density')\n\nax.set_title(\"Log Transformed Graph\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.528618Z","iopub.execute_input":"2022-07-14T01:52:13.529250Z","iopub.status.idle":"2022-07-14T01:52:13.812736Z","shell.execute_reply.started":"2022-07-14T01:52:13.529213Z","shell.execute_reply":"2022-07-14T01:52:13.810320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 그나마 정규분포 그래프와 비슷하다.","metadata":{}},{"cell_type":"code","source":"temp_df = temp_df.drop(labels=['datetime', 'atemp'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.815872Z","iopub.execute_input":"2022-07-14T01:52:13.816493Z","iopub.status.idle":"2022-07-14T01:52:13.838379Z","shell.execute_reply.started":"2022-07-14T01:52:13.816458Z","shell.execute_reply":"2022-07-14T01:52:13.836204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.843173Z","iopub.execute_input":"2022-07-14T01:52:13.843715Z","iopub.status.idle":"2022-07-14T01:52:13.875001Z","shell.execute_reply.started":"2022-07-14T01:52:13.843663Z","shell.execute_reply":"2022-07-14T01:52:13.872274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# count와 각종 독립변수의 관계를 상관계수로 보자\ncorrMat = temp_df[['count','season','holiday','workingday','weather',\n                   'temp', 'humidity', 'windspeed', 'hour','weekday', 'weekend']].corr()\ncorrMat","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.877095Z","iopub.execute_input":"2022-07-14T01:52:13.877525Z","iopub.status.idle":"2022-07-14T01:52:13.906929Z","shell.execute_reply.started":"2022-07-14T01:52:13.877488Z","shell.execute_reply":"2022-07-14T01:52:13.905767Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 상관계수 시각화\nfig, ax = plt.subplots(figsize=(15, 10))\nax = sns.heatmap(corrMat, annot=True, fmt = \".3g\")","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:13.908774Z","iopub.execute_input":"2022-07-14T01:52:13.909259Z","iopub.status.idle":"2022-07-14T01:52:14.461875Z","shell.execute_reply.started":"2022-07-14T01:52:13.909223Z","shell.execute_reply":"2022-07-14T01:52:14.459722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 그래프로 시각화","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows = 2, ncols = 2)\nfig.subplots_adjust(hspace = 0.3)\nfig.set_size_inches(15, 9)\n\nsns.barplot(x = 'year', y = 'count', data = temp_df, ax=ax[0,0])\nsns.barplot(x = 'month',y = 'count', data = temp_df, ax=ax[0,1])\nsns.barplot(x = 'day', y = 'count', data = temp_df, ax=ax[1,0])\nsns.barplot(x = 'hour', y = 'count', data = temp_df, ax=ax[1,1])\n\nax[0, 0].set_title(\"Rental Amounts by Year\", loc ='left',x = 0.03, y=0.9)\nax[0, 1].set_title(\"Rental Amounts by month\", loc='left',x = 0.03, y=0.9)\nax[1, 0].set_title(\"Rental Amounts by day\", loc='left',x = 0.03, y=0.9)\nax[1, 1].set_title(\"Rental Amounts by hour\",loc='left',x = 0.03, y=0.9)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:14.463422Z","iopub.execute_input":"2022-07-14T01:52:14.463852Z","iopub.status.idle":"2022-07-14T01:52:19.541695Z","shell.execute_reply.started":"2022-07-14T01:52:14.463811Z","shell.execute_reply":"2022-07-14T01:52:19.540458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install sidetable","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:19.543426Z","iopub.execute_input":"2022-07-14T01:52:19.544286Z","iopub.status.idle":"2022-07-14T01:52:48.792217Z","shell.execute_reply.started":"2022-07-14T01:52:19.544241Z","shell.execute_reply":"2022-07-14T01:52:48.788865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport sidetable as stb\nimport io\nimport requests\nprint(temp_df.stb.freq(['season']))","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:48.796711Z","iopub.execute_input":"2022-07-14T01:52:48.803439Z","iopub.status.idle":"2022-07-14T01:52:48.882806Z","shell.execute_reply.started":"2022-07-14T01:52:48.803389Z","shell.execute_reply":"2022-07-14T01:52:48.881533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(9,8))\nplt.rc('font', size = 20)\nplt.rc('legend',fontsize = 20)\nexplode = [0.03, 0.03, 0.03, 0.03]\nlabels = 'spring', 'summer','fall','winter'\ncolors = 'sandybrown', 'springgreen', 'deepskyblue', 'lavenderblush'\nwedgeprops={'width': 0.7, 'edgecolor': 'w', 'linewidth': 5}\n\ntemp_df.groupby('season').sum()['count'].plot.pie(startangle = 200, \n                                                  counterclock=False, \n                                                  explode = explode, \n                                                  colors=colors,\n                                                  shadow = True,\n                                                  wedgeprops=wedgeprops,\n                                                  autopct='%.2f%%');\nplt.title(\"Number of rented bikes share per season\");","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:59:43.116494Z","iopub.execute_input":"2022-07-14T01:59:43.116953Z","iopub.status.idle":"2022-07-14T01:59:43.306172Z","shell.execute_reply.started":"2022-07-14T01:59:43.116915Z","shell.execute_reply":"2022-07-14T01:59:43.304863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows = 3)\nfig.set_size_inches(15, 15)\nfig.tight_layout()\n\nsns.boxplot(x = 'weather',y = 'count', data = temp_df, ax=ax[0])\nsns.boxplot(x = 'windspeed', y = 'count', data = temp_df, ax=ax[1])\nsns.barplot(x = 'humidity', y = 'count', data = temp_df, ax=ax[2])\n\nax[0].set_title(\"Rental Amounts by Weather\", y=0.92)\nax[1].set_title(\"Rental Amounts by Windspeed\", y=0.92)\nax[2].set_title(\"Rental Amounts by Humidity\", y=0.92)\n\nax[1].tick_params(axis = 'x', labelrotation=15)\nax[2].tick_params(axis = 'x', labelrotation=90)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.036507Z","iopub.status.idle":"2022-07-14T01:52:49.048427Z","shell.execute_reply.started":"2022-07-14T01:52:49.048175Z","shell.execute_reply":"2022-07-14T01:52:49.048213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(ncols = 2)\nfig.set_size_inches(10, 6)\nfig.tight_layout()\n\nsns.boxplot(x = 'holiday', y = 'count', data = temp_df, ax = ax[0])\nsns.boxplot(x = 'weekend', y = 'count', data = temp_df, ax = ax[1])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.052270Z","iopub.status.idle":"2022-07-14T01:52:49.052658Z","shell.execute_reply.started":"2022-07-14T01:52:49.052465Z","shell.execute_reply":"2022-07-14T01:52:49.052483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig_dims = (12, 10)\nfig, ax = plt.subplots(nrows=3, figsize=fig_dims)\n\nfig.tight_layout()\n\nsns.regplot(x ='temp', y ='count', data = temp_df, scatter_kws = {'alpha' : 0.3}, marker = \"+\",line_kws = {'color' : 'blue'}, ax = ax[0])\nsns.regplot(x ='humidity', y ='count', data = temp_df, scatter_kws = {'alpha' : 0.3}, color = \"salmon\", marker = \"+\",line_kws = {'color' : 'blue'}, ax = ax[1])\nsns.regplot(x ='windspeed', y ='count', data = temp_df, scatter_kws = {'alpha' : 0.3}, color = \"purple\", marker = \"+\", line_kws = {'color' : 'blue'}, ax = ax[2])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.054002Z","iopub.status.idle":"2022-07-14T01:52:49.054362Z","shell.execute_reply.started":"2022-07-14T01:52:49.054187Z","shell.execute_reply":"2022-07-14T01:52:49.054204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows = 5)\nfig.set_size_inches(12, 18)\n\nsns.pointplot(x = 'hour', y = 'count', hue = 'workingday', data = temp_df, ax = ax[0])\nsns.pointplot(x = 'hour', y = 'count', hue = 'holiday', data = temp_df, ax = ax[1])\nsns.pointplot(x = 'hour', y = 'count', hue = 'weekend', data = temp_df, ax = ax[2])\nsns.pointplot(x = 'hour', y = 'count', hue = 'season', data = temp_df, ax = ax[3])\nsns.pointplot(x = 'hour', y = 'count', hue = 'weather', data = temp_df, ax = ax[4])\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.071988Z","iopub.status.idle":"2022-07-14T01:52:49.072835Z","shell.execute_reply.started":"2022-07-14T01:52:49.072412Z","shell.execute_reply":"2022-07-14T01:52:49.072436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Heavy Snow, Rain 데이터가 심상치 않다.\ntemp_df['weather'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.077704Z","iopub.status.idle":"2022-07-14T01:52:49.078213Z","shell.execute_reply.started":"2022-07-14T01:52:49.077953Z","shell.execute_reply":"2022-07-14T01:52:49.077974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 날씨 데이터의 Heavy Snow, Rain 데이터가 1개밖에 없어서 쓸모없는 수준의 데이터임. 고로 지워주자","metadata":{}},{"cell_type":"code","source":"temp_df = temp_df[temp_df['weather'] != 'Heavy Snow, Rain'].reset_index(drop = True)\ntemp_df['weather'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.090778Z","iopub.status.idle":"2022-07-14T01:52:49.091434Z","shell.execute_reply.started":"2022-07-14T01:52:49.091178Z","shell.execute_reply":"2022-07-14T01:52:49.091200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.092288Z","iopub.status.idle":"2022-07-14T01:52:49.092617Z","shell.execute_reply.started":"2022-07-14T01:52:49.092431Z","shell.execute_reply":"2022-07-14T01:52:49.092447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp_df = temp_df.drop(labels = ['count', 'date'], axis = 1)\ntemp_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.093519Z","iopub.status.idle":"2022-07-14T01:52:49.095191Z","shell.execute_reply.started":"2022-07-14T01:52:49.093779Z","shell.execute_reply":"2022-07-14T01:52:49.093803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"--------------------------------------------------------------------------------------------------------\n--------------------------------------------------------------------------------------------------------\n","metadata":{}},{"cell_type":"markdown","source":"## - 데이터 전처리 총정리\n1. train 데이터에서 'casual', 'registered', 'atemp' 컬럼 제거\n2. train 데이터에서 weather 4인 데이터는 삭제(이상치)\n3. windspeed 0인 데이터와 아닌 데이터로 나누어서 분석 후 0인 데이터에 랜덤포레스트 값 넣기\n    + windspeed가 0인 측정치가 이상하지만 누락시키는 대신 측정된 값 분을 0인 측정치로 대체\n4. 데이터 합치기\n5. 날짜 데이터 분리하기\n6. 계절 데이터 전처리\n7. 평일, 주말 분리해보기\n8. 합쳐진 데이터 셋에서 count 유무로 훈련과 테스트셋을 분리한 후 datetime으로 정렬\n9. 필요없는 컬럼들 지정 후 버리기\n10. 선형 회귀모델 평가\n11. XGBoost 써보기","metadata":{}},{"cell_type":"code","source":"# 데이터 합친 후 drop 하면 에러나서 train 데이터에서 필요없는 컬럼들 제거\n# 'casual', 'registered' 컬럼은 합치면 'count' 컬럼이기 때문에 중복되어서 삭제\n# 'atemp 컬럼도 'temp' 컬럼과 중복되어서 삭제\ntrain = train.drop(labels = ['casual', 'registered', 'atemp'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.105925Z","iopub.status.idle":"2022-07-14T01:52:49.106387Z","shell.execute_reply.started":"2022-07-14T01:52:49.106156Z","shell.execute_reply":"2022-07-14T01:52:49.106175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 상관계수가 낮아서 거의 상관은 없지만, 날씨가 4인 데이터를 이상치로 분류하여 삭제\ntrain = train[train['weather'] != 4].reset_index(drop=True)\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.107786Z","iopub.status.idle":"2022-07-14T01:52:49.108189Z","shell.execute_reply.started":"2022-07-14T01:52:49.107973Z","shell.execute_reply":"2022-07-14T01:52:49.107993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 라이브러리 불러오기\nfrom sklearn import linear_model\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn import ensemble \nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.ensemble import BaggingRegressor\nimport xgboost as xgb\nfrom sklearn.model_selection import train_test_split\n\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import cross_val_score, KFold\nfrom sklearn import preprocessing\n\nimport pickle\nimport re\nimport os\nimport warnings\nwarnings.filterwarnings('ignore') \nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.preprocessing import LabelEncoder\nle = LabelEncoder()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.108852Z","iopub.status.idle":"2022-07-14T01:52:49.110415Z","shell.execute_reply.started":"2022-07-14T01:52:49.109412Z","shell.execute_reply":"2022-07-14T01:52:49.109482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# windspeed 0인 데이터와 아닌 데이터로 나누어서 분석 후 0인 데이터에 랜덤포레스트 값 넣기\n# windspeed가 0인 측정치가 이상하지만 누락시키는 대신 측정된 값 분을 0인 측정치로 대체\nwindspeed_0 = train[train.windspeed==0]\nwindspeed_not0 = train[train.windspeed!=0]\nwindspeed_0_df = windspeed_0.drop(['windspeed', 'count', 'datetime'], axis = 1)\nwindspeed_not0_df = windspeed_not0.drop(['windspeed', 'count', 'datetime'], axis = 1)\n\n# 학습 시킬 Windspeed Series는 따로 빼기\nwindspeed_not0_series = windspeed_not0['windspeed']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.120953Z","iopub.status.idle":"2022-07-14T01:52:49.121638Z","shell.execute_reply.started":"2022-07-14T01:52:49.121401Z","shell.execute_reply":"2022-07-14T01:52:49.121421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 랜덤포레스트 회귀로 학습\nrf = RandomForestRegressor()\nrf.fit(windspeed_not0_df, windspeed_not0_series)\n\npredicted_windspeed_0 = rf.predict(windspeed_0_df)\nwindspeed_0['windspeed'] = predicted_windspeed_0","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.138260Z","iopub.status.idle":"2022-07-14T01:52:49.138914Z","shell.execute_reply.started":"2022-07-14T01:52:49.138607Z","shell.execute_reply":"2022-07-14T01:52:49.138629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 랜덤포레스트로 풍속 0 이었던 데이터에 풍속 0이 아닌 데이터 학습한 걸로 대체\ntrain = pd.concat([windspeed_0,windspeed_not0], axis = 0)\ntrain.datetime = pd.to_datetime(train.datetime,errors = 'coerce')\ntrain = train.sort_values(by=['datetime'])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.155779Z","iopub.status.idle":"2022-07-14T01:52:49.156457Z","shell.execute_reply.started":"2022-07-14T01:52:49.156250Z","shell.execute_reply":"2022-07-14T01:52:49.156270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 데이터 합치기\nall_data = pd.concat([train, test], axis = 0)\nall_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.157309Z","iopub.status.idle":"2022-07-14T01:52:49.157975Z","shell.execute_reply.started":"2022-07-14T01:52:49.157761Z","shell.execute_reply":"2022-07-14T01:52:49.157779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time \nimport datetime \n\n# 시간 테스트 \nstart_time = time.time()\n\nall_data['date'] = pd.to_datetime(all_data['datetime'])\nall_data['year'] = all_data['date'].dt.year\nall_data['month'] = all_data['date'].dt.month\nall_data['day'] = all_data['date'].dt.day\nall_data['hour'] = all_data['date'].dt.hour\n\nend_time = time.time() \ndt_ctime = end_time - start_time\nall_data[['datetime', 'year', 'month', 'day', 'hour']]\n\n# errors = 'coerce' 는 잘못된 분석치는 Nan으로 설정하라는 뜻\nall_data['year'] = pd.to_numeric(all_data.year, errors='coerce')\nall_data['month'] = pd.to_numeric(all_data.month, errors='coerce')\nall_data['day'] = pd.to_numeric(all_data.day, errors='coerce')\nall_data['hour'] = pd.to_numeric(all_data.hour, errors='coerce')\n\nall_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.186929Z","iopub.status.idle":"2022-07-14T01:52:49.188882Z","shell.execute_reply.started":"2022-07-14T01:52:49.188475Z","shell.execute_reply":"2022-07-14T01:52:49.188521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 위에서 만들었던 월별 계절 데이터 전처리 함수 사용\nall_data['season'] = all_data.month.apply(badToRight)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.202823Z","iopub.status.idle":"2022-07-14T01:52:49.210480Z","shell.execute_reply.started":"2022-07-14T01:52:49.210066Z","shell.execute_reply":"2022-07-14T01:52:49.210112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 요일 데이터 추출\nall_data['weekday'] = all_data['date'].dt.day_name()\n\n# 머신러닝 전에 문자를 숫자화\nall_data['weekday']= all_data.weekday.astype('category')\nprint(all_data['weekday'].cat.categories)\n\n#0:Sunday --> 6:Saturday\nall_data.weekday.cat.categories = ['5','1','6','0','4','2','3']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.211968Z","iopub.status.idle":"2022-07-14T01:52:49.212373Z","shell.execute_reply.started":"2022-07-14T01:52:49.212194Z","shell.execute_reply":"2022-07-14T01:52:49.212214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(all_data['weekday'])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.213839Z","iopub.status.idle":"2022-07-14T01:52:49.221966Z","shell.execute_reply.started":"2022-07-14T01:52:49.218559Z","shell.execute_reply":"2022-07-14T01:52:49.220947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 'Saturday'(6), 'Sunday'(0) 데이터를 'weekend' 데이터로 분리\ndef fx(week):\n    if week in ['6', '0']:\n        return 1\n    else:\n        return 0\n\nall_data['weekend'] = all_data['weekday'].apply(fx)\nprint(all_data['weekend'])","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.227451Z","iopub.status.idle":"2022-07-14T01:52:49.228040Z","shell.execute_reply.started":"2022-07-14T01:52:49.227776Z","shell.execute_reply":"2022-07-14T01:52:49.227799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 합친 데이터에서도 풍속 0인 데이터 랜덤포레스트값 넣는 작업 동일하게 진행\ndatawind0 = all_data[all_data['windspeed']==0]\ndatawindnot0 = all_data[all_data['windspeed']!=0]\n\ndatawind0.columns\n\ndatawind0_df = datawind0.drop(['datetime', 'windspeed', 'count', 'atemp','date'], axis = 1)\ndatawindnot0_df = datawindnot0.drop(['datetime', 'windspeed', 'count', 'atemp','date'], axis = 1)\n\ndatawindnot0_series = datawindnot0['windspeed']","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.236309Z","iopub.status.idle":"2022-07-14T01:52:49.240884Z","shell.execute_reply.started":"2022-07-14T01:52:49.237261Z","shell.execute_reply":"2022-07-14T01:52:49.237318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf2 = RandomForestRegressor()\nrf2.fit(datawindnot0_df,datawindnot0_series)\npredicted = rf2.predict(datawind0_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.250976Z","iopub.status.idle":"2022-07-14T01:52:49.251765Z","shell.execute_reply.started":"2022-07-14T01:52:49.251560Z","shell.execute_reply":"2022-07-14T01:52:49.251579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"datawind0['windspeed'] = predicted\n\nall_data = pd.concat([datawind0, datawindnot0], axis = 0)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.256813Z","iopub.status.idle":"2022-07-14T01:52:49.257265Z","shell.execute_reply.started":"2022-07-14T01:52:49.257066Z","shell.execute_reply":"2022-07-14T01:52:49.257086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 'weekday' 컬럼이 카테고리유형이기 때문에 모델링 오류가 났음 --> int 숫자형으로 변환\nall_data = all_data.astype({'weekday' : 'int'})\nall_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.258282Z","iopub.status.idle":"2022-07-14T01:52:49.258634Z","shell.execute_reply.started":"2022-07-14T01:52:49.258446Z","shell.execute_reply":"2022-07-14T01:52:49.258463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 데이터 합친 뒤 중복 데이터 삭제\nall_data = all_data.drop(labels = ['datetime', 'atemp','date','month','day'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.272940Z","iopub.status.idle":"2022-07-14T01:52:49.273449Z","shell.execute_reply.started":"2022-07-14T01:52:49.273241Z","shell.execute_reply":"2022-07-14T01:52:49.273258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# One-Hot Encoding 으로 오차 범위 줄이기\nall_data = pd.get_dummies(all_data).reset_index(drop=True)\nall_data.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.292612Z","iopub.status.idle":"2022-07-14T01:52:49.293434Z","shell.execute_reply.started":"2022-07-14T01:52:49.292924Z","shell.execute_reply":"2022-07-14T01:52:49.292944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 합친 데이터 분리\ntrain = all_data[~pd.isnull(all_data['count'])]\ntest = all_data[pd.isnull(all_data['count'])]\n\ntrain.shape, test.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.297708Z","iopub.status.idle":"2022-07-14T01:52:49.298527Z","shell.execute_reply.started":"2022-07-14T01:52:49.298006Z","shell.execute_reply":"2022-07-14T01:52:49.298321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train.info())\nprint(test.info())","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.300243Z","iopub.status.idle":"2022-07-14T01:52:49.300642Z","shell.execute_reply.started":"2022-07-14T01:52:49.300420Z","shell.execute_reply":"2022-07-14T01:52:49.300444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 종속변수가 정규분포가 아니기 때문에 비교적 정규분포에 가깝게 로그변환을 해준다\ny = train['count']\ny_log = np.log(y)\ntrain = train.drop(['count'], axis = 1)\ntest = test.drop(['count'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.315868Z","iopub.status.idle":"2022-07-14T01:52:49.316530Z","shell.execute_reply.started":"2022-07-14T01:52:49.316269Z","shell.execute_reply":"2022-07-14T01:52:49.316288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 선형회귀 모델\nfrom sklearn.linear_model import LinearRegression\n# train_test_split\n\nlr_model = LinearRegression()\nlr_model.fit(train, y_log)\n\n# 모형 예측\nlr_preds = lr_model.predict(test)\nlr_preds[:10]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.317564Z","iopub.status.idle":"2022-07-14T01:52:49.318202Z","shell.execute_reply.started":"2022-07-14T01:52:49.317975Z","shell.execute_reply":"2022-07-14T01:52:49.317995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# XGBoost 에러 열심히 잡고 하이퍼파라미터 만져봤지만 모델링 점수결과 3.08295.. 주석처리\n# xgr=xgb.XGBRegressor(max_depth=8,min_child_weight=6,gamma=0.4,colsample_bytree=0.6,subsample=0.6)\n# xgr.fit(train,y_log)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-14T01:52:49.332107Z","iopub.status.idle":"2022-07-14T01:52:49.332566Z","shell.execute_reply.started":"2022-07-14T01:52:49.332376Z","shell.execute_reply":"2022-07-14T01:52:49.332394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# y_output=xgr.predict(test)\n# y_output","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.334041Z","iopub.status.idle":"2022-07-14T01:52:49.334420Z","shell.execute_reply.started":"2022-07-14T01:52:49.334234Z","shell.execute_reply":"2022-07-14T01:52:49.334250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_final = np.exp(lr_preds)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.345989Z","iopub.status.idle":"2022-07-14T01:52:49.346589Z","shell.execute_reply.started":"2022-07-14T01:52:49.346314Z","shell.execute_reply":"2022-07-14T01:52:49.346334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(y_final)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.347876Z","iopub.status.idle":"2022-07-14T01:52:49.348290Z","shell.execute_reply.started":"2022-07-14T01:52:49.348092Z","shell.execute_reply":"2022-07-14T01:52:49.348112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission['count']=y_final","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.358086Z","iopub.status.idle":"2022-07-14T01:52:49.361435Z","shell.execute_reply.started":"2022-07-14T01:52:49.358386Z","shell.execute_reply":"2022-07-14T01:52:49.358406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T01:52:49.363887Z","iopub.status.idle":"2022-07-14T01:52:49.365412Z","shell.execute_reply.started":"2022-07-14T01:52:49.365079Z","shell.execute_reply":"2022-07-14T01:52:49.365114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 코드 참고 및 개념 참고\n- https://dsbook.tistory.com/328\n- https://www.kaggle.com/code/kwonyoung234/for-beginner\n- https://www.kaggle.com/code/suvinlee/kernel77d7dbc37e\n    + 휴일 만들기\n- https://www.kaggle.com/code/carolineecc/xg-boost-random-forest-ridge-lasso-regression\n- https://www.kaggle.com/code/suvinlee/bike-sharing-demand-note\n- https://www.kaggle.com/code/suvinlee/kernel77d7dbc37e\n    + 크리스마스도 분류한다는걸 참고\n","metadata":{}}]}