{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Machine Learning 프로젝트 수행을 위한 코드 구조화\n\n`(분류, 회귀 Task)`","metadata":{"id":"8f3119aa"}},{"cell_type":"markdown","source":"- ML project를 위해서 사용하는 템플릿 코드를 만듭니다.\n\n1. **필요한 라이브러리와 데이터를 불러옵니다.**\n\n\n2. **EDA를 수행합니다.** 이 때 EDA의 목적은 풀어야하는 문제를 위해서 수행됩니다.\n\n\n3. **전처리를 수행합니다.** 이 때 중요한건 **feature engineering**을 어떻게 하느냐 입니다.\n\n\n4. **데이터 분할을 합니다.** 이 때 train data와 test data 간의 분포 차이가 없는지 확인합니다.\n\n\n5. **학습을 진행합니다.** 어떤 모델을 사용하여 학습할지 정합니다. 성능이 잘 나오는 GBM을 추천합니다.\n\n\n6. **hyper-parameter tuning을 수행합니다.** 원하는 목표 성능이 나올 때 까지 진행합니다. 검증 단계를 통해 지속적으로 **overfitting이 되지 않게 주의**하세요.\n\n\n7. **최종 테스트를 진행합니다.** 데이터 분석 대회 포맷에 맞는 submission 파일을 만들어서 성능을 확인해보세요.","metadata":{"id":"9f7610ea"}},{"cell_type":"markdown","source":"## 1. 라이브러리, 데이터 불러오기","metadata":{"id":"bd2f7530"}},{"cell_type":"code","source":"# 데이터분석 4종 세트\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# 모델들, 성능 평가\n# (저는 일반적으로 정형데이터로 머신러닝 분석할 때는 이 2개 모델은 그냥 돌려봅니다. 특히 RF가 테스트하기 좋습니다.)\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.ensemble import RandomForestRegressor\nfrom lightgbm.sklearn import LGBMClassifier\nfrom lightgbm.sklearn import LGBMRegressor\n\n# 상관관계 분석, VIF : 다중공선성 제거\nfrom statsmodels.stats.outliers_influence import variance_inflation_factor\n\n# KFold(CV), partial : optuna를 사용하기 위함\nfrom sklearn.model_selection import KFold\nfrom functools import partial\n\n# hyper-parameter tuning을 위한 라이브러리, optuna\nimport optuna","metadata":{"id":"125fc348","execution":{"iopub.status.busy":"2022-08-05T04:21:48.778002Z","iopub.execute_input":"2022-08-05T04:21:48.778529Z","iopub.status.idle":"2022-08-05T04:21:51.599759Z","shell.execute_reply.started":"2022-08-05T04:21:48.778433Z","shell.execute_reply":"2022-08-05T04:21:51.598474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# flag setting\ndata_reducing = True ## memory reducing technique\nfeature_reducing = True ## feature extraction (curse of dimensionality)","metadata":{"id":"0e4d49ce","execution":{"iopub.status.busy":"2022-08-05T04:21:51.602466Z","iopub.execute_input":"2022-08-05T04:21:51.603009Z","iopub.status.idle":"2022-08-05T04:21:51.609383Z","shell.execute_reply.started":"2022-08-05T04:21:51.602974Z","shell.execute_reply":"2022-08-05T04:21:51.607942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 데이터를 불러옵니다.\ntrain_cat = pd.read_csv('../input/bosch-production-line-performance/train_categorical.csv.zip',\n                       nrows=1000)\ntrain_date = pd.read_csv('../input/bosch-production-line-performance/train_date.csv.zip',\n                         nrows=1000)\ntrain_num = pd.read_csv('../input/bosch-production-line-performance/train_numeric.csv.zip',\n                       nrows=1000)","metadata":{"id":"3615c24a","execution":{"iopub.status.busy":"2022-08-05T04:21:51.610750Z","iopub.execute_input":"2022-08-05T04:21:51.611325Z","iopub.status.idle":"2022-08-05T04:21:52.536686Z","shell.execute_reply.started":"2022-08-05T04:21:51.611291Z","shell.execute_reply":"2022-08-05T04:21:52.535602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 데이터를 일부를 잘라서(chunk) 불러온 뒤에, 확인하고 싶은 정보를 계산하고 다음 chunk을 가져오는 방식을 사용합니다.","metadata":{}},{"cell_type":"code","source":"missings = train_num.count() * 0\nsums = train_num.count() * 0\ndata_len = 0\n\nfor chunk in pd.read_csv('../input/bosch-production-line-performance/train_numeric.csv.zip',\n                        chunksize=100000): # (1184737 // 100000)+1 iterations\n    missings = missings + chunk.isnull().sum()\n    sums = sums + chunk.sum()\n    data_len = data_len + len(chunk)\n\nmissings = missings / data_len\nsums = sums / data_len","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:21:52.538166Z","iopub.execute_input":"2022-08-05T04:21:52.538850Z","iopub.status.idle":"2022-08-05T04:24:01.225578Z","shell.execute_reply.started":"2022-08-05T04:21:52.538802Z","shell.execute_reply":"2022-08-05T04:24:01.224466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## 결측치가 50%가 넘는 모든 numeric columns을 제거하고 싶습니다.\n\n## TO-DO : masking을 사용해서 결측치가 50%가 넘는 column을 제거해주세요!\n# train_num.drop(columns=train_num.columns[missings > 0.5])\n# train_num[train_num.columns[missings <= 0.5]]\n\nusecols = train_num.columns[missings <= 0.5] # 결측치가 50%를 넘지 않는 column들.\n\ntrain_num = pd.read_csv('../input/bosch-production-line-performance/train_numeric.csv.zip',\n                       usecols=usecols)\ntrain_num","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:24:01.228525Z","iopub.execute_input":"2022-08-05T04:24:01.228968Z","iopub.status.idle":"2022-08-05T04:24:53.413208Z","shell.execute_reply.started":"2022-08-05T04:24:01.228935Z","shell.execute_reply":"2022-08-05T04:24:53.411898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_num = train_num.fillna(0.0) # 비어있는 값들은 측정이 아예 안된, 즉 지나가지 않은 공정 과정이라서.","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:24:53.414929Z","iopub.execute_input":"2022-08-05T04:24:53.415688Z","iopub.status.idle":"2022-08-05T04:24:54.532960Z","shell.execute_reply.started":"2022-08-05T04:24:53.415636Z","shell.execute_reply":"2022-08-05T04:24:54.531851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def reduce_mem_usage(props, print_mode=False):\n    start_mem_usg = props.memory_usage().sum() / 1024**2 \n    if print_mode:\n        print(\"Memory usage of properties dataframe is :\",start_mem_usg,\" MB\")\n    NAlist = [] # Keeps track of columns that have missing values filled in.\n    for col in props.columns:\n        if props[col].dtype != object:  # Exclude strings\n            \n            # Print current column type\n            if print_mode:\n                print(\"******************************\")\n                print(\"Column: \",col)\n                print(\"dtype before: \",props[col].dtype)\n\n            # make variables for Int, max and min\n            IsInt = False\n            mx = props[col].max()\n            mn = props[col].min()\n           # Integer does not support NA, therefore, NA needs to be filled\n            if not np.isfinite(props[col]).all(): \n                NAlist.append(col)\n                props[col].fillna(mn-1,inplace=True)  \n                   \n            # test if column can be converted to an integer\n            asint = props[col].fillna(0).astype(np.int64)\n            result = (props[col] - asint)\n            result = result.sum()\n            if result > -0.01 and result < 0.01:\n                IsInt = True\n\n            \n            # Make Integer/unsigned Integer datatypes\n            if IsInt:\n                if mn >= 0:\n                    if mx < 255:\n                        props[col] = props[col].astype(np.uint8)\n                    elif mx < 65535:\n                        props[col] = props[col].astype(np.uint16)\n                    elif mx < 4294967295:\n                        props[col] = props[col].astype(np.uint32)\n                    else:\n                        props[col] = props[col].astype(np.uint64)\n                else:\n                    if mn > np.iinfo(np.int8).min and mx < np.iinfo(np.int8).max:\n                        props[col] = props[col].astype(np.int8)\n                    elif mn > np.iinfo(np.int16).min and mx < np.iinfo(np.int16).max:\n                        props[col] = props[col].astype(np.int16)\n                    elif mn > np.iinfo(np.int32).min and mx < np.iinfo(np.int32).max:\n                        props[col] = props[col].astype(np.int32)\n                    elif mn > np.iinfo(np.int64).min and mx < np.iinfo(np.int64).max:\n                        props[col] = props[col].astype(np.int64)    \n            \n            # Make float datatypes 32 bit\n            else:\n                props[col] = props[col].astype(np.float32)\n           \n            # Print new column type\n            if print_mode:\n                print(\"dtype after: \",props[col].dtype)\n                print(\"******************************\")\n\n    # Print final result\n    if print_mode:\n        print(\"___MEMORY USAGE AFTER COMPLETION:___\")\n        mem_usg = props.memory_usage().sum() / 1024**2\n        print(\"Memory usage is: \",mem_usg,\" MB\")\n        print(\"This is \",100*mem_usg/start_mem_usg,\"% of the initial size\")\n    return props, NAlist","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:24:54.534747Z","iopub.execute_input":"2022-08-05T04:24:54.535249Z","iopub.status.idle":"2022-08-05T04:24:54.554732Z","shell.execute_reply.started":"2022-08-05T04:24:54.535195Z","shell.execute_reply":"2022-08-05T04:24:54.552119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reducing DataFrame memory ~65%\nif data_reducing:\n    props, NAlist = reduce_mem_usage(train_num)\n    print(\"_________________\")\n    print(\"\")\n    if NAlist:\n        print(\"Warning: the following columns have missing values filled with 'df['column_name'].min() -1': \")\n        print(\"_________________\")\n        print(\"\")\n        print(NAlist)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:24:54.556403Z","iopub.execute_input":"2022-08-05T04:24:54.556926Z","iopub.status.idle":"2022-08-05T04:25:38.948850Z","shell.execute_reply.started":"2022-08-05T04:24:54.556888Z","shell.execute_reply":"2022-08-05T04:25:38.947559Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## 각 product가 처음으로 지나가는 station_name과 마지막으로 지나가는 station_name을 feature로 사용하고 싶습니다.\n## TO-DO: ~ 20:20까지\n## start_station, end_station이라는 feature를 만들어 봅시다!\n\n# 1. column에서 station 정보 뽑기\ncounts = train_date.count()\ndate_cols = counts.reset_index()[\"index\"].str.split('_', expand=True)[1].drop_duplicates().index\n\n# 2. 각 station별로 column 뽑기\ndate_cols = train_date.columns[date_cols]\ndate_cols","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:25:38.950168Z","iopub.execute_input":"2022-08-05T04:25:38.950465Z","iopub.status.idle":"2022-08-05T04:25:38.976571Z","shell.execute_reply.started":"2022-08-05T04:25:38.950434Z","shell.execute_reply":"2022-08-05T04:25:38.975349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 3. 뽑은 column을 기준으로 start, end station 구하기\n#train_date = train_date[date_cols]\ntrain_date = pd.read_csv('../input/bosch-production-line-performance/train_date.csv.zip',\n                        usecols=date_cols)\n\ntrain_date['start_station'] = -1\ntrain_date['end_station'] = -1\n\nfor col in train_date.drop(columns=[\"Id\", \"start_station\", \"end_station\"]).columns:\n    notnulls = (~train_date[col].isnull())\n    station_no = int(col.split(\"_\")[1][1:])\n    \n    train_date.loc[(notnulls) & (train_date.start_station == -1), \"start_station\"] = station_no\n    train_date.loc[(notnulls), \"end_station\"] = station_no\n    \ntrain_date","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:25:38.978569Z","iopub.execute_input":"2022-08-05T04:25:38.979360Z","iopub.status.idle":"2022-08-05T04:26:20.503117Z","shell.execute_reply.started":"2022-08-05T04:25:38.979309Z","shell.execute_reply":"2022-08-05T04:26:20.501766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_date_part = train_date[[\"Id\", \"start_station\", \"end_station\"]]\ntrain_date_part","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:26:20.504914Z","iopub.execute_input":"2022-08-05T04:26:20.505421Z","iopub.status.idle":"2022-08-05T04:26:20.527387Z","shell.execute_reply.started":"2022-08-05T04:26:20.505387Z","shell.execute_reply":"2022-08-05T04:26:20.526212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if data_reducing:\n    train_date_part, _ = reduce_mem_usage(train_date_part)\n    display(train_date_part.info())","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:26:20.529136Z","iopub.execute_input":"2022-08-05T04:26:20.529941Z","iopub.status.idle":"2022-08-05T04:26:20.612692Z","shell.execute_reply.started":"2022-08-05T04:26:20.529905Z","shell.execute_reply":"2022-08-05T04:26:20.611137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_date\nimport gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:26:20.615230Z","iopub.execute_input":"2022-08-05T04:26:20.615768Z","iopub.status.idle":"2022-08-05T04:26:20.773600Z","shell.execute_reply.started":"2022-08-05T04:26:20.615720Z","shell.execute_reply":"2022-08-05T04:26:20.772552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. EDA","metadata":{"id":"c9c9acb8"}},{"cell_type":"markdown","source":"- 데이터에서 찾아야 하는 기초적인 내용들을 확인합니다.\n\n\n- class imbalance, target distribution, outlier, correlation을 확인합니다.","metadata":{"id":"6fdf620b"}},{"cell_type":"code","source":"# ## On your Own\n# data.column.value_counts()\n# sns.countplot()\n# sns.histplot()\n# ...","metadata":{"id":"adb06474","execution":{"iopub.status.busy":"2022-08-05T04:26:20.780374Z","iopub.execute_input":"2022-08-05T04:26:20.780751Z","iopub.status.idle":"2022-08-05T04:26:20.786410Z","shell.execute_reply.started":"2022-08-05T04:26:20.780717Z","shell.execute_reply":"2022-08-05T04:26:20.785439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"이런 식으로 여러가지 그래프를 그려가며, 데이터에 대한 인사이트를 얻습니다!","metadata":{"id":"f2c602bd"}},{"cell_type":"markdown","source":"### 3. 전처리","metadata":{"id":"9dbb8802"}},{"cell_type":"markdown","source":"#### 결측치 처리","metadata":{"id":"b79a6f0a"}},{"cell_type":"code","source":"# 결측치가 있는 column\ndisplay(train_num[train_num.isnull().any(axis=1)])\ndisplay(train_date_part[train_date_part.isnull().any(axis=1)])","metadata":{"id":"bbafdcd0","execution":{"iopub.status.busy":"2022-08-05T04:26:20.787504Z","iopub.execute_input":"2022-08-05T04:26:20.787830Z","iopub.status.idle":"2022-08-05T04:26:21.621372Z","shell.execute_reply.started":"2022-08-05T04:26:20.787797Z","shell.execute_reply":"2022-08-05T04:26:21.620100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### feature extraction\n\n- 차원의 저주를 해결하거나, 데이터의 feature 조합을 이용하는 새로운 feature를 생성할 때, PCA를 사용합니다.\n\n- 분석에 사용할 feature를 선택하는 과정도 포함합니다.","metadata":{"id":"606493f0"}},{"cell_type":"code","source":"X = train_num.drop(columns=[\"Id\", \"Response\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:26:21.622868Z","iopub.execute_input":"2022-08-05T04:26:21.623241Z","iopub.status.idle":"2022-08-05T04:26:21.890197Z","shell.execute_reply.started":"2022-08-05T04:26:21.623207Z","shell.execute_reply":"2022-08-05T04:26:21.888664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# PCA 적용\nfrom sklearn.decomposition import PCA\n\nif feature_reducing:\n    pca = PCA(n_components=15) # PCA(n_components=6)\n    pca_data = pca.fit_transform(X)\n    pca_cols = [f\"PC{i}\" for i in range(1, pca_data.shape[1]+1)]\n    X = pd.DataFrame(data=pca_data, columns=pca_cols)\n    display(X)","metadata":{"id":"4a137c94","execution":{"iopub.status.busy":"2022-08-05T04:26:21.891598Z","iopub.execute_input":"2022-08-05T04:26:21.892141Z","iopub.status.idle":"2022-08-05T04:26:42.810071Z","shell.execute_reply.started":"2022-08-05T04:26:21.892088Z","shell.execute_reply":"2022-08-05T04:26:42.808314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X[\"Id\"] = train_num.Id\nX = pd.merge(X, train_date_part, on=\"Id\").drop(\"Id\", axis=1) # feature vector\ny = train_num.Response                                       # target value\nprint(X.shape, y.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:56:10.101280Z","iopub.execute_input":"2022-08-05T04:56:10.101717Z","iopub.status.idle":"2022-08-05T04:56:10.344802Z","shell.execute_reply.started":"2022-08-05T04:56:10.101682Z","shell.execute_reply":"2022-08-05T04:56:10.343000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:39:16.188118Z","iopub.execute_input":"2022-08-05T04:39:16.188633Z","iopub.status.idle":"2022-08-05T04:39:16.224510Z","shell.execute_reply.started":"2022-08-05T04:39:16.188597Z","shell.execute_reply":"2022-08-05T04:39:16.223440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:42:00.374593Z","iopub.execute_input":"2022-08-05T04:42:00.375077Z","iopub.status.idle":"2022-08-05T04:42:00.443634Z","shell.execute_reply.started":"2022-08-05T04:42:00.375022Z","shell.execute_reply":"2022-08-05T04:42:00.442221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del train_num\ndel train_date_part\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:26:43.558639Z","iopub.execute_input":"2022-08-05T04:26:43.559472Z","iopub.status.idle":"2022-08-05T04:26:43.773108Z","shell.execute_reply.started":"2022-08-05T04:26:43.559421Z","shell.execute_reply":"2022-08-05T04:26:43.771427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T04:51:36.104925Z","iopub.execute_input":"2022-08-05T04:51:36.105402Z","iopub.status.idle":"2022-08-05T04:51:36.123368Z","shell.execute_reply.started":"2022-08-05T04:51:36.105366Z","shell.execute_reply":"2022-08-05T04:51:36.122155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 4. 학습 데이터 분할\n\n- 현재 데이터는 class imbalance가 아주 심한 상황입니다.\n\n\n- Under sampling과 Over sampling을 사용하여 클래스의 비율을 맞춰봅니다.","metadata":{"id":"f497a2d8"}},{"cell_type":"code","source":"temp = pd.concat([X, y], axis=1)\n\nnormal = temp[temp.Response == 0] # 정상(majority class)\nabnormal = temp[temp.Response == 1] # 불량(minority class)\nprint(normal.shape, abnormal.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:44:45.292349Z","iopub.execute_input":"2022-08-05T06:44:45.292764Z","iopub.status.idle":"2022-08-05T06:44:45.419644Z","shell.execute_reply.started":"2022-08-05T06:44:45.292730Z","shell.execute_reply":"2022-08-05T06:44:45.417928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 1. Undersampling\nnormal.sample(n=len(abnormal), random_state=42)\n#X = pd.concat([normal, abnormal]) (6879x2, 18)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:46:57.017732Z","iopub.execute_input":"2022-08-05T06:46:57.018240Z","iopub.status.idle":"2022-08-05T06:46:57.098560Z","shell.execute_reply.started":"2022-08-05T06:46:57.018202Z","shell.execute_reply":"2022-08-05T06:46:57.096744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 2. Oversampling\nfrom imblearn.over_sampling import SMOTE\n\nsmote = SMOTE()\nX_train1, y_train1 = smote.fit_resample(X, y)\nprint(X_train1.shape, y_train1.shape) # 2353736 = 1176868 x 2","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:55:04.676212Z","iopub.execute_input":"2022-08-05T06:55:04.676678Z","iopub.status.idle":"2022-08-05T06:55:07.025997Z","shell.execute_reply.started":"2022-08-05T06:55:04.676641Z","shell.execute_reply":"2022-08-05T06:55:07.024488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 3. hybird\n#N = 50000\nN = len(abnormal) * 10\ntemp = normal.sample(n=N, random_state=42)\nsample = pd.concat([temp, abnormal])\nsample.shape\n\nX_sample = sample.drop(\"Response\", axis=1)\ny_sample = sample.Response\n\nX_, y_ = smote.fit_resample(X_sample, y_sample)\nprint(X_.shape, y_.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T07:04:00.236980Z","iopub.execute_input":"2022-08-05T07:04:00.237772Z","iopub.status.idle":"2022-08-05T07:04:01.736358Z","shell.execute_reply.started":"2022-08-05T07:04:00.237691Z","shell.execute_reply":"2022-08-05T07:04:01.734855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 첫번째 테스트용으로 사용하고, 실제 학습시에는 K-Fold CV를 사용합니다.\n# train : test = 8 : 2\nfrom sklearn.model_selection import train_test_split\n\n#X = X.drop(\"Id\", axis=1)\n\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2,\n                                                    random_state=42, stratify=y)\nprint(X_train.shape, X_test.shape, y_train.shape, y_test.shape)","metadata":{"id":"47306aaf","execution":{"iopub.status.busy":"2022-08-05T05:31:05.148621Z","iopub.execute_input":"2022-08-05T05:31:05.149115Z","iopub.status.idle":"2022-08-05T05:31:05.750926Z","shell.execute_reply.started":"2022-08-05T05:31:05.149052Z","shell.execute_reply":"2022-08-05T05:31:05.749523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train.value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:31:05.755284Z","iopub.execute_input":"2022-08-05T05:31:05.755648Z","iopub.status.idle":"2022-08-05T05:31:05.774450Z","shell.execute_reply.started":"2022-08-05T05:31:05.755615Z","shell.execute_reply":"2022-08-05T05:31:05.772661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test.value_counts(normalize=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:31:05.776491Z","iopub.execute_input":"2022-08-05T05:31:05.776843Z","iopub.status.idle":"2022-08-05T05:31:05.787759Z","shell.execute_reply.started":"2022-08-05T05:31:05.776809Z","shell.execute_reply":"2022-08-05T05:31:05.786758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 5. 학습 및 평가","metadata":{"id":"58056e51"}},{"cell_type":"code","source":"# 간단하게 LightGBM 테스트\n# 적당한 hyper-parameter 조합을 두었습니다. (항상 best는 아닙니다. 예시입니다.)\n\nmodel = LGBMClassifier()","metadata":{"id":"39fd2515","execution":{"iopub.status.busy":"2022-08-05T05:31:08.994051Z","iopub.execute_input":"2022-08-05T05:31:08.994499Z","iopub.status.idle":"2022-08-05T05:31:09.001517Z","shell.execute_reply.started":"2022-08-05T05:31:08.994462Z","shell.execute_reply":"2022-08-05T05:31:09.000141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"\\nFitting LightGBM...\")\nmodel.fit(X_train, y_train)","metadata":{"scrolled":true,"id":"ddffa474","execution":{"iopub.status.busy":"2022-08-05T05:31:09.307739Z","iopub.execute_input":"2022-08-05T05:31:09.308204Z","iopub.status.idle":"2022-08-05T05:31:15.490535Z","shell.execute_reply.started":"2022-08-05T05:31:09.308169Z","shell.execute_reply":"2022-08-05T05:31:15.489314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# metric은 그때마다 맞게 바꿔줘야 합니다.\nfrom sklearn.metrics import matthews_corrcoef\n\nevaluation_metric = matthews_corrcoef","metadata":{"id":"6c8b0259","execution":{"iopub.status.busy":"2022-08-05T05:31:15.493170Z","iopub.execute_input":"2022-08-05T05:31:15.493581Z","iopub.status.idle":"2022-08-05T05:31:15.499313Z","shell.execute_reply.started":"2022-08-05T05:31:15.493547Z","shell.execute_reply":"2022-08-05T05:31:15.497821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Prediction\")\npred_train = model.predict(X_train)\npred_test = model.predict(X_test)\n\n\ntrain_score = evaluation_metric(y_train, pred_train)\ntest_score = evaluation_metric(y_test, pred_test)\n\nprint(\"Train Score : %.4f\" % train_score)\nprint(\"Test Score : %.4f\" % test_score)","metadata":{"id":"a6b39be5","execution":{"iopub.status.busy":"2022-08-05T05:31:15.501134Z","iopub.execute_input":"2022-08-05T05:31:15.501495Z","iopub.status.idle":"2022-08-05T05:31:17.512853Z","shell.execute_reply.started":"2022-08-05T05:31:15.501460Z","shell.execute_reply":"2022-08-05T05:31:17.511372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 6. Hyper-parameter Tuning","metadata":{"id":"bc755b17"}},{"cell_type":"markdown","source":"> optuna를 사용해봅시다 !","metadata":{"id":"97e5302d"}},{"cell_type":"code","source":"def optimizer(trial, X, y, K):\n    # 조절할 hyper-parameter 조합을 적어줍니다.\n    n_estimators = \n    max_depth = \n    max_features = \n    \n    \n    # 원하는 모델을 지정합니다, optuna는 시간이 오래걸리기 때문에 저는 보통 RF로 일단 테스트를 해본 뒤에 LGBM을 사용합니다.\n    model = RandomForestRegressor(n_estimators=n_estimators,\n                                 max_depth=max_depth,\n                                 max_features=max_features)\n    \n    \n    # K-Fold Cross validation을 구현합니다.\n    folds = KFold(n_splits=K)\n    losses = []\n    \n    for train_idx, val_idx in folds.split(X, y):\n        X_train = X.iloc[train_idx, :]\n        y_train = y.iloc[train_idx]\n        \n        X_val = X.iloc[val_idx, :]\n        y_val = y.iloc[val_idx]\n        \n        model.fit(X_train, y_train)\n        preds = model.predict(X_val)\n        loss = mean_absolute_error(y_val, preds)\n        losses.append(loss)\n    \n    \n    # K-Fold의 평균 loss값을 돌려줍니다.\n    return np.mean(losses)","metadata":{"id":"34ce4986","execution":{"iopub.status.busy":"2022-08-05T04:26:43.814511Z","iopub.status.idle":"2022-08-05T04:26:43.815169Z","shell.execute_reply.started":"2022-08-05T04:26:43.814805Z","shell.execute_reply":"2022-08-05T04:26:43.814838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"K = # Kfold 수\nopt_func = partial(optimizer, X=X_train, y=y_train, K)\n\nstudy = optuna.create_study(direction=\"minimize\") # 최소/최대 어느 방향의 최적값을 구할 건지.\nstudy.optimize(opt_func, n_trials=5)","metadata":{"id":"7150b210","execution":{"iopub.status.busy":"2022-08-05T04:26:43.817628Z","iopub.status.idle":"2022-08-05T04:26:43.818549Z","shell.execute_reply.started":"2022-08-05T04:26:43.818027Z","shell.execute_reply":"2022-08-05T04:26:43.818091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# optuna가 시도했던 모든 실험 관련 데이터\nstudy.trials_dataframe()","metadata":{"id":"72d0a118","execution":{"iopub.status.busy":"2022-08-05T04:26:43.821408Z","iopub.status.idle":"2022-08-05T04:26:43.822251Z","shell.execute_reply.started":"2022-08-05T04:26:43.821702Z","shell.execute_reply":"2022-08-05T04:26:43.821793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Best Score: %.4f\" % study.best_value) # best score 출력\nprint(\"Best params: \", study.best_trial.params) # best score일 때의 하이퍼파라미터들","metadata":{"id":"a805da05","execution":{"iopub.status.busy":"2022-08-05T04:26:43.823571Z","iopub.status.idle":"2022-08-05T04:26:43.825282Z","shell.execute_reply.started":"2022-08-05T04:26:43.824909Z","shell.execute_reply":"2022-08-05T04:26:43.824943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 실험 기록 시각화\noptuna.visualization.plot_optimization_history(study)","metadata":{"id":"051ae1eb","execution":{"iopub.status.busy":"2022-08-05T04:26:43.827147Z","iopub.status.idle":"2022-08-05T04:26:43.827861Z","shell.execute_reply.started":"2022-08-05T04:26:43.827511Z","shell.execute_reply":"2022-08-05T04:26:43.827553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# hyper-parameter들의 중요도\noptuna.visualization.plot_param_importances(study)","metadata":{"id":"efbf8f65","execution":{"iopub.status.busy":"2022-08-05T04:26:43.830215Z","iopub.status.idle":"2022-08-05T04:26:43.830845Z","shell.execute_reply.started":"2022-08-05T04:26:43.830515Z","shell.execute_reply":"2022-08-05T04:26:43.830549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 7. 테스트 및 제출 파일 생성","metadata":{"id":"24b360ec"}},{"cell_type":"code","source":"## test 데이터 만들기\ntest_num = pd.read_csv('../input/bosch-production-line-performance/test_numeric.csv.zip',\n                      usecols=usecols.drop(\"Response\"))\ntest_date = pd.read_csv('../input/bosch-production-line-performance/test_date.csv.zip',\n                       usecols=date_cols)\n\nprint(test_num.shape, test_date.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:55:22.065322Z","iopub.execute_input":"2022-08-05T05:55:22.065760Z","iopub.status.idle":"2022-08-05T05:56:58.162149Z","shell.execute_reply.started":"2022-08-05T05:55:22.065725Z","shell.execute_reply":"2022-08-05T05:56:58.160554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 3. 뽑은 column을 기준으로 start, end station 구하기\n\ntest_date['start_station'] = -1\ntest_date['end_station'] = -1\n\nfor col in test_date.drop(columns=[\"Id\", \"start_station\", \"end_station\"]).columns:\n    notnulls = (~test_date[col].isnull())\n    station_no = int(col.split(\"_\")[1][1:])\n    \n    test_date.loc[(notnulls) & (test_date.start_station == -1), \"start_station\"] = station_no\n    test_date.loc[(notnulls), \"end_station\"] = station_no\n    \ntest_date","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:57:33.070444Z","iopub.execute_input":"2022-08-05T05:57:33.071283Z","iopub.status.idle":"2022-08-05T05:57:34.742179Z","shell.execute_reply.started":"2022-08-05T05:57:33.071246Z","shell.execute_reply":"2022-08-05T05:57:34.740863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_date_part = test_date[[\"Id\", \"start_station\", \"end_station\"]]\ntest_date_part","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:58:10.417235Z","iopub.execute_input":"2022-08-05T05:58:10.418195Z","iopub.status.idle":"2022-08-05T05:58:10.441024Z","shell.execute_reply.started":"2022-08-05T05:58:10.418158Z","shell.execute_reply":"2022-08-05T05:58:10.439676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_num = test_num.fillna(0)\nX_test = test_num.drop(\"Id\", axis=1)\nif feature_reducing:\n    X_test = pca.transform(X_test)\n    X_test = pd.DataFrame(data=X_test, columns=pca_cols)\n    display(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T05:59:59.173122Z","iopub.execute_input":"2022-08-05T05:59:59.173849Z","iopub.status.idle":"2022-08-05T06:00:02.375936Z","shell.execute_reply.started":"2022-08-05T05:59:59.173813Z","shell.execute_reply":"2022-08-05T06:00:02.374406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test = pd.concat([X_test, test_date_part], axis=1).drop(\"Id\", axis=1)\nX_test","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:00:02.385519Z","iopub.execute_input":"2022-08-05T06:00:02.391160Z","iopub.status.idle":"2022-08-05T06:00:02.717593Z","shell.execute_reply.started":"2022-08-05T06:00:02.391075Z","shell.execute_reply":"2022-08-05T06:00:02.716133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 메모리 정리\nif data_reducing:\n    X_test, _ = reduce_mem_usage(X_test)\n    X_test.info()\n    \ndel test_date\ndel test_num","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:14:14.686231Z","iopub.execute_input":"2022-08-05T06:14:14.686759Z","iopub.status.idle":"2022-08-05T06:14:15.071646Z","shell.execute_reply.started":"2022-08-05T06:14:14.686706Z","shell.execute_reply":"2022-08-05T06:14:15.070129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model = RandomForestRegressor(n_estimators=study.best_trial.params[\"n_estimators\"],\n#                                  max_depth=study.best_trial.params[\"max_depth\"],\n#                                  max_features=study.best_trial.params[\"max_features\"])\n\npreds = model.predict(X_test)\npreds","metadata":{"id":"a0daf54e","execution":{"iopub.status.busy":"2022-08-05T06:01:15.882694Z","iopub.execute_input":"2022-08-05T06:01:15.883157Z","iopub.status.idle":"2022-08-05T06:01:17.668686Z","shell.execute_reply.started":"2022-08-05T06:01:15.883122Z","shell.execute_reply":"2022-08-05T06:01:17.667735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.read_csv('../input/bosch-production-line-performance/sample_submission.csv.zip')\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:01:43.098310Z","iopub.execute_input":"2022-08-05T06:01:43.098952Z","iopub.status.idle":"2022-08-05T06:01:43.339709Z","shell.execute_reply.started":"2022-08-05T06:01:43.098916Z","shell.execute_reply":"2022-08-05T06:01:43.338337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[\"Response\"] = preds\nsubmission.to_csv(\"submission.csv\", index=False)","metadata":{"id":"55a2c13f","execution":{"iopub.status.busy":"2022-08-05T06:02:00.042775Z","iopub.execute_input":"2022-08-05T06:02:00.043598Z","iopub.status.idle":"2022-08-05T06:02:01.369123Z","shell.execute_reply.started":"2022-08-05T06:02:00.043560Z","shell.execute_reply":"2022-08-05T06:02:01.367344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}