{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":4117,"databundleVersionId":46665,"sourceType":"competition"},{"sourceId":3161014,"sourceType":"datasetVersion","datasetId":1917019}],"dockerImageVersionId":30157,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**I used papago for translation :)**","metadata":{}},{"cell_type":"markdown","source":"**이 노트북은 악성코드 데이터를 분석해보고 싶지만 너무 큰 데이터 용량에 가로막혀 분석을 진행하지 못하는 사람들을 위하여 제작되었습니다.**\n\n**용량의 한계가 있어 asm 파일을 제외하고 byte 파일만 사용하였으며 byte 파일만을 이용하여 public leader board score에서 0.02742를 달성하였습니다.**\n\n**byte파일만으로 데이터셋을 제작하여 용량에 제한을 받지 않고 따라할 수 있도록 하였습니다.**\n\n**여러 방법을 시도한 다음, 모델을 평가하는 과정은 중복되는 부분이 많아 제거하고 성능이 좋은 부분만 남겼습니다.**\n\n**용량의 제약이 없다면 asm 데이터를 사용하여 더 좋은 성능을 보이실 수 있습니다.**","metadata":{}},{"cell_type":"markdown","source":"**This code is designed for those who want to analyze malicious code data but are blocked by too much data capacity to proceed with the analysis.**\n\n**Due to the capacity limitation, only byte files except for asm files were used, and 0.02742 was achieved at the public leader board core using only byte files.**\n\n**We made a dataset with only byte files so that it can be followed without limitation in capacity.**\n\n**After attempting several methods, the process of evaluating the model removed many overlapping parts and left only the good performance.**\n\n**If you don't have a capacity limit, you can perform better using asm data.**","metadata":{}},{"cell_type":"markdown","source":"# **Import**","metadata":{}},{"cell_type":"markdown","source":"**pandas, numpy를 사용하고**\n\n**진행상황을 알기 위하여 tqdm을 사용합니다.**\n\n**파일을 읽기 위한 os와 경고 무시를 위한 warnings,**\n\n**cv를 이용한 stacking data 제작을 위한 stratified kfold,**\n\n**이외에 요즘 캐글에서 좋은 성능을 보이고 있는 boosting 모델들을 import 하였습니다.**","metadata":{}},{"cell_type":"markdown","source":"**use pandas, numpy**\n\n**We use tqdm to know the progress.**\n\n**os for reading files and warnings for ignoring warnings,**\n\n**Stratified kfold for producing stacking data using cv,**\n\n**In addition, we imported boosting models that are showing good performance in kaggle these days.**","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom tqdm import tqdm\nimport os\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nfrom sklearn.model_selection import StratifiedKFold\nfold = StratifiedKFold(n_splits=5, shuffle = True, random_state=62)\n\nfrom lightgbm import LGBMClassifier\nfrom xgboost import XGBClassifier\nfrom catboost import CatBoostClassifier","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T13:11:41.507478Z","iopub.execute_input":"2025-04-06T13:11:41.507915Z","iopub.status.idle":"2025-04-06T13:11:41.514956Z","shell.execute_reply.started":"2025-04-06T13:11:41.507869Z","shell.execute_reply":"2025-04-06T13:11:41.513892Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **데이터 불러오기(data load)**","metadata":{}},{"cell_type":"markdown","source":"**파일의 class를 적어놓은 trainlabels와 submission의 형식이 있는 sample submission을 읽어옵니다.**\n\n**Read the sample submission in the form of train labels and submission where the class of the file is written.**","metadata":{}},{"cell_type":"code","source":"train_labels=pd.read_csv(\"../input/malware-classification/trainLabels.csv\")\nsample = pd.read_csv(\"../input/malware-classification/sampleSubmission.csv\",index_col=\"Id\")\na = '\"0\",\"1\",\"2\",\"3\",\"4\",\"5\",\"6\",\"7\",\"8\",\"9\",\"0a\",\"0b\",\"0c\",\"0d\",\"0e\",\"0f\",\"10\",\"11\",\"12\",\"13\",\"14\",\"15\",\"16\",\"17\",\"18\",\"19\",\"1a\",\"1b\",\"1c\",\"1d\",\"1e\",\"1f\",\"20\",\"21\",\"22\",\"23\",\"24\",\"25\",\"26\",\"27\",\"28\",\"29\",\"2a\",\"2b\",\"2c\",\"2d\",\"2e\",\"2f\",\"30\",\"31\",\"32\",\"33\",\"34\",\"35\",\"36\",\"37\",\"38\",\"39\",\"3a\",\"3b\",\"3c\",\"3d\",\"3e\",\"3f\",\"40\",\"41\",\"42\",\"43\",\"44\",\"45\",\"46\",\"47\",\"48\",\"49\",\"4a\",\"4b\",\"4c\",\"4d\",\"4e\",\"4f\",\"50\",\"51\",\"52\",\"53\",\"54\",\"55\",\"56\",\"57\",\"58\",\"59\",\"5a\",\"5b\",\"5c\",\"5d\",\"5e\",\"5f\",\"60\",\"61\",\"62\",\"63\",\"64\",\"65\",\"66\",\"67\",\"68\",\"69\",\"6a\",\"6b\",\"6c\",\"6d\",\"6e\",\"6f\",\"70\",\"71\",\"72\",\"73\",\"74\",\"75\",\"76\",\"77\",\"78\",\"79\",\"7a\",\"7b\",\"7c\",\"7d\",\"7e\",\"7f\",\"80\",\"81\",\"82\",\"83\",\"84\",\"85\",\"86\",\"87\",\"88\",\"89\",\"8a\",\"8b\",\"8c\",\"8d\",\"8e\",\"8f\",\"90\",\"91\",\"92\",\"93\",\"94\",\"95\",\"96\",\"97\",\"98\",\"99\",\"9a\",\"9b\",\"9c\",\"9d\",\"9e\",\"9f\",\"a0\",\"a1\",\"a2\",\"a3\",\"a4\",\"a5\",\"a6\",\"a7\",\"a8\",\"a9\",\"aa\",\"ab\",\"ac\",\"ad\",\"ae\",\"af\",\"b0\",\"b1\",\"b2\",\"b3\",\"b4\",\"b5\",\"b6\",\"b7\",\"b8\",\"b9\",\"ba\",\"bb\",\"bc\",\"bd\",\"be\",\"bf\",\"c0\",\"c1\",\"c2\",\"c3\",\"c4\",\"c5\",\"c6\",\"c7\",\"c8\",\"c9\",\"ca\",\"cb\",\"cc\",\"cd\",\"ce\",\"cf\",\"d0\",\"d1\",\"d2\",\"d3\",\"d4\",\"d5\",\"d6\",\"d7\",\"d8\",\"d9\",\"da\",\"db\",\"dc\",\"dd\",\"de\",\"df\",\"e0\",\"e1\",\"e2\",\"e3\",\"e4\",\"e5\",\"e6\",\"e7\",\"e8\",\"e9\",\"ea\",\"eb\",\"ec\",\"ed\",\"ee\",\"ef\",\"f0\",\"f1\",\"f2\",\"f3\",\"f4\",\"f5\",\"f6\",\"f7\",\"f8\",\"f9\",\"fa\",\"fb\",\"fc\",\"fd\",\"fe\",\"ff\",\"??\",\"size\",\"Class\"'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T13:11:41.516531Z","iopub.execute_input":"2025-04-06T13:11:41.516769Z","iopub.status.idle":"2025-04-06T13:11:41.643722Z","shell.execute_reply.started":"2025-04-06T13:11:41.516733Z","shell.execute_reply":"2025-04-06T13:11:41.642731Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **전처리(preprocessing)**","metadata":{}},{"cell_type":"markdown","source":"```python\ntrain_np = np.zeros((len(files),len(a.split(\",\"))))\nfiles = os.listdir(\"../input/malware-only-byte/only_byte\")\nfiles.sort()\nk=0\nfor file in tqdm(files):\n    statinfo=os.stat(\"../input/malware-only-byte/only_byte/\"+file)\n    with open(\"../input/malware-only-byte/only_byte/\"+file,\"r\") as fp:\n        for lines in fp.readlines():\n            line=lines.rstrip().split(\" \")[1:]\n            for hex_code in line:\n                if hex_code=='??':\n                    train_np[k][256]+=1\n                else:\n                    train_np[k][int(hex_code,16)]+=1\n        train_np[k][257]=statinfo.st_size/(1024*1024)\n        train_np[k][258]=train_labels[train_labels[\"Id\"]==file.split('.')[0]][\"Class\"].tolist()[0]\n    fp.close()\n    k += 1\ntrain=pd.DataFrame(train_np, columns=a[1:-1].split('\",\"'))\ntrain.to_csv(\"train_data.csv\",index=False)\ntrain.head()\n```","metadata":{}},{"cell_type":"code","source":"files = os.listdir(\"../input/malware-only-byte/only_byte\")\nfiles.sort()\n\ntrain_np = np.zeros((len(files),len(a.split(\",\"))))\n\nk=0\nfor file in tqdm(files):\n    statinfo=os.stat(\"../input/malware-only-byte/only_byte/\"+file)\n    with open(\"../input/malware-only-byte/only_byte/\"+file,\"r\") as fp:\n        for lines in fp.readlines():\n            line=lines.rstrip().split(\" \")[1:]\n            for hex_code in line:\n                if hex_code=='??':\n                    train_np[k][256]+=1\n                else:\n                    train_np[k][int(hex_code,16)]+=1\n        train_np[k][257]=statinfo.st_size/(1024*1024)\n        train_np[k][258]=train_labels[train_labels[\"Id\"]==file.split('.')[0]][\"Class\"].tolist()[0]\n    fp.close()\n    k += 1\ntrain=pd.DataFrame(train_np, columns=a[1:-1].split('\",\"'))\ntrain.to_csv(\"train_data.csv\",index=False)\ntrain.head()  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-06T13:11:41.645108Z","iopub.execute_input":"2025-04-06T13:11:41.645483Z","execution_failed":"2025-04-06T16:58:27.428Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**위에 있는 코드는 파일명을 이용하여 파일을 읽어들이고 각 파일에 0,1,2,...,?? 등의 hex값의 수를 나타낸 코드입니다.**\n\n**https://www.kaggle.com/paulrohan2020/microsoft-malware-detection-log-loss-of-0-0070 를 참고하였습니다.**\n\n**실행시간이 길어서 markdown으로 적었습니다. 위 코드를 복사하시고 실행하실 수 있습니다.**","metadata":{}},{"cell_type":"markdown","source":"**The code above is a code that reads the file using the file name and represents the number of hex values such as 0, 1, 2, ..., ?? in each file.**\n\n**I referred to https://www.kaggle.com/paulrohan2020/microsoft-malware-detection-log-loss-of-0-0070**\n\n**I wrote down markdown because of the long execution time. You can copy and execute the code above.**","metadata":{}},{"cell_type":"code","source":"train=pd.read_csv(\"../input/malware-only-byte/train_data.csv\")","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.428Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"```python\nfile2=[]\nfor file in files:\n    file2.append(file.split(\".\")[0])\ntest_np = np.zeros((len(files),len(a.split(\",\"))))\nfiles = os.listdir(\"../input/malware-only-byte/test\")\nfiles.sort()\nk=0\nfor file in tqdm(files):\n    statinfo=os.stat(\"../input/malware-only-byte/test/\"+file)\n    with open(\"../input/malware-only-byte/test/\"+file,\"r\") as fp:\n        for lines in fp.readlines():\n            line=lines.rstrip().split(\" \")[1:]\n            for hex_code in line:\n                if hex_code=='??':\n                    test_np[k][256]+=1\n                else:\n                    test_np[k][int(hex_code,16)]+=1\n        test_np[k][257]=statinfo.st_size/(1024*1024)\n    fp.close()\n    k += 1\ntest=pd.DataFrame(test_np, columns=a[1:-1].split('\",\"'))\ntest[\"Id\"] = file2\ntest.to_csv(\"test_data.csv\",index=False)\ntest.head()\n```","metadata":{}},{"cell_type":"code","source":"file2=[]\nfor file in files:\n    file2.append(file.split(\".\")[0])\ntest_np = np.zeros((len(files),len(a.split(\",\"))))\nfiles = os.listdir(\"../input/malware-only-byte/test\")\nfiles.sort()\nk=0\nfor file in tqdm(files):\n    statinfo=os.stat(\"../input/malware-only-byte/test/\"+file)\n    with open(\"../input/malware-only-byte/test/\"+file,\"r\") as fp:\n        for lines in fp.readlines():\n            line=lines.rstrip().split(\" \")[1:]\n            for hex_code in line:\n                if hex_code=='??':\n                    test_np[k][256]+=1\n                else:\n                    test_np[k][int(hex_code,16)]+=1\n        test_np[k][257]=statinfo.st_size/(1024*1024)\n    fp.close()\n    k += 1\ntest=pd.DataFrame(test_np, columns=a[1:-1].split('\",\"'))\ntest[\"Id\"] = file2\ntest.to_csv(\"test_data.csv\",index=False)\ntest.head()","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**위 코드는 test byte 파일에 대하여 진행한 코드입니다.**\n\n**The code above is the code conducted for the test byte file.**","metadata":{}},{"cell_type":"code","source":"test=pd.read_csv(\"../input/malware-only-byte/test_data.csv\")","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **여러가지 model을 이용한 학습과 스태킹(modeling and stacking)**","metadata":{}},{"cell_type":"code","source":"X = train.drop(['Class'], axis=1)\ny = train['Class']","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**예측할 값을 제외한 값을 X로, 예측값을 y로 나누어 줍니다.**\n\n**Divide the values excluding the predicted values by X and y.**","metadata":{}},{"cell_type":"markdown","source":"**이후로 stacking을 위하여 train data에 cross validation을 진행하여 각 validation set에 대한 예측값으로 train data set의 예측값을 만듭니다.**\n\n**바로 X,y를 학습시키고 예측을 진행해도 좋지만 stacking을 진행한 쪽이 더 성능이 좋았습니다.**","metadata":{}},{"cell_type":"markdown","source":"**After that, cross validation is performed on the train data for stacking to create a predicted value of the train data set as a predicted value for each validation set.**\n\n**You can learn X and y right away and make predictions, but the one who did the stacking performed better.**","metadata":{}},{"cell_type":"code","source":"stack_df = pd.read_csv(\"../input/malware-only-byte/train_data.csv\")\nfor i,j in enumerate(fold.split(X,y)):\n    stack_train_X = X.iloc[j[0]]\n    stack_train_y = y.iloc[j[0]]\n    stack_test_X = X.iloc[j[1]]\n    stack_test_y = y.iloc[j[1]]\n    model = LGBMClassifier(learning_rate= 0.025, n_estimators = 850, min_child_weight = 1, boosting_type = \"gbdt\", min_child_samples=68,random_state = 62,objective = \"multi-class\",metric = \"multi_logloss\")\n    model.fit(stack_train_X,stack_train_y)\n    preds = model.predict_proba(stack_test_X)\n    preds = pd.DataFrame(preds)\n    stack_df[\"0\"].iloc[j[1]] = preds[0]\n    stack_df[\"1\"].iloc[j[1]] = preds[1]\n    stack_df[\"2\"].iloc[j[1]] = preds[2]\n    stack_df[\"3\"].iloc[j[1]] = preds[3]\n    stack_df[\"4\"].iloc[j[1]] = preds[4]\n    stack_df[\"5\"].iloc[j[1]] = preds[5]\n    stack_df[\"6\"].iloc[j[1]] = preds[6]\n    stack_df[\"7\"].iloc[j[1]] = preds[7]\n    stack_df[\"8\"].iloc[j[1]] = preds[8]\nlgbm_stack = stack_df[[\"0\",\"1\",\"2\",\"3\",\"4\",\"5\",\"6\",\"7\",\"8\"]]\nlgbm_stack.columns = [\"lgbm_Prediction1\",\"lgbm_Prediction2\",\"lgbm_Prediction3\",\"lgbm_Prediction4\",\"lgbm_Prediction5\",\"lgbm_Prediction6\",\"lgbm_Prediction7\",\"lgbm_Prediction8\",\"lgbm_Prediction9\"]","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stack_df = pd.read_csv(\"../input/malware-only-byte/train_data.csv\")\nfor i,j in enumerate(fold.split(X,y)):\n    stack_train_X = X.iloc[j[0]]\n    stack_train_y = y.iloc[j[0]]\n    stack_test_X = X.iloc[j[1]]\n    stack_test_y = y.iloc[j[1]]\n    model = XGBClassifier(booster=\"gbtree\",eta = 0.0975, min_child_weight = 2,random_state = 62, objective = \"multi:softmax\",eval_metric=\"logloss\")\n    model.fit(stack_train_X,stack_train_y)\n    preds = model.predict_proba(stack_test_X)\n    preds = pd.DataFrame(preds)\n    stack_df[\"0\"].iloc[j[1]] = preds[0]\n    stack_df[\"1\"].iloc[j[1]] = preds[1]\n    stack_df[\"2\"].iloc[j[1]] = preds[2]\n    stack_df[\"3\"].iloc[j[1]] = preds[3]\n    stack_df[\"4\"].iloc[j[1]] = preds[4]\n    stack_df[\"5\"].iloc[j[1]] = preds[5]\n    stack_df[\"6\"].iloc[j[1]] = preds[6]\n    stack_df[\"7\"].iloc[j[1]] = preds[7]\n    stack_df[\"8\"].iloc[j[1]] = preds[8]\nxgb_stack = stack_df[[\"0\",\"1\",\"2\",\"3\",\"4\",\"5\",\"6\",\"7\",\"8\"]]\nxgb_stack.columns = [\"xgb_Prediction1\",\"xgb_Prediction2\",\"xgb_Prediction3\",\"xgb_Prediction4\",\"xgb_Prediction5\",\"xgb_Prediction6\",\"xgb_Prediction7\",\"xgb_Prediction8\",\"xgb_Prediction9\"]","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stack_df = pd.read_csv(\"../input/malware-only-byte/train_data.csv\")\nfor i,j in enumerate(fold.split(X,y)):\n    stack_train_X = X.iloc[j[0]]\n    stack_train_y = y.iloc[j[0]]\n    stack_test_X = X.iloc[j[1]]\n    stack_test_y = y.iloc[j[1]]\n    model = CatBoostClassifier(verbose=0)\n    model.fit(stack_train_X,stack_train_y)\n    preds = model.predict_proba(stack_test_X)\n    preds = pd.DataFrame(preds)\n    stack_df[\"0\"].iloc[j[1]] = preds[0]\n    stack_df[\"1\"].iloc[j[1]] = preds[1]\n    stack_df[\"2\"].iloc[j[1]] = preds[2]\n    stack_df[\"3\"].iloc[j[1]] = preds[3]\n    stack_df[\"4\"].iloc[j[1]] = preds[4]\n    stack_df[\"5\"].iloc[j[1]] = preds[5]\n    stack_df[\"6\"].iloc[j[1]] = preds[6]\n    stack_df[\"7\"].iloc[j[1]] = preds[7]\n    stack_df[\"8\"].iloc[j[1]] = preds[8]\ncat_stack = stack_df[[\"0\",\"1\",\"2\",\"3\",\"4\",\"5\",\"6\",\"7\",\"8\"]]\ncat_stack.columns = [\"cat_Prediction1\",\"cat_Prediction2\",\"cat_Prediction3\",\"cat_Prediction4\",\"cat_Prediction5\",\"cat_Prediction6\",\"cat_Prediction7\",\"cat_Prediction8\",\"cat_Prediction9\"]","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stacking_X = pd.concat([xgb_stack,lgbm_stack,cat_stack],axis=1)","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**각 모델들을 이용하여 train set에 대하여 stacking을 진행한 다음 그 예측값을 새로운 학습 데이터로 만들었습니다.**\n\n**Stacking was performed on the train set using each model, and the predicted value was made into new learning data.**","metadata":{}},{"cell_type":"code","source":"model = LGBMClassifier(learning_rate= 0.025, n_estimators = 850, min_child_weight = 1, boosting_type = \"gbdt\", min_child_samples=68,random_state = 62,objective = \"multi-class\",metric = \"multi_logloss\")\nmodel.fit(X,y)\nlgbm_pred = model.predict_proba(test.drop(\"Id\",axis=1))\nlgbm_pred = pd.DataFrame(lgbm_pred)\nlgbm_pred.columns = [\"lgbm_Prediction1\",\"lgbm_Prediction2\",\"lgbm_Prediction3\",\"lgbm_Prediction4\",\"lgbm_Prediction5\",\"lgbm_Prediction6\",\"lgbm_Prediction7\",\"lgbm_Prediction8\",\"lgbm_Prediction9\"]","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = XGBClassifier(booster=\"gbtree\",eta = 0.0975, min_child_weight = 2,random_state = 62, objective = \"multi:softmax\",eval_metric=\"logloss\")\nmodel.fit(X,y)\nxgb_pred = model.predict_proba(test.drop(\"Id\",axis=1))\nxgb_pred = pd.DataFrame(xgb_pred)\nxgb_pred.columns = [\"xgb_Prediction1\",\"xgb_Prediction2\",\"xgb_Prediction3\",\"xgb_Prediction4\",\"xgb_Prediction5\",\"xgb_Prediction6\",\"xgb_Prediction7\",\"xgb_Prediction8\",\"xgb_Prediction9\"]","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = CatBoostClassifier(verbose=0)\nmodel.fit(X,y)\ncat_pred = model.predict_proba(test.drop(\"Id\",axis=1))\ncat_pred = pd.DataFrame(cat_pred)\ncat_pred.columns = [\"cat_Prediction1\",\"cat_Prediction2\",\"cat_Prediction3\",\"cat_Prediction4\",\"cat_Prediction5\",\"cat_Prediction6\",\"cat_Prediction7\",\"cat_Prediction8\",\"cat_Prediction9\"]","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_X = pd.concat([xgb_pred,lgbm_pred,cat_pred],axis=1)","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**test data에 대한 예측값을 생성하였습니다.**\n\n**A predicted value for test data was generated.**","metadata":{}},{"cell_type":"code","source":"model = CatBoostClassifier(verbose=0)\nmodel.fit(stacking_X,y)\npred = model.predict_proba(test_X)\nimsi_df = pd.DataFrame(pred)\nimsi_df = imsi_df.set_index(test[\"Id\"])\nsubmission = pd.concat([sample,imsi_df],axis=1)\nsubmission = submission.drop([\"Prediction1\",\"Prediction2\",\"Prediction3\",\"Prediction4\",\"Prediction5\",\"Prediction6\",\"Prediction7\",\"Prediction8\",\"Prediction9\"],axis=1)\nsubmission.columns = [\"Prediction1\",\"Prediction2\",\"Prediction3\",\"Prediction4\",\"Prediction5\",\"Prediction6\",\"Prediction7\",\"Prediction8\",\"Prediction9\"]\nsubmission.to_csv(\"submission.csv\")","metadata":{"trusted":true,"execution":{"execution_failed":"2025-04-06T16:58:27.429Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**예측값을 기반으로 실제값과 학습시킨다음, test data의 예측값을 이용하여 예측을 진행하였습니다.**\n\n**After learning the actual value based on the predicted value, the prediction was conducted using the predicted value of test data.**","metadata":{}}]}