{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7921029,"sourceType":"competition"},{"sourceId":7584174,"sourceType":"datasetVersion","datasetId":4414761}],"dockerImageVersionId":30664,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Referenced Notebooks\nhttps://www.kaggle.com/code/greysky/home-credit-baseline - GOAT! Thanks for your work \n\n\nhttps://www.kaggle.com/code/ravi20076/homecredit-starter-inference-v1 - \"MakeAgg\" was referenced.\n\n\nhttps://www.kaggle.com/code/peizhengwang/lb-0-57-mod-weight-pure-lgb - the LGBM hyperparameters were referenced.","metadata":{}},{"cell_type":"markdown","source":"## What I worked for\n1. Added code to reduce memory usage to the baseline.\n2. Added the cb_preprocessing code which is preprocessing static_cb\n\n    1) Dropped 'education_88M' and 'maritalst_893M' because static_cb contain columns that are completely redundant\n    \n    2-1) riskassesment_302T: Convert to probability values and change to 0.xx similar to riskassesment_940T.\n    \n    2-2) riskassesment_940T: Use MinMaxScaler to convert values between 0 and 1.\n    \n    ![image.png](attachment:dd67de6f-6604-40cd-b734-d1c4a7463e17.png)\n    \n    3) 3) For the A, L, D columns, combine columns with redundant meanings into a single column by merging values for each row into one row and one column. However, this process takes too long, so I do not recommend it.\n\n4. Use reduce_memory_usage_pl and filter_cols when reading files to reduce memory usage.\n    \n    1) For filter_cols, I also dropped columns containing \"a55475b1\" in more than 95% of their entries.\n    \n5. Changed hyperparameters\n    \n    ","metadata":{},"attachments":{"dd67de6f-6604-40cd-b734-d1c4a7463e17.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAL0AAABaCAYAAADtllcKAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAABkISURBVHhe7d0FtFVVtwfwhd2F3d2B3QV2oKAoqIiNhd0oBhZ+KnYnFiIYmNjdgd3dhYlgvudvfi7GefcB9+CH78Hd6z/GGfecffZeMWvNfc9/7tXsv/5EKiioEP6X0fv48ccfp6FDh6a55porjTvuuH998298++236eeff07TTTddGmeccf46Oubjjz/+SM2aNYvX8PD777/HnM1/ookmSuONN14cN9dffvklTTDBBPFy/a+//hrHvZ9kkkmibefUinL88cePV70YMmRIGjhwYFpiiSWizVr47ptvvgmZj0qb/98gD7JpaEMZvifz3377LU044YQxt6wf+iBj1/oOsi6c5xg9uLYWdNeYXYbR6zwbhYZuvfXW9Omnn6Ydd9wxGjcADXldeuml6bXXXktdu3ZNk046aVzj5XrtZGPxvvY68NlxE6mdHGTBNLwutwu5L983bMN75zs3zyUL/KeffkoPPPBAmnvuudMCCyww7LoMgh8wYEC65557oq0VVlghtWnTJoR8+umnp7fffjstvPDCqXPnziHU3r17p6effjr62mabbVLz5s3TNddck7788su4fooppkitW7dOq6666rC5Dw95bsbz3nvvpW7duqWjjz46xlkrzyeffDJdeOGF6aijjkozzTRTHDcv5+jPObWyAP065lzHsjwht10rc8fAsfy5Vob5c20b+TvH9ZHPz31z4tdffz1tvvnmw+yiFuzo4osvDoeeb7750vbbb5+mn376aIsubrzxxgi8+++/f/r888/T+eefH7ogH+c+++yz6cEHHwz9ap8zHHjggWnWWWf9q4fhY9w/BXnU119/ne644470yiuvpGeeeSaUqJF555033XLLLemuu+5Kn3zySZpjjjnSU089lT766KM0zzzzxHfTTDNNOEjfvn3TE088kSabbLI0+eSTp3vvvTe+f+utt9Iss8ySPvvss3T99denRx55JAQz44wzpueffz4mRjDTTjttGFT//v2H9WfCFH7//ffHBN944430xRdfhFMStjZeeOGFdMMNN6Q333wzzTbbbGnw4MHRhuPG4DzjO+aYY6Kf+eefP/qqhfkzelFWm65faqmloh/Ht9pqq3T33XdHpJ9qqqnSzTffnFq1apXeeeed6KNt27Zp4oknjvmYb7t27cK5pp566lD+8EAG2iZv11EY5S+33HIxXrIiTzIR6Sl3rbXWiv7IxvGbbropPn///fchY/K57rrrQlYCkrHWypP+yIe8Hn744WEyJJd+/fpFf+YhCNx5553ppZdeCtlr3/wZHPkZJ9ncd999ESTZgLn43l/jJ6vTTjstXXXVVaHHOeecM/rLEKG7d++evvrqq5Df7bffHjIhdwGAgRvzoEGD0qabbhpBRX/bbbddjPPHH39Miy++eMzz0UcfjXmsvvrqMT6yGRnCbXVscITIy1999dX0+OOPh6f27NkzFCKyffjhh3GRFOfcc8+NSU455ZQRRQ2CkP71r3+FIi+66KI0wwwzRFsMV/schiII3LlHHHFEtEVJvnfdKaecEpHCsQ8++CCcRGTVzqmnnpr69OmTXn755XT55ZdHGnbooYdGhBUZRENz0fdzzz0XhnLiiSfGmM2BQimjITi5KE4RBCjqMFjXM0JCZ8QMQBQ5/PDD02qrrRZjF3kJfJ111olA4Lp11103otGIDB7ImYEwDFGM45G/OV999dXxVwQTOHIUZawXXHBByNz8OLZUyLzJi+wZsHk89thjcU5D/V122WUR4BjuYYcdFgHsyiuvjO84l+8EEG2+++67YYznnHNOOHOPHj2iT+PkFNpluGzFOLVDF+SvDcbnnJwR1ML8zVmQJFvnsSNOxMHItUWLFnEdOdA/XZDtkksuGfOkp/XWWy/kv8gii6SNN9447KsxhNEbgNeee+6Ztt122/BInWvAEi9SiBIadx5DZMh77LFHnCNCig48nBApyYuxigIGlBXlusUWWyy9//77cS6v9le0s7rI2Xi0yGXiFChS7LTTTuEMlsoNN9wwIhanFG0JWFRgsPo1JtFZBLHCcD5jX3nllcMYG0K7hK9/8zMmzmNO0jvy8B2ZZEWefPLJMQYrSF7yRxXaW3rppdN+++2XZp999mH3CYzkoYceCnm2bNky2ucEHIRil1lmmbTooovGnAQAzi/ocDgGSLa+dw/QUH8cjazIjbMxJiuzQOT4KqusEv3RF5n7PPPMM6edd945ZOEaejQeOmO42iArwaFLly4RWMhMSsgppXlkVgvylD4bb6dOnWI85s4WrF5rrLHGsNSaM7IDsneOtnzOdjuqGKYtg2g4MAM+4IAD0rLLLhu5PMHoVBpBoPIxQhBNGCvhgO/kXKKfCHDFFVeEwXESExE5vvvuuxCOiCl/ZsyMXE5GqSK/yYM+vYAw83tpCmPVdocOHSK/9p1zKE5fkI+JZpbqhiDo8847L+bVvn37SKEYnHmLoBxV1Ft++eVj3PJ8kXO33XaLyEYBfxdWn1oYJ4Vbxo2DPH744YcwTCmVNMSYrLScXMDRBuVzcsbJEH3PYGr153zBRKqx0UYbhY4WWmihtOaaa4bhcXSrr5TCOCDLmjy995fcpRHrr79+2mGHHSIi658+fe8F2uDY5C51qUU21n333Tcde+yxacEFFww96pujSm+kUFI2QdCYyVyG4F7A5+Gt2vUgRsc4RJqsANGZAeYbQIp1c2cZkgowZh6tc15u0hQjQhkMr5d3iUAm41qeLL0hEJOjWIbK0x2noNwfZTN8CuFAnMkYRWmO6EaR8ixvjFS00wbHFQ0Yr7mI+K4RtYzBymDMxlf70jalMAyphcgkUlotKJgBmfcGG2wQK8eLL74YbcvtpVnGDcaq72wojcG8RF8wbvdQFGnplkaYp4Cgb+PhAGRNRs7l8GTsPbm4RoroOsGETGv1l28W6YkhOa7f3J/2rCTGxB7IM8vfuMgyp29kZjWyqnBIaZ9+XeNabdGhlM/qYlWslbkgqQ1pqYyAnDkRnUtxjz/++Jg7XbA3/xhgl2eccUYEGbowJo5lfMZVL+K/NwbAw1yoIdGMh+pE1LB0mgThe+98EyQ0wmbIIigBus4duM/aMTHphYHmqKgfhmjZEl0YiUm73jn6oBgC14ZrnM8ZjEmUIETjERW0YQz6hTwXY5F3EwrDdFzksELlSAZunjp27Bj9Mn7t6s+4rGTa0J4xmrtxGINXFrr+nUeh5puj3Yjg2tq5m4e+zMEYyME5Pue2nUeuxkA+8nJG5nv6AGPLbdJNlmfWHzhG9ubEecmmtj/Xkzv5669W/tqQEbALOnDcy2djEmy0xRHo05xcZ4WyWmYYt/sBY9eOsWi7Vm7mZ67kaWzapUNjNjayJyty1FeWQWOo3I9TBO8GNQvX9EVnuWvD9K5g9IADuQ/kGBnk796MAf9fo3JGX1Aw8jW4oKAJohh9QeVQjL6gcihGX1A5FKMvqByK0RdUDsXoCyqHYvQFYzxG909JxehHA/wU7lUw+oEXhOODSoGWgEuEa4W2kYFGgV6CzlAPitH/TYg++YW7notjao8X/GcgQyRAHCqcJkQ0xEJ1BagkOdiQv6oy5Ld6EJVTf70vqAOEjNqLTi3iYBEq3qAENF9Uag6Am479WEtsK6gfDL5Pnz5RMYVchmzG+FGRMU6R2xDd8Hqch5imNiGzVkeGEulHEZZQVUOqeLD6UI3x0VFdFWqIOshUKLGYhAV/D5ieKMVoyptsskmwTEXy3XffPQp3sDcBvRu3X/BxTT0oRj+KQJVWm6kwItcQMH5sQd95MXjU3nqVUDB8iOQow+jkmRWrZFQBTC4oEeVRqOlDJZc0qDEUox9F4JZLcRRIELRiGcUaIn6+uSrGPvqh9JB88em9rKJSRyWK0h4rMO4/3TSGYvSjiFyfqWKKcStMVh3mSRKKp1UQiU4FowfSG8atVNMK6nEgqsPWXnvtqL1WZ6tMUqWYp0WI+o2h8OlHI4iSI+S/BaMfAo5UJ6c7tahX7sXoCyqHkt4UVA7F6Asqh2L0BZVDs0GDBpWcvqAy8KiRZt9//30x+oLKwPN4yn9vCiqHktMXVA7F6Asqh2L0BZVDMfo6UG57/jPUK7/hnfdPyH6MupHFlPOscps2eBQ3gtHWW28dPAtMxoMPPji403a8sFsFyukhhxwShCQEMIUEdibx1zY0eNj5ScYjA/betddeG1veoLPusssu8Whrj5E+66yz4t9cnn1vexcbIyhb88jrvD2RJwLj1CNC2cnDJhIrrrjicDeAaGrwHH90X5s0rLTSSsF3zw/C9TRh+sR9Jw/kMAxJMlVkQ882pPBg17PPPjsegU6mCkXoQXGOSilks1133TUemW4DDvwbL/sc2DuA3nBuPJuf3hrDGBXpGbdKGEboxXgZtOIupWIqZjznnoAJjlETuGO+Uyup0MDmEapt6n0ibt7Nw4YEBIeiaqMGXO0jjzwyjnE6SkQnVkDie0buue4ETeiZ4mrczq8C7Nriic82qlAxxgky8NvtRWDXFobvRW6CGX3am8BjvNGzVTxxEOdrQyUa5xHg1CYwfs/nt3kGR0Ht9r39BhSVsAUFJfVgjDJ6z8a3I4mo4OH8IgBBqYm0VxLDBAbFKfxldKKCR3DjUjN8FF8RnoHWgjEq/FBkzEAz9MFxrB5enrHOsfShSspeUxxSsYi+FCErVrD3koIGivc57/DBCThdU4QIS4Ze3iuRtEmDegIrIvlkiNA2eVD1hG5tRVTSZxW3L5VNPlCBBSoF4LfddlvoU7CiD86kTRVpbIF9MHQbgSjkoQ+rB4egIw5TD8aoGlmGhT/NswmPEZswQxIReLUUg0GqnRRVHbetj2WRs9i2x44bojejtORmjjXBSE9U1IssNlMAXHhCldbYgcUeTRSkCt8Wl8bjWtvYiPJK02xlg78tjeKMIr1CB8+6dz4lN0V6MaO0S0jeF4y+zL9Xr16xX5nULoORcg7F3fRAPoo/yIWupSYMX1pDz1JZtbA++54utC+lEcwYOufSN5tg9DmV0g9nsMNNY3If425kpTMmbODSDCkF4SnLE51Fa7vLHXfccVFIoIAj73bCmEXb/OIItdFeoYGl1rUiVIZ7A06gn1wAohjEaqFtx8Htj3Plo5xSSiWKUYax+t7Y8/ibIsxZamLXRu/d39iyyG4uDF70JxMBiaGeeeaZIUcRnnzcA3ESTiBACGCuY8DknPfPEpCkj9riCEEf+FOm7pkEHQ6hD3m9cwVD7bKPxjBGPw2BEERrns54ebefkW3AJeKIFDY7E1nlgT5v/+cNJqPj9RzGPkbZkDmUNr2ck0HgUhY1l/JQEZ1SRH+F3pZahi1SEard9UR8jiK1kS6JWITfVI09g9yyDMlTamFFFBzIjzE6Rk7kRldWBKupvyKzbTrplNFuueWWYdCuccw2rVJcsrbRnrTHeYrEOQOjtyIIgtqiczqjbzqhp8YwxtMQ3NjmyMrIGFU+xljz7tCWNzm+fNCU5On+UkI9yDevIo0IZBm2slAc4VIYxYpkopuIT+nG4pVXiqrBf2MEByBv+mDwZOY+h6yyiZGhAEQ39EVmdOh7aaog5V4o/8dNu4KZz467x3JMOxmOCYj0lZ2qMRTuTUHlUH6cKqgcitEXVA7F6Asqh2Z/3vyVnL6gMoh/J//444/F6AsqA//qbDZ48OBi9AWVgV/nS05fUDkUoy+oHIrRF1QOY2VO70fkTENASEIb8HM3erGfoTH2cHRQVv0kjtCEIuAnbTmd6+qBn9hRE1yrDzdBSE1+XvdzO1oCCgL6A1qCn9aNx3F9ZM5PFYGPhEcDCGJYqxlDhgwJ/ZEhfSGUoSaggNAVegdagnPI0vX+0jGCGn3QMV2ijzjmvzJIhplROyL4ftyuXbuOddvv4L+olGKEqmZatGgRf5HGGCn6KQ5O7969w0gZn5dr0I7rMXqKUahCqMhSjBu344ILLghl6BvRDc/fX9wQ5/Xr1y+Uo39EqiqCfvr37x/EPIQ8+kDIY8yYl+ojevXqFQQxAYxeBBhFJbhVjL9Pnz5h+GQqmOBQXXjhhcG9QT3WB9ox2jKjR0yj+8YIZ5xsrE1vTBrJiHGK8KIvLrWyPbRhhic6oB77HtsPcy8T1BqDdnHy27RpE07F0DkAoYsyIjoGKAUpYcxAUlMHUM/mAE0VDJixd+jQIViUVtxMSkMQEyCaN28eMiJnRDObKyivRCemLxHZta1bt47CHvrkOB07dgwaszbJ2o4kGJgcyLW17NkRYaw0esZu8oSAWswAUYDVWeLKKz8TBXi+qCLCoKcSquhTWzUlNclMyVoutkitrE1ljuhu6cTZlzYpYhB9RCabMqC1cjoO1rlz56Az1+tcTRGM3kpoRfXy3jGw+rVs2XLYaos9ywEUAikM2mKLLSJYSXdy6sJhrAAcwMqgmAiN2+rNWZQP4vijntcTbMbK9IYACJLxM0LUUgZ+wgknhECV+FnmNtxwwyhewL2XK1oyGTdhSnOAcC+++OIwYArI1VRKDqVM++yzT6Q1or4IYzVRPOH8du3aRW0ohVp5FK0Yj+hvbFWFgCAtsUIyUrsxKv/MhSACjXpaK7LUhyx32mmnCDQq1Bi31MZ7EZ1elAKK4nTiegYuSClm6dKlS7RvNaf/kcl+rE5vaiHfEznkdYRtYqKAGx4CZ4iiPEeRijB6wgfRe6+99opowVEy09oynNthxDltIXgO4MZWmuO4VYAC8rVVh3sc8pB6SDmsggzRCtkwEpOflZic3exakaUyjjFsKSk+ve+vuOKKWMVtuyPQuFawoR/pkD45RGMYKyN9QzB4E1e3SmBSDukFgTB6211SBGMXQUSD/N8Exs+4IwLU5IOiDiehOIpo1apVtEmwVhXVVfp1HQX47JqCf1e8SVk8lUIkt+IKPIIFgyVzq4GVUSEJPal+EjTk8Yzcf9oYOMOWwtCNKilRn445j+vpRMrqmPMEo5GBvpoUDYHQcgQfnfin2m3qyCtfPbIbnoz/CblbEZpEepPxTxlmMfi/B3KrV3bDO++fknuTMvqCgnpQjL6gcihGX1A5NBs6dGj5P1tBZeCHr/IIkILKoaQ3BZVDMfqCyqEYfUHlUIy+oHIY629kcTDwOAD/Bu8DXwNPG7/G9HBj8GTcuTtfVY+fo+thQrpe4YNfB13jvZdKKXwQfSGn6ctf3+F36MvfpgocJHIkh4bAVcKnIfP8qyrSGT3h5JCb71HCcaLI2Ht6yg9h9ZlsgZ5wbnCn8jHyxo1yLT4VglpuuzGM0Y/qrgcEdckll8TjogmQMM4///ygoKKjIiPZngXw71FV+/btG/sdcZDG4KnFBx10UDiL67t16xaUV8rB0LQJBLoy6jJymkoeZKnME2+KQO6yuYXKNCSvWjD4Hj16BF3YTiQMGPNR1VPPnj2D/MdwfXaevbzopHv37vGceyxZ1O0zzjgj9KoCC5EQFZwe0cM96luAw5glf2Px3HukttqCnhGhSaQ3BMvTeTkmn5pLBsdQcenbtm0bUVgxAsFyhsyyHBlcT7FoySITBxPhMC5V+3imujpQGwKIaF7G0JS59AyYc4uyOepmiLp48qqmGKDPXgMHDgx5oQeTKUYlo6UT7ZEjnahEIzsUccVAPtv1REDRn5dtezgMZ6NzDnDSSSfF5mz49PVgrDd6KYe9iRgimipOfKdOnUJAIDp72D8jF/0zB175H4UAwYtelkmOQTHAgPfee+8oHgFtiO6qeUQ0SrGnEsfyWbWWqimbgFl1miLIxI4uAknDlVIxDyO06RmjBnIRqW2NlNMdNGwUYikgXUiTyE57NsNQ6wwCjUhOtzmNEfnx6vH06UmREEeg26ynxjDWGz3hmqzNzRiaaFKbSztOCcC48a/tcCFFkfeDXNEmXwqUCTUfh9oVQdscimKkPQzesipySaM4nYil8KGpFoUzWsbacDWzAtryUskkx7Daystx4nHqBSEOINjQT74+r5B5VWDIrgerhmMM28ptTzCbuukDt97KK6DRsUDDWWp1NyI0ifRmZOAUijsUG2+22WahDHmfG6u8PMu/RS57oLZv3z5WjxFBlZbVwA2cVYUSrR65sKFqYJRuUM2fYSuzFAhEYvm5gMLY3UfJ3QcMGPA/0iIOILBYJQQl92ZZP7bPVHPMyeiNoSvnpE+y1q6CHhveWX20UVmjz3f7cj6ClOerxxSJRWFlaxSUl2CRxft8TOSpheNyVLAUE66USuTXHkWp2MkRqgog27wljseeyMNtXO3mVEpo3yirntzbXr92aWSYCu0ZMRkLIGQmUtuvSpE/B5AKcRppUN6AWvrjXG3b6Fr0d0xVllJPN8Jqa4f336SGaJLcG1HAcku44D0lcQbfiUwMNn/fGFxPTM53vSjkekqgdJ8pSx9VATmIqgICGZBFTivl4l7ZuIH8/DtXUHGuzwJIvuGlE6tCTlG1CbU6skJYZR0T1LQjqEmjvHfPUI8OCuGsoHKoznpcUPAXitEXVA7F6Asqh1I5VVApxM314LL9TkGF0OSee1NQUA+K0RdUDsXoCyqGlP4bV3JsM2PIyvkAAAAASUVORK5CYII="}}},{"cell_type":"code","source":"import os\nimport gc\nfrom glob import glob\nfrom pathlib import Path\nfrom datetime import datetime\n\nimport numpy as np\nimport pandas as pd\nimport polars as pl\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom sklearn.base import BaseEstimator, RegressorMixin\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import MinMaxScaler\nfrom sklearn.metrics import roc_auc_score\n\nimport joblib\n\nimport lightgbm as lgb\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2024-03-21T02:52:13.775049Z","iopub.execute_input":"2024-03-21T02:52:13.775355Z","iopub.status.idle":"2024-03-21T02:52:20.402801Z","shell.execute_reply.started":"2024-03-21T02:52:13.775329Z","shell.execute_reply":"2024-03-21T02:52:20.402003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pre-Fitted Voting Model","metadata":{}},{"cell_type":"code","source":"class VotingModel(BaseEstimator, RegressorMixin):\n    def __init__(self, estimators):\n        super().__init__()\n        self.estimators = estimators\n        \n    def fit(self, X, y=None):\n        return self\n    \n    def predict(self, X):\n        y_preds = [estimator.predict(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)\n    \n    def predict_proba(self, X):\n        y_preds = [estimator.predict_proba(X) for estimator in self.estimators]\n        return np.mean(y_preds, axis=0)","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.404486Z","iopub.execute_input":"2024-03-21T02:52:20.404762Z","iopub.status.idle":"2024-03-21T02:52:20.411313Z","shell.execute_reply.started":"2024-03-21T02:52:20.404738Z","shell.execute_reply":"2024-03-21T02:52:20.410435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pipeline","metadata":{}},{"cell_type":"code","source":"class Pipeline:\n    @staticmethod\n    def set_table_dtypes(df):\n        for col in df.columns:\n            if col in [\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Int64))\n            elif col in [\"date_decision\"]:\n                df = df.with_columns(pl.col(col).cast(pl.Date))\n            elif col[-1] in (\"P\", \"A\"):\n                df = df.with_columns(pl.col(col).cast(pl.Float64))\n            elif col[-1] in (\"M\",):\n                df = df.with_columns(pl.col(col).cast(pl.String))\n            elif col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col).cast(pl.Date))            \n\n        return df\n    \n    @staticmethod\n    def handle_dates(df):\n        for col in df.columns:\n            if col[-1] in (\"D\",):\n                df = df.with_columns(pl.col(col) - pl.col(\"date_decision\"))\n                df = df.with_columns(pl.col(col).dt.total_days())\n                \n        df = df.drop(\"date_decision\", \"MONTH\")\n\n        return df\n    \n    @staticmethod\n    def filter_cols(df):\n        for col in df.columns:\n            if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n                isnull = df[col].is_null().mean()\n                if col[-1]=='M':\n                    specific_value_ratio = df.filter(pl.col(col) == \"a55475b1\").height / df.height\n                    if specific_value_ratio > 0.95:\n                        df = df.drop(col)\n                elif isnull > 0.95:\n                    df = df.drop(col)\n\n        for col in df.columns:\n            if (col not in [\"target\", \"case_id\", \"WEEK_NUM\"]) & (df[col].dtype == pl.String):\n                freq = df[col].n_unique()\n\n                if (freq == 1) | (freq > 200):\n                    df = df.drop(col)\n\n        return df\n\n    @staticmethod\n    def reduce_memory_usage_pl(df):\n        \"\"\" Reduce memory usage by polars dataframe {df} with name {name} by changing its data types.\n            Original pandas version of this function: https://www.kaggle.com/code/arjanso/reducing-dataframe-memory-size-by-65 \"\"\"\n        print(f\"Memory usage of dataframe is {round(df.estimated_size('mb'), 2)} MB\")\n        Numeric_Int_types = [pl.Int8,pl.Int16,pl.Int32,pl.Int64]\n        Numeric_Float_types = [pl.Float32,pl.Float64]    \n        for col in df.columns:\n            if col == 'case_id': continue\n            try:\n                col_type = df[col].dtype\n                if col_type == pl.Categorical:\n                    continue\n                c_min = df[col].min()\n                c_max = df[col].max()\n                if col_type in Numeric_Int_types:\n                    if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                        df = df.with_columns(df[col].cast(pl.Int8))\n                    elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                        df = df.with_columns(df[col].cast(pl.Int16))\n                    elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                        df = df.with_columns(df[col].cast(pl.Int32))\n                    elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                        df = df.with_columns(df[col].cast(pl.Int64))\n                elif col_type in Numeric_Float_types:\n                    if c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                        df = df.with_columns(df[col].cast(pl.Float32))\n                    else:\n                        pass\n                # elif col_type == pl.Utf8:\n                #     df = df.with_columns(df[col].cast(pl.Categorical))\n                else:\n                    pass\n            except:\n                pass\n        print(f\"Memory usage of dataframe became {round(df.estimated_size('mb'), 2)} MB\")\n        return df\n    \n    @staticmethod\n    def convert_to_prob(value):\n        if pd.isnull(value):\n            return np.nan  # NaN 값은 그대로 반환\n        low, high = value.replace('%', '').split(' - ')\n        return (float(low) + float(high)) / 200 \n    \n    @staticmethod\n    def cb_preprocessing(df, cols):\n        df = df.to_pandas()\n        \n        for key in cols.keys():\n            if key[0]=='T':\n                df['riskassesment_302T'] = df['riskassesment_302T'].apply(Pipeline.convert_to_prob)\n                scaler = MinMaxScaler()\n                df[['riskassesment_940T']] = scaler.fit_transform(df[['riskassesment_940T']])\n            elif key[0]=='M':\n                df = df.drop(columns=cols['M_cols'][0])\n            else:\n                for col_list in cols[key]:\n                    print(col_list)\n                    col_name = col_list[0].split(\"_\")[0]\n                    df[col_name] = df.apply(lambda row: row[col_list].dropna().mean(), axis=1)\n                    df = df.drop(columns=col_list)\n        df = pl.from_pandas(df)\n        return df","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.412547Z","iopub.execute_input":"2024-03-21T02:52:20.412813Z","iopub.status.idle":"2024-03-21T02:52:20.439815Z","shell.execute_reply.started":"2024-03-21T02:52:20.412781Z","shell.execute_reply":"2024-03-21T02:52:20.438844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Automatic Aggregation","metadata":{}},{"cell_type":"code","source":"class Aggregator:\n    @staticmethod    \n    def MakeAgg(df: pl.DataFrame):\n        \"\"\"\n        Code for aggregating files with a depth of 0 or more.\n\n        Note: \n        1. For all columns, create new columns by applying [max, min, first, last].\n        2. Apply the mean function to the [P, A, D] columns.\n        3. Apply the mode function to the [M] column.\n        \"\"\"\n        all_agg = []\n        df_cols = df.columns\n\n        all_agg.extend([method(col).alias(f\"{method.__name__}_{col}\") \\\n                        # for method in [pl.max, pl.min, pl.first, pl.last] \\\n                        for method in [pl.max] \\\n                        for col in df_cols if col[-1] in (\"P\", \"A\", \"D\", \"M\", \"T\", \"L\") or \"num_group\" in col\n                        ]\n                       )\n        all_agg.extend([pl.col(col).mean().alias(f\"mean_{col}\") for col in df_cols if col.endswith((\"P\", \"A\", \"D\"))])\n        all_agg.extend([pl.col(col).std().alias(f\"std_{col}\") for col in df_cols if col.endswith((\"P\", \"A\", \"D\"))])\n        # all_agg.extend([pl.col(col).drop_nulls().mode().first().alias(f\"mode_{col}\") for col in df_cols if col.endswith(\"M\")])\n        return df.group_by(\"case_id\").agg(all_agg)","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.442327Z","iopub.execute_input":"2024-03-21T02:52:20.442650Z","iopub.status.idle":"2024-03-21T02:52:20.455710Z","shell.execute_reply.started":"2024-03-21T02:52:20.442620Z","shell.execute_reply":"2024-03-21T02:52:20.454722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### File I/O","metadata":{}},{"cell_type":"code","source":"def read_file(path, depth=None): \n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    \n    if depth in [1, 2]:\n        # pl.max, pl.min, pl.first, pl.last, mean, mode 함수 MakeAgg에서 골라서 사용\n        df = Aggregator.MakeAgg(df)\n    \n    df = df.pipe(Pipeline.reduce_memory_usage_pl)\n    return df\n\ndef read_files(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        chunks.append(pl.read_parquet(path).pipe(Pipeline.set_table_dtypes))\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    if depth in [1, 2]:\n        # pl.max, pl.min, pl.first, pl.last, mean, mode 함수 MakeAgg에서 골라서 사용\n        df = Aggregator.MakeAgg(df)\n\n    df = df.pipe(Pipeline.reduce_memory_usage_pl)\n    return df\n\ndef read_train_file(path, depth=None): \n    df = pl.read_parquet(path)\n    df = df.pipe(Pipeline.set_table_dtypes)\n    \n    if depth in [1, 2]:\n        # pl.max, pl.min, pl.first, pl.last, mean, mode 함수 MakeAgg에서 골라서 사용\n        df = Aggregator.MakeAgg(df)\n    \n    df = df.pipe(Pipeline.reduce_memory_usage_pl)\n    df = df.pipe(Pipeline.filter_cols)\n    return df\n\ndef read_train_files(regex_path, depth=None):\n    chunks = []\n    for path in glob(str(regex_path)):\n        chunks.append(pl.read_parquet(path).pipe(Pipeline.set_table_dtypes))\n        \n    df = pl.concat(chunks, how=\"vertical_relaxed\")\n    if depth in [1, 2]:\n        # pl.max, pl.min, pl.first, pl.last, mean, mode 함수 MakeAgg에서 골라서 사용\n        df = Aggregator.MakeAgg(df)\n\n    df = df.pipe(Pipeline.reduce_memory_usage_pl)\n    df = df.pipe(Pipeline.filter_cols)\n    return df","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.456844Z","iopub.execute_input":"2024-03-21T02:52:20.458930Z","iopub.status.idle":"2024-03-21T02:52:20.471017Z","shell.execute_reply.started":"2024-03-21T02:52:20.458897Z","shell.execute_reply":"2024-03-21T02:52:20.470092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Engineering","metadata":{}},{"cell_type":"code","source":"cols = {\n    'T_cols' : [['riskassesment_302T','riskassesment_940T']],\n    'M_cols': [['education_88M','maritalst_893M']],\n    'D_cols': [\n        ['assignmentdate_238D', 'assignmentdate_4527235D', 'assignmentdate_4955616D'],\n        ['birthdate_574D', 'dateofbirth_337D', 'dateofbirth_342D'],\n        ['responsedate_1012D', 'responsedate_4527233D', 'responsedate_4917613D']\n    ],\n    'A_cols': [['pmtaverage_3A', 'pmtaverage_4527227A', 'pmtaverage_4955615A']],\n    'L_cols': [['pmtcount_4527229L', 'pmtcount_4955617L', 'pmtcount_693L']],\n    \n    \n}\ndef feature_eng(df_base, depth_0, depth_1, depth_2):\n    df_base = (\n        df_base\n        .with_columns(\n            month_decision = pl.col(\"date_decision\").dt.month(),\n            weekday_decision = pl.col(\"date_decision\").dt.weekday(),\n        )\n    )\n        \n    for i, df in enumerate(depth_0 + depth_1 + depth_2):\n        df_base = df_base.join(df, how=\"left\", on=\"case_id\", suffix=f\"_{i}\")\n        \n    df_base = df_base.pipe(Pipeline.handle_dates)\n\n    # Preprocessing code for CB files. Uncomment and use as needed.\n    # cb_list = [item for sublist in cols.values() for sublist_item in sublist for item in sublist_item]\n    # preprocessed_cb = Pipeline.cb_preprocessing(df_base[cb_list + ['case_id']], cols)\n    # df_base = df_base.drop(columns=cb_list)\n    # df_base = df_base.join(preprocessed_cb, how=\"left\", on=\"case_id\")\n    \n    return df_base","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.472271Z","iopub.execute_input":"2024-03-21T02:52:20.473119Z","iopub.status.idle":"2024-03-21T02:52:20.484678Z","shell.execute_reply.started":"2024-03-21T02:52:20.473087Z","shell.execute_reply":"2024-03-21T02:52:20.483722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def to_pandas(df_data, cat_cols=None):\n    df_data = df_data.to_pandas()\n    \n    if cat_cols is None:\n        cat_cols = list(df_data.select_dtypes(\"object\").columns)\n    \n    df_data[cat_cols] = df_data[cat_cols].astype(\"category\")\n    \n    return df_data, cat_cols","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.488164Z","iopub.execute_input":"2024-03-21T02:52:20.488499Z","iopub.status.idle":"2024-03-21T02:52:20.498197Z","shell.execute_reply.started":"2024-03-21T02:52:20.488467Z","shell.execute_reply":"2024-03-21T02:52:20.497177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configuration","metadata":{}},{"cell_type":"code","source":"ROOT            = Path(\"/kaggle/input/home-credit-credit-risk-model-stability\")\nTRAIN_DIR       = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/train\")\nTEST_DIR        = Path(\"/kaggle/input/home-credit-credit-risk-model-stability/parquet_files/test\")","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:52:20.499442Z","iopub.execute_input":"2024-03-21T02:52:20.499693Z","iopub.status.idle":"2024-03-21T02:52:20.508124Z","shell.execute_reply.started":"2024-03-21T02:52:20.499671Z","shell.execute_reply":"2024-03-21T02:52:20.507263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train Files Read & Feature Engineering","metadata":{}},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_train_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_train_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_train_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_train_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_train_files(TRAIN_DIR / \"train_credit_bureau_a_1_*.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_train_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_train_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n        # read_train_files(TRAIN_DIR / \"train_credit_bureau_a_2_*.parquet\", 2),\n        # read_train_file(TRAIN_DIR / \"train_person_2.parquet\", 2),\n        # read_train_file(TRAIN_DIR / \"train_applprev_2.parquet\", 2)\n    ]\n}","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:56:51.003471Z","iopub.execute_input":"2024-03-21T02:56:51.003829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = feature_eng(**data_store)\n\nprint(\"train data shape:\\t\", df_train.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del data_store\n\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Test Files Read & Feature Engineering","metadata":{}},{"cell_type":"code","source":"data_store = {\n    \"df_base\": read_file(TEST_DIR / \"test_base.parquet\"),\n    \"depth_0\": [\n        read_file(TEST_DIR / \"test_static_cb_0.parquet\"),\n        read_files(TEST_DIR / \"test_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TEST_DIR / \"test_applprev_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_a_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_tax_registry_c_1.parquet\", 1),\n        read_files(TEST_DIR / \"test_credit_bureau_a_1_*.parquet\", 1),\n        read_file(TEST_DIR / \"test_credit_bureau_b_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_other_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_person_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_deposit_1.parquet\", 1),\n        read_file(TEST_DIR / \"test_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TEST_DIR / \"test_credit_bureau_b_2.parquet\", 2),\n        # read_files(TEST_DIR / \"test_credit_bureau_a_2_*.parquet\", 2),\n        # read_file(TEST_DIR / \"test_person_2.parquet\", 2),\n        # read_file(TEST_DIR / \"test_applprev_2.parquet\", 2)\n    ]\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test = feature_eng(**data_store)\n\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del data_store\n\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Elimination","metadata":{}},{"cell_type":"code","source":"# df_train = df_train.pipe(Pipeline.filter_cols)\ndf_test = df_test.select([col for col in df_train.columns if col != \"target\"])\n\nprint(\"train data shape:\\t\", df_train.shape)\nprint(\"test data shape:\\t\", df_test.shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:55:09.934425Z","iopub.execute_input":"2024-03-21T02:55:09.934772Z","iopub.status.idle":"2024-03-21T02:55:09.943322Z","shell.execute_reply.started":"2024-03-21T02:55:09.934743Z","shell.execute_reply":"2024-03-21T02:55:09.942248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Pandas Conversion","metadata":{}},{"cell_type":"code","source":"df_train, cat_cols = to_pandas(df_train)\ndf_test, cat_cols = to_pandas(df_test, cat_cols)","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:55:10.318384Z","iopub.execute_input":"2024-03-21T02:55:10.319231Z","iopub.status.idle":"2024-03-21T02:55:27.767736Z","shell.execute_reply.started":"2024-03-21T02:55:10.319193Z","shell.execute_reply":"2024-03-21T02:55:27.766908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training K-Fold Voting Model","metadata":{}},{"cell_type":"code","source":"X = df_train.drop(columns=[\"target\", \"case_id\", \"WEEK_NUM\"]) # + deprecated_features)\ny = df_train[\"target\"]\nweeks = df_train[\"WEEK_NUM\"]","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:55:27.769282Z","iopub.execute_input":"2024-03-21T02:55:27.769591Z","iopub.status.idle":"2024-03-21T02:55:28.890731Z","shell.execute_reply.started":"2024-03-21T02:55:27.769564Z","shell.execute_reply":"2024-03-21T02:55:28.889906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_train\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:55:28.892098Z","iopub.execute_input":"2024-03-21T02:55:28.892783Z","iopub.status.idle":"2024-03-21T02:55:29.160837Z","shell.execute_reply.started":"2024-03-21T02:55:28.892744Z","shell.execute_reply":"2024-03-21T02:55:29.159950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv = StratifiedGroupKFold(n_splits=10, shuffle=False)\n\nparams = {\n    \"boosting_type\": \"gbdt\",\n    \"objective\": \"binary\",\n    \"metric\": \"auc\",\n    \"max_depth\": 10,\n    \"learning_rate\": 0.05,\n    \"max_bin\": 255,\n    \"n_estimators\": 5000,\n    \"colsample_bytree\": 0.8, \n    \"colsample_bynode\": 0.8,\n    \"verbose\": -1,\n    \"random_state\": 42,\n    \"reg_alpha\": 0.1, \n    \"reg_lambda\": 10, \n    \"extra_trees\":True,\n    'num_leaves':64\n    #\"device\": \"gpu\",  # Uncomment if you want to use GPU for training\n}\n\nfitted_models = []\n\nfor idx_train, idx_valid in cv.split(X, y, groups=weeks):\n    X_train, y_train = X.iloc[idx_train], y.iloc[idx_train]\n    X_valid, y_valid = X.iloc[idx_valid], y.iloc[idx_valid]\n\n    model = lgb.LGBMClassifier(**params)\n    model.fit(\n        X_train, y_train,\n        eval_set=[(X_valid, y_valid)],\n        callbacks=[lgb.log_evaluation(100), lgb.early_stopping(50)]\n    )\n\n    fitted_models.append(model)\n\nmodel = VotingModel(fitted_models)","metadata":{"execution":{"iopub.status.busy":"2024-03-21T02:55:29.162907Z","iopub.execute_input":"2024-03-21T02:55:29.163198Z","iopub.status.idle":"2024-03-21T02:56:21.568314Z","shell.execute_reply.started":"2024-03-21T02:55:29.163175Z","shell.execute_reply":"2024-03-21T02:56:21.566878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prediction","metadata":{}},{"cell_type":"code","source":"X_test = df_test.drop(columns=[\"WEEK_NUM\"])\nX_test = X_test.set_index(\"case_id\")\n\ny_pred = pd.Series(model.predict_proba(X_test)[:, 1], index=X_test.index)","metadata":{"execution":{"iopub.status.busy":"2024-03-20T12:44:09.512783Z","iopub.status.idle":"2024-03-20T12:44:09.513116Z","shell.execute_reply.started":"2024-03-20T12:44:09.512953Z","shell.execute_reply":"2024-03-20T12:44:09.512970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Submission","metadata":{}},{"cell_type":"code","source":"df_subm = pd.read_csv(ROOT / \"sample_submission.csv\")\ndf_subm = df_subm.set_index(\"case_id\")\n\ndf_subm[\"score\"] = y_pred","metadata":{"execution":{"iopub.status.busy":"2024-03-20T12:44:09.514590Z","iopub.status.idle":"2024-03-20T12:44:09.514940Z","shell.execute_reply.started":"2024-03-20T12:44:09.514774Z","shell.execute_reply":"2024-03-20T12:44:09.514792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Check null: \", df_subm[\"score\"].isnull().any())\n\ndf_subm.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-20T12:44:09.516283Z","iopub.status.idle":"2024-03-20T12:44:09.516649Z","shell.execute_reply.started":"2024-03-20T12:44:09.516486Z","shell.execute_reply":"2024-03-20T12:44:09.516503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_subm.to_csv(\"submission.csv\")","metadata":{"execution":{"iopub.status.busy":"2024-03-20T12:44:09.518381Z","iopub.status.idle":"2024-03-20T12:44:09.518711Z","shell.execute_reply.started":"2024-03-20T12:44:09.518546Z","shell.execute_reply":"2024-03-20T12:44:09.518562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}