{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# A beginner's approach to: Home Credit - Credit Risk Model Stability\n\nThis is an attempt to create a *beginner friendly Quick-start notebook* for this competetion.\n\nFoucs is on:\n\n* Why aspect (strategy and steps) of things.\n* Simple code, with lots of comments.\n\nSteps:\n\n* Look at available data, sanity + Base EDA\n* Define strategy, helper functions for data preperation and transformations\n* Fit to model and observe results\n* Find ways to improve the model (Feature Engineering, Parameter tuning etc.)\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## Hmm.. so, what do we have here?\nLet's take a quick look at availble files, and a quick-fix on path config....","metadata":{}},{"cell_type":"code","source":"# Utilizing the start function provided by kaggle, with some simplification and explaination\n# Load base libraries for file/ folder access and basic functions (Numpy for linear algebra, pandas forfile/ data processing)\n\nimport os, sys\nimport numpy as np\nimport pandas as pd\n\n# Input data files are available in the read-only \"../input/\" directory\n# Below function return filenames in given directory, which you can browse as well by going to folder structure\n# Check the files available in input directory by running this code\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:13.929708Z","iopub.execute_input":"2024-02-15T04:15:13.930240Z","iopub.status.idle":"2024-02-15T04:15:15.584030Z","shell.execute_reply.started":"2024-02-15T04:15:13.930193Z","shell.execute_reply":"2024-02-15T04:15:15.582443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nSo, from the above output we can observe there is folder segregation for:\n\n* training files (train folder)\n* test files (test folder)\n\nWhich is further divided based on file format (csv, parquet). \n\nThere are two reference files as well.\n\n1. sample_submission.csv\n2. feature_definitions.csv\n\n**It's a good practice to read all the context provided on competetion data and rules before starting our work. Reference files and rules is something we'll keep refering to in this notebook.**\n\n![image.png](attachment:5ac73624-e733-4a41-ae64-71b226eb7637.png)","metadata":{},"attachments":{"5ac73624-e733-4a41-ae64-71b226eb7637.png":{"image/png":"iVBORw0KGgoAAAANSUhEUgAAAWoAAADrCAYAAABAQ9wqAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAADbPSURBVHhe7d0PXBNXujfw31YDEl2CCha1+AdlcanWatEu0it2xb2l75V2ZfdV7yp9FduFtaCiLNai+IdaXSu6/ilq1W7FW7UVt2Ir9qpdsAoVEaQgWwvSCqJUUAmFgESc95zJBAIkIaDIAM/388mHOTOTM2dOwpOTk0meXwgMCCGEyNYT0l9CCCEyRYGaEEJkjgI1IYTIHAVqQgiROQrUhBAicxSoCSFE5ihQE0KIzFGgJoQQmaNATQghMtd8oK6tQYW6gt1qpBWEEEIep+YDdUkCVr36Kl56dSeypVWP1Z08XEw+j4sZxaCXCkJIVyT/qY/8BCx4+20s2H0eFdIqQgjpSloRqCuQfXArPtj6KbLv3EH2kZ1YFRiG6N0J+FEfSdUZOLKV7/MVflTn4eTGMMwNXYsjycXSDpy+nq1I+kFahQIkifdjdatZ8Yev8ME/M3Sbik7jH/r1hBDShbQqUP/49VF8dOQ04v4WiEAWPE9eYYH5fzZi7rKjuMN3qSpAyhG+z36sWxSIVZ9n4PuMrxD99lwsOFgg1lJfz1Fk3ZZW4TayxPt9jR+rePEyPkqW9lfn4Ih+PSGEdCEPMfWRh+8dF+PEqVP4au+f4cHW1GSdRoNBMwvbI+d9hrP/OoUT7/igP2pwced+pFgabN2DcXbDK7plt2DE/2sL/stRVySEkK7ioeaoPV54Hr26AVZD/wNew6WVDUzGpAm9xKVeE3zwXyq+lIrv66Y6CCGENOfRfZhoJf01qQ/6OvC/FfhZI64ghBBigUcXqI3KwPf6Kek7GUjJ4wtucB0srqnDr9PWLdxlYZwQQoihNg7UxdgWHoZtWzdi1fytSGJrrJ5/FR7iyNoRv3pGNy3y+d9mIezttxH4f7fgojg9YkBphf78b14CdtFVH4SQLqiNA/VkhM7pifTPE3CyuAa9XP+EHRG/hS48A7/601q86WoF1FQgJe0ahix8C68NlDbquf0RS30cYVWTh8+PfIiUfGk9IYR0EW2T3Lb4KMJmbkUKXsGOfwVjJP8aeo0VetlI21uD1VEDK1h1k8qEENJFtPGIWtLtIYM0x+qgIE0I6YraJlCzoNp3oCN+NVA/yUEIIaS12mbqgxBCyCPzeKY+CCGEtBoFakIIkTkK1IQQInMUqAkhROYoUBNCiMxRoCaEEJmjQE0IITJHgZoQQmSOAjUhhMgcBWpCCJE5CtSEECJzFKgJIUTmKFATQojMUaAmhBCZo0BNCCEy1+6/R537fSbiMhJxqPAi/Oyc8WNZOX7j6Ih/3i3ECNUQFN/7FkOrXfDUE3aI/fkiPLXDMWzQEIx9dhTGjX5GqoUQQjqvdg3UqRmpuHHxH7imGA1FWRGcflGIXzq7o+wX/ZBy4yp+iZ7oYdMNP2YUYpy2DN0h4Lwa6DmoHxTaQvhOD4L7mPFSbYQQ0jm169THt+e/hdVTL8LF0QHC9yW4kWeLkpJuKE9Ox/CsIvTJv44nrbpj2vO3YOc7GVd+2R9jVYCd+mfcKOqOj+PPSTWZcK8YOan5YLG967mbj8ysYmjZYknidkQfuiIuy4U6/wJyiniLtMg7tAE7Eot1G1rNwnpk9ZzQoiTrAq7dlYpmaIuykZlfLpXaGH/umDxWw35+9I8jMaZdA/WNX5SjRnsft+6Uo3LoMAhP9ccvqr7B8FE3MeR3z8B9vAY/97gOv9pabC88h8qa+7jT3Qa3NPdwt7Y7rhVdkmoy4W4G4nadQpFU7FJ+OMWCcwbKpGK9bMSFHUSeVGovRf+7F3Hpt6WSgVtJ2Bqd1HaBVFbPidvIPLQXZ36QimaUpcch+n+vSqWWy9sfjrgsqdAc/tyx8Fjt9jh2Me0aqGsKr8OuMh0O93PR3/Zn2PerQVkvTwRVPcAPNd1RqngRfapGwDerByb98ASs+nbDzaF98aBvFUa6atFrsINUkwU05dBUmBhT3tNAo9ZIhaa0Fey+96QCI5Y1Ruqq1bJ6Gu7bPC20/D7G6uN4uw3rk9qqrZXKhqTjaxtV5TBpPkKnu0Ihltg+d9kLpLhsgWbOifdFg+Px9prqS6P9rMDw6WEInOSoKz7QoOyOpoXt43U2qkfUTN8aMtafjVj8PNAf19TzwNxzsSVtNsr4/Wt4W41Vae7xYpqeo7F+5h7ycSRmtesc9caoVbj8ZCmmD/oNasrZE6wgBdXVv4DVdS2ytd3gMcQKhYU1sOvXGwX2A1GjuYy+NvawV1dAIzyBjL6/wnvBQVJtRhQnYE3EFQzxuIkLOYBVdTnK+vvi3bd84NCNba8tRsrfN2EPGzzY9ahCmcYRfqEL4eOi1N13mxoeT2Ugnm3Xqqsw5NXZmJhzGAd+4uVyOExciOWzdAFQk3UQ695PZi80NkBFFew838DSWSPBajKtiI864pADG/SsrgKG+SF8gRdrWwZi512A6qVCJJytgctLixD6ki0bFW3AunNq2PViI6wKFXz+Ega/UbojaNL3YsWOS6hRseNrVfCdZI/YNCe8F+UDHFuFJUW++GhaObZuOIrMu/eh7G3D6g1D8GR78f7GmDwnsW9uY7RDFhIKAC//9ZgxslFfwhkBrC89Boo1IWfPKkSnacW216imwKdPPNKcIrF8qiNydgQhbiBbHpKIJf9gx/uZ1cHOw4PV6zdKbEoTJeycdpSOhEN2Ent3MB5zN84C9PWwOvV9m/uEDay0rG8Hsb4NZX0rPicK4bf7Dbjxek5vxvrT9giMmIXhxh6sh3oelMNq1Cws/7MnVNLz7dS7a3Hwpg3rIy0UblPg/kM8NNNiMHss227y+aA7X/ExDBzDW2Wc0XN+GjnRGxCXywKutS3sBr2M5awfVBr2zmrtLiTcZW1RatnzyR4zIpbBmz9e6bvwWqoT5paeRJxaAStNw7bUPV6Gj10zj6NvxXYEfeGMd9nzUT+8ytsXih2KN/HeTGdpDTGJB+oWuZUuJJxKFs593fSW8HWeUC3tZondu3cIxw9vET5ePU84siBQ2B+5UPjn+gDh2Lr/JxzdECF8tWa+8Onm94SNBw8K/zwQKRzb/Vfh1OpFwmeRiwS/NxYIL66cI9Vkws3jwuqARcK+i5W68v2bwvG3A4WYszVisejTCCFwe6pQeV8sCpUXdwrBC/YIl/lJiPeNEA5/r9tX+P6AsJjXdUkq/3RcWBewTbjA961OFWIC3xGOF0jb2HFOrgwWYpKlslEFwvHwRcLuZLVUrhQubV8kLP20gC2nC/sCAoXV8Td1m5ia5G1CwMqjQpG+yuusfYHS8e8kChvZ8U9e120Saljd7Dz93z4u3GLFW/ErBf+YdN02se6dwmWpZBI/p9dXNjgnse/4OTXuV4b3ZXBMo74MPywUseWyU39jbWdtkbYJBUfZ/evP73KMwbnyuqV2myOe0wJ2HvVNMKinRriwOVDYkqR/3EuE5N3bhGTetWLbdedfdnabsPitWCHXoI4mWvo8aNRn/Hmw+gvdueV+tIj1UTp7pHUqk3cKgawf9l3kJd3zYd/5+ufD5Rj986HxY2iMmXNmeN/ojqNz6/weYeNeg7YkbRL8N59ltTDssfNv8H9TIiSuM/54mVpu8jjezxIOBK8UTkqbBeE7Vo5gfSUViVktn/pwGIMXVUmI9PfHnwxucz+6CqdnhsFa2s0S9+8/wINaAb1stFAOqkKv3k9A+TNbZ9Ufn9hW4Fs7Z5T1sUd/pQLqB0Pw/b1BON9rCDJVQ9DPzg5DywdLNZkzAs+NlYZK3RwxepQtSu7wObV8pJ2vgc9L46Dkox1GOXYG/Owv4OJ3ujLghF+76CYM4OKK0byu0VK5nyMGoBRld9moKi0VKSM8McG2SjdNUKHE6OedkJKZzXYsRMqevYituyXhBr9/fipO13rCx8OWlxh2n79E490/OEllR3iM07+91CIzNRvuXhNgJ75VZbdeYzBxSDb+fZltzclG5qgputEQp3CCt89IqdAc4+0Tz+lpb3g7SefL+s4nKgaBHlLZsF/FvgS8J7qKo0jePgybgAnV2cgp1iI3Ox8ePtK7GM7JBz4mRspNqLORYNi+I9lsfC4ZMQ5uRt+yKGDH3jHkZmbgBp9i6GYPj4D58NB3LVNybjvWfKEyGEmbeJxELXgeNOoz72njceNCBtSsj3KyesJn6pi6d1lKD2/49JYK/PnA3hl4sC4UH1/1ffZOcCwqL2WjRNpFT5MZb9DOvUjI5B/8NX/OhhzGz0XoHN4W3VRJGb9PCfsrbYfjFLxU939jDy/2f3ItnZ9HK3UbCfcxapw+W6grZ11Aku0EuJtoH2moVXPU1u5L8fknARiuL0+ORBx76/m0PuZYSLC2xhO9+kHxqwk4qeqJew5OeDDud3jC9Tn8TtsDPS/n4hfffIO78V+gx8Vv4JB/GU8WfY8+PxWg58934ag18iFGMxRK/b+JGuq7StQVRbZwsAdu3CyVypYpu8P2/+4o1kS9gxXSbf3pUjj04v+wKgwYNQrP1N2c0FO8022UWCvFt8vNu40y9t+a9tmGuvpXRG3CsVJbsNcw3fGf6C7tq6OwsaxmU+0T67S4fbwvS1lArT//FVGxSFOooHhC13arBs1TQGnpK3oPeww2bJ+LvUVtGj5zEWYqkhAdHoJ585cj+jAL8HXz0BmI/fgKyjSGc6gmHqcWMNpnDn0xoJg91mIfdYdC/2IlYs8/fT/w58OdZOyo6z92258Fqz5KWEm76CnsnQ3aOQqD7W3E9ebPuZGic9gTyfeLwFtR2xCf1egqj6cc66YoRE/2x2AWyx9mznm4l6f4YsZfAPPS0jGElVvwKVOX1uoPE/XB+ulWBmlOq63CzZtfo7rkOibksWfUXQHawovo9fMdDFF3x7CBKti/+DQ8p/XDk55KjP0/A2A/tg/uDx2KAU8/hV86j5Bqag0VVL01YP+rBspRwv7XBvQ3PW9rjB0b9WPUTLy3YX3D2yw+qrXF4PHjMLru5syOzO/UFw73NGw8Y4m+sGPPaI//blQ/u/E5XPH4jerS3LV07GO8fcbqNI33pT18Qxu3byG8+unaXlllWFM51JY2z9oRbobtG+Vo2YsHe1fhEbgM722Pwe6oWRh8aTt2JOqD0RgEb1mH0BHfIfqDc9II3cTj1AJG+6zkNm44ssfa6PNNjVJ9k/jzwd4b4Q36j934fLK0i55i4EiDdrJ3FQOlHjF7zobKkbJ/PzQTo7B7Oz/OMgROaTS0LShsOJK/UYhr7DCNXzRaxHk8JuMSMvOzkZbRHxPHtyJodFEPddWHGKw/aF2Q5hTaEpTWWqOy3BYq9r646tp1ZF+/g1s//Btf2alQOngguhdeQk7Bk0iqGob4EgcUa53xQPszKu/+gPzr9W9OW84Z7s9bIeHEhbpRhyb9IOJKPeFh6ayBROHO3rJmnURCofQvWluOzD1sRHPCzDWl/Enb7Rziz7BXBo7f5/1QvHVYemvYgIL9Q45ESkI8buijgPoCYpduQFKRdPzLp3BKf3xtPhJO5uuWjbqJa81cn6avMz5XiizaQiREBGFHSoMwJOF9CSQcq+9LFMYjOmw/cu7p2p6WkFDf9tzjrF5p2ZhbhfX7tgo7/4i17BhSJbZ9oeqhW9RTdFPC7fW/YkbZIWw89miu/W3yOPAPD4+kwuUFTxZsjTzfzrB99YGbPx8eJOJEqj6wanGDvYNa8mGGhS+WzZ/zjwWNzlMhvc3hz73Uuvk+nVusLelS4/h5HMuAywR+Hi3Q5HF0grtnd5zedQinXCbBvaWvhF1Yu16eN8hpBMYMcIddfztU1/aCY7da/LIKuHTlNhx+uIbi9ALkFCmQlpWBXT9n4GzWZeRdu4HUK1eRedsGZbX3pZpaZ8DvFyDgicMIeTMUSxaHIOR/NJgZNgvDG7w9tYD1OPiHuuLf6xYjaHE4qy8csWpvzJ7S+BImQ07wWeIHxbHVmLcwHEtCIhBb/QpCf2980k7hMQfLn76CNSEhWBIWinlhh1E2ZTa8+Lw0P36wK1LWhojHD1p4EKrfmLo6YAxemqZAXGQQgvZdkdYZwesMG4lrm5fqzilkA1JGzId/3Rx1Q7wv/VHfl/PWZmCw/zS4sbf2vO1LR2RgxXzWdrYt6JAKE/lVDsY4esHv+avYGhSE6NPGRoOWcIbXdGekRS9GSBh/PFbj9C9nYPakRiMKPoccNgdDTv8Nsfqg9DCkPqt/HqzFmWFvIHCy7rgDfs/6r1rqo4UhWJHrDJ+6p4ju+VDzaYTu+cD6akWaEwKnj7HsHUQz5+w21RdWX67CayH7kcfePXjM8oXmgK6dSxa/g28xSNyvjrsfnstYhRDxPFbhhMMsBP62BSMyE4+jA3vRcigthbv7sxaeF+Ha9fK8j1nAscV95PXpjW8HahDo8juUVlZj1d0vENLbC862T2HpzTj4lg7GQKchOPdTLrJuXMOL1i7Iu/4T7tRq8MWmv0u1PQQte0vKRn5KcU754fDrTmFtC0VLquIfEHaztWzell83XKGFUtVgcl3CPxjSQmF0W+u16JzM9SXfpmXb+MT6Y9Kqx+MRMHtcfj254fx0Yy15Phhh8TmLz6UqKJRm9mVt0SoeYf/dPYX1kcXw29SKAVEX1q6B+tS/voRQU47yggKo7qhRVVYJzZ1ydLMWYNPDBvdGuaF3jwyc7D1M/Oef9uR4bP7+CwystUG/wlL06TkE8xeESrURQuSLf13+EjITD+B0vzfxLhv9E8u1a6D+9tt0HE4/jHHW9rhsb4W7d9UQ8nKhFhwxaOhAONy6h24PrkNlUwZY9YLVwP/Ez/fvI6b0K/T5WYvlvw3EuKeflWojhMjWvStI2n8OpYM84fNb17pLYoll2jVQcxfyMlF98waSinLR/8FlsHE1yiu64ZM+PRDRzQ322lJYda9FZvF1KLqNRFGpBk52tZjwHy/hmWc8pFoIIaTzavdATQghxLx2veqDEEJI8yhQE0KIzFGgJoQQmaNATQghMkeBmhBCZI4CNSGEyBwFakIIkbnOcR115XUkJ3yJlMJyaQXTczhe+dNUDG/pjwq3k5LT27H1RAE0Q32xOsAepzafgsNr8+HhWI5rqQVQjh4Jh1b+9gMhpGPrBIG6HCf+OgFBh41kEh2zCHHbZ2F44wDXzRq2tjKKerd4Hr98eEfNxmiVLZS4gvi6QM3zJ8ajf1QkvM39GB8hpNPqBFMf96AuNZHuOWMT/CY8h9HPNbo9OxIjXj/WJMWRaWYyQ/NfhGPbGmf+bpEH7NbPGcP7Sb+YZu0K33AepMWtRphpj9lthJCOqBOMqEtwaO4ELL29FHH/+GNdejBzLm55DnM/moW9/47Ei80NrM1lhuYZrI8UgOfDqtEALn9chOBJjrhxeDlWlE7D7rqM0eVIWR+ONI+NCJ7Y6CdIsw42yNg8WMwMXmgwim40ojbTHpOZt8UDEUI6qnYYUZcg/ctEJJ9NaXr7MgV5JgbHzbOGbW9bi24qi2c9CpGw9SiUf4hCzMb1eG/7OszucRTR/yxkA/lziDsA+G2IFlMmbYl8GQ5ZV8V8cAMme8El/QJy9NlO1BlIuToSHs8b+Z3oUTPwXpgPBvfzEdMwBU82lwZM1x676fXtmauU2sNG0mmHDwKvrMMWnsJp4wr4qbKRZyxhDCGkQ2mHQO2AsS+okLiqYRbzPwXuRZ7T003nk9uTuUzh1vawUxYgM60QGv7i0s8LM4I9MYDv1ns8fjMoA2dSddMP6tRzyBk7Ae4Pe25mM1W3LAs1IaTjaJ856p5jsOyzT/D6MKmsnIRVn2yEv1sLUv08DmYzhbtiRtgrUJzZhsXBQQhaugHxYtp+zhYTXhyJtLRLbJxbjkvnS+HtaSo1Vgs0k6m6RVmoCSEdRvt9mKgP1m4yDdJcc5nCnbwQELEeMbu2YJ2/EzK37kSKlF1b4T4e7jmpyCxKxTe3xsJ9lG79Q2kuU7XFWagJIR1J+171wYP15x/IM0hz5jKF58djxao4XBOjuAIqe1XDVPrW4zBxdDbitiahbKKXRR9yNstspurms1ATQjqm9g3UsmcmU7izN/yHZWNNSKguK3hkIlQzZ8PDIAW+m6cnyliMn/j8o5ooNpep2sLM24SQDqcTXJ53D+nrXoLfrutS2UKTInEyxsiXYUwxlRnabFZwA8X8Sy3xyJOKddzfwEd1l/G1gJlM1e2VeZsQ0jY6x1fI799DybUc5N6sllY0o0c/uIwaRl/JJoR0CJ0jUBNCSCdGc9SEECJzFKgJIUTmKFATQojMUaAmhBCZo0BNCCEyR4GaEEJkjgI1IYTIHAVqQgiROfrCC9cJkuMSQjovCtRtnRz3VhK27gf89T9F+hDy9ocjc/R6+D2Kn0wlhHQYHXjq4x4Kr1qenta0Nk6O+0CDsjsa1EjFOvzHnHiWFmOHNpEwt6aC7W/yx7EJIZ1VBx5R65LafvLCCcTN1aeKaY02TI7bKHGth79uNKxh69e9z9b3sgEqqmDn+QaWzhoJ/vt7xhPmdkdS9AbE5bJAbW0Lu0Evi8kCCCFdBA/UHdMt4eCc4cKQocOFaXvypHWtIdXjGyvkSmuac/FdftyVwlfV0gpzbh4XVr99nB1FUp0qxAS+IxwvqNGV798UTq4MFmKSWbn6rBATsElIrtRtEn5KFA5sOSsUScXLMYHCvotSgRDSZXSKqz7So16C396rUknetGmpSBnhiQm2VbqpjwolRj/vhJTMbPMJcwkhXVYnuTxvGEYO7hiZTMrulALfHcUagwS160+XwqEX/5V/cwlzCSFdVScI1MPg/0EsVk12kMryZtfHHhg1s2FyWn6bNVK3g5mEuYSQrqmDB+oOEqRvFeKGdLUGz07ukXUSCYXSCp4wd89yRJ8obj5hLvNjAduPENKldOBA7YBXNx+Vf5B29ILf81exNSgI0ad5LsNx8A91xb/XLUbQYp6ENhyxam/MnuLYbMJct6m+sPpyFV4L2a9bQQjpEugLL48rOa4RJpPQWpowlxDSJVCg5ig5LiFExihQE0KIzHWSy/MIIaTzokBNCCEyR4GaEEJkjgI1IYTIHAVqQgiROQrUhBAicxSoCSFE5ihQE0KIzFGgJoQQmaNvJnIyzkJekrgdsT95I3i6Kxr/JAghpGugQC3zLOQUqAkhHThQ8yzk5XAa9rA/cyolt02UihaynhyNrz+YimaPXpyANduAwCifJvs2+fU8Dc8y3t2CX83TQquugraHLZT0w1CEdHodOFB30CzkDjxw38ZohywkFABebN0Ml2zErd2FhLs2sFNqUVZhjxkRy+A9kLXu2CosKfLFR4FjxOUdFeMxJDMRmbVsv7v34TJtGcJfdpQOSAjpjDr8h4mPLrGtNWx721p0U1k6ih01A++F+WBwPx+Eb2BBepS0vjgd6v+IRMxGFqRHs2CcnYprw+Zgy3aelisaW2YqEXv4HBs3N5V3vhDPrdDttzt8Aso+O44caRshpHPqFFd9dKQs5Doj8NzY+ukNh/FzETpnDJTilEY5G1GzEF3C/krbGxgxDm76u7q4wu1BIW5Qdi5COrVOcnlex8lCblTROeyJDMG8+RF4K2ob4rMo8zghpF4nCNQdKwt5U+VI2b8fmolR2C1OfSxD4BQnaRshhHT4QN3xspCbpOiu+8uzkqd+p1smhBCmAwfqDpqFvAlbeMzyheaALiv5ksXv4FsMkrYRQgh94YVpvyzkDYiZx6ugUBrJSk4I6dIoUHOUhZwQImMUqAkhROY6yeV5hBDSeVGgJoQQmaNATQghMkeBmhBCZI4CNSGEyBwFakIIkTkK1IQQInMUqAkhROY6xxdeZJycVnZqi5GybSfirmvw61cj4av9sC4nI4qykXNvEEY7d+CfjCWkE+oEgbqNk9N2EHn7w5E52iCLjAnq42uxONcT7/6/cbDrpUTZ1/XJc8sM0n4RQuSjE0x93IO61EiQ5jI2wW/Ccxj9XKPbsyMx4vVjKJF2s8g9DTRqjVRoRMu3lUNr4qdMeRLbBtt4EluexaU1+H2NtKOGHUNjQZU1tVoMdnaFg0oJRTfAYdJ8hJrLcC6dt7ZWKhsyt40Q8sh0ghF1GyanRQZi511A/1kanPj0JtCjCmVwRkDoQngM5Ns1bCS7AevO3BaT19ZUVGHgVCnZrJh9vFES25HFOPXuWhy8aQO7Hloo3KbA/Yd4aKbFYPZYfqx49I+KhLeUq1ZMZos/Y/lUtoJPWfx9E/ZcBbuvYTtKkRS9AXG5LFBb28Ju0MtYHuoFla6KBvioe0cKT/HF9lP2x9SwhXA73zB5bv2IWjq3c2o28gbKKlTw+UsYG7HzPGC6bRtTNOjJipUaJXzmL4OvG/3sHyFtggfqx+uWcPHEv4RzXyc3vZ1IFnKrpd0sdks4OGe4MMQ3VsiV1jTn4rts/6Erha+aPVa6sC8gUAiOSRcqpTWVSduEgLePs6MyP6UK+9bFCpfrNp4VtgRsEy7wem8eF1YHLBL2XdRvFITcjxY1rCt5pxDI6t93kZf4sVYKJ2+Km0S34lcKq+N1K4o+jWD3TRUq74tFofLiTiE4/LBQpCsKl2P09ZhnWCfHy/6sTY2Xa5LZea48KhTViEVBuM7OJ1A6t6uHhcVLDgg/Sm0Rvj0sxBzMkgqEkEetHaY+HDD2BRUSV/njT/4Gt8C9yHN6+tH9vvMjYw+fqTzxrI5yoje8yzOQeYsV+o3D7PBZumSzfEriNruhFGV3xV0ZwyS2+cjJ6tmwLg9v+PSWCmblI+084D3RFeBTHOpyYNgETKjORk6bJLbVIjM1G+5eE2AnTrWwW68xmDgkG/++zDb3tccAdT4yr5TqpnRG+SFw+kjxnoSQR6995qh7jsGyzz7B68OksnISVn2yEf5ucrzawAmDxWkOPRXsbaug4dPEtcVI27EKQUEhWBL5DrYeyTaeOVykhvpud3FeuJ4SSotemPh9S5Gw5x2siNLfYpGmUEHRJo/gbZSVAGmfbTA43iYcK7WFks9uqLwQON8VeR+vFc89JGoXUgpbOedOCGlW+32YqA/WbnIO0lwhrhVJi1xtKUrLbaBkw2JNYiy2VkzCxm1b8N6G9Qhf4A03abemVFD11ugCfB01q0taNEJbtzO/rz18Q3nyW8PbQnj1k3Z5pPrCzgHw+O/Gx6u/qkQ52g+hUdHYvXMjlntosW/DId0GQsgj175XffBg/fkHMg7SHBvJHrsAjXRlQ8nxeJyyHYPRdQGyOyCNkjWpF3BBt2iEM9yft0LCifq6NGdO4VRd4O6P/o7FSPmmUFe8lYT4FH0U5/dFg3agMB7RYfuRY3DBy48Fj2oeRIHR40ciJSG+Pimv+gJil25AEnvR0pzbhSXvn9O1pZsCDvY2un0IIW2ifQN1h+AIr4FZWMMTzy4MwVuJ9ghc6AOeUlc5aTZmVx9CyHyelDYUa7IAF92djBrw+/nwrz6MkDdDxbpW5DrDR7rCgx9nwnQvaL9ci9fmBWHeLg08JtVtZPddAH9I910cgnlrMzDYfxrcpKkTt6m+sPpyFV4L2Y883aqHovCYg+VPX8GakBAsCQvFvLDDKJsym/UFO+/f+MKvOh4hIfy8WVt2FcL79WnSPQkhj1onuDyvLZPTGlwy11cDzT0FlL2aXoImXif9BJ8OsfDyNH79sTg/XYxTEatwU7w8T9rWHH7Ntol26OXsCML6NKlQxxGzDS79s5iYdFcLpUr/EagBC9pCCHl4neMr5G2WnLbptc2PVisCNSGky+kcgbrNFCJlTwZUf/CFm7FvkDy0cuQcOQz12LnwGCKtIoSQRihQE0KIzNGHiYQQInMUqAkhROYoUBNCiMxRoCaEEJmjQE0IITJHgZoQQmSOAjUhhMgcBWpCCJE5+sILR1nMCSEyRoG6rbOY30rC1v2Av4k8hi1haaZxQkjn0oGnPu6h8GqL8oib0MZZzB9oUHZHgxqpWIf/Kh1PcWXs0CaymluaaZwQ0rl04BG1Lvv4Jy+cQNxcfU6v1mjDLOZZB7HkH8ko+xlilnIPf91oWMPWr3ufre9lA1RUwc7zDSydNVLMpVhyejPWHykAz3lVowFc/rgIwZO6N8k0/h4boRNCuggeqDsmKfv40OHCtD150rrWaMss5gzPRq7PWs5Vpwoxge8Ixwuk9N73bwonVwYLMcmsXH1WiAnYJCTr05T/lCgc2HK2xZnGCSGdS6e46iM96iX47b0qleRNm5aKlBGemMAT5PKpjwolRj/vhJTMbMDaHnbKAmSmFeqmRPp5YUawJwbo7koI6aI6yeV5wzBysJzzLtYru1MKfHcUa+qye7+D9adL4SBmSXHFjLBXoDizDYuDgxC0dAPiM81kvyWEdAmdIFAPg/8HsVg1mWcxlD+7PvbAqJlNsnu/N2ukbgcnLwRErEfMri1Y5++EzK07kaLWbSKEdE0dPFB3kCB9q7Aum7fCfTw8sk4ioVBaUVuOzD3LEX2iGMiPx4pVcbgmblJAZa+ClbhTvUeXaZwQ0lF04EDtgFc3H5V/kHb0gt/zV7E1KAjRp8sB63HwD3XFv9ctRtDicIS8GY5YtTdmT3EEnL3hPywba0JCdZm/IxOhmjkbHtIF2IaZxgkhXQd94aVNs5ibx7OXw9oWisZJvM1l/iaEdDkUqLk2y2JOCCEPjwI1IYTIXCe5PI8QQjovCtSEECJzFKgJIUTmKFATQojMUaAmhBCZo0BNCCEyR4GaEEJkjgI1IYTIHAVqQgiROfpmIifjLOQlidsR+5M3gqe7ovFPghBCugYK1DLPQk6BmhDSgQM1z0JeDqdhD/szp1Jy20SpaCHrydH4+oOpaPboxQlYsw0IjPJpsm+TX8/T8Czj3S341TwttOoqaHvYQkk/DEVIp9eBA3UHzULuwAP3bYx2yEJCAeDF1s1wyUbc2l1IuGsDO6UWZRX2mBGxDN4DWeuOrcKSIl98FDhGXN5RMR5DMhORWcv2u3sfLtOWIfxlR+mAhJDOqMN/mPjoEttaw7a3rUU3laWj2FEz8F6YDwb380H4BhakR0nri9Oh/o9IxGxkQXo0C8bZqbg2bA62bOdpuaKxZaYSsYfPsXFzU3nnC/HcCt1+u8MnoOyz48iRthFCOqdOcdVHR8pCrjMCz42tn95wGD8XoXPGQClOaZSzETUL0SXsr7S9gRHj4Ka/q4sr3B4U4gZl5yKkU+skl+d1nCzkRhWdw57IEMybH4G3orYhPosyjxNC6nWCQN2xspA3VY6U/fuhmRiF3eLUxzIETnGSthFCSIcP1B0vC7lJiu66vzwreep3umVCCGE6cKDuoFnIm7CFxyxfaA7ospIvWfwOvsUgaRshhNAXXpj2y0LegJh5vAoKpZGs5ISQLo0CNUdZyAkhMkaBmhBCZK6TXJ5HCCGdFwVqQgiROQrUhBAicxSoCSFE5ihQE0KIzFGgJoQQmaNATQghMkeBmhBCZK5TfOGlPOtTHDp1HS35cVCHZ/6I6ZOfQrNfLqwtRsq2nYi7rsGvX41EgGdzabLkTouSrEvQPDUOg3tLq8woOb0dW08UQDPUF6un3ce+j0rhvdCvma/Oa5F3aDNOPTkbgZOMZJ8pTsIOi+ohhHCdYERdgoRNy7B2z6f4V2IiDu15H9uk5RP72fLW/ThhZDny9T1INpLPtjH1l3uxB5MQHhEJ/988XJDO2x+OuCyp0G5uI/PQXpz5QSqacysBOw4BPmFvY3WAJ1p99jzBb3QS1FKRENIynWfqY9hfsPnobiwer18+ivV/5Bt8sewTY8uWqanVYrCzKxxUSii6SSs5nohWrZEKjWg1bFs5tI1+2rSmgievlQotwJPg1telywKjaWlFvL08c4xRJup8wG79nDG8n5RE19ELgeFGRsFN6lZg+PSw+tH0Aw3K7mhQoyuZqEdqg6kXT/6jVUb6VHSP97cG2lqpbIpUh9FjWFoHIe2gE0x96JPTRuLk0f/ExbrlWShf5wK/XboktKpNxpfNJaflI+AdKTwlli3slP0xNWwhvPoWI+Xvm7DnKmDXo4ptc0ZA6EJ4DOT30LD7bMC6M7fFZLY1FVUYOJUnn+2OpOgNiMtlQcKa1TXoZSwP9ULRjiDEDYzE8qlSQDPMWC4uN0qCa89HpnHIgQ16VlexFyQ/hC/wgoPhC0hjtcU49e5aHLxpw9qrhcJtCtx/iIdmWgxmj2Xbi0zUmdMwMe/gl8IQ/PQFrIkohN/uN+Amtk8Nj6GXceK7Gt2L0EBfvPsWaztrT47+3IYkGk/wq6+HNUGTdRDr3mf79LIBWD1Wo2Zh+Z89oZLqSXH0hfrrRNwA6++fbeAVGMnazsf3uv7emKJBT1as1CjhM38ZfN2a/vxgw2NUwc7zDSydNZK9SzBRx7BL2BF8HMNXR8JbP3uTsx9BH3ZHzIYZ0gpCHhMeqDu2W8LBOcOFIb6xQm6DZUG4+C5bHrpS+Kra9HJzbsWvFFbH35RKglD0aYQQHJMqVN7XlSsv7hSCww8LRbzwU6qwb12scLlS3MQ2nhW2BGwTLkjHuRwTKOy7qFvmeNmwbuHmcWH128fZWUjLAYvY/vrKCoTj4ax8Xi2VK9n9FwlLPy2QysblfrSItTed7a1TmbxTCAzQt6OZOg3bw4lt2ilcrls2aF/NVeFwOKv3kq7Y4NzM1VOdKsS8vlI4XlAjbhLu3xROrgwWVn+huy+vJ2Alu6++v5O2ieUyXrh6WFi85IDwo7RN+PawEHMwq+5c6xg5xvG3A4WYZFY2U0fu3kUNHh/el831NyFtga76aJF8pJ0HvCe6iiM//jYawyZgQnU2cniC2X7jMDt8li75LJ8OuM1uKEXZXfHOrWCQBDc/FacxHh7s0OLbd/V9DPEYi8pL2ew9hSn5yMnqCZ+pPHGujtLDGz76DxFbVachg/YpnPHrIcCPBS3LtKtNS0XK097wdpJGwd0c4T1tPG5cyKib0x48dkzduwblr5wxuLAQRbzQ1x4D1PnIvFKqmxIZ5YfA6XyU3JCxY/hExSDQg5XN1DHcfSxunDvHRvJMbTbS0lSY/AKlSSOPHwXqFlFDfbcUCXvewYoo/S0WaQoVFLwna4uRtmMVgoJCsCTyHWw9km08k3hrlN1GyZ1k7Kg7Lrvtz4JVHyWspF2a4u3t3nBunYUgcb6Za1Wdj1bZnVLAWokGkxUOfTGgmLVNKpqk8kLgfFfkfbxW7POQqF1IKWw6iW30GHrm6hg1AT61l5BZyJZzLiGpnyfGGbmIhZC2RoG6RVRQ9baHbyhPQmt4WwivfmxUmhiLrRWTsHHbFnF9+AJvcQ7WYhqN6Ssj7PrCwd4b4Q2Oy26hXqxVpvD2ani1BtQo1V/H2Ko6Hy27PvbiB3kNwmvJbdxwZG2TiuYoR/shNCoau3duxHIPLfZtOIQ8aZue0WMYMF2HM9yfB86cz0deWjpcPMY/tn4hxBAF6hbR/eMmHLsAjf7qgMJ4RIftR07dlQTd2Vtr3ZIm9QIu6BbrGE4NOAx0xLX0VF3iW20pUr44Z3oE7jwekx8k4kSqPspqceOzDVjyYYbJAKRrrxUSTtS3V3PmFE7pA3er6mwlEwl+Fe7j4XGZtUk/iuUffh5JhcsLns0GRc25XVjy/jnduXVTsBcdG90GRpOVgLjT+eJ56I8RnyuduLYQCRFB2JGiNVsHN+CFCVCe/xB7Uodh4gu20lpCHi8K1C004PcL4I/DCHkzFEsWh2De2gwM9p8GN2s2Mps0G7OrDyFkPk9SG4o1WYCLdD/ObaovrL5chddC9osjNofJfvDRnsJbQUFs3S5UPj8Jg3W7GuEEnyV+qPk0AvMWsvrnh2BFmhMCp48x/pZeMuD38+FfLbV3IbtPrjN86t6+t67OFjOX4Nd6HPzDRuLf63TJfUPeXIszw95A4OTmg6LyN77wq45HSAjvb/ZY7CqE9+vTMJxtK8k8ifiTGbrpE+kY1zYv1R0jZANSRrB+8VCYrUPk6ImJqlKUuI3HaDNXCBHSljrF5XlHX5+AhaelosWavzzPLH6t9D32j96raUgTr3t+wgZK5SMNd/X4B5XdpGubmZJjq7DkaNMP8bz/Il2Cx/HrhA3npxtrVGd74P0G61Yk9zXzWDRm8hgtqIOQx61z5Ey8exXpl2/BwtS0oh79n8bYYfRWlhAif50jUBNCSCdGc9SEECJzFKgJIUTmKFATQojMUaAmhBCZo0BNCCEyR4GaEEJkjgI1IYTIHAVqQgiROQrUhBAic53im4ltmoW8nanzL6DI+lm4DXy8v0GhLcpGLlwf/rj3riB+8yk4vDYfHq36LedmMpoT0gV0W8lIyx1UCT57678R+fmPqCwvQerp40jM0C1fOXscJ5Ku4E7VHeQ1Wj7ySS3GBk7C0O5SNTKVfygKcZXj4eXaS1rzeNxJikHsTbeHP27tbVz5Jh89nx0Pp1ZV9QB3Ln+D/F6j4T7k8fYBIXLReaY+2igLuchEVnER/9U5U9nIGf5rbYZZr8WyqQziTbJ5m9DMMZtqZeZyUTPZwXlbDOttfA7WrvANbzSaNtefTbKBN8poriftZ5KlfdnkePXEx8qSOghpY5SFvJm5j5LTm7H+SAGgVKCGxQWXPy5CMA8ammzErd2FhLs2sFNqUVZhjxkRy+DNs5HrM3Q/lYH4q+wfXl2FIa/OxsScwzjwEy+Xw2HiQiyf5QqFuK8G3iMuYV9aDXryTNu92Dksm4HhSoNs3jxTea25DOgmSFnGc5+wgZW2Chjkh/BQLzggA7Hz4tE/qj7LNv+51B34s3gscVkzHi5pJ5HGgiXPqF6fuZvf9wLsXylFQqJaDIpW42dhdrdT2JPJOomVMWwGli9pehyT/Wkmo7jZPtA4wo/1gY8Lu1MzmdEbkjLGn1PDjg3UyypU8PlLGPxGsXqkx/ZUlY34eFTaeGHxcj8MxwXKTE7aBw/UHVsbZiGvPivEBGwSkvVprX9KFA5sOStmHL91fo+wca9Bdu+kTYL/5rOCmOdazLIdIRz+Xsp6/f0BYTHP2H1JKv90XFinz04u7muYjbxG+PHjCCEwJl2syzCbt9kM6EbVCBc2BwpbkqRW3i8RkndvE5LFRNrpwr6AlcLJ+iTbDTKu82X/198RTl4Xi+xg3wkHeMbyi7xV/L6BQkySlL1czLYeKKz7okRXrskS9gX/VTh6lRcMjmOmP81lA2/cB4HbG/XBgj3C5bq+NJ0Z3VBN8jYh4O2jQpH0kAjXjwqrX9c9JkUH/yos/lhsPFMj5H66TTj6ra5OykxO2gNd9WGOtT0bLRcgM61Q99a/nxdmBHtiAFt0GD8XoXN4dm/d1EAZf4tcwv6Kd+Sc8GsX6YM4F1eM5hm7R0vlfo6sDsPs5GPg+7J+iKbA4Je94ZKegVxpjU4zGdCNUsCutw1yMzNwg7evmz08AubDw8JE2oqxPrp3CJzSFT4T7ZGUcUVa4Yjhv5J+z1s5DL92dMRzY+11ZYU9+tuWQ904r5iZ/rQsozjvgxr4vDQOSn1W8rEz4Gd/ARe/05Uty4yuRWZqNtx9fDBA/1kpG3kv3zUf7uwdlp29Pcq+z0DeLT61osDwP8yHLx9pM5SZnLQHCtRmuWJG2CtQnNmGxcFBCFq6AfGZ0rUlReewJzIE8+ZH4K2obYjPask1J404OmGA4VtzlS0cHlQ1mhduJgO6CcNnLsJMRRKiw3lblyP6cHZ9vsdmDB7YX1rSUTnasyD6MHO2ZvrToozivA+UUDaI3qyv2OvDjZulUtkSt1FWAvS00UfphpST/ozQp/Oxb+1SvBYUivU7kurzPVJmctIOKFA3x8kLARHrEbNrC9b5OyFz606ksNFsyv790EyMwu7tPHP3MgROeYhRVXEhbhgGz7u3cYOn8mowf24+A7pJCid4BC7De9tjsDtqFgZf2o4dicZfVLQN05Ujr5BHo3olRTehaHGerEaM9qduU/MZxY1lVS9HCYvRA/pLo3mL9IWdA1BZZeJFp5st3P4QhtWbt+CjzX/FxHtHseaA/p0EZSYnjx8FanPy47FiVRyuif/PCqjsVbASN0gU0rV9teXsrXTde+9WyED8MX1Q1CDn0HH8OHZMg8S4lmVAbywfCRFrEZ8rBSTbvlD10C0C/dHfsRgp30jHvZWE+JRGAfzSSZwqkpY1GThxWg3v8SOlFa1gpj+bywauYySrevpBxJV6wqPZZmmQd+wgUvL5wRUYzc4jLYG1RR/0WV+ueWM70u5pkPZ+OPackfrCmvVZo6sCKTM5edwoUJvj7A3/YdlYExKKJWGhmBeZCNXM2fBQ2cJjli80B3SZs5csfgffYpB0p9YYA3erWISImcAXI5oFnqX+TTOBm8uAbpwzvKY7Iy16MULCeIbv1Tj9yxmYPYkHF0dMmO4F7Zdr8dq8IMzbpYFHo0vghr88BfgwVLxv0MIPceOFN+E7StrYGib7k42mm8sGLuF9EPBEfR+E/I8GM8NmYXiTqzoaK0bmafZilK6br1Z4zMHyZwuxjtUhtoX1pVvwHLhbK+H+ysvQHI3QPbbs8dhaMB6L/+Aq3k9EmcnJY0ZZyC35R6vVitfTKlUNP9rSra+CQtmKzNl6/JKyiEL47X4Dbvya3lpF89nLG2fMFuuIbzRNwLi/gY8Cx4iLJrNvW0BbwYadNkoomg2GFjLVn5yl2cAfVdZwc23h12K3c2Z2QjjKQt7eDAO1tIoQQgxRFvL2ps5GwuHbGB3gpbtMjRBCGqFATQghMkcfJhJCiMxRoCaEEJmjQE0IITJHgZoQQmSOAjUhhMgcBWpCCJE14P8D53rdcPbalDoAAAAASUVORK5CYII="}}},{"cell_type":"code","source":"# Let's do some basic path configuration for ease of file accesibility\n## We will create few path variables as we have many files in different folders and format\n## It's a good practice as it prevent any code error/ mistake due to lengthy file urls and make code more readable\n\nPath = \"/kaggle/input/home-credit-credit-risk-model-stability/\" #Base Folder\n\n# Training files\nPath_train = Path +\"csv_files/train/\"\nPath_train_par = Path +\"parquet_files/train/\"\n\n#Testing files\nPath_test  = Path +\"csv_files/test/\"\nPath_test_par = Path +\"parquet_files/test/\"\n\n\n# Let's do a basic file count and availability check, will modify the earlier file function and use it\n## Below function is not a necessity but a good practice, we should use functions if we are asking the same question multiple times. \n\n\n# ------------------------------------- Helper function to see files and get filecount --------------------------------------\n\ndef file_count(file_path):\n    file_count=0 #Initiate filecount with 0\n    f_list = []\n    for dirname, _, filenames in os.walk(file_path):\n        print(\"\\n Files present in dir: \", dirname)\n        for filename in filenames:\n            file_count += 1                                  #Increment filecount for each file found\n            f_list.append(filename)\n            print(\"\\n\", file_count , \" :\", filename)\n        print(\"\\n Total Filecount: \", file_count)\n    return f_list\n\n\n# ---------------------------------------- Basic Sanity check on file availability --------------------------------------------\n\n# Path & Function definition will pay-off now in simplicity of execution\n\n# Check for training files, par and csv types\n\ntrain_f_list =file_count(Path_train) #Saving filenames to a list, will be usefull later\nfile_count(Path_train_par)\n\n# Check for test files, par and csv types\n\ntest_f_list = file_count(Path_test) #Saving filenames to a list, will be usefull later\nfile_count(Path_test_par)\n\n# ---------- Notes and Observations-------------------------\n\n## We can see filecount is same and by description files seems same as well (csv and Par) for both train and test\n## We have same data in csv and parquet, currently we move ahead with csv in this notebook\n## Ideally, we should take a decision based on format's performance (read, write time) and ease of execution. Maybe, we will revisit this point later.\n\n# However, Train and Test have different filecount, this could be problematic while creating train and test datasets (Both should have same cols)\n# Our observations will help us while defining a good strategy on Train, Test dataset creation","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:15.586739Z","iopub.execute_input":"2024-02-15T04:15:15.588309Z","iopub.status.idle":"2024-02-15T04:15:15.614627Z","shell.execute_reply.started":"2024-02-15T04:15:15.588254Z","shell.execute_reply":"2024-02-15T04:15:15.613199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# File/ column Definition and depth rules\nRead comptetion provided details: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data\n\n## Summary\n\n### We have two base files..\n\n* **train_base.csv**\n* **test_base.csv**\n\n*(Note, the hidden test_base.csv contains approximately 90% of the numbers of case_id values of train_base.csv)*\n\n### ..and three depth levels of supplementary information.\n\n**Depth values:**\n\n* depth=0 - These are static features directly tied to a specific case_id.\n* depth=1 - Each case_id has an associated historical record, indexed by num_group1.\n* depth=2 - Each case_id has an associated historical record, indexed by both num_group1 and num_group2.\n\nYou can read more about Credit bureau (CB) here https://en.wikipedia.org/wiki/Credit_bureau.\n\n\n### Some more useful information which can come handy..\n\n## Columns\n\n**Special columns:**\n\n* case_id - This is the unique identifier for each credit case. You'll need this ID to join relevant tables to the base table.\n* date_decision - This refers to the date when a decision was made regarding the approval of the loan.\n* WEEK_NUM - This is the week number used for aggregation. In the test sample, WEEK_NUM continues sequentially from the last training value of WEEK_NUM.\n* MONTH - This column represents the month and is intended for aggregation purposes.\n* target - This is the target value, determined after a certain period based on whether or not the client defaulted on the specific credit case (loan).\n* num_group1 - This is an indexing column used for the historical records of case_id in both depth=1 and depth=2 tables.\n* num_group2 - This is the second indexing column for depth=2 tables' historical records of case_id. The order of num_group1 and num_group2 is important and will be clarified in feature definitions.\n\nAll other raw columns in the tables serve as predictors. Their definitions can be found in the file feature_definitions.csv. For depth=0 tables, predictors can be directly used as features. However, for tables with depth>0, you may need to employ aggregation functions that will condense the historical records associated with each case_id into a single feature. In case num_group1 or num_group2 stands for person index (this is clear with predictor definitions) the zero index has special meaning. When num_groupN=0 it is the applicant (the person who applied for a loan).\n\n**Various predictors were transformed, therefore we have the following notation for similar groups of transformations**\n\n* P - Transform DPD (Days past due)\n* M - Masking categories\n* A - Transform amount\n* D - Transform date\n* T - Unspecified Transform\n* L - Unspecified Transform\n\n*Please note that transformations within a group are denoted by a capital letter at the end of the predictor name (e.g., maxdbddpdtollast6m_4187119P). We hope that this will simplify the manipulation with predictors.*\n\n## Strategy for Train dataset\n\n**We have information at various levels, a basic strategy will be to start with base file and then go through each depth level to enrich the train dataset with supplementary information. Also, to help ourself work on transformation functions for common operations like col treatment, datetime etc.**","metadata":{}},{"cell_type":"code","source":"# Let's start by reading the base files (Train/ Test)\n\ntrain= pd.read_csv(Path_train + \"train_base.csv\", index_col=False)\nprint(\"First 10 rows of train df:\\n\", train.head(10))\n\ntest= pd.read_csv(Path_test + \"test_base.csv\", index_col=False)\nprint(\"\\n First 10 rows of test df:\\n\",test.head(10))\n# Test file doesn't have Y col (Target), which is normal\n\n\n## ------------------------------ Basic EDA -----------------------------------------------------\n\n# Check shape\n\nprint(\"\\n Shape of train df: \\n\", train.shape)\nprint(\"\\n Shape of test df: \\n\", test.shape)\n\n## That's good news! We have same columns (except \"Target\") in train and test base files. We have good number of train records.\n\n# Check NULL value\n\nprint(\"\\n Null check on train df: \\n\", train.isnull().sum())\n\nprint(\" \\n Null check on test df: \\n\", train.isnull().sum())\n\n#No NULLs in train, test base files\n\n# Other info\n\nprint(\"\\n Basic info on train df: \\n\", train.info())\n\nprint(\" \\n Basic info on test df: \\n\", train.info())\n\n# integer and object datatype, need to correct for the date columns.\n# We will need to check and treat datatypes in other files too. Good approach would be to define a function to do so...","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:15.617103Z","iopub.execute_input":"2024-02-15T04:15:15.617640Z","iopub.status.idle":"2024-02-15T04:15:18.071558Z","shell.execute_reply.started":"2024-02-15T04:15:15.617598Z","shell.execute_reply":"2024-02-15T04:15:18.070489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# Helper function for dtype treatment\n# Inspired by https://www.kaggle.com/code/greysky/home-credit-baseline\n# Always a good practice to read other notebooks and borrow good practices.\n\ndef tr_dtypes(df):\n   \n    for col in df.columns:\n        if col == \"date_decision\":\n            df[col] = pd.to_datetime(df[col])\n        elif col in (\"case_id\", \"WEEK_NUM\", \"num_group1\", \"num_group2\"):   \n            df[col] = df[col].astype(\"int64\")\n        elif col[-1] in (\"A\", \"P\"):\n            df[col] = df[col].astype(\"float64\")\n        elif col[-1] in (\"D\", ):\n            df[col] = pd.to_datetime(df[col])\n        elif col[-1] in (\"M\", \"L\", \"T\"):\n            df[col] = df[col].astype(\"category\")\n        \n    return df\n\n# -------------- Transform the Train, Test datasets and see the results ------------------------\n\ndf_train = tr_dtypes(train)\ndf_test = tr_dtypes(test)\n\ndf_train.info()","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:18.075115Z","iopub.execute_input":"2024-02-15T04:15:18.076057Z","iopub.status.idle":"2024-02-15T04:15:18.507889Z","shell.execute_reply.started":"2024-02-15T04:15:18.075943Z","shell.execute_reply":"2024-02-15T04:15:18.506519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Time to load some other libraries for visualization and transformation\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\n\n#to ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\n## --------------------- A simple lineplot showing target(mean) over time (WEEK_NUM) ----------------------------------------------\n\ndf_viz = df_train.groupby('WEEK_NUM')\ndf_viz  = df_viz['target'].mean()\nsns.lineplot(data=df_viz)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:18.509628Z","iopub.execute_input":"2024-02-15T04:15:18.510987Z","iopub.status.idle":"2024-02-15T04:15:20.900454Z","shell.execute_reply.started":"2024-02-15T04:15:18.510928Z","shell.execute_reply":"2024-02-15T04:15:20.899350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Distribution of target variable\n\nsns.displot(df_train['target'])\n\n# Have a % visual to understand distribution of target variable\n\n# Create a pie chart of the distribution of target column values\n#sns.set_style('whitegrid')\n# p = plt.pie(data=df_train, x='target', autopct='%.1f%%')\n# p.set_title('Distribution of Target Column Values')\n# p.set_ylabel('Percentage')\n# plt.show()\n\n# Note for later steps (modelling):  It seems to be very imbalanced dataset, for good modelling we should treat this later","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:20.902184Z","iopub.execute_input":"2024-02-15T04:15:20.903438Z","iopub.status.idle":"2024-02-15T04:15:23.183994Z","shell.execute_reply.started":"2024-02-15T04:15:20.903391Z","shell.execute_reply":"2024-02-15T04:15:23.182666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preperation \n\nOk, so far...\n\n- We took a sneak-peak into available files, and developed some understanding about folder structure, files, depth and column logic\n- Took some decisions (going ahead with csv), developed some helper functions (filecrawler, dtypes transformation)\n\nnow with our base files (train_base.csv, test_base.csv) are treated with dtypes transformation and available in respective data frames (df_train, df_test)\n\nNext step is to add features to them from other files, we will have to treat files differently at different depth level. \n\nA simple flow would look like:\n\n- Identify the depth and merge logic\n- Read the file - apply the transformations - save to datafram\n- Merge dataframe with base (Train, Test)\n\n**Let's start with Depth 0 files (information provided in Data tab: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/data)\n**\n\n\n**Train Files:**\n\n* train_static_0_0.csv\n* train_static_0_1.csv\n* train_static_cb_0.csv\n\n**Test Files:**\n\n* test_static_0_0.csv\n* test_static_0_1.csv\n* test_static_0_2.csv\n* test_static_cb_0.csv","metadata":{}},{"cell_type":"code","source":"## Before we start merging, let's do a quick check on each file\n\n# Train files (depth 0): \"train_static_0_0.csv\", \"train_static_0_1.csv\", \"train_static_cb_0.csv\"\n# Test files (depth 0): \"test_static_0_0.csv\", \"test_static_0_1.csv\", \"test_static_0_2.csv\",\"test_static_cb_0.csv\"\n\n# Check train file\ntemp_df1 = pd.read_csv(Path_train + \"train_static_0_0.csv\", index_col=False)\nprint(temp_df1.shape)\nprint(temp_df1.head(10))\n\n\n# Check test file\ntemp_df2 = pd.read_csv(Path_test + \"test_static_0_0.csv\", index_col=False)\nprint(temp_df2.shape)\nprint(temp_df2.head(10))\n\n\n# Good news, both files have same column counts. We should also validate at col level before we can merge them to respective base data (Train, Test)\n# Also we might need to repeat this kind of comparison multiple times - better to have a helper function to compare two dfs at col level\n\n\n#### A helper function to compare train and test files to see if there is any column mismatch----------------------------\n\ndef col_compare(df1, df2):\n    \n    #Save columnnames from input dfs as lists \n    lst1 = list(df1)\n    lst2 = list(df2)\n   \n    #Initiate empty lists to save missing cols if any\n    col_mis_test=[]\n    col_mis_train =[]\n    \n    for colname in lst1:\n        if colname not in lst2:\n            col_mis_test.append(colname)\n        \n    print(\"Columns available in Train but not in Test: \\n\", col_mis_test)\n    \n    for colname in lst2:\n        if colname not in lst1:\n            col_mis_train.append(colname)\n        \n    print(\"Columns available in Test but not in Train: \\n\", col_mis_train)\n    print(\"No output means all columns are present in both train and test datasets\")\n    \n    \n#------------------------------------------------------------------------------------------------------------------------\n\n# Compare cols for *static_0_0 files (Train and test)\n\ncol_compare(temp_df1, temp_df2)\n\n# Ok, so we validated that we have same data at col level. Good to merge with base, but before that let's perform some basic eda on these files as well","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:23.185514Z","iopub.execute_input":"2024-02-15T04:15:23.185948Z","iopub.status.idle":"2024-02-15T04:15:57.837648Z","shell.execute_reply.started":"2024-02-15T04:15:23.185913Z","shell.execute_reply":"2024-02-15T04:15:57.836515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ### Let's package the complete process (reading csv, shaper comparison and col comparison) in a function\n# ### A helper function to read csv file and do a basic check for train and test\n\n# def file_compare (file_suffix)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:57.839387Z","iopub.execute_input":"2024-02-15T04:15:57.839799Z","iopub.status.idle":"2024-02-15T04:15:57.846775Z","shell.execute_reply.started":"2024-02-15T04:15:57.839763Z","shell.execute_reply":"2024-02-15T04:15:57.844668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's analyze files at depth 0\n\n## ------------------------------ Basic checks --------------------------------------------------------------------\n\n# Check shape\n\nprint(\"\\n Shape of train df: \\n\", temp_df1.shape)\nprint(\"\\n Shape of test df: \\n\", temp_df2.shape)\n\n## That's good news! We have same columns in train and test base files.\n\n# Check NULL value\n\nprint(\"\\n Null check on train df: \\n\", temp_df1.isnull().sum())\n\nprint(\" \\n Null check on test df: \\n\", temp_df2.isnull().sum())\n\n##### We will have to treat null values, but before that let's check _static_0_1.csv files as well\n\n\n","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:15:57.849236Z","iopub.execute_input":"2024-02-15T04:15:57.849682Z","iopub.status.idle":"2024-02-15T04:16:01.157449Z","shell.execute_reply.started":"2024-02-15T04:15:57.849645Z","shell.execute_reply":"2024-02-15T04:16:01.156019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check train file\ntemp_df1 = pd.read_csv(Path_train + \"train_static_0_1.csv\", index_col=False)\nprint(temp_df1.shape)\nprint(temp_df1.head(10))\n\n\n# Check test file\ntemp_df2 = pd.read_csv(Path_test + \"test_static_0_1.csv\", index_col=False)\nprint(temp_df2.shape)\nprint(temp_df2.head(10))\n\n# Compare cols for *static_0_1 files (Train and test)\n\ncol_compare(temp_df1, temp_df2)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:16:01.162218Z","iopub.execute_input":"2024-02-15T04:16:01.162696Z","iopub.status.idle":"2024-02-15T04:16:21.104863Z","shell.execute_reply.started":"2024-02-15T04:16:01.162655Z","shell.execute_reply":"2024-02-15T04:16:21.103519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check train file\ntemp_df1 = pd.read_csv(Path_train + \"train_static_cb_0.csv\", index_col=False)\nprint(temp_df1.shape)\nprint(temp_df1.head(10))\n\n\n# Check test file\ntemp_df2 = pd.read_csv(Path_test + \"test_static_cb_0.csv\", index_col=False)\nprint(temp_df2.shape)\nprint(temp_df2.head(10))\n\n# Compare cols for *static_cb_0 files (Train and test)\n\ncol_compare(temp_df1, temp_df2)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:16:21.107502Z","iopub.execute_input":"2024-02-15T04:16:21.108102Z","iopub.status.idle":"2024-02-15T04:16:33.978684Z","shell.execute_reply.started":"2024-02-15T04:16:21.108047Z","shell.execute_reply":"2024-02-15T04:16:33.977292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"def tr_df(base, filenames, file_loc):\n    df_merge = base\n    for filename in filenames:\n        temp_df = pd.read_csv(file_loc + filename, index_col=False).pipe(tr_dtypes)\n        print(\"Successfully read and applied dtypes transform on file: \"+filename+\"\\n attempting merge now...\")\n        df_merge = pd.merge(df_merge, temp_df, on=\"case_id\", how=\"left\", validate=\"one_to_many\")\n        print(\"\\n Successfully merged file: \"+filename)\n        print(df_merge.head(5))\n    print (\"\\n Merge operation successful.\")   \n    return df_merge","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:16:33.981025Z","iopub.execute_input":"2024-02-15T04:16:33.981555Z","iopub.status.idle":"2024-02-15T04:16:33.990380Z","shell.execute_reply.started":"2024-02-15T04:16:33.981505Z","shell.execute_reply":"2024-02-15T04:16:33.988898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base = df_train\nfilenames = [\"train_static_0_0.csv\", \"train_static_0_1.csv\", \"train_static_cb_0.csv\"]\nfile_loc = Path_train\n\ntrain_merge = tr_df(base, filenames, file_loc)\ntrain_merge.shape\n","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:15:56.245304Z","iopub.execute_input":"2024-02-15T05:15:56.247511Z","iopub.status.idle":"2024-02-15T05:17:09.475689Z","shell.execute_reply.started":"2024-02-15T05:15:56.247447Z","shell.execute_reply":"2024-02-15T05:17:09.474548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base = df_test\nfilenames = [\"test_static_0_0.csv\", \"test_static_0_1.csv\", \"test_static_cb_0.csv\"] #,\"test_static_0_2.csv\"\nfile_loc = Path_test\n\ntest_merge = tr_df(base, filenames, file_loc)\ntest_merge.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:17:09.478043Z","iopub.execute_input":"2024-02-15T05:17:09.479084Z","iopub.status.idle":"2024-02-15T05:17:09.847989Z","shell.execute_reply.started":"2024-02-15T05:17:09.479044Z","shell.execute_reply":"2024-02-15T05:17:09.846526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_merge.shape)\nprint(test_merge.shape)\n","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:17:09.849793Z","iopub.execute_input":"2024-02-15T05:17:09.850213Z","iopub.status.idle":"2024-02-15T05:17:09.857996Z","shell.execute_reply.started":"2024-02-15T05:17:09.850178Z","shell.execute_reply":"2024-02-15T05:17:09.856625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Null/ Missing/ Duplicate value treatment strategy\n\nWe will do a basic duplicate check on unique keys like \"case id\"\n\nWe have a lot many null values in our base file. We can treat null/ missing values in many ways -> removing them, replacing them with mean, median, mode or some other calculated value based on context.\n\nImportant thing to notice is, we have to balance those operations on both train and test.\n\nFor now, as first step we will just remove nulls where we have ","metadata":{}},{"cell_type":"code","source":"## Drop if null count of column higher than 80%\n\ndef tr_cols (df_train, df_test):\n    col_rm_list=[]\n    for col in df_train.columns:\n        if col not in [\"target\", \"case_id\", \"WEEK_NUM\"]:\n            isnull = df_train[col].isna().mean()\n            if isnull > 0.8:\n                col_rm_list.append(col)\n                \n    return col_rm_list","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:17:09.861312Z","iopub.execute_input":"2024-02-15T05:17:09.861731Z","iopub.status.idle":"2024-02-15T05:17:09.870448Z","shell.execute_reply.started":"2024-02-15T05:17:09.861697Z","shell.execute_reply":"2024-02-15T05:17:09.868960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rm_col = tr_cols(train_merge, test_merge)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:17:09.872898Z","iopub.execute_input":"2024-02-15T05:17:09.874938Z","iopub.status.idle":"2024-02-15T05:17:10.649581Z","shell.execute_reply.started":"2024-02-15T05:17:09.874883Z","shell.execute_reply":"2024-02-15T05:17:10.648452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in rm_col:\n    train_merge = train_merge.drop(col, axis=1)\n    test_merge = test_merge.drop(col, axis=1)\n    \nprint(train_merge.shape)\nprint(test_merge.shape)    ","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:17:10.651242Z","iopub.execute_input":"2024-02-15T05:17:10.653072Z","iopub.status.idle":"2024-02-15T05:18:19.474017Z","shell.execute_reply.started":"2024-02-15T05:17:10.653020Z","shell.execute_reply":"2024-02-15T05:18:19.472419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_merge.info()","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:19.475473Z","iopub.execute_input":"2024-02-15T05:18:19.475847Z","iopub.status.idle":"2024-02-15T05:18:19.556763Z","shell.execute_reply.started":"2024-02-15T05:18:19.475816Z","shell.execute_reply":"2024-02-15T05:18:19.555747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_merge.info()","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:19.558391Z","iopub.execute_input":"2024-02-15T05:18:19.559118Z","iopub.status.idle":"2024-02-15T05:18:19.633452Z","shell.execute_reply.started":"2024-02-15T05:18:19.559081Z","shell.execute_reply":"2024-02-15T05:18:19.631948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_merge.isna().mean()","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:19.634991Z","iopub.execute_input":"2024-02-15T05:18:19.635393Z","iopub.status.idle":"2024-02-15T05:18:20.169538Z","shell.execute_reply.started":"2024-02-15T05:18:19.635358Z","shell.execute_reply":"2024-02-15T05:18:20.168058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model & Results\n\nIf we reached this far.....\nMaybe, we should try making one prediction and make our first submission.\n\nWe are yet very far from an ideal situation, we haven't completed EDA, data is not updated with depth 1,2 files. Still, I think we should now make our first attempt, even it will be a poor score because:\n\n- An end-to-end run will give us good ideas about generic approach\n- Clear out some first time hurdles, and give us a positive mental boost (specially for beginners)\n- We will be able to think more broadly about the problem, and thanks to this knowledge can improve on our mistakes (there are many) and our solution\n\nSo, with strong desire to fail-fast :)","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# Prepare dataset for model fitting\n\n### For now, let's drop catagorical + Date columns\n### Replace 'NaN' with Median for numeric columns\n### Replacing Inf, -Inf with 0\n\n# Delete the dtype category column\ntrain_merge = train_merge.select_dtypes(exclude=['category','datetime64[ns]' ])\ntest_merge = test_merge.select_dtypes(exclude=['category', 'datetime64[ns]'])\n\n# Check NaN values and replace with Median\nprint(train_merge.isna().sum())\ntrain_merge = train_merge.fillna(train_merge.median())\ntest_merge = test_merge.fillna(train_merge.median())\n\n# Replacing Inf, -Inf with 0\ntrain_merge = train_merge.replace([np.inf, -np.inf], np.nan)\ntrain_merge = train_merge.fillna(0)\n\ntest_merge = test_merge.replace([np.inf, -np.inf], np.nan)\ntest_merge = test_merge.fillna(0)\n\n\n# Validate if we have any NaN remaining in train dataset\nprint(train_merge.isna().sum())\nprint(test_merge.isna().sum())","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:20.175626Z","iopub.execute_input":"2024-02-15T05:18:20.176702Z","iopub.status.idle":"2024-02-15T05:18:26.291797Z","shell.execute_reply.started":"2024-02-15T05:18:20.176652Z","shell.execute_reply":"2024-02-15T05:18:26.290399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# A quick dimention validation for our Train/ test sets\n\nprint(train_merge.shape)\nprint(test_merge.shape)  \n\nprint(train_merge[\"case_id\"].nunique()) # To check if all case ids are unique","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:26.293223Z","iopub.execute_input":"2024-02-15T05:18:26.293623Z","iopub.status.idle":"2024-02-15T05:18:26.357718Z","shell.execute_reply.started":"2024-02-15T05:18:26.293589Z","shell.execute_reply":"2024-02-15T05:18:26.356205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split, StratifiedGroupKFold, GridSearchCV\nfrom sklearn.preprocessing import StandardScaler, MinMaxScaler\nfrom sklearn.linear_model import LogisticRegression, LogisticRegressionCV\nfrom sklearn.metrics import auc, accuracy_score, roc_auc_score\n","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:26.359865Z","iopub.execute_input":"2024-02-15T05:18:26.360414Z","iopub.status.idle":"2024-02-15T05:18:26.368457Z","shell.execute_reply.started":"2024-02-15T05:18:26.360366Z","shell.execute_reply":"2024-02-15T05:18:26.366923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Identify X and y\nX = train_merge.drop(\"target\", axis=1)\ny = train_merge[\"target\"]\n\n#Create train/ test split\n\nX_train, X_test, y_train, y_test = train_test_split(\n  X,y , random_state=42,test_size=0.3, shuffle=True)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:26.370424Z","iopub.execute_input":"2024-02-15T05:18:26.370931Z","iopub.status.idle":"2024-02-15T05:18:30.007451Z","shell.execute_reply.started":"2024-02-15T05:18:26.370879Z","shell.execute_reply":"2024-02-15T05:18:30.006008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #Create BASE df for gini stability scoring\n\n# def base_df (m,tr,ts):\n#     base = pd.DataFrame({})\n#     for each cid in tr[\"case_id\"]:\n#         new_row = {\"case_id\":cid, \"WEEK_NUM\":tr[\"WEEK_NUM\"], \"target\":m[\"target\"]}\n#         base.append(new_row)\n        \n#     retrun base_df\n\n# base_merge_train = train_merge[\"case_id\", \"WEEK_NUM\", \"target\"]\n\n\n# def base_df (m,tr):\n\n#     for cid in tr[\"case_id\"]:\n#         base = m.loc[:, \"case_id\":cid]\n#     return base\n\n# base_train = base_df(train_merge, X_train)\n# base_test = base_df(train_merge, X_test)\n\n# print(base_test)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:30.008810Z","iopub.execute_input":"2024-02-15T05:18:30.009249Z","iopub.status.idle":"2024-02-15T05:18:30.016300Z","shell.execute_reply.started":"2024-02-15T05:18:30.009213Z","shell.execute_reply":"2024-02-15T05:18:30.014871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()  \n#scaler = MinMaxScaler()    #StandardScaler gives error due to datetime cols\nscaler.fit(X_train)\nX_train = scaler.transform(X_train)\nX_test = scaler.transform(X_test)\n\n#Testing dataset\ntest_X = scaler.transform(test_merge)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:30.017825Z","iopub.execute_input":"2024-02-15T05:18:30.018426Z","iopub.status.idle":"2024-02-15T05:18:32.112765Z","shell.execute_reply.started":"2024-02-15T05:18:30.018386Z","shell.execute_reply":"2024-02-15T05:18:32.111379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # LogReg classifier\n\n# LGR = LogisticRegression(random_state=42, solver=\"liblinear\", verbose=3)\n\n# # Base model without Hyperparameter tuning and cross-validation\n\n# LGR.fit(X_train, y_train)\n\n\n# # Predicted values\n# y_pred = LGR.predict(X_test)\n\n# # roc_auc_score\n# score = roc_auc_score(y_test, LGR.predict_proba(X_test)[:, 1])\n\n\n# print(Score)\n\n\n# param_grid = { \n#     'C': [0.001, 0.01, 0.1, 1, 10, 100, 1000],\n#     'penalty': ['l1', 'l2' ],\n#     'solver': ['newton-cg', 'lbfgs', 'liblinear', 'sag', 'saga']\n# }","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:32.115223Z","iopub.execute_input":"2024-02-15T05:18:32.115869Z","iopub.status.idle":"2024-02-15T05:18:32.122650Z","shell.execute_reply.started":"2024-02-15T05:18:32.115820Z","shell.execute_reply":"2024-02-15T05:18:32.121204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# y_pred.sum()","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:32.125004Z","iopub.execute_input":"2024-02-15T05:18:32.125732Z","iopub.status.idle":"2024-02-15T05:18:32.134132Z","shell.execute_reply.started":"2024-02-15T05:18:32.125685Z","shell.execute_reply":"2024-02-15T05:18:32.132746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Use logRegCV classifier for a simple first model, get score for test\n\nlr_cv = LogisticRegressionCV(solver=\"lbfgs\",\n                             multi_class=\"ovr\",\n                             max_iter=1000,\n                             n_jobs = -1,\n                             cv=5,\n                            verbose=0)\n\n# Fit model on Training Dataset\nlr_cv.fit(X_train, y_train)\n\n# Testing model on Validation dataset\ntrain_pred = lr_cv.predict_proba(X_test)[:, 1]\nscore = roc_auc_score(y_test, train_pred)\n\nprint(\"Training score: \", score)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:18:32.136071Z","iopub.execute_input":"2024-02-15T05:18:32.136606Z","iopub.status.idle":"2024-02-15T05:22:32.813651Z","shell.execute_reply.started":"2024-02-15T05:18:32.136562Z","shell.execute_reply":"2024-02-15T05:22:32.812108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Applying model on complete Train dataset\n\nvalid_pred = lr_cv.predict_proba(X)[:, 1]\nscore = roc_auc_score(y, valid_pred)\n\nprint(\"Validation score: \", score)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:22:32.816949Z","iopub.execute_input":"2024-02-15T05:22:32.817588Z","iopub.status.idle":"2024-02-15T05:22:33.883642Z","shell.execute_reply.started":"2024-02-15T05:22:32.817529Z","shell.execute_reply":"2024-02-15T05:22:33.881797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Finally, Applying model on Test dataset\n\ntest_pred = lr_cv.predict_proba(test_X)[:, 1]\n\n# Prepare submission DF\n\nsubmission = pd.DataFrame(test_merge.case_id)\nsubmission[\"score\"] = pd.Series(test_pred)\n#submission.set_index('case_id', inplace=True)\nprint(submission.head(10))\nprint(submission.info())","metadata":{"execution":{"iopub.status.busy":"2024-02-15T07:02:50.078666Z","iopub.execute_input":"2024-02-15T07:02:50.079136Z","iopub.status.idle":"2024-02-15T07:02:50.097836Z","shell.execute_reply.started":"2024-02-15T07:02:50.079102Z","shell.execute_reply":"2024-02-15T07:02:50.096320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Compare \"Submission file\" structure with \"Sample\"\n\nsample = pd.read_csv(Path+\"sample_submission.csv\")\nprint(sample.head())\nprint(sample.info())\n# /kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\n# /kaggle/input/home-credit-credit-risk-model-stability/feature_definitions.csv","metadata":{"execution":{"iopub.status.busy":"2024-02-15T07:06:23.169522Z","iopub.execute_input":"2024-02-15T07:06:23.170087Z","iopub.status.idle":"2024-02-15T07:06:23.195241Z","shell.execute_reply.started":"2024-02-15T07:06:23.170051Z","shell.execute_reply":"2024-02-15T07:06:23.193931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Write csv file\nsubmission.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T05:56:24.931908Z","iopub.execute_input":"2024-02-15T05:56:24.932392Z","iopub.status.idle":"2024-02-15T05:56:24.941696Z","shell.execute_reply.started":"2024-02-15T05:56:24.932355Z","shell.execute_reply":"2024-02-15T05:56:24.940563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# GridSearch with cross-validation\n\n# CV_LGR = GridSearchCV(estimator=LGR, param_grid=param_grid, cv= 5, n_jobs=-1, verbose=1, scoring=\"f1\")\n\n# CV_LGR.fit(X_train, y_train)\n\n#Get Best model from Gridsearch\n# LGR_best = CV_LGR.best_estimator_\n\n# print(\"Best logistic regression model auc:\", CV_LGR.best_score_)\n# print(\"Best logistic regression parameters:\", CV_LGR.best_params_)\n# print(\"Best logistic regression Classifier:\", LGR_best)","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:19:42.252964Z","iopub.status.idle":"2024-02-15T04:19:42.253525Z","shell.execute_reply.started":"2024-02-15T04:19:42.253253Z","shell.execute_reply":"2024-02-15T04:19:42.253276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Stability score function\n# # This competetion has an accuracy calculation provided\n# # Borrowing the calculation function from : https://www.kaggle.com/code/jetakow/home-credit-2024-starter-notebook\n\n# def gini_stability(base, w_fallingrate=88.0, w_resstd=-0.5):\n#     gini_in_time = base.loc[:, [\"WEEK_NUM\", \"target\", \"score\"]]\\\n#         .sort_values(\"WEEK_NUM\")\\\n#         .groupby(\"WEEK_NUM\")[[\"target\", \"score\"]]\\\n#         .apply(lambda x: 2*roc_auc_score(x[\"target\"], x[\"score\"])-1).tolist()\n    \n#     x = np.arange(len(gini_in_time))\n#     y = gini_in_time\n#     a, b = np.polyfit(x, y, 1)\n#     y_hat = a*x + b\n#     residuals = y - y_hat\n#     res_std = np.std(residuals)\n#     avg_gini = np.mean(gini_in_time)\n#     return avg_gini + w_fallingrate * min(0, a) + w_resstd * res_std","metadata":{"execution":{"iopub.status.busy":"2024-02-15T04:19:42.254893Z","iopub.status.idle":"2024-02-15T04:19:42.255466Z","shell.execute_reply.started":"2024-02-15T04:19:42.255194Z","shell.execute_reply":"2024-02-15T04:19:42.255217Z"},"trusted":true},"execution_count":null,"outputs":[]}]}