{
  "id": 473950,
  "title": "Understanding completion data",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/473950",
  "author_name": "SSS",
  "post_date": "2024-02-06T15:36:00.605000",
  "votes": 375,
  "comment_count": 79,
  "views": 0,
  "content": "<p>There are <strong>32</strong> files in <code>train</code> csv_files folder. The same files provided in <code>.parquet</code> format in the respective folder. So the competition data might be confusing for some of us. Some csvs are greater than <strong><code>2 GB</code></strong> in size. <code>train_base.csv</code> has <strong>1526659</strong> unique <code>case_id</code> which equals the length of this file.</p>\n<h2>Explanation and schema</h2>\n<p>There are <strong>465</strong> features and <strong>436</strong> respective descriptions in <code>feature_definitions.csv</code>. There are no missing descriptions so it means some feature might have same descriptions (for example description <code>Number of tax deductions</code> for features: <code>pmtcount_4527229L</code>, <code>pmtcount_4955617L</code>, <code>pmtcount_693L</code>).</p>\n<p>We are going to use <code>case_id</code> to merge underlying tables. We got <strong>internal</strong> and <strong>external</strong> data sources and this is how the schema looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F901c94931ce7f953983680c1155f5747%2Fschema.png?generation=1707250963705097&amp;alt=media\"></p>\n<p>We got tables classified by depth, where:</p>\n<ul>\n<li>depth=0 - These are static features directly tied to a specific <code>case_id</code>.</li>\n<li>depth=1 - Each <code>case_id</code> has an associated historical record, indexed by <code>num_group1</code>.</li>\n<li>depth=2 - Each <code>case_id</code> has an associated historical record, indexed by both <code>num_group1</code> and <code>num_group2</code>.</li>\n</ul>\n<p>Various predictors were transformed, so to have the following notation for similar groups of transformations: <strong><code>P M A D T L</code></strong></p>\n<h2>Train File sizes and Null Values</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0fdebbea69c5090f5fb890de72f23e4e%2Fscaled_nulls_updated.png?generation=1707271157104551&amp;alt=media\"></p>\n<p><strong>How to read it</strong>:</p>\n<ul>\n<li>The sizes of squares are scaled across all <code>train</code> files, where the biggest area is the biggest file;</li>\n<li>Red squares are normalized total counts of all records across all columns in the respective <code>train</code> file;</li>\n<li>Black squares are normalized total counts of <strong>null</strong> records across all columns in the respective <code>train</code> file.</li>\n</ul>\n<p>For example, It means the <code>train</code> files are  <strong>&gt;30%</strong> empty on average. The <code>credit_bureau</code> files are the biggest ones.</p>\n<h1>Feature selection</h1>\n<p>TODO</p>\n<h2>Have Fun!</h2>\n<p><a href=\"https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission\" target=\"_blank\">Code</a> - slowly put flesh on the bone.<br>\n<a href=\"https://www.kaggle.com/datasets/sergiosaharovskiy/2024-home-credit-public-repo\" target=\"_blank\">Repo</a> - external scripts</p>",
  "messages": [
    {
      "id": 2638964,
      "postDate": "2024-02-06T15:36:00.607Z",
      "content": "<p>There are <strong>32</strong> files in <code>train</code> csv_files folder. The same files provided in <code>.parquet</code> format in the respective folder. So the competition data might be confusing for some of us. Some csvs are greater than <strong><code>2 GB</code></strong> in size. <code>train_base.csv</code> has <strong>1526659</strong> unique <code>case_id</code> which equals the length of this file.</p>\n<h2>Explanation and schema</h2>\n<p>There are <strong>465</strong> features and <strong>436</strong> respective descriptions in <code>feature_definitions.csv</code>. There are no missing descriptions so it means some feature might have same descriptions (for example description <code>Number of tax deductions</code> for features: <code>pmtcount_4527229L</code>, <code>pmtcount_4955617L</code>, <code>pmtcount_693L</code>).</p>\n<p>We are going to use <code>case_id</code> to merge underlying tables. We got <strong>internal</strong> and <strong>external</strong> data sources and this is how the schema looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F901c94931ce7f953983680c1155f5747%2Fschema.png?generation=1707250963705097&amp;alt=media\"></p>\n<p>We got tables classified by depth, where:</p>\n<ul>\n<li>depth=0 - These are static features directly tied to a specific <code>case_id</code>.</li>\n<li>depth=1 - Each <code>case_id</code> has an associated historical record, indexed by <code>num_group1</code>.</li>\n<li>depth=2 - Each <code>case_id</code> has an associated historical record, indexed by both <code>num_group1</code> and <code>num_group2</code>.</li>\n</ul>\n<p>Various predictors were transformed, so to have the following notation for similar groups of transformations: <strong><code>P M A D T L</code></strong></p>\n<h2>Train File sizes and Null Values</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0fdebbea69c5090f5fb890de72f23e4e%2Fscaled_nulls_updated.png?generation=1707271157104551&amp;alt=media\"></p>\n<p><strong>How to read it</strong>:</p>\n<ul>\n<li>The sizes of squares are scaled across all <code>train</code> files, where the biggest area is the biggest file;</li>\n<li>Red squares are normalized total counts of all records across all columns in the respective <code>train</code> file;</li>\n<li>Black squares are normalized total counts of <strong>null</strong> records across all columns in the respective <code>train</code> file.</li>\n</ul>\n<p>For example, It means the <code>train</code> files are  <strong>&gt;30%</strong> empty on average. The <code>credit_bureau</code> files are the biggest ones.</p>\n<h1>Feature selection</h1>\n<p>TODO</p>\n<h2>Have Fun!</h2>\n<p><a href=\"https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission\" target=\"_blank\">Code</a> - slowly put flesh on the bone.<br>\n<a href=\"https://www.kaggle.com/datasets/sergiosaharovskiy/2024-home-credit-public-repo\" target=\"_blank\">Repo</a> - external scripts</p>",
      "rawMarkdown": "There are **32** files in `train` csv_files folder. The same files provided in `.parquet` format in the respective folder. So the competition data might be confusing for some of us. Some csvs are greater than **`2 GB`** in size. `train_base.csv` has **1526659** unique `case_id` which equals the length of this file.\n\n##Explanation and schema\nThere are **465** features and **436** respective descriptions in `feature_definitions.csv`. There are no missing descriptions so it means some feature might have same descriptions (for example description `Number of tax deductions` for features: `pmtcount_4527229L`, `pmtcount_4955617L`, `pmtcount_693L`).\n\nWe are going to use `case_id` to merge underlying tables. We got **internal** and **external** data sources and this is how the schema looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F901c94931ce7f953983680c1155f5747%2Fschema.png?generation=1707250963705097&alt=media)\n\nWe got tables classified by depth, where:\n<ul>\n<li>depth=0 - These are static features directly tied to a specific <code>case_id</code>.</li>\n<li>depth=1 - Each <code>case_id</code> has an associated historical record, indexed by <code>num_group1</code>.</li>\n<li>depth=2 - Each <code>case_id</code> has an associated historical record, indexed by both <code>num_group1</code> and <code>num_group2</code>.</li>\n</ul>\n\nVarious predictors were transformed, so to have the following notation for similar groups of transformations: **`P M A D T L`**\n\n## Train File sizes and Null Values\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0fdebbea69c5090f5fb890de72f23e4e%2Fscaled_nulls_updated.png?generation=1707271157104551&alt=media)\n\n**How to read it**:\n- The sizes of squares are scaled across all `train` files, where the biggest area is the biggest file;\n- Red squares are normalized total counts of all records across all columns in the respective `train` file;\n- Black squares are normalized total counts of **null** records across all columns in the respective `train` file.\n\nFor example, It means the `train` files are  **>30%** empty on average. The `credit_bureau` files are the biggest ones.\n\n#Feature selection\nTODO\n\n##Have Fun!\n\n[Code](https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission) - slowly put flesh on the bone.\n[Repo](https://www.kaggle.com/datasets/sergiosaharovskiy/2024-home-credit-public-repo) - external scripts",
      "votes": 367
    },
    {
      "id": 2644886,
      "postDate": "2024-02-09T18:31:42.917Z",
      "content": "<p>A few interesting features based on SHAP output. (I'll rerun and colour by \"future week\", which should make clear features that change over time).</p>\n<p>Common sense:</p>\n<ul>\n<li>Higher Annuity (annuity_780A) -&gt; more likely to default</li>\n<li>Younger individuals (birth_259D) -&gt; more likely to default</li>\n<li>recently started employment (empl_employedfrom_271D) -&gt; more likely to default.</li>\n<li>number of months without payments (cntpmts24_3658933L), higher -&gt; more likely to default</li>\n<li>Males are more likely to default.</li>\n</ul>\n<p>Less obvious</p>\n<ul>\n<li>More individuals sharing a mobile number (mobilephncnt_593L) -&gt; more likely to default</li>\n<li>Weekday of decision -&gt; less likely to default if made at the weekend. Odd.</li>\n<li>Number of loan payments made (pmtnum_254L) -&gt; least likely to default ~12 </li>\n</ul>\n<p>On the y-axis +ve =&gt; evidence in favour, -ve =&gt; evidence against. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2Fc76f12979967d54fbcbd51f8cfee6476%2Fhc_shap.png?generation=1707502502893240&amp;alt=media\"></p>",
      "rawMarkdown": "A few interesting features based on SHAP output. (I'll rerun and colour by \"future week\", which should make clear features that change over time).\n\nCommon sense:\n- Higher Annuity (annuity_780A) -> more likely to default\n- Younger individuals (birth_259D) -> more likely to default\n- recently started employment (empl_employedfrom_271D) -> more likely to default.\n- number of months without payments (cntpmts24_3658933L), higher -> more likely to default\n- Males are more likely to default.\n\nLess obvious\n- More individuals sharing a mobile number (mobilephncnt_593L) -> more likely to default\n- Weekday of decision -> less likely to default if made at the weekend. Odd.\n- Number of loan payments made (pmtnum_254L) -> least likely to default ~12 \n\nOn the y-axis +ve => evidence in favour, -ve => evidence against. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2Fc76f12979967d54fbcbd51f8cfee6476%2Fhc_shap.png?generation=1707502502893240&alt=media)",
      "votes": 69,
      "replies": [
        {
          "id": 2644968,
          "postDate": "2024-02-09T19:50:17.333Z",
          "content": "<p><a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> good stuff! glad to see you here, it's been a while.</p>",
          "rawMarkdown": "@paddykb good stuff! glad to see you here, it's been a while.",
          "votes": 3
        },
        {
          "id": 2645105,
          "postDate": "2024-02-10T00:26:33.937Z",
          "content": "<p>Very nice!</p>",
          "rawMarkdown": "Very nice!",
          "votes": 2
        },
        {
          "id": 2650090,
          "postDate": "2024-02-13T08:38:39.650Z",
          "content": "<p>Thank you for good information! I am curious about the code related to the above. Is it possible to share? </p>",
          "rawMarkdown": "Thank you for good information! I am curious about the code related to the above. Is it possible to share? ",
          "votes": 2,
          "replies": [
            {
              "id": 2650262,
              "postDate": "2024-02-13T10:42:58.750Z",
              "content": "<p>Different data, but the code is the same as cell 12 here: <a href=\"https://www.kaggle.com/code/paddykb/lgbm-mapie-birth-weight-oh-my\" target=\"_blank\">https://www.kaggle.com/code/paddykb/lgbm-mapie-birth-weight-oh-my</a></p>\n<p>During the CV loop, I used one fold to calculate <a href=\"https://christophm.github.io/interpretable-ml-book/shapley.html\" target=\"_blank\">shap-values</a> on the validation data:</p>\n<pre><code> shap\nshap_vl = vl.copy()\nshap_values = shap.TreeExplainer(model).shap_values(shap_vl)\n</code></pre>\n<p>For this diagram, I subsampled 20,000 rows (shap_vl = vl.sample(20000)), so it ran in a reasonable time.</p>\n<p>And plotted the \"most important\" variables, again based on shap:</p>\n<pre><code>= pd.DataFrame(columns = \n\npd.DataFrame({\n    :  \n    :     .sort_values([], ascending=[False])\n</code></pre>",
              "rawMarkdown": "Different data, but the code is the same as cell 12 here: https://www.kaggle.com/code/paddykb/lgbm-mapie-birth-weight-oh-my\n\nDuring the CV loop, I used one fold to calculate [shap-values](https://christophm.github.io/interpretable-ml-book/shapley.html) on the validation data:\n\n```python\nimport shap\nshap_vl = vl.copy()\nshap_values = shap.TreeExplainer(model).shap_values(shap_vl)\n```\nFor this diagram, I subsampled 20,000 rows (shap_vl = vl.sample(20000)), so it ran in a reasonable time.\n\nAnd plotted the \"most important\" variables, again based on shap:\n\n```\nshap_df = pd.DataFrame(shap_values[1], columns = shap_vl.columns)\n\n# variable importance:\npd.DataFrame({\n    'imp': shap_df.abs().mean(),  # this is how the shap library defines importance!\n    'col': shap_df.columns})\n    .sort_values(['imp'], ascending=[False])\n```\n\n",
              "votes": 9
            },
            {
              "id": 2650508,
              "postDate": "2024-02-13T13:50:10.447Z",
              "content": "<p>Thanks to paddykb's detailed explanation, I had the opportunity to study shap. <br>\nI also enjoyed reviewing your code. thank you 😀</p>\n<p>In particular, -(-len_features // N_X) was fun! </p>\n<p>math.ceil(len_features / N_X)<br>\nAbout expressing this by using minus twice</p>\n<p>I'm curious as to whether your method is a common method. 👍</p>",
              "rawMarkdown": "Thanks to paddykb's detailed explanation, I had the opportunity to study shap. \nI also enjoyed reviewing your code. thank you 😀\n\nIn particular, -(-len_features // N_X) was fun! \n\nmath.ceil(len_features / N_X)\nAbout expressing this by using minus twice\n\nI'm curious as to whether your method is a common method. 👍",
              "votes": 2
            },
            {
              "id": 2650619,
              "postDate": "2024-02-13T15:25:35.383Z",
              "content": "<p>Probably just a bad habit when math isn't imported 😀</p>",
              "rawMarkdown": "Probably just a bad habit when math isn't imported 😀",
              "votes": 2
            }
          ]
        },
        {
          "id": 2751463,
          "postDate": "2024-04-14T09:07:35.697Z",
          "content": "<p>Thank you for that!</p>",
          "rawMarkdown": "Thank you for that!"
        }
      ]
    },
    {
      "id": 2652740,
      "postDate": "2024-02-15T02:03:56.220Z",
      "content": "<p>Thanks alott for sharing , its very helpfullll !!!!</p>",
      "rawMarkdown": "Thanks alott for sharing , its very helpfullll !!!!",
      "votes": 14
    },
    {
      "id": 2796893,
      "postDate": "2024-05-06T13:01:18.950Z",
      "content": "<p>Thanks Sergey for the great work. What is the package/software that you used for generating the visualization of the NULL values%? Thanks! </p>",
      "rawMarkdown": "Thanks Sergey for the great work. What is the package/software that you used for generating the visualization of the NULL values%? Thanks! ",
      "votes": 3
    },
    {
      "id": 2841951,
      "postDate": "2024-05-28T19:07:43.860Z",
      "content": "<p>Thanks alott for sharing! Well explained Really, helpful to startoff thanks.</p>",
      "rawMarkdown": "Thanks alott for sharing! Well explained Really, helpful to startoff thanks.",
      "votes": 1
    },
    {
      "id": 2646243,
      "postDate": "2024-02-10T18:28:33.933Z",
      "content": "<p>Well explained </p>",
      "rawMarkdown": "Well explained ",
      "votes": 3
    },
    {
      "id": 2645587,
      "postDate": "2024-02-10T10:42:35.407Z",
      "content": "<p>Clear overview! Your detailed insights into file sizes, feature descriptions, and schema provide a solid understanding of the competition data. The visualization aids comprehension. Thanks!</p>",
      "rawMarkdown": "Clear overview! Your detailed insights into file sizes, feature descriptions, and schema provide a solid understanding of the competition data. The visualization aids comprehension. Thanks!",
      "votes": 3
    },
    {
      "id": 2644329,
      "postDate": "2024-02-09T11:54:06.960Z",
      "content": "<p>Thanks for the awesome visualization, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. </p>\n<p>I have one question though. As nulls are still a part of the training data, shouldn't the black squares be always inside of the red ones? It would be automatic for pie charts, but I'm not sure what determines position of blacks with respect to reds on your infographics.</p>",
      "rawMarkdown": "Thanks for the awesome visualization, @sergiosaharovskiy. \n\nI have one question though. As nulls are still a part of the training data, shouldn't the black squares be always inside of the red ones? It would be automatic for pie charts, but I'm not sure what determines position of blacks with respect to reds on your infographics.",
      "votes": 4,
      "replies": [
        {
          "id": 2644349,
          "postDate": "2024-02-09T12:08:13.910Z",
          "content": "<p><a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> Correct, by default it centered and the black squares are the subset of red. I moved black squares to the left top corner on purpose, so it can be seen in case of small squares. This particular chart should not be read as venn diagram of any sort.</p>",
          "rawMarkdown": "@kononenko Correct, by default it centered and the black squares are the subset of red. I moved black squares to the left top corner on purpose, so it can be seen in case of small squares. This particular chart should not be read as venn diagram of any sort.",
          "votes": 4
        }
      ]
    },
    {
      "id": 2745033,
      "postDate": "2024-04-10T10:15:25.013Z",
      "content": "<p>The data is a little confusing, but thanks for the clarification. I wonder what does \"transform\" mean in this context, and why there are numbers before transform type</p>",
      "rawMarkdown": "The data is a little confusing, but thanks for the clarification. I wonder what does \"transform\" mean in this context, and why there are numbers before transform type",
      "votes": 1,
      "replies": [
        {
          "id": 2746511,
          "postDate": "2024-04-11T09:20:02.367Z",
          "content": "<p>The numbers are just our internal IDs used for naming the columns. There are cases when the name is still same, so the IDs can separate them. There is no hidden meaning to those numbers.</p>",
          "rawMarkdown": "The numbers are just our internal IDs used for naming the columns. There are cases when the name is still same, so the IDs can separate them. There is no hidden meaning to those numbers.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2796506,
      "postDate": "2024-05-06T09:14:32.957Z",
      "content": "<p>Thanks for leading me to try such competition, especially the visualization (they are too clear to use!!!) and since I am a beginner in the Risk Management Industry.</p>",
      "rawMarkdown": "Thanks for leading me to try such competition, especially the visualization (they are too clear to use!!!) and since I am a beginner in the Risk Management Industry.\n",
      "votes": 2
    },
    {
      "id": 2790341,
      "postDate": "2024-05-03T06:29:52.273Z",
      "content": "<p>Helpful thread.</p>",
      "rawMarkdown": "Helpful thread.",
      "votes": 2
    },
    {
      "id": 2775310,
      "postDate": "2024-04-25T15:51:11.920Z",
      "content": "<p>I am a beginner. This is my first time participating in a competition. Thank you for your valuable information.</p>",
      "rawMarkdown": "I am a beginner. This is my first time participating in a competition. Thank you for your valuable information.",
      "votes": 2
    },
    {
      "id": 2674237,
      "postDate": "2024-02-29T07:23:16.397Z",
      "content": "<p>Thanks a lot for sharing the insight on the dataset!</p>",
      "rawMarkdown": "Thanks a lot for sharing the insight on the dataset!",
      "votes": 1
    },
    {
      "id": 2666293,
      "postDate": "2024-02-24T09:46:55.143Z",
      "content": "<p>The scheme is very useful</p>",
      "rawMarkdown": "The scheme is very useful",
      "votes": 1
    },
    {
      "id": 2646210,
      "postDate": "2024-02-10T17:51:06.497Z",
      "content": "<p>Really Helpful Solid understanding of the competition data. The visualization comprehension also very good. Thanks.</p>",
      "rawMarkdown": "Really Helpful Solid understanding of the competition data. The visualization comprehension also very good. Thanks.",
      "votes": 1
    },
    {
      "id": 2644173,
      "postDate": "2024-02-09T10:07:55.167Z",
      "content": "<p>Really, helpful to startoff.<br>\nThank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.</p>\n<p>Keep it up and motivate us so that we can also do more better like you </p>",
      "rawMarkdown": "Really, helpful to startoff.\nThank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.\n\nKeep it up and motivate us so that we can also do more better like you ",
      "votes": 1
    },
    {
      "id": 2641669,
      "postDate": "2024-02-07T15:48:51.787Z",
      "content": "<p>Really helpful!</p>",
      "rawMarkdown": "Really helpful!",
      "votes": 1
    },
    {
      "id": 2640550,
      "postDate": "2024-02-07T01:24:38.497Z",
      "content": "<p>Really Commendable job!! Your visualization work makes the entire structure clear</p>",
      "rawMarkdown": "Really Commendable job!! Your visualization work makes the entire structure clear",
      "votes": 1
    },
    {
      "id": 2638966,
      "postDate": "2024-02-06T15:36:55.303Z",
      "content": "<p>VERY HELPFUL!!!</p>",
      "rawMarkdown": "VERY HELPFUL!!!",
      "votes": 1
    },
    {
      "id": 2649258,
      "postDate": "2024-02-12T17:39:27.233Z",
      "content": "<p>Very nice!</p>",
      "rawMarkdown": "Very nice!",
      "votes": 2
    },
    {
      "id": 2648782,
      "postDate": "2024-02-12T12:19:57.013Z",
      "content": "<p>well explained</p>",
      "rawMarkdown": "well explained",
      "votes": 2
    },
    {
      "id": 2648201,
      "postDate": "2024-02-12T06:14:19.747Z",
      "content": "<p>Got it so well! Done with my first task.</p>",
      "rawMarkdown": "Got it so well! Done with my first task.",
      "votes": 2
    },
    {
      "id": 2646392,
      "postDate": "2024-02-10T21:01:39.767Z",
      "content": "<p>Awesome, the initial sets are a bit overwhelming. Clear and concise!</p>",
      "rawMarkdown": "Awesome, the initial sets are a bit overwhelming. Clear and concise!",
      "votes": 2
    },
    {
      "id": 2645167,
      "postDate": "2024-02-10T02:24:45.897Z",
      "content": "<p>Test path have 4 additional files which are not for train. Why is that?</p>\n<pre><code> ,\n ,\n ,\n \n</code></pre>",
      "rawMarkdown": "Test path have 4 additional files which are not for train. Why is that?\n\n```python\n 'applprev_1_2.parquet',\n 'credit_bureau_a_1_4.parquet',\n 'credit_bureau_a_2_11.parquet',\n 'static_0_2.parquet'\n```",
      "votes": 2,
      "replies": [
        {
          "id": 2645600,
          "postDate": "2024-02-10T10:57:23.743Z",
          "content": "<p>That's because I have split the files for test and train in order to get smaller files. The number of files is not the same in every case for train and test.</p>",
          "rawMarkdown": "That's because I have split the files for test and train in order to get smaller files. The number of files is not the same in every case for train and test.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2644645,
      "postDate": "2024-02-09T15:48:33.130Z",
      "content": "<p>Thanks so much! This is super helpful!<br>\nYour notebook is very educational!</p>",
      "rawMarkdown": "Thanks so much! This is super helpful!\nYour notebook is very educational!",
      "votes": 2
    },
    {
      "id": 2642750,
      "postDate": "2024-02-08T11:36:45.757Z",
      "content": "<p>Really, helpful to startoff. Thank you!! </p>",
      "rawMarkdown": "Really, helpful to startoff. Thank you!! ",
      "votes": 2
    },
    {
      "id": 2641717,
      "postDate": "2024-02-07T16:23:20.277Z",
      "content": "<p>Outstanding work to visualize the problem clearly. Thanks!</p>",
      "rawMarkdown": "Outstanding work to visualize the problem clearly. Thanks!",
      "votes": 2
    },
    {
      "id": 2639089,
      "postDate": "2024-02-06T17:05:23.077Z",
      "content": "<p>Thank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.</p>",
      "rawMarkdown": "Thank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.",
      "votes": 2
    },
    {
      "id": 2816474,
      "postDate": "2024-05-16T10:39:36.527Z",
      "content": "<p>Hello, can I know whether the data for depth equals 1 and equals 2, do I need to group them into like different clusters before joining the data into the train base data or I just have to aggregate all the features together in the train base data. Which one will be a better or a more suitable approach? Thank you 😄😄</p>\n<p>Besides, for the transform part is it just how the data is named in the file, because I can't understand what it is related to. Hope someone can clear my doubts. Thank you.</p>",
      "rawMarkdown": "Hello, can I know whether the data for depth equals 1 and equals 2, do I need to group them into like different clusters before joining the data into the train base data or I just have to aggregate all the features together in the train base data. Which one will be a better or a more suitable approach? Thank you 😄😄\n\nBesides, for the transform part is it just how the data is named in the file, because I can't understand what it is related to. Hope someone can clear my doubts. Thank you."
    },
    {
      "id": 2755643,
      "postDate": "2024-04-16T15:58:05.020Z",
      "content": "<p>greatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreat </p>",
      "rawMarkdown": "greatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreat "
    },
    {
      "id": 2739664,
      "postDate": "2024-04-07T07:28:47.340Z",
      "content": "<p>Thanks for sharing!! Your schema is very usefull.</p>",
      "rawMarkdown": "Thanks for sharing!! Your schema is very usefull."
    },
    {
      "id": 2728554,
      "postDate": "2024-04-02T10:03:11.607Z",
      "content": "<p>Really helpful!</p>",
      "rawMarkdown": "Really helpful!"
    },
    {
      "id": 2695535,
      "postDate": "2024-03-13T18:06:38.820Z",
      "content": "<p>What a great work! Thank you very much for sharing your work, it helps me a lot to understand the whole dataset!</p>",
      "rawMarkdown": "What a great work! Thank you very much for sharing your work, it helps me a lot to understand the whole dataset!"
    },
    {
      "id": 2694469,
      "postDate": "2024-03-13T05:16:12.810Z",
      "content": "<p>Thanks a lot for your overall explanation, very much appreciated to easily get into the data</p>",
      "rawMarkdown": "Thanks a lot for your overall explanation, very much appreciated to easily get into the data"
    },
    {
      "id": 2664483,
      "postDate": "2024-02-23T04:00:29.620Z",
      "content": "<p>Very nice!</p>",
      "rawMarkdown": "Very nice!"
    },
    {
      "id": 2659937,
      "postDate": "2024-02-20T08:16:31.930Z",
      "content": "<p>Confused by the features for a while, thanks for the clarification!</p>",
      "rawMarkdown": "Confused by the features for a while, thanks for the clarification!"
    },
    {
      "id": 2658753,
      "postDate": "2024-02-19T11:34:35.477Z",
      "content": "<p>thanksss, it really gives a favor.4</p>",
      "rawMarkdown": "thanksss, it really gives a favor.4"
    },
    {
      "id": 2656187,
      "postDate": "2024-02-17T14:25:03.750Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>, this will serve as a good starter to understand data. </p>",
      "rawMarkdown": "Thank you @sergiosaharovskiy, this will serve as a good starter to understand data. "
    },
    {
      "id": 2654622,
      "postDate": "2024-02-16T09:55:23.710Z",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> Thank you for sharing your insights, This is super helpful🫡</p>",
      "rawMarkdown": "@sergiosaharovskiy Thank you for sharing your insights, This is super helpful🫡"
    },
    {
      "id": 2654427,
      "postDate": "2024-02-16T06:23:33.123Z",
      "content": "<p>Helpful thread!</p>",
      "rawMarkdown": "Helpful thread!"
    },
    {
      "id": 2651177,
      "postDate": "2024-02-14T02:33:24.470Z",
      "content": "<p>Amazing visualizations thanks for indepth explanation</p>",
      "rawMarkdown": "Amazing visualizations thanks for indepth explanation"
    },
    {
      "id": 2650662,
      "postDate": "2024-02-13T15:56:21.483Z",
      "content": "<p>Helpful thread</p>",
      "rawMarkdown": "Helpful thread"
    },
    {
      "id": 2650399,
      "postDate": "2024-02-13T12:47:27.353Z",
      "content": "<p>Code was simple and well implemented.</p>",
      "rawMarkdown": "Code was simple and well implemented."
    },
    {
      "id": 2722838,
      "postDate": "2024-03-29T19:28:44.470Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2658769,
      "postDate": "2024-02-19T11:45:47.990Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2667136,
      "postDate": "2024-02-24T22:34:38.317Z",
      "content": "<p>Thanks! 😁</p>",
      "rawMarkdown": "Thanks! 😁",
      "votes": 1
    },
    {
      "id": 2662661,
      "postDate": "2024-02-22T04:00:21.967Z",
      "content": "<p>Thank you for the explanation!!</p>",
      "rawMarkdown": "Thank you for the explanation!!",
      "votes": 1
    },
    {
      "id": 2661175,
      "postDate": "2024-02-21T04:37:51.620Z",
      "content": "<p>Thanks so much! This is very helpfulll</p>",
      "rawMarkdown": "Thanks so much! This is very helpfulll",
      "votes": 1
    },
    {
      "id": 2656431,
      "postDate": "2024-02-17T17:35:05.530Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 1
    },
    {
      "id": 2656080,
      "postDate": "2024-02-17T12:46:39.817Z",
      "content": "<p>Very helpful,  thanks a lot!!!!!</p>",
      "rawMarkdown": "Very helpful,  thanks a lot!!!!!",
      "votes": 1
    },
    {
      "id": 2652136,
      "postDate": "2024-02-14T15:22:56.703Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": 1
    },
    {
      "id": 2648659,
      "postDate": "2024-02-12T10:49:46.877Z",
      "content": "<p>This is very helpful Thank!</p>",
      "rawMarkdown": "This is very helpful Thank!",
      "votes": 1
    },
    {
      "id": 2647923,
      "postDate": "2024-02-11T21:49:59.100Z",
      "content": "<p>Good explanation thanks.</p>",
      "rawMarkdown": "Good explanation thanks.",
      "votes": 1
    },
    {
      "id": 2647434,
      "postDate": "2024-02-11T14:36:31.583Z",
      "content": "<p>Good explanation, thanks</p>",
      "rawMarkdown": "Good explanation, thanks",
      "votes": 1
    },
    {
      "id": 2646693,
      "postDate": "2024-02-11T06:32:08.130Z",
      "content": "<p>It was easy to understand. Thank you.😀</p>",
      "rawMarkdown": "It was easy to understand. Thank you.😀",
      "votes": 1
    },
    {
      "id": 2643982,
      "postDate": "2024-02-09T08:07:04.797Z",
      "content": "<p>thanks! really helpful!</p>",
      "rawMarkdown": "thanks! really helpful!",
      "votes": 1
    },
    {
      "id": 2643681,
      "postDate": "2024-02-09T03:21:18.210Z",
      "content": "<p>Really, helpful thanks for clarifying</p>",
      "rawMarkdown": "Really, helpful thanks for clarifying",
      "votes": 1
    },
    {
      "id": 2642914,
      "postDate": "2024-02-08T13:57:59.023Z",
      "content": "<p>Really, Helpful sir . Thank you</p>",
      "rawMarkdown": "Really, Helpful sir . Thank you\n",
      "votes": 1
    },
    {
      "id": 2642655,
      "postDate": "2024-02-08T10:23:44.987Z",
      "content": "<p>Really helpful thanks for clarifying</p>",
      "rawMarkdown": "Really helpful thanks for clarifying\n",
      "votes": 1
    },
    {
      "id": 2642001,
      "postDate": "2024-02-07T20:20:08.573Z",
      "content": "<p>Thank you! Incredible work!</p>",
      "rawMarkdown": "Thank you! Incredible work!",
      "votes": 1
    },
    {
      "id": 2640883,
      "postDate": "2024-02-07T06:07:18.127Z",
      "content": "<p>Thanks so much! This is super helpful!</p>",
      "rawMarkdown": "Thanks so much! This is super helpful!",
      "votes": 1
    },
    {
      "id": 2640547,
      "postDate": "2024-02-07T01:20:40.207Z",
      "content": "<p>Thank you. Very helpful.</p>",
      "rawMarkdown": "Thank you. Very helpful.",
      "votes": 1
    },
    {
      "id": 2790045,
      "postDate": "2024-05-03T01:47:57.447Z",
      "content": "<p>Thank you for sharing this!</p>",
      "rawMarkdown": "Thank you for sharing this!"
    },
    {
      "id": 2751562,
      "postDate": "2024-04-14T10:43:34.037Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 2732616,
      "postDate": "2024-04-03T08:21:47.223Z",
      "content": "<p>Thank you for good information!!</p>",
      "rawMarkdown": "Thank you for good information!!"
    },
    {
      "id": 2681462,
      "postDate": "2024-03-04T16:50:06.137Z",
      "content": "<p>Awesome!! Thank you!</p>",
      "rawMarkdown": "Awesome!! Thank you!"
    },
    {
      "id": 2669737,
      "postDate": "2024-02-26T13:17:04.447Z",
      "content": "<p>Really helpful explanation. Thanks!</p>",
      "rawMarkdown": "Really helpful explanation. Thanks!"
    },
    {
      "id": 2669677,
      "postDate": "2024-02-26T12:21:39.233Z",
      "content": "<p>Thanks! very well explained</p>",
      "rawMarkdown": "Thanks! very well explained"
    },
    {
      "id": 2664543,
      "postDate": "2024-02-23T05:32:57.103Z",
      "content": "<p>thanks! very nice</p>",
      "rawMarkdown": "thanks! very nice"
    },
    {
      "id": 2653627,
      "postDate": "2024-02-15T15:00:24.647Z",
      "content": "<p>Very clear, thanks!</p>",
      "rawMarkdown": "Very clear, thanks!",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2644886,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "2024-02-09T18:31:42.917000",
      "content": "<p>A few interesting features based on SHAP output. (I'll rerun and colour by \"future week\", which should make clear features that change over time).</p>\n<p>Common sense:</p>\n<ul>\n<li>Higher Annuity (annuity_780A) -&gt; more likely to default</li>\n<li>Younger individuals (birth_259D) -&gt; more likely to default</li>\n<li>recently started employment (empl_employedfrom_271D) -&gt; more likely to default.</li>\n<li>number of months without payments (cntpmts24_3658933L), higher -&gt; more likely to default</li>\n<li>Males are more likely to default.</li>\n</ul>\n<p>Less obvious</p>\n<ul>\n<li>More individuals sharing a mobile number (mobilephncnt_593L) -&gt; more likely to default</li>\n<li>Weekday of decision -&gt; less likely to default if made at the weekend. Odd.</li>\n<li>Number of loan payments made (pmtnum_254L) -&gt; least likely to default ~12 </li>\n</ul>\n<p>On the y-axis +ve =&gt; evidence in favour, -ve =&gt; evidence against. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2Fc76f12979967d54fbcbd51f8cfee6476%2Fhc_shap.png?generation=1707502502893240&amp;alt=media\"></p>",
      "votes": 69,
      "replies": [
        {
          "id": 2644968,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-02-09T19:50:17.333000",
          "content": "<p><a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> good stuff! glad to see you here, it's been a while.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2645105,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-10T00:26:33.937000",
          "content": "<p>Very nice!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2650090,
          "author_name": "faeqsu10",
          "author_url": "",
          "post_date": "2024-02-13T08:38:39.650000",
          "content": "<p>Thank you for good information! I am curious about the code related to the above. Is it possible to share? </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2650262,
              "author_name": "paddykb",
              "author_url": "",
              "post_date": "2024-02-13T10:42:58.750000",
              "content": "<p>Different data, but the code is the same as cell 12 here: <a href=\"https://www.kaggle.com/code/paddykb/lgbm-mapie-birth-weight-oh-my\" target=\"_blank\">https://www.kaggle.com/code/paddykb/lgbm-mapie-birth-weight-oh-my</a></p>\n<p>During the CV loop, I used one fold to calculate <a href=\"https://christophm.github.io/interpretable-ml-book/shapley.html\" target=\"_blank\">shap-values</a> on the validation data:</p>\n<pre><code> shap\nshap_vl = vl.copy()\nshap_values = shap.TreeExplainer(model).shap_values(shap_vl)\n</code></pre>\n<p>For this diagram, I subsampled 20,000 rows (shap_vl = vl.sample(20000)), so it ran in a reasonable time.</p>\n<p>And plotted the \"most important\" variables, again based on shap:</p>\n<pre><code>= pd.DataFrame(columns = \n\npd.DataFrame({\n    :  \n    :     .sort_values([], ascending=[False])\n</code></pre>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 2650508,
              "author_name": "faeqsu10",
              "author_url": "",
              "post_date": "2024-02-13T13:50:10.447000",
              "content": "<p>Thanks to paddykb's detailed explanation, I had the opportunity to study shap. <br>\nI also enjoyed reviewing your code. thank you 😀</p>\n<p>In particular, -(-len_features // N_X) was fun! </p>\n<p>math.ceil(len_features / N_X)<br>\nAbout expressing this by using minus twice</p>\n<p>I'm curious as to whether your method is a common method. 👍</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2650619,
              "author_name": "paddykb",
              "author_url": "",
              "post_date": "2024-02-13T15:25:35.383000",
              "content": "<p>Probably just a bad habit when math isn't imported 😀</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2751463,
          "author_name": "Evgeny Nevodnich",
          "author_url": "",
          "post_date": "2024-04-14T09:07:35.697000",
          "content": "<p>Thank you for that!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2652740,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-15T02:03:56.220000",
      "content": "<p>Thanks alott for sharing , its very helpfullll !!!!</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 2796893,
      "author_name": "faithk7u",
      "author_url": "",
      "post_date": "2024-05-06T13:01:18.950000",
      "content": "<p>Thanks Sergey for the great work. What is the package/software that you used for generating the visualization of the NULL values%? Thanks! </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2841951,
      "author_name": "Zahid Feroze",
      "author_url": "",
      "post_date": "2024-05-28T19:07:43.860000",
      "content": "<p>Thanks alott for sharing! Well explained Really, helpful to startoff thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2646243,
      "author_name": "Disha Agarwal",
      "author_url": "",
      "post_date": "2024-02-10T18:28:33.933000",
      "content": "<p>Well explained </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2645587,
      "author_name": "Devang Giri Goswami",
      "author_url": "",
      "post_date": "2024-02-10T10:42:35.407000",
      "content": "<p>Clear overview! Your detailed insights into file sizes, feature descriptions, and schema provide a solid understanding of the competition data. The visualization aids comprehension. Thanks!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2644329,
      "author_name": "Oleksiy Kononenko",
      "author_url": "",
      "post_date": "2024-02-09T11:54:06.960000",
      "content": "<p>Thanks for the awesome visualization, <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a>. </p>\n<p>I have one question though. As nulls are still a part of the training data, shouldn't the black squares be always inside of the red ones? It would be automatic for pie charts, but I'm not sure what determines position of blacks with respect to reds on your infographics.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2644349,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-02-09T12:08:13.910000",
          "content": "<p><a href=\"https://www.kaggle.com/kononenko\" target=\"_blank\">@kononenko</a> Correct, by default it centered and the black squares are the subset of red. I moved black squares to the left top corner on purpose, so it can be seen in case of small squares. This particular chart should not be read as venn diagram of any sort.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2745033,
      "author_name": "Roy Wei",
      "author_url": "",
      "post_date": "2024-04-10T10:15:25.013000",
      "content": "<p>The data is a little confusing, but thanks for the clarification. I wonder what does \"transform\" mean in this context, and why there are numbers before transform type</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2746511,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-04-11T09:20:02.367000",
          "content": "<p>The numbers are just our internal IDs used for naming the columns. There are cases when the name is still same, so the IDs can separate them. There is no hidden meaning to those numbers.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2796506,
      "author_name": "xinning chi",
      "author_url": "",
      "post_date": "2024-05-06T09:14:32.957000",
      "content": "<p>Thanks for leading me to try such competition, especially the visualization (they are too clear to use!!!) and since I am a beginner in the Risk Management Industry.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2790341,
      "author_name": "Heloise2103",
      "author_url": "",
      "post_date": "2024-05-03T06:29:52.273000",
      "content": "<p>Helpful thread.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2775310,
      "author_name": "HEKI",
      "author_url": "",
      "post_date": "2024-04-25T15:51:11.920000",
      "content": "<p>I am a beginner. This is my first time participating in a competition. Thank you for your valuable information.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2674237,
      "author_name": "jiarui_z",
      "author_url": "",
      "post_date": "2024-02-29T07:23:16.397000",
      "content": "<p>Thanks a lot for sharing the insight on the dataset!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2666293,
      "author_name": "Baffalo",
      "author_url": "",
      "post_date": "2024-02-24T09:46:55.143000",
      "content": "<p>The scheme is very useful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2646210,
      "author_name": "Zahid Feroze",
      "author_url": "",
      "post_date": "2024-02-10T17:51:06.497000",
      "content": "<p>Really Helpful Solid understanding of the competition data. The visualization comprehension also very good. Thanks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2644173,
      "author_name": "Tanishq dublish",
      "author_url": "",
      "post_date": "2024-02-09T10:07:55.167000",
      "content": "<p>Really, helpful to startoff.<br>\nThank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.</p>\n<p>Keep it up and motivate us so that we can also do more better like you </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2641669,
      "author_name": "Ashok Kumar",
      "author_url": "",
      "post_date": "2024-02-07T15:48:51.787000",
      "content": "<p>Really helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2640550,
      "author_name": "Subhasish Sinha",
      "author_url": "",
      "post_date": "2024-02-07T01:24:38.497000",
      "content": "<p>Really Commendable job!! Your visualization work makes the entire structure clear</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2638966,
      "author_name": "Simon Beck",
      "author_url": "",
      "post_date": "2024-02-06T15:36:55.303000",
      "content": "<p>VERY HELPFUL!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2649258,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-12T17:39:27.233000",
      "content": "<p>Very nice!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2648782,
      "author_name": "punnam neha",
      "author_url": "",
      "post_date": "2024-02-12T12:19:57.013000",
      "content": "<p>well explained</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2648201,
      "author_name": "Kusum Sunil Choudhary",
      "author_url": "",
      "post_date": "2024-02-12T06:14:19.747000",
      "content": "<p>Got it so well! Done with my first task.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2646392,
      "author_name": "Markus.JM",
      "author_url": "",
      "post_date": "2024-02-10T21:01:39.767000",
      "content": "<p>Awesome, the initial sets are a bit overwhelming. Clear and concise!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2645167,
      "author_name": "Moreno",
      "author_url": "",
      "post_date": "2024-02-10T02:24:45.897000",
      "content": "<p>Test path have 4 additional files which are not for train. Why is that?</p>\n<pre><code> ,\n ,\n ,\n \n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2645600,
          "author_name": "Daniel Herman",
          "author_url": "",
          "post_date": "2024-02-10T10:57:23.743000",
          "content": "<p>That's because I have split the files for test and train in order to get smaller files. The number of files is not the same in every case for train and test.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2644645,
      "author_name": "uec_2010670",
      "author_url": "",
      "post_date": "2024-02-09T15:48:33.130000",
      "content": "<p>Thanks so much! This is super helpful!<br>\nYour notebook is very educational!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2642750,
      "author_name": "Kishan Vavdara",
      "author_url": "",
      "post_date": "2024-02-08T11:36:45.757000",
      "content": "<p>Really, helpful to startoff. Thank you!! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2641717,
      "author_name": "Joaquin Fernandez",
      "author_url": "",
      "post_date": "2024-02-07T16:23:20.277000",
      "content": "<p>Outstanding work to visualize the problem clearly. Thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2639089,
      "author_name": "Evgeny Bratkovsky",
      "author_url": "",
      "post_date": "2024-02-06T17:05:23.077000",
      "content": "<p>Thank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2816474,
      "author_name": "EVAN WONG KUAN SENG",
      "author_url": "",
      "post_date": "2024-05-16T10:39:36.527000",
      "content": "<p>Hello, can I know whether the data for depth equals 1 and equals 2, do I need to group them into like different clusters before joining the data into the train base data or I just have to aggregate all the features together in the train base data. Which one will be a better or a more suitable approach? Thank you 😄😄</p>\n<p>Besides, for the transform part is it just how the data is named in the file, because I can't understand what it is related to. Hope someone can clear my doubts. Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2755643,
      "author_name": "whatnamedoyouhave",
      "author_url": "",
      "post_date": "2024-04-16T15:58:05.020000",
      "content": "<p>greatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreat </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2739664,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-07T07:28:47.340000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2728554,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-02T10:03:11.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2695535,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-13T18:06:38.820000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2694469,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-13T05:16:12.810000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2664483,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-23T04:00:29.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2659937,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-20T08:16:31.930000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2658753,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-19T11:34:35.477000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2656187,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-17T14:25:03.750000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2654622,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-16T09:55:23.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2654427,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-16T06:23:33.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2651177,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-14T02:33:24.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2650662,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-13T15:56:21.483000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2650399,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-13T12:47:27.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2722838,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-29T19:28:44.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2658769,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-19T11:45:47.990000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2667136,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-24T22:34:38.317000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2662661,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-22T04:00:21.967000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2661175,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-21T04:37:51.620000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2656431,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-17T17:35:05.530000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2656080,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-17T12:46:39.817000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2652136,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-14T15:22:56.703000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2648659,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-12T10:49:46.877000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2647923,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-11T21:49:59.100000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2647434,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-11T14:36:31.583000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2646693,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-11T06:32:08.130000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2643982,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-09T08:07:04.797000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2643681,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-09T03:21:18.210000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2642914,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-08T13:57:59.023000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2642655,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-08T10:23:44.987000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2642001,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-07T20:20:08.573000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2640883,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-07T06:07:18.127000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2640547,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-07T01:20:40.207000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2790045,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-03T01:47:57.447000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2751562,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-14T10:43:34.037000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2732616,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-03T08:21:47.223000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2681462,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-04T16:50:06.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2669737,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-26T13:17:04.447000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2669677,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-26T12:21:39.233000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2664543,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-23T05:32:57.103000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2653627,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-15T15:00:24.647000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2638964": "There are **32** files in `train` csv_files folder. The same files provided in `.parquet` format in the respective folder. So the competition data might be confusing for some of us. Some csvs are greater than **`2 GB`** in size. `train_base.csv` has **1526659** unique `case_id` which equals the length of this file.\n\n##Explanation and schema\nThere are **465** features and **436** respective descriptions in `feature_definitions.csv`. There are no missing descriptions so it means some feature might have same descriptions (for example description `Number of tax deductions` for features: `pmtcount_4527229L`, `pmtcount_4955617L`, `pmtcount_693L`).\n\nWe are going to use `case_id` to merge underlying tables. We got **internal** and **external** data sources and this is how the schema looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F901c94931ce7f953983680c1155f5747%2Fschema.png?generation=1707250963705097&alt=media)\n\nWe got tables classified by depth, where:\n<ul>\n<li>depth=0 - These are static features directly tied to a specific <code>case_id</code>.</li>\n<li>depth=1 - Each <code>case_id</code> has an associated historical record, indexed by <code>num_group1</code>.</li>\n<li>depth=2 - Each <code>case_id</code> has an associated historical record, indexed by both <code>num_group1</code> and <code>num_group2</code>.</li>\n</ul>\n\nVarious predictors were transformed, so to have the following notation for similar groups of transformations: **`P M A D T L`**\n\n## Train File sizes and Null Values\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F0fdebbea69c5090f5fb890de72f23e4e%2Fscaled_nulls_updated.png?generation=1707271157104551&alt=media)\n\n**How to read it**:\n- The sizes of squares are scaled across all `train` files, where the biggest area is the biggest file;\n- Red squares are normalized total counts of all records across all columns in the respective `train` file;\n- Black squares are normalized total counts of **null** records across all columns in the respective `train` file.\n\nFor example, It means the `train` files are  **>30%** empty on average. The `credit_bureau` files are the biggest ones.\n\n#Feature selection\nTODO\n\n##Have Fun!\n\n[Code](https://www.kaggle.com/code/sergiosaharovskiy/home-credit-crms-2024-eda-and-submission) - slowly put flesh on the bone.\n[Repo](https://www.kaggle.com/datasets/sergiosaharovskiy/2024-home-credit-public-repo) - external scripts",
    "2644886": "A few interesting features based on SHAP output. (I'll rerun and colour by \"future week\", which should make clear features that change over time).\n\nCommon sense:\n- Higher Annuity (annuity_780A) -> more likely to default\n- Younger individuals (birth_259D) -> more likely to default\n- recently started employment (empl_employedfrom_271D) -> more likely to default.\n- number of months without payments (cntpmts24_3658933L), higher -> more likely to default\n- Males are more likely to default.\n\nLess obvious\n- More individuals sharing a mobile number (mobilephncnt_593L) -> more likely to default\n- Weekday of decision -> less likely to default if made at the weekend. Odd.\n- Number of loan payments made (pmtnum_254L) -> least likely to default ~12 \n\nOn the y-axis +ve => evidence in favour, -ve => evidence against. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2Fc76f12979967d54fbcbd51f8cfee6476%2Fhc_shap.png?generation=1707502502893240&alt=media)",
    "2652740": "Thanks alott for sharing , its very helpfullll !!!!",
    "2796893": "Thanks Sergey for the great work. What is the package/software that you used for generating the visualization of the NULL values%? Thanks! ",
    "2841951": "Thanks alott for sharing! Well explained Really, helpful to startoff thanks.",
    "2646243": "Well explained ",
    "2645587": "Clear overview! Your detailed insights into file sizes, feature descriptions, and schema provide a solid understanding of the competition data. The visualization aids comprehension. Thanks!",
    "2644329": "Thanks for the awesome visualization, @sergiosaharovskiy. \n\nI have one question though. As nulls are still a part of the training data, shouldn't the black squares be always inside of the red ones? It would be automatic for pie charts, but I'm not sure what determines position of blacks with respect to reds on your infographics.",
    "2745033": "The data is a little confusing, but thanks for the clarification. I wonder what does \"transform\" mean in this context, and why there are numbers before transform type",
    "2796506": "Thanks for leading me to try such competition, especially the visualization (they are too clear to use!!!) and since I am a beginner in the Risk Management Industry.\n",
    "2790341": "Helpful thread.",
    "2775310": "I am a beginner. This is my first time participating in a competition. Thank you for your valuable information.",
    "2674237": "Thanks a lot for sharing the insight on the dataset!",
    "2666293": "The scheme is very useful",
    "2646210": "Really Helpful Solid understanding of the competition data. The visualization comprehension also very good. Thanks.",
    "2644173": "Really, helpful to startoff.\nThank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.\n\nKeep it up and motivate us so that we can also do more better like you ",
    "2641669": "Really helpful!",
    "2640550": "Really Commendable job!! Your visualization work makes the entire structure clear",
    "2638966": "VERY HELPFUL!!!",
    "2649258": "Very nice!",
    "2648782": "well explained",
    "2648201": "Got it so well! Done with my first task.",
    "2646392": "Awesome, the initial sets are a bit overwhelming. Clear and concise!",
    "2645167": "Test path have 4 additional files which are not for train. Why is that?\n\n```python\n 'applprev_1_2.parquet',\n 'credit_bureau_a_1_4.parquet',\n 'credit_bureau_a_2_11.parquet',\n 'static_0_2.parquet'\n```",
    "2644645": "Thanks so much! This is super helpful!\nYour notebook is very educational!",
    "2642750": "Really, helpful to startoff. Thank you!! ",
    "2641717": "Outstanding work to visualize the problem clearly. Thanks!",
    "2639089": "Thank you, Sergey! Your picture will make it much easier for me to remember the entire structure, and I won't need to take as many notes in my notebook.",
    "2816474": "Hello, can I know whether the data for depth equals 1 and equals 2, do I need to group them into like different clusters before joining the data into the train base data or I just have to aggregate all the features together in the train base data. Which one will be a better or a more suitable approach? Thank you 😄😄\n\nBesides, for the transform part is it just how the data is named in the file, because I can't understand what it is related to. Hope someone can clear my doubts. Thank you.",
    "2755643": "greatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreatgreat ",
    "2739664": "Thanks for sharing!! Your schema is very usefull.",
    "2728554": "Really helpful!",
    "2695535": "What a great work! Thank you very much for sharing your work, it helps me a lot to understand the whole dataset!",
    "2694469": "Thanks a lot for your overall explanation, very much appreciated to easily get into the data",
    "2664483": "Very nice!",
    "2659937": "Confused by the features for a while, thanks for the clarification!",
    "2658753": "thanksss, it really gives a favor.4",
    "2656187": "Thank you @sergiosaharovskiy, this will serve as a good starter to understand data. ",
    "2654622": "@sergiosaharovskiy Thank you for sharing your insights, This is super helpful🫡",
    "2654427": "Helpful thread!",
    "2651177": "Amazing visualizations thanks for indepth explanation",
    "2650662": "Helpful thread",
    "2650399": "Code was simple and well implemented.",
    "2722838": "",
    "2658769": "",
    "2667136": "Thanks! 😁",
    "2662661": "Thank you for the explanation!!",
    "2661175": "Thanks so much! This is very helpfulll",
    "2656431": "thanks for sharing",
    "2656080": "Very helpful,  thanks a lot!!!!!",
    "2652136": "Thanks a lot!",
    "2648659": "This is very helpful Thank!",
    "2647923": "Good explanation thanks.",
    "2647434": "Good explanation, thanks",
    "2646693": "It was easy to understand. Thank you.😀",
    "2643982": "thanks! really helpful!",
    "2643681": "Really, helpful thanks for clarifying",
    "2642914": "Really, Helpful sir . Thank you\n",
    "2642655": "Really helpful thanks for clarifying\n",
    "2642001": "Thank you! Incredible work!",
    "2640883": "Thanks so much! This is super helpful!",
    "2640547": "Thank you. Very helpful.",
    "2790045": "Thank you for sharing this!",
    "2751562": "Thank you!",
    "2732616": "Thank you for good information!!",
    "2681462": "Awesome!! Thank you!",
    "2669737": "Really helpful explanation. Thanks!",
    "2669677": "Thanks! very well explained",
    "2664543": "thanks! very nice",
    "2653627": "Very clear, thanks!"
  }
}