{
  "id": 475773,
  "title": "Meaning of a55475b1",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/475773",
  "author_name": "",
  "post_date": "2024-02-09T18:46:00.789313200Z",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi,<br>\nwhat's the meaning of value a55475b1?</p>",
  "messages": [
    {
      "id": "2644894",
      "postDate": "02/09/2024 18:46:00",
      "content": "<p>Hi,<br>\nwhat's the meaning of value a55475b1?</p>",
      "rawMarkdown": "Hi,\nwhat's the meaning of value a55475b1?",
      "votes": null
    },
    {
      "id": "2645114",
      "postDate": "02/10/2024 01:08:51",
      "content": "<p>yes! and that of P94_109_143 etc etc the whole cancelreason_3545846M column basically. And other column as well, and I've only looked into the very first document I looked at train_applprev_1_0.csv so I guess this is like that throughout the whole training set without explanation about what they mean? Or am I missing something? (I hope I do)</p>",
      "rawMarkdown": "yes! and that of P94_109_143 etc etc the whole cancelreason_3545846M column basically. And other column as well, and I've only looked into the very first document I looked at train_applprev_1_0.csv so I guess this is like that throughout the whole training set without explanation about what they mean? Or am I missing something? (I hope I do)",
      "votes": null
    },
    {
      "id": "2647871",
      "postDate": "02/11/2024 20:31:59",
      "content": "<p>We share your confusion. A trouble shared is a trouble halved. But this does not do us much good…. 😀</p>",
      "rawMarkdown": "We share your confusion. A trouble shared is a trouble halved. But this does not do us much good.... 😀",
      "votes": null
    },
    {
      "id": "2648460",
      "postDate": "02/12/2024 09:09:29",
      "content": "<p>That is masked category and it is masked on purpose :) </p>",
      "rawMarkdown": "That is masked category and it is masked on purpose :)",
      "votes": null
    },
    {
      "id": "2648717",
      "postDate": "02/12/2024 11:25:47",
      "content": "<p>Hi,<br>\nRegarding cancelreason_3545846M, from feature name and feature definition it's clear that this column contains reason, why given application was cancelled. </p>\n<p>Due to business reason, we can't disclose exact reason - we can't and don't want to provide full details about underwriting process and rules, this is out internal know how (many of those details are not even shared outside risk departments). Therefore, data is masked…</p>\n<p>However, from masked data you can see how many cancellation reasons we have and test whether some cancellation reason category has predictive power or not.</p>",
      "rawMarkdown": "Hi,\nRegarding cancelreason_3545846M, from feature name and feature definition it's clear that this column contains reason, why given application was cancelled. \n\nDue to business reason, we can't disclose exact reason - we can't and don't want to provide full details about underwriting process and rules, this is out internal know how (many of those details are not even shared outside risk departments). Therefore, data is masked...\n\nHowever, from masked data you can see how many cancellation reasons we have and test whether some cancellation reason category has predictive power or not.",
      "votes": null
    },
    {
      "id": "2649399",
      "postDate": "02/12/2024 19:53:45",
      "content": "<p>Hi Daniel and Thomas,</p>\n<p>Thank you for your swift explanation.</p>\n<p>Can you confirm:</p>\n<ul>\n<li>if a55475b1 is found multiple times in one and the same column, it is a masked version of the same original value in this column,</li>\n<li>if a55475b1 is found in another column, it may be the masked version of a different value in this other colums.</li>\n</ul>\n<p>Example:<br>\nin column x, a55475b1 may be the masked version of \"cancel code 1\" while in column y, a55475b1 is the masked version of \"university level education\".</p>\n<p>If this assumption is correct, then imho this is not the most obvious way to mask data and a note in the competition intro is justified (and if it is not correct, then even more so).</p>\n<p>KR, Ruud. </p>",
      "rawMarkdown": "Hi Daniel and Thomas,\n\nThank you for your swift explanation.\n\nCan you confirm:\n- if a55475b1 is found multiple times in one and the same column, it is a masked version of the same original value in this column,\n- if a55475b1 is found in another column, it may be the masked version of a different value in this other colums.\n\nExample:\nin column x, a55475b1 may be the masked version of \"cancel code 1\" while in column y, a55475b1 is the masked version of \"university level education\".\n\nIf this assumption is correct, then imho this is not the most obvious way to mask data and a note in the competition intro is justified (and if it is not correct, then even more so).\n\nKR, Ruud.",
      "votes": null
    },
    {
      "id": "2649553",
      "postDate": "02/13/2024 00:37:34",
      "content": "<p>Given the frequency and the correlation with <code>null</code> values in other numeric columns, I think it <code>a55475b1</code> likely just means <code>null</code>.</p>",
      "rawMarkdown": "Given the frequency and the correlation with `null` values in other numeric columns, I think it `a55475b1` likely just means `null`.",
      "votes": null
    },
    {
      "id": "2650066",
      "postDate": "02/13/2024 08:22:28",
      "content": "<p>Hi Ruud,<br>\nsame masked value = same original value. However if you have same values in different columns like cancellation reason and education, it most probably mean that original values have been already encoded (e.g., \"category 1\", \"category 2\",…) =&gt; they look same, but meaning is different…</p>",
      "rawMarkdown": "Hi Ruud,\nsame masked value = same original value. However if you have same values in different columns like cancellation reason and education, it most probably mean that original values have been already encoded (e.g., \"category 1\", \"category 2\",...) => they look same, but meaning is different...",
      "votes": null
    },
    {
      "id": "2657131",
      "postDate": "02/18/2024 10:29:58",
      "content": "<p>Thx Thomas. That explains well. Kind regards, Ruud </p>",
      "rawMarkdown": "Thx Thomas. That explains well. Kind regards, Ruud",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2645114,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "02/10/2024 01:08:51",
      "content": "<p>yes! and that of P94_109_143 etc etc the whole cancelreason_3545846M column basically. And other column as well, and I've only looked into the very first document I looked at train_applprev_1_0.csv so I guess this is like that throughout the whole training set without explanation about what they mean? Or am I missing something? (I hope I do)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2648717,
          "author_name": "tomasjeline2",
          "author_url": "",
          "post_date": "02/12/2024 11:25:47",
          "content": "<p>Hi,<br>\nRegarding cancelreason_3545846M, from feature name and feature definition it's clear that this column contains reason, why given application was cancelled. </p>\n<p>Due to business reason, we can't disclose exact reason - we can't and don't want to provide full details about underwriting process and rules, this is out internal know how (many of those details are not even shared outside risk departments). Therefore, data is masked…</p>\n<p>However, from masked data you can see how many cancellation reasons we have and test whether some cancellation reason category has predictive power or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2647871,
      "author_name": "ruudkapteijn",
      "author_url": "",
      "post_date": "02/11/2024 20:31:59",
      "content": "<p>We share your confusion. A trouble shared is a trouble halved. But this does not do us much good…. 😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2648460,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/12/2024 09:09:29",
      "content": "<p>That is masked category and it is masked on purpose :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2649399,
      "author_name": "ruudkapteijn",
      "author_url": "",
      "post_date": "02/12/2024 19:53:45",
      "content": "<p>Hi Daniel and Thomas,</p>\n<p>Thank you for your swift explanation.</p>\n<p>Can you confirm:</p>\n<ul>\n<li>if a55475b1 is found multiple times in one and the same column, it is a masked version of the same original value in this column,</li>\n<li>if a55475b1 is found in another column, it may be the masked version of a different value in this other colums.</li>\n</ul>\n<p>Example:<br>\nin column x, a55475b1 may be the masked version of \"cancel code 1\" while in column y, a55475b1 is the masked version of \"university level education\".</p>\n<p>If this assumption is correct, then imho this is not the most obvious way to mask data and a note in the competition intro is justified (and if it is not correct, then even more so).</p>\n<p>KR, Ruud. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2650066,
          "author_name": "tomasjeline2",
          "author_url": "",
          "post_date": "02/13/2024 08:22:28",
          "content": "<p>Hi Ruud,<br>\nsame masked value = same original value. However if you have same values in different columns like cancellation reason and education, it most probably mean that original values have been already encoded (e.g., \"category 1\", \"category 2\",…) =&gt; they look same, but meaning is different…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2657131,
              "author_name": "ruudkapteijn",
              "author_url": "",
              "post_date": "02/18/2024 10:29:58",
              "content": "<p>Thx Thomas. That explains well. Kind regards, Ruud </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2649553,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "02/13/2024 00:37:34",
      "content": "<p>Given the frequency and the correlation with <code>null</code> values in other numeric columns, I think it <code>a55475b1</code> likely just means <code>null</code>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2644894": "Hi,\nwhat's the meaning of value a55475b1?",
    "2645114": "yes! and that of P94_109_143 etc etc the whole cancelreason_3545846M column basically. And other column as well, and I've only looked into the very first document I looked at train_applprev_1_0.csv so I guess this is like that throughout the whole training set without explanation about what they mean? Or am I missing something? (I hope I do)",
    "2647871": "We share your confusion. A trouble shared is a trouble halved. But this does not do us much good.... 😀",
    "2648460": "That is masked category and it is masked on purpose :)",
    "2648717": "Hi,\nRegarding cancelreason_3545846M, from feature name and feature definition it's clear that this column contains reason, why given application was cancelled. \n\nDue to business reason, we can't disclose exact reason - we can't and don't want to provide full details about underwriting process and rules, this is out internal know how (many of those details are not even shared outside risk departments). Therefore, data is masked...\n\nHowever, from masked data you can see how many cancellation reasons we have and test whether some cancellation reason category has predictive power or not.",
    "2649399": "Hi Daniel and Thomas,\n\nThank you for your swift explanation.\n\nCan you confirm:\n- if a55475b1 is found multiple times in one and the same column, it is a masked version of the same original value in this column,\n- if a55475b1 is found in another column, it may be the masked version of a different value in this other colums.\n\nExample:\nin column x, a55475b1 may be the masked version of \"cancel code 1\" while in column y, a55475b1 is the masked version of \"university level education\".\n\nIf this assumption is correct, then imho this is not the most obvious way to mask data and a note in the competition intro is justified (and if it is not correct, then even more so).\n\nKR, Ruud.",
    "2649553": "Given the frequency and the correlation with `null` values in other numeric columns, I think it `a55475b1` likely just means `null`.",
    "2650066": "Hi Ruud,\nsame masked value = same original value. However if you have same values in different columns like cancellation reason and education, it most probably mean that original values have been already encoded (e.g., \"category 1\", \"category 2\",...) => they look same, but meaning is different...",
    "2657131": "Thx Thomas. That explains well. Kind regards, Ruud"
  },
  "source": "meta"
}