{
  "id": 476463,
  "title": "Analysis of birthday",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/476463",
  "author_name": "chumajin",
  "post_date": "2024-02-12T13:32:40.373000",
  "votes": 101,
  "comment_count": 13,
  "views": 0,
  "content": "<p>The data contains multiple descriptions related to birthdays.<br>\nI attempted to compare these birthdays after joining train_static_cb_0.csv and train_person_1.csv(extract person index 0) to train_base.csv</p>\n<p>I observed that the internal data, birth_259D, which has no NaN data and might be accurate in these columns.</p>\n<h2>1. \"birth\" in columns</h2>\n<ul>\n<li><p>train_person_1.csv : Properties: depth=1, <strong>internal data source</strong> (extract person index 0)</p>\n<ul>\n<li>'birth_259D' : Date of birth of the person.</li>\n<li>'birthdate_87D : Birth date of the person.</li></ul></li>\n<li><p>train_static_cb_0.csv : Properties: depth=0, <strong>external data source</strong></p>\n<ul>\n<li>'birthdate_574D' : Client's date of birth (credit bureau data).</li>\n<li>'dateofbirth_337D' : Client's date of birth.</li>\n<li>'dateofbirth_342D' : Client's date of birth.</li></ul></li>\n</ul>\n<h2>2. The number, rate of nan, Match rate with \"birth_259D\"</h2>\n<table>\n<thead>\n<tr>\n<th>columns</th>\n<th>file</th>\n<th>depth</th>\n<th>source</th>\n<th>number of nan</th>\n<th>nan rate</th>\n<th>Match Rate with birth_259D</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>birth_259D</td>\n<td>train_person_1.csv</td>\n<td>1</td>\n<td>internal</td>\n<td>0</td>\n<td>0.0%</td>\n<td>100.0000%</td>\n</tr>\n<tr>\n<td>birthdate_87D</td>\n<td>train_person_1.csv</td>\n<td>1</td>\n<td>internal</td>\n<td>1553993</td>\n<td>99.2%</td>\n<td>100.0000%</td>\n</tr>\n<tr>\n<td>birthdate_574D</td>\n<td>train_static_cb_0.csv</td>\n<td>0</td>\n<td>external</td>\n<td>942362</td>\n<td>60.2%</td>\n<td>99.9992%</td>\n</tr>\n<tr>\n<td>dateofbirth_337D</td>\n<td>train_static_cb_0.csv</td>\n<td>0</td>\n<td>external</td>\n<td>143084</td>\n<td>9.1%</td>\n<td>99.8521%</td>\n</tr>\n<tr>\n<td>dateofbirth_342D</td>\n<td>train_static_cb_0.csv</td>\n<td>0</td>\n<td>external</td>\n<td>1529328</td>\n<td>97.6%</td>\n<td>99.8070%</td>\n</tr>\n<tr>\n<td>min</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0</td>\n<td>0.0%</td>\n<td>99.9293%</td>\n</tr>\n<tr>\n<td>max</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0</td>\n<td>0.0%</td>\n<td>99.9339%</td>\n</tr>\n</tbody>\n</table>\n<p>※ min, max is calculated, df.min(axis=1) like that.<br>\n※ Match Rate with birth_259D is calculated without nan data.</p>\n<ul>\n<li>There are no NaN data in birth_259D.</li>\n<li>Excluding NaN data, birth_259D and birthdate_87D are exactly the same.</li>\n<li>The match rate rank for birth259D is as follows.<br>\nbirth_259D = birthdate_87D &gt; birthdate_574D &gt; dateofbirth_337D &gt; dateofbirth_342D</li>\n</ul>\n<h2>3. Visualization of Differences in Min, Max Values</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F9440bc0653174d37926b6e8460b50ec9%2FClipboard01.jpg?generation=1707709919603323&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F598c58f0a2fdb2375483c33f35279445%2FClipboard02.jpg?generation=1707710216037407&amp;alt=media\"></p>\n<p>※ Intuitively, I drew red lines where the mistakes seemed to be significant.</p>\n<ul>\n<li>It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.</li>\n<li>dateofbirth_337D often deviates by about a year or a month from birthdate_574D or birthdate_259D.</li>\n<li>dateofbirth_337D occasionally has significant errors to birth_259D. It might have been misread from handwritten entries.</li>\n<li>dateofbirth_342D also has cases where it differs from birth_259D, with a notable NaN rate of 97.6%.</li>\n</ul>\n<h2>4. Summary</h2>\n<ul>\n<li>It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.</li>\n<li>Comparing with the mode might provide more insight, but I stopped because the calculation takes a long time. <br>\nPerhaps it might be okay to drop columns other than birth259D.</li>\n<li>Furthermore, although it might be slight, there is a non-zero possibility that dateofbirth_337D or others could be correct and birth_259D might be incorrect, so caution is needed in this aspect as well.</li>\n</ul>",
  "messages": [
    {
      "id": 2648878,
      "postDate": "2024-02-12T13:32:40.373Z",
      "content": "<p>The data contains multiple descriptions related to birthdays.<br>\nI attempted to compare these birthdays after joining train_static_cb_0.csv and train_person_1.csv(extract person index 0) to train_base.csv</p>\n<p>I observed that the internal data, birth_259D, which has no NaN data and might be accurate in these columns.</p>\n<h2>1. \"birth\" in columns</h2>\n<ul>\n<li><p>train_person_1.csv : Properties: depth=1, <strong>internal data source</strong> (extract person index 0)</p>\n<ul>\n<li>'birth_259D' : Date of birth of the person.</li>\n<li>'birthdate_87D : Birth date of the person.</li></ul></li>\n<li><p>train_static_cb_0.csv : Properties: depth=0, <strong>external data source</strong></p>\n<ul>\n<li>'birthdate_574D' : Client's date of birth (credit bureau data).</li>\n<li>'dateofbirth_337D' : Client's date of birth.</li>\n<li>'dateofbirth_342D' : Client's date of birth.</li></ul></li>\n</ul>\n<h2>2. The number, rate of nan, Match rate with \"birth_259D\"</h2>\n<table>\n<thead>\n<tr>\n<th>columns</th>\n<th>file</th>\n<th>depth</th>\n<th>source</th>\n<th>number of nan</th>\n<th>nan rate</th>\n<th>Match Rate with birth_259D</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>birth_259D</td>\n<td>train_person_1.csv</td>\n<td>1</td>\n<td>internal</td>\n<td>0</td>\n<td>0.0%</td>\n<td>100.0000%</td>\n</tr>\n<tr>\n<td>birthdate_87D</td>\n<td>train_person_1.csv</td>\n<td>1</td>\n<td>internal</td>\n<td>1553993</td>\n<td>99.2%</td>\n<td>100.0000%</td>\n</tr>\n<tr>\n<td>birthdate_574D</td>\n<td>train_static_cb_0.csv</td>\n<td>0</td>\n<td>external</td>\n<td>942362</td>\n<td>60.2%</td>\n<td>99.9992%</td>\n</tr>\n<tr>\n<td>dateofbirth_337D</td>\n<td>train_static_cb_0.csv</td>\n<td>0</td>\n<td>external</td>\n<td>143084</td>\n<td>9.1%</td>\n<td>99.8521%</td>\n</tr>\n<tr>\n<td>dateofbirth_342D</td>\n<td>train_static_cb_0.csv</td>\n<td>0</td>\n<td>external</td>\n<td>1529328</td>\n<td>97.6%</td>\n<td>99.8070%</td>\n</tr>\n<tr>\n<td>min</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0</td>\n<td>0.0%</td>\n<td>99.9293%</td>\n</tr>\n<tr>\n<td>max</td>\n<td></td>\n<td></td>\n<td></td>\n<td>0</td>\n<td>0.0%</td>\n<td>99.9339%</td>\n</tr>\n</tbody>\n</table>\n<p>※ min, max is calculated, df.min(axis=1) like that.<br>\n※ Match Rate with birth_259D is calculated without nan data.</p>\n<ul>\n<li>There are no NaN data in birth_259D.</li>\n<li>Excluding NaN data, birth_259D and birthdate_87D are exactly the same.</li>\n<li>The match rate rank for birth259D is as follows.<br>\nbirth_259D = birthdate_87D &gt; birthdate_574D &gt; dateofbirth_337D &gt; dateofbirth_342D</li>\n</ul>\n<h2>3. Visualization of Differences in Min, Max Values</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F9440bc0653174d37926b6e8460b50ec9%2FClipboard01.jpg?generation=1707709919603323&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F598c58f0a2fdb2375483c33f35279445%2FClipboard02.jpg?generation=1707710216037407&amp;alt=media\"></p>\n<p>※ Intuitively, I drew red lines where the mistakes seemed to be significant.</p>\n<ul>\n<li>It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.</li>\n<li>dateofbirth_337D often deviates by about a year or a month from birthdate_574D or birthdate_259D.</li>\n<li>dateofbirth_337D occasionally has significant errors to birth_259D. It might have been misread from handwritten entries.</li>\n<li>dateofbirth_342D also has cases where it differs from birth_259D, with a notable NaN rate of 97.6%.</li>\n</ul>\n<h2>4. Summary</h2>\n<ul>\n<li>It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.</li>\n<li>Comparing with the mode might provide more insight, but I stopped because the calculation takes a long time. <br>\nPerhaps it might be okay to drop columns other than birth259D.</li>\n<li>Furthermore, although it might be slight, there is a non-zero possibility that dateofbirth_337D or others could be correct and birth_259D might be incorrect, so caution is needed in this aspect as well.</li>\n</ul>",
      "rawMarkdown": "The data contains multiple descriptions related to birthdays.\nI attempted to compare these birthdays after joining train_static_cb_0.csv and train_person_1.csv(extract person index 0) to train_base.csv\n\nI observed that the internal data, birth_259D, which has no NaN data and might be accurate in these columns.\n\n\n## 1. \"birth\" in columns\n\n* train_person_1.csv : Properties: depth=1, **internal data source** (extract person index 0)\n    -  'birth_259D' : Date of birth of the person.\n    -  'birthdate_87D : Birth date of the person.\n\n* train_static_cb_0.csv : Properties: depth=0, **external data source**\n    - 'birthdate_574D' : Client's date of birth (credit bureau data).\n    - 'dateofbirth_337D' : Client's date of birth.\n    - 'dateofbirth_342D' : Client's date of birth.\n\n## 2. The number, rate of nan, Match rate with \"birth_259D\"\n| columns          | file                  | depth | source   | number of nan | nan rate | Match Rate with birth_259D |\n|------------------|-----------------------|-------|----------|---------------|----------|----------------------------|\n| birth_259D       | train_person_1.csv    | 1     | internal | 0             | 0.0%     | 100.0000%                  |\n| birthdate_87D    | train_person_1.csv    | 1     | internal | 1553993       | 99.2%    | 100.0000%                  |\n| birthdate_574D   | train_static_cb_0.csv | 0     | external | 942362        | 60.2%    | 99.9992%                   |\n| dateofbirth_337D | train_static_cb_0.csv | 0     | external | 143084        | 9.1%     | 99.8521%                   |\n| dateofbirth_342D | train_static_cb_0.csv | 0     | external | 1529328       | 97.6%    | 99.8070%                   |\n| min              |                       |       |          | 0             | 0.0%     | 99.9293%                   |\n| max              |                       |       |          | 0             | 0.0%     | 99.9339%                   |\n\n\n※ min, max is calculated, df.min(axis=1) like that.\n※ Match Rate with birth_259D is calculated without nan data.\n\n* There are no NaN data in birth_259D.\n* Excluding NaN data, birth_259D and birthdate_87D are exactly the same.\n* The match rate rank for birth259D is as follows.\nbirth_259D = birthdate_87D > birthdate_574D > dateofbirth_337D > dateofbirth_342D\n\n## 3. Visualization of Differences in Min, Max Values\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F9440bc0653174d37926b6e8460b50ec9%2FClipboard01.jpg?generation=1707709919603323&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F598c58f0a2fdb2375483c33f35279445%2FClipboard02.jpg?generation=1707710216037407&alt=media)\n\n※ Intuitively, I drew red lines where the mistakes seemed to be significant.\n\n* It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.\n* dateofbirth_337D often deviates by about a year or a month from birthdate_574D or birthdate_259D.\n* dateofbirth_337D occasionally has significant errors to birth_259D. It might have been misread from handwritten entries.\n* dateofbirth_342D also has cases where it differs from birth_259D, with a notable NaN rate of 97.6%.\n\n## 4. Summary\n\n* It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.\n* Comparing with the mode might provide more insight, but I stopped because the calculation takes a long time. \nPerhaps it might be okay to drop columns other than birth259D.\n* Furthermore, although it might be slight, there is a non-zero possibility that dateofbirth_337D or others could be correct and birth_259D might be incorrect, so caution is needed in this aspect as well.",
      "votes": 101
    },
    {
      "id": 2652777,
      "postDate": "2024-02-15T02:23:20.720Z",
      "content": "<p>Thanks for sharing , i learn a lott !!!!</p>",
      "rawMarkdown": "Thanks for sharing , i learn a lott !!!!",
      "votes": 4
    },
    {
      "id": 2715003,
      "postDate": "2024-03-25T08:06:39.820Z",
      "content": "<p>Great work.<br>\nThat's interesting.</p>",
      "rawMarkdown": "Great work.\nThat's interesting.",
      "votes": 1
    },
    {
      "id": 2677431,
      "postDate": "2024-03-02T07:29:47.137Z",
      "content": "<p>Brilliant work.</p>",
      "rawMarkdown": "Brilliant work.",
      "votes": 1
    },
    {
      "id": 2655705,
      "postDate": "2024-02-17T06:48:27.347Z",
      "content": "<p>Thanks for sharing this it seems to be very helpful <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "rawMarkdown": "Thanks for sharing this it seems to be very helpful @chumajin ",
      "votes": 1
    },
    {
      "id": 2654610,
      "postDate": "2024-02-16T09:29:41.953Z",
      "content": "<p>Your insights have been very helpful. Thank you:)</p>",
      "rawMarkdown": "Your insights have been very helpful. Thank you:)",
      "votes": 1
    },
    {
      "id": 2649927,
      "postDate": "2024-02-13T05:55:21.987Z",
      "content": "<p>ty for your work <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "rawMarkdown": "ty for your work @chumajin ",
      "votes": 1
    },
    {
      "id": 2653054,
      "postDate": "2024-02-15T06:46:16.770Z",
      "content": "<p>I see two possible answers:<br>\nA. External source got it wrong somehow<br>\nB. The applicant was lying - showing yourself as younger, intuitively increases your credit standing for older people. But this is very easily verifiable and from your analysis sometimes mistakes are in the other direction.</p>\n<p>IMO in case of B., it could be a signal to the model, but needs to be verified.</p>",
      "rawMarkdown": "I see two possible answers:\nA. External source got it wrong somehow\nB. The applicant was lying - showing yourself as younger, intuitively increases your credit standing for older people. But this is very easily verifiable and from your analysis sometimes mistakes are in the other direction.\n\nIMO in case of B., it could be a signal to the model, but needs to be verified.",
      "votes": 2,
      "replies": [
        {
          "id": 2653524,
          "postDate": "2024-02-15T13:49:02.807Z",
          "content": "<p>Deep insight! Thank you for your opinion!</p>",
          "rawMarkdown": "Deep insight! Thank you for your opinion!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2649088,
      "postDate": "2024-02-12T15:53:12.850Z",
      "content": "<p>Nice work!</p>",
      "rawMarkdown": "Nice work!",
      "votes": 2
    },
    {
      "id": 2796528,
      "postDate": "2024-05-06T09:34:17.493Z",
      "content": "<p>So maybe 259D is the most up-to-date version of birthday information? Wonder if it could be any kind of signal given such high matching rate, or maybe the nan rate itself maybe informing. Worth exploiting, thanks~</p>",
      "rawMarkdown": "So maybe 259D is the most up-to-date version of birthday information? Wonder if it could be any kind of signal given such high matching rate, or maybe the nan rate itself maybe informing. Worth exploiting, thanks~"
    },
    {
      "id": 2824188,
      "postDate": "2024-05-19T16:17:48.777Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2656920,
      "postDate": "2024-02-18T06:30:39.580Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2654312,
      "postDate": "2024-02-16T04:12:28.237Z",
      "content": "<p>Thanks a lot! 😄</p>",
      "rawMarkdown": "Thanks a lot! :smile:",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2652777,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-15T02:23:20.720000",
      "content": "<p>Thanks for sharing , i learn a lott !!!!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2715003,
      "author_name": "shun takinami",
      "author_url": "",
      "post_date": "2024-03-25T08:06:39.820000",
      "content": "<p>Great work.<br>\nThat's interesting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2677431,
      "author_name": "Aniruddha Pal",
      "author_url": "",
      "post_date": "2024-03-02T07:29:47.137000",
      "content": "<p>Brilliant work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2655705,
      "author_name": "Tanishq dublish",
      "author_url": "",
      "post_date": "2024-02-17T06:48:27.347000",
      "content": "<p>Thanks for sharing this it seems to be very helpful <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2654610,
      "author_name": "Why_Be",
      "author_url": "",
      "post_date": "2024-02-16T09:29:41.953000",
      "content": "<p>Your insights have been very helpful. Thank you:)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2649927,
      "author_name": "Anil",
      "author_url": "",
      "post_date": "2024-02-13T05:55:21.987000",
      "content": "<p>ty for your work <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2653054,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2024-02-15T06:46:16.770000",
      "content": "<p>I see two possible answers:<br>\nA. External source got it wrong somehow<br>\nB. The applicant was lying - showing yourself as younger, intuitively increases your credit standing for older people. But this is very easily verifiable and from your analysis sometimes mistakes are in the other direction.</p>\n<p>IMO in case of B., it could be a signal to the model, but needs to be verified.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2653524,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-02-15T13:49:02.807000",
          "content": "<p>Deep insight! Thank you for your opinion!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2649088,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-02-12T15:53:12.850000",
      "content": "<p>Nice work!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2796528,
      "author_name": "TkiJoestar",
      "author_url": "",
      "post_date": "2024-05-06T09:34:17.493000",
      "content": "<p>So maybe 259D is the most up-to-date version of birthday information? Wonder if it could be any kind of signal given such high matching rate, or maybe the nan rate itself maybe informing. Worth exploiting, thanks~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2824188,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-19T16:17:48.777000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2656920,
      "author_name": "Joser",
      "author_url": "",
      "post_date": "2024-02-18T06:30:39.580000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2654312,
      "author_name": "Analyticity",
      "author_url": "",
      "post_date": "2024-02-16T04:12:28.237000",
      "content": "<p>Thanks a lot! 😄</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2648878": "The data contains multiple descriptions related to birthdays.\nI attempted to compare these birthdays after joining train_static_cb_0.csv and train_person_1.csv(extract person index 0) to train_base.csv\n\nI observed that the internal data, birth_259D, which has no NaN data and might be accurate in these columns.\n\n\n## 1. \"birth\" in columns\n\n* train_person_1.csv : Properties: depth=1, **internal data source** (extract person index 0)\n    -  'birth_259D' : Date of birth of the person.\n    -  'birthdate_87D : Birth date of the person.\n\n* train_static_cb_0.csv : Properties: depth=0, **external data source**\n    - 'birthdate_574D' : Client's date of birth (credit bureau data).\n    - 'dateofbirth_337D' : Client's date of birth.\n    - 'dateofbirth_342D' : Client's date of birth.\n\n## 2. The number, rate of nan, Match rate with \"birth_259D\"\n| columns          | file                  | depth | source   | number of nan | nan rate | Match Rate with birth_259D |\n|------------------|-----------------------|-------|----------|---------------|----------|----------------------------|\n| birth_259D       | train_person_1.csv    | 1     | internal | 0             | 0.0%     | 100.0000%                  |\n| birthdate_87D    | train_person_1.csv    | 1     | internal | 1553993       | 99.2%    | 100.0000%                  |\n| birthdate_574D   | train_static_cb_0.csv | 0     | external | 942362        | 60.2%    | 99.9992%                   |\n| dateofbirth_337D | train_static_cb_0.csv | 0     | external | 143084        | 9.1%     | 99.8521%                   |\n| dateofbirth_342D | train_static_cb_0.csv | 0     | external | 1529328       | 97.6%    | 99.8070%                   |\n| min              |                       |       |          | 0             | 0.0%     | 99.9293%                   |\n| max              |                       |       |          | 0             | 0.0%     | 99.9339%                   |\n\n\n※ min, max is calculated, df.min(axis=1) like that.\n※ Match Rate with birth_259D is calculated without nan data.\n\n* There are no NaN data in birth_259D.\n* Excluding NaN data, birth_259D and birthdate_87D are exactly the same.\n* The match rate rank for birth259D is as follows.\nbirth_259D = birthdate_87D > birthdate_574D > dateofbirth_337D > dateofbirth_342D\n\n## 3. Visualization of Differences in Min, Max Values\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F9440bc0653174d37926b6e8460b50ec9%2FClipboard01.jpg?generation=1707709919603323&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F598c58f0a2fdb2375483c33f35279445%2FClipboard02.jpg?generation=1707710216037407&alt=media)\n\n※ Intuitively, I drew red lines where the mistakes seemed to be significant.\n\n* It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.\n* dateofbirth_337D often deviates by about a year or a month from birthdate_574D or birthdate_259D.\n* dateofbirth_337D occasionally has significant errors to birth_259D. It might have been misread from handwritten entries.\n* dateofbirth_342D also has cases where it differs from birth_259D, with a notable NaN rate of 97.6%.\n\n## 4. Summary\n\n* It has been confirmed that the internal data, birth_259D, and the external data, birthdate_574D, show a high match rate. Therefore, birth_259D, which has no NaN data, is likely to be of high accuracy.\n* Comparing with the mode might provide more insight, but I stopped because the calculation takes a long time. \nPerhaps it might be okay to drop columns other than birth259D.\n* Furthermore, although it might be slight, there is a non-zero possibility that dateofbirth_337D or others could be correct and birth_259D might be incorrect, so caution is needed in this aspect as well.",
    "2652777": "Thanks for sharing , i learn a lott !!!!",
    "2715003": "Great work.\nThat's interesting.",
    "2677431": "Brilliant work.",
    "2655705": "Thanks for sharing this it seems to be very helpful @chumajin ",
    "2654610": "Your insights have been very helpful. Thank you:)",
    "2649927": "ty for your work @chumajin ",
    "2653054": "I see two possible answers:\nA. External source got it wrong somehow\nB. The applicant was lying - showing yourself as younger, intuitively increases your credit standing for older people. But this is very easily verifiable and from your analysis sometimes mistakes are in the other direction.\n\nIMO in case of B., it could be a signal to the model, but needs to be verified.",
    "2649088": "Nice work!",
    "2796528": "So maybe 259D is the most up-to-date version of birthday information? Wonder if it could be any kind of signal given such high matching rate, or maybe the nan rate itself maybe informing. Worth exploiting, thanks~",
    "2824188": "",
    "2656920": "Thanks for sharing!",
    "2654312": "Thanks a lot! :smile:"
  }
}