{
  "id": 169963,
  "title": "Imputing patient_id",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169963",
  "author_name": "",
  "post_date": "2020-07-25T21:44:17.781870200Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Has anyone done imputation of <code>patient_id</code> where you assign a random unused value <code>IP_XXXXXX</code>?  If someone is doing this and can share, I think that will be my preferred way of handling the missing patient ID's.  If I get something working I will post it here as well.</p>\n\n<p>The reason it is important is for doing KFolds CV.  The patient_id is a \"group\", and you don't want to be training on say <code>patient_id</code> 1234 and at the same time validating on <code>patient_id</code> 1234, because that is a serious leak.  If you take a large number of samples and give them all the SAME <code>patient_id</code> (0, 1234, etc), that seriously limits your CV because if you have a constraint like I have, where you force all members of a \"group\" (same <code>patient_id</code>) to be in the same Fold, then you aren't getting the best distribution you could hope for.</p>",
  "messages": [
    {
      "id": "945492",
      "postDate": "07/25/2020 21:44:17",
      "content": "<p>Has anyone done imputation of <code>patient_id</code> where you assign a random unused value <code>IP_XXXXXX</code>?  If someone is doing this and can share, I think that will be my preferred way of handling the missing patient ID's.  If I get something working I will post it here as well.</p>\n\n<p>The reason it is important is for doing KFolds CV.  The patient_id is a \"group\", and you don't want to be training on say <code>patient_id</code> 1234 and at the same time validating on <code>patient_id</code> 1234, because that is a serious leak.  If you take a large number of samples and give them all the SAME <code>patient_id</code> (0, 1234, etc), that seriously limits your CV because if you have a constraint like I have, where you force all members of a \"group\" (same <code>patient_id</code>) to be in the same Fold, then you aren't getting the best distribution you could hope for.</p>",
      "rawMarkdown": "Has anyone done imputation of `patient_id` where you assign a random unused value `IP_XXXXXX`?  If someone is doing this and can share, I think that will be my preferred way of handling the missing patient ID's.  If I get something working I will post it here as well.\n\nThe reason it is important is for doing KFolds CV.  The patient_id is a \"group\", and you don't want to be training on say `patient_id` 1234 and at the same time validating on `patient_id` 1234, because that is a serious leak.  If you take a large number of samples and give them all the SAME `patient_id` (0, 1234, etc), that seriously limits your CV because if you have a constraint like I have, where you force all members of a \"group\" (same `patient_id`) to be in the same Fold, then you aren't getting the best distribution you could hope for.",
      "votes": null
    },
    {
      "id": "945558",
      "postDate": "07/26/2020 00:26:27",
      "content": "<p>Here is what I ended up doing, which is quite messy but works:</p>\n\n<p>```\ntrain_df['patient_id'] = train_df['patient_id'].replace(-1,np.NaN, regex=True)</p>\n\n<p>pid_list = train_df['patient_id'].tolist()\nall_list = []\nunused_list = []\nreplace_list = []</p>\n\n<p>for index in range(len(pid_list)):\n    if isinstance(pid_list[index], str):\n        pid_list[index] = int(pid_list[index][3:])\npid_list=list(set(pid_list))\nfor index in range(1,10000000):\n    all_list.append(index)\nunused_list = list(set(all_list) - set(pid_list))\nreplace_list = random.sample(unused_list, len(train_df))\nfor index in range(len(replace_list)):\n    replace_list[index] = \"IP_\" + str(replace_list[index])</p>\n\n<p>train_df['patient_id'] = train_df['patient_id'].fillna(pd.Series(replace_list))\n```</p>",
      "rawMarkdown": "Here is what I ended up doing, which is quite messy but works:\n\n```\ntrain_df['patient_id'] = train_df['patient_id'].replace(-1,np.NaN, regex=True)\n\npid_list = train_df['patient_id'].tolist()\nall_list = []\nunused_list = []\nreplace_list = []\n\nfor index in range(len(pid_list)):\n    if isinstance(pid_list[index], str):\n        pid_list[index] = int(pid_list[index][3:])\npid_list=list(set(pid_list))\nfor index in range(1,10000000):\n    all_list.append(index)\nunused_list = list(set(all_list) - set(pid_list))\nreplace_list = random.sample(unused_list, len(train_df))\nfor index in range(len(replace_list)):\n    replace_list[index] = \"IP_\" + str(replace_list[index])\n    \ntrain_df['patient_id'] = train_df['patient_id'].fillna(pd.Series(replace_list))\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 945558,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "07/26/2020 00:26:27",
      "content": "<p>Here is what I ended up doing, which is quite messy but works:</p>\n\n<p>```\ntrain_df['patient_id'] = train_df['patient_id'].replace(-1,np.NaN, regex=True)</p>\n\n<p>pid_list = train_df['patient_id'].tolist()\nall_list = []\nunused_list = []\nreplace_list = []</p>\n\n<p>for index in range(len(pid_list)):\n    if isinstance(pid_list[index], str):\n        pid_list[index] = int(pid_list[index][3:])\npid_list=list(set(pid_list))\nfor index in range(1,10000000):\n    all_list.append(index)\nunused_list = list(set(all_list) - set(pid_list))\nreplace_list = random.sample(unused_list, len(train_df))\nfor index in range(len(replace_list)):\n    replace_list[index] = \"IP_\" + str(replace_list[index])</p>\n\n<p>train_df['patient_id'] = train_df['patient_id'].fillna(pd.Series(replace_list))\n```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "945492": "Has anyone done imputation of `patient_id` where you assign a random unused value `IP_XXXXXX`?  If someone is doing this and can share, I think that will be my preferred way of handling the missing patient ID's.  If I get something working I will post it here as well.\n\nThe reason it is important is for doing KFolds CV.  The patient_id is a \"group\", and you don't want to be training on say `patient_id` 1234 and at the same time validating on `patient_id` 1234, because that is a serious leak.  If you take a large number of samples and give them all the SAME `patient_id` (0, 1234, etc), that seriously limits your CV because if you have a constraint like I have, where you force all members of a \"group\" (same `patient_id`) to be in the same Fold, then you aren't getting the best distribution you could hope for.",
    "945558": "Here is what I ended up doing, which is quite messy but works:\n\n```\ntrain_df['patient_id'] = train_df['patient_id'].replace(-1,np.NaN, regex=True)\n\npid_list = train_df['patient_id'].tolist()\nall_list = []\nunused_list = []\nreplace_list = []\n\nfor index in range(len(pid_list)):\n    if isinstance(pid_list[index], str):\n        pid_list[index] = int(pid_list[index][3:])\npid_list=list(set(pid_list))\nfor index in range(1,10000000):\n    all_list.append(index)\nunused_list = list(set(all_list) - set(pid_list))\nreplace_list = random.sample(unused_list, len(train_df))\nfor index in range(len(replace_list)):\n    replace_list[index] = \"IP_\" + str(replace_list[index])\n    \ntrain_df['patient_id'] = train_df['patient_id'].fillna(pd.Series(replace_list))\n```"
  },
  "source": "meta"
}