{
  "id": 253108,
  "title": "Difference between 'id' column in train_study_level.csv and 'StudyInstanceUID' column in train_image_level.csv",
  "url": "/competitions/siim-covid19-detection/discussion/253108",
  "author_name": "",
  "post_date": "2021-07-15T04:44:26.048539700Z",
  "votes": -2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>What is the difference between the 'id' column in train_study_level.csv and the 'StudyInstanceUID' column in train_image_level.csv ? Because I found that after removing \"_study\" from the id's in the train_study_level.csv and comparing those with the values in the 'StudyInstanceUID' column , most of the values matched but 723 of those didn't match - <br>\nThis is the code - </p>\n<p>len(set(train_study['id'].str.strip(\"_study\")).difference(set(train_image['StudyInstanceUID'].unique())))</p>\n<p>this is the output - <br>\n723 . </p>",
  "messages": [
    {
      "id": "1388571",
      "postDate": "07/15/2021 04:44:26",
      "content": "<p>What is the difference between the 'id' column in train_study_level.csv and the 'StudyInstanceUID' column in train_image_level.csv ? Because I found that after removing \"_study\" from the id's in the train_study_level.csv and comparing those with the values in the 'StudyInstanceUID' column , most of the values matched but 723 of those didn't match - <br>\nThis is the code - </p>\n<p>len(set(train_study['id'].str.strip(\"_study\")).difference(set(train_image['StudyInstanceUID'].unique())))</p>\n<p>this is the output - <br>\n723 . </p>",
      "rawMarkdown": "What is the difference between the 'id' column in train_study_level.csv and the 'StudyInstanceUID' column in train_image_level.csv ? Because I found that after removing \"_study\" from the id's in the train_study_level.csv and comparing those with the values in the 'StudyInstanceUID' column , most of the values matched but 723 of those didn't match - \nThis is the code - \n\nlen(set(train_study['id'].str.strip(\"_study\")).difference(set(train_image['StudyInstanceUID'].unique())))\n\nthis is the output - \n723 .",
      "votes": null
    },
    {
      "id": "1389665",
      "postDate": "07/15/2021 23:37:47",
      "content": "<p>They are exaclty the same.  Both have 6054 unique study ids. StudyInstanceUID  repeats some to result in 6334 ids. </p>",
      "rawMarkdown": "They are exaclty the same.  Both have 6054 unique study ids. StudyInstanceUID  repeats some to result in 6334 ids.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1389665,
      "author_name": "rpsantosakaggle",
      "author_url": "",
      "post_date": "07/15/2021 23:37:47",
      "content": "<p>They are exaclty the same.  Both have 6054 unique study ids. StudyInstanceUID  repeats some to result in 6334 ids. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1388571": "What is the difference between the 'id' column in train_study_level.csv and the 'StudyInstanceUID' column in train_image_level.csv ? Because I found that after removing \"_study\" from the id's in the train_study_level.csv and comparing those with the values in the 'StudyInstanceUID' column , most of the values matched but 723 of those didn't match - \nThis is the code - \n\nlen(set(train_study['id'].str.strip(\"_study\")).difference(set(train_image['StudyInstanceUID'].unique())))\n\nthis is the output - \n723 .",
    "1389665": "They are exaclty the same.  Both have 6054 unique study ids. StudyInstanceUID  repeats some to result in 6334 ids."
  },
  "source": "meta"
}