{
  "id": 84132,
  "title": "How to work with WSI ids",
  "url": "/competitions/histopathologic-cancer-detection/discussion/84132",
  "author_name": "",
  "post_date": "2019-03-14T19:45:12.937580400Z",
  "votes": 18,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Complete code for train/cv split based on WSI from this discussion <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760</a></p>\n\n<p>At the end you will have arrays with ids and labels for training and cross validation, and information about the generated split (how many images went to training/cv and how many of them are positive/negative). The only thing left for you is to change the way you load the data. </p>\n\n<p>The resulting arrays are: train ids, train labels and cv ids and cv labels</p>\n\n<p>If you have any questions, just let me know. \n```\nimport pandas as pd\nimport numpy as np</p>\n\n<p>def return_tumor_or_not(dic, one_id):\n    return dic[one_id]</p>\n\n<p>def create_dict():\n    df = pd.read_csv(\"../input/train_labels.csv\")\n    result_dict = {}\n    for index in range(df.shape[0]):\n        one_id = df.iloc[index,0]\n        tumor_or_not = df.iloc[index,1]\n        result_dict[one_id] = int(tumor_or_not)\n    return result_dict</p>\n\n<p>def find_missing(train_ids, cv_ids):\n    all_ids = set(pd.read_csv(\"../input/train_labels.csv\")['id'].values)\n    wsi_ids = set(train_ids + cv_ids)</p>\n\n<pre><code>missing_ids = list(all_ids-wsi_ids)\nreturn missing_ids\n</code></pre>\n\n<p>def generate_split():\n    ids = pd.read_csv(\"../input/patch_id_wsi.csv\")\n    wsi_dict = {}\n    for i in range(ids.shape[0]):\n        wsi = ids.iloc[i,1]\n        train_id = ids.iloc[i,0]\n        wsi_array = wsi.split('_')\n        number = int(wsi_array[3])\n        if wsi_dict.get(number) is None:\n            wsi_dict[number] = [train_id]\n        else:\n            wsi_dict[number].append(train_id)</p>\n\n<pre><code>wsi_keys = list(wsi_dict.keys())\nnp.random.seed()\nnp.random.shuffle(wsi_keys)\namount_of_keys = len(wsi_keys)\n\nkeys_for_train = wsi_keys[0:int(amount_of_keys*0.8)]\nkeys_for_cv = wsi_keys[int(amount_of_keys*0.8):]\ntrain_ids = []\ncv_ids = []\n\nfor key in keys_for_train:\n    train_ids += wsi_dict[key]\n\nfor key in keys_for_cv:\n    cv_ids += wsi_dict[key]\n\ndic = create_dict()\n\nmissing_ids = find_missing(train_ids, cv_ids)\nmissing_ids_total = len(missing_ids)\nnp.random.seed()\nnp.random.shuffle(missing_ids)\n\ntrain_missing_ids = missing_ids[0:int(missing_ids_total*0.8)]\ncv_missing_ids = missing_ids[int(missing_ids_total*0.8):]\n\ntrain_ids += train_missing_ids\ncv_ids += cv_missing_ids\n\ntrain_labels = []\ncv_labels = []\n\ntrain_tumor = 0\nfor one_id in train_ids:\n    temp = return_tumor_or_not(dic, one_id)\n    train_tumor += temp\n    train_labels.append(temp)\n\ncv_tumor = 0\nfor one_id in cv_ids:\n    temp = return_tumor_or_not(dic, one_id)\n    cv_tumor += temp\n    cv_labels.append(temp)\ntotal = len(train_ids) + len(cv_ids)\n\nprint(\"Amount of train labels: {}, {}/{}\".format(len(train_ids), train_tumor, len(train_ids)-train_tumor))\nprint(\"Amount of cv labels: {}, {}/{}\".format(len(cv_ids), cv_tumor, len(cv_ids) - cv_tumor))\nprint(\"Percentage of cv labels: {}\".format(len(cv_ids)/total))\n\nreturn train_ids, cv_ids, train_labels, cv_labels\n</code></pre>\n\n<p>train_ids, cv_ids, train_labels, cv_labels = generate_split()\n```</p>",
  "messages": [
    {
      "id": "490721",
      "postDate": "03/14/2019 19:45:12",
      "content": "<p>Complete code for train/cv split based on WSI from this discussion <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760</a></p>\n\n<p>At the end you will have arrays with ids and labels for training and cross validation, and information about the generated split (how many images went to training/cv and how many of them are positive/negative). The only thing left for you is to change the way you load the data. </p>\n\n<p>The resulting arrays are: train ids, train labels and cv ids and cv labels</p>\n\n<p>If you have any questions, just let me know. \n```\nimport pandas as pd\nimport numpy as np</p>\n\n<p>def return_tumor_or_not(dic, one_id):\n    return dic[one_id]</p>\n\n<p>def create_dict():\n    df = pd.read_csv(\"../input/train_labels.csv\")\n    result_dict = {}\n    for index in range(df.shape[0]):\n        one_id = df.iloc[index,0]\n        tumor_or_not = df.iloc[index,1]\n        result_dict[one_id] = int(tumor_or_not)\n    return result_dict</p>\n\n<p>def find_missing(train_ids, cv_ids):\n    all_ids = set(pd.read_csv(\"../input/train_labels.csv\")['id'].values)\n    wsi_ids = set(train_ids + cv_ids)</p>\n\n<pre><code>missing_ids = list(all_ids-wsi_ids)\nreturn missing_ids\n</code></pre>\n\n<p>def generate_split():\n    ids = pd.read_csv(\"../input/patch_id_wsi.csv\")\n    wsi_dict = {}\n    for i in range(ids.shape[0]):\n        wsi = ids.iloc[i,1]\n        train_id = ids.iloc[i,0]\n        wsi_array = wsi.split('_')\n        number = int(wsi_array[3])\n        if wsi_dict.get(number) is None:\n            wsi_dict[number] = [train_id]\n        else:\n            wsi_dict[number].append(train_id)</p>\n\n<pre><code>wsi_keys = list(wsi_dict.keys())\nnp.random.seed()\nnp.random.shuffle(wsi_keys)\namount_of_keys = len(wsi_keys)\n\nkeys_for_train = wsi_keys[0:int(amount_of_keys*0.8)]\nkeys_for_cv = wsi_keys[int(amount_of_keys*0.8):]\ntrain_ids = []\ncv_ids = []\n\nfor key in keys_for_train:\n    train_ids += wsi_dict[key]\n\nfor key in keys_for_cv:\n    cv_ids += wsi_dict[key]\n\ndic = create_dict()\n\nmissing_ids = find_missing(train_ids, cv_ids)\nmissing_ids_total = len(missing_ids)\nnp.random.seed()\nnp.random.shuffle(missing_ids)\n\ntrain_missing_ids = missing_ids[0:int(missing_ids_total*0.8)]\ncv_missing_ids = missing_ids[int(missing_ids_total*0.8):]\n\ntrain_ids += train_missing_ids\ncv_ids += cv_missing_ids\n\ntrain_labels = []\ncv_labels = []\n\ntrain_tumor = 0\nfor one_id in train_ids:\n    temp = return_tumor_or_not(dic, one_id)\n    train_tumor += temp\n    train_labels.append(temp)\n\ncv_tumor = 0\nfor one_id in cv_ids:\n    temp = return_tumor_or_not(dic, one_id)\n    cv_tumor += temp\n    cv_labels.append(temp)\ntotal = len(train_ids) + len(cv_ids)\n\nprint(\"Amount of train labels: {}, {}/{}\".format(len(train_ids), train_tumor, len(train_ids)-train_tumor))\nprint(\"Amount of cv labels: {}, {}/{}\".format(len(cv_ids), cv_tumor, len(cv_ids) - cv_tumor))\nprint(\"Percentage of cv labels: {}\".format(len(cv_ids)/total))\n\nreturn train_ids, cv_ids, train_labels, cv_labels\n</code></pre>\n\n<p>train_ids, cv_ids, train_labels, cv_labels = generate_split()\n```</p>",
      "rawMarkdown": "Complete code for train/cv split based on WSI from this discussion https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\n\nAt the end you will have arrays with ids and labels for training and cross validation, and information about the generated split (how many images went to training/cv and how many of them are positive/negative). The only thing left for you is to change the way you load the data. \n\nThe resulting arrays are: train ids, train labels and cv ids and cv labels\n\nIf you have any questions, just let me know. \n```\nimport pandas as pd\nimport numpy as np\n\ndef return_tumor_or_not(dic, one_id):\n    return dic[one_id]\n\ndef create_dict():\n    df = pd.read_csv(\"../input/train_labels.csv\")\n    result_dict = {}\n    for index in range(df.shape[0]):\n        one_id = df.iloc[index,0]\n        tumor_or_not = df.iloc[index,1]\n        result_dict[one_id] = int(tumor_or_not)\n    return result_dict\n\ndef find_missing(train_ids, cv_ids):\n    all_ids = set(pd.read_csv(\"../input/train_labels.csv\")['id'].values)\n    wsi_ids = set(train_ids + cv_ids)\n\n    missing_ids = list(all_ids-wsi_ids)\n    return missing_ids\n\n\ndef generate_split():\n    ids = pd.read_csv(\"../input/patch_id_wsi.csv\")\n    wsi_dict = {}\n    for i in range(ids.shape[0]):\n        wsi = ids.iloc[i,1]\n        train_id = ids.iloc[i,0]\n        wsi_array = wsi.split('_')\n        number = int(wsi_array[3])\n        if wsi_dict.get(number) is None:\n            wsi_dict[number] = [train_id]\n        else:\n            wsi_dict[number].append(train_id)\n\n    wsi_keys = list(wsi_dict.keys())\n    np.random.seed()\n    np.random.shuffle(wsi_keys)\n    amount_of_keys = len(wsi_keys)\n\n    keys_for_train = wsi_keys[0:int(amount_of_keys*0.8)]\n    keys_for_cv = wsi_keys[int(amount_of_keys*0.8):]\n    train_ids = []\n    cv_ids = []\n\n    for key in keys_for_train:\n        train_ids += wsi_dict[key]\n\n    for key in keys_for_cv:\n        cv_ids += wsi_dict[key]\n\n    dic = create_dict()\n\n    missing_ids = find_missing(train_ids, cv_ids)\n    missing_ids_total = len(missing_ids)\n    np.random.seed()\n    np.random.shuffle(missing_ids)\n\n    train_missing_ids = missing_ids[0:int(missing_ids_total*0.8)]\n    cv_missing_ids = missing_ids[int(missing_ids_total*0.8):]\n\n    train_ids += train_missing_ids\n    cv_ids += cv_missing_ids\n\n    train_labels = []\n    cv_labels = []\n\n    train_tumor = 0\n    for one_id in train_ids:\n        temp = return_tumor_or_not(dic, one_id)\n        train_tumor += temp\n        train_labels.append(temp)\n\n    cv_tumor = 0\n    for one_id in cv_ids:\n        temp = return_tumor_or_not(dic, one_id)\n        cv_tumor += temp\n        cv_labels.append(temp)\n    total = len(train_ids) + len(cv_ids)\n\n    print(\"Amount of train labels: {}, {}/{}\".format(len(train_ids), train_tumor, len(train_ids)-train_tumor))\n    print(\"Amount of cv labels: {}, {}/{}\".format(len(cv_ids), cv_tumor, len(cv_ids) - cv_tumor))\n    print(\"Percentage of cv labels: {}\".format(len(cv_ids)/total))\n\n    return train_ids, cv_ids, train_labels, cv_labels\n\ntrain_ids, cv_ids, train_labels, cv_labels = generate_split()\n```",
      "votes": null
    },
    {
      "id": "491069",
      "postDate": "03/15/2019 06:44:42",
      "content": "<p>Thank you Ivan!</p>",
      "rawMarkdown": "Thank you Ivan!",
      "votes": null
    },
    {
      "id": "491075",
      "postDate": "03/15/2019 06:57:54",
      "content": "<p>Glad to be helpful :)</p>",
      "rawMarkdown": "Glad to be helpful :)",
      "votes": null
    },
    {
      "id": "491100",
      "postDate": "03/15/2019 07:39:55",
      "content": "<p>Must be the sum of (train_ids + cv_ids) equal to the total number of ids given in train_labels or in wsi?</p>",
      "rawMarkdown": "Must be the sum of (train_ids + cv_ids) equal to the total number of ids given in train_labels or in wsi?",
      "votes": null
    },
    {
      "id": "491140",
      "postDate": "03/15/2019 08:43:27",
      "content": "<p>Total number of ids given in the train set, of course. Because train ids + cv ids = wsi ids + missing ids (all ids - wsi ids) </p>",
      "rawMarkdown": "Total number of ids given in the train set, of course. Because train ids + cv ids = wsi ids + missing ids (all ids - wsi ids)",
      "votes": null
    },
    {
      "id": "491285",
      "postDate": "03/15/2019 12:29:41",
      "content": "<p>Thank you Ivan! It help me improve my results.\nI have a question,  you differ wsi through wsi id(wsi_array[3]).\nBut  in the patch_id_wsi.csv, i found it have tumor_096 and normal_096, they share the same wsi id number but they are total different slide.\nmaybe you can try to use wsi_array[2]+wsi_array[3] to differ data. </p>",
      "rawMarkdown": "Thank you Ivan! It help me improve my results.\nI have a question,  you differ wsi through wsi id(wsi_array[3]).\nBut  in the patch_id_wsi.csv, i found it have tumor_096 and normal_096, they share the same wsi id number but they are total different slide.\nmaybe you can try to use wsi_array[2]+wsi_array[3] to differ data.",
      "votes": null
    },
    {
      "id": "491310",
      "postDate": "03/15/2019 13:02:27",
      "content": "<p>Happy it helped you. </p>\n\n<p>That is unnecessarily. You see, if the WSI label is normal, then all patches are normal. But if WSI label is tumor, then some patches still could be normal (I even found one, where all patches are normal). In other words, there is no need for us to do extra work in order to better separate slides. </p>",
      "rawMarkdown": "Happy it helped you. \n\nThat is unnecessarily. You see, if the WSI label is normal, then all patches are normal. But if WSI label is tumor, then some patches still could be normal (I even found one, where all patches are normal). In other words, there is no need for us to do extra work in order to better separate slides.",
      "votes": null
    },
    {
      "id": "491335",
      "postDate": "03/15/2019 13:44:18",
      "content": "<p>Nice job!! Upvoted!</p>",
      "rawMarkdown": "Nice job!! Upvoted!",
      "votes": null
    },
    {
      "id": "492432",
      "postDate": "03/17/2019 09:24:44",
      "content": "<p>This approach didn't improve my score. Or perhaps I did something wrong. Will check it.</p>",
      "rawMarkdown": "This approach didn't improve my score. Or perhaps I did something wrong. Will check it.",
      "votes": null
    },
    {
      "id": "492578",
      "postDate": "03/17/2019 13:28:08",
      "content": "<p>I still get auc=0.99 in validation set using this method but 0.96+ in LB.  This is still unavoidable over-fitting？</p>",
      "rawMarkdown": "I still get auc=0.99 in validation set using this method but 0.96+ in LB.  This is still unavoidable over-fitting？",
      "votes": null
    },
    {
      "id": "492580",
      "postDate": "03/17/2019 13:29:52",
      "content": "<p>Most likely. I no longer can achieve 0.99+ on validation set, but my LB score got higher. </p>",
      "rawMarkdown": "Most likely. I no longer can achieve 0.99+ on validation set, but my LB score got higher.",
      "votes": null
    },
    {
      "id": "492618",
      "postDate": "03/17/2019 14:11:30",
      "content": "<p>Hi, Ivan. With this new splitting, my densenet169 can get 0.9761 than 0.9742, but densenet121 is worse get 0.9666 now (it should be more than 0.9700 before. Btw, the val auc is around 0.98 (both 2 models.</p>",
      "rawMarkdown": "Hi, Ivan. With this new splitting, my densenet169 can get 0.9761 than 0.9742, but densenet121 is worse get 0.9666 now (it should be more than 0.9700 before. Btw, the val auc is around 0.98 (both 2 models.",
      "votes": null
    },
    {
      "id": "493100",
      "postDate": "03/18/2019 09:15:47",
      "content": "<p>Thanks Ivan!! I hope it helps me with my score.</p>",
      "rawMarkdown": "Thanks Ivan!! I hope it helps me with my score.",
      "votes": null
    },
    {
      "id": "494385",
      "postDate": "03/19/2019 19:30:40",
      "content": "<p>Does this line split the train and val data?</p>\n\n<p>train_ids, cv_ids, train_labels, cv_labels = generate_split()</p>",
      "rawMarkdown": "Does this line split the train and val data?\n\ntrain_ids, cv_ids, train_labels, cv_labels = generate_split()",
      "votes": null
    },
    {
      "id": "494392",
      "postDate": "03/19/2019 19:33:49",
      "content": "<p>Yes, exactly. </p>",
      "rawMarkdown": "Yes, exactly.",
      "votes": null
    },
    {
      "id": "494590",
      "postDate": "03/20/2019 02:55:20",
      "content": "<p>Thank you! I was just trying to figure out how update this kernel. </p>\n\n<p><a href=\"https://www.kaggle.com/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb\">https://www.kaggle.com/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb</a></p>",
      "rawMarkdown": "Thank you! I was just trying to figure out how update this kernel. \n\nhttps://www.kaggle.com/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb",
      "votes": null
    },
    {
      "id": "494593",
      "postDate": "03/20/2019 03:03:22",
      "content": "<p>Got it. But why are you interested in this specific kernel? Btw, would you like to discuss this competition in Telegram? </p>",
      "rawMarkdown": "Got it. But why are you interested in this specific kernel? Btw, would you like to discuss this competition in Telegram?",
      "votes": null
    },
    {
      "id": "495087",
      "postDate": "03/20/2019 16:08:40",
      "content": "<p>We can discuss in Telegram. But first what is Telegram.  My private kernel is setup similar to this kernel. I made some minor changes to it based on Umberto Colab kernel.  </p>",
      "rawMarkdown": "We can discuss in Telegram. But first what is Telegram.  My private kernel is setup similar to this kernel. I made some minor changes to it based on Umberto Colab kernel.",
      "votes": null
    },
    {
      "id": "495089",
      "postDate": "03/20/2019 16:12:18",
      "content": "<p>Sent you an email. </p>",
      "rawMarkdown": "Sent you an email.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 491069,
      "author_name": "robotdreams",
      "author_url": "",
      "post_date": "03/15/2019 06:44:42",
      "content": "<p>Thank you Ivan!</p>",
      "votes": null,
      "replies": [
        {
          "id": 491075,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/15/2019 06:57:54",
          "content": "<p>Glad to be helpful :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 491100,
      "author_name": "robotdreams",
      "author_url": "",
      "post_date": "03/15/2019 07:39:55",
      "content": "<p>Must be the sum of (train_ids + cv_ids) equal to the total number of ids given in train_labels or in wsi?</p>",
      "votes": null,
      "replies": [
        {
          "id": 491140,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/15/2019 08:43:27",
          "content": "<p>Total number of ids given in the train set, of course. Because train ids + cv ids = wsi ids + missing ids (all ids - wsi ids) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 491285,
      "author_name": "lvguofeng",
      "author_url": "",
      "post_date": "03/15/2019 12:29:41",
      "content": "<p>Thank you Ivan! It help me improve my results.\nI have a question,  you differ wsi through wsi id(wsi_array[3]).\nBut  in the patch_id_wsi.csv, i found it have tumor_096 and normal_096, they share the same wsi id number but they are total different slide.\nmaybe you can try to use wsi_array[2]+wsi_array[3] to differ data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 491310,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/15/2019 13:02:27",
          "content": "<p>Happy it helped you. </p>\n\n<p>That is unnecessarily. You see, if the WSI label is normal, then all patches are normal. But if WSI label is tumor, then some patches still could be normal (I even found one, where all patches are normal). In other words, there is no need for us to do extra work in order to better separate slides. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 491335,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "03/15/2019 13:44:18",
      "content": "<p>Nice job!! Upvoted!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 492432,
      "author_name": "robotdreams",
      "author_url": "",
      "post_date": "03/17/2019 09:24:44",
      "content": "<p>This approach didn't improve my score. Or perhaps I did something wrong. Will check it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 492578,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "03/17/2019 13:28:08",
          "content": "<p>I still get auc=0.99 in validation set using this method but 0.96+ in LB.  This is still unavoidable over-fitting？</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 492580,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/17/2019 13:29:52",
          "content": "<p>Most likely. I no longer can achieve 0.99+ on validation set, but my LB score got higher. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 492618,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "03/17/2019 14:11:30",
          "content": "<p>Hi, Ivan. With this new splitting, my densenet169 can get 0.9761 than 0.9742, but densenet121 is worse get 0.9666 now (it should be more than 0.9700 before. Btw, the val auc is around 0.98 (both 2 models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 493100,
      "author_name": "abhinav0208",
      "author_url": "",
      "post_date": "03/18/2019 09:15:47",
      "content": "<p>Thanks Ivan!! I hope it helps me with my score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 494385,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "03/19/2019 19:30:40",
      "content": "<p>Does this line split the train and val data?</p>\n\n<p>train_ids, cv_ids, train_labels, cv_labels = generate_split()</p>",
      "votes": null,
      "replies": [
        {
          "id": 494392,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/19/2019 19:33:49",
          "content": "<p>Yes, exactly. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494590,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "03/20/2019 02:55:20",
          "content": "<p>Thank you! I was just trying to figure out how update this kernel. </p>\n\n<p><a href=\"https://www.kaggle.com/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb\">https://www.kaggle.com/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494593,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/20/2019 03:03:22",
          "content": "<p>Got it. But why are you interested in this specific kernel? Btw, would you like to discuss this competition in Telegram? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495087,
          "author_name": "dskswu",
          "author_url": "",
          "post_date": "03/20/2019 16:08:40",
          "content": "<p>We can discuss in Telegram. But first what is Telegram.  My private kernel is setup similar to this kernel. I made some minor changes to it based on Umberto Colab kernel.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495089,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/20/2019 16:12:18",
          "content": "<p>Sent you an email. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "490721": "Complete code for train/cv split based on WSI from this discussion https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\n\nAt the end you will have arrays with ids and labels for training and cross validation, and information about the generated split (how many images went to training/cv and how many of them are positive/negative). The only thing left for you is to change the way you load the data. \n\nThe resulting arrays are: train ids, train labels and cv ids and cv labels\n\nIf you have any questions, just let me know. \n```\nimport pandas as pd\nimport numpy as np\n\ndef return_tumor_or_not(dic, one_id):\n    return dic[one_id]\n\ndef create_dict():\n    df = pd.read_csv(\"../input/train_labels.csv\")\n    result_dict = {}\n    for index in range(df.shape[0]):\n        one_id = df.iloc[index,0]\n        tumor_or_not = df.iloc[index,1]\n        result_dict[one_id] = int(tumor_or_not)\n    return result_dict\n\ndef find_missing(train_ids, cv_ids):\n    all_ids = set(pd.read_csv(\"../input/train_labels.csv\")['id'].values)\n    wsi_ids = set(train_ids + cv_ids)\n\n    missing_ids = list(all_ids-wsi_ids)\n    return missing_ids\n\n\ndef generate_split():\n    ids = pd.read_csv(\"../input/patch_id_wsi.csv\")\n    wsi_dict = {}\n    for i in range(ids.shape[0]):\n        wsi = ids.iloc[i,1]\n        train_id = ids.iloc[i,0]\n        wsi_array = wsi.split('_')\n        number = int(wsi_array[3])\n        if wsi_dict.get(number) is None:\n            wsi_dict[number] = [train_id]\n        else:\n            wsi_dict[number].append(train_id)\n\n    wsi_keys = list(wsi_dict.keys())\n    np.random.seed()\n    np.random.shuffle(wsi_keys)\n    amount_of_keys = len(wsi_keys)\n\n    keys_for_train = wsi_keys[0:int(amount_of_keys*0.8)]\n    keys_for_cv = wsi_keys[int(amount_of_keys*0.8):]\n    train_ids = []\n    cv_ids = []\n\n    for key in keys_for_train:\n        train_ids += wsi_dict[key]\n\n    for key in keys_for_cv:\n        cv_ids += wsi_dict[key]\n\n    dic = create_dict()\n\n    missing_ids = find_missing(train_ids, cv_ids)\n    missing_ids_total = len(missing_ids)\n    np.random.seed()\n    np.random.shuffle(missing_ids)\n\n    train_missing_ids = missing_ids[0:int(missing_ids_total*0.8)]\n    cv_missing_ids = missing_ids[int(missing_ids_total*0.8):]\n\n    train_ids += train_missing_ids\n    cv_ids += cv_missing_ids\n\n    train_labels = []\n    cv_labels = []\n\n    train_tumor = 0\n    for one_id in train_ids:\n        temp = return_tumor_or_not(dic, one_id)\n        train_tumor += temp\n        train_labels.append(temp)\n\n    cv_tumor = 0\n    for one_id in cv_ids:\n        temp = return_tumor_or_not(dic, one_id)\n        cv_tumor += temp\n        cv_labels.append(temp)\n    total = len(train_ids) + len(cv_ids)\n\n    print(\"Amount of train labels: {}, {}/{}\".format(len(train_ids), train_tumor, len(train_ids)-train_tumor))\n    print(\"Amount of cv labels: {}, {}/{}\".format(len(cv_ids), cv_tumor, len(cv_ids) - cv_tumor))\n    print(\"Percentage of cv labels: {}\".format(len(cv_ids)/total))\n\n    return train_ids, cv_ids, train_labels, cv_labels\n\ntrain_ids, cv_ids, train_labels, cv_labels = generate_split()\n```",
    "491069": "Thank you Ivan!",
    "491075": "Glad to be helpful :)",
    "491100": "Must be the sum of (train_ids + cv_ids) equal to the total number of ids given in train_labels or in wsi?",
    "491140": "Total number of ids given in the train set, of course. Because train ids + cv ids = wsi ids + missing ids (all ids - wsi ids)",
    "491285": "Thank you Ivan! It help me improve my results.\nI have a question,  you differ wsi through wsi id(wsi_array[3]).\nBut  in the patch_id_wsi.csv, i found it have tumor_096 and normal_096, they share the same wsi id number but they are total different slide.\nmaybe you can try to use wsi_array[2]+wsi_array[3] to differ data.",
    "491310": "Happy it helped you. \n\nThat is unnecessarily. You see, if the WSI label is normal, then all patches are normal. But if WSI label is tumor, then some patches still could be normal (I even found one, where all patches are normal). In other words, there is no need for us to do extra work in order to better separate slides.",
    "491335": "Nice job!! Upvoted!",
    "492432": "This approach didn't improve my score. Or perhaps I did something wrong. Will check it.",
    "492578": "I still get auc=0.99 in validation set using this method but 0.96+ in LB.  This is still unavoidable over-fitting？",
    "492580": "Most likely. I no longer can achieve 0.99+ on validation set, but my LB score got higher.",
    "492618": "Hi, Ivan. With this new splitting, my densenet169 can get 0.9761 than 0.9742, but densenet121 is worse get 0.9666 now (it should be more than 0.9700 before. Btw, the val auc is around 0.98 (both 2 models.",
    "493100": "Thanks Ivan!! I hope it helps me with my score.",
    "494385": "Does this line split the train and val data?\n\ntrain_ids, cv_ids, train_labels, cv_labels = generate_split()",
    "494392": "Yes, exactly.",
    "494590": "Thank you! I was just trying to figure out how update this kernel. \n\nhttps://www.kaggle.com/CVxTz/cnn-starter-nasnet-mobile-0-9709-lb",
    "494593": "Got it. But why are you interested in this specific kernel? Btw, would you like to discuss this competition in Telegram?",
    "495087": "We can discuss in Telegram. But first what is Telegram.  My private kernel is setup similar to this kernel. I made some minor changes to it based on Umberto Colab kernel.",
    "495089": "Sent you an email."
  },
  "source": "meta"
}