{
  "id": 85185,
  "title": "Is WSI really working?",
  "url": "/competitions/histopathologic-cancer-detection/discussion/85185",
  "author_name": "",
  "post_date": "2019-03-22T06:38:52.197679700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I split the dataset according to WSI like this:</p>\n\n<p>```python\n    # Data loading code\n    df = pd.read_csv(os.path.join(args.data, 'train_labels.csv'))\n    df_wsi = pd.read_csv(os.path.join(args.data, 'patch_id_wsi.csv'))\n    df_test = pd.read_csv(os.path.join(args.data, 'sample_submission.csv'))\n    df_total = pd.merge(df, df_wsi, on='id', how='outer')\n    df_total.fillna('unknown', inplace=True)\n    df_total = shuffle(df_total, random_state=0)\n    df_total['wsi'] = df_total['wsi'].apply(lambda x: x if x == 'unkown' else x.split('_')[-1])\n    df_total = df_total.sort_values(by='wsi')</p>\n\n<pre><code>kf = KFold(n_splits=5, shuffle=False, random_state=0)\nfor idx, (train_idx, val_idx) in enumerate(kf.split(df_total, df_total['wsi'])):\n    if idx &amp;gt; args.split:\n        tr_wsi = df_total.iloc[train_idx]['wsi'].unique()\n        vl_wsi = df_total.iloc[val_idx]['wsi'].unique()\n        print(f\"Training WSI: {tr_wsi}, {len(tr_wsi)}, {len(train_idx)}\")\n        print(f\"Validating WSI: {vl_wsi}, {len(vl_wsi)}, {len(val_idx)}\")\n        break\n</code></pre>\n\n<p>```\nWhich will give printout:</p>\n\n<p><code>bash\nTraining WSI: ['001' '002' '003' '004' '005' '006' '007' '008' '009' '010' '011' '012' '014' '015' '016' '017' '018' '020' '021' '022' '023' '024' '025' '026' '027' '028' '029' '030' '032' '033' '034' '061' '062' '063' '064' '065' '066' '068' '069' '070' '071' '072' '073' '074' '075' '076' '077' '078'\n '080' '081' '082' '083' '084' '086' '087' '088' '089' '090' '091' '092'\n '093' '094' '095' '096' '097' '098' '099' '100' '101' '102' '103' '104'\n '105' '106' '107' '108' '110' '111' '112' '113' '114' '115' '116' '117'\n '118' '119' '120' '121' '122' '123' '124' '125' '126' '127' '128' '130'\n '131' '133' '134' '135' '136' '138' '139' '140' '141' '144' '145' '146'\n '147' '149' '150' '151' '153' '154' '155' '156' '157' '158' '159' '160'\n 'unknown'], 121, 176020\nValidating WSI: ['034' '035' '036' '037' '038' '039' '040' '041' '042' '043' '044' '045'\n '046' '047' '048' '049' '050' '051' '052' '053' '054' '055' '056' '058'\n '059' '060' '061'], 27, 44005\n</code>\nBut still got validation auc 0.99, but submission as 0.9745</p>",
  "messages": [
    {
      "id": "496408",
      "postDate": "03/22/2019 06:38:52",
      "content": "<p>I split the dataset according to WSI like this:</p>\n\n<p>```python\n    # Data loading code\n    df = pd.read_csv(os.path.join(args.data, 'train_labels.csv'))\n    df_wsi = pd.read_csv(os.path.join(args.data, 'patch_id_wsi.csv'))\n    df_test = pd.read_csv(os.path.join(args.data, 'sample_submission.csv'))\n    df_total = pd.merge(df, df_wsi, on='id', how='outer')\n    df_total.fillna('unknown', inplace=True)\n    df_total = shuffle(df_total, random_state=0)\n    df_total['wsi'] = df_total['wsi'].apply(lambda x: x if x == 'unkown' else x.split('_')[-1])\n    df_total = df_total.sort_values(by='wsi')</p>\n\n<pre><code>kf = KFold(n_splits=5, shuffle=False, random_state=0)\nfor idx, (train_idx, val_idx) in enumerate(kf.split(df_total, df_total['wsi'])):\n    if idx &amp;gt; args.split:\n        tr_wsi = df_total.iloc[train_idx]['wsi'].unique()\n        vl_wsi = df_total.iloc[val_idx]['wsi'].unique()\n        print(f\"Training WSI: {tr_wsi}, {len(tr_wsi)}, {len(train_idx)}\")\n        print(f\"Validating WSI: {vl_wsi}, {len(vl_wsi)}, {len(val_idx)}\")\n        break\n</code></pre>\n\n<p>```\nWhich will give printout:</p>\n\n<p><code>bash\nTraining WSI: ['001' '002' '003' '004' '005' '006' '007' '008' '009' '010' '011' '012' '014' '015' '016' '017' '018' '020' '021' '022' '023' '024' '025' '026' '027' '028' '029' '030' '032' '033' '034' '061' '062' '063' '064' '065' '066' '068' '069' '070' '071' '072' '073' '074' '075' '076' '077' '078'\n '080' '081' '082' '083' '084' '086' '087' '088' '089' '090' '091' '092'\n '093' '094' '095' '096' '097' '098' '099' '100' '101' '102' '103' '104'\n '105' '106' '107' '108' '110' '111' '112' '113' '114' '115' '116' '117'\n '118' '119' '120' '121' '122' '123' '124' '125' '126' '127' '128' '130'\n '131' '133' '134' '135' '136' '138' '139' '140' '141' '144' '145' '146'\n '147' '149' '150' '151' '153' '154' '155' '156' '157' '158' '159' '160'\n 'unknown'], 121, 176020\nValidating WSI: ['034' '035' '036' '037' '038' '039' '040' '041' '042' '043' '044' '045'\n '046' '047' '048' '049' '050' '051' '052' '053' '054' '055' '056' '058'\n '059' '060' '061'], 27, 44005\n</code>\nBut still got validation auc 0.99, but submission as 0.9745</p>",
      "rawMarkdown": "I split the dataset according to WSI like this:\n\n\n```python\n    # Data loading code\n    df = pd.read_csv(os.path.join(args.data, 'train_labels.csv'))\n    df_wsi = pd.read_csv(os.path.join(args.data, 'patch_id_wsi.csv'))\n    df_test = pd.read_csv(os.path.join(args.data, 'sample_submission.csv'))\n    df_total = pd.merge(df, df_wsi, on='id', how='outer')\n    df_total.fillna('unknown', inplace=True)\n    df_total = shuffle(df_total, random_state=0)\n    df_total['wsi'] = df_total['wsi'].apply(lambda x: x if x == 'unkown' else x.split('_')[-1])\n    df_total = df_total.sort_values(by='wsi')\n\n    kf = KFold(n_splits=5, shuffle=False, random_state=0)\n    for idx, (train_idx, val_idx) in enumerate(kf.split(df_total, df_total['wsi'])):\n        if idx &gt; args.split:\n            tr_wsi = df_total.iloc[train_idx]['wsi'].unique()\n            vl_wsi = df_total.iloc[val_idx]['wsi'].unique()\n            print(f\"Training WSI: {tr_wsi}, {len(tr_wsi)}, {len(train_idx)}\")\n            print(f\"Validating WSI: {vl_wsi}, {len(vl_wsi)}, {len(val_idx)}\")\n            break\n```\nWhich will give printout:\n\n```bash\nTraining WSI: ['001' '002' '003' '004' '005' '006' '007' '008' '009' '010' '011' '012' '014' '015' '016' '017' '018' '020' '021' '022' '023' '024' '025' '026' '027' '028' '029' '030' '032' '033' '034' '061' '062' '063' '064' '065' '066' '068' '069' '070' '071' '072' '073' '074' '075' '076' '077' '078'\n '080' '081' '082' '083' '084' '086' '087' '088' '089' '090' '091' '092'\n '093' '094' '095' '096' '097' '098' '099' '100' '101' '102' '103' '104'\n '105' '106' '107' '108' '110' '111' '112' '113' '114' '115' '116' '117'\n '118' '119' '120' '121' '122' '123' '124' '125' '126' '127' '128' '130'\n '131' '133' '134' '135' '136' '138' '139' '140' '141' '144' '145' '146'\n '147' '149' '150' '151' '153' '154' '155' '156' '157' '158' '159' '160'\n 'unknown'], 121, 176020\nValidating WSI: ['034' '035' '036' '037' '038' '039' '040' '041' '042' '043' '044' '045'\n '046' '047' '048' '049' '050' '051' '052' '053' '054' '055' '056' '058'\n '059' '060' '061'], 27, 44005\n```\nBut still got validation auc 0.99, but submission as 0.9745",
      "votes": null
    },
    {
      "id": "496608",
      "postDate": "03/22/2019 11:39:40",
      "content": "<p>This small difference can be explained here: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84964\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84964</a>. (but I did not get validation auc 0.99 at all). Did you actually submit your 0.98 to LB? If so, you wouldn't have 147th place right now?</p>",
      "rawMarkdown": "This small difference can be explained here: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84964. (but I did not get validation auc 0.99 at all). Did you actually submit your 0.98 to LB? If so, you wouldn't have 147th place right now?",
      "votes": null
    },
    {
      "id": "496701",
      "postDate": "03/22/2019 13:38:55",
      "content": "<p>A typo sorry.</p>",
      "rawMarkdown": "A typo sorry.",
      "votes": null
    },
    {
      "id": "496768",
      "postDate": "03/22/2019 15:05:50",
      "content": "<p>wsi splits does nothing for me, but I could be doing it wrong. </p>",
      "rawMarkdown": "wsi splits does nothing for me, but I could be doing it wrong.",
      "votes": null
    },
    {
      "id": "497853",
      "postDate": "03/24/2019 02:11:14",
      "content": "<p>unbelievable score!\nAcc is not equal to score. By the way,what is wsi?</p>",
      "rawMarkdown": "unbelievable score!\nAcc is not equal to score. By the way,what is wsi?",
      "votes": null
    },
    {
      "id": "497872",
      "postDate": "03/24/2019 02:33:13",
      "content": "<p>whole slide image, the larger full microscope slide images from which the patches were sampled.</p>\n\n<p>The 1.000 is not really a score since it used file hashes to match test images with their labels, it is actually a leak that invalidates the test set altogether. A legit score of 1 is probably not possible given the number of images that lack enough information for a good prediction.</p>",
      "rawMarkdown": "whole slide image, the larger full microscope slide images from which the patches were sampled.\n\nThe 1.000 is not really a score since it used file hashes to match test images with their labels, it is actually a leak that invalidates the test set altogether. A legit score of 1 is probably not possible given the number of images that lack enough information for a good prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496608,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "03/22/2019 11:39:40",
      "content": "<p>This small difference can be explained here: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84964\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84964</a>. (but I did not get validation auc 0.99 at all). Did you actually submit your 0.98 to LB? If so, you wouldn't have 147th place right now?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496701,
          "author_name": "xyzbinzhou",
          "author_url": "",
          "post_date": "03/22/2019 13:38:55",
          "content": "<p>A typo sorry.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496768,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "03/22/2019 15:05:50",
      "content": "<p>wsi splits does nothing for me, but I could be doing it wrong. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 497853,
      "author_name": "",
      "author_url": "",
      "post_date": "03/24/2019 02:11:14",
      "content": "<p>unbelievable score!\nAcc is not equal to score. By the way,what is wsi?</p>",
      "votes": null,
      "replies": [
        {
          "id": 497872,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "03/24/2019 02:33:13",
          "content": "<p>whole slide image, the larger full microscope slide images from which the patches were sampled.</p>\n\n<p>The 1.000 is not really a score since it used file hashes to match test images with their labels, it is actually a leak that invalidates the test set altogether. A legit score of 1 is probably not possible given the number of images that lack enough information for a good prediction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "496408": "I split the dataset according to WSI like this:\n\n\n```python\n    # Data loading code\n    df = pd.read_csv(os.path.join(args.data, 'train_labels.csv'))\n    df_wsi = pd.read_csv(os.path.join(args.data, 'patch_id_wsi.csv'))\n    df_test = pd.read_csv(os.path.join(args.data, 'sample_submission.csv'))\n    df_total = pd.merge(df, df_wsi, on='id', how='outer')\n    df_total.fillna('unknown', inplace=True)\n    df_total = shuffle(df_total, random_state=0)\n    df_total['wsi'] = df_total['wsi'].apply(lambda x: x if x == 'unkown' else x.split('_')[-1])\n    df_total = df_total.sort_values(by='wsi')\n\n    kf = KFold(n_splits=5, shuffle=False, random_state=0)\n    for idx, (train_idx, val_idx) in enumerate(kf.split(df_total, df_total['wsi'])):\n        if idx &gt; args.split:\n            tr_wsi = df_total.iloc[train_idx]['wsi'].unique()\n            vl_wsi = df_total.iloc[val_idx]['wsi'].unique()\n            print(f\"Training WSI: {tr_wsi}, {len(tr_wsi)}, {len(train_idx)}\")\n            print(f\"Validating WSI: {vl_wsi}, {len(vl_wsi)}, {len(val_idx)}\")\n            break\n```\nWhich will give printout:\n\n```bash\nTraining WSI: ['001' '002' '003' '004' '005' '006' '007' '008' '009' '010' '011' '012' '014' '015' '016' '017' '018' '020' '021' '022' '023' '024' '025' '026' '027' '028' '029' '030' '032' '033' '034' '061' '062' '063' '064' '065' '066' '068' '069' '070' '071' '072' '073' '074' '075' '076' '077' '078'\n '080' '081' '082' '083' '084' '086' '087' '088' '089' '090' '091' '092'\n '093' '094' '095' '096' '097' '098' '099' '100' '101' '102' '103' '104'\n '105' '106' '107' '108' '110' '111' '112' '113' '114' '115' '116' '117'\n '118' '119' '120' '121' '122' '123' '124' '125' '126' '127' '128' '130'\n '131' '133' '134' '135' '136' '138' '139' '140' '141' '144' '145' '146'\n '147' '149' '150' '151' '153' '154' '155' '156' '157' '158' '159' '160'\n 'unknown'], 121, 176020\nValidating WSI: ['034' '035' '036' '037' '038' '039' '040' '041' '042' '043' '044' '045'\n '046' '047' '048' '049' '050' '051' '052' '053' '054' '055' '056' '058'\n '059' '060' '061'], 27, 44005\n```\nBut still got validation auc 0.99, but submission as 0.9745",
    "496608": "This small difference can be explained here: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84964. (but I did not get validation auc 0.99 at all). Did you actually submit your 0.98 to LB? If so, you wouldn't have 147th place right now?",
    "496701": "A typo sorry.",
    "496768": "wsi splits does nothing for me, but I could be doing it wrong.",
    "497853": "unbelievable score!\nAcc is not equal to score. By the way,what is wsi?",
    "497872": "whole slide image, the larger full microscope slide images from which the patches were sampled.\n\nThe 1.000 is not really a score since it used file hashes to match test images with their labels, it is actually a leak that invalidates the test set altogether. A legit score of 1 is probably not possible given the number of images that lack enough information for a good prediction."
  },
  "source": "meta"
}