{
  "id": 203413,
  "title": "PyTorch Baseline",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/203413",
  "author_name": "",
  "post_date": "2020-12-15T06:49:07.386407400Z",
  "votes": 40,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I prepared PyTorch Baseline.</p>\n<p>Version1</p>\n<ul>\n<li>Model: resnext50_32x4d</li>\n<li>Split: GroupKFold 5 folds</li>\n<li>Size: 320x320</li>\n<li>CV: 0.959, LB: 0.971 (old metric)</li>\n<li>CV: 0.893, LB: 0.923 (new metric)</li>\n<li>training: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49348116\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49348116</a></li>\n<li>inference: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49367987\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49367987</a></li>\n</ul>\n<p>Version2</p>\n<ul>\n<li>Model: resnext50_32x4d</li>\n<li>Split: GroupKFold 5 folds</li>\n<li>Size: 448x448</li>\n<li>CV: 0.9281, LB: 0.943</li>\n<li>training: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49697928\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49697928</a></li>\n<li>inference: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49707639\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49707639</a></li>\n</ul>\n<p>Version3</p>\n<ul>\n<li>Model: resnext50_32x4d</li>\n<li>Split: GroupKFold 4 folds</li>\n<li>Size: 600x600</li>\n<li>CV: 0.9337, LB: 0.948</li>\n<li>training: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49722999\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49722999</a></li>\n<li>inference: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49757551\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49757551</a></li>\n</ul>\n<p>Hope this helps, happy kaggling!</p>",
  "messages": [
    {
      "id": "1113069",
      "postDate": "12/15/2020 06:49:07",
      "content": "<p>I prepared PyTorch Baseline.</p>\n<p>Version1</p>\n<ul>\n<li>Model: resnext50_32x4d</li>\n<li>Split: GroupKFold 5 folds</li>\n<li>Size: 320x320</li>\n<li>CV: 0.959, LB: 0.971 (old metric)</li>\n<li>CV: 0.893, LB: 0.923 (new metric)</li>\n<li>training: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49348116\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49348116</a></li>\n<li>inference: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49367987\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49367987</a></li>\n</ul>\n<p>Version2</p>\n<ul>\n<li>Model: resnext50_32x4d</li>\n<li>Split: GroupKFold 5 folds</li>\n<li>Size: 448x448</li>\n<li>CV: 0.9281, LB: 0.943</li>\n<li>training: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49697928\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49697928</a></li>\n<li>inference: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49707639\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49707639</a></li>\n</ul>\n<p>Version3</p>\n<ul>\n<li>Model: resnext50_32x4d</li>\n<li>Split: GroupKFold 4 folds</li>\n<li>Size: 600x600</li>\n<li>CV: 0.9337, LB: 0.948</li>\n<li>training: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49722999\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49722999</a></li>\n<li>inference: <a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49757551\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49757551</a></li>\n</ul>\n<p>Hope this helps, happy kaggling!</p>",
      "rawMarkdown": "I prepared PyTorch Baseline.\n\nVersion1\n- Model: resnext50_32x4d\n- Split: GroupKFold 5 folds\n- Size: 320x320\n- CV: 0.959, LB: 0.971 (old metric)\n- CV: 0.893, LB: 0.923 (new metric)\n- training: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49348116\n- inference: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49367987\n\nVersion2\n- Model: resnext50_32x4d\n- Split: GroupKFold 5 folds\n- Size: 448x448\n- CV: 0.9281, LB: 0.943\n- training: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49697928\n- inference: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49707639\n\nVersion3\n- Model: resnext50_32x4d\n- Split: GroupKFold 4 folds\n- Size: 600x600\n- CV: 0.9337, LB: 0.948\n- training: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49722999\n- inference: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49757551\n\nHope this helps, happy kaggling!",
      "votes": null
    },
    {
      "id": "1113072",
      "postDate": "12/15/2020 06:51:49",
      "content": "<p>I'm not sure we can share resized dataset (refer to <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203342\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203342</a>)<br>\nHere I put the code of preparing 384x384 resized dataset used in this notebook.</p>\n<pre><code>import os\nimport cv2\nimport zipfile\nimport numpy as np\nimport pandas as pd\nfrom tqdm.auto import tqdm\nfrom matplotlib import pyplot as plt\n\ntrain = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\n\nTRAIN_PATH = '../input/ranzcr-clip-catheter-line-classification/train/'\n\nwith zipfile.ZipFile(f'train.zip', 'w') as img_out:\n    for uid in tqdm(train['StudyInstanceUID'].values):\n        image = cv2.imread(TRAIN_PATH + f'{uid}.jpg')\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        image = cv2.resize(image, (384, 384))\n        image = cv2.imencode('.png', image)[1]\n        img_out.writestr(f\"{uid}.png\", image)\n</code></pre>",
      "rawMarkdown": "I'm not sure we can share resized dataset (refer to https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203342)\nHere I put the code of preparing 384x384 resized dataset used in this notebook.\n\n```\nimport os\nimport cv2\nimport zipfile\nimport numpy as np\nimport pandas as pd\nfrom tqdm.auto import tqdm\nfrom matplotlib import pyplot as plt\n\ntrain = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\n\nTRAIN_PATH = '../input/ranzcr-clip-catheter-line-classification/train/'\n\nwith zipfile.ZipFile(f'train.zip', 'w') as img_out:\n    for uid in tqdm(train['StudyInstanceUID'].values):\n        image = cv2.imread(TRAIN_PATH + f'{uid}.jpg')\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        image = cv2.resize(image, (384, 384))\n        image = cv2.imencode('.png', image)[1]\n        img_out.writestr(f\"{uid}.png\", image)\n```",
      "votes": null
    },
    {
      "id": "1113346",
      "postDate": "12/15/2020 11:37:36",
      "content": "<p>Wow this seems very high for a baseline model :( Not much room to improve on that I guess.</p>",
      "rawMarkdown": "Wow this seems very high for a baseline model :( Not much room to improve on that I guess.",
      "votes": null
    },
    {
      "id": "1113365",
      "postDate": "12/15/2020 11:54:58",
      "content": "<p>As a baseline, I haven't used any augmentations &amp; segmentation annotations during training.<br>\nAnd as discussed in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203375\" target=\"_blank\">this thread</a>, AUC averaging column wise is better for competition metric…</p>",
      "rawMarkdown": "As a baseline, I haven't used any augmentations & segmentation annotations during training.\nAnd as discussed in [this thread](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203375), AUC averaging column wise is better for competition metric...",
      "votes": null
    },
    {
      "id": "1113372",
      "postDate": "12/15/2020 12:03:45",
      "content": "<p>very good work! kagglers are fast.</p>\n<p>i am stuck at other competitions … i will come back later.<br>\nAll the best to you!</p>",
      "rawMarkdown": "very good work! kagglers are fast.\n\ni am stuck at other competitions ... i will come back later.\nAll the best to you!",
      "votes": null
    },
    {
      "id": "1113393",
      "postDate": "12/15/2020 12:14:15",
      "content": "<p>Thanks!<br>\nWe have many competitions these days…</p>",
      "rawMarkdown": "Thanks!\nWe have many competitions these days...",
      "votes": null
    },
    {
      "id": "1113808",
      "postDate": "12/15/2020 17:48:41",
      "content": "<p>Maggie confirmed that it's fine to have preprocessing notebooks/datasets as long as it's clearly specified and attached to your model notebooks.</p>",
      "rawMarkdown": "Maggie confirmed that it's fine to have preprocessing notebooks/datasets as long as it's clearly specified and attached to your model notebooks.",
      "votes": null
    },
    {
      "id": "1115762",
      "postDate": "12/16/2020 14:58:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>\n<p>Your notebook has a score of 0.923. Another public notebook which is identical to yours (Kaggle shows no lines added or  removed) has a score of 0.971</p>\n<p>Can you please help me understand why there should be a difference in scores with same code and same data? Am i missing something… Thanks for explaining!</p>",
      "rawMarkdown": "Hi @yasufuminakama \n\nYour notebook has a score of 0.923. Another public notebook which is identical to yours (Kaggle shows no lines added or  removed) has a score of 0.971\n\nCan you please help me understand why there should be a difference in scores with same code and same data? Am i missing something... Thanks for explaining!",
      "votes": null
    },
    {
      "id": "1115872",
      "postDate": "12/16/2020 16:42:53",
      "content": "<p>I'm not sure but it seems my notebook shows new metric score, forked notebook shows old metric score.</p>",
      "rawMarkdown": "I'm not sure but it seems my notebook shows new metric score, forked notebook shows old metric score.",
      "votes": null
    },
    {
      "id": "1115900",
      "postDate": "12/16/2020 17:00:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>",
      "rawMarkdown": "Thanks @yasufuminakama",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1113072,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "12/15/2020 06:51:49",
      "content": "<p>I'm not sure we can share resized dataset (refer to <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203342\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203342</a>)<br>\nHere I put the code of preparing 384x384 resized dataset used in this notebook.</p>\n<pre><code>import os\nimport cv2\nimport zipfile\nimport numpy as np\nimport pandas as pd\nfrom tqdm.auto import tqdm\nfrom matplotlib import pyplot as plt\n\ntrain = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\n\nTRAIN_PATH = '../input/ranzcr-clip-catheter-line-classification/train/'\n\nwith zipfile.ZipFile(f'train.zip', 'w') as img_out:\n    for uid in tqdm(train['StudyInstanceUID'].values):\n        image = cv2.imread(TRAIN_PATH + f'{uid}.jpg')\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        image = cv2.resize(image, (384, 384))\n        image = cv2.imencode('.png', image)[1]\n        img_out.writestr(f\"{uid}.png\", image)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1113808,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "12/15/2020 17:48:41",
          "content": "<p>Maggie confirmed that it's fine to have preprocessing notebooks/datasets as long as it's clearly specified and attached to your model notebooks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1113346,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "12/15/2020 11:37:36",
      "content": "<p>Wow this seems very high for a baseline model :( Not much room to improve on that I guess.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1113365,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "12/15/2020 11:54:58",
          "content": "<p>As a baseline, I haven't used any augmentations &amp; segmentation annotations during training.<br>\nAnd as discussed in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203375\" target=\"_blank\">this thread</a>, AUC averaging column wise is better for competition metric…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1113372,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/15/2020 12:03:45",
      "content": "<p>very good work! kagglers are fast.</p>\n<p>i am stuck at other competitions … i will come back later.<br>\nAll the best to you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1113393,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "12/15/2020 12:14:15",
          "content": "<p>Thanks!<br>\nWe have many competitions these days…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1115762,
      "author_name": "kmldas",
      "author_url": "",
      "post_date": "12/16/2020 14:58:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>\n<p>Your notebook has a score of 0.923. Another public notebook which is identical to yours (Kaggle shows no lines added or  removed) has a score of 0.971</p>\n<p>Can you please help me understand why there should be a difference in scores with same code and same data? Am i missing something… Thanks for explaining!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1115872,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "12/16/2020 16:42:53",
          "content": "<p>I'm not sure but it seems my notebook shows new metric score, forked notebook shows old metric score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115900,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "12/16/2020 17:00:54",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1113069": "I prepared PyTorch Baseline.\n\nVersion1\n- Model: resnext50_32x4d\n- Split: GroupKFold 5 folds\n- Size: 320x320\n- CV: 0.959, LB: 0.971 (old metric)\n- CV: 0.893, LB: 0.923 (new metric)\n- training: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49348116\n- inference: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49367987\n\nVersion2\n- Model: resnext50_32x4d\n- Split: GroupKFold 5 folds\n- Size: 448x448\n- CV: 0.9281, LB: 0.943\n- training: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49697928\n- inference: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49707639\n\nVersion3\n- Model: resnext50_32x4d\n- Split: GroupKFold 4 folds\n- Size: 600x600\n- CV: 0.9337, LB: 0.948\n- training: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-training?scriptVersionId=49722999\n- inference: https://www.kaggle.com/yasufuminakama/ranzcr-resnext50-32x4d-starter-inference?scriptVersionId=49757551\n\nHope this helps, happy kaggling!",
    "1113072": "I'm not sure we can share resized dataset (refer to https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203342)\nHere I put the code of preparing 384x384 resized dataset used in this notebook.\n\n```\nimport os\nimport cv2\nimport zipfile\nimport numpy as np\nimport pandas as pd\nfrom tqdm.auto import tqdm\nfrom matplotlib import pyplot as plt\n\ntrain = pd.read_csv('../input/ranzcr-clip-catheter-line-classification/train.csv')\n\nTRAIN_PATH = '../input/ranzcr-clip-catheter-line-classification/train/'\n\nwith zipfile.ZipFile(f'train.zip', 'w') as img_out:\n    for uid in tqdm(train['StudyInstanceUID'].values):\n        image = cv2.imread(TRAIN_PATH + f'{uid}.jpg')\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        image = cv2.resize(image, (384, 384))\n        image = cv2.imencode('.png', image)[1]\n        img_out.writestr(f\"{uid}.png\", image)\n```",
    "1113346": "Wow this seems very high for a baseline model :( Not much room to improve on that I guess.",
    "1113365": "As a baseline, I haven't used any augmentations & segmentation annotations during training.\nAnd as discussed in [this thread](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203375), AUC averaging column wise is better for competition metric...",
    "1113372": "very good work! kagglers are fast.\n\ni am stuck at other competitions ... i will come back later.\nAll the best to you!",
    "1113393": "Thanks!\nWe have many competitions these days...",
    "1113808": "Maggie confirmed that it's fine to have preprocessing notebooks/datasets as long as it's clearly specified and attached to your model notebooks.",
    "1115762": "Hi @yasufuminakama \n\nYour notebook has a score of 0.923. Another public notebook which is identical to yours (Kaggle shows no lines added or  removed) has a score of 0.971\n\nCan you please help me understand why there should be a difference in scores with same code and same data? Am i missing something... Thanks for explaining!",
    "1115872": "I'm not sure but it seems my notebook shows new metric score, forked notebook shows old metric score.",
    "1115900": "Thanks @yasufuminakama"
  },
  "source": "meta"
}